EDBT 2026 Demo / reviewers in the wild / expert
Xiaolin Wu 0001
dblp:w/XiaolinWu
· DBLP profile ↗
301ranked-venue papers
76as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 237 · 60 first-author · 14 since 2021Databases, data management, data science and information retrieval · 41 · 17 first-authorArtificial intelligence and machine learning · 24 · 6 first-author · 9 since 2021Theory of computation · 20 · 7 first-authorComputer networks · 13 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 13 · 3 first-authorSystems, architecture and hardware · 5Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving 3D Gaussian Splatting Compression by Scene-Adaptive Lattice Vector Quantizationabstract3D Gaussian Splatting (3DGS) is rapidly gaining popularity for its photorealistic rendering quality and real-time performance, but it generates massive amounts of data. Hence compressing 3DGS data is necessary for the cost effectiveness of 3DGS models. Recently, several anchor-based neural compression methods have been proposed, achieving good 3DGS compression performance. However, they all rely on uniform scalar quantization (USQ) due to its simplicity. A tantalizing question is whether more sophisticated quantizers can improve the current 3DGS compression methods with very little extra overhead and minimal change to the system. The answer is yes by replacing USQ with lattice vector quantization (LVQ). To better capture scene-specific characteristics, we optimize the lattice basis for each scene, improving LVQ's adaptability and R-D efficiency. This scene-adaptive LVQ (SALVQ) strikes a balance between the R-D efficiency of vector quantization and the low complexity of USQ. SALVQ can be seamlessly integrated into existing 3DGS compression architectures, enhancing their R-D performance with minimal modifications and computational overhead. Moreover, by scaling the lattice basis vectors, SALVQ can dynamically adjust lattice density, enabling a single model to accommodate multiple bit rate targets. This flexibility eliminates the need to train separate models for different compression levels, significantly reducing training time and memory consumption. Hao Xu 0051, Xiaolin Wu 0001, Xi Zhang 0019 |
IEEE Trans. Image Process. | 2 |
| 2025 | Multirate Neural Image Compression with Adaptive Lattice Vector QuantizationabstractRecent research has explored integrating lattice vector quantization (LVQ) into learned image compression models. Due to its more efficient Voronoi covering of vector space than scalar quantization (SQ), LVQ achieves better rate-distortion (R-D) performance than SQ, while still retaining the low complexity advantage of SQ. However, existing LVQ-based methods have two shortcomings: 1) lack of a multirate coding mode, hence incapable to operate at different rates; 2) the use of a fixed lattice basis, hence nonadaptive to changing source distributions. To overcome these shortcomings, we propose a novel adaptive LVQ method, which is the first among LVQ-based methods to achieve both rate and domain adaptations. By scaling the lattice basis vector, our method can adjust the density of lattice points to achieve various bit rate targets, achieving superior R-D performance to current SQ-based variable rate models. Additionally, by using a learned invertible linear transformation between two different input domains, we can reshape the predefined lattice cell to better represent the target domain, further improving the R-D performance. To our knowledge, this paper represents the first attempt to propose a unified solution for rate adaptation and domain adaptation through quantizer design. Hao Xu 0051, Xiaolin Wu 0001, Xi Zhang 0019 |
CVPR | 2 |
| 2025 | Learning Grouped Lattice Vector Quantizers for Low-Bit LLM CompressionabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities but typically require extensive computational resources and memory for inference. Post-training quantization (PTQ) can effectively reduce these demands by storing weights in lower bit-width formats. However, standard uniform quantization often leads to notable performance degradation, particularly in low-bit scenarios. In this work, we introduce a Grouped Lattice Vector Quantization (GLVQ) framework that assigns each group of weights a customized lattice codebook, defined by a learnable generation matrix. To address the non-differentiability of the quantization process, we adopt Babai rounding to approximate nearest-lattice-point search during training, which enables stable optimization of the generation matrices. Once trained, decoding reduces to a simple matrix-vector multiplication, yielding an efficient and practical quantization pipeline. Experiments on multiple benchmarks show that our approach achieves a better trade-off between model size and accuracy compared to existing post-training quantization baselines, highlighting its effectiveness in deploying large models under stringent resource constraints. Our source code is available on GitHub repository: https://github.com/xzhang9308/GLVQ. Xi Zhang 0019, Xiaolin Wu 0001, Jiamang Wang, Weisi Lin |
NeurIPS | 2 |
| 2025 | Group Image Compression for Dual Use of Machine and Human VisionabstractFaces in a scene of human group, if coded with sufficient precision, can be computer analyzed for machine vision tasks involving faces. But this requires storing and communicating them at a very high bit rate. Traditional ROI-based image compression methods are ill suited to code many faces at high precision against a complex background. In this work, we propose a novel group image compression neural network (GICNet) of two layers: 1) the face layer dedicated to machine analysis, in which face bounding boxes are first cropped out of the background and converted to a compression-friendly canonical sketch-guided representation of fixed resolution for compact coding and facilitating downstream tasks without additional preprocessing; 2) the background layer dedicated to overall human vision perceptual quality, in which face residuals and background elements are coded and appended to the code stream. Experimental results demonstrate the effectiveness of our proposed GICNet, conserving up to 13%-57% bitrate for machine vision applications while maintaining competitive perceptual quality. Xiaolin Wu 0001, Fan Li 0003, Yiping Duan, Xiaoming Tao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Fast Point Cloud Geometry Compression with Context-Based Residual Coding and INR-Based Refinement
Hao Xu 0051, Xi Zhang 0019, Xiaolin Wu 0001 |
ECCV (68) | 3 |
| 2024 | Low-complexity ℓ∞-compression of light field images with a deep-decompression stage
M. Umair Mukati, Xi Zhang 0019, Xiaolin Wu 0001, Søren Forchhammer |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Deep Lossy Plus Residual Coding for Lossless and Near-Lossless Image CompressionabstractLossless and near-lossless image compression is of paramount importance to professional users in many technical fields, such as medicine, remote sensing, precision engineering and scientific research. But despite rapidly growing research interests in learning-based image compression, no published method offers both lossless and near-lossless modes. In this paper, we propose a unified and powerful deep lossy plus residual (DLPR) coding framework for both lossless and near-lossless image compression. In the lossless mode, the DLPR coding system first performs lossy compression and then lossless coding of residuals. We solve the joint lossy and residual compression problem in the approach of VAEs, and add autoregressive context modeling of the residuals to enhance lossless compression performance. In the near-lossless mode, we quantize the original residuals to satisfy a given ℓ∞error bound, and propose a scalable near-lossless compression scheme that works for variable ℓ∞bounds instead of training multiple networks. To expedite the DLPR coding, we increase the degree of algorithm parallelization by a novel design of coding context, and accelerate the entropy coding with adaptive residual interval. Experimental results demonstrate that the DLPR coding system achieves both the state-of-the-art lossless and near-lossless image compression performance with competitive coding speed. Yuanchao Bai, Xianming Liu 0005, Kai Wang 0070, Xiangyang Ji, Xiaolin Wu 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | KBStyle: Fast Style Transfer Using a 200 KB Network With Symmetric Knowledge DistillationabstractConvolutional Neural Networks (CNNs) have achieved remarkable progress in arbitrary artistic style transfer. However, the model size of existing state-of-the-art (SOTA) style transfer algorithms is immense, leading to enormous computational costs and memory demand. It makes real-time and high resolution hard for GPUs with limited memory and limits the application on mobile devices. This paper proposes a novel arbitrary artistic style transfer algorithm, KBStyle, whose model size is only 200 KB. Firstly, we design a style transfer network where the style encoder, content encoder, and corresponding decoder are custom designed to guarantee low computational cost and high shape retention. Besides, the weighted style loss function is presented to improve the performance of style migration. Then, we propose a novel knowledge distillation method (Symmetric Knowledge Distillation, SKD) for encoder-decoder-based style transfer models, which redefines the knowledge and symmetrically compresses the encoder and decoder. With the SKD, the proposed style transfer network is further compressed by 14 times to achieve the KBStyle. Experimental results demonstrate that the proposed SKD method achieves comparable results with other SOTA knowledge distillation algorithms for style transfer. Besides, the proposed KBStyle achieves high-quality stylized images. And the inference time of the KBStyle on an Nvidia TITAN RTX GPU is only 20 ms when the resolutions of the content image and style image are both 2k-resolution ( 2048×1080 ). Moreover, the 200 KB model size of KBStyle is much smaller than the SOTA models and facilitates style transfer on mobile devices. Wenshu Chen, Mingyu Wang 0001, Xiaolin Wu 0001, Xiaoyang Zeng |
IEEE Trans. Image Process. | 4 |
| 2023 | LVQAC: Lattice Vector Quantization Coupled with Spatially Adaptive Companding for Efficient Learned Image CompressionabstractRecently, numerous end-to-end optimized image compression neural networks have been developed and proved themselves as leaders in rate-distortion performance. The main strength of these learnt compression methods is in powerful nonlinear analysis and synthesis transforms that can be facilitated by deep neural networks. However, out of operational expediency, most of these end-to-end methods adopt uniform scalar quantizers rather than vector quantizers, which are information-theoretically optimal. In this paper, we present a novel Lattice Vector Quantization scheme coupled with a spatially Adaptive Companding (LVQAC) mapping. LVQ can better exploit the inter-feature dependencies than scalar uniform quantization while being computationally almost as simple as the latter. Moreover, to improve the adaptability of LVQ to source statistics, we couple a spatially adaptive companding (AC) mapping with LVQ. The resulting LVQAC design can be easily embedded into any end-to-end optimized image compression system. Extensive experiments demonstrate that for any end-to-end CNN image compression models, replacing uniform quantiter by LVQAC achieves better rate-distortion performance without significantly increasing the model complexity. Xi Zhang 0019, Xiaolin Wu 0001 |
CVPR | 2 |
| 2023 | AND: Adversarial Neural Degradation for Learning Blind Image Super-ResolutionabstractLearnt deep neural networks for image super-resolution fail easily if the assumed degradation model in training mismatches that of the real degradation source at the inference stage. Instead of attempting to exhaust all degradation variants in simulation, which is unwieldy and impractical, we propose a novel adversarial neural degradation (AND) model that can, when trained in conjunction with a deep restoration neural network under a minmax criterion, generate a wide range of highly nonlinear complex degradation effects without any explicit supervision. The AND model has a unique advantage over the current state of the art in that it can generalize much better to unseen degradation variants and hence deliver significantly improved restoration performance on real-world images. Fangzhou Luo, Xiaolin Wu 0001 |
NeurIPS | 2 |
| 2023 | Multi-Modality Deep Restoration of Extremely Compressed Face VideosabstractArguably the most common and salient object in daily video communications is the talking head, as encountered in social media, virtual classrooms, teleconferences, news broadcasting, talk shows, etc. When communication bandwidth is limited by network congestions or cost effectiveness, compression artifacts in talking head videos are inevitable. The resulting video quality degradation is highly visible and objectionable due to high acuity of human visual system to faces. To solve this problem, we develop a multi-modality deep convolutional neural network method for restoring face videos that are aggressively compressed. The main innovation is a new DCNN architecture that incorporates known priors of multiple modalities: the video-synchronized speech signal and semantic elements of the compression code stream, including motion vectors, code partition map and quantization parameters. These priors strongly correlate with the latent video and hence they are able to enhance the capability of deep learning to remove compression artifacts. Ample empirical evidences are presented to validate the superior performance of the proposed DCNN method on face videos over the existing state-of-the-art methods. Xi Zhang 0019, Xiaolin Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Test-Time Adaptation for Optical Flow Estimation Using Motion VectorsabstractDue to the prohibitive cost as well as technical challenges in annotating ground-truth optical flow for large-scale realistic video datasets, the existing deep learning models for optical flow estimation mostly rely on synthetic data for training, which in turn may lead to significant performance degradation under test-data distribution shift in real-world environments. In this work, we propose the methodology to tackle this important problem. We design a self-supervised learning task for adjusting the optical flow estimation model at test time. We exploit the fact that most videos are stored in compressed formats, from which compact information on motion, in the form of motion vectors and residuals, can be made readily available. We formulate the self-supervised task as motion vector prediction, and link this task to optical flow estimation. To the best of our knowledge, our Test-Time Adaption guided with Motion Vectors (TTA-MV), is the first work to perform such adaptation for optical flow. The experimental results demonstrate that TTA-MV can improve the generalization capability of various well-known deep learning methods for optical flow estimation, such as FlowNet, PWCNet, and RAFT. Seyed Mehdi Ayyoubzadeh, Irina Kezele, Yuanhao Yu, Xiaolin Wu 0001, Yang Wang 0003, Jin Tang 0005 |
IEEE Trans. Image Process. | 5 |
| 2023 | TSDN: Two-Stage Raw Denoising in the DarkabstractDenoising is one of the most significant procedures in the image processing pipeline. Nowadays, deep-learning-based algorithms have achieved superior denoising quality than traditional algorithms. However, the noise becomes severe in the dark environment, where even the SOTA algorithms fail to achieve satisfactory performance. Besides, the high computational complexity of deep-learning-based denoising algorithms makes them hardware unfriendly and difficult to process high-resolution images in real-time. To address these issues, a novel low-light RAW denoising algorithm Two-Stage-Denoising (TSDN), is proposed in this paper. In TSDN, denoising consists of two procedures: noise removal and image restoration. Firstly, in the noise-removal stage, most noise is removed from the image, and an intermediate image that is easier for the network to recover the clean image is obtained. Then, in the restoration stage, the clean image is restored from the intermediate image. The TSDN is designed to be light-weight for real-time and hardware friendly. However, the tiny network will be insufficient for satisfactory performance if directly trained from scratch. Therefore, we present an Expand-Shrink-Learning (ESL) method to train the TSDN. In the ESL method, firstly, the tiny network is expanded to a larger one with similar architecture but more channels and layers, which enhances the learning ability of the network because of more parameters. Secondly, the larger network is shrunk and restored to the original small network in fine-grained learning procedures, including Channel-Shrink-Learning (CSL) and Layer-Shrink-Learning (LSL). Experimental results demonstrate that the proposed TSDN achieves better performance (PSNR and SSIM) than other SOTA algorithms in the dark environment. Besides, the model size of TSDN is one-eighth of that of the U-Net for denoising (a classical denoising network). Wenshu Chen, Mingyu Wang 0001, Xiaolin Wu 0001, Xiaoyang Zeng |
IEEE Trans. Image Process. | 4 |
| 2022 | Data Acquisition and Preparation for Dual-Reference Deep Learning of Image Super-ResolutionabstractThe performance of deep learning based image super-resolution (SR) methods depend on how accurately the paired low and high resolution images for training characterize the sampling process of real cameras. Low and high resolution (LR ∼ HR) image pairs synthesized by degradation models (e.g., bicubic downsampling) deviate from those in reality; thus the synthetically-trained DCNN SR models work disappointingly when being applied to real-world images. To address this issue, we propose a novel data acquisition process to shoot a large set of LR ∼ HR image pairs using real cameras. The images are displayed on an ultra-high quality screen and captured at different resolutions. The resulting LR ∼ HR image pairs can be aligned at very high sub-pixel precision by a novel spatial-frequency dual-domain registration method, and hence they provide more appropriate training data for the learning task of super-resolution. Moreover, the captured HR image and the original digital image offer dual references to strengthen supervised learning. Experimental results show that training a super-resolution DCNN by our LR ∼ HR dataset achieves higher image quality than training it by other datasets in the literature. Moreover, the proposed screen-capturing data collection process can be automated; it can be carried out for any target camera with ease and low cost, offering a practical way of tailoring the training of a DCNN SR model separately to each of the given cameras. Xiaolin Wu 0001, Xiao Shu |
IEEE Trans. Image Process. | 2 |
| 2021 | Attention-Guided Image Compression by Deep Reconstruction of Compressive Sensed Saliency SkeletonabstractWe propose a deep learning system for attention-guided dual-layer image compression (AGDL). In the AGDL compression system, an image is encoded into two layers, a base layer and an attention-guided refinement layer. Unlike the existing ROI image compression methods that spend an extra bit budget equally on all pixels in ROI, AGDL employs a CNN module to predict those pixels on and near a saliency sketch within ROI that are critical to perceptual quality. Only the critical pixels are further sampled by compressive sensing (CS) to form a very compact refinement layer. Another novel CNN method is developed to jointly decode the two compression layers for a much refined reconstruction, while strictly satisfying the transmitted CS constraints on perceptually critical pixels. Extensive experiments demonstrate that the proposed AGDL system advances the state of the art in perception-aware image compression. Xi Zhang 0019, Xiaolin Wu 0001 |
CVPR | 2 |
| 2021 | Functional Neural Networks for Parametric Image Restoration ProblemsabstractAlmost every single image restoration problem has a closely related parameter, such as the scale factor in super-resolution, the noise level in image denoising, and the quality factor in JPEG deblocking. Although recent studies on image restoration problems have achieved great success due to the development of deep neural networks, they handle the parameter involved in an unsophisticated way. Most previous researchers either treat problems with different parameter levels as independent tasks, and train a specific model for each parameter level; or simply ignore the parameter, and train a single model for all parameter levels. The two popular approaches have their own shortcomings. The former is inefficient in computing and the latter is ineffective in performance. In this work, we propose a novel system called functional neural network (FuncNet) to solve a parametric image restoration problem with a single model. Unlike a plain neural network, the smallest conceptual element of our FuncNet is no longer a floating-point variable, but a function of the parameter of the problem. This feature makes it both efficient and effective for a parametric problem. We apply FuncNet to super-resolution, image denoising, and JPEG deblocking. The experimental results show the superiority of our FuncNet on all three parametric image restoration tasks over the state of the arts. Fangzhou Luo, Xiaolin Wu 0001 |
NeurIPS | 2 |
| 2021 | High Frequency Detail Accentuation in CNN Image RestorationabstractGiven its nature of statistical inference, machine learning methods incline to downplay relatively rare events. But in many applications statistical outliers carry disproportional significance; they can, if being left without special treatment as of now, cause CNNs to perform unsatisfactorily on instances of interests. This is the reason why existing CNN image restoration methods all suffer from the problem of blurred details. To overcome this weakness, we advocate a new training methodology to sensitize the CNNs to desired events even they are atypical. Specifically for image restoration, we propose a so-called high frequency feature accentuation space that promotes image sharpness and clarity by maximally discriminating the ground truth image and the CNN-restored image in atypical but semantically important features. Then we force the restored image to agree with the ground truth image in the feature accentuation space by including an auxiliary loss term in the training process. This aims at a high degree of agreement of the two images on high frequency constructs such as sharp edges and fine textures, i.e., penalizes image blurs. The new CNN design method is implemented and tested for tasks of image super-resolution and denoising. Experimental results demonstrate the achievement of our design objective. Seyed Mehdi Ayyoubzadeh, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Real-Time Deep Image Retouching Based on Learnt Semantics Dependent Global TransformsabstractAlthough artists' actions in photo retouching appear to be highly nonlinear in nature and very difficult to characterize analytically, we find that the net effects of interactively editing a mundane image to a desired appearance can be modeled, in most cases, by a parametric monotonically non-decreasing global tone mapping function in the luminance axis and by a global affine transform in the chrominance plane that are weighted by saliency. This allows us to simplify the machine learning problem of mimicking artists in photo retouching to constructing a deep artful image transform (DAIT) using convolutional neural networks (CNN). The CNN design of DAIT aims to learn the image-dependent parameters of the luminance tone mapping function and the affine chrominance transform, rather than learning the end-to-end pixel level mapping as in the mainstream methods of image restoration and enhancement. The proposed DAIT approach reduces the computation complexity of the neural network by two orders of magnitude, which also, as a side benefit, improves the robustness and generalization capability at the inference stage. The high throughput and robustness of DAIT lend itself readily to real-time video enhancement as well after a simple temporal processing. Experiments and a Turing-type test are conducted to evaluate the proposed method and its competitors. Qifan Gao, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Ultra High Fidelity Deep Image Decompression With l∞-Constrained CompressionabstractWe propose a novel asymmetric image compression system of light ℓ∞-constrained predictive encoding and heavy-duty CNN-based soft decoding. The system achieves superior rate-distortion performances over the best of existing image compression methods, including BPG, WebP, FLIF and recent CNN codecs, in both ℓ2and ℓ∞error metrics, for bit rates near or above the threshold of perceptually transparent reconstruction. These remarkable coding gains are made by deep learning for compression artifact removal. A restoration CNN is designed to map a lossy compressed image to its original. Its unique strength is to enforce a tight error bound on a per pixel basis. As such, no small distinctive structures of the original image can be dropped or distorted, even if they are statistical outliers that are otherwise sacrificed by mainstream CNN restoration methods. Xi Zhang 0019, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsabstractClose-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity of human visual system to faces, compression distortions in talking heads videos are highly visible and annoying. To address this problem, we present a novel deep convolutional neural network (DCNN) method for very low bit rate video reconstruction of talking heads. The key innovation is a new DCNN architecture that can exploit the audio-video correlations to repair compression defects in the face region. We further improve reconstruction quality by embedding into our DCNN the encoder information of the video compression standards and introducing a constraining projection module in the network. Extensive experiments demonstrate that the proposed DCNN method outperforms the existing state-of-the-art methods on videos of talking heads. Xi Zhang 0019, Xiaolin Wu 0001, Xinliang Zhai, Xianye Ben, Chengjie Tu |
CVPR | 2 |
| 2020 | Classification of Depth and Surface Edges with Deep FeaturesabstractEdges in 2D images fall into two categories: depth edges and surface edges, depending on if the edge corresponds to an abrupt change in depth (the distance from the camera). This edge type is an efficient, robust, and effective information in many applications. In this paper we study the problem of automatic classification of the two types of edges. We use features discovered by deep convolutional neural networks to predict the edge type. The labeled sample edges for training our classifiers are semiautomatically generated via a technique of the edge-segment geometrical duality. Experiments are carried out to demonstrate the effectiveness of the proposed edge type classification methods. Zhenhao Li 0001, Xiaolin Wu 0001 |
ICASSP | 2 |
| 2020 | Exaggerated Learning For Clean-And-Sharp Image RestorationabstractDeep learning has become a methodology of choice for image restoration tasks, including denoising, super-resolution, deblurring, exposure correction, etc., because of its superiority to traditional methods in reconstruction quality. However, the published deep learning methods still have not solve the old dilemma between low noise level and detail sharpness. We propose a new CNN design strategy, called exaggerated deep learning, to reconcile two mutually conflicting objectives: noise free and detail sharpness. The idea is to deliberately overshoot for the desired attributes in the CNN optimization objective function; the cleanness or sharpness is overemphasized according to different semantic contexts. The exaggerated learning approach is experimented on the restoration tasks of super-resolution and low light correction. Its effectiveness and advantages have been empirically affirmed. Qifan Gao, Xiaolin Wu 0001 |
ICIP | 3 |
| 2020 | Maximum a Posteriori on a Submanifold: a General Image Restoration Method with GANabstractWe propose a general method for various image restoration problems, such as denoising, deblurring, super-resolution and inpainting. The problem is formulated as a constrained optimization problem. Its objective is to maximize a posteriori probability of latent variables, and its constraint is that the image generated by these latent variables must be the same as the degraded image. We use a Generative Adversarial Network (GAN) as our density estimation model. Convincing results are obtained on MNIST dataset. Fangzhou Luo, Xiaolin Wu 0001 |
IJCNN | 2 |
| 2020 | Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosabstractWe propose a novel deep multi-modality neural network for restoring very low bit rate videos of talking heads. Such video contents are very common in social media, teleconferencing, distance education, tele-medicine, etc., and often need to be transmitted with limited bandwidth. The proposed CNN method exploits the correlations among three modalities, video, audio and emotion state of the speaker, to remove the video compression artifacts caused by spatial down sampling and quantization. The deep learning approach turns out to be ideally suited for the video restoration task, as the complex non-linear cross-modality correlations are very difficult to model analytically and explicitly. The new method is a video post processor that can significantly boost the perceptual quality of aggressively compressed talking head videos, while being fully compatible with all existing video compression standards. Xi Zhang 0019, Xiaolin Wu 0001 |
ACM Multimedia | 3 |
| 2020 | On Numerosity of Deep Neural NetworksabstractRecently, a provocative claim was published that number sense spontaneously emerges in a deep neural network trained merely for visual object recognition. This has, if true, far reaching significance to the fields of machine learning and cognitive science alike. In this paper, we prove the above claim to be unfortunately incorrect. The statistical analysis to support the claim is flawed in that the sample set used to identify number-aware neurons is too small, compared to the huge number of neurons in the object recognition network. By this flawed analysis one could mistakenly identify number-sensing neurons in any randomly initialized deep neural networks that are not trained at all. With the above critique we ask the question what if a deep convolutional neural network is carefully trained for numerosity? Our findings are mixed. Even after being trained with number-depicting images, the deep learning approach still has difficulties to acquire the abstract concept of numbers, a cognitive task that preschoolers perform with ease. But on the other hand, we do find some encouraging evidences suggesting that deep neural networks are more robust to distribution shift for small numbers than for large numbers. Xi Zhang 0019, Xiaolin Wu 0001 |
NeurIPS | 2 |
| 2019 | Cognitive Deficit of Deep Learning in NumerosityabstractSubitizing, or the sense of small natural numbers, is an innate cognitive function of humans and primates; it responds to visual stimuli prior to the development of any symbolic skills, language or arithmetic. Given successes of deep learning (DL) in tasks of visual intelligence and given the primitivity of number sense, a tantalizing question is whether DL can comprehend numbers and perform subitizing. But somewhat disappointingly, extensive experiments of the type of cognitive psychology demonstrate that the examples-driven black box DL cannot see through superficial variations in visual representations and distill the abstract notion of natural number, a task that children perform with high accuracy and confidence. The failure is apparently due to the learning method not the CNN computational machinery itself. A recurrent neural network capable of subitizing does exist, which we construct by encoding a mechanism of mathematical morphology into the CNN convolutional kernels. Also, we investigate, using subitizing as a test bed, the ways to aid the black box DL by cognitive priors derived from human insight. Our findings are mixed and interesting, pointing to both cognitive deficit of pure DL, and some measured successes of boosting DL by predetermined cognitive implements. This case study of DL in cognitive computing is meaningful for visual numerosity represents a minimum level of human intelligence. Xiaolin Wu 0001, Xi Zhang 0019, Xiao Shu |
AAAI | 1 |
| 2019 | Near-Lossless ℓ∞-Constrained Image Decompression via Deep Neural NetworkabstractRecently a number of CNN-based techniques were proposed to remove image compression artifacts. As in other restoration applications, these techniques all learn a mapping from decompressed patches to the original counterparts under the ubiquitous L2 metric. However, this approach is incapable of restoring distinctive image details which may be statistical outliers but have high semantic importance (e.g., tiny lesions in medical images). To overcome this weakness, we propose to incorporate an ℓ∞fidelity criterion in the design of neural network so that no small, distinctive structures of the original image can be dropped or distorted. Experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in ℓ∞error metric and perceptual quality, while being competitive in L2 error metric as well. It can restore subtle image details that are otherwise destroyed or missed by other algorithms. Our research suggests a new machine learning paradigm of ultra high fidelity image compression that is ideally suited for applications in medicine, space, and sciences. Xi Zhang 0019, Xiaolin Wu 0001 |
DCC | 2 |
| 2019 | Nonlinear Prediction of Multidimensional Signals via Deep Regression with Applications to Image CodingabstractDeep convolutional neural networks (DCNN) have enjoyed great successes in many signal processing applications because they can learn complex, non-linear causal relationships from input to output. In this light, DCNNs are well suited for the task of sequential prediction of multidimensional signals, such as images, and have the potential of improving the performance of traditional linear predictors. In this research we investigate how far DCNNs can push the envelop in terms of prediction precision. We propose, in a case study, a two-stage deep regression DCNN framework for nonlinear prediction of two-dimensional image signals. In the first-stage regression, the proposed deep prediction network (PredNet) takes the causal context as input and emits a prediction of the present pixel. Three PredNets are trained with the regression objectives of minimizing l1, l2and l∞norms of prediction residuals, respectively. The second-stage regression combines the outputs of the three PredNets to generate an even more precise and robust prediction. The proposed deep regression model is applied to lossless predictive image coding, and it outperforms the state-of-the-art linear predictors by appreciable margin. Xi Zhang 0019, Xiaolin Wu 0001 |
ICASSP | 2 |
| 2019 | Deep Restoration of Vintage Photographs From Scanned Halftone PrintsabstractA great number of invaluable historical photographs unfortunately only exist in the form of halftone prints in old publications such as newspapers or books. Their original continuous-tone films have long been lost or irreparably damaged. There have been attempts to digitally restore these vintage halftone prints to the original film quality or higher. However, even using powerful deep convolutional neural networks, it is still difficult to obtain satisfactory results. The main challenge is that the degradation process is complex and compounded while little to no real data is available for properly training a data-driven method. In this research, we adopt a novel strategy of two-stage deep learning, in which the restoration task is divided into two stages: the removal of printing artifacts and the inverse of halftoning. The advantage of our technique is that only the simple first stage requires unsupervised training in order to make the combined network generalize on real halftone prints, while the more complex second stage of inverse halftoning can be easily trained with synthetic data. Extensive experimental results demonstrate the efficacy of the proposed technique for real halftone prints; the new technique significantly outperforms the existing ones in visual quality. Qifan Gao, Xiao Shu, Xiaolin Wu 0001 |
ICCV | 3 |
| 2019 | High Joint Spectral-Spatial Resolution Imaging via Nanostructured Random Broadband FilteringabstractIt is a challenge to acquire a snapshot image of very high resolutions in both spectral and spatial domain via a single short exposure. In this setting one cannot trade time for spectral resolution, such as via spectral bands scanning. Cameras of color filter arrays (CFA) (e.g., the Bayer mosaic) cannot obtain high spectral resolution. To overcome these difficulties, we propose a new multispectral imaging system that makes random linear broadband measurements of the spectrum via a nanostructured multispectral filter array (MSFA). These MS-FA random measurements can be used by sparsity-based recovery algorithms to achieve much higher spectral resolution than conventional CFA cameras, without sacrificing spatial resolution. The key innovation is to jointly exploit both spatial and spectral sparsity properties that are inherent to spectral reflectance of natural objects. Experimental results establish the superior performance of the proposed multispectral imaging system over existing ones. Xiaolin Wu 0001, Dahua Gao, Kaiwei Zhang |
ICIP | 1 |
| 2019 | Joint Demosaicking and Blind Deblurring Using Deep Convolutional Neural NetworkabstractDespite extensive research efforts, blind image deblurring remains a challenge without general robust solutions. A long-overlooked problem of existing deblurring methods is that they are all designed to work on fully sampled RGB input images for simplicity. But, in practice, most RGB color images are reconstructed from Bayer mosaic data hence riddled with various high-frequency demosaicking artifacts, such as zippering and moiré patterns, which can easily derail a deblurring algorithm. In this paper, we propose a novel multi-scale deep convolutional neural network to solve demosaicking and deblurring jointly. By processing Bayer raw images directly, our method is free of the interference of demosaicking artifacts. Extensive experiments show that the joint approach greatly outperforms the simple cascade of state-of-art demosaicking and deblurring methods. Zhixiang Chi, Xiao Shu, Xiaolin Wu 0001 |
ICIP | 3 |
| 2018 | Fast Screening Algorithm for Rotation Invariant Template MatchingabstractThis paper presents a generic screening algorithm for expediting conventional template matching techniques. The algorithm can rule out regions with no possible matches with minimum computational efforts; moreover, the match is robust against rotation changes. The computational efficiency and robustness against rotation are gained by using a novel octagonal star shaped query template and the inclusion-exclusion principle to extract and compare patch features. Extensive experiments demonstrate that the proposed algorithm greatly reduces the search space without adversely affecting the matching accuracy. Bolin Liu, Xiao Shu, Xiaolin Wu 0001 |
ICIP | 3 |
| 2018 | Multispectral Image Restoration via Inter- and Intra-Block Sparse Estimation Based on Physically-Induced Joint Spatiospectral StructuresabstractExisting low-level vision algorithms (e.g., those for superresolution, denoising, deblurring etc.) were primarily motivated and optimized for precision in spatial domain. However, high precision in spectral domain is of importance for many applications in scientific and technical fields, such as spectral analysis, recognition, and classification. In quest for both high spectral and spatial fidelity we introduce previously-unexplored, physically-induced, joint spatiospectral sparsities to improve existing methods for multispectral image restoration. The bidirectional image formation model is used to reveal that the discontinuities of a multispectral image tend to align spatially across different spectral bands; in other words, the 2D Laplacians of different bands are not only sparse each, but they also agree with one the other in significance positions. Such strongly structured sparsities give rise to a new inter-and intra-block sparse estimation approach. The estimation is performed on 3D spatiospectral sample blocks, rather than on separate 2D patches, one per spectral band or per luminance and chrominance component as in current practice. Moreover, intra-block and inter-block sparsity priors are combined via an intra-block ℓ1,2-norm minimization term and an inter-block low rank term, strengthening the regularization of the underlying inverse problem. The new approach is tested and evaluated on two concrete applications: superresolving and denoising multispectral images; its validity and advantages over the current state of the art are established by empirical results. Dahua Gao, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Learning-Based Restoration of Backlit ImagesabstractBacklighting is a commonly encountered ill illumination condition that can cause serious degradation of image quality. In this paper, we propose a learning-based spatially adaptive technique of optimal tone mapping to restore backlit images. Object surfaces illuminated from behind in a scene are detected by a soft binary classifier that is constructed via supervised learning. Two optimal tone mapping functions, one for backlit regions and the other for the remainder of the image, are used and their outputs are fused to restore illegible surface details in backlit regions and at the same time improve contrast in overexposed regions, if any. Experimental results demonstrate the superior performance of the proposed new technique over existing image enhancement techniques on backlit photographs. Zhenhao Li 0001, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Locally Adaptive Rank-Constrained Optimal Tone MappingabstractHigh dynamic range (HDR) tone mapping is formulated as an optimization problem of maximizing perceivable spatial details given the limited dynamic range of display devices. This objective can be attained, as supported by our results, by a novel image display methodology called locally adaptive rank-constrained optimal tone mapping (LARCOTM). The scientific basis for LARCOTM is that the maximum discrimination power of human vision system can only be achieved in a relatively small locality of an image. LARCOTM is fundamentally different from existing HDR tone mapping techniques in that the former can preserve pixel value order statistics within localities in which human foveal vision retains maximum sensitivity, while the latter cannot. As a result, images enhanced by LARCOTM are free of artifacts such as halos and double edges that plague other HDR methods. Xiao Shu, Xiaolin Wu 0001 |
ACM Trans. Graph. | 2 |
| 2017 | Content Adaptive Embedded CompressionabstractImage resolution in modern video processing and display systems are rocketing up in recent years. With the main stream video quality evolving from standard definition to high-definition, and further towards the emerging super high-definition, the bandwidth and power consumption of external memory are becoming serious bottlenecks. In this paper, frequency domain analysis is made on image down-sampling, which gives birth to the optimal sampling strategy for high-frequency energy protection. On that basis, we develop a new embedded compression (EC) technique, which encodes image frames through content-adaptive down-sampling and decodes them using side-information aided up-sampling. Apart from the low encoding complexity and even lower decoding complexity, the proposed EC algorithm also features high fidelity for images with sharp edges. Comparisons with some existing EC algorithms show the advantage of the new technique over its counterparts on a wide class of testing images. Yuxiang Shen, Xiaolin Wu 0001, Xiao Shu |
DCC | 2 |
| 2017 | A model-based approach for human head-and-shoulder segmentationabstractObject boundary extraction has long been a fundamental research topic, as well as an essential component in many visual computing and communication algorithms, such as computer vision, robotics, pattern recognition and video compression. Under this topic, human head-and-shoulder segmentation is of particular meaning, given the ubiquity of head-and-shoulder type of videos in social media, teleconferencing, and entertainment. Although human visual system can easily detect and recognize the head and upper body of a person, this seemingly simple task still poses a challenge to computers. In this paper, an effective and efficient segmentation method is proposed. This method consists of a novel human body descriptor in polar coordinates and a Markov chain based boundary model, which work together to generate precise boundary results. Moreover, dynamic programming is employed in this work, so as to accelerate the segmentation process. Comparisons with other algorithms are made in the experimental part, which clearly exhibits the advantage of our proposed method over some of its precedents. Xiaowei Deng, Yuxiang Shen, Xiaolin Wu 0001 |
ICIP | 3 |
| 2017 | Multichannel guided image filterabstractGuided image filter (GIF) is an edge-preserving filtering technique that smooths the fine texture of an input image with the guide of a second image. One shortcoming of GIF and all its existing variants, such as, weighed guided image filter and gradient domain guided image filter, is that they only use one grayscale image as the guide and are consequently unable to fully utilize the rich information offered by multichannel images. In this paper, we extend GIF for multichannel guidance image and propose a novel correlation detection technique for retaining sharp edges with opposite gradient directions in the different channels of the same guidance image. Experimental results show that the proposed method preserves the details from both the input and guidance images better than existing GIF techniques. Xiaolin Wu 0001, Xiao Shu |
ICIP | 2 |
| 2017 | Fovea weighting of multiview computational displays for enhanced user experienceabstractA challenge for multiview displays is how to concurrently exhibit many different views on the same physical medium without sacrificing image quality for individual viewers. To address the above challenge, we propose a novel scheme of fovea weighting in the framework of the TPVM multiview computational display to enhance users' visual experiences. Underlying our new design is an optimization problem of nonnegative matrix factorization. This seemingly difficult problem turns out to be solvable by efficient algorithms. The effectiveness of the proposed multiview fovea weighting is validated by simulation results. Fangzhou Luo, Xiaolin Wu 0001 |
ICIP | 2 |
| 2017 | A study on quantization effects of DCT based compressionabstractQuantization in discrete cosine transform (DCT) domain is a widely used lossy compression technique in international multimedia compression standards from audio (e.g., MP3) to image (e.g., JPEG) to video (e.g., H.264). Unlike many degradation sources, like sensor noises, quantization errors of DCT coefficients are signal dependent and difficult to isolate and remove, causing serious artifacts in decompressed and post-processed signals. In this research, quantization errors in the DCT domain are analytically assessed. Our analysis exposes complex behaviors of the DCT quantization errors after being mapped back into the temporal or spatial domain. These behaviors are highly sensitive to quantization precision, the amplitude and phase of the input signal. Based on these observations, we develop a DCT-domain error model to predict and quantify the quantization effects in cases where artifacts are most perceivable to humans, and offer some insights into possible strategies for further suppressing compression noises. Xiao Shu, Xiaolin Wu 0001, Bolin Liu |
ICIP | 2 |
| 2017 | Anti-camera LED LightingabstractThis work is concerned with the protection of intellectual property rights and privacy against unpermitted uses of digital cameras. A technique of multispectral coded illumination (MSCI) with LED lights is proposed to defeat cameras capturing indoor scenes by inducing annoying color artifacts into the acquired images or video frames. The main idea of MSCI is to temporally modulate LED lights of different colors at certain frequencies so that they interfere with the rolling shutter of the camera, but at the same time the coded illumination appears to human eyes the same as steady white lighting. Xiao Shu, Xiaolin Wu 0001, Qifan Gao |
ACM Multimedia | 2 |
| 2017 | Illumination invariant feature based on neighboring radiance ratioabstractIn many object recognition applications, especially in face recognition, varying illuminations can adversely affect the robustness of the object recognition system. In this paper, we propose a novel illumination invariant feature called Neighboring Radiance Ratio (NRR) which is insensitive to both intensity and direction of light. NRR is derived and analyzed based on a physical image formation model. The computation of NRR does not need any prior information or any training data and NRR is far less sensitive to the border of shadows than most existing methods. The analysis of the illumination invariance of NRR is also presented. The proposed NRR feature is tested on Extended Yale B and CMU-PIE databases and compared with several previous methods. The experimental results corroborate our analysis and demonstrate that NRR is highly robust image feature against illumination changes. Xi Zhang 0019, Xiaolin Wu 0001 |
VCIP | 2 |
| 2017 | Random Walk Graph Laplacian-Based Smoothness Prior for Soft Decoding of JPEG ImagesabstractGiven the prevalence of joint photographic experts group (JPEG) compressed images, optimizing image reconstruction from the compressed format remains an important problem. Instead of simply reconstructing a pixel block from the centers of indexed discrete cosine transform (DCT) coefficient quantization bins (hard decoding), soft decoding reconstructs a block by selecting appropriate coefficient values within the indexed bins with the help of signal priors. The challenge thus lies in how to define suitable priors and apply them effectively. In this paper, we combine three image priors-Laplacian prior for DCT coefficients, sparsity prior, and graph-signal smoothness prior for image patches-to construct an efficient JPEG soft decoding algorithm. Specifically, we first use the Laplacian prior to compute a minimum mean square error initial solution for each code block. Next, we show that while the sparsity prior can reduce block artifacts, limiting the size of the overcomplete dictionary (to lower computation) would lead to poor recovery of high DCT frequencies. To alleviate this problem, we design a new graph-signal smoothness prior (desired signal has mainly low graph frequencies) based on the left eigenvectors of the random walk graph Laplacian matrix (LERaG). Compared with the previous graph-signal smoothness priors, LERaG has desirable image filtering properties with low computation overhead. We demonstrate how LERaG can facilitate recovery of high DCT frequencies of a piecewise smooth signal via an interpretation of low graph frequency components as relaxed solutions to normalized cut in spectral clustering. Finally, we construct a soft decoding algorithm using the three signal priors with appropriate prior weights. Experimental results show that our proposal outperforms the state-of-the-art soft decoding algorithms in both objective and subjective evaluations noticeably. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001, Debin Zhao |
IEEE Trans. Image Process. | 3 |
| 2016 | Multispectral image super-resolution with ℓ1, 2-norm regularization of spatially-aligned LaplaciansabstractIn quest for high spectral fidelity in spatial superresolution of multispectral images, we explore physically-induced, joint spectral-spatial sparsities. The bichromatic image formation model is used to reveal that the discontinuities of a multi-spectral image tend to align spatially across different spectral bands; in other words, the 2D Laplacians of different bands are not only sparse but also agree with one the other in positions of significance. This strong prior of natural images can be incorporated, as an ℓ1,2-norm regularization term, into an inverse problem formulation for superresolution of multispectral images. Experiments show that exploiting the newly discovered joint spectral-spatial sparsities can improve the performance of existing methods, especially in spectral fidelity. Xiaolin Wu 0001, Dahua Gao |
ICIP | 1 |
| 2016 | Blind quality assessment of compressed images via pseudo structural similarityabstractBlock-based compression causes severe pseudo structures. We find that the pseudo structures of images compressed by different levels show some degree of similarity. So we propose to evaluate the quality of compressed images via the similarity between pseudo structures of two images. To obtain a “reference” image, we introduce the most distorted image (MDI), which is derived from the distorted image and suffers from the highest degree of compression. The proposed pseudo structural similarity (PSS) model calculates the similarity between pseudo structures of the distorted image and MDI. Pseudo structures of the distorted image become similar to the MDI's under the condition of severe compression. Via comparative tests, the proposed PSS model, on one hand, is shown to be comparable to state-of-the-art competitors, and on the other hand, it is not only good at assessing natural scene images but also performs the best in the hotly-researched screen content image (SCI) database. It deserves to mention that PSS is able to boost the performance of mainstream general-purpose no-reference (NR) quality measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yuming Fang 0001, Xiaokang Yang 0001, Xiaolin Wu 0001, Jiantao Zhou 0001, Xianming Liu 0005 |
ICME | 6 |
| 2016 | Frame Untangling for Unobtrusive Display-Camera Visible Light CommunicationabstractPairing displays and cameras can open up convenient and "free" visible light communication channels. But in realistic settings, the synchronization between displays (transmitters) and cameras (receivers) can be far more involved than assumed in the literature. This study aims to analyze and model the temporal behaviors of displays and cameras to make the visible light communication channel between the two more robust, while maintaining perceptual transparency of the transmitted data. Xiao Shu, Xiaolin Wu 0001 |
ACM Multimedia | 2 |
| 2016 | Single image haze removal using Gaussian mixture model and sparse optimizationabstractSingle image haze removal is an underdetermined inverse problem whose solution hinges on valid image priors or models. In this work, robust priors drawn from outdoor scene statistics are explored. Specifically, a Gaussian mixture model of chrominance distribution is proposed toward transmittance estimation and its physical validity is justified. In addition, a new sparsity-based optimization approach for transmittance image super-resolution/restoration is proposed, which makes a solid assumption that most outdoor object surfaces are piece-wise linear and thus the corresponding depth image is sparse in Laplacian space. Experimental results are given in proof of the remarkably improved visual quality of our new haze removal technique over its predecessors. Yuxiang Shen, Xiaolin Wu 0001 |
VCIP | 2 |
| 2016 | Data-Driven Soft Decoding of Compressed Images in Dual Transform-Pixel DomainabstractIn the large body of research literature on image restoration, very few papers were concerned with compression-induced degradations, although in practice, the most common cause of image degradation is compression. This paper presents a novel approach to restoring JPEG-compressed images. The main innovation is in the approach of exploiting residual redundancies of JPEG code streams and sparsity properties of latent images. The restoration is a sparse coding process carried out jointly in the DCT and pixel domains. The prowess of the proposed approach is directly restoring DCT coefficients of the latent image to prevent the spreading of quantization errors into the pixel domain, and at the same time, using online machine-learned local spatial features to regulate the solution of the underlying inverse problem. Experimental results are encouraging and show the promise of the new approach in significantly improving the quality of DCT-coded images. Xianming Liu 0005, Xiaolin Wu 0001, Jiantao Zhou 0001, Debin Zhao |
IEEE Trans. Image Process. | 2 |
| 2016 | Image Enhancement by Entropy Maximization and Quantization Resolution UpconversionabstractThis article introduces a new contrast enhancement algorithm of tone-preserving entropy maximization. Its design objective is to present the maximal amount of information content in the enhanced image, or being optimal in an information theoretical sense, while preventing the loss of tone continuity. The resulting optimization problem can be graph theoretically modeled as the construction of the K-edges maximum-weight path, and it can be solved efficiently by dynamic programming. Moreover, the proposed algorithm is made more effective by being combined with a preprocess of image restoration that aims to correct quantization errors caused by the analog-to-digital conversion of image signals. Empirical evidences are provided to demonstrate the superior visual quality obtained by the new image enhancement algorithm. Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2015 | Data-driven sparsity-based restoration of JPEG-compressed images in dual transform-pixel domainabstractArguably the most common cause of image degradation is compression. This papers presents a novel approach to restoring JPEG-compressed images. The main innovation is in the approach of exploiting residual redundancies of JPEG code streams and sparsity properties of latent images. The restoration is a sparse coding process carried out jointy in the DCT and. pixel domains. The prowess of the proposed approach is directly restoring DCT coefficients of the latent image to prevent the spreading of quantization errors into the pixel domain, and at the same time using on-line machine-learnt local spatial features to regulate the solution of the underlying inverse problem. Experimental results are encouraging and show the promise of the new approach in significantly improving the quality of DCT-coded images. Xianming Liu 0005, Xiaolin Wu 0001, Jiantao Zhou 0001, Debin Zhao |
CVPR | 2 |
| 2015 | Joint denoising and contrast enhancement of images using graph laplacian operatorabstractImages and videos are often captured in poor light conditions, resulting in low-contrast images that are corrupted by acquisition noise. To recreate a high-quality image for visual observation, the captured image must be denoised and contrastenhanced. Conventional methods perform these two tasks in two separate stages: an image is first denoised, followed by an enhancement procedure. In this paper, we propose to jointly denoise and enhance an image in one unified optimization framework. The crux of the optimization rests on the definition of the enhancement operator, described by a graph Laplacian matrix H. The operator must enhance the high frequency details of the original image without amplifying additive noise. We propose a graph-based low-pass filtering approach to denoise edge weights in the graph, resulting in a more robust estimate of H. Experimental results show that our proposed joint approach can outperform the separate approach in demonstrable image quality. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001 |
ICASSP | 3 |
| 2015 | Combining information display and visible light wireless communicationabstractThis paper opens up an unforeseen and intriguing application area of our recent pioneer work of temporal psychovisual modulation: wireless optical communication. It is demonstrated how a high-speed optoelectronic display functions as a 2D array of optical transmitters and at the same exhibits conventional images as usual. At the receiving end, digital cameras can download data via the optical MIMO link while user(s) can work and read the display as they are accustomed. The said unification of information display and wireless optical communication is made possible by psychovisually based image processing. Xiaolin Wu 0001, Xiao Shu |
ICASSP | 1 |
| 2015 | Video Restoration Against Yin-Yang PhasingabstractA common video degradation problem, which is largely untreated in literature, is what we call Yin-Yang Phasing (YYP). YYP is characterized by involuntary, dramatic flip-flop in the intensity and possibly chromaticity of an object as the video plays. Such temporal artifacts occur under ill illumination conditions and are triggered by object or/and camera motions, which mislead the settings of camera's auto-exposure and white point. In this paper, we investigate the problem and propose a video restoration technique to suppress YYP artifacts and retain temporal consistency of objects appearance via inter-frame, spatially-adaptive, optimal tone mapping. The video quality can be further improved by a novel image enhancer designed in Weber's perception principle and by exploiting the second-order statistics of the scene. Experimental results are encouraging, pointing to an effective, practical solution for a common but surprisingly understudied problem. Xiaolin Wu 0001, Zhenhao Li 0001, Xiaowei Deng |
ICCV | 1 |
| 2015 | Sparsity-based depth image restoration using surface priors and RGB-D correlationsabstractIn this paper we propose a sparsity-based, directional variational approach for upsampling depth images, aided by an accompanying optical (in RGB) image of higher spatial resolution. Compared to previously published works on RGB-D superresolution, the main innovations of this work are: 1. performing depth image restoration in an overcomplete sparsity space derived from the directionalities of the RGB image; 2. refining the regularization term of the underlying inverse problem by a cross-validating spatial discontinuities in the optical and depth images. By integrating these new techniques the proposed depth image superresolution method delivers very competitive performance against existing ones. Xiaowei Deng, Xiaolin Wu 0001 |
ICIP | 2 |
| 2015 | Inter-block consistent soft decoding of JPEG images with sparsity and graph-signal smoothness priorsabstractGiven the prevalence of JPEG compressed images on the Internet, image reconstruction from the compressed format remains an important and practical problem. Instead of simply reconstructing a pixel block from the centers of assigned DCT coefficient quantization bins (hard decoding), we propose to jointly reconstruct a neighborhood group of pixel patches using two image priors while satisfying the quantization bin constraints. First, we assume that a pixel patch can be approximated as a sparse linear combination of atoms from an offline-learned over-complete dictionary. Second, we assume that a patch, when interpreted as a graph-signal, is smooth with respect to an appropriately defined graph that captures the estimated structure of the target image. Finally, neighboring patches in the optimization have sufficient overlaps and are forced to be consistent, so that blocking artifacts typical in JPEG decoded images are avoided. To find the optimal group of patches, we formulate a constrained optimization problem and propose a fast alternating algorithm to find locally optimal solutions. Experimental results show that our proposed algorithm outperforms state-of-the-art soft decoding algorithms by up to 1.47dB in PSNR. Xianming Liu 0005, Gene Cheung, Xiaolin Wu 0001, Debin Zhao |
ICIP | 3 |
| 2015 | Analysis on spectral effects of dark-channel prior for haze removalabstractIn solving the inverse problem of haze removal, the most commonly used prior in the literature is perhaps that of dark channel, which assumes that at least one pixel in a small patch has a zero or near zero intensity level in one of the RGB color channels. However, this assumption is not physically based; it can be significantly off from the reality because most colors in outdoor natural scenes are unsaturated (e.g., the sky). Chances are that none of the R, G, B values in a patch of the latent image is close to zero. This paper offers detailed analysis on the effects of invalid dark channel assumption on dehazed images; in particular, it reveals the causes and behavior of spectral distortions that are inherent to the dark channel type of dehazing methods. Yuxiang Shen, Xiaolin Wu 0001, Xiaowei Deng |
ICIP | 2 |
| 2015 | Down-sampling based embedded compression in video systemsabstractWith rapid increase of image resolution in modern video processing and display systems, the bandwidth and power consumption of external memory are becoming serious bottlenecks. This problem can be alleviated by high-fidelity embedded compression (EC) techniques for video frame buffers. Classic lossless or near-lossless coding methods like CALIC are ill suited for embedded systems due to their high complexity. In this work, a new, simple infra-frame EC technique based on downsampling and side-information aided upsampling is developed. Through a study of a family of downsampling schemes, an optimal one is found and analyzed for EC. This downsampling scheme gives birth to the new EC technique. The main idea is to first split an image into blocks, and then adaptively choose different down sampling patterns and upsampling methods to code/decode these blocks. For a memory bandwidth reduction of 60%, the proposed EC system can achieve PSNR above 40dB, while allowing very simple, low-cost real-time hardware realization. A noteworthy novelty of this work is compression without entropy coding. The resulting code stream is of fixed-rate, supporting random access to pixel blocks. Yuxiang Shen, Xiaolin Wu 0001 |
ISCAS | 2 |
| 2015 | Soft binary segmentation-based backlit image enhancementabstractThis paper is concerned with the enhancement of backlit images by compensating for abnormal illumination conditions. The underexposed (backlit) or/and overexposed regions in a backlit image are identified by a soft binary segmentation process that is driven by a Gaussian mixture model. Optimal tone-mapping is performed on the under- and over-exposed regions separately to improve the visual quality. Experimental results demonstrate the efficacy of the proposed restoration method and its advantages over existing image enhancement algorithms in perceptual quality. Zhenhao Li 0001, Xiaolin Wu 0001 |
MMSP | 3 |
| 2015 | On display-camera synchronization for visible light communicationabstractRecently, Wu and Shu proposed an intriguing technique that transmits data from regular information display devices to smart phone cameras without using any additional hardware or interfering the original content intended to display. This preliminary work reveals the possibility of conducting high-speed data transmission with displays in transparency to human eyes; however, in practice, it could introduce visible artifacts and the display-camera synchronization remains an issue that affects data throughput. In this paper, we examine this technique in more realistic application scenarios and put forward a practical solution that hides the data more effectively and greatly improves the robustness of the transmission. Experimental results show that, using the proposed solution, viewers can hardly notice any artifacts caused by the data embedded in videos, and the bit error rate is acceptable under common use conditions. Xiaolin Wu 0001, Xiao Shu |
VCIP | 2 |
| 2014 | Multiple Description Image Coding with Local Random MeasurementsabstractIn this paper, an effective multiple description image coding technique is developed to achieve competitive coding efficiency at low encoder complexity, while being standard compliant. The new technique is particularly suitable for visual communication over packet-switched networks and with resource-deficient wireless devices. To keep the encoder simple and standard compliant, multiple descriptions are produced by quincunx spatial multiplexing. Each side description is a polyphase down sampled version of the input image, but the conventional low-pass filter prior to downsampling is replaced by a local random binary convolution kernel. The pixels of each resulting side description are local random measurements and placed in the original spatial configuration. The advantages of local random measurements are two folds: 1) preservation of high-frequency image features that are otherwise discarded by low-pass filtering, 2) each side description remains a conventional image and can therefore be coded by any standardized codec to remove statistical redundancy of larger scales. The decoder performs joint upsampling of received description(s) and recovers the image from local random measurements in a framework of compressive sensing. Experimental results demonstrate that the proposed multiple description image codec is competitive in rate-distortion performance compared with existing methods, with a unique strength of recovering fine details and sharp edges at low bit rates. Xianming Liu 0005, Xiaolin Wu 0001, Debin Zhao |
DCC | 2 |
| 2014 | Sparsity fine tuning in wavelet domain with application to compressive image reconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this work, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have non-zero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and close-to-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
ICASSP | 2 |
| 2014 | Contrast enhancement with chromaticity error boundabstractAll existing contrast enhancement methods focus on heightening of spatial details in the luminance channel, with no or little consideration of the color fidelity of the processed images; they can introduce highly noticeable distortions in hue and saturation. This long-time much overlooked problem is addressed by this paper. A new algorithm based on Wu's optimal contrast-tone mapping (OCTM) is proposed to make maximal gain of contrast while keeping the hue component the same and at the same time respecting an upper bound on the distortion of saturation. Experimental results demonstrate the superior perceptual quality of the new chromaticity-conserving contrast enhancement algorithm over existing methods. Zhenhao Li 0001, Xiaolin Wu 0001 |
ICIP | 2 |
| 2014 | Image enhancement by entropy maximization and quantization resolution upconversionabstractThis article introduces a new contrast enhancement algorithm of tone-preserving entropy maximization. Its design objective is to present the maximal amount of information content in the enhanced image, or being optimal in an information theoretical sense, while preventing the loss of tone continuity. The resulting optimization problem can be graph-theoretically modeled as the construction of K-edge maximum-weight path, and it can be solved efficiently by dynamic programming. Moreover, the proposed algorithm is made more effective by being combined with a preprocess of image restoration that aims to correct quantization errors caused by the analog-to-digital conversion of image signals. Empirical evidences are provided to demonstrate the superior visual quality obtained by the new image enhancement algorithm. Xiaolin Wu 0001, Guangming Shi |
ICIP | 2 |
| 2014 | Demo: DLP based anti-piracy display systemabstractCamcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. To realize anti-piracy, we uses a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. Based on this difference, we build a prototype system on the platform of DLP® LightCrafter 4500™ which features high speed pattern display. The display system serves as a proof-of-concept of anti-piracy system. Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Chunjia Hu |
VCIP | 3 |
| 2014 | DLP based anti-piracy display systemabstractCamcorder piracy has great impact on the movie industry. Although there are many methods to prevent recording in theatre, no recognized technology satisfies the need of defeating camcorder piracy as well as having no effect on the audience. This paper presents a new projector display technique to defeat camcorder piracy in the theatre using a new paradigm of information display technology, called temporal psychovisual modulation (TPVM). TPVM exploits the difference in image formation mechanisms of human eyes and imaging sensors. The images formed in human vision is continuous integration of the light field while discrete sampling is used in digital video acquisition which has "blackout" period in each sampling cycle. Based on this difference, we can decompose a movie into a set of display frames and broadcast them out at high speed so that the audience can not notice any disturbance, while the video frames captured by camcorder will contain highly objectionable artifacts. The proposed prototype system built on the platform of DLP® LightCrafter 4500™ serves as a proof-of-concept of anti-piracy system. Zhongpai Gao, Guangtao Zhai, Xiaolin Wu 0001, Xiongkuo Min, Cheng Zhi |
VCIP | 3 |
| 2014 | GPU-aided real-time image/video super resolution based on error feedbackabstractSuper resolution is a process to generate high-resolution images from their low-resolution versions. In many applications such as super-HD (4K) TV, super resolution has to be performed in real time. In this paper we propose a real-time image/video super-resolution algorithm, which achieves good performance at low computational cost via off-line learning of interpolation errors in different pixel contexts. The proposed algorithm consists of three stages: fast edge-guided interpolation to generate an initial HR estimation, GPU-aided de-convolution, and error feedback compensation. All three stages can be implemented with GPU to support real-time applications. Experiments demonstrate the competitive performance of the new real-time super-resolution algorithm in both PSNR and visual quality. Yuxiang Shen, Xiaolin Wu 0001, Xiaowei Deng |
VCIP | 2 |
| 2014 | Sparsity Fine Tuning in Wavelet Domain With Application to Compressive Image ReconstructionabstractIn compressive sensing, wavelet space is widely used to generate sparse signal (image signal in particular) representations. In this paper, we propose a novel approach of statistical context modeling to increase the level of sparsity of wavelet image representations. It is shown, contrary to a widely held assumption, that high-frequency wavelet coefficients have nonzero mean distributions if conditioned on local image structures. Removing this bias can make wavelet image representations sparser, i.e., having a greater number of zero and closeto-zero coefficients. The resulting unbiased probability models can significantly improve the performance of existing wavelet-based compressive image reconstruction methods in both PSNR and visual quality. An efficient algorithm is presented to solve the compressive image recovery (CIR) problem using the refined models. Experimental results on both simulated compressive sensing (CS) image data and real CS image data show that the new CIR method significantly outperforms existing CIR methods in both PSNR and visual quality. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2014 | Coded Acquisition of High Frame Rate VideoabstractHigh frame rate video (HFV) is an important investigational tool in sciences, engineering, and military. In ultrahigh speed imaging, the obtainable temporal, spatial, and spectral resolutions are limited by the sustainable throughput of in-camera mass memory, lower bound of exposure time, and illumination conditions. To break these bottlenecks, we propose a new coded video acquisition framework that employs K ≥ 2 cameras, each of which makes random measurements of the video signal in both temporal and spatial domains. For each of the K cameras, this multicamera strategy greatly relaxes the stringent requirements in memory speed, shutter speed, and illumination strength. The recovery of HFV from these random measurements is posed and solved as a large-scale l1 minimization problem by exploiting joint temporal and spatial sparsities of the 3D signal. Three coded video acquisition techniques of varied tradeoffs between performance and hardware complexity are developed: 1) framewise coded acquisition; 2) pixelwise coded acquisition; and 3) columnwise-rowwise coded acquisition. The performances of these techniques are analyzed in relation to the sparsity of the underlying video signal. Simulations of these new HFV capture techniques are carried out and experimental results are reported. Reza Pournaghi, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | On Two-Stage Sequential Coding of Correlated SourcesabstractWe study the problem of two-stage sequential coding (TSSC), which is an extension of sequential coding of correlated sources. Let X and Y be dependent random variables. The network contains two encoders and two decoders: 1) a Y encoder with input Y; 2) an X encoder with inputs X and Y; 3) a Y decoder that reconstructs Y; and 4) an X decoder that reconstructs X. The first stage is traditional sequential coding, where the Y encoder describes Y to both the X decoder and Y decoder, and the X encoder describes X and Y to the X decoder. At the second stage, the Y encoder refines the description of Y, and the X encoder refines the description of X. The TSSC model is a theoretical abstraction of scalable video coding; here, Y and X represent successive frames of a video sequence, and the two stages together give an embedded description that allows the video to be decoded at two distinct rates. We give an inner bound on the rate distortion region for this TSSC model. The tight bound on the rate distortion region is derived when Y must be reconstructed losslessly (in the usual Shannon sense) in the second stage. We also study the minimum total rate of the TSSC model and show that the minimum total rate of one-stage sequential coding cannot be achieved at both stages for jointly Gaussian sources. This theoretical result can shed light on the rate-distortion performance behavior of scalable video coding widely noted by practitioners. Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu |
IEEE Trans. Inf. Theory | 2 |
| 2014 | Generalized Equalization Model for Image EnhancementabstractIn this paper, we propose a generalized equalization model for image enhancement. Based on our analysis on the relationships between image histogram and contrast enhancement/white balancing, we first establish a generalized equalization model integrating contrast enhancement and white balancing into a unified framework of convex programming of image histogram. We show that many image enhancement tasks can be accomplished by the proposed model using different configurations of parameters. With two defining properties of histogram transform, namely contrast gain and nonlinearity, the model parameters for different enhancement applications can be optimized. We then derive an optimal image enhancement algorithm that theoretically achieves the best joint contrast enhancement and white balancing result with trading-off between contrast enhancement and tonal distortion. Subjective and objective experimental results show favorable performances of the proposed algorithm in applications of image enhancement, white balancing and tone correction. Computational complexity of the proposed method is also analyzed. Hongteng Xu, Guangtao Zhai, Xiaolin Wu 0001, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | ℓ2 optimized predictive image coding with ℓ∞ boundabstractIn many scientific, medical and defense applications of image/video compression, an ℓ∞error bound is required. However, pure ℓ∞-optimized image coding, colloquially known as near-lossless image coding, is prone to structured errors such as contours and speckles if the bit rate is not sufficiently high; moreover, previous ℓ∞-based image coding methods suffer from poor rate control. In contrast, the ℓ2error metric aims for average fidelity and hence preserves the subtlety of smooth waveforms better than the ℓ∞error metric and it offers fine granularity in rate control; but pure ℓ2-based image coding methods (e.g., JPEG 2000) cannot bound individual errors as the ℓ∞-based methods can. This paper presents a new compression approach to retain the benefits and circumvent the pitfalls of the two error metrics. Sceuchin Chuah, Sorina Dumitrescu, Xiaolin Wu 0001 |
ICASSP | 3 |
| 2013 | Sparsity-based soft decoding of compressed images in transform domainabstractWe propose a sparsity-based soft decoding approach to restore compressed images directly in the transform domain of compression (DCT domain specifically examined in this paper). Restoring transform coefficients rather than pixel values prevents the propagation of quantization errors in the image domain. As natural images are statistically non-stationary with spatially varying sparse representations, we develop an adaptive block-wise sparsity-based restoration method that learns and exploits local statistics. Specially, for each DCT block, we collect sample blocks via non-local patch grouping to learn a compact dictionary based on principal component analysis. The resulting block-specific dictionary is used to estimate the corresponding DCT coefficients by a technique of collaborative sparse coding, in which the similarity between sample DCT patches used in dictionary construction is further considered. Experimental results are encouraging and demonstrate that the proposed soft decoding approach performs competitively on restoring compressed images against existing methods. Xianming Liu 0005, Xiaolin Wu 0001, Debin Zhao |
ICIP | 2 |
| 2013 | Image enhancement revisited: From first order to second order statisticsabstractThis paper proposes a new image enhancement algorithm in a recently published framework of optimal contrast-tone mapping (OCTM). The new algorithm represents a fundamental departure from traditional histogram-based image enhancement techniques (i.e., histogram equalization and all of its variants), in that second-order rather than first-order statistics is used. Perceptual quality attributes, such as contrast and tone, are quantified by joint distribution of the values of spatially adjacent pixels instead of histograms as of today. The problem of image enhancement is then formulated as one of linear programming, at the heart of which is a joint distribution-based objective function that can accommodate various psychovisual properties related to image quality. The new linear program algorithm for image enhancement is implemented and its superior performance in visual quality is empirically verified, corroborating with our analysis. Xiao Shu, Xiaolin Wu 0001 |
ICIP | 2 |
| 2013 | OPtimal backlight scanning for 3D crosstalk reduction in LCD TVabstractThis work presents a method to determine the optimal backlight scanning signals to minimize crosstalk for time-sequential stereoscopic 3D on LCD TV with active shutter glasses. The solution is obtained through optimization of the variables defined by a model of backlight scanning that considers important aspects like liquid crystal transitions and light diffusion, subject to constraints that ensure the rendition of a uniform backlight. Compared with basic backlight scanning, the proposed method can increase luminance at a given crosstalk level or reduce crosstalk at a given luminance level. Nino Burini, Xiao Shu, Liangbao Jiao, Søren Forchhammer, Xiaolin Wu 0001 |
ICME | 5 |
| 2013 | Low bit-rate image coding via local random down-samplingabstractA common practice in low bit-rate image/video compression is uniform spatial down-sampling at the encoder and upsampling at the decoder. The down-sampling is performed in conjunction with deterministic low-pass filtering (e.g., Gaussian or the alike) to prevent aliasing. The down-sampled image is compressed and decompressed as usual; the upsampling is treated as an image restoration problem. In this paper, we show that the rate-distortion performance of the above low bit-rate image coding system can be improved, if the deterministic low-pass down-sampling filter is replaced by a random convolution kernel. The resulting down-sampled image is a two-dimensional array of local random measurements; this smaller image is still compressible in most cases. Accordingly, the decoder recovers the image from these local random measurements in the framework of compressive sensing. Theoretical analysis is conducted to support the superior performance of the proposed new method over its predecessors, and it is corroborated by our simulation results. At low to medium bit rates, the new method outperforms not only JPEG 2000 but also our earlier low bit-rate image codec CADU, with clear advantages over the competing methods in the reconstruction of high frequency features. In addition, the new method retains the system advantages of low encoder complexity and standard compliance as in CADU. Reza Pournaghi, Xiaolin Wu 0001, Xianming Liu 0005 |
PCS | 2 |
| 2013 | A learning-based method for compressive image recovery
Weisheng Dong, Guangming Shi, Xiaolin Wu 0001, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Joint segmentation and pairing of multispectral chromosome images
Yongqiang Zhao 0001, Xiaolin Wu 0001, Seong G. Kong, Lei Zhang 0006 |
Pattern Anal. Appl. | 2 |
| 2013 | LED Backlight Adjustment for Backward-Compatible Stereoscopic DisplayabstractIt was recently shown that a high-speed optoelectronic display, via a novel signal processing technique called temporal psychovisual modulation (TPVM), can exhibit stereoscopic images to viewers wearing 3-D glasses and clean 2-D images to those without glasses all at the same time. This research aims to improve the above backward-compatible stereoscopy method by adjusting the backlight of today's common LED-lit liquid crystal display systems. Visual quality of the system can be enhanced by jointly optimizing the backlight intensity and the image signal at a negligible extra computational cost. For real-time applications of low-cost consumer electronics, this work also provides a low-complexity solution of backward-compatible stereoscopic display, in which highest 3-D quality is ensured with a small compromise of the 2-D quality. Liangbao Jiao, Xiao Shu, Xiaolin Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2013 | ℓ2 Optimized Predictive Image Coding With ℓ∞ BoundabstractIn many scientific, medical, and defense applications of image/video compression, an [Symbol: see text]∞ error bound is required. However, pure[Symbol: see text]∞-optimized image coding, colloquially known as near-lossless image coding, is prone to structured errors such as contours and speckles if the bit rate is not sufficiently high; moreover, most of the previous [Symbol: see text]∞-based image coding methods suffer from poor rate control. In contrast, the [Symbol: see text]2 error metric aims for average fidelity and hence preserves the subtlety of smooth waveforms better than the ∞ error metric and it offers fine granularity in rate control, but pure [Symbol: see text]2-based image coding methods (e.g., JPEG 2000) cannot bound individual errors as the [Symbol: see text]∞-based methods can. This paper presents a new compression approach to retain the benefits and circumvent the pitfalls of the two error metrics. A common approach of near-lossless image coding is to embed into a DPCM prediction loop a uniform scalar quantizer of residual errors. The said uniform scalar quantizer is replaced, in the proposed new approach, by a set of context-based [Symbol: see text]2-optimized quantizers. The optimization criterion is to minimize a weighted sum of the [Symbol: see text]2 distortion and the entropy while maintaining a strict [Symbol: see text]∞ error bound. The resulting method obtains good rate-distortion performance in both [Symbol: see text]2 and [Symbol: see text]∞ metrics and also increases the rate granularity. Compared with JPEG 2000, the new method not only guarantees lower [Symbol: see text]∞ error for all bit rates, but also it achieves higher PSNR for relatively high bit rates. Sceuchin Chuah, Sorina Dumitrescu, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2013 | Optimal Local Dimming for LC Image Formation With Controllable BacklightingabstractLight emitting diode (LED)-backlit liquid crystal displays (LCDs) hold the promise of improving image quality while reducing the energy consumption with signal-dependent local dimming. However, most existing local dimming algorithms are mostly motivated by simple implementation, and they often lack concern for visual quality. To fully realize the potential of LED-backlit LCDs and reduce the artifacts that often occur in current systems, we propose a novel local dimming technique that can achieve the theoretical highest fidelity of intensity reproduction in either l(1) or l(2) metrics. Both the exact and fast approximate versions of the optimal local dimming algorithm are proposed. Simulation results demonstrate superior performances of the proposed algorithm in terms of visual quality and power consumption. Xiao Shu, Xiaolin Wu 0001, Søren Forchhammer |
IEEE Trans. Image Process. | 2 |
| 2012 | Context Modeling and Correction of Quantization Errors in Prediction LoopabstractIn lossy predictive coding of Differential Pulse Code Modulation (DPCM) type, quantization performed in the prediction loop induces propagation of quantization errors, resulting in biased predictions of the subsequent samples. In this work, we aim to alleviate the negative effect of quantization errors on the robustness of prediction. We propose some practical techniques for context modeling of quantization errors and cancelation of estimation biases in the DPCM reconstruction. The resulting refined estimates are fed into the prediction to improve coding efficiency. When applied to 1D audio and 2D image signals, the proposed techniques can reduce the bit rate and at the same time improve the PSNR performance significantly. Jiantao Zhou 0001, Xiaolin Wu 0001 |
DCC | 2 |
| 2012 | On monotonicity of image quality metricsabstractPerceptual image quality assessment (IQA) is an important research topic of visual signal processing both in its own right and for its utility in designing various optimal image processing and coding algorithms. This work is concerned with an issue that has been largely overlooked by the research community of IQA, that is, the monotonicity, or lack of it, between the subjective scores and the predictions of image quality metrics (IQM) for images with compression artifacts. We analyze the data of several well-known databases for IQA and expose among them a large number of instances of non-monotonicity between subjective and objective quality scores. Further, we observe that a nonlinear dynamical model of 3D cusp catastrophe can well explain the intricate relationship between the subjective and objective quality scores. Our findings identify an inherent flaw of current signal-distance or fidelity-based IQMs, which neglect the psycho-physiological aspect of human visual perception. This research suggests a new direction of IQA research and it also sheds light on the design of subjective quality evaluation process. Guangtao Zhai, Xiaolin Wu 0001 |
ICASSP | 2 |
| 2012 | Image dependent energy-constrained local backlight dimmingabstractIn this work, we consider and propose two extensions to an optimization-based image dependent backlight dimming algorithm. The first extension introduces error weighting based on human perception of luminance, aiming to improve the perceived image quality; the second extension adds an adjustable term for power consumption to the cost function, allowing flexible power management. Experimental results show that the proposed solution can achieve better results than other algorithms at several power consumption levels. Nino Burini, Ehsan Nadernejad, Jari Korhonen, Søren Forchhammer, Xiaolin Wu 0001 |
ICIP | 5 |
| 2012 | Low bit-rate video coding via mode-dependent adaptive regression for wireless visual communicationsabstractIn this paper, a practical video coding scheme is developed to realize state-of-the-art video coding efficiency with lower encoder complexity at low bit-rate, while supporting standard compliance and error resilience. Such an architecture is particularly attractive for wireless visual communications. At the encoder, multiple descriptions of a video sequence are generated in the spatio-temporal domain by temporal multiplexing and spatial adaptive downsampling. The resulting side descriptions are interleaved with each other in temporal domain, and still with conventional square sample grids in spatial domain. As such, each side description can be compressed without any change to existing video coding standards. At the decoder, each side description is first decompressed, and then reconstructed to original resolution with the help of the other side description. In this procedure, the decoder recover the original video sequence in a constrained least squares regression process, using 2D or 3D piecewise autoregressive model according to different prediction modes. In this way, the spatial and temporal correlation is sufficiently explored to achieve superior quality. Experiment results demonstrate the proposed video coding scheme outperforms H.264 in rate-distortion performance at low bit-rates and achieves superior visual quality at medium bit-rates as well. Xianming Liu 0005, Xiaolin Wu 0001, Xinwei Gao, Debin Zhao, Wen Gao 0001 |
VCIP | 2 |
| 2012 | Color demosaicking with an image formation model and adaptive PCA
Dahua Gao, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Model-based adaptive resolution upconversion of degraded images
Xiaolin Wu 0001, Xiangjun Zhang, Guangming Shi |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Image reconstruction with locally adaptive sparsity and nonlocal robust regularization
Weisheng Dong, Guangming Shi, Xin Li 0005, Lei Zhang 0006, Xiaolin Wu 0001 |
Signal Process. Image Commun. | 5 |
| 2012 | Edge-Based Perceptual Image CodingabstractWe develop a novel psychovisually motivated edge-based low-bit-rate image codec. It offers a compact description of scale-invariant second-order statistics of natural images, the preservation of which is crucial to the perceptual quality of coded images. Although being edge based, the codec does not explicitly code the edge geometry. To save bits on edge descriptions, a background layer of the image is first coded and transmitted, from which the decoder estimates the trajectories of significant edges. The edge regions are then refined by a residual coding technique based on edge dilation and sequential scanning in the edge direction. Experimental results show that the new image coding technique outperforms the existing ones in both objective and perceptual quality, particularly at low bit rates. Xiaolin Wu 0001, Guangming Shi, Xiaotian Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Binned Progressive Quantization for Compressive SensingabstractCompressive sensing (CS) has been recently and enthusiastically promoted as a joint sampling and compression approach. The advantages of CS over conventional signal compression techniques are architectural: the CS encoder is made signal independent and computationally inexpensive by shifting the bulk of system complexity to the decoder. While these properties of CS allow signal acquisition and communication in some severely resource-deprived conditions that render conventional sampling and coding impossible, they are accompanied by rather disappointing rate-distortion performance. In this paper, we propose a novel coding technique that rectifies, to a certain extent, the problem of poor compression performance of CS and, at the same time, maintains the simplicity and universality of the current CS encoder design. The main innovation is a scheme of progressive fixed-rate scalar quantization with binning that enables the CS decoder to exploit hidden correlations between CS measurements, which was overlooked in the existing literature. Experimental results are presented to demonstrate the efficacy of the new CS coding technique. Encouragingly, on some test images, the new CS technique matches or even slightly outperforms JPEG. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Image Process. | 2 |
| 2012 | Model-Assisted Adaptive Recovery of Compressed Sensing with Imaging ApplicationsabstractIn compressive sensing (CS), a challenge is to find a space in which the signal is sparse and, hence, faithfully recoverable. Since many natural signals such as images have locally varying statistics, the sparse space varies in time/spatial domain. As such, CS recovery should be conducted in locally adaptive signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary, existing CS reconstruction methods use a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem, we propose a new framework for model-guided adaptive recovery of compressive sensing (MARX) and show how a 2-D piecewise autoregressive model can be integrated into the MARX framework to make CS recovery adaptive to spatially varying second order statistics of an image. In addition, MARX offers a mechanism of characterizing and exploiting structured sparsities of natural images, greatly restricting the CS solution space. Simulation results over a wide range of natural images show that the proposed MARX technique can improve the reconstruction quality of existing CS methods by 2-7 dB. Xiaolin Wu 0001, Weisheng Dong, Xiangjun Zhang, Guangming Shi |
IEEE Trans. Image Process. | 1 |
| 2012 | A Psychovisual Quality Metric in Free-Energy PrincipleabstractIn this paper, we propose a new psychovisual quality metric of images based on recent developments in brain theory and neuroscience, particularly the free-energy principle. The perception and understanding of an image is modeled as an active inference process, in which the brain tries to explain the scene using an internal generative model. The psychovisual quality is thus closely related to how accurately visual sensory data can be explained by the generative model, and the upper bound of the discrepancy between the image signal and its best internal description is given by the free energy of the cognition process. Therefore, the perceptual quality of an image can be quantified using the free energy. Constructively, we develop a reduced-reference free-energy-based distortion metric (FEDM) and a no-reference free-energy-based quality metric (NFEQM). The FEDM and the NFEQM are nearly invariant to many global systematic deviations in geometry and illumination that hardly affect visual quality, for which existing image quality metrics wrongly predict severe quality degradation. Although with very limited or even without information on the reference image, the FEDM and the NFEQM are highly competitive compared with the full-reference SSIM image quality metric on images in the popular LIVE database. Moreover, FEDM and NFEQM can measure correctly the visual quality of some model-based image processing algorithms, for which the competing metrics often contradict with viewers' opinions. Guangtao Zhai, Xiaolin Wu 0001, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | l2 Restoration of l∞-Decoded Images Via Soft-Decision EstimationabstractThe l(∞)-constrained image coding is a technique to achieve substantially lower bit rate than strictly (mathematically) lossless image coding, while still imposing a tight error bound at each pixel. However, this technique becomes inferior in the l(2) distortion metric if the bit rate decreases further. In this paper, we propose a new soft decoding approach to reduce the l(2) distortion of l(∞)-decoded images and retain the advantages of both minmax and least-square approximations. The soft decoding is performed in a framework of image restoration that exploits the tight error bounds afforded by the l(∞)-constrained coding and employs a context modeler of quantization errors. Experimental results demonstrate that the l(∞)-constrained hard decoded images can be restored to gain more than 2 dB in peak signal-to-noise ratio PSNR, while still retaining tight error bounds on every single pixel. The new soft decoding technique can even outperform JPEG 2000 (a state-of-the-art encoder-optimized image codec) for bit rates higher than 1 bpp, a critical rate region for applications of near-lossless image compression. All the coding gains are made without increasing the encoder complexity as the heavy computations to gain coding efficiency are delegated to the decoder. Jiantao Zhou 0001, Xiaolin Wu 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2011 | Progressive Quantization of Compressive Sensing MeasurementsabstractCompressive sensing (CS) is recently and enthusiastically promoted as a joint sampling and compression approach. The advantages of CS over conventional signal compression techniques are architectural: the CS encoder is made signal independent and computationally inexpensive by shifting the bulk of system complexity to the decoder. While these properties of CS allow signal acquisition and communication in some severely resource-deprived conditions that render conventional sampling and coding impossible, they are accompanied by rather disappointing rate-distortion performance. In the present work we propose a novel coding technique that rectifies, to certain extent, the problem of poor compression performance of CS and at the same time maintains the simplicity and universality of the current CS encoder design. The main innovation is a scheme of progressive fixed-rate scalar quantization with binning that enables the CS decoder to exploit hidden correlations between CS measurements, which was overlooked in the existing literature. Experimental results are presented to demonstrate the efficacy of the new CS coding technique. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
DCC | 2 |
| 2011 | High-Fidelity Image Compression for High-Throughput and Energy-Efficient CamerasabstractWe propose a new encoder-friendly image compression strategy for high-throughput cameras and other scenarios of resource-constrained encoders. The encoder performs $\ell_infty$-constrained predictive coding (DPCM coupled with uniform scalar quantizer), while the decoder solves an inverse problem of $\ell_2$ restoration of $\ell_\infty$-coded images. Although designed for minimum encoder complexity, the new codec outperforms the state-of-the-art encoder-centralized image codecs such as JPEG 2000 in PSNR for bit rates higher than 1.2 bpp, while maintaining much tighter $\ell_\infty$ error bounds as well. This is achieved through exploiting the tight error bound on each pixel naturally offered by the $\ell_\infty$-constrained encoder and by locally adaptive image modeling. Xiaolin Wu 0001, Jiantao Zhou 0001 |
DCC | 1 |
| 2011 | A hidden Markov model-based methodology for intra-field video deinterlacingabstractThis paper presents a new technique of hidden Markov model (HMM) for video deinterlacing. Existing deinterlacing algorithms estimate missing pixels of an absent row in an interlaced frame on a sample-by-sample basis. In contrast, the proposed HMM-based deinterlacing technique adopts an approach of sequence estimation and makes a joint decision on the row of missing pixels as a whole. This allows a more thorough exploitation of the spatial correlation of the image signal. The HMM-based sequence estimation technique is coupled with a number of existing spatial deinterlacing algorithms in the literature to boost their performance. Experimental results show that HMM can significantly improve the deinterlacing results in both PSNR measure and subjective visual quality. Amin Behnad, Konstantinos N. Plataniotis, Xiaolin Wu 0001 |
ICIP | 3 |
| 2011 | On sparse representations of color imagesabstractWe investigate an intrinsic and useful form of sparsity of color images that was largely overlooked in the literature of image/video processing. This sparsity of multispectral images is revealed and formulated by modeling the image formation process. The underlying new sparse representations of color images are general and can be exploited to improve the performance of existing image restoration algorithms, such as denoising, deblurring, and resolution upconversion. Xiaolin Wu 0001, Guangtao Zhai |
ICIP | 1 |
| 2011 | Hybrid parametric-nonparametric modeling with application to natural image upsamplingabstractLinear autoregressive (AR) model is widely used in signal processing. Usually the AR models are solved by classical least square (LS) method. An important issue with the LS solution of the AR model, which has been seemingly overlooked, is its numerical stability. The issue is related to the rank condition of the design matrix. We observed, in case of natural images, that the probability of numerical rank deficiency is rather high, roughly thirty-five per cent, due to discrete nature and structures of the digital images. Without care numerical rank deficiency can adversely affect the parameter estimation of the AR model. In this paper we use the rank revealing QR (RRQR) factorization to select optimal subset from the design matrix so as to effectively lower the condition number of the system. By removing the ill conditioned part of the right orthogonal matrix of the RRQR decomposition, we obtain a robust truncated solution to the linear system. On the other hand, for natural images, the unselected data tend to highly correlate with the pixel being modeled, and their exclusion from the modeling process waste valuable information. To avoid this loss we recycle the data including those discard by the parametric AR estimator into a nonparametrgic model of nonlocal type. Interestingly, the data that cause ill condition to the parametric AR model are of high quality for the non-local nonparametric modeling. Therefore, an approach of hybrid parametric-nonparametric modeling can make the best use of data and improve the model performance. The hybrid modeling approach is applied to image resolution upconversion, and it greatly improves the performance of the state-of-the-art image interpolator, achieving a gain of 3dB or more in PSNR in some cases. Guangtao Zhai, Xiaolin Wu 0001 |
ICIP | 2 |
| 2011 | Noise estimation using statistics of natural imagesabstractWe develop a framework for estimating noises of natural images using two important properties of natural image statistics: high kurtosis and scale invariance of natural images in certain transform domains. We examine the effects of additive independent noise on the third and fourth moments of the transformed image signal (skewness and kurtosis). By exploring the said priors of high kurtosis and scale invariance of natural image statistics in 2D discrete cosine transform domain and random unitary transform domain, we derive constrained nonlinear optimization algorithms for accurate estimation of noise variance. Simulation and comparative study show that the proposed approach is capable of estimating the variance of Gaussian additive noise with a relative error as low as one percent. Moreover, the new estimation approach is shown to be effective on multiplicative-additive compound noises as well. This work can significantly improve the performance of existing denoising techniques that require the noise variance as a critical parameter. Guangtao Zhai, Xiaolin Wu 0001 |
ICIP | 2 |
| 2011 | L2 restoration of L∞-decoded images with context modelingabstractThe L∞-constrained image coding is a technique to achieve substantially lower bit rate than strictly (mathematically) lossless image coding while still imposing a tight error bound at each pixel (colloquially referred to as near-lossless image coding). However, this technique becomes inferior in the L2distortion metric if the bit rate decreases further. We propose a new soft decoding approach to reduce the L2distortion of L∞-coded images, benefiting from the advantages of both minmax and mean square approximations. This is made possible by context modeling of quantization distortions and by exploiting the L∞bound inherent to near-lossless coding in a framework of image restoration. In addition, the proposed soft decoding approach offers an asymmetric high-fidelity image compression solution: the encoder is of low complexity with heavy computations of gaining coding efficiency performed by the decoder. Experimental results demonstrate that the new soft decoding approach can improve the PSNR of L∞-decoded images by more than 1 dB, and it can even outperform JPEG 2000 (a state-of-the-art encoder-optimized image codec) for bit rates higher than 1.17 bpp, while achieving much tighter L∞error bound. Jiantao Zhou 0001, Xiaolin Wu 0001 |
ICIP | 2 |
| 2011 | A psychovisually tuned image codecabstractA psychovisual quality driven image codec exploiting the psychological and neurological process of visual perception is proposed in this paper. Recent findings in brain theory and neuroscience suggest that visual perception is a process of fitting brain's internal generative model to the outside retina stimuli. And the psychovisual quality is related to how accurately visual sensory data can be explained by the internal generative model. Therefore, the design criterion of our psychovisually tuned image compression system is to find a compact description of the optimal generative model from the input image on the encoding end, which is then used to regenerate the output image on the decoding end. By noting an important finding from empirical natural image statistics that natural images have scale invariant features in the pixels' high order statistics, the generative model can be efficiently compressed through model preserving spatial downsampling on the encoder. And the decoder can reverse the process with a model preserving upsampling module to generate the decoded image. The proposed system is fully standard complaint because the downsampled image can be compressed with any exiting codec (JPEG2000 in this work). The proposed algorithm is shown to systematically outperform JPEG2000 in a wide bit rate range in terms of both subjective and objective qualities. Guangtao Zhai, Xiaolin Wu 0001 |
MMSP | 2 |
| 2011 | Multiple description video coding against both erasure and bit errors by compressive sensingabstractWe propose a novel multiple description video coding (MDVC) technique for robust video transmission via lossy networks of both packet erasure and bit errors. The new MDVC technique is designed to meet two objectives: ultra fine description granularity and low encoder complexity, which allow resource-deprived video transmitters (e.g., smart phones) to operate in time-varying adverse network conditions. These design goals are met by the signal acquisition and coding strategy of compressive sensing, and by locally adaptive sparse representation of video signals. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
VCIP | 2 |
| 2011 | On Computation of Performance Bounds of Optimal Index AssignmentabstractChannel-optimized index assignment of source codewords is arguably the simplest way of improving transmission error resilience, while keeping the source and/or channel codes intact. But optimal design of index assignment is an instance of quadratic assignment problem (QAP), one of the hardest optimization problems in the NP-complete class. In this work we make a progress in the research of index assignment optimization. We apply some recent results of QAP research to compute the strongest lower bounds so far for channel distortion of BSC among all index assignments. The strength of the resulting lower bounds is validated by comparing them against the upper bounds produced by heuristic index assignment algorithms. Xiaolin Wu 0001, Hans D. Mittelmann, Jia Wang 0004 |
IEEE Trans. Commun. | 1 |
| 2011 | Image Deblurring and Super-Resolution by Adaptive Sparse Domain Selection and Adaptive RegularizationabstractAs a powerful statistical image modeling technique, sparse representation has been successfully used in various image restoration applications. The success of sparse representation owes to the development of the l(1)-norm optimization techniques and the fact that natural images are intrinsically sparse in some domains. The image restoration quality largely depends on whether the employed sparse domain can represent well the underlying image. Considering that the contents can vary significantly across different images or different patches in a single image, we propose to learn various sets of bases from a precollected dataset of example image patches, and then, for a given patch to be processed, one set of bases are adaptively selected to characterize the local sparse domain. We further introduce two adaptive regularization terms into the sparse representation framework. First, a set of autoregressive (AR) models are learned from the dataset of example image patches. The best fitted AR models to a given patch are adaptively selected to regularize the image local structures. Second, the image nonlocal self-similarity is introduced as another regularization term. In addition, the sparsity regularization parameter is adaptively estimated for better image restoration performance. Extensive experiments on image deblurring and super-resolution validate that by using adaptive sparse domain selection and adaptive regularization, the proposed method achieves much better results than many state-of-the-art algorithms in terms of both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2011 | A Linear Programming Approach for Optimal Contrast-Tone MappingabstractThis paper proposes a novel algorithmic approach of image enhancement via optimal contrast-tone mapping. In a fundamental departure from the current practice of histogram equalization for contrast enhancement, the proposed approach maximizes expected contrast gain subject to an upper limit on tone distortion and optionally to other constraints that suppress artifacts. The underlying contrast-tone optimization problem can be solved efficiently by linear programming. This new constrained optimization approach for image enhancement is general, and the user can add and fine tune the constraints to achieve desired visual effects. Experimental results demonstrate clearly superior performance of the new approach over histogram equalization and its variants. Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Adaptive Sequential Prediction of Multidimensional Signals With Applications to Lossless Image CodingabstractWe investigate the problem of designing adaptive sequential linear predictors for the class of piecewise autoregressive multidimensional signals, and adopt an approach of minimum description length (MDL) to determine the order of the predictor and the support on which the predictor operates. The design objective is to strike a balance between the bias and variance of the prediction errors in the MDL criterion. The predictor design problem is particularly interesting and challenging for multidimensional signals (e.g., images and videos) because of the increased degree of freedom in choosing the predictor support. Our main result is a new technique of sequentializing a multidimensional signal into a sequence of nested contexts of increasing order to facilitate the MDL search for the order and the support shape of the predictor, and the sequentialization is made adaptive on a sample by sample basis. The proposed MDL-based adaptive predictor is applied to lossless image coding, and its performance is empirically established to be the best among all the results that have been published till present. Xiaolin Wu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Layered Multicast With Inter-Layer Network Coding for Multimedia StreamingabstractMultirate multicast is a powerful methodology of multimedia communication in heterogenous networks. A variant of multirate multicast motivated by scalable multimedia streaming is layered multicast, where the transmitted signal is presented in successive data layers. With recent advances of network coding theory, many layered multicast schemes using network coding have been proposed to improve the performance of traditional routing-based layered multicast. They divide the network into different layers and construct a unirate multicast network code for each layer. However, these schemes do not perform network coding between data layers, and consequently cannot realize the full potential of network coding. In this paper, we propose a novel approach to layered multicast that allows network coding of data in different layers. This relaxation lends the proposed scheme greater flexibility in optimizing the data flow than previous layered solutions, and thus achieves higher throughput. Mingkai Shao, Sorina Dumitrescu, Xiaolin Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2010 | Information Flows in Video CodingabstractWe study information theoretical performance of common video coding methodologies at the frame level. Via an abstraction of consecutive video frames as correlated random variables, many existing video coding techniques, including the baseline of MPEG-x and H.26x, the scalable coding and the distributed video coding, can have corresponding information theoretical models. The theoretical achievable rate distortion regions have been completely solved for some systems while for others remain open. We show that the achievable rate region of sequential coding equals to that of predictive coding for Markov sources. We give a theoretical analysis of the coding efficiency of B frames in the popular hybrid video coding architecture, bringing new understanding of the current practice. We also find that distributed sequential video coding generally incurs a performance loss if the source is not Markov. Jia Wang 0004, Xiaolin Wu 0001 |
DCC | 2 |
| 2010 | On Computation of Performance Bounds of Optimal Index AssignmentabstractChannel-optimized index assignment of source codewords is arguably the simplest way of improving transmission error resilience, while keeping the source and/or channel codes intact. But optimal design of index assignment is an instance of quadratic assignment problem (QAP), one of the hardest optimization problems in the NP-complete class. In this paper we make a progress in the research of index assignment optimization. We apply some recent results of QAP research to compute the strongest lower bounds so far for channel distortion of BSC among all index assignments. The strength of the resulting lower bounds is validated by comparing them against the upper bounds produced by heuristic index assignment algorithms. Xiaolin Wu 0001, Hans D. Mittelmann, Jia Wang 0004 |
DCC | 1 |
| 2010 | A New Algorithmic Approach for Contrast Enhancement
Xiaolin Wu 0001 |
ECCV (6) | 1 |
| 2010 | Image interpolation with hidden Markov modelabstractWe propose an adaptive image interpolation technique based on hidden Markov modeling (HMM) and maximum a posterior (MAP) estimation. The HMM incorporates the statistics of high resolution (HR) images into the interpolation process and the MAP estimation exploits high-order statistical dependency between pixels. Experimental results show that the HMM-based image interpolation technique can reproduce cleaner and sharper image details than its predecessors, while suppressing common interpolation artifacts such as ringing, jaggies, and blurring. Amin Behnad, Xiaolin Wu 0001 |
ICASSP | 2 |
| 2010 | Adaptive structured recovery of compressive sensing via piecewise autoregressive modelingabstractIn compressive sensing (CS) a challenge is to find a space in which the signal is sparse and hence recoverable faithfully and efficiently. Given the nonstationarity of many natural signals such as images, the sparse space varies in time/spatial domain. As such, CS recovery should be conducted in locally adaptive, signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary most CS methods seek for a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem we propose a new framework for model-guided adaptive recovery of compressive sensing (MARX), and show how a piecewise autoregressive model can be integrated into the MARX framework to adapt to changing second order statistics of a signal in CS recovery. In addition, MARX offers a powerful mechanism of characterizing and exploiting structured sparsities of a signal, greatly restricting the CS solution space. A case study on CS-acquired images shows that the proposed MARX technique can increase the reconstruction quality by up to 8 dB over existing methods. Xiaolin Wu 0001, Xiangjun Zhang |
ICASSP | 1 |
| 2010 | Directional image interpolation with ANOVA methodologyabstractThis paper proposes a directional image interpolation technique based on analysis of variance (ANOVA). ANOVA, a hypothesis testing methodology, is used to decide on the interpolation direction. Experimental results show the proposed ANOVA-based image interpolation technique preserves edges and fine structures of an image and suppresses common interpolation artifacts (e.g. ringing, blurring, and jaggies). Amin Behnad, Konstantinos N. Plataniotis, Xiaolin Wu 0001 |
ICIP | 3 |
| 2010 | Edge-based image coding at low bit-rateabstractWe propose a novel edge-based low bit-rate image codec. Although being edge based, the codec does not explicitly code the edge geometry. To avoid spending high bit budget on edges, a background layer of the image is first coded and transmitted, from which the decoder estimates the trajectories of significant edges. The edge regions are then refined by a residual coding technique based on edge dilation and sequential scanning in the edge direction. Experimental results show that the new image coding technique outperforms existing ones in both objective and perceptual quality, particularly at low bit rates. Xiaolin Wu 0001, Guangming Shi, Xiaotian Wang 0001 |
ICIP | 2 |
| 2010 | Color demosaicking with sparse representationsabstractColor demosaicking is an ill-posed inverse problem of image restoration. The performance of a color demosaicking algorithm depends on how thoroughly it can exploit domain knowledge to confine the solution space for the underlying true color image. We propose a sparsity-based ℓ1minimization technique for color demosaicking that exploits both interband and intra-band sparse representations of natural images. In some of most challenging cases of color demosaicking, the proposed technique outperforms those published in the literature by a significant margin in both PSNR and visual quality. Xiaolin Wu 0001, Dahua Gao, Guangming Shi, Danhua Liu |
ICIP | 1 |
| 2010 | High frame rate video capture by multiple cameras with coded exposureabstractWe propose novel techniques for multiple conventional video cameras to capture high frame rate (HFR) video without sacrificing spatial resolution. The coded exposure technology is used to make random measurements of the HFR video in temporal domain. The recovery of the HFR video from the random measurements is based on sparsities of the 3D video signal. Xiaolin Wu 0001, Reza Pournaghi |
ICIP | 1 |
| 2010 | Improvement of H.264 SVC by model-based adaptive resolution upconversionabstractH.264 SVC extension, as the state of art scalable video coding standard, can offer a single code stream to serve diverse communication bandwidths and display resolutions. However, the rate-distortion performance of H.264 SVC is still inferior to the non-scalable H.264 AVC. To reduce the performance gap between H.264 SVC and H.264 AVC, we propose a model-based adaptive resolution upconversion algorithm to improve the precision of the H.264 SVC inter-layer prediction. The new algorithm treats the up-sampling of video frames as an inverse problem of initial H.264 SVC down-sampling operation, and it significantly improves the performance of current H.264 SVC by optimally reversing the down-sampling filter. Xiaolin Wu 0001, Mingkai Shao, Xiangjun Zhang |
ICIP | 1 |
| 2010 | Joint color decrosstalk and demosaicking for CFA camerasabstractThe problem of crosstalk between different color bands was seemingly overlooked by existing color demosaicking algorithms. In this paper we propose a new joint demosaicking and decrosstalk technique that corrects channel crosstalks by adaptive least-squares inverse filtering. The new technique integrates the operations of deconvolution for crosstalk removal and interpolation for color demosaicking, and it introduces a general framework in which any spatially varying crosstalks and varying spatial-spectral correlations can be modeled and factored into the color reproduction. Simulation results show that the proposed technique is highly effective and capable to obtain both high color fidelity and sharp, clean spatial details. Xiaolin Wu 0001, Xiangjun Zhang |
ICIP | 1 |
| 2010 | Live demonstration: Spatial-temporal color video reproduction from noisy CFA sequence track: Digital signal processingabstractThis demonstration shows a spatial-temporal denoising and demosaicking scheme for noisy CFA videos. This scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. The experimental results showed that this scheme achieves promising color video reproduction in terms of both PSNR and visual perception. Lei Zhang 0006, Weisheng Dong, Chiu-Wai Hui, Xiaolin Wu 0001, Guangming Shi |
ISCAS | 4 |
| 2010 | Video super-resolution for dual-mode digital cameras via scene-matched learningabstractMany consumer digital cameras support dual shooting mode of both low-resolution (LR) video and high-resolution (HR) image. By periodically switching between the video and image modes, this type of cameras make it possible to super-resolve the LR video with the assistance of neighboring HR still images. We propose a model-based video super-resolution (VSR) technique for the above dual-mode cameras. A HR video frame is modeled as a 2D piecewise autoregressive (PAR) process. The PAR model parameters are learnt from the HR still images inserted between LR video frames. By registering the LR video frames and the HR still images, we base the learning on sample statistics that matches the scene to be constructed. The resulting PAR model is more accurate and robust than if the model parameters are estimated from the LR video frames without referring to the HR images or from a training set. Aided by the powerful scene-matched model the LR video frame is upsampled to the resolution of the HR image via adaptive interpolation. As such, the proposed VSR technique does not require explicit motion estimation of subpixel precision nor the solution of a large-scale inverse problem. The new VSR technique is competitive in visual quality against existing techniques with a fraction of the computational cost. Guangtao Zhai, Xiaolin Wu 0001 |
MMSP | 2 |
| 2010 | Super-resolution with nonlocal regularized sparse representationabstractThe reconstruction of a high resolution (HR) image from its low resolution (LR) counterpart is a challenging problem. The recently developed sparse representation (SR) techniques provide new solutions to this inverse problem by introducing the l1-norm sparsity prior into the super-resolution reconstruction process. In this paper, we present a new SR based image super-resolution by optimizing the objective function under an adaptive sparse domain and with the nonlocal regularization of the HR images. The adaptive sparse domain is estimated by applying principal component analysis to the grouped nonlocal similar image patches. The proposed objective function with nonlocal regularization can be efficiently solved by an iterative shrinkage algorithm. The experiments on natural images show that the proposed method can reconstruct HR images with sharp edges from degraded LR images. Weisheng Dong, Guangming Shi, Lei Zhang 0006, Xiaolin Wu 0001 |
VCIP | 4 |
| 2010 | Achieving the rate-distortion bound with low-density generator matrix codesabstractIt is shown that binary low-density generator matrix codes can achieve the rate-distortion bound of discrete memoryless sources with general distortion measure via multilevel quantization. A practical encoding scheme based on the survey-propagation algorithm is proposed. The effectiveness of the proposed scheme is verified through simulation. Zhibin Sun, Mingkai Shao, Jun Chen 0005, Kon Max Wong, Xiaolin Wu 0001 |
IEEE Trans. Commun. | 5 |
| 2010 | Index assignment optimization for joint source-channel MAP decodingabstractChannel-optimized quantizer index assignment and maximum a posteriori (MAP) decoding have been extensively studied for error-resilient communications. An interesting and largely untreated problem is how to optimize the index assignment with respect to joint source-channel MAP decoding. In this paper we formulate the above problem as one of quadratic assignment, and discuss its solutions from very general to some special cases. For highly correlated Gaussian Markov sources and Hamming distortion, we can construct the optimal index assignment analytically. For general cases, simulated annealing algorithm is adopted to search for the optimal index assignment. Experimental results are presented to demonstrate the performance improvement of the index assignments optimized for MAP decoding over those designed for hard-decision decoding (e.g. Gray code). The reduction of symbol error rate and mean squared error can be as large as 40% and 50% respectively for highly correlated Gaussian Markov sources. Xiaolin Wu 0001 |
IEEE Trans. Commun. | 2 |
| 2010 | Spatial-Temporal Color Video Reconstruction From Noisy CFA SequenceabstractSingle-sensor digital video cameras use a color filter array (CFA) to capture video and a color demosaicking (CDM) procedure to reproduce the full color sequence. The reproduced video frames suffer from the inevitable sensor noise introduced in the video acquisition process. This paper presents a spatial-temporal denoising and demosaicking scheme that works without explicit motion estimation. We first perform patch based denoising on the mosaic CFA video. For each CFA patch to be denoised, similar patches are selected within a local spatial-temporal neighborhood. The principal component analysis is performed on the selected patches to remove noise. We then apply an initial single-frame CDM to the denoised CFA data, and subsequently post-process the demosaicked frames by exploiting the spatial-temporal redundancy to reduce the color artifacts. The experimental results on simulated and real noisy CFA sequences demonstrate that the proposed spatial-temporal CFA video denoising and demosaicking scheme can significantly reduce the noise-caused color artifacts and effectively preserve the image edge structures. Lei Zhang 0006, Weisheng Dong, Xiaolin Wu 0001, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | Joint Color Decrosstalk and Demosaicking for CFA CamerasabstractIn interest of low cost, low power consumption, and compact size, most digital cameras adopt a design of single sensor array coupled with a color filter array. This design inevitably suffers, due to physical characteristics of the optical and semiconductor components and the imperfection of manufacturing, from the problem of crosstalk between different color channels. Channel crosstalk can desaturate colors and blur image details, but the problem was seemingly overlooked by existing color demosaicking algorithms. To rectify this deficiency we propose a new joint demosaicking and decrosstalk technique that counters the effects of channel crosstalks by adaptive least-squares inverse filtering. The new technique integrates the operations of deconvolution for crosstalk removal and interpolation for color demosaicking, and it introduces a general framework in which any spatially varying crosstalks and varying spatial-spectral correlations can be modeled and factored into the color reproduction. Simulation results show that the proposed technique is highly effective and capable to obtain both high color fidelity and sharp, clean spatial details. Xiaolin Wu 0001, Xiangjun Zhang |
IEEE Trans. Image Process. | 1 |
| 2010 | On Linfinity Properties of Multiresolution Scalar QuantizersabstractWe investigate the max-norm (L∞) properties of multiresolution scalar quantizers (MRSQ). The multiresolution requirement imposes nontrivial constraints on the maximum distortion at each level of the quantizer. To quantify these constraints, we define the overall multiresolutionL∞distortion of an MRSQ to be a weighted sum ofL∞distortions over all refinement levels of the MRSQ. We then seek MRSQ constructions that minimize this average-max distortion measure. An interesting relationship between this problem and the structure of Huffman code trees is established. Lower bounds for the average-max distortion are derived based on this relationship. The derivation of these lower bounds also lead to efficient dynamic programming heuristic solutions. Nima Sarshar, Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2010 | On the sum rate of Gaussian multiterminal source coding: new proofs and resultsabstractWe show that the lower bound on the sum rate of the direct and indirect Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the minimum mean squared error (MMSE) estimation and the so-called reduced optimal linear estimation, through which an intimate connection between the lower bound and the Berger-Tung upper bound is revealed. We give a new proof of the minimum sum rate of the indirect Gaussian multiterminal source coding problem (i.e., the Gaussian CEO problem). For the direct Gaussian multiterminal source coding problem, we derive a general lower bound on the sum rate and establish a set of sufficient conditions under which the lower bound coincides with the Berger-Tung upper bound. We show that the sufficient conditions are satisfied for a class of sources and distortion constraints; in particular, they hold for arbitrary positive definite source covariance matrices in the high-resolution regime. In contrast with the existing proofs, the new method does not rely on Shannon's entropy power inequality. Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 3 |
| 2009 | Model-Guided Adaptive Recovery of Compressive SensingabstractFor the new signal acquisition methodology of compressive sensing (CS) a challenge is to find a space in which the signal is sparse and hence recoverable faithfully. Given the nonstationarity of many natural signals such as images, the sparse space is varying in time or spatial domain. As such, CS recovery should be conducted in locally adaptive, signal-dependent spaces to counter the fact that the CS measurements are global and irrespective of signal structures. On the contrary existing CS reconstruction methods use a fixed set of bases (e.g., wavelets, DCT, and gradient spaces) for the entirety of a signal. To rectify this problem we propose a new model-based framework to facilitate the use of adaptive bases in CS recovery. In a case study we integrate a piecewise stationary autoregressive model into the recovery process for CS-coded images, and are able to increase the reconstruction quality by 2 ~ 7dB over existing methods. The new CS recovery framework can readily incorporate prior knowledge to boost reconstruction quality. Xiaolin Wu 0001, Xiangjun Zhang, Jia Wang 0004 |
DCC | 1 |
| 2009 | Model-based non-linear estimation for adaptive image restorationabstractWe propose a new image restoration algorithm that is driven by an adaptive piecewise autoregressive model (PAR). The strength of the new algorithm is its ability to preserve spatial structures better than its predecessors. The high adaptability is achieved by locally fitting 2D image waveform to the PAR model in moving windows. The problem is posed as one of nonlinear least-square estimation of both PAR parameters and original pixels, constrained by the degradation function. Robust solutions of the underlying underdetermined inverse problem are obtained by an innovative use of multiple PAR models that circumvent the issue of model overfitting, and by applying a structured total least-square technique. Xiaolin Wu 0001, Xiangjun Zhang |
ICASSP | 1 |
| 2009 | Context-based bias removal of statistical models of wavelet coefficients for image denoisingabstractExisting wavelet-based image denoising techniques all assume a probability model of wavelet coefficients that has zero mean, such as zero-mean Laplacian, Gaussian, or generalized Gaussian distributions. While such a zero-mean probability model fits a wavelet subband well, in areas of edges and textures the distribution of wavelet coefficients exhibits a significant bias. We propose a context modeling technique to estimate the expectation of each wavelet coefficient conditioned on the local signal structure. The estimated expectation is then used to shift the probability model of wavelet coefficient back to zero. This bias removal technique can significantly improve the performance of existing wavelet-based image denoisers. Weisheng Dong, Xiaolin Wu 0001, Guangming Shi, Lei Zhang 0006 |
ICIP | 2 |
| 2009 | Nonlocal back-projection for adaptive image enlargementabstractThis paper presents a novel non-local iterative back-projection (NLIBP) algorithm for image enlargement. The iterative back-projection (IBP) technique iteratively reconstructs a high resolution (HR) image from its blurred and downsampled low resolution (LR) counterpart. However, the conventional IBP methods often produce many ¿jaggy¿ and ¿ringing¿ artifacts because the reconstruction errors are back projected into the reconstructed image isotropically and locally. In natural images, usually there exist many non-local redundancies which can be exploited to improve the image reconstruction quality. Therefore, we propose to incorporate adaptively the non-local information into the IBP process so that the reconstruction errors can be reduced. Experimental results demonstrated that the proposed NLBP can reconstruct faithfully the HR images with sharp edges and texture structures. It outperforms the state-of-the-art methods in both PSNR and visual perception. Weisheng Dong, Lei Zhang 0006, Guangming Shi, Xiaolin Wu 0001 |
ICIP | 4 |
| 2009 | Compressive-uniform hybrid sensing for image acquisition and communicationabstractWe propose a new image acquisition and recovery strategy of hybrid sensing (HS) that combines random sampling of compressive sensing (CS) and uniform down sampling. HS lets the two sampling schemes complement each other so that one can have the best of both worlds: signal-independent sparse sampling that is the hallmark of CS, and locally adaptive signal reconstruction that is afforded by uniform sampling. We suggest a few important applications of HS in image acquisition and communication, such as multispectral imaging, multiple description image coding, multiview video, and ultra-high throughput imaging (e.g., functional medical imaging). Xiaolin Wu 0001, Xiangjun Zhang |
ICIP | 1 |
| 2009 | MDL context modeling of images with application to denoisingabstractThe lately popularized patch-based nonlocal (NL) image processing approach is cast into a framework of statistical context modeling, a thoroughly studied topic in data compression and information theory. The adaptation of image patch (context) to local waveform is crucial to the performance of NL-type of image processing but yet lacks a rigorous study. In this paper we propose a minimum description length (MDL) approach for choosing the size and spatial configuration of the context in which a degraded pixel is to be restored. The MDL criterion of context formation aims to strike an optimal balance between the variance and bias of the errors in fitting a 2D piecewise autoregressive (PAR) model to input image signal. To exemplify the use of the proposed context modeling technique in image processing, an MDL-guided context-based image denoiser is derived and its performance evaluated. Empirical results show that the new context-based denoiser is highly competitive against the current state of the art. Guangtao Zhai, Xiaolin Wu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 2 |
| 2009 | Layered Multicast with Inter-Layer Network CodingabstractMultirate multicast is a powerful methodology of multimedia communication in heterogenous networks. A variant of multirate multicast motivated by scalable multimedia streaming is layered multicast, where the transmitted signal is presented in successive data layers. With recent advances of network coding theory, many layered multicast schemes using network coding have been proposed to improve the performance of traditional routing based layered multicast. They divide the network into different layers and construct a unirate multicast network code for each layer. However, these schemes do not perform network coding between data layers, and consequently cannot realize the full potential of network coding. In this paper, we propose a novel approach to layered multicast that allows network coding of data in different layers. This relaxation lends the proposed scheme greater flexibility in optimizing the data flow than previous layered solutions, and thus achieves higher throughput. Sorina Dumitrescu, Mingkai Shao, Xiaolin Wu 0001 |
INFOCOM | 3 |
| 2009 | On the minimum sum rate of Gaussian multiterminal source coding: New proofsabstractWe show that the minimum sum rate of the Gaussian multiterminal source coding problems can be derived in a unified manner by exploiting the semidefinite partial order of the distortion covariance matrices associated with the MMSE estimation and the so-called reduced optimal linear estimation. In contrast to the existing proofs, the new method does not rely on Shannon's entropy power inequality. Furthermore, this new method leads to a direct proof of the minimum sum rate of the Gaussian two-terminal source coding problem without coupling it to a Gaussian CEO problem. Jia Wang 0004, Jun Chen 0005, Xiaolin Wu 0001 |
ISIT | 3 |
| 2009 | Rate-distortion optimized network communication using general MDCabstractEfficient delivery scheme for multimedia content is critical for the quality of various multimedia network services. It has been shown that the joint design of the source coding techniques and network communication strategies for multimedia streaming can outperform the traditional separate design in a rate-distortion measure. In this paper, we show that current joint design framework in which a data flow optimization followed by a balanced multiple description coding (MDC) can be further improved by allowing more general MDC techniques. A new general MDC construction algorithm is proposed, and the simulation results show a significant performance improvement over the previous balanced MDC based algorithm. Mingkai Shao, Sorina Dumitrescu, Xiaolin Wu 0001 |
ITW | 3 |
| 2009 | GPU-aided directional image/video interpolation for real time resolution upconversionabstractImage/video spatial resolution upconversion aims to obtain a high resolution output from the original low resolution input. Fast resolution upconversion is desired in many applications. In this paper, we develop a GPU-friendly two-pass directional image/video resolution upconversion algorithm and present a GPU implementation of the method, using the NVIDIA CUDA (Compute Unified Device Architecture) technology. Design considerations to speed up the execution are discussed, by taking full advantage of the properties of the CUDA framework and the upconversion scheme. Experimental results show that using a mid-range GPU card, the GPU-optimized resolution upconversion implementation can be more than five times as fast as the original method. Ming-Chao Che, Xiaolin Wu 0001, Jie Liang 0001 |
MMSP | 3 |
| 2009 | Edge-based dynamic ROI coding with standard complianceabstractAn edge-based region-of-interest (ROI) image coding technique is proposed for low bit-rate visual communication. An image is compactly coded into two semantic levels: background sketch and object textures. The background sketch is a down-sampled version of the input image. The decoder reconstructs the background first by adaptive interpolation and lets users identify the ROI, and then requests the ROI textures to be transmitted. A distinct advantage of the proposed ROI technique is a compact edge-based descriptor of natural object boundaries. The expensive ROI geometry computations are carried out at the encoder only on demand, keeping decoder complexity low to benefit wireless devices. The new system outperforms the dynamic ROI coding of JPEG 2000 in both visuality and PSNR. Furthermore, both background and texture coding can be made compliant with any existing compression standard. Xiaolin Wu 0001, Guangming Shi |
MMSP | 2 |
| 2009 | Learning-based recovery of compressive sensing with application in multiple description codingabstractThe recently proposed compressive sensing (CS) theory provides a new solution for multiple description coding (MDC) with fine granularity, by treating each random CS measurement as a description. The performance of CS-based MDC (CS-MDC) depends on the efficacy of the CS recovery algorithm. Existing CS recovery algorithms recover the signal in a fixed space (e.g., Wavelet, DCT, and gradient spaces) for the entire duration of the signal, even though a typical multimedia signal exhibits sparsity in time/space variant spaces. To rectify this problem and develop a better CS recovery algorithm for CSMDC, we propose a learning-based framework to conduct the CS recovery in locally adaptive spaces, and carry out a case study on image MDC. A set of prior image models are learned offline from a training set to facilitate the CS recovery in local adaptive bases. Experiments show that the learning-based CS recovery algorithm can significantly improve the performance of the previous CS-MDC technique in both PSNR and visual quality. Guangming Shi, Weisheng Dong, Xiaolin Wu 0001 |
MMSP | 4 |
| 2009 | On explicit formulas for bandwidth and antibandwidth of hypercubes
Xiaolin Wu 0001, Sorina Dumitrescu |
Discret. Appl. Math. | 2 |
| 2009 | Context-based entropy coding in AVS video coding standard
Li Zhang 0006, Qiang Wang 0011, Ning Zhang 0023, Debin Zhao, Xiaolin Wu 0001, Wen Gao 0001 |
Signal Process. Image Commun. | 5 |
| 2009 | Low Bit-Rate Image Compression via Adaptive Down-Sampling and Constrained Least Squares UpconversionabstractRecently, many researchers started to challenge a long-standing practice of digital photography: oversampling followed by compression and pursuing more intelligent sparse sampling techniques. In this paper, we propose a practical approach of uniform down sampling in image space and yet making the sampling adaptive by spatially varying, directional low-pass prefiltering. The resulting down-sampled prefiltered image remains a conventional square sample grid, and, thus, it can be compressed and transmitted without any change to current image coding standards and systems. The decoder first decompresses the low-resolution image and then upconverts it to the original resolution in a constrained least squares restoration process, using a 2-D piecewise autoregressive model and the knowledge of directional low-pass prefiltering. The proposed compression approach of collaborative adaptive down-sampling and upconversion (CADU) outperforms JPEG 2000 in PSNR measure at low to medium bit rates and achieves superior visual quality, as well. The superior low bit-rate performance of the CADU approach seems to suggest that oversampling not only wastes hardware resources and energy, and it could be counterproductive to image quality given a tight bit budget. Xiaolin Wu 0001, Xiangjun Zhang |
IEEE Trans. Image Process. | 1 |
| 2009 | PCA-Based Spatially Adaptive Denoising of CFA Images for Single-Sensor Digital CamerasabstractSingle-sensor digital color cameras use a process called color demosiacking to produce full color images from the data captured by a color filter array (CAF). The quality of demosiacked images is degraded due to the sensor noise introduced during the image acquisition process. The conventional solution to combating CFA sensor noise is demosiacking first, followed by a separate denoising processing. This strategy will generate many noise-caused color artifacts in the demosiacking process, which are hard to remove in the denoising process. Few denoising schemes that work directly on the CFA images have been presented because of the difficulties arisen from the red, green and blue interlaced mosaic pattern, yet a well-designed "denoising first and demosiacking later" scheme can have advantages such as less noise-caused color artifacts and cost-effective implementation. This paper presents a principle component analysis (PCA)-based spatially-adaptive denoising algorithm, which works directly on the CFA data using a supporting window to analyze the local image statistics. By exploiting the spatial and spectral correlations existing in the CFA image, the proposed method can effectively suppress noise while preserving color edges and details. Experiments using both simulated and real CFA images indicate that the proposed scheme outperforms many existing approaches, including those sophisticated demosiacking and denoising schemes, in terms of both objective measurement and visual evaluation. Lei Zhang 0006, Rastislav Lukac, Xiaolin Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2009 | Robust Color Demosaicking With Adaptation to Varying Spectral CorrelationsabstractAlmost all existing color demosaicking algorithms for digital cameras are designed on the assumption of high correlation between red, green, blue (or some other primary color) bands. They exploit spectral correlations between the primary color bands to interpolate the missing color samples, but in areas of no or weak spectral correlations, these algorithms are prone to large interpolation errors. Such demosaicking errors are visually objectionable because they tend to correlate with object boundaries and edges. This paper proposes a remedy to the above problem that has long been overlooked in the literature. The main contribution of this work is a hybrid demosaicking approach that supplements an existing color demosaicking algorithm by combining its results with those of adaptive intraband interpolation. This is formulated as an optimal data fusion problem, and two solutions are proposed: one is based on linear minimum mean-square estimation and the other based on support vector regression. Experimental results demonstrate that the new hybrid approach is more robust and eliminates the worst type of color artifacts of existing color demosaicking methods. Fan Zhang 0080, Xiaolin Wu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 2 |
| 2009 | On properties of locally optimal multiple description scalar quantizers with convex cellsabstractIt is known that the generalized Lloyd method is applicable to locally optimal multiple description scalar quantizer (MDSQ) design. However, it remains unsettled when the resulting MDSQ is also globally optimal. We partially answer the above question by proving that for a fixed index assignment there is a unique locally optimal fixed-rate MDSQ of convex cells under Trushkin's sufficient conditions for the uniqueness of locally optimal fixed-rate single description scalar quantizer. This result holds for fixed-rate multiresolution scalar quantizer (MRSQ) of convex cells as well. Thus, the well-known log-concave probability density function (pdf) condition can be extended to the multiple description and multiresolution cases. Sorina Dumitrescu, Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2008 | Can Lower Resolution Be Better?abstractRecently, many researchers started to question a long-standing paradox in the engineering practice of digital photography: oversampling followed by compression, and pursue more intelligent sparse sampling techniques. In this research we take a practical approach of uniform down sampling in image space, and the sampling is made adaptive by a spatially varying directional low-pass prefiltering. Since the down-sampled prefiltered image is a low-resolution image of conventional square sample grid, it can be compressed and transmitted without any change to current image coding standards and systems. The decoder first decompresses the low-resolution image and then upsamples it to the original resolution by least-square estimation using a 2D piecewise autoregressive model and the knowledge of directional low-pass filter. The proposed joint adaptive down-sampling and up-sampling technique outperforms JPEG 2000 (the state-of-the-art in lossy image coding) in PSNR measure at low to modest bit rates and achieves superior visual quality at all bit rates. This work shows that oversampling not only increases cost and energy consumption, but it could, even when coupled with a sophisticated rate-distortion optimized compression scheme, cause inferior image quality at certain bit rates. Xiangjun Zhang, Xiaolin Wu 0001 |
DCC | 2 |
| 2008 | Video denoising using 3-D Hybrid Wavelets and Directional filter banksabstractWe propose a new family of nonredundant 3D directional transforms that are useful for video signals. In our construction, taking into account the correlation amongst frames of video, we first decorrelate the temporal data using a stage of 1D wavelet transform and then we employ the efficient 2D Hybrid Wavelets and Directional filter banks (HWD) transform family to the resulting spatial data where we achieve 3D HWD transform family. We construct translation-invariant version of the proposed family and show its efficiency in video denoising compared to other wavelet-based methods through several experiments. Ramin Eslami, Xiaolin Wu 0001 |
ICIP | 2 |
| 2008 | Automatic tonal harmonization for multi-spectral mosaicsabstractWhen producing a mosaic of multiple multi-spectral images one needs to harmonize the colours so that the tone transition is smooth from one image to the other. Given two images Imaand Imb, a transform T is sought to map Imbto an image that is harmonious in multi-spectral appearance to Ima. We give the above problem of tonal harmonization an analytical framework, in which both ideal and practical solutions of the problem are studied. Using a physically motivated image formation model, we prove that a perfect tonal harmonizing operator cannot in general be found, but that whenever such an operator exists it is linear. In the latter case, finding the optimal harmonizing transformation can be cast as a linear programme (LP), which is a type of problem that can be efficiently solved using known techniques. Finally, strong empirical evidence is provided for the efficacy of the proposed solution. Pouya Dehghani Tafti, Xiaolin Wu 0001 |
ICIP | 2 |
| 2008 | Trellis quantization for L∞-constrained compression with integer waveletsabstractAlthough integer wavelets have been successfully used in lossless signal compression, they have not been generalized to near-lossless (L∞ constrained) coding. This paper proposes a new technique of trellis quantization in the framework of lifting integer wavelet as a promising way of near-lossless signal compression. This new technique achieves continuous scalability of the wavelet code stream from highly lossy (low bit rates) to near-lossless (respecting a tight error bound on each decoded sample) reconstruction. Xiaolin Wu 0001, Guangming Shi |
ISIT | 1 |
| 2008 | Toward the optimal multirate multicast for lossy packet networkabstractMultirate multicast is a powerful methodology of multimedia communication in heterogenous networks. With recent advances of network coding theory, many multirate multicast techniques using network coding have been proposed and they hold promises of improving information multicast efficiency over the traditional application-layer multicast protocols. However, all these network coding based multirate multicast solutions adopted a layered coding technique and involves computationally expensive optimization to decide the amount of information transmitted in each multicast layer. In this paper, we propose a novel multirate multicast framework for lossy networks, which uses both the uneven erasure protection (UEP) technique and linear network codes. The new multicast framework not only simplifies the construction of the multicast codes, but also guarantees the strict fairness among clients of different bandwidths. Mingkai Shao, Sorina Dumitrescu, Xiaolin Wu 0001 |
ACM Multimedia | 3 |
| 2008 | A compressive sensing approach of multiple descriptions for network multimedia communicationabstractA new multiple description coding (MDC) approach is proposed based on the theory of compressive sensing (CS). The CS theory allows a signal to be reconstructed from a small number of its random measurements if the signal is sparse in some space. An attractive property of CS for MDC applications is that the reconstruction error only depends on the number but not on which of the transmitted measurements that are received. By treating each CS measurement as a description, we have a balanced MDC scheme with fine description granularity and low encoding complexity. Another advantage of the new MDC approach is that all signals can be coded the same but decoded in different spaces for better sparse reconstruction. Liangjun Wang, Xiaolin Wu 0001, Guangming Shi |
MMSP | 2 |
| 2008 | Standard-compliant multiple description image coding by spatial multiplexing and constrained least-squares restorationabstractWe propose a practical standard-compliant multiple description (MD) image coding technique. Multiple descriptions of an image are generated in the spatial domain by an adaptive prefiltering and uniform down sampling process. The resulting side descriptions are conventional square sample grids that are interleaved with one the other. As such each side description can be coded by any of the existing image compression standards. A side decoder reconstructs the input image by first decompressing the down-sampled image and then solving a least-squares inverse problem, guided by a two-dimensional windowed piecewise autoregressive model. The central decoder is algorithmically similar to the side decoder, but it improves the reconstruction quality by using received side descriptions as additional constraints when solving the underlying inverse problem. Compared with its predecessors the proposed image MD technique offers the lowest encoder complexity, complete standard compliance, competitive rate-distortion performance, and superior subjective quality. Xiangjun Zhang, Xiaolin Wu 0001 |
MMSP | 2 |
| 2008 | On the complexity of joint source-channel decoding of Markov sequences over memoryless channelsabstractWe investigate the complexity of joint source- channel maximum a posteriori (MAP) decoding of a Markov sequence which is first encoded by a source code, then encoded by a convolutional code, and sent through a noisy memoryless channel. As established previously the MAP decoding can be performed by a Viterbi-like algorithm on a trellis whose states are triples of the states of the Markov source, source coder and convolutional coder. The large size of the product space (in the order of K2N, where K is the number of source symbols and N is the number of states of the convolutional coder) appears to prohibit such a scheme. We show that for finite impulse response convolutional codes, the state space size of joint source-channel decoding can be reduced to O(K2+N log N), hence the decoding time becomes O(TK2+TN log N), where T is the length in bits of the decoded bitstream. We further prove that an additional complexity reduction can be achieved when K > N, if the logarithm of the source transition probabilities satisfy the so- called Monge property. This decrease becomes more significant as the tree structure of the source code is more unbalanced. The reduction factor ranges between O(K/N) (for a fixed-length source code) and O(K / log N) (for Golomb-Rice code). Sorina Dumitrescu, Xiaolin Wu 0001 |
IEEE Trans. Commun. | 2 |
| 2008 | Efficient Multiple-Description Image Coding Using Directional Lifting-Based TransformabstractThis paper proposes an efficient two-description image coding technique. The two side descriptions of an image are generated by quincunx subsampling. The decoding from any side description is done by an interpolation process that exploits sample correlation. Although the quincunx subsampling is a natural choice for the best use of sample correlations in image multiple-description coding (MDC), each side description is not amenable to existing image coding techniques because the pixels are not aligned rectilinearly. We show how this difficulty can be overcome by an adaptive directional lifting (ADL) transform that is particularly suitable for decorrelating samples on the quincunx lattice. The ADL transform can be embedded into JPEG 2000 to construct a practical MD image encoder. Experimental results demonstrate that the proposed image MDC scheme can achieve good coding performance. Nan Zhang 0006, Yan Lu 0001, Feng Wu 0001, Xiaolin Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Image Interpolation by Adaptive 2-D Autoregressive Modeling and Soft-Decision EstimationabstractThe challenge of image interpolation is to preserve spatial details. We propose a soft-decision interpolation technique that estimates missing pixels in groups rather than one at a time. The new technique learns and adapts to varying scene structures using a 2-D piecewise autoregressive model. The model parameters are estimated in a moving window in the input low-resolution image. The pixel structure dictated by the learnt model is enforced by the soft-decision estimation process onto a block of pixels, including both observed and estimated. The result is equivalent to that of a high-order adaptive nonseparable 2-D interpolation filter. This new image interpolation approach preserves spatial coherence of interpolated images better than the existing methods, and it produces the best results so far over a wide range of scenes in both PSNR measure and subjective visual quality. Edges and textures are well preserved, and common interpolation artifacts (blurring, ringing, jaggies, zippering, etc.) are greatly reduced. Xiangjun Zhang, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2008 | Rainbow Network Flow of Multiple Description CodesabstractThis paper is an enquiry into the interaction between multiple description coding (MDC) and network routing. We are mainly concerned with rate-distortion optimized network flow of a multiple description (MD) source from multiple servers to multiple sinks. We aim at maximizing a collective metric of the quality of source reconstruction at all sinks, by optimally routing the MD source streams from the server nodes to the sinks. This problem turns out to be very different from conventional maximum network flow. The objective function involves not only the flow volume but also the diversity of the flow contents (i.e., distinction of descriptions), hence, the term rainbow network flow (RNF). For a general network topology, a general fidelity function, and an arbitrary distribution of MDC descriptions on the servers, we prove the RNF problem to be Max-SNP-hard. However, the problem becomes tractable in many practical scenarios, such as when MDC is balanced with descriptions of the same length and importance, when all source nodes have the complete set of MDC descriptions, and when the network topology is a tree or has only one sink. Polynomial-time RNF algorithms are developed for these cases. Xiaolin Wu 0001, Bin Ma 0002, Nima Sarshar |
IEEE Trans. Inf. Theory | 1 |
| 2007 | On Multi-Stage Sequential Coding of Correlated SourcesabstractWe study the problem of multi-stage sequential coding (MSSC), which is an extension of sequential coding of correlated sources. Consider two correlated random variables X and Y to be coded in two stages. The first stage is sequential coding as referred to in the existing literature. At the second stage, the Y encoder refines the information of Y without any knowledge of X, and X encoder refines the information of X with the knowledge of Y, while all previous outputs are known at the decoder. As the sequential coding problem provides a theoretical abstraction of video coding, the MSSC model is a theoretical abstraction of scalable video coding, which is an important application of network communications. We give an achievable region for the MSSC system. The given achievable region is tight when Y is required to be reconstructed perfectly in the usual Shannon sense at the second stage. We also study the minimum total rate MSSC problem, and derive the minimum total rate for Gaussian sources. This result disproves the possibility that the minimum total rate of one stage sequential coding can be achieved at both stages even for correlated Gaussian sources. Thus we offer a theoretical explanation for the performance loss of scalable video coding widely noted by practitioners Jia Wang 0004, Xiaolin Wu 0001, Jun Sun 0005, Songyu Yu |
DCC | 2 |
| 2007 | Image Coding on Quincunx Lattice with Adaptive Lifting and InterpolationabstractConsidering that quincunx lattice is a more efficient spatial sampling scheme than square lattice, we investigate a new approach of image coding for quincunx sample arrangement. The key findings are: 1) adaptive directional lifting is particularly suited to decorrelate samples on quincunx lattice, and 2) quincunx samples can be processed by a 2D piecewise autoregressive model to reproduce the image of conventional square pixel grid, while preserving high frequency spatial features well. By incorporating these two techniques into the encoder and decoder respectively, we are able to improve the performance of JPEG 2000 image codec at low to modest bit rates. Since an image can be easily split into quincunx segments, this work has significance for multiple description image/video coding as well Xiangjun Zhang, Xiaolin Wu 0001, Feng Wu 0001 |
DCC | 2 |
| 2007 | Fast Generalized Motion Estimation and SuperresolutionabstractWe propose a new superresolution algorithm based on a fast motion estimation technique. Two stages of this algorithm, namely, motion estimation and high-resolution reconstruction, rely on an area-based interpolation scheme that involves intersecting two pixel grids in arbitrary orientation, displacement, and scaling. We develop a fast approximate solution of the above problem, whose exact solution is prohibitively expensive. Also, gradient descent algorithm is used for fast convergence of the motion estimation algorithm. Experimental results demonstrate the good performance of the proposed superresolution algorithm as well its robustness against noise. Abhijit Sinha, Xiaolin Wu 0001 |
ICIP (5) | 2 |
| 2007 | Structure Preserving Image Interpolation via Adaptive 2D Autoregressive ModelingabstractThe performance of image interpolation depends on an image model that can adapt to nonstationary statistics of natural images when estimating the missing pixels. However, the construction of such an adaptive model needs the knowledge of every pixels that are absent. We resolve this dilemma by a new piecewise 2D autoregressive technique that builds the model and estimates the missing pixels jointly. This task is formulated as a non-linear optimization problem. Although computationally demanding, the new non-linear approach produces superior results than current methods in both PSNR and subjective visual quality. Moreover, in quest for a practical solution, we break the non-linear optimization problem into two subproblems of linear least-squares estimation. This linear approach proves very effective in our experiments. Xiangjun Zhang, Xiaolin Wu 0001 |
ICIP (4) | 2 |
| 2007 | Edge-Guided Perceptual Image Coding via Adaptive InterpolationabstractWe propose a new image compression technique based on adaptive edge-guided interpolation. The design criterion of the new image compression technique is a perceptual one. We aim to preserve the spatial correlation structures of image edges because human visual system is highly sensitive to distortions of spatial coherence of edges. Experimental results show that the proposed technique indeed achieves superior visual quality than JPEG 2000 at low to modest bit rates. For situations when the encoder has limited computing resources and/or power supply (e.g., wireless communications and sensor networks), the new technique has the desirable property of low encoding complexity. Xiangjun Zhang, Xiaolin Wu 0001 |
ICME | 2 |
| 2007 | Rate-Distortion Optimized Network CommunicationabstractNetwork information multicast has been considered extensively, either as a routing problem or more recently in the context of network coding. Most of the Internet bandwidth, however, is consumed by multimedia contents that are amenable to lossy reconstruction. In this paper, we investigate the following fundamental question: How does one communicate a media content from nodes (servers) that observe/supply the content to a set of sink nodes (clients) to realize the best possible reconstruction of the content in a rate-distortion sense? While this problem remains essentially open, this paper takes the first step by exploring the intricate entanglement of source coding and network communication, within an optimization framework. In particular, we investigate the joint optimization of network communication strategies (e.g., routing or network coding) and common source coding schemes (e.g., progressive coding or more general multiple description coding). We formulate several such problems for which we are able to develop efficient polynomial time solutions. In particular, we consider layered multicast of progressively encoded source code streams using network coding and optimal routing of balanced multiple description codes. Finally, the improvement in the overall quality of source reconstruction by using the proposed schemes is verified through simulations. Nima Sarshar, Xiaolin Wu 0001 |
INFOCOM | 2 |
| 2007 | Context-based Arithmetic Coding Reexamined for DCT Video CompressionabstractThis paper presents a new context modeling technique for arithmetic coding of DCT coefficients in video compression. A key feature of the new technique is the inclusion of all previously coded coefficient magnitudes in a DCT block in context modeling. This enables adaptive arithmetic coding to exploit the redundancy of the high-order Markov process in the DCT domain with a few conditioning states. In addition, a context weighting technique is used to further improve the coding efficiency. The complexity of the new arithmetic coding scheme is slightly lower than that of context-based adaptive binary arithmetic coding (CABAC) of H.264. Moreover, the scheme is made compatible to the AVS baseline profile. It achieves on average 13% improvement in compression ratio over context-based two dimension variable length coding (C2DVLC) designed for the DCT domain, and a similar coding efficiency as the CABAC technique in H.264. Li Zhang 0006, Xiaolin Wu 0001, Ning Zhang 0023, Wen Gao 0001, Qiang Wang 0011, Debin Zhao |
ISCAS | 2 |
| 2007 | Context Quantization by Minimum Adaptive Code LengthabstractContext quantization is a technique to deal with the issue of context dilution in high-order conditional entropy coding. We investigate the problem of context quantizer design under the criterion of minimum adaptive code length. A property of such context quantizers is derived for binary symbols. A fast context quantizer design algorithm for conditioning binary symbols is presented and its complexity analyzed. It is conjectured that this algorithm is optimal. The context quantization is performed in what may be perceived as a probability simplex space rather than in the space of context instances. Søren Forchhammer, Xiaolin Wu 0001 |
ISIT | 2 |
| 2007 | Multiple Descriptions with Side Informations Also Known At the EncoderabstractWe propose a new scheme of multiple descriptions with side information (SI). The two side decoders of the system use two different SI streams. Both SI streams are available to the central decoder and to the encoder. We give an inner bound for this system for general source and SI. The tight bound is obtained for the quadratic Gaussian case. This result is compared with our previous result of the MDWZ (multiple descriptions in the Wyner-Ziv setting) problem in which none of the side information is known at the encoder. It is shown that when side information is absent at the encoder, there is a performance loss. The proposed scheme and its achievable region have practical significance. It offers theoretical insight into the multiple description video coding (MDVC) and suggests an optimal coding strategy. Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005 |
ISIT | 2 |
| 2007 | On Design of Linear Minimum-Entropy PredictorabstractLinear predictors for lossless data compression should ideally minimize the entropy of prediction errors. But in current practice predictors of least-square type are used instead. In this paper, we formulate and solve the linear minimum-entropy predictor design problem as one of convex or quasiconvex programming. The proposed minimum-entropy design algorithms are derived from the well-known fact that prediction errors of most signals obey generalized Gaussian distribution. Empirical results and analysis are presented to demonstrate the superior performance of the linear minimum-entropy predictor over the traditional least-square counterpart for lossless coding. Xiaolin Wu 0001 |
MMSP | 2 |
| 2007 | Alphabet Partitioning Techniques for Semiadaptive Huffman Coding of Large AlphabetsabstractPractical applications that employ entropy coding for large alphabets often partition the alphabet set into two or more layers, and encode each symbol by using some suitable prefix coding for each layer. In this paper, we formulate the problem of finding an alphabet partitioning for the design of a two-layer semiadaptive code as an optimization problem, and give a solution based on dynamic programming. However, the complexity of the dynamic programming approach can be quite prohibitive for a long sequence and a very large alphabet size. Hence, we also give a simple greedy heuristic algorithm whose running time is linear in the length of the input sequence, irrespective of the underlying alphabet size. Although our dynamic programming and greedy algorithms do not provide a globally optimal solution for the alphabet partitioning problem, experimental results demonstrate that superior prefix coding schemes for large alphabets can be designed using our new approach Yi-Jen Chiang, Nasir Memon, Xiaolin Wu 0001 |
IEEE Trans. Commun. | 4 |
| 2007 | Adaptive Directional Lifting-Based Wavelet Transform for Image CodingabstractWe present a novel 2-D wavelet transform scheme of adaptive directional lifting (ADL) in image coding. Instead of alternately applying horizontal and vertical lifting, as in present practice, ADL performs lifting-based prediction in local windows in the direction of high pixel correlation. Hence, it adapts far better to the image orientation features in local windows. The ADL transform is achieved by existing 1-D wavelets and is seamlessly integrated into the global wavelet transform. The predicting and updating signals of ADL can be derived even at the fractional pixel precision level to achieve high directional resolution, while still maintaining perfect reconstruction. To enhance the ADL performance, a rate-distortion optimized directional segmentation scheme is also proposed to form and code a hierarchical image partition adapting to local features. Experimental results show that the proposed ADL-based image coding technique outperforms JPEG 2000 in both PSNR and visual quality, with the improvement up to 2.0 dB on images with rich orientation features. Wenpeng Ding, Feng Wu 0001, Xiaolin Wu 0001, Shipeng Li 0001, Houqiang Li |
IEEE Trans. Image Process. | 3 |
| 2007 | On Rate-Distortion Models for Natural Images and Wavelet Coding PerformanceabstractOperational rate-distortion (RD) functions of most natural images, when compressed with state-of-the-art wavelet coders, exhibit a power-law behavior D alpha R(-gamma) at moderately high rates, with gamma being a constant depending on the input image, deviating from the well-known exponential form of the RD function D alpha 2(-xiR) for bandlimited stationary processes. This paper explains this intriguing observation by investigating theoretical and operational RD behavior of natural images. We take as our source model the fractional Brownian motion (fBm), which is often used to model nonstationary behaviors in natural images. We first establish that the theoretical RD function of the fBm process (both in 1-D and 2-D) indeed follows a power law. Then we derive operational RD function of the fBm process when wavelet encoded based on water-filling principle. Interestingly, both the operational and theoretical RD functions behave as D alpha R(-gamma). For natural images, the values of gamma are found to be distributed around 1. These results lend an information theoretical support to the merit of multiresolution wavelet compression of self-similar processes and, in particular, natural images that can be modelled by such processes. They may also prove useful in predicting performance of RD optimized image coders. Nima Sarshar, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2007 | Color Reproduction From Noisy CFA Data of Single Sensor Digital CamerasabstractSingle sensor digital color still/video cameras capture images using a color filter array (CFA) and require color interpolation (demosaicking) to reconstruct full color images. The color reproduction has to combat sensor noises which are channel dependent. If untreated in demosaicking, sensor noises can cause color artifacts that are hard to remove later by a separate denoising process, because the demosaicking process complicates the noise characteristics by blending noises of different color channels. This paper presents a joint demosaicking-denoising approach to overcome this difficulty. The color image is restored from noisy mosaic data in two steps. First, the difference signals of color channels are estimated by linear minimum mean square-error estimation. This process exploits both spectral and spatial correlations to simultaneously suppress sensor noise and interpolation error. With the estimated difference signals, the full resolution green channel is recovered. The second step involves in a wavelet-based denoising process to remove the CFA channel-dependent noises from the reconstructed green channel. The red and blue channels are subsequently recovered. Simulated and real CFA mosaic data are used to evaluate the performance of the proposed joint demosaicking-denoising scheme and compare it with many recently developed sophisticated demosaicking and denoising schemes. Lei Zhang 0006, Xiaolin Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2007 | Lagrangian Optimization of Two-Description Scalar QuantizersabstractIn this paper, we study the problem of optimal design of balanced two-description fixed-rate scalar quantizer (2DSQ) under the constraint of convex codecells. Using a graph-based approach to model the problem, we show that the minimum expected distortion of the 2DSQ is a convex function of the number of codecells in the side quantizers. This property allows the problem to be solved by Lagrangian minimization for which the optimal Lagrangian multiplier exists. Given a trial multiplier, we exploit a monotonicity of the objective function, and develop a simple and fast dynamic programming technique to solve the parameterized problem. To further improve the algorithm efficiency, we propose an RD-guided search strategy to find the optimal Lagrangian multiplier. In our experiments on distributions of interest for signal compression applications the proposed algorithm improves the speed of the fastest algorithm so far, by a factor ofO(K/logK), whereKis the number of codecells in each side quantizer. We also assess the impact on the optimality of the convex codecell constraint. Using a published performance analysis of 2DSQ at high rates, we show that asymptotically this constraint does not preclude optimality forL2distortion measure, when channels have a higher than 0.12 loss rate. Sorina Dumitrescu, Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2007 | Efficient Algorithms for Optimal Uneven Protection of Single and Multiple Scalable Code Streams Against Packet Erasuresabstractwe study algorithmic approaches for rate-fidelity optimal packetization of a single and multiple scalable source code streams with uneven erasure protection (UEP). A new algorithm is developed to obtain the globally optimal solution for scalable source codes of convex rate-fidelity function and for a wide class of erasure channels, including channels for which the probability of losing packets is monotonically nonincreasing in , and independent erasure channels with packet erasure rate smaller than 0.5. This is achieved at linear space complexity and near-linear time complexity in the transmission budget, representing significant improvement over the known globally optimal algorithm. When applied to SPIHT compressed images, the results of the proposed algorithm are virtually the same as the global optima. The above success is also extended to UEP packetization of multiple scalable code streams. We improve the existing algorithms in both speed and performance. Sorina Dumitrescu, Xiaolin Wu 0001, Zhe Wang 0022 |
IEEE Trans. Multim. | 2 |
| 2006 | Optimal Index Assignment for Multiple Description Lattice Vector QuantizationabstractOptimal index assignment of multiple description lattice vector quantizer (MDLVQ) can be posed as a large-scale linear assignment problem. But is this expensive algorithmic approach necessary? This paper presents a simple index assignment algorithm for high-resolution MDLVQ of K /spl ges/ 2 balanced descriptions in any dimensions. Despite its simplicity, the new algorithm is optimal for a large family of lattices encountered in theory and practice, in terms of minimizing the expected distortion for any side description loss rate and any side entropy rate. This work offers exact combinatoric constructions of optimal index assignments, rather than arguing for the optimality asymptotically. Consequently, the optimality holds for all values of sublattice index N (i.e., over all trade-offs between the central and side distortions), rather than for very large N only. Furthermore, the time complexity of the new algorithm is O(N) as opposed to O(N/sup 6/) for a current linear assignment-based method. New and improved closed form expressions of the expected distortion as the function of N and K are also presented. Thus the optimal values of N and K can be computed. Xiaolin Wu 0001 |
DCC | 2 |
| 2006 | A Practical Approach to Joint Network-Source CodingabstractWe are interested in how to best communicate a real valued source to a number of destinations (sinks) over a network with capacity constraints in a collective fidelity metric over all the sinks, a problem which we call joint network-source coding. It is demonstrated that multiple description codes in conjunction with proper diversity routing provide a powerful solution to joint network-source coding. A systematic optimization approach is proposed. It consists of optimizing the network routing given a multiple description code and designing optimal multiple description code for the corresponding optimized routes. Nima Sarshar, Xiaolin Wu 0001 |
DCC | 2 |
| 2006 | Joint Source-Channel Decoding of Multiple Description Quantized Markov SequencesabstractThis paper proposes a framework for joint source-channel decoding of Markov sequences that are coded by a fixed-rate multiple description quantizer (MDQ), and transmitted via a lossy network. This framework is suited for lossy networks of primitive energy-deprived source encoders. Our technical approach is one of maximum a posteriori probability (MAP) sequence estimation that exploits both the source memory and the correlation between different MDQ descriptions. We solve the MAP estimation problem by computing the longest path in a weighted directed acyclic graph, at a complexity of O(L/sup 2/NK), where N is the number of source symbols in the input sequence, K is the number of MDQ descriptions, and L is the number of codewords of the central quantizer. If the source sequence is Gaussian Markovian, the decoder complexity can be reduced to O(LNK). For MDQ-compressed Markov sequences impaired by both bit errors and erasure errors, the performance of joint source-channel MAP decoder can be 6 dB higher than the conventional hard-decision decoder. Furthermore, the new MDQ decoding technique unifies the treatments of different subsets of the K descriptions available at the decoder, circumventing the thorny issue of requiring up to 2/sup K/ - 1 MDQ side decoders. Xiaolin Wu 0001 |
DCC | 1 |
| 2006 | Efficient Algorithm for Globally Optimal Uneven Erasure-Protected Packetization of Scalable Code StreamsabstractA new algorithm is presented for rate-fidelity optimal packetization of scalable source bit streams with uneven erasure protection. It provides the globally optimal solution for input sources of convex rate-fidelity function and for a wide class of erasure channels, including channels for which the probability of losing n packets is monotonically decreasing in n, and independent erasure channels with packet erasure rate smaller than 0.5. The time and space complexities of the new algorithm are both O(NL), where N is the number of packets and L is the packet payload size, comparing to the O(NL2) time and space complexities of the existing globally optimal solution. When applied to SPIHT compressed images, the results of the proposed algorithm are virtually the same as the globally optima Sorina Dumitrescu, Xiaolin Wu 0001, Zhe Wang 0022 |
ICME | 2 |
| 2006 | Error Resilient Multiple Description Compression of Vector GraphicsabstractThis research is motivated by the needs of robust streaming of vector graphics contents over the Internet, wireless and other lossy networks. We present a multiple description coding (MDC) technique for error resilient compression and transmission of 2D vector graphics contents. An object is coded into two or more so-called co-descriptors, which are transmitted in separate data packets and generally via different network routes from a server to a client. Each co-descriptor can autonomously provide an approximation of the input object, and it can collaborate with other co-descriptors, if also available at the decoder, to refine the approximation Martin Röder, Xiaolin Wu 0001, Sorina Dumitrescu |
ICME | 2 |
| 2006 | Joint Source-Channel Decoding of Multiple Description Quantized and Variable Length Coded Markov SequencesabstractThis paper proposes a framework for joint source-channel decoding of Markov sequences that are encoded by an entropy coded multiple description quantizer (MDQ), and transmitted via a lossy network. This framework is particularly suited for lossy networks of inexpensive energy-deprived mobile source encoders. Our approach is one of maximum aposteriori probability (MAP) sequence estimation that exploits both the source memory and the correlation between different MDQ descriptions. The MAP problem is modeled and solved as one of the longest path in a weighted directed acyclic graph. For MDQ-compressed Markov sequences impaired by both bit errors and erasure errors, the proposed joint source-channel MAP decoder can achieve 5 dB higher SNR than the conventional hard-decision decoder. Furthermore, the new MDQ decoding technique unifies the treatments of different subsets of the K descriptions available at the decoder, circumventing the thorny issue of requiring up to 2K-1 MDQ side decoders Xiaolin Wu 0001 |
ICME | 2 |
| 2006 | Index assignment design for three-description lattice vector quantizationabstractIn this paper, we propose a new index assignment scheme for the 3-description case, which aims to find a 3-tuple sublattice points to represent a fine lattice point. The design is made such that each edge of the triangle formed by the 3-tuple points needs to be as short as possible while the gravity center of the triangle is as close as possible to the fine lattice point. With a delicate sublattice partition, a well designed construction and mapping is developed to minimize the expected distortion. Minglei Liu, Ce Zhu, Xiaolin Wu 0001 |
ISCAS | 3 |
| 2006 | Multiple Descriptions in the Wyner-Ziv SettingabstractWe propose a new scheme of multiple descriptions in the Wyner-Ziv setting (MD-WZ). The two side decoders of MD-WZ use two different side information (SI) streams. Both SI streams are available to the central decoder, but none to the encoder. We derive an achievable region (inner bound) for this MD-WZ system for general source and SI. If the source and SI are correlated Gaussian and for quadratic distortion metric, the tight bound is obtained. Our result is an extension of Ozarow's result on multiple descriptions of Gaussian source without SI. The MD-WZ coding scheme is shown to have a property of practical significance. For symmetric case where the joint distributions of the source and the two SI are the same and the two channels are balanced, interchanging the two channels causes no performance loss for Gaussian source. Considering that the existing multi-description video coding methods suffer from the notorious drifting problem induced by channels interchange, this work lends a theoretical support to distributed multi-description video coding in the Wyner-Ziv setting Jia Wang 0004, Xiaolin Wu 0001, Songyu Yu, Jun Sun 0005 |
ISIT | 2 |
| 2006 | On Optimal Index Assignment for MAP Decoding of Markov SequencesabstractIndex assignment and maximum a posteriori (MAP) decoding are two well-known techniques for error-resilient multimedia communications. If the two techniques are used in tandem, how they interact with each other will greatly affect the system performance. An important problem in this regard is, which has seemingly evaded attention, the design of index assignment to achieve the best possible performance of joint source-channel MAP decoding, given the source and channel statistics and given a distortion metric. In a first attempt on this design challenge, we pose the index assignment of a scalar quantizer for MAP decoding of Markov sequences coded by this quantizer as a quadratic programming problem. For Gaussian Markov sequences we derive a locally optimal index assignment by exploring some properties of the objective function. Experimental results show that the proposed scheme can find optimal or near-optimal solutions. The optimized index assignment can achieve much lower average symbol error rate than conventional schemes Xiaolin Wu 0001, Sorina Dumitrescu |
ISIT | 2 |
| 2006 | Lossless Geometry Compression for Steady-State and Time-Varying Irregular GridsabstractIn this paper we investigate the problem of lossless geometry compression of irregular-grid volume data represented as a tetrahedral mesh. We propose a novel lossless compression technique that effectively predicts, models, and encodes geometry data for both steady-state (i.e., with only a single time step) and time-varying datasets. Our geometry coder is truly lossless and also does not need any connectivity information. Moreover, it can be easily integrated with a class of the best existing connectivity compression techniques for tetrahedral meshes with a small amount of overhead information. We present experimental results which show that our technique achieves superior compression ratios, with reasonable encoding times and fast (linear) decoding times. Yi-Jen Chiang, Nasir Memon, Xiaolin Wu 0001 |
EuroVis | 4 |
| 2006 | On Interpolation and Resampling of Discrete DataabstractThis letter introduces a new representation of discrete signals based on the mathematical notions of functionals and continuous dual spaces. A new and more general sampling theorem is also suggested. Next, the problems of interpolating and resampling discrete signals are addressed; and a general solution using functional interpolation-which is applicable to many different settings-is proposed. Families of resampling filters dubbed de Boor-Ron filters that use de Boor-Ron interpolation are introduced, and their numerical realization is discussed. Some applications of this research are suggested Pouya Dehghani Tafti, Shahram Shirani, Xiaolin Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2006 | Length-Constrained MAP Decoding of Variable-Length Encoded Markov SequencesabstractIn this paper, we consider the problem of length-constrained maximum a posteriori decoding of a Markov sequence that is variable-length encoded and transmitted over a binary symmetric channel. We convert this problem into one of a maximum-weight k-link path in a weighted directed acyclic graph. The induced graph-optimization problem can be solved by a fast parameterized search algorithm that finds either the optimal solution with high probability, or a good approximate solution otherwise. The proposed algorithm has lower complexity and superior performance than the previous approximation algorithms Zhe Wang 0022, Xiaolin Wu 0001 |
IEEE Trans. Commun. | 2 |
| 2006 | Temporal color video demosaicking via motion estimation and data fusionabstractColor demosaicking of charge-coupled device (CCD) data has been thoroughly studied for single-sensor still digital cameras. However, there has seemingly been little research on color demosaicking techniques for single-sensor video digital cameras. The temporal dimension of a color mosaic image sequence can reveal new information on the missing color components due to the mosaic subsampling, which is otherwise unavailable in the spatial domain of individual frames. This paper proposes a temporal approach to color demosaicking. A pixel of the current frame is matched to another in a reference frame via motion analysis, such that the CCD sensor samples different color components of the same object position in the two frames. The resulting inter-frame estimates of missing color components are fused with suitable intra-frame estimates to achieve a more robust color restoration. Our experimental results demonstrate clear advantages of the presented temporal color demosaicking approach over its intra-frame counterparts in reducing the color artifacts. Xiaolin Wu 0001, Lei Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | Improvement of Color Video Demosaicking in Temporal DomainabstractColor demosaicking is critical to the image quality of digital still and video cameras that use a single-sensor array. Limited by the mosaic sampling pattern of the color filter array (CFA), color artifacts may occur in a demosaicked image in areas of high-frequency and/or sharp color transition structures. However, a color digital video camera captures a sequence of mosaic images and the temporal dimension of the color signals provides a rich source of information about the scene via camera and object motions. This paper proposes an inter-frame demosaicking approach to take advantage of all three forms of pixel correlations: spatial, spectral, and temporal. By motion estimation and statistical data fusion between adjacent mosaic frames, the new approach can remove much of the color artifacts that survive intra-frame demosaicking and also improve tone reproduction accuracy. Empirical results show that the proposed inter-frame demosaicking approach consistently outperforms its intra-frame counterparts both in peak signal-to-noise measure and subjective visual quality. Xiaolin Wu 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 1 |
| 2006 | Context quantization by kernel Fisher discriminantabstractOptimal context quantizers for minimum conditional entropy can be constructed by dynamic programming in the probability simplex space. The main difficulty, operationally, is the resulting complex quantizer mapping function in the context space, in which the conditional entropy coding is conducted. To overcome this difficulty, we propose new algorithms for designing context quantizers in the context space based on the multiclass Fisher discriminant and the kernel Fisher discriminant (KFD). In particular, the KFD can describe linearly nonseparable quantizer cells by projecting input context vectors onto a high-dimensional curve, in which these cells become better separable. The new algorithms outperform the previous linear Fisher discriminant method for context quantization. They approach the minimum empirical conditional entropy context quantizer designed in the probability simplex space, but with a practical implementation that employs a simple scalar quantizer mapping function rather than a large lookup table. Mantao Xu, Xiaolin Wu 0001, Pasi Fränti |
IEEE Trans. Image Process. | 2 |
| 2006 | Lossless compression of color mosaic imagesabstractLossless compression of color mosaic images poses a unique and interesting problem of spectral decorrelation of spatially interleaved R, G, B samples. We investigate reversible lossless spectral-spatial transforms that can remove statistical redundancies in both spectral and spatial domains and discover that a particular wavelet decomposition scheme, called Mallat wavelet packet transform, is ideally suited to the task of decorrelating color mosaic data. We also propose a low-complexity adaptive context-based Golomb-Rice coding technique to compress the coefficients of Mallat wavelet packet transform. The lossless compression performance of the proposed method on color mosaic images is apparently the best so far among the existing lossless image codecs. Ning Zhang 0023, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2006 | An edge-guided image interpolation algorithm via directional filtering and data fusionabstractPreserving edge structures is a challenge to image interpolation algorithms that reconstruct a high-resolution image from a low-resolution counterpart. We propose a new edge-guided nonlinear interpolation technique through directional filtering and data fusion. For a pixel to be interpolated, two observation sets are defined in two orthogonal directions, and each set produces an estimate of the pixel value. These directional estimates, modeled as different noisy measurements of the missing pixel are fused by the linear minimum mean square-error estimation (LMMSE) technique into a more robust estimate, using the statistics of the two observation sets. We also present a simplified version of the LMMSE-based interpolation algorithm to reduce computational cost without sacrificing much the interpolation performance. Experiments show that the new interpolation techniques can preserve edge sharpness and reduce ringing artifacts. Lei Zhang 0006, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2005 | Optimized Prediction for Geometry Compression of Triangle MeshesabstractIn this paper we propose a novel geometry compression technique for 3D triangle meshes. We focus on a commonly used technique for predicting vertex positions via a flipping operation using the parallelogram rule. We show that the efficiency of the flipping operation is dependent on the order in which triangles are traversed and vertices are predicted accordingly. We formulate the problem of optimally (traversing triangles and) predicting the vertices via flippings as a combinatorial optimization problem of constructing a constrained minimum spanning tree. We give heuristic solutions for this problem and show that we can achieve prediction efficiency within 17.4% on average as compared to the unconstrained minimum spanning tree which is an unachievable lower bound. We also show significant improvements over previous techniques in the literature that strive to find good traversals that also attempt to minimize prediction errors obtained by a sequence of flipping operations, albeit using a different approach. Yi-Jen Chiang, Nasir Memon, Xiaolin Wu 0001 |
DCC | 4 |
| 2005 | On Global Optimality of Gradient Descent Algorithms for Fixed-Rate Scalar Multiple Description Quantizer DesignabstractWe prove that Trushkin's (1982) sufficient conditions for the global optimality of a locally optimal fixed-rate scalar quantizer also ensure the global optimality of a locally optimal fixed-rate multiple description scalar quantizer of convex codecells, with respect to a fixed index assignment. This result also holds for the fixed-rate multiresolution scalar quantizer of convex codecells. As a consequence the well-known log-concave pdf condition can be extended to the multiple description and multiresolution case. Sorina Dumitrescu, Xiaolin Wu 0001 |
DCC | 2 |
| 2005 | Steganalysis of halftone imagesabstractWe present a novel steganalysis technique for halftone images without knowledge of the original cover image. We first convert halftone images into grayscale-like images by low-pass filtering. The low-pass-filtered image is then decomposed using quadrature mirror filters, and a set of subband coefficients are generated at different scales and orientations. Next, a set of statistical features are computed from the subband coefficients and their predicated errors. Using Fisher linear discriminant analysis, a statistical classifier is designed to detect marked images. Experimental results demonstrate the effectiveness and accuracy of the proposed technique. Ming Jiang 0006, Edward K. Wong, Nasir Memon, Xiaolin Wu 0001 |
ICASSP (2) | 4 |
| 2005 | Multi-dimensional average-interpolating refinement on arbitrary latticesabstractMulti-dimensional datasets containing local averages of a function arise in many applications such as processing of CCD captures and medical images. Motivated by this fact we introduce multi-dimensional average-interpolating refinement on arbitrary lattices in arbitrary dimensions. Our refinement algorithm results in smooth scaling functions of compact support. This method forms a basis for multi-dimensional multi-resolution analysis and subdivision on datasets obtained by locally averaging a smooth function. As an example, we present two-dimensional polynomial average-interpolating subdivision on the quincunx lattice and show that the resulting scaling functions are highly regular in the sense of Sobolev. Pouya Dehghani Tafti, Shahram Shirani, Xiaolin Wu 0001 |
ICASSP (4) | 3 |
| 2005 | On cross correlation based-discrete time delay estimationabstractThe cross correlation function (CCF) is a powerful tool in time delay estimation and parabola functions are widely used as parametric models of it. However, no study has been done on the accuracy of the parabola approximation of CCF. In this paper, we analyze the CCF of multi-sensors and derive the analytic forms of CCF for the stationary processes of the exponential auto-correlation function with respect to two important types of sensor kernels. We demonstrate that the Gaussian function is a better and more robust approximation of CCF than the parabola in these cases. This new approach leads to higher precision in time delay estimation using the CCF peak locating strategy. Lei Zhang 0006, Xiaolin Wu 0001 |
ICASSP (4) | 2 |
| 2005 | Image interpolation using texture orientation map and kernel Fisher discriminantabstractWe propose a new non-linear approach of high order context to image interpolation. A global texture orientation map is generated by directional Gabor filters to estimate the edge directions in subpixel precision. The interpolation direction is further refined by a kernel Fisher discriminant that exploits prior knowledge gained from a training set. Experiments show that the proposed method can preserve edge sharpness and subdue the ringing artifacts better than the existing methods, and it obtains higher PSNR as well. Xiaolin Wu 0001, Xiangjun Zhang |
ICIP (1) | 1 |
| 2005 | Globally Optimal Uneven Erasure-Protected Multi-Group Packetization of Scalable CodesabstractWe study the problem of rate-distortion optimal packetization with uneven erasure protection (UEP) of scalable source sequence, into multiple groups of packets. The grouping of packets is needed when the length of the channel code, hence the number of packets, has to be modest for low decoding complexity. The problem was previously addressed in the literature but only locally optimal solution was proposed. We develop an algorithm for globally optimal solution and show that it has the same complexity as optimal UEP packetization into a single group of packets, i. e., quadratic in transmission budget. Sorina Dumitrescu, Xiaolin Wu 0001 |
ICME | 2 |
| 2005 | Optimal packetization of VLC and convolution coded Markov sequencesabstractWe consider the problem of packetizing a variable length coded Markov sequence into fixed length packets, while being protected by variable rate channel code. Given the total transmission bit budget, a joint source-channel coding problem is how to partition the input sequence and how to determine the coding rates of individual packets for minimum expected distortion when the sequence is sent over binary symmetric channel. Three methods are proposed to estimate the performance of a sequence when transmitted through the system, based on which we convert the joint source-channel coding problem into a shortest path problem in a weighted directed acyclic graph which can be solved by using dynamic programming. Simulation shows that the overall performance of the system can be improved by 10-30% compared with the performance of the fixed rate packetization scheme. Xiaolin Wu 0001 |
ICME | 2 |
| 2005 | Rainbow network problems and multiple description codingabstractIn packet switched networks receivers can get packets of a multiple description code (MDC) from different sources for enhanced QoS and robust transmission. The quality achieved by a decoder increases in the number of distinct rather than the total number of packets received. This property makes the problems of optimizing network flows and transmission strategies for MDC, called rainbow network problems, very different from those of conventional network flow and management. Two interesting problems: rainbow network flow and rainbow multicast, are formulated and treated. The rainbow network flow problem of maximizing the number of distinct packets received, constrained by edge capacities, is shown to be NP-hard in multisource-multisink setting. But it can be reduced to conventional maximum network flow problem in the case of single sink, hence becomes solvable in polynomial time. Rainbow multicast problem is about coordinating multiple servers for minimum expected distortion at one or a set of clients. Although being seemingly intractable in general, some variants of the problem have analytical solutions Xiaolin Wu 0001, Bin Ma 0002, Nima Sarshar |
ISIT | 1 |
| 2005 | On the complexity of joint source-channel decoding of markov sequences over memoryless channelsabstractWe investigate the complexity of joint source-channel maximum a posteriori (MAP) decoding of a Markov sequence which is first encoded by a source code, then encoded by a convolutional code, and sent through a noisy memoryless channel. As established previously the MAP decoding can be performed by a Viterbi-like algorithm on a trellis whose states are triples of the states of the Markov source, the source coder and convolutional coder. The large size of the product space (in the order of K2N, where K is the number of source symbols and N is the number of states of the convolutional coder) appears to prohibit such a scheme. We show that in the case of finite impulse response convolutional codes the state space size can be reduced to O(K2+ NlogN), hence the decoding time becomes O(TK2+ TNlogN), where T is the length in bits of the decoded bitstream. We further show that an additional complexity reduction can be achieved when K > N, if the source satisfies a certain property, which is the case for a scalar quantized Gaussian-Markov source. This decrease becomes more significant as the tree structure of the source code is more unbalanced. The reduction factor ranges between O(K/N) (for a fixed-length source code) and O(K/logN) (for Golomb-Rice code) Sorina Dumitrescu, Xiaolin Wu 0001 |
ISIT | 2 |
| 2005 | Steganalysis of Degraded Document ImagesabstractIn this paper, a steganalysis technique using compression bit rate as a distinguishing statistic is presented to detect secret messages embedded in document images that are degraded in quality by printing, photocopying, and/or scanning processes. We consider embedding techniques that flip pixels in binary document images that contain characters and symbols. Noise introduced by printing, photocopying, and/or scanning can be modeled by a local optical distortion process. Steganographic embedding is modeled as an additive noise process and we use compression bit rate as a distinguishing statistic to discriminate between stego images and unmarked images. Experimental results showed that the proposed technique can detect stego images with reasonably good accuracy, given the inherent difficulty of the problem Ming Jiang 0006, Edward K. Wong, Nasir Memon, Xiaolin Wu 0001 |
MMSP | 4 |
| 2005 | Optimal Two-Description Scalar Quantizer Design
Sorina Dumitrescu, Xiaolin Wu 0001 |
Algorithmica | 2 |
| 2005 | Canny Edge Detection Enhancement by Scale MultiplicationabstractThe technique of scale multiplication is analyzed in the framework of Canny edge detection. A scale multiplication function is defined as the product of the responses of the detection filter at two scales. Edge maps are constructed as the local maxima by thresholding the scale multiplication results. The detection and localization criteria of the scale multiplication are derived. At a small loss in the detection criterion, the localization criterion can be much improved by scale multiplication. The product of the two criteria for scale multiplication is greater than that for a single scale, which leads to better edge detection performance. Experimental results are presented. Paul Bao, Lei Zhang 0006, Xiaolin Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Joint Classification and Pairing of Human ChromosomesabstractWe reexamine the problems of computer-aided classification and pairing of human chromosomes, and propose to jointly optimize the solutions of these two related problems. The combined problem is formulated into one of optimal three-dimensional assignment with an objective function of maximum likelihood. This formulation poses two technical challenges: 1) estimation of the posterior probability that two chromosomes form a pair and the pair belongs to a class and 2) good heuristic algorithms to solve the three-dimensional assignment problem which is NP-hard. We present various techniques to solve these problems. We also generalize our algorithms to cases where the cell data are incomplete as often encountered in practice. Pravesh Biyani, Xiaolin Wu 0001, Abhijit Sinha |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2005 | Multiscale LMMSE-Based Image Denoising With Optimal Wavelet SelectionabstractIn this paper, a wavelet-based multiscale linear minimum mean square-error estimation (LMMSE) scheme for image denoising is proposed, and the determination of the optimal wavelet basis with respect to the proposed scheme is also discussed. The overcomplete wavelet expansion (OWE), which is more effective than the orthogonal wavelet transform (OWT) in noise reduction, is used. To explore the strong interscale dependencies of OWE, we combine the pixels at the same spatial location across scales as a vector and apply LMMSE to the vector. Compared with the LMMSE within each scale, the interscale model exploits the dependency information distributed at adjacent scales. The performance of the proposed scheme is dependent on the selection of the wavelet bases. Two criteria, the signal information extraction criterion and the distribution error criterion, are proposed to measure the denoising performance. The optimal wavelet that achieves the best tradeoff between the two criteria can be determined from a library of wavelet bases. To estimate the wavelet coefficient statistics precisely and adaptively, we classify the wavelet coefficients into different clusters by context modeling, which exploits the wavelet intrascale dependency and yields a local discrimination of images. Experiments show that the proposed scheme outperforms some existing denoising methods. Lei Zhang 0006, Paul Bao, Xiaolin Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2005 | On multirate optimality of JPEG2000 code streamabstractArguably, the most important and defining feature of the JPEG2000 image compression standard is its R-D optimized code stream of multiple progressive layers. This code stream is an interleaving of many scalable code streams of different sample blocks. In this paper, we reexamine the R-D optimality of JPEG2000 scalable code streams under an expected multirate distortion measure (EMRD), which is defined to be the average distortion weighted by a probability distribution of operational rates in a given range, rather than for one or few fixed rates. We prove that the JPEG2000 code stream constructed by embedded block coding of optimal truncation is almost optimal in the EMRD sense for uniform rate distribution function, even if the individual scalable code streams have nonconvex operational R-D curves. We also develop algorithms to optimize the JPEG2000 code stream for exponential and Laplacian rate distribution functions while maintaining compatibility with the JPEG2000 standard. Both of our analytical and experimental results lend strong support to JPEG2000 as a near-optimal scalable image codec in a fairly general setting. Xiaolin Wu 0001, Sorina Dumitrescu, Ning Zhang 0023 |
IEEE Trans. Image Process. | 1 |
| 2005 | Color demosaicking via directional linear minimum mean square-error estimationabstractDigital cameras sample scenes using a color filter array of mosaic pattern (e.g., the Bayer pattern). The demosaicking of the color samples is critical to the image quality. This paper presents a new color demosaicking technique of optimal directional filtering of the green-red and green-blue difference signals. Under the assumption that the primary difference signals (PDS) between the green and red/blue channels are low pass, the missing green samples are adaptively estimated in both horizontal and vertical directions by the linear minimum mean square-error estimation (LMMSE) technique. These directional estimates are then optimally fused to further improve the green estimates. Finally, guided by the demosaicked full-resolution green channel, the other two color channels are reconstructed from the LMMSE filtered and fused PDS. The experimental results show that the presented color demosaicking technique outperforms the existing methods both in PSNR measure and visual perception. Lei Zhang 0006, Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2005 | Wavelet coding of volumetric medical images for high throughput and operabilityabstractThis paper presents a new three-dimensional (3-D) wavelet-based scalable lossless coding scheme for compression of volumetric medical images. Aiming to improve the productivity of radiologists and the cost-effectiveness of the system, we strive to achieve high decoder throughput, random access to coded data volume, progressive transmission, and high compression ratio in a balanced design approach. These desirable functionalities are realized by a modified 3-D dyadic wavelet transform tailored to volumetric medical images and an optimized Rice code of very low complexity. Xiaolin Wu 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2005 | Directly operable image representation of multiscale primal sketchabstractIn this paper, we propose a versatile semantics-driven image representation that can support many common operations in visual computing and communications, in addition to being itself an efficient image coding scheme. The proposed image representation is based on a semantically meaningful construct called multiscale primal sketch (MPS). The MPS consists of edges that are extracted and organized successively from fine to coarse scales. The edges are further classified into two types: pulse edge and step edge. MPS is an intermediate-level image representation, which is between pixel-based (low-level) and object-based (high-level) descriptions. The MPS image representation reaches a good compromise between its construction cost and descriptive power. It has a compact form, and hence is amenable to image compression. Furthermore, because the new representation consists of semantically meaningful primitives - edges of different scales and types, and background - many common image operations, such as classification, restoration, detection, and content-based retrieval, can be performed directly in the MPS framework, without first converting the coded image back to the spatial domain. Xiaohui Xue, Xiaolin Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2004 | Fast Algorithms for Optimal Two-description Scalar Quantizer DesignabstractNew efficient algorithms are presented to design globally optimal two-description quantizers of fixed rate. The optimization objective is to minimize the expected distortion at the receiver side. We formulate the problem as one of shortest path in a directed acyclic graph. The fixed rate requirement puts constraints on the number and type of edges of the shortest path, which leads to an O(K/sub 1/K/sub 2/N/sup 3/) time design algorithm, where N is the cardinality of the source alphabet, and K/sub 1/, K/sub 2/ are the number of codecells, respectively, of the two side quantizers. This complexity is reduced to O(K/sub 1/K/sub 2/N/sup 2/) by exploiting a so-called Monge property of the objective function. Furthermore, if K/sub 1/=K/sub 2/=K and the two descriptions are subject to the same channel statistics, then the optimal description quantizer design problem can be solved in O(KN/sup 2/) time. Sorina Dumitrescu, Xiaolin Wu 0001, Gaurav Bahl |
Data Compression Conference | 2 |
| 2004 | Minimax Multiresolution Scalar QuantizationabstractWe consider the problem of design and analysis of optimal L/sub /spl infin// (minmax) multiresolution scalar quantizers (MRSQ). The overall multiresolution L/sub /spl infin// distortion of an MRSQ is denned to be a weighted sum of L/sub /spl infin// distortions over all refinement levels of the MRSQ. The weight for a refinement level usually denotes the probability that the MRSQ will operate at that level (rate). An interesting relation of the problem to the design of optimal binary prefix codes under a code cell contiguity constraint is established: Lower bounds for the overall multiresolution L/sub /spl infin// distortion are derived based on this relation. Provably optimal as well as fast, near optimal algorithms are also developed for practically interesting scenarios. Furthermore, the performance penalty incurred by making a scalable quantizer embedded (progressively refinable) is analyzed. It is shown that constraining the quantizers to be embedded would on average increase the L/sub /spl infin// quantization error by at least 44%. Nima Sarshar, Xiaolin Wu 0001 |
Data Compression Conference | 2 |
| 2004 | On Wavelet Compression of Self-Similar ProcessesabstractSelf-similar stochastic processes are stochastic counterparts of deterministic fractals. Fractional Brownian motion (fBm) is a self-similar nonstationary Gaussian process originally proposed to model power-law behavior of power spectrum of long-range dependant (LRD) natural processes. Multiscale nature of wavelets make them natural candidates for analysis and synthesis of fractional Brownian motions. Despite wavelet compression being the method of choice for image compression, the performance of wavelet compression schemes are investigated for compressing fractional Brownian motions. Theoretical rate-distortion function of fBm is explicitly derived. Nima Sarshar, Xiaolin Wu 0001 |
Data Compression Conference | 2 |
| 2004 | A simple technique for estimating message lengths for additive noise steganographyabstractWe propose a practical steganalysis method to detect hidden information and estimate embedding rate. By modelling steganographic embedding as an additive noise process, we exploit the fact that the mean and variance of stegosignal is a function of embedding rate. This can be used to estimate the embedding rate without knowledge of the original cover object. We present simulation results to demonstrate that the proposed method can estimate embedding rate with reasonable accuracy. The technique we propose is general and can be used for steganalysis of images, video, and audio etc. Ming Jiang 0006, Edward K. Wong, Nasir Memon, Xiaolin Wu 0001 |
ICARCV | 4 |
| 2004 | Fast length-constrained MAP decoding of variable length coded Markov sequences over noisy channelabstractThe problem of maximum a posterior probability (MAP) decoding of a Markov sequence that is variable length coded and transmitted over a binary symmetric channel (BSC) is considered. The number of source symbols in the sequence, if made known to the decoder, can improve MAP decoding performance. But adding a sequence length constraint to MAP decoding problem increases its complexity and also converts the length-constrained MAP decoding problem into one of maximum-weight k-link path in a weighted directed acyclic graph. The corresponding graph optimization problem can be solved by a fast parameterized search algorithm that finds either the exact solution with high probability or a good approximate solution otherwise. The proposed algorithm has lower complexity and superior performance than the previous heuristic algorithms. Zhe Wang 0022, Xiaolin Wu 0001, Sorina Dumitrescu |
ICC | 2 |
| 2004 | Multiple-description geometry compression for networked interactive 3D graphicsabstractAn existing technique for robust streaming of 3D graphics contents over lossy networks is multi-resolution coding of 3D geometry. An advantage of this approach is that it uses refinement layers and therefore multiple clients with different bandwidths can be served by a single unified code stream. However, there is a dependency between refinement layers, called prefix condition. Decoding of a given layer requires the knowledge of all the previous layers. A problem in the base layer reception interrupts the streaming all together and voids the remaining layers even though they are received perfectly. To overcome this drawback, this paper proposes an alternative approach to multi-resolution geometry coding, called multiple-description coding of 3D geometry. Instead of organizing code stream into embedded layers, MDC generates several separate descriptions of a geometric object, called co-descriptions. Each co-description of MDC can be independently decoded without any knowledge of other co-descriptions. Each extra successfully received co-description improves the fidelity of reconstructed geometry regardless of what has been received so far or in what order. Pavel Jaromersky, Xiaolin Wu 0001, Yi-Jen Chiang, Nasir Memon |
ICIG | 2 |
| 2004 | Quantitative steganalysis of binary imagesabstractWe propose a quantitative steganalysis method to detect hidden information embedded by flipping pixels along boundaries in binary images. We model steganographic embedding as an additive noise process and use compression rate as a distinguishing statistic that aids in discriminating between stego-images and cover-images. We specifically use the JBIG 2 binary image compression algorithm to derive a quantitative relation between compression rate and embedding rate. Based on this relationship, a practical steganalysis technique is proposed by examining the change of compression rate as embedding rate increases. Experiments conducted show that the proposed technique can reliably detect a steganographic embedding process that flips boundary pixels. Furthermore, it can estimate embedding rate with reasonable accuracy. Ming Jiang 0006, Nasir Memon, Edward K. Wong, Xiaolin Wu 0001 |
ICIP | 4 |
| 2004 | A classification approach to color demosaickingabstractColor demosaicking for CCD cameras is a task of estimating missing data. Based on different hypotheses on the image structures multiple estimates can be made. Choosing the best estimate becomes a statistical decision problem. We propose an optimal classification technique based on Fisher's discriminant to solve this decision problem. The new technique is more robust and obtains superior visual quality of demosaicked images than existing methods. Cindy Kwan, Xiaolin Wu 0001 |
ICIP | 2 |
| 2004 | Temporal color video demosaicking via motion estimation and optimal data fusingabstractDemosaicking of the color CCD sensor data is an important task in digital image/video acquisition. Due to the Nyquist frequency limit of color filter array (CFA), it is impossible to faithfully reconstruct some high frequency structures in the original scene if one is limited to the mosaic samples of a still frame. However, a color digital video camera captures a sequence of mosaic images and the temporal dimension provides a wealth of information about the scene via camera and object motions. With the help of adjacent CCD frames by motion estimation and statistical data fusion, we can significantly improve the demosaicking quality of the video sequence. This paper presents a general approach of temporal demosaicking. Our experimental results demonstrate clear advantages of the presented temporal color demosaicking approach over its intra-frame counterparts both in PSNR measure and subjective visual quality. Xiaolin Wu 0001, Lei Zhang 0006 |
ICIP | 1 |
| 2004 | Lossless compression of color mosaic imagesabstractWe present a low complexity algorithm for lossless compression of color mosaic images generated by a Bayer CCD color filter array. This algorithm is based on an interesting use of the integer wavelet transform followed by a fast adaptive context-based Golomb-Rice coding. The lossless compression performance of the proposed algorithm is apparently the best reported in the literature so far for color mosaic images. Ning Zhang 0023, Xiaolin Wu 0001 |
ICIP | 2 |
| 2004 | Steganalysis of boundary-based steganography using autoregressive model of digital boundariesabstractIn this paper, we present a novel technique for the steganalysis of digital documents, when the secret message is embedded along the boundaries of text characters or other symbols in the documents. The proposed technique uses an auto-regressive model to detect marked documents, as well as to estimate the relative length of the embedded messages. Experimental results demonstrate the effectiveness and accuracy of the proposed technique Ming Jiang 0006, Xiaolin Wu 0001, Edward K. Wong, Nasir Memon |
ICME | 2 |
| 2004 | Buffer size reduction through buffer sharing for streaming applicationsabstractMany multimedia streaming applications have to buffer a number of different source streams for playback of a single multimedia composition. Multiple buffers (one for each stream) have to be deployed at the decoder to make continuous, almost real-time, playback of the multimedia content possible. We propose a novel implementation of two buffers in one array that efficiently reduces the overall memory dedicated to buffering by allowing for a common space where data for both buffers can be stored. An algorithm for finding optimal parameters of the shared buffer and calculating the reduction in the buffer size by buffer sharing is proposed. This involves finding level sets of solutions to some well studied 2D partial differential equations of mathematical physics with simple boundary conditions. Nima Sarshar, Xiaolin Wu 0001 |
ICME | 2 |
| 2004 | Length-constrained MAP decoding revisitedabstractWe consider the problem of length-constrained maximum a posteriori (MAP) decoding of a Markov sequence that is variable length encoded and transmitted over a binary symmetric channel. We convert this problem into one of maximum-weight k-link path in a weighted directed acyclic graph. The induced graph optimization problem can be solved by a fast parameterized search algorithm that finds either the optimal solution with high probability or a good approximate solution otherwise. The proposed algorithm has lower complexity and superior performance than previous heuristic algorithms. Zhe Wang 0022, Xiaolin Wu 0001, Sorina Dumitrescu |
ICME | 2 |
| 2004 | Lagrangian global optimization of two-description scalar quantizersabstractWe develop an efficient Lagrangian-type algorithm for optimal two-description fixed-rate scalar quantizer design, for a very large class of distortion measures. Our key result is the discovery that the Lagrangian multiplier for the globally optimal solution exists. Although Lagrangian optimization is a method of choice for quantizer design, none of the previous algorithms using this method was shown to guarantee the global optimality for any instance of the problem Sorina Dumitrescu, Xiaolin Wu 0001 |
ISIT | 2 |
| 2004 | Broadcasting with fidelity criteriaabstractConsider the problem of broadcasting an i.i.d. source sequence X = {X/sub i/} /sub i=1//sup N/ (possibly N /spl rarr/ /spl infin/) to n listeners over a discrete broadcast channel, consisting of n channels with capacities C/sub 1/ = C/sub max/ /spl ges/ C/sub 2/ /spl ges/.../spl ges/ C/sub n/ = C/sub min/. Let the tuple D = (D/sub 1/, D/sub 2/,...,D/sub n/) represent the average distortion in reconstructing sources at the n listeners. The problem of characterizing all achievable tuples D is still open for a general case. For a fairly general class of discrete channels, we prove the achievability of the tuple n(/spl rho//sub 1/,/spl rho//sub 2/,...,/spl rho//sub n/) = (D/sub X/(/spl rho//sub 1/C/sub 1/ /spl zeta/), D/sub X/(/spl rho//sub 2/C/sub 2/ - /spl zeta/),...,D/sub X/(/spl rho//sub n/C/sub n/ - /spl zeta/)), provided that /spl lambda//sub i/ = (/spl rho//sub i/C/sub i/ - /spl rho//sub i/+/sub 1/C/sub i+1/)/C/sub i/ > 0, for 1 /spl les/ i /spl les/ n $1, /spl lambda//sub n/ = /spl rho//sub n/ and /spl Sigma//sub i=1//sup n-1/ /spl lambda//sub i/ /spl les/ 1, where D/sub X/ (R) is the distortion rate function of X. The penalty term /spl zeta/ = 1/2 for a general source with real alphabets and is /spl zeta/ = 0 if X is progressively refinable. The factor 00, we find examples of channels for which /sup 3/(2/3+/spl delta/,2/3+/spl delta/,2/3+ /spl delta/) is not achievable. Nima Sarshar, Xiaolin Wu 0001 |
ITW | 2 |
| 2004 | Optimal unequal channel protection of multiple-description product codes for multimedia communications over fast fading channelsabstractIn recent literature a powerful multiple description product coding scheme for protection of progressively encoded source streams has been devised that disperses information evenly between all description packets. Also, techniques were proposed to protect these packets equally by an optimal channel coder (found by exhaustive search). The contribution of this paper is to show that equal protection of all descriptions is suboptimal when the channel varies with time despite the fact that all descriptions have equal importance. We propose a theoretical framework for computing the globally optimal channel protection assignment for a given set of available channel coders under some idealized assumptions. For more practical scenarios we propose an optimized uneven packet protection scheme that outperforms equal protection schemes in terms of the expected distortion of received sources. Simulations of an image transmission system that resembles a 3G high bitrate link is provided where our unequal protection scheme improves the average PSNR of the received images by more than 1.3dB. Nima Sarshar, Xiaolin Wu 0001 |
WCNC | 2 |
| 2004 | Optimal context quantization in lossless compression of image data sequencesabstractIn image compression context-based entropy coding is commonly used. A critical issue to the performance of context-based image coding is how to resolve the conflict of a desire for large templates to model high-order statistic dependency of the pixels and the problem of context dilution due to insufficient sample statistics of a given input image. We consider the problem of finding the optimal quantizer Q that quantizes the K-dimensional causal context Ct = (Xt-t1,Xt-t2,...,X t-tK) of a source symbol Xt into one of a set of conditioning states. The optimality of context quantization is defined to be the minimum static or minimum adaptive code length of given a data set. For a binary source alphabet an optimal context quantizer can be computed exactly by a fast dynamic programming algorithm. Faster approximation solutions are also proposed. In case of m-ary source alphabet a random variable can be decomposed into a sequence of binary decisions, each of which is coded using optimal context quantization designed for the corresponding binary random variable. This optimized coding scheme is applied to digital maps and alpha-plane sequences. The proposed optimal context quantization technique can also be used to establish a lower bound on the achievable code length, and hence is a useful tool to evaluate the performance of existing heuristic context quantizers. Søren Forchhammer, Xiaolin Wu 0001, Jakob Dahl Andersen |
IEEE Trans. Image Process. | 2 |
| 2004 | Primary-consistent soft-decision color demosaicking for digital cameras (patent pending)abstractColor mosaic sampling schemes are widely used in digital cameras. Given the resolution of CCD sensor arrays, the image quality of digital cameras using mosaic sampling largely depends on the performance of the color demosaicking process. A common problem with existing color demosaicking algorithms is an inconsistency of sample interpolations in different primary color channels, which is the cause of the most objectionable color artifacts. To cure the problem, we propose a new primary-consistent soft-decision framework (PCSD) of color demosaicking. In the PCSD framework, we make multiple estimates of a missing color sample under different hypotheses on edge or texture directions. The estimates are made via a primary consistent interpolation, meaning that all three primary components of a color are interpolated in the same direction. The final estimate of a color sample is obtained by testing different interpolation hypotheses in the reconstructed full-resolution color image and selecting the best via an optimal statistical decision or inference process. A concrete color demosaicking method of the PCSD framework is presented. This new method eliminates certain types of color artifacts of existing color demosaicking methods. Extensive experimental results demonstrate that the PCSD approach can significantly improve the image quality of digital cameras in both subjective and objective measures. In some instances, our gain over the competing methods can be as much as 7 dB. Xiaolin Wu 0001, Ning Zhang 0023 |
IEEE Trans. Image Process. | 1 |
| 2004 | Monotonicity-based fast algorithms for MAP estimation of Markov sequences over noisy channelsabstractIn this correspondence, we study algorithmic approach to solving the problem of maximum a posteriori (MAP) estimation of Markov sequences transmitted over noisy channels, which is also known as the MAP decoding problem. For the class of memoryless binary channels that produce independent substitution and erasure errors, the MAP sequence estimation problem can be formulated and solved as one of the longest path in a weighted directed acyclic graph. But for algorithm efficiency, we transform the graph problem to one of matrix search. If the underlying matrix is totally monotone, then the complexity of MAP sequence estimation can be greatly reduced. We give a sufficient condition for the matrix induced by MAP sequence estimation to be totally monotone, which is indeed the case if the input sequence is Gaussian Markov. Under this condition, the complexity of MAP decoding can be reduced from O(N/sup 2/M) to O(NM), where N is the size of source alphabet and M is the length of input sequence. Furthermore, for Markov sequences of fixed-length code we propose a block parsing strategy to reduce the complexity of MAP sequence estimation to O(M+N/sup 2/M/logM) or to O(M+NM/logM), depending on if the total monotonicity holds. Another significance of this correspondence lies in the applicability of the presented algorithmic approach, which has been thoroughly studied in computer science literature, to many other discrete optimization problems encountered in both source and channel coding, ranging from optimal multiresolution and multiple-description quantizer design, to context quantization for minimum conditional entropy, and to optimal packetization with uneven error protection. Xiaolin Wu 0001, Sorina Dumitrescu, Zhe Wang 0022 |
IEEE Trans. Inf. Theory | 1 |
| 2004 | Wavelet Estimation of Fractional Brownian Motion Embedded in a Noisy EnvironmentabstractThis correspondence proposes a wavelet-based fractional Brownian motion (fBm) signal estimation scheme. Despite the fact that wavelet transform approximately whitens the fBm processes, it is observed that statistical dependencies still exist across adjacent wavelet scales and between neighboring wavelet coefficients. These dependencies can be exploited to improve the estimation of fBm signals embedded into noise. The idea is to reorganize the wavelet coefficients into a scale-time mixture model and then carry out the minimum mean-square-error estimation (MMSE) using the model. Experiments show that the proposed scheme obtains better estimates than Wornell and Oppenheim's algorithm, in which the wavelet dependencies are not utilized. Lei Zhang 0006, Paul Bao, Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 3 |
| 2004 | Globally optimal uneven error-protected packetization of scalable code streamsabstractIn this paper, we present a family of new algorithms for rate-fidelity optimal packetization of scalable source bit streams with uneven error protection. In the most general setting where no assumption is made on the probability function of packet loss or on the rate-fidelity function of the scalable code stream, one of our algorithms can find the globally optimal solution to the problem in O(N/sup 2/L/sup 2/) time, compared to a previously obtained O(N/sup 3/L/sup 2/) complexity, where N is the number of packets and L is the packet payload size. If the rate-fidelity function of the input is convex, the time complexity can be reduced to O(NL/sup 2/) for a class of erasure channels, including channels for which the probability function of losing n packets is monotonically decreasing in n and independent erasure channels with packet erasure rate no larger than N/2(N + 1). Furthermore, our O(NL/sup 2/) algorithm for the convex case can be modified to rind an approximation solution for the general case. All of our algorithms do away with the expediency of fractional bit allocation, a limitation of some existing algorithms. Sorina Dumitrescu, Xiaolin Wu 0001, Zhe Wang 0022 |
IEEE Trans. Multim. | 2 |
| 2003 | Optimal Alphabet Partitioning for Semi-Adaptive Coding of Sources of Unknown Sparse DistributionsabstractPractical applications that employ entropy coding for large alphabets often partition the alphabet set into two or more layers. Each symbol was encoded using suitable prefix coding for each layer. The problem of optimal alphabet partitioning was formulated for the design of a two layer semi-adaptive code and the given solution was based on dynamic programming. However, the complexity of the dynamic programming approach can be quite prohibitive for a long sequence and very large alphabet size. Hence, a simple greedy heuristic algorithm whose running time is linear in the number of symbols being encoded was given, irrespective of the underlying alphabet size. The given experimental results demonstrated the fact that superior prefix coding schemes for large alphabets can be designed using this approach as opposed to the typically ad-hoc partitioning approach applied in the literature. Yi-Jen Chiang, Nasir Memon, Xiaolin Wu 0001 |
DCC | 4 |
| 2003 | Optimal Variable Rate Multiplexing of Scalable Code StreamsabstractSummary form only given. The problem of multiplexing of several rate-distortion scalable multimedia code streams to be transmitted via a variable-rate channel is investigated. The aim is to minimize the expected distortion at the receiver, weighted by the probability distribution of the truncation point, P(l). Firstly, a very simple algorithm is presented which solves the problem in the case when the distribution of the truncation point is uniform or exponential. Secondly, the problem of constructing the optimal code stream is treated with the side information included. In order to reduce the amount of side information, a trade off between the optimal order of the /spl alpha/-atom strings described and a predefined order is proposed. Finally, a dynamic programming algorithm is presented to determine the optimal number of refinement stages. The time complexity of this algorithm is quadratic in the total number of /spl alpha/-atom strings. Sorina Dumitrescu, Xiaolin Wu 0001 |
DCC | 2 |
| 2003 | Searchable Compressed Representations of Very Sparse Bitmaps (extended abstract)abstractVery sparse bitmaps are used in a wide variety of applications, ranging from adjacency matrices in representation of large sparse graphs, representation of sparse space occupancy to book-keeping in databases. A method based on pruning of the binary space partition (BSP) tree in the minimal description length (MDL) principle for coding very sparse bitmaps was proposed. This new method for coding of sparse bitmaps meets seemingly competing objectives of good compression, the ability of conducting queries directly in the compression domain, and simple and fast decoding. Steven Pigeon, Xiaolin Wu 0001 |
DCC | 2 |
| 2003 | On optimality of JPEG2000 code streamabstractArguably the most important and defining feature of JPEG 2000 image compression standard is its R-D optimized code stream of multiple progressive layers. This code stream is an interleaving of many scalable code streams of different sample blocks. In this paper we investigate the algorithms of optimizing the interleaving to minimize the expected distortion weighted by the probability distribution of operational rates in a given range, rather than for one or few fixed rates. We prove that the JPEG 2000 code stream constructed by EBCOT (embedded block coding of optimal truncation) is indeed optimal for uniform rate distribution function even if the individual scalable code streams have non-convex operational R-D curves. We also develop algorithms to optimize the JPEG 2000 code stream for exponential and Laplacian rate distribution functions. Sorina Dumitrescu, Xiaolin Wu 0001 |
ICIP (3) | 2 |
| 2003 | Primary-consistent soft-decision color demosaic for digital camerasabstractBayer color mosaic sampling scheme is widely used in digital cameras. Given the resolution of CCD sensor arrays, the image quality of digital cameras using Bayer sampling mosaic largely depends on the performance of the color demosaic process. A common and serious weakness shared by all existing color demosaic algorithms is an inconsistency of sample interpolations in different primary color components, which is the culprit for the most objectionable color artifacts. To cure the problem we propose a primary-consistent color demosaic algorithm. The performance of this algorithm is further enhanced by a soft-decision sample interpolation scheme. Experiments demonstrate that the proposed framework of primary-consistent soft-decision color demosaic can significantly improve the image quality of digital cameras. Xiaolin Wu 0001, Ning Zhang 0023 |
ICIP (1) | 1 |
| 2003 | MAP decoding of variable length code with substitution, insertion and deletionabstractWe propose a soft decision decoding algorithm for a variable-length encoded Markov source transmitted via a noisy channel that has all three types of errors - substitution, deletion, and insertion. The decoder aims to find a sequence, X, that maximizes the posterior probability, given a received sequence, Y. First, we assume the channel is a binary symmetric channel (BSC), and the input of the channel is a first order Markov source. We convert the MAP (maximum a posteriori) problem to one of finding a single-source longest path in a directed acyclic graph, which can be solved by dynamic programming. We also present a generalization of the BSC to include both insertion and deletion errors, and the development of a soft decision MAP decoding algorithm for this more challenging case, which is our main contribution. Zhe Wang 0022, Xiaolin Wu 0001 |
ITW | 2 |
| 2003 | Hybrid inter- and intra-wavelet scale image restoration
Lei Zhang 0006, Paul Bao, Xiaolin Wu 0001 |
Pattern Recognit. | 3 |
| 2003 | Lossy-to-Lossless Compression of Medical Volumetric Data Using Three-dimensional Integer Wavelet TransformsabstractWe study lossy-to-lossless compression of medical volumetric data using three-dimensional (3-D) integer wavelet transforms. To achieve good lossy coding performance, it is important to have transforms that are unitary. In addition to the lifting approach, we first introduce a general 3-D integer wavelet packet transform structure that allows implicit bit shifting of wavelet coefficients to approximate a 3-D unitary transformation. We then focus on context modeling for efficient arithmetic coding of wavelet coefficients. Two state-of-the-art 3-D wavelet video coding techniques, namely, 3-D set partitioning in hierarchical trees (Kim et al., 2000) and 3-D embedded subband coding with optimal truncation (Xu et al., 2001), are modified and applied to compression of medical volumetric data, achieving the best performance published so far in the literature-both in terms of lossy and lossless compression. Zixiang Xiong, Xiaolin Wu 0001, Samuel Cheng 0001, Jianping Hua |
IEEE Trans. Medical Imaging | 2 |
| 2002 | Globally Optimal Uneven Error-Protected Packetization of Scalable Code StreamsabstractIn this extended abstract we present a family of new algorithms for rate-fidelity optimal packetization of scalable source bit stream with uneven error protection. In the most general setting where no assumption is made on the probability function of packet loss or on the rate-fidelity function of the scalable code stream, one of our algorithms can find the globally optimal solution to the problem in O(N/sup 2/L/sup 2/) time, compared to a previously claimed O(N/sup 3/L/sup 2/) complexity, where N is the number of packets and L is the packet payload size. The time complexity can be reduced to O(NL/sup 2/) if the rate-fidelity function of the input is convex and under the reasonable assumption that the probability function of packet loss is monotonically decreasing. In the convex case the algorithm of Mohr et al. (2000) has complexity O(N/sup 2/L log N). Furthermore, our O(NL/sup 2/) algorithm for the convex case can be modified to find an approximation solution for the general case that is better than the results of other algorithms in the prior literature. All of our algorithms do away with the expediency of fractional redundancy allocation, a limitation of some existing algorithms. To our best knowledge this work offers for the first time globally optimal solutions to the important problem of optimal UEP packetization. Sorina Dumitrescu, Xiaolin Wu 0001, Zhe Wang 0022 |
DCC | 2 |
| 2002 | On Optimal Multi-resolution Scalar QuantizationabstractAny scalar quantizer of 2/sup h/ bins, where h is a positive integer, can be structured by a balanced binary quantizer tree T of h levels. Any pruned subtree /spl tau/ of T corresponds to an operational rate R(/spl tau/) and distortion D(/spl tau/) pair. Denote by S/sub n/ the set of all pruned subtrees of n leaf nodes, 1/spl les/n/spl les/2/sup h/. We consider the problem of designing a 2/sup h/-bin scalar quantizer that minimizes the weighted average distortion D~=/spl Sigma//sub n=1//sup 2(h)/ D(/spl tau/)W(n), where W(n) is a weighting function in the size of pruned subtrees (or the resolution of the underlying quantizer). We present an O(hN/sup 3/) algorithm to solve the underlying optimization problem (N is the number of points of the histogram that represents the source probability mass function), and call the resulting quantizer optimal multi-resolution scalar quantizer in the sense that it minimizes a global distortion measure averaged over all quantization resolutions of T. Interestingly, a similar quantizer design problem studied by Brunk et al. (1996) is a special case of our formulation, and can thus be solved exactly and efficiently using our algorithm. Furthermore, we present an algorithm to generate a sequence of 2/sup h/ nested pruned subtrees of T, from the root of T to the entire tree T itself, which minimizes an expected distortion over a range of operational rates. The resulting nested pruned subtree sequence generates an optimized embedded (rate-distortion scalable) code stream with maximum granularity of 2/sup h/ quantization stages, as opposed to existing successively refinable quantizers, such as the popular bit-plane coding scheme, which offer only h stages. Xiaolin Wu 0001, Sorina Dumitrescu |
DCC | 1 |
| 2002 | On steganalysis of random LSB embedding in continuous-tone imagesabstractWe present an LSB steganalysis technique that can detect the existence of hidden messages that are randomly embedded in the least significant bits of natural continuous-tone images. The technique is inspired by the recent work of J. Fridrich et al. (see Proc. ACM Workshop on Multimedia and Security, p.27-30, 2001) and just like their work, it can also precisely measure the length of the embedded message, even when the hidden message is very short relative to the image size. The key to our success is the formation of some subsets of pixels whose cardinalities change with LSB embedding, and such changes can be precisely quantified under the assumption that the embedded bits are randomly scattered. Interestingly, our study on steganalysis of LSB embedding sheds light on the work of Fridrich et al. on the detection of LSB embedding, and offers an analytical proof of an observation made by them. Sorina Dumitrescu, Xiaolin Wu 0001, Nasir Memon |
ICIP (3) | 2 |
| 2002 | Steganalysis of LSB embedding in multimedia signalsabstractWe present a new, principled steganalytic approach to estimate the length of messages, if any, hidden in the least significant bits of digitized continuous signals such as images and audio. Our estimates of LSB-embedded messages are based on some intrinsic statistical properties of sample pairs drawn from the signal. We also give bounds on the estimation errors. Furthermore, possible attacks on our new steganalytic approach and our countermeasures are also discussed. Sorina Dumitrescu, Xiaolin Wu 0001 |
ICME (1) | 2 |
| 2002 | Optimal multiresolution quantization for scalable multimedia codingabstractWe have investigated the problem of designing an optimal entropy-constrained multiresolution scalar quantizer under the criterion of minimizing the expected distortion weighted by the probability of transmission rate and for arbitrary probability mass function of signal amplitude. An O(LN/sup 3/) algorithm is proposed to solve the optimization problem, where L is the number of refinement stages, and N is the size of input symbol alphabet. The proposed algorithm is globally optimal, and furthermore, it is more general than its locally optimal predecessors in terms of quantizer structures. Sorina Dumitrescu, Xiaolin Wu 0001 |
ITW | 2 |
| 2001 | Lossless Image Data Sequence Compression Using Optimal Context QuantizationabstractContext based entropy coding often faces the conflict of a desire for large templates and the problem of context dilution. We consider the problem of finding the quantizer Q that quantizes the K-dimensional causal context C/sub i/=(X(i-t/sub 1/), X(i-t/sub 2/), ..., X(i-t/sub K/)) of a source symbol X/sub i/ into one of M conditioning states. A solution giving the minimum adaptive code length for a given data set is presented (when the cost of the context quantizer is neglected). The resulting context quantizers can be used for sequential coding of the sequence X/sub 0/, X/sub 1/, X/sub 2/, .... A coding scheme based on binary decomposition and context quantization for coding the binary decisions is presented and applied to digital maps and /spl alpha/-plane sequences. The optimal context quantization is also used to evaluate existing heuristic context quantizations. Søren Forchhammer, Xiaolin Wu 0001, Jakob Dahl Andersen |
Data Compression Conference | 2 |
| 2001 | Joint UEP and layered source coding with application to transmission of JPEG-2000 coded imagesabstractThis paper presents a joint source-channel coding framework based on layered source coding and Reed-Solomon channel coding for unequal error protection. An iterative procedure is described to search for the best source coding rate and the optimal UEP of layered bitstreams. We apply our JSCC technique to transmission of JPEG-2000 coded images over binary symmetric channels. Compared to results reported in the literature, our UEP based approach gives better results while having lower complexity. Tianli Chu, Zhongmin Liu, Zixiang Xiong, Xiaolin Wu 0001 |
GLOBECOM | 4 |
| 2001 | High-performance 3-D embedded wavelet video (EWV) codingabstractThis paper presents a rate-distortion (R-D) optimized 3-D embedded wavelet video (EWV) coder by extending the concept of EBCOT from 2-D to 3-D. After a lifting based 3-D wavelet transform, different subbands are coded independently using bit plane coding with different context models to provide flexible scalability in both spatial and temporal domain. A global R-D optimization procedure is used to generate an embedded bitstream for a target bit rate. Experiments show that, even without motion estimation, the EWV coder outperforms both MPEG-4 and 3-D ESCOT for most low motion video sequences. Jianping Hua, Zixiang Xiong, Xiaolin Wu 0001 |
MMSP | 3 |
| 2001 | On packetization of embedded multimedia bitstreamsabstractWe study the problem of packetizing embedded multimedia bitstreams to improve the error resilience of source (compression) codes. This problem is important because of the increasing popularity of embedded compression methodology and its suitability for scalable streaming media over IP or/and mobile IP. We study various packetization schemes against packet erasure at both low and high bit rates. Maximizing packetization efficiency for embedded bitstreams is formulated as a discrete optimization problem and globally optimal packetization (OP) algorithms are proposed under different settings. Suboptimal packetization algorithms are also devised to reduce the complexity of the OP algorithms. In order to assess their effectiveness, the proposed packetization algorithms are used to packetize embedded image and video bitstreams with simulated packet loss. Experimental results show that our OP algorithms slightly outperforms suboptimal ones. In addition to confirming the superiority of the OP algorithms, these results also provide justification of heuristic packetization methods published in the literature. Xiaolin Wu 0001, Samuel Cheng 0001, Zixiang Xiong |
IEEE Trans. Multim. | 1 |
| 2000 | Optimal Packetization of Embedded BitstreamsabstractSummary form only given. To achieve error resilience, digital communication systems typically partition a data file into blocks of samples in the time domain which represent small cohesive segments of the input source, be it image, video or audio. We consider the problem of packing a set of embedded bitstreams B/sub k/ into M packets of payload L. Optimal packetization is to select ML bits to fill in M packets while satisfying certain alignment constraints that are imposed by error resilience designs, and at the same time minimizing the distortion. Just as source and channel coding have conflicting objectives of removing and adding redundancy, error resilience via packetization will somewhat reduce the rate distortion performance of the compression code when the transmission is error free. Our goal is to minimize such losses of coding efficiency. Xiaolin Wu 0001, Zixiang Xiong |
Data Compression Conference | 1 |
| 2000 | Scalable Lossy to Lossless Video Coding via Adaptive 3D Wavelet Transform and Context ModelingabstractWe investigated scalable compression of image sequences using 3D wavelet transform and adaptive arithmetic coding driven by 3D context modeling. The interplays between motion compensation, 3D wavelet transform, and entropy coding were studied empirically. The best compression, in both lossy and lossless cases, was achieved by a non-dyadic 3D wavelet decomposition and high-order Markov modeling of transform coefficients, but without motion compensation. The experimental results are very encouraging. On the mother-and-daughter and salesman QCIF sequences, the proposed method outperforms H.263 by up to 3 dB. If integer wavelets are used, our method also improves the lossless performance of other 3D integer wavelet methods by an appreciable margin. Xiaolin Wu 0001, Zhenchu Xiao |
ICIP | 2 |
| 2000 | Linfinity constrained high-fidelity image compression via adaptive context modelingabstractIn this paper, we study high-fidelity image compression with a given tight L(infinity) bound. We propose some practical adaptive context modeling techniques to correct prediction biases caused by quantizing prediction residues, a problem common to the existing DPCM-type predictive near-lossless image coders. By incorporating the proposed techniques into the near-lossless version of CALIC that is considered by many as the state-of-the-art algorithm, we were able to increase its PSNR by 1 dB or more and/or reduce its bit rate by 10% or more, more encouragingly, at bit rates around 1.25 bpp or higher, our method obtained competitive PSNR results against the best L(2)-based wavelet coders, while obtaining much smaller L(infinity) bound. Xiaolin Wu 0001, Paul Bao |
IEEE Trans. Image Process. | 1 |
| 2000 | Context-based lossless interband compression-extending CALICabstractThis paper proposes an interband version of CALIC (context-based, adaptive, lossless image codec) which represents one of the best performing, practical and general purpose lossless image coding techniques known today. Interband coding techniques are needed for effective compression of multispectral images like color images and remotely sensed images. It is demonstrated that CALIC's techniques of context modeling of DPCM errors lend themselves easily to modeling of higher-order interband correlations that cannot be exploited by simple interband linear predictors alone. The proposed interband CALIC exploits both interband and intraband statistical redundancies, and obtains significant compression gains over its intrahand counterpart. On some types of multispectral images, interband CALIC can lead to a reduction in bit rate of more than 20% as compared to intraband CALIC. Interband CALIC only incurs a modest increase in computational cost as compared to intraband CALIC. Xiaolin Wu 0001, Nasir Memon |
IEEE Trans. Image Process. | 1 |
| 1999 | Resynchronization Properties of Arithmetic CodingabstractSummary form only given. Arithmetic coding is a popular and efficient lossless compression technique that maps a sequence of source symbols to an interval of numbers between zero and one. We consider the important problem of decoding an arithmetic code stream when an initial segment of that code stream is unknown. We call decoding under these conditions resynchronizing an arithmetic code. This problem has importance in both error resilience and cryptology. If an initial segment of the code stream is corrupted by channel noise, then the decoder must attempt to determine the original source sequence without full knowledge of the code stream. In this case, the ability to resynchronize helps the decoder to recover from the channel errors. But in the situation of encryption one would like to have very high time complexity for resynchronization. We consider the problem of resynchronizing simple arithmetic codes. This research lays the groundwork for future analysis of arithmetic codes with high-order context models. In order for the decoder to achieve full resynchronization, the unknown, initial b bits of the code stream must be determined exactly. When the source is approximately IID, the search complexity associated with choosing the correct sequence is at least O(2/sup b/2/). Therefore, when b is 100 or more, the time complexity required to achieve full resynchronization is prohibitively high. To partially resynchronize, the decoder must determine the coding interval after b bits have been output by the encoder. For a stationary source and a finite-precision static binary arithmetic coder, the complexity of determining the code interval is O(2/sup 2s/), where the precision is s bits. Peter W. Moo, Xiaolin Wu 0001 |
Data Compression Conference | 2 |
| 1999 | Context Quantization with Fisher Discriminant for Adaptive Embedded Wavelet Image CodingabstractRecent progress in context modeling and adaptive entropy coding of wavelet coefficients has probably been the most important catalyst for the rapidly maturing area of wavelet image compression technology. In this paper we identify statistical context modeling of wavelet coefficients as the determining factor of rate-distortion performance of wavelet codecs. We propose a new context quantization algorithm for minimum conditional entropy. The algorithm is a dynamic programming process guided by Fisher's linear discriminant. It facilitates high-order context modeling and adaptive entropy coding of embedded wavelet bit streams, and leads to superb compression performance in both lossy and lossless cases. Xiaolin Wu 0001 |
Data Compression Conference | 1 |
| 1999 | Low Complexity High-Order Context Modeling of Embedded Wavelet Bit StreamsabstractIn the past three or so years, particularly during the JPEG 2000 standardization process that was launched last year, statistical context modeling of embedded wavelet bit streams has received a lot of attention from the image compression community. High-order context modeling has been proven to be indispensable for high rate-distortion performance of wavelet image coders. However, if care is not taken in algorithm design and implementation, the formation of high-order modeling contexts can be both CPU and memory greedy, creating a computation bottleneck for wavelet coding systems. In this paper we focus on the operational aspect of high-order statistical context modeling, and introduce some fast algorithm techniques that can drastically reduce both time and space complexities of high-order context modeling in the wavelet domain. Xiaolin Wu 0001 |
Data Compression Conference | 1 |
| 1999 | Resynchronization Properties of Arithmetic CodingabstractThis paper considers decoding an arithmetic code stream when an initial portion of the code stream is unknown. Full resynchronization is hypothesized to have complexity that is exponential in the length of the initial portion. Experimental results specify the time complexity of determining the current arithmetic code interval, which is the important task in partial resynchronization. Peter W. Moo, Xiaolin Wu 0001 |
ICIP (2) | 2 |
| 1999 | Image Compression Based on Multi-Scale Edge CompensationabstractIn this paper we introduce a new image model called MSEC (Multi-Scale Edge Compensation) and implement an image compression system based on MSEC. MSEC consists of semantically meaningful components including edges of different types and scales, and the background. The MSEC-based image coding technique permits some image analysis operations such as edge detection, enhancement, and content-based retrieval to be carried out directly in the compression domain. Xiaohui Xue, Xiaolin Wu 0001 |
ICIP (3) | 2 |
| 1999 | Wavelet image coding using trellis coded space-frequency quantizationabstractThe progress in wavelet image coding have brought the field into its maturity. Major developments in the process are rate-distortion (R-D) based wavelet packet transformation, zerotree quantization, subband classification and trellis-coded quantization, and sophisticated context modeling in entropy coding. Drawing from past experience and recent in sights, we propose a new wavelet image coding technique with trellis coded space-frequency quantization (TCSFQ). TCSFQ aims to explore space-frequency characterizations of wavelet image representations via R-D optimized zerotree pruning, trellis-coded quantization, and context modeling in entropy coding. Experiments indicate that the TCSFQ coder achieves twice as much compression as the baseline JPEG coder does at the same peak signal to noise ratio (PSNR), making it better than all other coders described in the literature. Zixiang Xiong, Xiaolin Wu 0001 |
IEEE Signal Process. Lett. | 2 |
| 1999 | Conditional entropy coding of VQ indexes for image compressionabstractBlock sizes of practical vector quantization (VQ) image coders are not large enough to exploit all high-order statistical dependencies among pixels. Therefore, adaptive entropy coding of VQ indexes via statistical context modeling can significantly reduce the bit rate of VQ coders for given distortion. Address VQ was a pioneer work in this direction. In this paper we develop a framework of conditional entropy coding of VQ indexes (CECOVI) based on a simple Bayesian-type method of estimating probabilities conditioned on causal contexts, CECOVI is conceptually cleaner and algorithmically more efficient than address VQ, with address-VQ technique being its special case. It reduces the bit rate of address VQ by more than 20% for the same distortion, and does so at only a tiny fraction of address VQ's computational cost. Xiaolin Wu 0001, Jiang Wen, Wing Hung Wong |
IEEE Trans. Image Process. | 1 |
| 1998 | Hybrid Image Compression Scheme Based on Wavelet Transform and Adaptive Context ModelingabstractSummary form only given. We propose a hybrid image compression scheme based on wavelet transform, HVS thresholding and L/sub /spl infin//-constrained adaptive context modelling. This hybrid system combines the strengths of the wavelet transform, the HVS thresholding and the adaptive context modelling to result in a near optimal compression scheme. The wavelet transform is very powerful in localizing the global spatial and frequency correlation. The HVS model-based thresholding is designed to exploit and eliminate the wavelet coefficients insensitive to the human visual system. The context-based modelling is superior in decorrelating the local redundancy. In the scheme, the image is first decomposed into the multiresolution subimages using the orthogonal wavelet transform; each subimage corresponds to a octave band in the wavelet decomposition. The coefficients in the high-pass octave bands of the wavelet transform are then quantized through HVS frequency- and spatial model-based thresholding and vector quantization into wavelet decomposition with only significant coefficients to the HVS retained. In this HVS quantized wavelet decomposition, the coefficients insignificant to the human visual system are normalized to zero and the global spatial and frequency correlation are exploited and removed. Then the quantized subimages in the low-pass band and the remaining high-pass octave bands of each octave level are processed using the L/sub /spl infin//-constrained CALIC to de-correlate the local redundancy. It is demonstrated that the hybrid scheme is one of the best compression schemes in achieving the excellent compression rates and competitive PSNR while maintaining a small visual distortion. In comparing with the original CALIC, we were able to increase the PSNR by 0.65 dB or more and obtain bit rates 15 percent lower than the latter. We were also able to obtain competitive PSNR results against the best wavelet coders, while maintaining a smaller visual distortion. In particular, the wavelet CALIC was able to obtain 1.34 to 7.84 dB higher PSNR on the standard ISO test benchmarks than the SPIHT, one of the best wavelet coder. Paul Bao, Xiaolin Wu 0001 |
Data Compression Conference | 2 |
| 1998 | Analysis of Trellis Quantization for Near-Lossless Image CodingabstractSummary form only given. We discuss several variations to the original algorithm proposed by Ke and Marcellin (see Proc. IEEE ICIP, Washington DC, 1995). We have extended the trellis quantization (TQ) scheme by performing two-row joint optmizations instead of optimizing row by row. Unfortunately, while increasing the computation time quite a bit, this has lead only to marginal coding gains. A progressive probability update scheme has lead to much better convergence and to a 0.3 bpp gain over the original fixed scheme. When using lossy plus near-lossless coding the lossy version can be used for better context modelling without increasing the computational complexity of the near-lossless residual coding. Improvements of 0.1-0.2 bpp were observed. Since it is computationally infeasible to include more pixels to be determined by the TQ process, one has the choice of either using better prediction/context-modelling or doing TQ. Our tests indicate that the preference should be given to sophisticated prediction/modelling. Hannes Hartenstein, Xiaolin Wu 0001 |
Data Compression Conference | 2 |
| 1998 | Lossless Interframe Image Compression via Context ModelingabstractIn this paper, we present an interband version of CALIC (context-based adaptive lossless image codec), a lossless image coding technique. It is demonstrated that CALIC's techniques of context-based modeling of images lend themselves easily to modeling of image sequences. The generalized interframe CALIC can exploit both interframe and intraframe statistical redundancies, and obtain significant compression gains over intraframe CALIC. The advantage of interframe CALIC is demonstrated by experimental results on different types of multispectral images. Xiaolin Wu 0001, Wai Kin Choi, Nasir Memon |
Data Compression Conference | 1 |
| 1998 | Improved Techniques for Lossless Image Compression with Reversible Integer Wavelet Transforms
Nasir Memon, Xiaolin Wu 0001, Boon-Lock Yeo |
ICIP (3) | 2 |
| 1998 | Adaptation to Nonstationarity of Embedded Wavelet Code StreamabstractWe address a problem that affects the performance of all embedded wavelet image codecs-how to track time varying statistics of the coefficient code stream. We propose a family of techniques for rapid adaptation of arithmetic coding to changing statistics of the wavelet coefficients. These techniques improve the compression performance of existing wavelet image codecs, particularly at high bit rates and for lossless coding. Xiaolin Wu 0001, Kai Uwe Barthel, Gerhard Ruhl |
ICIP (2) | 1 |
| 1998 | Piecewise 2D Autoregression for Predictive Image Coding
Xiaolin Wu 0001, Kai Uwe Barthel |
ICIP (3) | 1 |
| 1998 | Order statistics preserving near-lossless image codingabstractWe introduce a new concept of near-lossless image compression called order statistics preserving (OSP) near-lossless coding. Unlike ubiquitous L/sub 2/ (PSNR) and common near-lossless criterion of L/sub /spl infin// the OSP is a context-based fidelity measure that can meet more stringent requirements of high-end users in medical, space, and scientific communities. Xiaolin Wu 0001, Xuehong Li |
MMSP | 1 |
| 1998 | Wavelet image coding using trellis coded space-frequency quantizationabstractWe propose a new wavelet image coding technique with trellis coded space-frequency quantization (TCSFQ). Experiments indicate that the TCSFQ coder is better than all other coders described in the literature. Zixiang Xiong, Xiaolin Wu 0001 |
MMSP | 2 |
| 1998 | Progressive coding of medical volumetric data using three-dimensional integer wavelet packet transformabstractWe examine progressive lossy to lossless compression of medical volumetric data using three-dimensional (3D) integer wavelet packet transforms and set partitioning in hierarchical trees (SPIHT). To achieve good lossy coding performance, we describe a 3D integer wavelet packet transform that allows implicit bit shifting of wavelet coefficients to approximate a 3D unitary transformation. We also address context modeling for efficient entropy coding within the SPIHT framework. Both lossy and lossless coding performance are better than those previously reported. Zixiang Xiong, Xiaolin Wu 0001, David Y. Y. Yun, William A. Pearlman |
MMSP | 2 |
| 1998 | Linfin-Constrained near-lossless image compression using weighted finite automata encoding
Paul Bao, Xiaolin Wu 0001 |
Comput. Graph. | 2 |
| 1997 | Linfty-Constrained High-Fidelity Image Compression via Adaptive Context ModelingabstractWe study high-fidelity image compression with a given tight bound on the maximum error magnitude. We propose some practical adaptive context modeling techniques to correct prediction biases caused by quantizing prediction residues, a problem common to the current DPCM like predictive nearly-lossless image coders. By incorporating the proposed techniques into the nearly-lossless version of CALIC, we were able to increase its PSNR by 1 dB or more and/or reduce its bit rate by ten per cent or more. More encouragingly, at bit rates around 1.25 bpp our method obtained competitive PSNR results against the best wavelet coders, while obtaining much smaller maximum error magnitude. Xiaolin Wu 0001, Wai Kin Choi, Paul Bao |
Data Compression Conference | 1 |
| 1997 | Conditional Entropy Coding of VQ Indexes for Image CompressionabstractVector quantization (VQ) is a source coding methodology with provable rate-distortion optimality. However, despite more than two decades of intensive research, VQ theoretical promise is yet to be fully realized in image compression practice. Restricted by high VQ complexity in dimensions and due to high-order sample correlations in images, block sizes of practical VQ image coders are hardly large enough to achieve the rate-distortion optimality. Among the large number of VQ variants in the literature, a technique called address VQ (A-VQ) by Nasrabadi and Feng (1990) achieved the best rate-distortion performance so far to the best of our knowledge. The essence of A-VQ is to effectively increase VQ dimensions by a lossless coding of a group of 16-dimensional VQ codewords that are spatially adjacent. From a different perspective, we can consider a signal source that is coded by memoryless basic VQ to be just another signal source whose samples are the indices of the memoryless VQ codewords, and then induce the problem of lossless compression of the VQ-coded source. If the memoryless VQ is not rate-distortion optimal (often the case in practice), then there must exist hidden structures between the samples of VQ-coded source (VQ codewords). Therefore, an alternative way of approaching the rate-distortion optimality is to model and utilize these inter-codewords structures or correlations by context modeling and conditional entropy coding of VQ indexes. Xiaolin Wu 0001, Jiang Wen, Wing Hung Wong |
Data Compression Conference | 1 |
| 1997 | Context modeling and entropy coding of wavelet coefficients for image compressionabstractIn this paper we study the problem of context modeling and entropy coding of the symbol streams generated by the well-known EZW image coder (embedded image coding using zerotrees of wavelet coefficients). We present some simple context modeling techniques that can squeeze out more statistical redundancy in the wavelet coefficients of EZW-type image coders and hence lead to improved coding efficiency. Xiaolin Wu 0001 |
ICASSP | 1 |
| 1997 | Non-Embeded Wavelet Image Coding SchemeabstractA novel wavelet image coding scheme is presented. Significant discrete wavelet transformed coefficients are first selected with a threshold. Then, these significant coefficients are quantized with a uniform quantizer and decomposed into two parts: the most significant bit and the residual bits for entropy encoding. Simple but effective context modeling schemes are proposed for better squeezing of the redundancy lying in the significant map symbol stream determined by the threshold operation and the MSB symbol stream decomposed from the quantized significant coefficients. With these innovations, the proposed coding scheme outperforms all zerotree-structured embedded wavelet coding schemes and is competitive with other advanced coding algorithms reported in the literature. Tian-Hu Yu, Xiaolin Wu 0001 |
ICIP (1) | 3 |
| 1997 | VQ index coding for high-fidelity medical image compressionabstractIn order to obtain the high-fidelity medical compressed images, a new compression scheme is proposed. Based on the stringent requirements on lossy medical image compression, we refine the context modeling for a given class of medical images and utilize the conditional entropy coding of the VQ index (CECOVI) scheme to code the MR head images. The experimental results show that the image-type-dependent CECOVI can achieve better rate-distortion performance than the state-of-art wavelet image coder SPIHT. This also implies that incorporating the conditional entropy coding strategy into the VQ process is an appropriate way for high-fidelity medical image compression. Jiang Wen, Xiaolin Wu 0001, Wai-Yin Ng |
ICIP (3) | 2 |
| 1997 | Recent Developments in Context-Based Predictive Techniques for Lossless Image Compression
Nasir Memon, Xiaolin Wu 0001 |
Comput. J. | 2 |
| 1997 | On Minimizing the Lengths of Checking SequencesabstractA general model for constructing minimal length checking sequences employing a distinguishing sequence is proposed. The model is based on characteristics of checking sequences and a set of state recognition sequences. Some existing methods are shown to be special cases of the proposed model and are proven to construct checking sequences. The minimality of the resulting checking sequences is discussed and a heuristic algorithm for the construction of minimal length checking sequences is given. Hasan Ural, Xiaolin Wu 0001, Fan Zhang 0001 |
IEEE Trans. Computers | 2 |
| 1997 | Context-based, adaptive, lossless image codingabstractWe propose a context-based, adaptive, lossless image codec (CALIC). The codec obtains higher lossless compression of continuous-tone images than other lossless image coding techniques in the literature. This high coding efficiency is accomplished with relatively low time and space complexities. The CALIC puts heavy emphasis on image data modeling. A unique feature of the CALIC is the use of a large number of modeling contexts (states) to condition a nonlinear predictor and adapt the predictor to varying source statistics. The nonlinear predictor can correct itself via an error feedback mechanism by learning from its mistakes under a given context in the past. In this learning process, the CALIC estimates only the expectation of prediction errors conditioned on a large number of different contexts rather than estimating a large number of conditional error probabilities. The former estimation technique can afford a large number of modeling contexts without suffering from the context dilution problem of insufficient counting statistics as in the latter approach, nor from excessive memory use. The low time and space complexities are also attributed to efficient techniques for forming and quantizing modeling contexts. Xiaolin Wu 0001, Nasir Memon |
IEEE Trans. Commun. | 1 |
| 1997 | Lossless compression of continuous-tone images via context selection, quantization, and modelingabstractContext modeling is an extensively studied paradigm for lossless compression of continuous-tone images. However, without careful algorithm design, high-order Markovian modeling of continuous-tone images is too expensive in both computational time and space to be practical. Furthermore, the exponential growth of the number of modeling states in the order of a Markov model can quickly lead to the problem of context dilution; that is, an image may not have enough samples for good estimates of conditional probabilities associated with the modeling states. New techniques for context modeling of DPCM errors are introduced that can exploit context-dependent DPCM error structures to the benefit of compression. New algorithmic techniques of forming and quantizing modeling contexts are also developed to alleviate the problem of context dilution and reduce both time and space complexities. By innovative formation, quantization, and use of modeling contexts, the proposed lossless image coder has a highly competitive compression performance and yet remains practical. Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 1997 | Optimal binary vector quantization via enumeration of covering codesabstractBinary vector quantization (BVQ) refers to block coding of binary vectors under a fidelity measure. Covering codes were studied as a means of lattice BVQ. But a further source coding problem hidden in the equivalence of covering codes has seemingly eluded attention. Given a d-dimensional hypercube (code space), equivalent covering codes of the same covering radius but of different codewords have different expected BVQ distortions for a general probability mass function. Thus one can minimize, within the code equivalence, the expected distortion over all different covering codes. This leads a two-stage optimization scheme for BVQ design. First we use an optimal covering code to minimize the maximum per-vector distortion at a given rate. Then under the minmax constraint, we minimize the expected quantization distortion. This minmax constrained BVQ method (MCBVQ) controls both the maximum and average distortions, and hence improves subjective quality of compressed binary images, MCBVQ also avoids poor local minima that may trap the generalized Lloyd method. The [7,4] Hamming code and [8,4] extended Hamming code are found to be particularly suitable for MCBVQ on binary images. An efficient and simple algorithm is introduced to enumerate all distinct [7,4] Hamming/[8,4] extended Hamming codes and compute the corresponding expected distortions in optimal MCBVQ design. Furthermore, MCBVQ using linear covering codes has a compact codebook and a fast syndrome-encoding algorithm. Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 1996 | An Algorithmic Study on Lossless Image CompressionabstractThe author discusses how to reduce model complexity for improving both coding and computational efficiency. Some of the topics discussed are: prediction via context error modeling, context selection for entropy coding, context hierarchy and selection and performance evaluation. Xiaolin Wu 0001 |
Data Compression Conference | 1 |
| 1996 | CALIC-a context based adaptive lossless image codecabstractWe propose a context-based, adaptive, lossless image codec (CALIC). CALIC obtains higher lossless compression of continuous-tone images than other techniques reported in the literature. This high coding efficiency is accomplished with relatively low time and space complexities. CALIC puts heavy emphasis on image data modeling. A unique feature of CALIC is the use of a large number of modeling contexts to condition a non-linear predictor and make it adaptive to varying source statistics. The non-linear predictor adapts via an error feedback mechanism. In this adaptation process, CALIC only estimates the expectation of prediction errors conditioned on a large number of contexts rather than estimating a large number of conditional error probabilities. The former estimation technique can afford a large number of modeling contexts without suffering from the sparse context problem. The low time and space complexities of CALIC are attributed to efficient techniques for forming and quantizing modeling contexts. Xiaolin Wu 0001, Nasir Memon |
ICASSP | 1 |
| 1996 | Fast Algorithms for Minimum Matrix Norm with Application in Computer Graphics
Shouwen Tang, Kaizhong Zhang, Xiaolin Wu 0001 |
Algorithmica | 3 |
| 1996 | YIQ vector quantization in a new color palette architectureabstractMany CRT displays use color palettes to save highspeed frame buffer memory and to support many interactive, real-time graphics and imaging operations. Existing quantization algorithms for color palettes cluster colors as 3-D vectors in a color space. Thus, they fail to remove the statistical redundancy in the image space. This weakness prevents a more efficient use of frame buffer capacity. We propose a simple product vector quantization (VQ) technique for frame buffers that exploits redundancies in both image and color spaces. The new VQ technique can reduce the frame buffer use of a color image by half from the current algorithms with comparable image quality. The NTSC YIQ rather than RGB color space is used in our product VQ scheme. Consequently, we propose to modify the current architecture of the color palette to suit the proposed product VQ algorithm in the YIQ mode. The modified color palette is nearly ten times smaller than the current RGB palette and is more flexible. The improved space efficiency in both frame buffer and color palette is achieved while neither complicating the control or logic of the frame buffer architecture nor increasing computational complexity of color quantization. Xiaolin Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 1995 | Adaptive binary vector quantization using Hamming codesabstractHamming codes are studied as a means of adaptive vector quantization of binary images. The idea is to minimize, within the equivalence class of a Hamming code, the expected quantization distortion, while bounding the maximum distortion per vector to prevent burst quantization errors in a binary image. Some interesting and useful relationships between distinct Hamming codes are presented. These findings can lead to efficient algorithms for designing adaptive binary vector quantizers whose codebooks can adapt to sources of smoothly changing statistics. Xiaolin Wu 0001 |
ICIP (3) | 1 |
| 1995 | A segmentation-based predictive multiresolution image coderabstractRecursive rectilinear tessellations like quadtree are widely used in image coding, but a regular tessellation, despite simple geometry, may not suit image compression because it is too rigid to reflect the scene structure of an image. The paper presents a new image pyramid formed by adaptive, tree-structured segmentation to be a framework of a predictive multiresolution image coder. Subjectively appealing compression results are obtained at different resolutions by scene-adaptive, tree-structured segmentation and by exploiting the statistical dependency between the layers of the image pyramid. The adaptive segmentation-based image coder is constructed by recursive, least-squares piecewise functional approximation. The seemingly expensive encoding process can be made efficient by an incremental least-squares computation technique. The decoding is simple and can be done in real time if assisted by existing hardware technology. Xiaolin Wu 0001, Yonggang Fang |
IEEE Trans. Image Process. | 1 |
| 1994 | Matching with Matrix Norm Minimization
Shouwen Tang, Kaizhong Zhang, Xiaolin Wu 0001 |
CPM | 3 |
| 1994 | A Subjective Distortion Measure for Vector QuantizationabstractThe authors present some preliminary results of their ongoing study on subjective VQ distortion measure in the time/spatial domain. They first propose a context based distortion measure between two vectors. The new measure is intuitively appealing, and they include some empirical evidence for its subjective significance. Although the measure is formulated as a matrix norm, it is computationally no more difficult than the mean-squares error. This measure quantifies the quantization distortion in the context (shape) of the signal waveform, but it is amplitude-invariant. So they combine the context distortion measure with a weighted mean distortion measure to obtain a unified subjective distortion measure D. They show that D is a distance measure and can be easily computed. Moreover, the process of computing the centroid of a set of training vectors and designing the VQ codebook under the new subjective distortion measure D is as simple as the conventional VQ. Specifically, the LBG algorithm can be applied to design the subjective VQ codebook after a simple linear transformation of the vector space in which signal samples are originally taken. They also analytically relate their subjective distortion measure to the ubiquitous mean-squares measure, and demonstrate that the latter is only a special case of the former. They also observe that the mean-removed VQ in a sense clusters training vectors under the proposed context distortion measure.> Xiaolin Wu 0001, Kaizhong Zhang |
Data Compression Conference | 1 |
| 1994 | Progressive Image Coding by Hierarchical Linear Approximation
Xiaolin Wu 0001, Yonggang Fang |
Inf. Process. Manag. | 1 |
| 1994 | Acceleration of the LBG algorithmabstractA concentric spherical search technique is proposed to speed up the clustering process in VQ design. A linear data structure is incorporated into the LBG algorithm to keep and update the information about the proximity among the codewords. This proximity information can significantly reduce the number of candidate codewords to be the closest to a given training vector. An improved k-means type VQ design algorithm is proposed based on the new search technique and the supporting data structure. The new algorithm is simple to implement, valid for general error metric, and demonstrated by present experiments to be considerably faster than previous algorithms.> Xiaolin Wu 0001, Lian Guan |
IEEE Trans. Commun. | 1 |
| 1993 | Adaptive Split-and-Merge Segmentation Based on Piecewise Least-Square ApproximationabstractThe performance of the classic split-and-merge segmentation algorithm is severely hampered by its rigid split-and-merge processes, which are insensitive to the image semantics. The author proposes efficient algorithms and data structures to optimize the split-and-merge processes by piecewise least-square approximation of image intensity functions. This optimization aims at the unification of segment finding and edge detection. The optimized split-and-merge algorithm is shown to be adaptive to the image semantics and, hence, improves the segmentation validity of the previous algorithms. This algorithm also appears to work well on noisy sources. Since the optimization is done within the split-and-merge framework, the better segmentation performance is achieved at the same order of time complexity as the previous algorithms.> Xiaolin Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1993 | Quantizer monotonicities and globally optimal scalar quantizer designabstractNew monotonicity properties of optimal scalar quantizers are discussed. These monotonicities reveal a globally optimal scalar quantizer structure depending on the probability mass functions and on the number of quantizer levels. By incorporating the monotone quantizer structure into a dynamic programming process, the time complexities of previous algorithms for designing globally optimal scalar quantizers can be significantly reduced for very general classes of distortion measures.> Xiaolin Wu 0001, Kaizhong Zhang |
IEEE Trans. Inf. Theory | 1 |
| 1993 | Visual coding by optimal graph-coloring
Xiaolin Wu 0001, Yonggang Fang |
Vis. Comput. | 1 |
| 1992 | Vector Quantizer Design by Constrained Global OptimizationabstractCentral to vector quantization is the design of optimal code book. The construction of a globally optimal code book has been shown to be NP-complete. However, if the partition halfplanes are restricted to be orthogonal to the principal direction of the training vectors, then the globally optimal K-partition of a set of N D-dimensional data points can be computed in O((N+KM/sup 2/)D) time by dynamic programming, where M is the intensity resolution. This constrained optimization strategy improves the performance of vector quantizer over the classic LBG algorithm and the popular methods of tree-structured recursive greedy bipartition of the training data set.> Xiaolin Wu 0001 |
Data Compression Conference | 1 |
| 1992 | On convergence of Lloyd's method IabstractAlthough Lloyd's method I for optimal quantization was proposed more than thirty years ago and has been frequently referred to in the literature, its convergence has so far not been shown. This correspondence proves that Lloyd's method I converges for a large class of error measures, if the density function is continuous, positive, and defined on a finite interval. The proof is done by modeling the behavior of a continuous optimization algorithm by a finite state machine.> Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 1992 | Image coding by adaptive tree-structured segmentationabstractA new algorithmic approach to segmentation-based image coding is proposed. A good compromise is achieved between segmentation by quadtree-based decomposition and by free region-growing in terms of time complexity and scene adaptability. Encoding is to recursively partition an image into convex n-gons, 3> Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 1992 | Color Quantization by Dynamic Programming and Principal AnalysisabstractColor quantization is a process of choosing a set of K representative colors to approximate the N colors of an image, K < N , such that the resulting K -color image looks as much like the original N -color image as possible. This is an optimization problem known to be NP-complete in K . However, this paper shows that by ordering the N colors along their principal axis and partitioning the color space with respect to this ordering, the resulting constrained optimization problem can be solved in O ( N + KM 2 ) time by dynamic programming (where M is the intensity resolution of the device). Traditional color quantization algorithms recursively bipartition the color space. By using the above dynamic-programming algorithm, we can construct a globally optimal K -partition, K >2, of a color space in the principal direction of the input data. This new partitioning strategy leads to smaller quantization error and hence better image quality. Other algorithmic issues in color quantization such as efficient statistical computations and nearest-neighbor searching are also studied. The interplay between luminance and chromaticity in color quantization with and without color dithering is investigated. Our color quantization method allows the user to choose a balance between the image smoothness and hue accuracy for a given K . Xiaolin Wu 0001 |
ACM Trans. Graph. | 1 |
| 1991 | Image Coding by Adaptive Tree-Structured SegmentationabstractA new segmentation-based image coding method is proposed. The encoder recursively partitions an image into convex n-gons, 3> Xiaolin Wu 0001, C. Yao |
Data Compression Conference | 1 |
| 1991 | A Better Tree-Structured Vector QuantizerabstractA new vector quantizer permits logarithmic-time encoding and yet performs better than the locally optimal quantizers generated by the LBG algorithm. The success is credited to an elaborated tree-structured optimization process in the codebook design.> Xiaolin Wu 0001, Kaizhong Zhang |
Data Compression Conference | 1 |
| 1991 | An efficient antialiasing techniqueabstractAn intuitive concept of antialiasing is developed into very efficient antialiased line and circle generators that require even less amount of integer arithmetic than Bresenham's line and circle algorithms. Unlike its predecessors, the new antialiasing technique is derived in spatial domain (raster plane) under a subjectively meaningful error measure to preserve the dynamics of curve and object boundaries. A formal analysis of the new antialiasing technique in frequency domain is also conducted. It is shown that our antialiasing technique computes the same antialiased images as Fujimoto-Iwata's algorithm but at a fraction of the latter's computational cost. The simplicities of the new antialiased line and circle generators also mean their easy hardware implementations. Xiaolin Wu 0001 |
SIGGRAPH | 1 |
| 1991 | Optimal bi-level quantization and its application to multilevel quantizationabstractA special form of optimal quantization, optimal bilevel quantization, is studied. A fixed-point method is embedded in a search scheme to find all the locally optimal bilevel quantizers, resulting in an algorithm for computing the globally optimal bilevel quantizer. Some interesting relations between the optimal bilevel quantizer, and the mean and the median of the density function p(x) are explored. Efficient algorithms for computing optimal bilevel quantizers are proposed. The application of these results to optimal quantization in general is discussed.> Xiaolin Wu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 1990 | A tree-structured locally optimal vector quantizerabstractA tree-structured VQ (vector quantizer) that performs the nearest-neighbor encoding based on a locally optimal codebook generated by the algorithm of Y. Linde, A. Buzo and R.M. Gray (1980) is proposed. A design method is given to organize the code words by a quasi-voronoi tree. This tree structure allows the nearest-neighbor encoding without an exhaustive search. For a codebook of size K, encoding an input vector takes an expected number of O(log K) distortion evaluations for dimensionalities below eight; that time complexity is O(K/sup 1/2/) in practice for higher dimensionalities. The tree-structured VQ achieves a good compromise between optimality and encoding speed.> Xiaolin Wu 0001 |
ICPR (2) | 1 |
| 1990 | Fast line scan-conversionabstractA major bottleneck in many graphics displays is the time required to scan-convert straight line segments. Most manufacturers use hardware based on Bresenham's [5] line algorithm. In this paper an algorithm is developed based on the original Bresenham scan-conversion together with the symmetry first noted by Gardner [18] and a recent double-step technique [31]. This results in a speed-up of scan-conversion by a factor of approximately 4 as compared to the original Bresenham algorithm. Hardware implementations are simple and efficient since the property of using only shift and increment operations is preserved. Jon G. Rokne, Brian Wyvill, Xiaolin Wu 0001 |
ACM Trans. Graph. | 3 |
| 1989 | On Properties of Discretized Convex CurvesabstractThe connection between a continuous convex curve and its discrete image is investigated using an appropriate definition of discrete convexity. It is shown that the discrete image of continuous convex curve may not be convex; however, its deviation from convexity (in the discrete definition) is bounded by a small constant. The actual pixel patterns that are obtained by discretizing convex curves are studied. Certain constraints on the context of discrete images of continuous convex curves were discovered.> Xiaolin Wu 0001, Jon G. Rokne |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1987 | Double-step incremental generation of lines and circles
Xiaolin Wu 0001, Jon G. Rokne |
Comput. Vis. Graph. Image Process. | 1 |