EDBT 2026 Demo / reviewers in the wild / expert
Pamela C. Cosman
dblp:07/6197
· DBLP profile ↗
163ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-4012-0176ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 117 · 6 first-author · 8 since 2021Computer networks · 35 · 2 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-authorTheory of computation · 3Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh CompressionabstractThe compression of real-world scanned 3D human dynamic meshes is an emerging research area, driven by applications such as telepresence, virtual reality, and 3D digital streaming. Unlike synthesized dynamic meshes with fixed topology, scanned dynamic meshes often not only have varying topology across frames but also scan defects such as holes and outliers, increasing the complexity of prediction and compression. Additionally, human meshes often combine rigid and non-rigid motions, making accurate prediction and encoding significantly more difficult compared to objects that exhibit purely rigid motion. To address these challenges, we propose a compression method designed for real-world scanned human dynamic meshes, leveraging embedded key nodes. The temporal motion of each vertex is formulated as a distance-weighted combination of transformations from neighboring key nodes, requiring the transmission of solely the key nodes' transformations. To enhance the quality of the KeyNode-driven prediction, we introduce an octree-based residual coding scheme and a Dual-direction prediction mode, which uses I-frames from both directions. Extensive experiments demonstrate that our method achieves significant improvements over the state-of-the-art, with an average bitrate savings of 58.43% across the evaluated sequences, particularly excelling at low bitrates. Huong Hoang, Truong Q. Nguyen, Pamela C. Cosman |
IEEE Trans. Multim. | 3 |
| 2025 | Efficient Progressive Image Compression with Variance-Aware MaskingabstractLearned progressive image compression is gaining momentum as it allows improved image reconstruction as more bits are decoded at the receiver. We propose a progressive image compression method in which an image is first represented as a pair of base-quality and top-quality latent representations. Next, a residual latent representation is encoded as the element-wise difference between the top and base representations. Our scheme enables progressive image compression with element-wise granularity by introducing a masking system that ranks each element of the residual latent representation from most to least important, dividing it into complementary components, which can be transmitted separately to the decoder in order to obtain different reconstruction quality. The masking system does not add further parameters or complexity. At the receiver, any elements of the top latent representation excluded from the transmitted components can be independently replaced with the mean predicted by the hyperprior architecture, ensuring reliable reconstructions at any intermediate quality level. We also in-troduced Rate Enhancement Modules (REMs), which refine the estimation of entropy parameters using already decoded components. We obtain results competitive with state-of-the-art competitors, while significantly reducing computational complexity, decoding time, and number of parameters. Alberto Presta, Enzo Tartaglione, Attilio Fiandrotti, Marco Grangetto, Pamela C. Cosman |
WACV | 5 |
| 2025 | Slippage-robust linear features for eye trackingabstractRegression-based gaze estimation for wearable eye-tracking headsets inevitably suffers from headset slippage. Despite good calibration accuracy (< 2°), over time, slippage can cause the gaze estimation accuracy to deteriorate significantly, ranging from 2° to 40°. Existing corrective measures typically cater to large viewing depths (> 1 meter). However, these measures are unsuitable for lower depth, such as industrial or office settings, where objects of interest lie within arm’s reach. For a dataset with natural slippage, we propose a set of slippage-robust features along with a linear regression function to improve gaze estimation accuracy. We also propose metrics that are more informative about gaze estimation accuracy over the entire field of view. Our findings demonstrate that the proposed features improve the gaze estimation accuracy by nearly 30%. These contributions pave the way for improved performance and reliability of eye-tracking headsets, enabling their use in diverse research domains and real-world applications. Tawaana Gustad Homavazir, V. S. Raghu Parupudi, Surya L. S. R. Pilla, Pamela C. Cosman |
Expert Syst. Appl. | 4 |
| 2024 | Performance Analysis for Underwater Video Transmission with Imperfect ResamplingabstractLarge Doppler shifts are a challenge in underwater wireless acoustic communication compared to terrestrial wireless radio communication. Resampling and Doppler shift compensation are used at the receiver to counteract the effect of the Doppler shift introduced during transmission. Imperfect resampling can degrade system performance. While the conventional choice of Doppler compensation factor minimizes the intercarrier interference (ICI) around the central subcarrier, we show that using a Doppler compensation factor that treats lower subcarriers preferentially in terms of ICI can significantly improve the system performance in the case of imperfect resampling. Two ad-hoc schemes are proposed in the paper to find suitable Doppler compensation factors for a given Doppler factor and percentage of error in the resampling factor. The knowledge of these parameters facilitates the performance analysis of the schemes, rather than focusing on design aspects. The proposed schemes are tested for various relative vehicular speeds. In addition, different Doppler factors for the paths are considered. In all cases, our proposed schemes achieve higher video peak-signal-to-noise ratio compared to conventional schemes. Rana D. Hegazy, Laurence B. Milstein, Pamela C. Cosman |
ICCCN | 3 |
| 2024 | PriFU: Capturing Task-Relevant Information Without Adversarial LearningabstractAs machine learning advances, machine learning as a service (MLaaS) in the cloud brings convenience to human lives but also privacy risks, as powerful neural networks used for generation, classification or other tasks can also become privacy snoopers. This motivates privacy preservation in the inference phase. Many approaches for preserving privacy in the inference phase introduce multi-objective functions, training models to remove specific private information from users' uploaded data. Although effective, these adversarial learning-based approaches suffer not only from convergence difficulties, but also from limited generalization beyond the specific privacy for which they are trained. To address these issues, we propose a method for privacy preservation in the inference phase by removing task-irrelevant information, which requires no knowledge of the privacy attacks nor introduction of adversarial learning. Specifically, we introduce a metric to distinguish task-irrelevant information from task-relevant information, and achieve more efficient metric estimation to remove task-irrelevant features. The experiments demonstrate the potential of our method in several tasks. Our code will be available at: https://github.com/iwhoyoung/PriFU. Xiuli Bi, Bo Liu 0047, Weisheng Li 0001, Pamela C. Cosman, Bin Xiao 0002 |
ACM Multimedia | 5 |
| 2023 | End-to-End Blind Video Quality Assessment Based on Visual and Memory Attention ModelingabstractDeveloping an objective quality assessment model for user-generated content (UGC) videos is significant for multimedia applications, and also a challenge due to the diversity of video content and unpredictability of distortions. To predict the perceived quality, it is necessary to consider the human visual system, in which attention in visual and memory domains is an essential component. With the idea that the stimulus-driven bottom-up mechanism and cognition-driven top-down mechanism work in synergy to generate quality-aware attention, we propose an end-to-end blind video quality assessment (VQA) algorithm based on visual and memory attention modeling. First, a quality-aware visual attention module is established to obtain spatial-temporal attention-guided representations for frame-level quality perception. Specifically, an attention selection and confluence method is developed by circularly integrating the quality-aware attention information to spatial-temporal content features. Then, with the aid of a quality-aware memory attention module, the video-level attention-guided features are inferred through the dimension and attention reshaping of frame-level representations. The video quality is predicted with the guidance of frame-level visual attention and video-level memory attention in an end-to-end structure. Experimental results on five UGC-VQA databases (CVD2014, LIVE-Qualcomm, KoNViD-1 k, LIVE-VQC and Youtube-UGC) demonstrate the effectiveness of our modules. Xiaodi Guan, Fan Li 0003, Yangfan Zhang, Pamela C. Cosman |
IEEE Trans. Multim. | 4 |
| 2022 | Human-Machine Interaction-Oriented Image Coding for Resource-Constrained Visual Monitoring in IoTabstractVisual monitoring supported by the Internet of Things (IoT) increasingly relies on analyzing a mass of image data with human–machine interactive mechanisms. However, maintaining the efficiency of such a monitoring system in complex environments with energy or bandwidth constraints poses challenges. While both human perception and machine analysis performance should be satisfied, transmitting extra information to satisfy both should be avoided for efficient resource utilization. To this end, we propose a human–machine interaction-oriented image coding (HMI-IC) framework based on deep learning. In this framework, machines should provide early monitoring messages consisting of analysis results and preview images, and then humans can additionally request high-quality images of objects of interest. Each collected image is compressed into a layered data stream by HMI-IC to fulfill the demands of analysis, preview visualization, and high-quality reconstruction. Adaptive coding transmission can fit different demands in two stages, according to resource constraints. Experimental results show that both accuracy and inference speed on compressed images are improved by our method, with entire coding efficiency comparable to JPEG2000. To validate HMI-IC’s efficiency in practical terms, we provide two use cases (energy constrained and bandwidth constrained) for visual monitoring. Fan Li 0003, Jing Xu 0003, Pamela C. Cosman |
IEEE Internet Things J. | 4 |
| 2022 | Tile-Based Wireless Streaming of 360-Degree Video With Rate Adaptation Using Viewport EstimationabstractThis paper presents tile-based wireless streaming of 360-degree videos with rate adaptation using viewport estimation. We propose a probabilistic model for viewport location to design importance weights of tiles. Based on the tile weights, we improve the streaming performance by dynamically adjusting quantization parameters and forward error correction rates across tiles. Yeohee Im, Tian Qiu 0001, Laurence B. Milstein, Pamela C. Cosman |
IEEE Signal Process. Lett. | 4 |
| 2022 | Learning-Based Rate Control for Video-Based Point Cloud CompressionabstractDue to limited transmission resources and storage capacity, efficient rate control is important in Video-based Point Cloud Compression (V-PCC). In this paper, we propose a learning-based rate control method to improve the rate-distortion (RD) performance of V-PCC. A low-latency synchronous rate control structure is designed to reduce the overhead of pre-coding. The basic unit (BU) parameters are predicted accurately based on our proposed CNN-LSTM neural network, instead of the online updating approach, which can be inaccurate due to low consistency between adjacent 2D frames in V-PCC. When determining the quantization parameters for the BU, a patch-based clipping method is proposed to avoid unnecessary clipping. This approach is able to improve the RD performance and subjective dynamic point cloud quality. Experiments show that our proposed rate control method outperforms present approaches. Taiyu Wang, Fan Li 0003, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2021 | MMMNet: An End-to-End Multi-Task Deep Convolution Neural Network With Multi-Scale and Multi-Hierarchy Fusion for Blind Image Quality AssessmentabstractAs the evaluation of image quality depends on the human visual system (HVS), many existing image quality assessment (IQA) methods focus on modeling the HVS to account for subjective perception. The visual attention of the HVS makes humans more sensitive to distortion on the attended regions than on regions which are not the focus of attention. Therefore, we propose an end-to-end multi-task deep convolution neural network with multi-scale and multi-hierarchy fusion (MMMNet), in which the IQA and saliency subtasks are jointly optimized to improve saliency-guided IQA performance. Particularly, the incorporation of saliency information is achieved by fusing saliency features with IQA features hierarchically to progressively improve the IQA features over network depth. A multi-scale feature extraction module (MSFE) is proposed to provide effective saliency features for the IQA network. Based on the saliency fusion, MMMNet introduces an auxiliary saliency task, achieving the multi-task learning to improve the generalization of the IQA task. Experimental results show that MMMNet achieves state-of-the-art performance and strong generalization ability on IQA databases. Fan Li 0003, Yangfan Zhang, Pamela C. Cosman |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Low-Complexity Error Resilient HEVC Video Coding: A Deep Learning ApproachabstractIntra/inter switching-based error resilient video coding effectively enhances the robustness of video streaming when transmitting over error-prone networks. But it has a high computation complexity, due to the detailed end-to-end distortion prediction and brute-force search for rate-distortion optimization. In this article, a Low Complexity Mode Switching based Error Resilient Encoding (LC-MSERE) method is proposed to reduce the complexity of the encoder through a deep learning approach. By designing and training multi-scale information fusion-based convolutional neural networks (CNN), intra and inter mode coding unit (CU) partitions can be predicted by the networks rapidly and accurately, instead of using brute-force search and a large number of end-to-end distortion estimations. In the intra CU partition prediction, we propose a spatial multi-scale information fusion based CNN (SMIF-Intra). In this network a shortcut convolution architecture is designed to learn the multi-scale and multi-grained image information, which is correlated with the CU partition. In the inter CU partition, we propose a spatial-temporal multi-scale information fusion-based CNN (STMIF-Inter), in which a two-stream convolution architecture is designed to learn the spatial-temporal image texture and the distortion propagation among frames. With information from the image, and coding and transmission parameters, the networks are able to accurately predict CU partitions for both intra and inter coding tree units (CTUs). Experiments show that our approach significantly reduces computation time for error resilient video encoding with acceptable quality decrement. Taiyu Wang, Fan Li 0003, Xiaoya Qiao, Pamela C. Cosman |
IEEE Trans. Image Process. | 4 |
| 2020 | Analyzing Gaze Behavior Using Object Detection and Unsupervised ClusteringabstractGaze behavior is important in early development, and atypical gaze behavior is among the first symptoms of autism. Here we describe a system that quantitatively assesses gaze behavior using eye-tracking glasses. Objects in the subject’s field of view are detected using a deep learning model on the video captured by the glasses’ world-view camera, and a stationary frame of reference is estimated using the positions of the detected objects. The gaze positions relative to the new frame of reference are subjected to unsupervised clustering to obtain the time sequence of looks. The clustering method increases the accuracy of look detection on test videos compared against a previous algorithm, and is considerably more robust on videos with poor calibration. Pranav Venuprasad, Enoch Huang, Andrew Gilman, Leanne Chukoskie, Pamela C. Cosman |
ETRA | 6 |
| 2020 | Efficient Debanding Filtering for Inverse Tone Mapped High Dynamic Range VideosabstractWhen a low or standard dynamic range video is inverse tone mapped to high dynamic range (HDR), there can be banding artifacts in the output HDR video. We design a highly integrated architecture that can detect and alleviate banding artifacts while preserving true edges and details in an extremely efficient way. This is achieved by a 7-tap edge-aware selective sparse filter. Coding artifacts, such as blocky artifacts, can also be reduced by this filter. The filter includes some parameters that depend on the strengths of the banding artifacts. A parameter selection mechanism is presented which considers smoothness of the banding regions and fidelity of the filtering output. The filter yields significant PSNR gain at the regions of artifacts. Subjective tests demonstrate the great quality improvement achieved by the proposed filter, compared to the quality before filtering. The visual quality provided by the filter is better than or similar to that of algorithms which are far more complex. Qing Song 0005, Guan-Ming Su, Pamela C. Cosman |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Characterizing joint attention behavior during real world interactions using automated object and gaze detectionabstractJoint attention is an essential part of the development process of children, and impairments in joint attention are considered as one of the first symptoms of autism. In this paper, we develop a novel technique to characterize joint attention in real time, by studying the interaction of two human subjects with each other and with multiple objects present in the room. This is done by capturing the subjects' gaze through eye-tracking glasses and detecting their looks on predefined indicator objects. A deep learning network is trained and deployed to detect the objects in the field of vision of the subject by processing the video feed of the world view camera mounted on the eye-tracking glasses. The looking patterns of the subjects are determined and a real-time audio response is provided when a joint attention is detected, i.e., when their looks coincide. Our findings suggest a trade-off between the accuracy measure (Look Positive Predictive Value) and the latency of joint look detection for various system parameters. For more accurate joint look detection, the system has higher latency, and for faster detection, the detection accuracy goes down. Pranav Venuprasad, Tushar Dobhal, Anurag Paul, Tu N. M. Nguyen, Andrew Gilman, Pamela C. Cosman, Leanne Chukoskie |
ETRA | 6 |
| 2019 | Training Efficient Saliency Prediction Models with Knowledge DistillationabstractRecently, deep learning-based saliency prediction methods have achieved significant accuracy improvements. However, they are hard to embed in practical multimedia applications due to large memory consumption and running time caused by complicated architectures. In addition, most methods are fine-tuned from pre-trained models for classification tasks, and networks cannot flexibly be transferred for a new task. In this paper, a condensed and randomly initialized student network is employed to achieve higher efficiency by transferring knowledge from complicated and well-trained teacher networks. This is the first use of knowledge distillation for efficient pixel-wise saliency prediction. Instead of directly minimizing Euclidean distance between feature maps, we propose two statistical representations of feature maps (i.e., first-order and second-order statistics) as knowledge. We conduct experiments on three kinds of teacher networks and four benchmark datasets to verify the effectiveness of the proposed method. Compared with the teacher networks, the student networks achieve an acceleration ratio of 4.56-4.73. Compared with state-of-the-art approaches, the proposed model achieves competitive accuracy with faster running speed (up to 4.38 times) and smaller model size (up to 93.27% reduction). We further embedded the proposed saliency prediction model into a video captioning application. The saliency-embedded approaches improve video captioning on all test metrics with a small complexity cost. The student-model embedded approach achieves 25% time saving with similar performance to the teacher embedded one. Peng Zhang 0024, Li Su 0003, Liang Li 0003, Bing-Kun Bao, Pamela C. Cosman, Guorong Li, Qingming Huang |
ACM Multimedia | 5 |
| 2019 | No-Reference Video Quality Assessment Based on Ensemble of Knowledge and Data-Driven Models
Li Su 0003, Pamela C. Cosman, Qihang Peng |
MMM (2) | 2 |
| 2019 | Joint rate adaptation and resource allocation for real-time H.265/HEVC video transmission over uplink OFDMA systems
Fan Li 0003, Taiyu Wang, Pamela C. Cosman |
Multim. Tools Appl. | 3 |
| 2019 | Optimal Sensing Disruption: A Generalized Framework for a Power-Limited AdversaryabstractA generalized framework of spectrum sensing disruption for a power-limited adversary is proposed in this paper. In the literature, a conventional sensing attack typically assumes that the adversary has perfect knowledge of the spectral usage status. The framework in this paper considers a more general case where there are uncertainties in the estimates at the adversary. These uncertainties are modeled utilizing the probability of detection and the probability of false alarm. Then, the sum of the conditional probabilities of false detection at the secondary within the spectral range of interest, conditioned on the adversary's estimated spectrum usage status, is maximized. It is shown that the optimal sensing attack, given perfect estimation is a special case of the proposed framework. When the adversary has perfect spectrum usage information, this framework reduces to a previously demonstrated optimal sensing disruption. When the adversary has imperfect information on the spectral status, the proposed framework is significantly more robust than conventional sensing attacks. Further, when the adversary's power budget increases, it asymptotically approaches the sensing disruption performance upper bound. Qihang Peng, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2019 | Multicarrier DS-CDMA System Under Fast Rician Fading and Partial-Time Partial-Band JammingabstractThe impact of joint partial-time, partial-band jamming on a multicarrier (MC) asynchronous direct-sequence code-division multiple access (DS-CDMA) system in a fast fading environment is studied in conjunction with two different subcarrier combining and decoding schemes. An easy-to-evaluate upper bound using the Chernoff bound is provided and compared to simulation results. Simulation results suggest that for soft-decision decoding systems, under Rayleigh fading, full-time, full-band jamming is most effective. In contrast to the Rayleigh case, when a sufficiently strong line-of-sight component exists in the channel, the jammer's optimal strategy of attacking in time or frequency depends on the strength and the type of error correction that the system is deploying for that dimension. For hard-decision decoding systems, in Rayleigh fading, partial-band jamming is recommended. For the coding strategies examined, in Rician fading, the jammer should switch from full-time, partial-band jamming to a strategy that jams a higher percentage of the more heavily protected dimension as the jamming power increases. Furthermore, for AWGN channels, the results from the system show that the jammer should always jam a higher percentage of the more heavily protected dimension. Kanke Wu, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2018 | Scene-Aware Soccer Video QoE Assessment - A Compressed-Domain ApproachabstractThe small screen of mobile devices and bandwidth limitations of communication networks greatly affect users' quality of experience (QoE), especially for soccer video, which is characterized by rapid movement and small objects. In this paper, a Compressed-domain Soccer Video Quality assessment Model (CSVQM) is proposed based on the fact that soccer video includes three distinct scene types, which cause different concerns for viewers. To reduce complexity and operate in real-time, all model parameters are derived from the compressed video stream without resorting to complete video decoding. The validation shows that CSVQM significantly outperforms conventional models in terms of accuracy, consistency, and complexity. Fan Li 0003, Yixin Mei, Pamela C. Cosman |
ICME | 4 |
| 2018 | Joint Energy Optimization of Video Encoding and TransmissionabstractDisposable wireless video sensors have many potential applications but are subject to stringent energy constraints. We studied the minimization of end-to-end distortion under an total energy constraint, by means of optimizing FEC code rate, number of source bits, and energy allocation between video encoding and wireless transmission. A two-step approach is employed. First, the FEC rate is optimized by exhaustive search. Then a binary-search-based algorithm is proposed to optimize the energy allocation and number of source bits. Experiments show that the algorithm achieves a PSNR gain up to 1dB over some reasonable baselines. A simpler suboptimal algorithm is also tested and exhibits similar performance. Ziyu Ye, Rana D. Hegazy, Pamela C. Cosman, Laurence B. Milstein |
PCS | 4 |
| 2018 | High-Speed Railway Fastener Detection Based on a Line Local Binary PatternabstractTraditional image features are not able to effectively represent railway fasteners under varied illumination and conditions. We propose the line local binary pattern encoding method that considers the relationship between the center point and its upper and lower neighborhoods. The method can effectively represent the key components of fasteners. In comparison with several state-of-the-art methods, the proposed method has good performance on detecting the completely missing and partly missing fasteners on real data sets, especially when the illumination and background are not ideal. Pamela C. Cosman, Bailin Li |
IEEE Signal Process. Lett. | 2 |
| 2018 | Cross-Layer Resource Allocation Using Video Slice Header Information for Wireless Transmission Over LTEabstractIn this paper, cross-layer resource allocation methods for wireless video transmission are proposed. We propose two practical metrics to measure the relative importance of each forward error correction (FEC) block to the user's quality of experience. They can be obtained from the header information, and require no overhead bits or amendment of the header format of the existing standard video codec. We discuss adaptive modulation and coding scheme allocation, and adaptive energy allocation for the FEC blocks of a group of pictures. We also provide low-complexity resource allocation solutions. In both approaches, we consider two types of slice packetization, which affect the priority of FEC blocks. The proposed cross-layer resource allocation methods have significant performance gain over equal error protection, and similar performance to more complicated algorithms. Young-Ho Jung, Qing Song 0005, Kyung-Ho Kim, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Generalization of the Dark Channel Prior for Single Image RestorationabstractImages degraded by light scattering and absorption, such as hazy, sandstorm, and underwater images, often suffer color distortion and low contrast because of light traveling through turbid media. In order to enhance and restore such images, we first estimate ambient light using the depth-dependent color change. Then, via calculating the difference between the observed intensity and the ambient light, which we call the scene ambient light differential, scene transmission can be estimated. Additionally, adaptive color correction is incorporated into the image formation model (IFM) for removing color casts while restoring contrast. Experimental results on various degraded images demonstrate the new method outperforms other IFM-based methods subjectively and objectively. Our approach can be interpreted as a generalization of the common dark channel prior (DCP) approach to image restoration, and our method reduces to several DCP variants for different special cases of ambient lighting and turbid medium conditions. Yan-Tsung Peng, Keming Cao, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2018 | Luminance Enhancement and Detail Preservation of Images and Videos Adapted to Ambient IlluminationabstractWhen images and videos are displayed on a mobile device in bright ambient illumination, fewer details can be perceived than in the dark. The detail loss in dark areas of the images/videos is usually more severe. The reflected ambient light and the reduced sensitivity of viewer's eyes are the major factors. We propose two tone mapping operators to enhance the contrast and details in images/videos. One is content independent and thus can be applied to any image/video for the given device and the given ambient illumination. The other tone mapping operator uses simple statistics of the content. Display contrast and human visual adaptation are considered to construct the tone mapping operators. Both operators can be solved efficiently. Subjective tests and objective measurement show the improved quality achieved by the proposed methods. Qing Song 0005, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2017 | Sensing Disruption with Estimation UncertaintyabstractIn this paper, the sensing disruption for a power limited adversary with estimation uncertainty is formulated and analyzed. The estimation uncertainty of the adversary is modeled in terms of its probability of false alarm in vacant bands and the probability of detection in busy bands. The strategy for the adversary is obtained by maximizing the sum of the conditional probabilities of false detection within the spectral range of interest, conditioned on the adversary's estimated spectrum usage status. The proposed algorithm is shown to be significantly more robust than conventional algorithms. It is shown in simulation results that, as the adversary's power budget increases, the proposed algorithm asymptotically approaches the performance upper bound when the adversary has perfect information on the spectrum usage status. Qihang Peng, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 2 |
| 2017 | Sparsity regularized Principal Component PursuitabstractWe study the problem of low-rank and sparse decomposition from possibly noisy observations. We propose a novel objective function with nuclear norm on the low-rank term and ℓ0-`norm' on the sparse term, as well as ℓ1-norm on the additive noise term. When there is no dense inlier noise, the proposed method shares the same theoretical guarantee as the Principal Component Pursuit (PCP), i.e., it can recover the low-rank component and sparse component exactly with high probability. Simulations in the noisy case demonstrate that the proposed method outperforms existing state-of-the-art methods. Results on a surveillance video application further verify the effectiveness of the proposed method. Jing Liu 0009, Pamela C. Cosman, Bhaskar D. Rao |
ICASSP | 2 |
| 2017 | Resource Allocation for Multicarrier Device-to-Device Video Transmission: Symbol Error Rate Analysis and Algorithm DesignabstractIn resource allocation for a device-to-device (D2D) video transmission system, the performance improvement by applying the exact symbol error rate (SER) is compared with the conventional signal-to-interference-plus-noise-ratio-based SER evaluation method that uses a Gaussian approximation (GA) for the aggregated interference. An analytical SER expression for a D2D system using multicarrier bandlimited QAM is derived and then used in the resource allocation algorithm. We consider centralized resource allocation for the D2D system, given knowledge of the channel state information and the rate distortion information of the video streams, and propose an iterative algorithm for subcarrier assignment and power allocation. Bit-level simulations for different numbers of D2D pairs demonstrate a considerable improvement on user capacity and video peak signal-to-noise ratio by incorporating the proposed SER expression compared with the GA. By invoking the conditions under which the central limit theorem holds, and comparing these conditions with the number of interferers and the power ratio of the dominant interferer in the simulated D2D system, we also study why the GA for the interference degrades performance. Peizhi Wu, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2017 | Underwater Image Restoration Based on Image Blurriness and Light AbsorptionabstractUnderwater images often suffer from color distortion and low contrast, because light is scattered and absorbed when traveling through water. Such images with different color tones can be shot in various lighting conditions, making restoration and enhancement difficult. We propose a depth estimation method for underwater scenes based on image blurriness and light absorption, which can be used in the image formation model (IFM) to restore and enhance underwater images. Previous IFM-based image restoration methods estimate scene depth based on the dark channel prior or the maximum intensity prior. These are frequently invalidated by the lighting conditions in underwater images, leading to poor restoration results. The proposed method estimates underwater scene depth more accurately. Experimental results on restoring real and synthesized underwater images demonstrate that the proposed method outperforms other IFM-based underwater image restoration methods. Yan-Tsung Peng, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2017 | Subcarrier Assignment and Power Allocation for Device-to-Device Video Transmission in Rayleigh Fading ChannelsabstractSubcarrier assignment and power allocation for device-to-device (D2D) video transmission using a filter bank multicarrier waveform in a Rayleigh fading environment are investigated. We analyze the co-channel interference between D2D pairs, and propose a cross-layer algorithm with a subcarrier assignment outer loop and a power allocation inner loop, which aims to optimize the overall video quality. Unlike the non-convexity in physical layer power allocation for maximizing the total throughput, the cross-layer power allocation problem is convex under certain conditions, so a high quality solution for power allocation can be efficiently found. Simulation results demonstrate a higher overall video quality by the proposed cross-layer algorithm compared with baseline algorithms. Peizhi Wu, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Wirel. Commun. | 2 |
| 2016 | Resource Allocation for Multicarrier D2D Video Transmission Based on Exact Symbol Error RateabstractWe consider centralized cross-layer resource allocation for a device-to-device (D2D) video transmission system, given knowledge of the channel state information and the rate distortion information of the video streams, and propose an iterative algorithm for subcarrier assignment and power allocation. In the resource allocation, the performance improvement by applying the exact symbol error rate (SER) is compared with the conventional signal-to-interference-plus- noise-ratio (SINR) based SER evaluation method that uses a Gaussian approximation (GA) for the aggregated interference. An exact SER expression for a D2D system using multicarrier bandlimited QAM is derived and then used in the resource allocation algorithm. Bit-level simulations for different numbers of D2D pairs demonstrate a considerable improvement on user capacity and video peak signal-to-noise ratio by incorporating the proposed SER expression compared to the GA. Peizhi Wu, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 2 |
| 2016 | Single image restoration using scene ambient light differentialabstractIn this paper, we restore images degraded by scattering and absorption such as hazy, sandstorm, and underwater images. By calculating the difference between the observed intensity and the ambient light in a degraded image scene, which we call the scene ambient light differential, we estimate the transmission map. In the restoration process, we first enhance the degraded images based on the proposed transmission estimation using the image formation model, and then use an adaptive color correction method to restore color. Experimental results on various degraded images demonstrate the proposed method outperforms other enhancement and restoration methods. Yan-Tsung Peng, Pamela C. Cosman |
ICIP | 2 |
| 2016 | Hardware-efficient debanding and visual enhancement filter for inverse tone mapped high dynamic range images and videosabstractIf a low dynamic range (LDR) image or video is inverse tone mapped to a higher dynamic range, there can be banding artifacts in the output high dynamic range (HDR) image or video. We design a selective sparse filter to remove the banding artifacts and at the same time preserve edges and details. The filter is able to reduce other artifacts, such as blocky artifacts which are due to the compression of the LDR image/video. The filter is computationally efficient. Qing Song 0005, Guan-Ming Su, Pamela C. Cosman |
ICIP | 3 |
| 2016 | Joint error-resilient video source coding and FEC code rate optimization for an AWGN channelabstractA joint source-channel rate-distortion (RD) optimization is proposed for video communication systems. The source coding and channel coding options are optimized by seeking the best trade-off between the estimated end-to-end distortion of a video packet and the sum of the number of source bits and forward error correction bits used to encode that packet. The proposed RD algorithm controls the total bit rate by using a Lagrange multiplier. Compared to conventional RD optimization schemes, which only optimize over the source coding modes of macroblocks, our proposed RD algorithm achieves superior performance over an AWGN channel. Qing Song 0005, Arash Vosoughi, Pamela C. Cosman, Laurence B. Milstein |
ICIP | 3 |
| 2016 | A low complexity model for predicting slice loss distortion for prioritizing H.264/AVC video
Seethal Paluri, Kashyap K. R. Kambhatla, Barbara A. Bailey, Pamela C. Cosman, John D. Matyjas, Sunil Kumar 0001 |
Multim. Tools Appl. | 4 |
| 2016 | Disruptive Attacks on Video Tactical Cognitive Radio DownlinksabstractWe consider video transmission over a mobile cognitive radio (CR) system operating in a hostile environment where an intelligent adversary tries to disrupt communications. We investigate the optimal strategy for spoofing, desynchronizing, and jamming a cluster-based CR network with a Gaussian noise signal over a slow Rayleigh fading channel. The adversary can limit access for secondary users (SUs) by either transmitting a spoofing signal in the sensing interval, or a desynchronizing signal in the code acquisition interval. By jamming the network during the transmission interval, the adversary can reduce the rate of successful transmission. We show how the adversary can optimally allocate its energy across subcarriers during sensing, code acquisition, and transmission intervals. We determine a worst-case optimal energy allocation for spoofing, desynchronizing, and jamming, which gives an upper bound to the received video distortion of SUs. We also propose cross-layer resource allocation algorithms and evaluate their performance under disruptive attacks. Madushanka Soysa, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2015 | Single underwater image enhancement using depth estimation based on blurrinessabstractIn this paper, we propose to use image blurriness to estimate the depth map for underwater image enhancement. It is based on the observation that objects farther from the camera are more blurry for underwater images. Adopting image blurriness with the image formation model (IFM), we can estimate the distance between scene points and the camera and thereby recover and enhance underwater images. Experimental results on enhancing such images in different lighting conditions demonstrate the proposed method performs better than other IFM-based enhancement methods. Yan-Tsung Peng, Xiangyun Zhao, Pamela C. Cosman |
ICIP | 3 |
| 2015 | Multiview coding and error correction coding for 3D video over noisy channels
Arash Vosoughi, Vanessa Testoni, Pamela C. Cosman, Laurence B. Milstein |
Signal Process. Image Commun. | 3 |
| 2015 | Joint Source-Channel Coding and Unequal Error Protection for Video Plus DepthabstractWe consider the joint source-channel coding (JSCC) problem of 3-D stereo video transmission in video plus depth format over noisy channels. Full resolution and downsampled depth maps are considered. The proposed JSCC scheme yields the optimum color and depth quantization parameters as well as the optimum forward error correction code rates used for unequal error protection (UEP) at the packet level. Different coding scenarios are compared and the UEP gain over equal error protection is quantified for flat Rayleigh fading channels. Arash Vosoughi, Pamela C. Cosman, Laurence B. Milstein |
IEEE Signal Process. Lett. | 2 |
| 2015 | Resource Allocation and Performance Analysis for Multiuser Video Transmission Over Doubly Selective ChannelsabstractWe consider an uplink multicarrier system with multiple video users who want to send compressed video data to the base station. In the time domain, we model the time-varying channel using Jakes' model, and in the frequency domain, each subcarrier is assumed to be independently fading. The video is scalably coded in units of a group of pictures (GOP), and users have different video rate distortion (RD) functions. At the beginning of the GOP, the base station collects both the RD information and the instantaneous channel state information (CSI) for subcarrier allocation purposes. We design a cross-layer resource allocation algorithm to assign subcarriers to users based on both the demand of the video and the quality of the channel. Once the resource allocation decision is made, the users then periodically adapt the modulation format of the subcarriers allocated according to the evolution of the CSI for the duration of the GOP. We show that our cross-layer resource allocation robustly outperforms two baseline algorithms, each of which uses only one layer of information for resource allocation. Dawei Wang 0010, Laura Toni, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Wirel. Commun. | 3 |
| 2014 | Metal artifact reduction for CT-based luggage screeningabstractIn aviation security, checked luggage is screened by computed tomography (CT) scanning, followed by automatic target recognition from the CT images. Metal objects in the bags cause image artifacts that degrade object representation, leading to increased false alarms. We develop a new method, which isolates and reduces artifacts in an intermediate image, based on a numerical optimization that de-emphasizes metal and has a novel constraint for beam hardening and scatter. Results on test bags showed excellent artifact reduction, even for multiple metal objects. Seemeen Karimi, Pamela C. Cosman, Harry E. Martz |
ICASSP | 2 |
| 2014 | Optimized Spoofing and Jamming a Cognitive RadioabstractWe examine the performance of a cognitive radio system in a hostile environment where an intelligent adversary tries to disrupt communications by minimizing the system throughput. We investigate the optimal strategy for spoofing and jamming a cognitive radio network with a Gaussian noise signal over a Rayleigh fading channel. We analyze a cluster-based network of secondary users (SUs). The adversary may attack during the sensing interval to limit access for SUs by transmitting a spoofing signal. By jamming the network during the transmission interval, the adversary may reduce the rate of successful transmission. We present how the adversary can optimally allocate power across subcarriers during sensing and transmission intervals with knowledge of the system, using a simple optimization approach specific to this problem. We determine a worst-case optimal energy allocation for spoofing and jamming, which gives a lower bound to the overall information throughput of SUs under attack. Madushanka Soysa, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2014 | Double-Layer Video Transmission Over Decode-and-Forward Wireless Relay Networks Using Hierarchical ModulationabstractWe consider a wireless relay network with a single source, a single destination, and a multiple relay. The relays are half-duplex and use the decode-and-forward protocol. The transmit source is a layered video bitstream, which can be partitioned into two layers, a base layer (BL) and an enhancement layer (EL), where the BL is more important than the EL in terms of the source distortion. The source broadcasts both layers to the relays and the destination using hierarchical 16-QAM. Each relay detects and transmits successfully decoded layers to the destination using either hierarchical 16-QAM or QPSK. The destination can thus receive multiple signals, each of which can include either only the BL or both the BL and the EL. We derive the optimal linear combining method at the destination, where the uncoded bit error rate is minimized. We also present a suboptimal combining method with a closed-form solution, which performs very close to the optimal. We use the proposed double-layer transmission scheme with our combining methods for transmitting layered video bitstreams. Numerical results show that the double-layer scheme can gain 2-2.5 dB in channel signal-to-noise ratio or 5-7 dB in video peak signal-to-noise ratio, compared with the classical single-layer scheme using conventional modulation. Tu V. Nguyen, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2014 | Iterative Pricing-Based Rate Allocation for Video Streams With Fluctuating Bandwidth AvailabilityabstractWe consider rate allocation for video users in the case where the available bandwidth fluctuates. Simply minimizing the objective distortion or optimizing the stability of video qualities does not optimize subjective quality. We formulate a utility-based solution, considering that a user's preference of video quality often varies over a range with upper and lower thresholds of quality. Our iterative pricing-based resource allocation procedure reallocates the bandwidth not only between different users within a time slot but also between different time slots, such that no user suffers quality degradation on average by participating in the multiplexing process. Experimental results show that, compared with equal resource allocation and existing rate allocation solutions, the subjective result becomes increasingly better with the increase of bandwidth fluctuation rate or bandwidth fluctuation range. Moreover, as the number of users increases, the results improve. Meng Yang 0002, Theodore Groves, Nanning Zheng 0001, Pamela C. Cosman |
IEEE Trans. Multim. | 4 |
| 2013 | Joint source-channel coding of 3D video using multiview codingabstractWe consider the joint source-channel coding problem of a 3D video transmitted over an AWGN channel. The goal is to minimize the total number of bits, which is the sum of the number of source bits and the number of forward error correction bits, under two constraints: the quality of the primary view and the quality of the secondary view must be greater than or equal to a predetermined threshold at the receiver. The quality is measured in terms of the expected PSNR of an entire decoded group of pictures. A MVC (multiview coding) encoder is used as the source encoder, and rate compatible punctured turbo codes are utilized for protection of the encoded 3D video over the noisy channel. Equal error protection and unequal error protection are compared for various 3D video sequences and noise levels. Arash Vosoughi, Vanessa Testoni, Pamela C. Cosman, Laurence B. Milstein |
ICASSP | 3 |
| 2013 | Classification based fast mode decision for stereo video codingabstractWe propose a classification based fast mode decision scheme in stereo video coding. By treating mode decision as a classification problem, our scheme employs a decision tree classifier to separate out SKIP mode, which is the major and most computationally efficient mode in stereo video coding. Thus, we can pre-decide whether the current macroblock is coded as SKIP mode without going through the exhaustive mode decision process. Experimental results show that this scheme provides 30~75% of time saving over a wide range of quantization parameter values. Yuan Zhang 0013, Pamela C. Cosman |
ICIP | 3 |
| 2013 | Minimization of Expected Distortion with Layer-Selective Relaying of Two-Layer Superposition CodingabstractThis paper considers a relay system using two-layer superposition coding to minimize the expected distortion of a Gaussian source at the destination node. For the system, we propose two types of layer- selective relaying based on the local decoding result at the relay and the decoding result at the destination node fed back to the relay. One type of the proposed scheme uses decode-and-forward (DF) in the design of the relay signals, while the other type uses both DF and amplify-and-forward (AF). For the proposed scheme, we analyze the outage probabilities and evaluate the expected distortion according to the relay location. The results reveal that the proposed scheme improves the finite SNR performance, in particular when the relay node is closer to the source node than it is to the destination node. Jin Soo Wang, Yun Hee Kim, Pamela C. Cosman, Laurence B. Milstein |
VTC Spring | 3 |
| 2013 | Cooperative Relaying of Superposition Coding with Simple Feedback for Layered Source TransmissionabstractWe consider a relay network that delivers a Gaussian source by employing successive refinement source coding and superposition coding of layers at the source node and successive decoding at the relay and destination nodes. For the network, making use of the decoding results at the relay and destination nodes of the first transmission, an efficient relaying strategy of layers is proposed to minimize the expected distortion (ED) when only the average channel state information is available at the source node. Three types of the proposed scheme, defined as Prop-DF, using decode-and-forward signals, Prop-AF, using amplify-and-forward signals, and Prop-MF, using mixed-forward signals, are addressed and analyzed in terms of the outage probability and distortion exponent. Unlike other studies, we have also taken the relay location into account in deriving the distortion exponent showing the high SNR behavior of the ED. The results show that the proposed scheme increases the distortion exponent up to twice that of the conventional relaying schemes when the relay is close to the source node, and that Prop-MF provides the best performance for most relay locations. Jin Soo Wang, Yun Hee Kim, Iickho Song, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 4 |
| 2013 | Uplink Resource Management for Multiuser OFDM Video Transmission Systems: Analysis and Algorithm DesignabstractWe consider a multiuser OFDM system in which users want to transmit videos via a base station. The base station knows the channel state information (CSI) as well as the rate distortion (RD) information of the video streams and tries to allocate power and spectrum resources to the users according to both physical layer CSI and application layer RD information. We derive and analyze a condition for the optimal resource allocation solution in a continuous frequency response setting. The optimality condition for this cross layer optimization scenario is similar to the equal slope condition for conventional video multiplexing resource allocation. Based on our analysis, we design an iterative subcarrier assignment and power allocation algorithm for an uplink system, and provide numerical performance analysis with different numbers of users. Comparing to systems with either only physical layer or only application layer information available at the base station, our results show that the user capacity and the video PSNR performance can be increased significantly by using cross layer design. Bit-level simulations which take into account the imperfection of the video coding rate control, the variation of RD curve fitting, as well as channel errors, are presented. Dawei Wang 0010, Laura Toni, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 3 |
| 2013 | Motion-Compensated Scalable Video Transmission Over MIMO Wireless ChannelsabstractWe study motion compensated fine granular scalable (MC-FGS) video transmission over multiinput multioutput (MIMO) wireless channels applicable to video streaming, where leaky and partial prediction schemes are applied in the enhancement layer of MC-FGS to exploit the tradeoff between error propagation and coding efficiency. For reliable transmission, we propose unequal error protection (UEP) by considering a tradeoff between reliability and data rates, which are controlled by forward error correction and MIMO mode selection to minimize the average distortion. In a high Doppler environment, where it is hard to get an accurate channel estimate, we investigate the performance of the proposed MC-FGS video transmission scheme with joint control of both the leaky and partial prediction parameters, and the UEP. In a slow fading channel, where the channel throughput can be estimated at the transmitter, adaptive control of prediction parameters is considered. Hobin Kim, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | H.264/AVC video packet aggregation and unequal error protection for noisy channelsabstractThe quality of H.264/AVC compressed video delivery over wireless channels is affected by packet losses. Aggregating H.264/AVC slices to form video packets with sizes adaptive to their importance can improve transmission reliability. Larger packets are more likely to be in error but smaller packets cause more overhead. A second method is assigning stronger channel code rates to more important slices. We use cross-layer dynamic programming to address both adaptive packet formation as well as RCPC-channel code rate allocation simultaneously, to improve received video quality. Simulation results show the advantages of the proposed scheme. Kashyap K. R. Kambhatla, Sunil Kumar 0001, Pamela C. Cosman, John D. Matyjas |
ICIP | 3 |
| 2012 | Predicting slice loss distortion in H.264/AVC video for low complexity data prioritizationabstractWe propose a low complexity Generalized Linear Model (GLM) for prioritizing slices during real-time H.264/AVC compressed video streaming. We train the GLM over a video database to predict the Cumulative Mean Square Error (CMSE) corresponding to individual slice losses by using a combination of efficient video parameters which can be easily extracted during the encoding of a frame. We prioritize the slices generated within a Group of Pictures (GOP) based on the predicted CMSE by using a Quartile Based Prioritization (QBP) scheme. For comparison, we also perform QBP on measured CMSE values from individual pre-encoded slice losses and analyze the priority misclassifications of the slices. We validate our model by applying Unequal Error Protection (UEP) using RCPC codes to the different prioritized bitstreams and evaluating their performance over noisy channels. Simulation results show that predicted CMSE schemes achieve PSNR performance close to that of the measured CMSE schemes for different slice sizes, video bitrates and over different channel SNRs. Seethal Paluri, Kashyap K. R. Kambhatla, Sunil Kumar 0001, Barbara A. Bailey, Pamela C. Cosman, John D. Matyjas |
ICIP | 5 |
| 2012 | Frame loss visibility modeling of stereoscopic video for H.264/AVC-MVCabstractWe develop a visibility model which can predict the visibility of frame losses in compressed 3D video. The 3D video is encoded using the MVC (multiview coding) extension of the H.264/AVC standard. Frame losses both in the left view (base view) and right view (enhancement view) of the stereoscopic video are considered. A subjective test is conducted to identify which types of frame losses are perceptually noticeable. Several features are extracted from the encoded frames, and then support vector machines are employed to build visibility models based on these features. Results show that our model can predict the visibility of frame losses in stereoscopic video with good accuracy. Arash Vosoughi, Pamela C. Cosman |
ICIP | 2 |
| 2012 | Depth-assisted error concealment for intra frame slices in 3D videoabstractWe propose a depth-assisted error concealment method for slice loss in intra frames of 2D+depth video sequence. Intra frames in the 2D view sequence are offset from intra frames in the depth sequence to guarantee the corresponding frame in the other sequence is not also intra mode. Then for a slice loss in an intra frame in the 2D view sequence, the motion information is extracted from the depth sequence to conceal the slice loss using boundary matching. Experimental results show that the proposed method provides improved performance over existing methods both for PSNR results and computational complexity at the decoder. Meng Yang 0002, Yuhong Yang 0004, Pamela C. Cosman |
ICIP | 3 |
| 2012 | Optimization of generalized LT codes for progressive image transferabstractRateless codes allow a user to incrementally send additional redundancy, so they can be useful for heterogeneous and time-varying networks for which the choice of redundancy level in advance is difficult. Rateless codes are an attractive application layer forward error correction solution due to their flexibility and capacity-approaching performance. The original rateless codes were developed for the delivery of equally important information. In many multimedia applications, some data symbols are more important than others. Unequal error protection (UEP) designs are attractive solutions for such transmissions. However previous UEP rateless code designs were aimed for coarsely layered sources and might exhibit poor performance for fine-grained progressive coding. The main objective of this paper is to introduce a more generalized coding scheme, parameters of which can be tailored for progressive multimedia transmission. We present the optimization of a generalized rateless code using two different progressive source transmission protocols. Proposed coding scheme is shown to exhibit better unequal protection and recovery time properties than other published results. Suayb S. Arslan, Pamela C. Cosman, Laurence B. Milstein |
VCIP | 2 |
| 2012 | Adaptive rate control for Wyner-Ziv video codingabstractIn Wyner-Ziv video coding architectures, the available bit budget to each GOP is shared between key frames and Wyner-Ziv frames. In this work, we first propose a model to express the relationship between quantization step size of key and WZ frames based on their motion activity. Then we apply this model to propose an adaptive algorithm adjusting the quantization step size of key and WZ frames to achieve and maintain a target bit rate. We evaluate the rate distortion performance of the proposed method and compare to a common method in the literature. Ghazaleh Esmaili, Pamela C. Cosman |
VCIP | 2 |
| 2012 | Channel Coding Optimization Based on Slice Visibility for Transmission of Compressed Video over OFDM ChannelsabstractOptimization of multimedia transmissions over wireless channels should be aimed at maximizing the video quality perceived by the final user. For transmission of video sequences over an orthogonal frequency division multiplexing (OFDM) system in a slowly varying Rayleigh faded environment, we develop a cross-layer technique, based on a slice loss visibility (SLV) model used to evaluate the visual importance of each slice. In particular, taking into account the visibility scores available from the bitstream, depending on the scenario, we optimize the mapping of video slices within a 2-D time-frequency resource block and/or the channel code rates, in order to better protect more visually important slices. The proposed algorithm is investigated for several scenarios, with different levels of information about the channel available in the optimization process. Results demonstrate that, for different physical environments and different video sequences, the proposed algorithm outperforms baseline ones which do not take into account either the SLV or the CSI in the video transmission. Laura Toni, Pamela C. Cosman, Laurence B. Milstein |
IEEE J. Sel. Areas Commun. | 2 |
| 2012 | Concatenated Block Codes for Unequal Error Protection of Embedded Bit StreamsabstractA state-of-the-art progressive source encoder is combined with a concatenated block coding mechanism to produce a robust source transmission system for embedded bit streams. The proposed scheme efficiently trades off the available total bit budget between information bits and parity bits through efficient information block size adjustment, concatenated block coding, and random block interleavers. The objective is to create embedded codewords such that, for a particular information block, the necessary protection is obtained via multiple channel encodings, contrary to the conventional methods that use a single code rate per information block. This way, a more flexible protection scheme is obtained. The information block size and concatenated coding rates are judiciously chosen to maximize system performance, subject to a total bit budget. The set of codes is usually created by puncturing a low-rate mother code so that a single encoder-decoder pair is used. The proposed scheme is shown to effectively enlarge this code set by providing more protection levels than is possible using the code rate set directly. At the expense of complexity, average system performance is shown to be significantly better than that of several known comparison systems, particularly at higher channel bit error rates. Suayb S. Arslan, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2012 | Generalized Unequal Error Protection LT Codes for Progressive Data TransmissionabstractThe original design of standard digital fountain codes assumes that the coded information symbols are equally important. In many applications, some source symbols are more important than others, and they must be recovered prior to the rest. Unequal Error Protection (UEP) designs are attractive solutions for such source transmissions. In this study, we introduce a more generalized design for the first universal fountain code design, LT codes, that makes it particularly suited for progressive bit stream transmissions. We apply the generalized LT codes to a progressive source and show that it has better UEP properties than other published results in the literature. For example, using the proposed generalization, we obtained up to 1.7dB PSNR gain in a progressive image transmission scenario over the two major UEP fountain code designs. Suayb S. Arslan, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2012 | Iterative Channel Decoding of FEC-Based Multiple-Description CodesabstractMultiple description coding has been receiving attention as a robust transmission framework for multimedia services. This paper studies the iterative decoding of FEC-based multiple description codes. The proposed decoding algorithms take advantage of the error detection capability of Reed-Solomon (RS) erasure codes. The information of correctly decoded RS codewords is exploited to enhance the error correction capability of the Viterbi algorithm at the next iteration of decoding. In the proposed algorithm, an intradescription interleaver is synergistically combined with the iterative decoder. The interleaver does not affect the performance of noniterative decoding but greatly enhances the performance when the system is iteratively decoded. We also address the optimal allocation of RS parity symbols for unequal error protection. For the optimal allocation in iterative decoding, we derive mathematical equations from which the probability distributions of description erasures can be generated in a simple way. The performance of the algorithm is evaluated over an orthogonal frequency-division multiplexing system. The results show that the performance of the multiple description codes is significantly enhanced. Seok-Ho Chang, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2012 | Network-Based H.264/AVC Whole-Frame Loss Visibility Model and Frame Dropping MethodsabstractWe examine the visual effect of whole frame loss by different decoders. Whole frame losses are introduced in H.264/AVC compressed videos which are then decoded by two different decoders with different common concealment effects: frame copy and frame interpolation. The videos are seen by human observers who respond to each glitch they spot. We found that about 39% of whole frame losses of B frames are not observed by any of the subjects, and over 58% of the B frame losses are observed by 20% or fewer of the subjects. Using simple predictive features which can be calculated inside a network node with no access to the original video and no pixel level reconstruction of the frame, we developed models which can predict the visibility of whole B frame losses. The models are then used in a router to predict the visual impact of a frame loss and perform intelligent frame dropping to relieve network congestion. Dropping frames based on their visual scores proves superior to random dropping of B frames. Yueh-Lun Chang, Ting-Lan Lin, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2012 | Optimized Unequal Error Protection Using Multiplexed Hierarchical ModulationabstractWith progressive image or scalable video encoders, as more bits are received, the source can be reconstructed with progressively better quality. These progressive codes have gradual differences of importance in their bitstreams, which necessitates multiple levels of unequal error protection (UEP). One practical method of achieving UEP is based on a constellation of nonuniformly spaced signal points, or hierarchical constellations. However, hierarchical modulation can achieve only a limited number of UEP levels for a given constellation size. Though hierarchical modulation has been intensively studied for digital broadcasting or multimedia transmission, most work has considered only two layered source coding, and methods of achieving a large number of UEP levels for progressive transmission have rarely been studied. In this paper, we propose a multilevel UEP system using multiplexed hierarchical quadrature amplitude modulation (QAM). We show that multiple levels of UEP are achieved by the proposed multiplexing method. When the BER is dominated by the minimum Euclidian distance, we derive an optimal multiplexing approach which minimizes both the average and peak powers. We next propose an asymmetric hierarchical QAM which reduces the peak-to-average power ratio (PAPR) of the proposed UEP system without any performance loss. Numerical results show that the performance of progressive transmission over Rayleigh fading channels is significantly enhanced by the proposed methods. Seok-Ho Chang, Minjoong Rim, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Inf. Theory | 3 |
| 2012 | Wireless H.264 Video Quality Enhancement Through Optimal Prioritized Packet FragmentationabstractWe introduce a cross-layer priority-aware packet fragmentation scheme at the MAC layer to enhance the quality of pre-encoded H.264/AVC compressed bitstreams over bit-rate limited error-prone links in wireless networks. The H.264 slices are classified in four priorities at the encoder based on their cumulative mean square error (CMSE) contribution towards the received video quality. The slices of a priority class in each frame are aggregated into video packets of corresponding priority. We derive the optimal fragment size for each priority class which achieves the maximum expected weighted goodput at different encoded video bit rates, slice sizes and bit error rates. Priority-aware packet fragmentation invokes slice discard in the buffer due to channel bit rate constraints on allocating fragment header bits. We propose a slice discard scheme using frame importance and slice CMSE contribution to control error propagation effects. Packet fragmentation is extended to slice fragmentation by modifying the conventional H.264 decoder to handle partial slice decoding. Priority-aware slice fragmentation combined with the proposed slice discard scheme provides considerable PSNR and VQM gains as compared to priority-agnostic fragmentation. Kashyap K. R. Kambhatla, Sunil Kumar 0001, Seethal Paluri, Pamela C. Cosman |
IEEE Trans. Multim. | 4 |
| 2011 | Cross Layer Resource Allocation Design for Uplink Video OFDMA Wireless SystemsabstractWe study an uplink video communication system with multiple users in a centralized wireless cell. The multiple access scheme is Orthogonal Frequency Division Multiple Access (OFDMA). Both physical layer channel state information (CSI) and application layer rate distortion (RD) information of video streams are collected by the base station. With the goal of minimizing the average video distortion across all the users in the system, we design an iterative resource allocation algorithm for subcarrier assignment and power allocation. Based on the physical layer resource allocation decision, the user will adapt the application layer video source coding rate. To show the advantage of this cross layer algorithm, numerical results are compared with two baseline resource allocation algorithms using only physical layer information or only application layer information. Bit-level simulation results are presented which take into account the imperfection of the video coding rate control, as well as channel errors. Dawei Wang 0010, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 2 |
| 2011 | Prioritized packet fragmentation for H.264 videoabstractWe introduce a cross-layer priority-aware packet fragmentation scheme to enhance the quality of H.264 compressed bitstreams over bit-rate limited error-prone links in packet networks. The H.264 slices are prioritized in the encoder based on their cumulative mean square error (CMSE) contribution towards the received video quality. Specifically we derive the optimal fragment size for each priority level which achieves the maximum expected weighted goodput at different encoded video bit rates and slice sizes. The packet fragmentation scheme uses slice discard in the buffer. Simulation results show that the proposed scheme provides considerably improved video quality. Kashyap K. R. Kambhatla, Sunil Kumar 0001, Pamela C. Cosman |
ICIP | 3 |
| 2011 | Fast mode decision for H.264 video coding in packet loss environmentabstractIn this paper, we propose a fast mode decision scheme for H.264 video coding to address the requirements of both low complexity and error resilience in realtime video communications. Traditional fast mode decision schemes are usually designed based on feature analysis on the source videos. However, the rate-distortion behaviors of the coding modes change when channel errors are involved. Therefore, the existing fast algorithms may not be applicable in error-resilient video coding. We first study the end-to-end rate-distortion behaviors of various coding modes, and then derive a hierarchical mode decision scheme. Different from the traditional fast algorithms that separate skip and inter modes at the beginning, we propose a fast estimation of the coding costs of skip and intra modes in a packet-loss environment, and then narrow the mode decision into one of the two paths: non-intra and non-skip. Testing shows the significant time savings of the proposed algorithm. Yuan Zhang 0013, Pamela C. Cosman |
ICIP | 2 |
| 2011 | Unequal error protection based on slice visibility for transmission of compressed video over OFDM channelsabstractWe address channel code rate optimization for transmission of non-scalable coded video sequences over orthogonal frequency division multiplexing networks. A slice loss visibility (SLV) model is used to evaluate the visual importance of each H.264 slice. Based on both the SLV model and the frequency diversity order available from the channel, we propose a cross-layer technique to allocate video slices within a 2-D time-frequency resource block, and optimize the unequal channel code rate profile, in order to better protect more visually important slices. The proposed algorithm outperforms baseline ones which do not take into account the SLV. Laura Toni, Pamela C. Cosman, Laurence B. Milstein |
ICME | 2 |
| 2011 | Spoofing or Jamming: Performance Analysis of a Tactical Cognitive Radio AdversaryabstractThe tradeoff between spoofing and jamming a cognitive radio network by an intelligent adversary is analyzed in this paper. Due to the vulnerabilities of spectrum sensing noted in recent studies, a cognitive radio can be attacked during the sensing interval by an adversary who puts spoofing signals in unused bands. Further, once secondary users access unused bands, the adversary can use traditional jamming to interfere with them during transmission. For an energy-constrained intelligent adversary, a two step procedure is formulated to distribute the energy between spoofing and jamming, such that the average sum throughput of the secondary users is minimized. That is, we optimally spoof in the sensing duration and then optimally jam in the transmission slot. In a cluster-based cognitive radio network, when the number of spectral vacancies required by secondary users increases, the optimal attack for the intelligent adversary will shift from jamming only, to a combination of spoofing and jamming, to spoofing only. Qihang Peng, Pamela C. Cosman, Laurence B. Milstein |
IEEE J. Sel. Areas Commun. | 2 |
| 2011 | Chernoff-Type Bounds for the Gaussian Error FunctionabstractWe study single-term exponential-type bounds (also known as Chernoff-type bounds) on the Gaussian error function. This type of bound is analytically the simplest such that the performance metrics in most fading channel models can be expressed in a concise closed form. We derive the conditions for a general single-term exponential function to be an upper or lower bound on the Gaussian error function. We prove that there exists no tighter single-term exponential upper bound beyond the Chernoff bound employing a factor of one-half. Regarding the lower bound, we prove that the single-term exponential lower bound of this letter outperforms previous work. Numerical results show that the tightness of our lower bound is comparable to that of previous work employing eight exponential terms. Seok-Ho Chang, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2011 | Superposition MIMO Coding for the Broadcast of Layered SourcesabstractWe propose superposition multiple-input multiple-output (MIMO) coding for the transmission of unequally important sources in a point-to-multipoint system. First, a tradeoff between Alamouti code and spatial multiplexing (V-BLAST) is analyzed in terms of the average bit error rate (BER), where the maximum data rates of both MIMO schemes are set to be equal. The results show that for a given target bit error rate, Alamouti code is preferable for a low data rate, and spatial multiplexing is preferable for a high data rate. For layered sources such as scalable video, the more important component typically has lower data rate than does the less important component. Based on these, we construct a superposition MIMO scheme where two different MIMO techniques are hierarchically combined such that important data is Alamouti coded, less important data is spatially multiplexed, and then two unequally important data symbols are superposed. Seok-Ho Chang, Minjoong Rim, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 3 |
| 2011 | Performance Analysis of n -Channel Symmetric FEC-Based Multiple Description Coding for OFDM NetworksabstractRecently, multiple description source coding has emerged as an attractive framework for robust multimedia transmission over packet erasure channels. In this paper, we mathematically analyze the performance of n-channel symmetric FEC-based multiple description coding for a progressive mode of transmission over orthogonal frequency division multiplexing (OFDM) networks in a frequency-selective slowly-varying Rayleigh faded environment. We derive the expressions for the bounds of the throughput and distortion performance of the system in an explicit closed form, whereas the exact performance is given by an expression in the form of a single integration. Based on this analysis, the performance of the system can be numerically evaluated. Our results show that at high SNR, the multiple description encoder does not need to fine-tune the optimization parameters of the system due to the correlated nature of the subcarriers. It is also shown that, despite the bursty nature of the errors in a slow fading environment, FEC-based multiple description coding without temporal coding provides a greater advantage for smaller description sizes. Seok-Ho Chang, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2011 | Wyner-Ziv Video Coding With Classified Correlation Noise Estimation and Key Frame Coding Mode SelectionabstractWe improve the overall rate-distortion performance of distributed video coding by efficient techniques of correlation noise estimation and key frame encoding. In existing transform-domain Wyner-Ziv video coding methods, blocks within a frame are treated uniformly to estimate the correlation noise even though the success of generating side information is different for each block. We propose a method to estimate the correlation noise by differentiating blocks within a frame based on the accuracy of the side information. Simulation results show up to 2 dB improvement over conventional methods without increasing encoder complexity. Also, in traditional Wyner-Ziv video coding, the intercorrelation of key frames is not exploited since they are simply intracoded. In this paper, we propose a frequency band coding mode selection for key frames to exploit similarities between adjacent key frames at the decoder. Simulation results show significant improvement especially for low-motion and high frame rate sequences. Furthermore, the advantage of applying both schemes in a hierarchical order is investigated. This method achieves additional improvement. Ghazaleh Esmaili, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2011 | Bit-Rate Allocation for Multiple Video Streams Using a Pricing-Based MechanismabstractWe consider the problem of bit-rate allocation for multiple video users sharing a common transmission channel. Previously, overall quality of multiple users was improved by exploiting relative video complexity. Users with high-complexity video benefit at the expense of video quality reduction for other users with simpler videos. The quality of all users can be improved by collectively allocating the bit rate in a centralized fashion which requires sharing video information with a central controller. In this paper, we present an informationally decentralized bit-rate allocation for multiple users where a user only needs to inform his demand to an allocator. Each user separately calculates his bit-rate demand based on his video complexity and bit-rate price, where the bit-rate price is announced by the allocator. The allocator adjusts the bit-rate price for the next period based on the bit rate demanded by the users and the total available bit-rate supply. Simulation results show that all users improve their quality by the pricing-based decentralized bit-rate allocation method compared with their allocation when acting individually. The results of our proposed method are comparable to the centralized bit-rate allocation. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2010 | Packet Dropping for Widely Varying Bit Reduction Rates Using a Network-Based Packet Loss Visibility ModelabstractWe propose a packet dropping algorithm for various packet loss rates. A network-based packet loss visibility model is used to evaluate the visual importance of each H.264 packet inside the network. During network congestion, based on the estimated loss visibility of each packet, we drop the least visible frames and/or the least visible packets until the required bit reduction rate is achieved. Based on a computable perceptually-based metric, our algorithm performs better than an existing approach (dropping B packets or frames). Ting-Lan Lin, Jihyun Shin, Pamela C. Cosman |
DCC | 3 |
| 2010 | Network-Based Model for Video Packet Importance Considering Both Compression Artifacts and Packet LossesabstractIndividual packet losses can have differing impact on video quality. Simple factors such as packet size, average motion, and DCT coefficient energy can be extracted from an individual compressed video packet inside the network without any inverse transforms or pixel-level decoding. Using only such factors that are self-contained within packets, we aim to predict the impact on quality as measured by VQM (video quality metric) that the loss of this packet would entail. In the context of both compression artifacts and packet loss artifacts, we develop generalized linear models to predict VQM scores and our final model gives a good performance on objective evaluation of packet importance. Ting-Lan Lin, Pamela C. Cosman |
GLOBECOM | 3 |
| 2010 | Network-based packet loss visibility model for SDTV and HDTV for H.264 videosabstractWe conduct subjective experiments on visual quality following packet loss, and then construct models to predict these visual importance scores. The models are fully self-contained at the packet level, meaning that they use only information within one packet to predict the importance of that packet, requiring no frame-level reconstruction nor any information on the reference frame. Models are created for SDTV and HDTV resolutions, and the differences in the important factors between them are discussed. Ting-Lan Lin, Pamela C. Cosman |
ICASSP | 2 |
| 2010 | Classification of MPEG-2 Transport Stream packet loss visibilityabstractWe classify the visibility of TS (Transport Stream) packet losses for SDTV and HDTV MPEG-2 compressed video streams. TS packet losses can cause various temporal and spatial losses. The visual effect of a TS packet loss depends on many factors, in particular whether the loss causes a whole frame loss or partial frame loss. We develop models for predicting loss visibility for both SDTV and HDTV resolutions for frame loss and partial frame loss cases. We compare the dominant predictive factors and the results for the two resolutions. We achieve more than 85% classification accuracy. Jihyun Shin, Pamela C. Cosman |
ICASSP | 2 |
| 2010 | Variance-Aware Adaptive Modulation for OFDM-Based Multiple Description Progressive Image TransmissionabstractAverage throughput maximization has commonly been used as an alternative to image or video average distortion minimization. However, as the throughput-distortion curve is non-linear, throughput maximization can be very different from distortion minimization if the variance of the throughput is large. Under high channel quality variability, the variance of the throughput will be high. Multiple description coding with unequal error protection is a promising technique for the transmission of progressive images when there exists a high variability of the channel conditions. In this paper, we consider both the variance and the average of the throughput when deciding the constellation size for adaptive modulation in an OFDM system used for transmitting progressively-coded images with multiple description coding. Simulation results show that in an adaptive modulated image transmission system, taking into account both moments of the throughput achieves better performance than a system considering average throughput alone. S. S. Tan, Minjoong Rim, Pamela C. Cosman, Laurence B. Milstein |
ICC | 3 |
| 2010 | Throughput and Delay Analysis for Real-Time Applications in Ad-Hoc Cognitive NetworksabstractWe consider a simple ad-hoc cognitive scenario with two data up-links, one licensed to use the spectral resource (primary) and the other unlicensed (secondary or cognitive). It is assumed that the cognitive link accesses the channel only when the channel is sensed idle. An ON-OFF channel model is used for the primary link, where traffic statistical characteristics are taken into account. A closed-form expression for the signal-to noise-plus interference (SINR) statistics of the cognitive nodes is derived that can be used for estimating the network performance. Moreover, a M/G/1 queueing model is exploited for deriving a simple expression for the average packet delay. Finally, a MAC strategy based on a channel-and-queue aware scheduling is introduced. Diego Piazza, Pamela C. Cosman, Laurence B. Milstein, Guido Tartara |
WCNC | 2 |
| 2010 | Efficient Optimal RCPC Code Rate Allocation With Packet Discarding for Pre-Encoded Compressed VideoabstractIn an error-prone communication channel, more important video packets should be assigned stronger channel codes. With various packet sizes and distortions for each packet, we use the subgradient method to search in the dual domain for the optimal RCPC channel code rate allocation for each packet, to minimize the end-to-end video quality degradation for an AWGN channel. We exploit the advantage of not sending or not coding packets of lower importance. Ting-Lan Lin, Pamela C. Cosman |
IEEE Signal Process. Lett. | 2 |
| 2010 | Source-channel rate optimization for progressive image transmission over block fading relay channels [Transactions Papers]abstractIn this paper, we are concerned with the design and analysis of joint source-channel coding schemes for block fading channels with relay-assisted distributed spatial diversity. Assuming a progressive image coder with a constraint on the transmission bandwidth, we formulate a joint source-channel rate allocation scheme that maximizes the expected source throughput. Specifically, using Gaussian as well as BPSK inputs on flat Rayleigh fading channels, we lower bound the average packet error rate by the corresponding mutual information outage probability, and derive the average throughput expression as a function of channel code rates as well as channel SNR for both a frequency-division multiplexing-based baseline system without relaying, and a half-duplex relay system with a decode-and- forward protocol. At high signal-to-noise ratio (SNR), for the systems considered in this paper, we show that our rate optimization problem is a convex function of the channel code rates, and we show that a known recursive algorithm can be used to predict the performance of both systems. Hobin Kim, Ramesh Annavajjala, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 3 |
| 2010 | A Versatile Model for Packet Loss Visibility and its Application to Packet PrioritizationabstractIn this paper, we propose a generalized linear model for video packet loss visibility that is applicable to different group-of-picture structures. We develop the model using three subjective experiment data sets that span various encoding standards (H.264 and MPEG-2), group-of-picture structures, and decoder error concealment choices. We consider factors not only within a packet, but also in its vicinity, to account for possible temporal and spatial masking effects. We discover that the factors of scene cuts, camera motion, and reference distance are highly significant to the packet loss visibility. We apply our visibility model to packet prioritization for a video stream; when the network gets congested at an intermediate router, the router is able to decide which packets to drop such that visual quality of the video is minimally impacted. To show the effectiveness of our visibility model and its corresponding packet prioritization method, experiments are done to compare our perceptual-quality-based packet prioritization approach with existing Drop-Tail and Hint-Track-inspired cumulative-MSE-based prioritization methods. The result shows that our prioritization method produces videos of higher perceptual quality for different network conditions and group-of-picture structures. Our model was developed using data from high encoding-rate videos, and designed for high-quality video transported over a mostly reliable network; however, the experiments show the model is applicable to different encoding rates. Ting-Lan Lin, Sandeep Kanumuri, Yuan Zhi, David Poole 0003, Pamela C. Cosman, Amy R. Reibman |
IEEE Trans. Image Process. | 5 |
| 2010 | Competitive Equilibrium Bitrate Allocation for Multiple Video StreamsabstractWe consider the problem of simultaneous bitrate allocation for multiple video streams. Current methods for multiplexing video streams often rely on identifying the relative complexity of the video streams to improve the combined overall quality. In such methods, not all the videos benefit from the multiplexing process. Typically, the quality of high motion videos is improved at the expense of a reduction in the quality of low motion videos. In our approach, we use a competitive equilibrium allocation of bitrate to improve the quality of all the video streams by finding trades between videos across time. A central controller collects rate-distortion information from each video user and makes a joint bitrate allocation decision. Each user encodes and transmits his video at the allocated bitrate through a shared channel. The proposed method uses information about not only the differing complexity of the video streams at every moment but also the differing complexity of each stream over time. Using the competitive equilibrium bitrate allocation approach for multiple video streams, simulation results show that all the video streams perform better or at least as well as with individual encoding. The results of this research will be useful both for ad hoc networks that employ a cluster head model and for cellular architectures. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2010 | Delay Constrained Multiplexing of Video Streams Using Dual-Frame Video CodingabstractWe consider the multiplexing problem of transmitting multiple video source streams from a server over a shared channel. We use dual-frame video coding with high-quality Long-Term Reference (LTR) frames and propose multiplexing methods to reduce the sum of mean squared error for all the video streams. This paper makes several improvements to dual-frame video coding. A simple motion activity detection algorithm is used to choose the location of LTR frames as well as the number of bits given to such frames. An adaptive buffer-constrained rate-control algorithm is devised to accommodate the extra bits of the high-quality LTR frames. Multiplexing of video streams is studied under the constraint of a video encoder delay buffer. Using H.264/AVC, the results show considerable improvement over baseline schemes such as H.264 rate control when the video streams are encoded individually and over multiplexing methods proposed previously in the literature. The high-quality LTR frames are offset in time among different video streams. This provides the benefits of dual-frame coding with high-quality LTR frames while still fitting under the constraint of an output delay buffer. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
IEEE Trans. Image Process. | 3 |
| 2009 | Low Complexity Spatio-Temporal Key Frame Encoding for Wyner-Ziv Video CodingabstractIn most Wyner-Ziv video coding approaches, the temporal correlation of key frames is not exploited since they are simply intra encoded and decoded. In this paper, using the previously decoded key frame as the side information for the key frame to be decoded, we propose new methods of coding key frames in order to improve the rate distortion performance. These schemes which are based on switching between intra mode and Wyner-Ziv mode for a given block or a given frequency band attempt to make use of both spatial and temporal correlation of key frames while satisfying the low complexity encoding requirement of distributed video coding (DVC). Simulation results show that using the proposed methods, one can achieve up to 5 dB improvement over conventional intra coding for relatively low motion sequences and up to 1.3 dB improvement for relatively high motion sequences. Ghazaleh Esmaili, Pamela C. Cosman |
DCC | 2 |
| 2009 | Progressive Source Transmissions Using Joint Source-Channel Coding and Hierarchical Modulation in Packetized NetworksabstractWith Unequal Error Protection (UEP), more important symbols are given greater protection against channel errors than are less important symbols. The protection can be accomplished by various methods, including joint source-channel coding (JSCC) and hierarchical modulation. In this paper, a robust progressive source transmission system using forward error correction (FEC) and a bits-to-symbol assignment methodology that provides UEP is proposed. UEP is not only provided by hierarchical modulation but also by the packetization methodology combined with channel coding. It is demonstrated by simulation that our system improves performance compared to an Equal Error Protection (EEP) technique and to a baseline JSCC-only mechanism. Suayb S. Arslan, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 2 |
| 2009 | Optimal Multiplexed Hierarchical Modulation for Unequal Error Protection of Progressive Bit StreamsabstractProgressive image and scalable video have gradual differences of importance in their bitstreams, which can benefit from multiple levels of unequal error protection (UEP). Though hierarchical modulation has been intensively studied as an UEP approach for digital broadcasting and multimedia transmission, methods of achieving a large number of UEP levels have rarely been studied. In this paper, we propose a multilevel UEP system using multiplexed hierarchical quadrature amplitude modulation (QAM) for progressive transmission over mobile radio channels. We suggest a specific way of multiplexing, and prove that multiple levels of UEP are achieved by the suggested method. When the BER is dominated by the minimum Euclidian distance, we derive an optimal multiplexing approach which minimizes both the average and peak powers. An asymmetric hierarchical QAM which reduces the peak-to-average power ratio (PAPR) without performance loss is also proposed. Numerical results show that the performance of progressive transmission over Rayleigh fading channels is significantly enhanced by the proposed UEP systems. Seok-Ho Chang, Minjoong Rim, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 3 |
| 2009 | Motion-Compensated Scalable Video Transmission over MIMO Wireless Channels under Imperfect Channel EstimationabstractWe study motion compensated fine granular scalable (MC-FGS) video transmission over multi-input multi-output (MIMO) wireless channels, where leaky and partial prediction schemes are applied in the enhancement layer of MC-FGS to exploit the tradeoff between error propagation and coding efficiency. For reliable transmission, we propose unequal error protection (UEP) by considering a tradeoff between reliability and data rates, which are controlled by forward error correction (FEC) and MIMO mode selection to minimize the average distortion. In a high Doppler environment where it is hard to get an accurate channel estimate, we investigate the performance of the proposed MC-FGS video transmission scheme with joint control of the leaky and partial prediction parameters and the UEP. Hobin Kim, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 3 |
| 2009 | Correlation noise classification based on matching success for transform domain Wyner-Ziv video codingabstractDistributed source coding strongly depends on the knowledge of statistical dependency between source and side information. In transform domain Wyner-Ziv video coding (TDWZ) this statistical dependency (also known as correlation noise) has been usually modeled by a unique Laplacian distribution for each frequency band. In this paper, we propose a method to define different classes of correlation noise for each frequency band based on the accuracy of the side information. With this approach the correlation between source and side information is estimated separately for each frequency band of each class. Therefore, the decoder can discriminate blocks in order to estimate the correlation noise of their frequency bands. Simulation results show that applying the proposed method improves rate-distortion performance. Ghazaleh Esmaili, Pamela C. Cosman |
ICASSP | 2 |
| 2009 | Perceptual quality based packet dropping for generalized video GOP structuresabstractOur work builds a general visibility model of video packets which is applicable to various types of GOP (group of pictures). The data used for analysis and building the model come from three subjective experiment sets with different encoding and decoding parameters on H.264 and MPEG-2 videos. We consider factors not only within a packet but also across its vicinity to account for possible temporal and spatial masking effects. This model can be useful for an intermediate router in a congested network to drop less visible packets to maintain overall video quality. Experiments are done to compare our perceptual-quality-based packet dropping approach with existing drop-tail and hint-track-inspired cumulative-MSE-based dropping methods. The result shows that our dropping method produces videos of higher perceptual quality for different network conditions and GOP structures. Ting-Lan Lin, Yuan Zhi, Sandeep Kanumuri, Pamela C. Cosman, Amy R. Reibman |
ICASSP | 4 |
| 2009 | Pricing-based decentralized rate allocation for multiple video streamsabstractWe consider rate allocation for multiple video users sharing a constant bitrate channel. Previously, overall quality of multiple users was improved by exploiting relative complexity. Users with high complexity video benefit at the expense of video quality reduction for other users with simpler videos. The quality of all users can be improved by collectively allocating the bitrate which requires sharing video information with a central controller. In this paper, we present an informationally decentralized rate allocation for multiple users where a user only needs to inform its demand to an allocator based on its video complexity and bitrate price. Simulation results show that all users improve their quality by our pricing-based decentralized rate allocation method compared to their allocation when acting individually and the results are comparable to the centralized rate allocation. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
ICIP | 3 |
| 2009 | Frequency Band Coding Mode Selection for Key Frames of Wyner-Ziv Video CodingabstractIn most Wyner-Ziv video coding approaches, the temporal correlation of key frames is not exploited, since they are simply intra encoded and decoded. In previous work, by using the previously decoded key frames as the side information, we proposed to divide the frequency bands of each block into two classes. Wyner-Ziv coding was used for the low frequency bands of each block, while high frequency bands were intra coded. In this paper, we improve this approach with an efficient coding mode selection technique. Frequency bands are grouped as low and high bands and an appropriate method of coding is selected for them based on the correlation characteristics of each frame with the past. This method does not add complexity to the encoder. Simulation results show that using the proposed method, one can achieve up to 4 dB improvement over prior work. Ghazaleh Esmaili, Pamela C. Cosman |
ISM | 2 |
| 2009 | A resource allocation algorithm for real-time streaming in cognitive networksabstractCognitive radios have been proposed as a means to implement efficient reuse of the licensed spectrum. Commonly, wireless networks are characterized by a fixed spectrum assignment policy. The limited available spectrum and the inefficiency in the spectrum usage necessitate a new communication paradigm to exploit the existing wireless spectrum opportunistically. We consider a simple single-cell scenario with two data up-links, one licensed to use the spectral resource (primary) and the other unlicensed (secondary or cognitive). It is assumed that the cognitive user accesses the channel only when the channel is sensed idle. An ON-OFF channel model is used for the primary link, where traffic statistical characteristics are taken into account. We study a practical resource allocation algorithm that assigns the uplink to the secondary users according to a channel-and-queues aware scheduler when primary link OFF periods are sensed. We fit the resource allocation algorithm to the widely investigated orthogonal frequency division multiple access (OFDMA) scheme and we exploit multiuser diversity by applying a smart power allocation within independent OFDMA subchannels. A video encoder rate control is introduced in order to limit the video frame loss due to overflow that trades the video frame loss probability with the overall encoding quality. Lastly, the performance of the cognitive network model is investigated under the proposed resource allocation algorithm. Diego Piazza, Pamela C. Cosman, Laurence B. Milstein, Guido Tartara |
WCNC | 2 |
| 2009 | Channel Coding for Progressive Images in a 2-D Time-Frequency OFDM Block With Channel Estimation ErrorsabstractCoding and diversity are very effective techniques for improving transmission reliability in a mobile wireless environment. The use of diversity is particularly important for multimedia communications over fading channels. In this work, we study the transmission of progressive image bitstreams using channel coding in a 2-D time-frequency resource block in an OFDM network, employing time and frequency diversities simultaneously. In particular, in the frequency domain, based on the order of diversity and the correlation of individual subcarriers, we construct symmetric n -channel FEC-based multiple descriptions using channel erasure codes combined with embedded image coding. In the time domain, a concatenation of RCPC codes and CRC codes is employed to protect individual descriptions. We consider the physical channel conditions arising from various coherence bandwidths and coherence times, leading to a range of orders of diversities available in the time and frequency domains. We investigate the effects of different error patterns on the delivered image quality due to various fade rates. We also study the tradeoffs and compare the relative effectiveness associated with the use of erasure codes in the frequency domain and convolutional codes in the time domain under different physical environments. Both the effects of intercarrier interference and channel estimation errors are included in our study. Specifically, the effects of channel estimation errors, frequency selectivity and the rate of the channel variations are taken into consideration for the construction of the 2-D time-frequency block. We provide results showing the gain that the proposed model achieves compared to a system without temporal coding. In one example, for a system experiencing flat fading, low Doppler, and imperfect CSI, we find that the increase in PSNR compared to a system without time diversity is as much as 9.4 dB. Laura Toni, Yee Sin Chan, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 3 |
| 2008 | Adaptive Modulation for OFDM-Based Multiple Description Progressive Image TransmissionabstractThis paper addresses the use of adaptive modulation in progressive image transmission with multiple description coding in conjunction with an Orthogonal Frequency Division Multiplexing (OFDM) system. Specifically, two adaptive systems are considered: variable rate with fixed power, and variable rate with variable power. An algorithm is proposed to allocate power and constellation size at each subchannel by maximizing the throughput. Simulation results confirm that cross-layer optimization with adaptive modulation enhances system performance. S. S. Tan, Minjoong Rim, Pamela C. Cosman, Laurence B. Milstein |
GLOBECOM | 3 |
| 2008 | Multiplexing video streams using dual-frame video codingabstractWe consider the transmission of multiple video source streams over a shared channel from a server. Using the rate-distortion curves and dual-frame video coding with high quality long-term reference (LTR) frames, we propose a method to reduce the sum of mean squared error for all the video streams. A simple motion activity detection algorithm was used to determine the amount of high quality given to the LTR frames. Using H.264/AVC, the results show considerable improvement over a baseline scheme where each video stream is provided with equal bitrate. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
ICASSP | 3 |
| 2008 | Perceptual impact of burthy versus isolated packet losses in H.264 compressed videoabstractWhen video packets are lost in congested networks, one loss pattern creates a different visual impact than another. We conduct a subjective experiment with H.264 videos and conclude that isolated losses are better than bursty losses in terms of perceptual video quality. A network-implementable video quality model is developed for a router to drop packets so as to achieve good visual quality. Ting-Lan Lin, Pamela C. Cosman, Amy R. Reibman |
ICIP | 2 |
| 2008 | Buffer constrained rate control for low bitrate dual-frame video codingabstractIn dual-frame video coding, one long-term reference (LTR) and one short-term reference (STR) frames are used for motion estimation and compensation. In previous work, it was shown that the performance of video coding can be improved by pulsing the quality of LTR frames in dual-frame video coding, but this increases the encoder delay buffer size. Also, buffer constrained real-time video transmission requires an efficient rate control algorithm to meet the delay requirement. In this paper, we propose a rate control algorithm for dual- frame video coding under a delay buffer constraint. With the proposed rate control algorithm and motion activity detection for determining the LTR quality, simulation results using H.264/AVC show a significant PSNR improvement over H.264 rate control and other rate control algorithms for dual- frame video coding. Mayank Tiwari 0001, Theodore Groves, Pamela C. Cosman |
ICIP | 3 |
| 2008 | Selection of Long-Term Reference Frames in Dual-Frame Video Coding Using Simulated AnnealingabstractIn dual-frame video coding, both encoder and decoder store a short-term reference (STR) and a long-term reference (LTR) frame for motion compensation. In past work, LTR frames at regular intervals were assigned higher quality than the other frames to improve overall video quality. In this letter, we present a method of LTR frame selection using simulated annealing, and we show that PSNR is improved compared to the case of evenly spaced LTR frames. To reduce delay and computational complexity, we consider a constraint on the size of the look-ahead window. Mayank Tiwari 0001, Pamela C. Cosman |
IEEE Signal Process. Lett. | 2 |
| 2008 | Dual Frame Motion Compensation With Uneven Quality AssignmentabstractVideo codecs that use motion compensation have shown PSNR gains from the use of multiple frame prediction, in which more than one past reference frame is available for motion estimation. In dual frame motion compensation, one short-term reference frame and one long-term reference frame are available for prediction. In this paper, we explore using dual frame motion compensation in two contexts. We first show that using a single fixed long-term reference frame in the context of a rate switching network can enhance video quality. Next, by periodically creating high-quality long-term reference frames, we show that the performance is superior to a standard dual frame technique that has the same average rate but no high-quality frames. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | A Multiple Description Coding and Delivery Scheme for Motion-Compensated Fine Granularity Scalable VideoabstractMotion-compensated fine-granularity scalability (MC-FGS) with leaky prediction has been shown to provide an efficient tradeoff between compression gain and error resilience, facilitating the transmission of video over dynamic channel conditions. In this paper, we propose an n-channel symmetric motion-compensated multiple description (MD) coding and transmission scheme for the delivery of scalable video over orthogonal frequency division multiplexed systems, utilizing the concepts of partial and leaky predictions. We investigate the proposed MD coding and transmission scheme using a cross-layer design perspective. In particular, we construct the symmetric motion-compensated MD codes based on the diversity order of the channel, defined as the ratio of the overall bandwidth of the system to the coherence bandwidth of the channel. We show that knowing the diversity order of a physical channel can assist an MC-FGS video coder in selecting the motion-compensation prediction point, as well as on the use of leaky prediction. More importantly, we illustrate how the side information can reduce the drift management problem associated with the construction of symmetric motion-compensated MD codes. We provide results based on both an information-theoretic approach and simulations. Yee Sin Chan, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2007 | Flicker Suppression in JPEG2000 using Segmentation-Based Adjustment of Block Truncation LengthsabstractFlickering is a temporal visual artifact that affects compressed video. It is prominent in intra-frame video coders and is largely the result of content variations and quantization. We concentrate on flickering due to quantization. JPEG2000 uses post-compression quantization which is applied through the EBCOT algorithm. EBCOT has been found, however, to cause significant flickering in the reconstructed video. In this work, we evaluate existing flicker metrics, investigate the causes of flicker, and propose a new rate-distortion optimal algorithm that suppresses flicker. The proposed algorithm suppresses temporal flicker at a negligible cost in spatial image quality. Athanasios Leontaris, Yoshihide Tonomura, Takayuki Nakachi, Pamela C. Cosman |
ICASSP (1) | 4 |
| 2007 | Statistical channel knowledge-based optimum power allocation for relaying protocols in the high SNR regimeabstractWe are concerned with transmit power optimization in a wireless relay network with various cooperation protocols. With statistical channel knowledge (in the form of knowledge of the fading distribution and the path loss information across all the nodes) at the transmitters and perfect channel state information at the receivers, we derive the optimal power allocation that minimizes high signal-to-noise ratio (SNR) approximations of the outage probability of the mutual information (MI) with amplify-and-forward (AF), decode-and-forward (DF) and distributed space-time coded (DSTC) relaying protocols operating over Rayleigh fading channels. We demonstrate that the high SNR approximation-based outage probability expressions are convex functions of the transmit power vector, and the nature of the optimal power allocation depends on whether or not a direct link between the source and the destination exists. Interestingly, for AF and DF protocols, this allocation depends only on the ratio of mean channel power gains (i.e., the ratio of the source-relay gain to the relay-destination gain), whereas with a DSTC protocol this allocation also depends on the transmission rate when a direct link exists. In addition to the immediate benefits of improved outage behavior, our results show that optimal power allocation brings impressive coding gains over equal power allocation. Furthermore, our analysis reveals that the coding gain gap between the AF and DF protocols can also be reduced by the optimal power allocation. Ramesh Annavajjala, Pamela C. Cosman, Laurence B. Milstein |
IEEE J. Sel. Areas Commun. | 2 |
| 2007 | Compression Efficiency and Delay Tradeoffs for Hierarchical B-Pictures and Pulsed-Quality FramesabstractReal-time video applications require tight bounds on end-to-end delay. Hierarchical bidirectional prediction requires buffering frames in the encoder input buffer, thereby contributing to encoder input delay. Long-term frame prediction with pulsed quality requires buffering at the encoder output, increasing the output buffer delay. Both hierarchical B-pictures and pulsed-quality coders involve uneven bit-rate allocation. Both the encoder and decoder buffering requirements depend on the rate allocation. We derive an efficient rate allocation for hierarchical B-pictures using the power spectral density of a wide-sense stationary process. In addition, we discuss important aspects of hierarchical predictive coding, such as the effect of the temporal prediction distance and delay tradeoffs for prediction branch truncation. Finally, we investigate experimentally the tradeoff between delay and compression efficiency. Athanasios Leontaris, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2007 | Quality Evaluation of Motion-Compensated Edge Artifacts in Compressed VideoabstractLittle attention has been paid to an impairment common in motion-compensated video compression: the addition of high-frequency (HF) energy as motion compensation displaces blocking artifacts off block boundaries. In this paper, we employ an energy-based approach to measure this motion-compensated edge artifact, using both compressed bitstream information and decoded pixels. We evaluate the performance of our proposed metric, along with several blocking and blurring metrics, on compressed video in two ways. First, ordinal scales are evaluated through a series of expectations that a good quality metric should satisfy: the objective evaluation. Then, the best performing metrics are subjectively evaluated. The same subjective data set is finally used to obtain interval scales to gain more insight. Experimental results show that we accurately estimate the percentage of the added HF energy in compressed video. Athanasios Leontaris, Pamela C. Cosman, Amy R. Reibman |
IEEE Trans. Image Process. | 2 |
| 2007 | Performance Analysis of Linear Modulation Schemes With Generalized Diversity Combining on Rayleigh Fading Channels With Noisy Channel EstimatesabstractGeneralized diversity combining (GDC), also known as hybrid selection/maximal ratio combining or generalized selection combining, is a low-complexity diversity combining technique by which a fixed subset of a large number of available diversity channels is chosen and then combined using the rules of maximal ratio combining. In this paper, we analyze the performance of GDC on time-correlated Rayleigh fading channels with noisy channel estimates. We derive expressions for the probability of error for various linear modulation schemes with coherent detection, and discuss the conditions under which the analysis can be extended to noncoherent and differentially coherent receiver structures. Throughout the paper, using a fundamental approach to obtain the decision statistic at the combiner output, a number of new expressions for the error probabilities are obtained in a rigorous way, along with a presentation of their performance with channel estimation errors. The final expressions have roughly the same complexity of evaluation as that for the channel with only additive Gaussian noise. Our results correct various inaccuracies in the literature, and show that coherent receivers based on imperfectly estimated channel knowledge incur a significant performance loss. Ramesh Annavajjala, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Inf. Theory | 2 |
| 2006 | Dual Frame Video Coding with Pulsed Quality and a Lookahead WindowabstractIn dual frame video coding, one short-term reference frame and one long-term reference frame are available for motion compensation. In prior research, it was shown that overall video quality was improved by allocating bits unevenly among frames so as to periodically create a high-quality frame that could serve as the long-term reference frame for some time. We extend this work to a cognitive radio scenario where pulses of extra bandwidth can be rented by the second, but rental agreements can be cancelled on short notice if the legacy user of that spectrum returns. With a scalable video coder, this pulse of extra bandwidth can be used to improve the quality of the current frame being encoded, or of the past long term frame, or can be spent on future frames if the encoder has access to them in advance. We compare these various uses to explore the advantages of allocating some of the available bandwidth to past or future frames Mayank Tiwari 0001, Pamela C. Cosman |
DCC | 2 |
| 2006 | Predicting H.264 Packet Loss Visibility using a Generalized Linear ModelabstractWe consider modeling the visibility of individual and multiple packet losses in H.264 videos. We propose a model for predicting the visibility of multiple packet losses and demonstrate its performance on dual losses (two nearby packet losses). We extract the factors affecting visibility using a reduced-reference method. We predict the probability that a loss is visible using a generalized linear model. We achieve MSE values (between actual and predicted probabilities) of 0.0253 and 0.0398 for individual and dual losses respectively. We also examine the effect of various factors on visibility. Sandeep Kanumuri, Sitaraman G. Subramanian, Pamela C. Cosman, Amy R. Reibman |
ICIP | 3 |
| 2006 | End-to-End Delay for Hierarchical B-Pictures and Pulsed Quality Dual Frame Video CodersabstractReal-time video applications require tight bounds on end-to-end delay. Hierarchical bi-directional prediction requires buffering frames in the encoder input buffer, thereby contributing to encoder input delay. Long-term frame prediction with pulsed quality requires buffering at the encoder output, increasing the output buffer delay. We compare the end-to-end delay of these two approaches using simulations to determine the delay vs. compression efficiency trade-off. Athanasios Leontaris, Pamela C. Cosman |
ICIP | 2 |
| 2006 | First-order Markov Models for Packet Transmission on Rayleigh Fading Channels with DPSK/NCFSK ModulationabstractIn this paper, we develop first-order Markov models that characterize the packet error processes on Rayleigh fading channels considering binary DPSK/NCFSK modulation. Such models available in the literature so far consider only the fading process ignoring the underlying modulation used. Our contribution in this paper is that we consider first-order Markov models for binary DPSK/NCFSK modulation. To derive the Markov model parameters, we first derive expressions for the second-order statistics of the channel error process (specifically, the auto-correlation function of the bit error process as well as the packet error process), and obtain the Markov model parameters, in closed-form, as a function of normalized Doppler bandwidth, average received SNR and packet length. We also verify the accuracy of the proposed Markov model by deriving closed-form expressions for the mutual information of the channel error process Ramesh Annavajjala, Ananthanarayanan Chockalingam, Pamela C. Cosman, Laurence B. Milstein |
ISIT | 3 |
| 2006 | A cross-Layer diversity technique for multicarrier OFDM multimedia networksabstractDiversity can be used to combat multipath fading and improve the performance of wireless multimedia communication systems. In this work, by considering transmission of an embedded bitstream over an orthogonal frequency division multiplexing (OFDM) system in a slowly varying Rayleigh faded environment, we develop a cross-layer diversity technique which takes advantage of both multiple description coding and frequency diversity techniques. More specifically, assuming a frequency-selective channel, we study the packet loss behavior of an OFDM system and construct multiple independent descriptions using an FEC-based strategy. We provide some analysis of this cross-layer approach and demonstrate its superior performance using the set partitioning in hierarchical trees image coder. Yee Sin Chan, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2006 | Drift-resistant SNR scalable video codingabstractWe address the problem of enhancement layer drift estimation for fine granular scalable video. An optimal per-pixel drift estimation algorithm is introduced. The encoder assumes that there is some truncation of the enhancement layer, which does not allow the enhancement layer reference to be properly reconstructed, and the encoder recursively estimates the associated drift and chooses coding modes accordingly. The approach yields performance gains of about 1 dB across low to medium rates. In addition, we investigate dual frame prediction, for both base and enhancement layer, with pulsed-quality allocation in the base Athanasios Leontaris, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2006 | Video coding with fixed-length packetization for a tandem channelabstractA robust scheme is presented for the efficient transmission of packet video over a tandem wireless Internet channel. This channel is assumed to have bit errors (due to noise and fading on the wireless portion of the channel) and packet erasures (due to congestion on the wired portion). First, we propose an algorithm to optimally switch between intracoding and intercoding for a video coder that operates on a packet-switched network with fixed-length packets. Different re-synchronization schemes are considered and compared. This optimal mode selection algorithm is integrated with an efficient channel encoder, a cyclic redundancy check outer coder concatenated with an inner rate-compatible punctured convolutional coder. The system performance is both analyzed and simulated. Last, the framework is extended to operate on a time-varying wireless Internet channel with feedback information from the receiver. Both instantaneous feedback and delayed feedback are evaluated, and an improved method of refined distortion estimation for encoding is presented and simulated for the case of delayed feedback. Yushi Shen, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2006 | Error-Resilient Video Communications Over CDMA Networks With a Bandwidth ConstraintabstractWe present an adaptive video transmission scheme for use in a code-division multiple-access network, which incorporates efficient bandwidth allocation among source coding, channel coding, and spreading under a fixed total bandwidth constraint. We derive the statistics of the received signal, as well as a theoretical bound on the packet drop rate at the receiver. Based on these results, a bandwidth allocation algorithm is proposed at the packet level, which incorporates the effects of both the changing channel conditions and the dynamics of the source content. Detailed simulations are done to evaluate the performance of the system, and the sensitivity of the system to estimation error is presented. Yushi Shen, Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Image Process. | 2 |
| 2006 | Modeling packet-loss visibility in MPEG-2 videoabstractWe consider the problem of predicting packet loss visibility in MPEG-2 video. We use two modeling approaches: CART and GLM. The former classifies each packet loss as visible or not; the latter predicts the probability that a packet loss is visible. For each modeling approach, we develop three methods, which differ in the amount of information available to them. A reduced reference method has access to limited information based on the video at the encoder's side and has access to the video at the decoder's side. A no-reference pixel-based method has access to the video at the decoder's side but lacks access to information at the encoder's side. A no-reference bitstream-based method does not have access to the decoded video either; it has access only to the compressed video bitstream, potentially affected by packet losses. We design our models using the results of a subjective test based on 1080 packet losses in 72 minutes of video. Sandeep Kanumuri, Pamela C. Cosman, Amy R. Reibman, Vinay A. Vaishampayan |
IEEE Trans. Multim. | 2 |
| 2005 | Error Concealment for Dual Frame Video Coding with Uneven QualityabstractWhen losses occur in a transmission of compressed video, the decoder can attempt to conceal the loss by using spatial or temporal methods to estimate the missing macroblocks. We consider a multi-frame error concealment approach which exploits the uneven quality in the two reference frames to provide good concealment candidates. A binary decision tree is used to decide among various error concealment choices. The uneven quality of the reference frames provides an advantage for error concealment. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
DCC | 2 |
| 2005 | Video Coding for a Time Varying Tandem Channel with FeedbackabstractSummary form only given. A robust scheme for the efficient transmission of packet video over a tandem wireless Internet channel is extended to a time varying scenario with a feedback channel. This channel is assumed to have bit errors (due to noise and fading on the wireless portion of the channel) and packet erasures (due to congestion on the wired portion). Simulation results showed that refined estimation can dramatically improve the performance for varying channel conditions, and that combined feedback of both channel conditions and ACK/NACK information can further improve system performance compared with the feedback of just one type of information. Yushi Shen, Pamela C. Cosman, Laurence B. Milstein |
DCC | 2 |
| 2005 | Multimedia communication over OFDM mobile wireless networks: a cross-layer diversity approachabstractDiversity can be used to combat multipath fading and improve the performance of mobile wireless multimedia communication systems. In this work, by considering transmission of an embedded bitstream over a slow varying Rayleigh faded environment, we develop a cross-layer diversity technique which takes advantage of both multiple description source coding and frequency diversity techniques. More specifically, assuming a frequency-selective channel, we study the packet loss behavior of an OFDM system and construct multiple independent descriptions using an FEC-based strategy. We demonstrate the superior performance of this approach using the set partitioning in hierarchical trees (SPIHT) coder. Yee Sin Chan, Pamela C. Cosman, Laurence B. Milstein |
ICC | 2 |
| 2005 | Measuring the added high frequency energy in compressed videoabstractA major focus of video quality assessment research has been to quantify the amount of blocking, blurring, and ringing impairments. However, little attention has been paid to another impairment common in motion-compensated video compression systems: the addition of high frequency (HF) energy as motion compensation moves blocking artifacts off block boundaries. In this paper, we employ an energy-based approach to measure this motion-compensated edge artifact (MCEA) impairment, using both compressed bitstream information and decoded pixels. Experimental results show that we can accurately estimate the percentage of this energy in compressed video. Athanasios Leontaris, Pamela C. Cosman, Amy R. Reibman |
ICIP (2) | 2 |
| 2005 | Optimal allocation of bandwidth for minimum battery consumptionabstractIn general, a power amplifier utilizes battery energy more efficiently with a higher transmission power. For a given message, a given bandwidth constraint and a given performance constraint, different allocations of the bandwidth among source coding, channel coding and modulation result in different amounts of battery usage. We propose a method to optimize the bandwidth allocation and minimize the battery consumption due to transmission. Our results show that with the optimal allocation, significant reduction in battery consumption can be achieved without sacrificing the system performance. Pamela C. Cosman, Laurence B. Milstein |
WCNC | 2 |
| 2005 | On source coding, channel coding and spreading tradeoffs in a DS-CDMA system operating over frequency selective fading channels with narrowband interferenceabstractFor a fixed total bandwidth expansion factor, we consider the problem of optimal bandwidth allocation among the source coder, the channel coder, and the spread-spectrum unit for a direct-sequence code-division multiple-access system operating over a frequency-selective fading channel with narrowband interference. Assuming a Gaussian source with the optimum scalar quantizer, and a binary convolutional code with soft-decision decoding, and further assuming that the self-interference is negligible, we obtain both a lower and an upper bound on the end-to-end average source distortion. The joint three-way constrained optimization of the source code rate, the channel code rate, and the spreading factor can be simplified into an unconstrained optimization problem over two variables. Upon fixing the channel code rate, we show that both upper and lower bound-based distortion functions are convex functions of the source code rate. Because an explicit solution for the optimum source code rate, i.e., one that minimizes the average distortion, is difficult to obtain, computer-based search techniques are employed. Numerical results are presented for the optimum source code rate and spreading factor, parameterized by the channel code rate and code constraint length. The optimal bandwidth allocation, in general, depends on the system and the channel conditions, such as the total number of active users, the average jammer-to-signal power ratio, and the number of resolved multipath components together with their power delay profile. Ramesh Annavajjala, Pamela C. Cosman, Laurence B. Milstein |
IEEE J. Sel. Areas Commun. | 2 |
| 2005 | Joint source/channel coding and MAP decoding of arithmetic codesabstractIn this paper, a novel maximum a posteriori (MAP) estimation approach is employed for error correction of arithmetic codes with a forbidden symbol. The system is founded on the principle of joint source channel coding, which allows one to unify the arithmetic decoding and error correction tasks into a single process, with superior performance compared to traditional separated techniques. The proposed system improves the performance in terms of error correction with respect to a separated source and channel coding approach based on convolutional codes, with the additional great advantage of allowing complete flexibility in adjusting the coding rate. The proposed MAP decoder is tested in the case of image transmission across the additive white Gaussian noise channel and compared against standard forward error correction techniques in terms of performance and complexity. Both hard and soft decoding are taken into account, and excellent results in terms of packet error rate and decoded image quality are obtained. Marco Grangetto, Pamela C. Cosman, Gabriella Olmo |
IEEE Trans. Commun. | 2 |
| 2004 | Dual Frame Motion Compensation with Uneven Quality AssignmentabstractVideo codecs that use motion compensation have shown PSNR gains from the use of multiple frame prediction, in which more than one past reference frame is available for motion estimation. In dual frame motion compensation, one short-term reference frame and one long-term reference frame are available for prediction. In this paper, we propose a dual frame motion compensation technique that allocates bits unevenly among frames to periodically create a high-quality frame that serves as the long-term reference frame for some time. By modifying an MPEG-4 encoder to use this technique on a set of video sequences, we show that it outperforms a normal dual frame motion compensation scheme in which the long-term reference frames are regular frames that are not allocated any extra rate. Vijay Chellappa, Pamela C. Cosman, Geoffrey M. Voelker |
Data Compression Conference | 2 |
| 2004 | Optimal per-pixel estimation for scalable video coding
Athanasios Leontaris, Pamela C. Cosman |
ICIP | 2 |
| 2004 | Visibility of individual packet losses in MPEG-2 videoabstractThe ability of a human to visually detect whether a packet has been lost during the transport of compressed video depends heavily on the location of the packet loss and the content of the video. In this paper, we explore when humans can visually detect the error caused by individual packet losses. Using the results of a subjective test based on 1080 packet losses in 72 minutes of video, we design a classifier that uses objective factors extracted from the video to predict the visibility of each error. Our classifier achieves over 93% accuracy. Amy R. Reibman, Sandeep Kanumuri, Vinay A. Vaishampayan, Pamela C. Cosman |
ICIP | 4 |
| 2004 | Optimal mode selection for a pulsed-quality dual-frame video coderabstractA dual-frame video coder employs two past reference frames for motion compensated prediction. Compared to conventional single frame prediction, the dual-frame encoder can have advantages both in distortion-rate performance and in error resilience. In previous work, it was shown that optimal mode selection can enhance the performance of a dual-frame encoder. In another strand of previous work, it was shown that uneven assignment of quality to frames, to create high-quality (HQ) long-term reference frames, can enhance the performance of a dual-frame encoder. In this letter, we combine these two strands and demonstrate the performance advantages of optimal mode selection among HQ frames for video transmission over noisy channels. Athanasios Leontaris, Vijay Chellappa, Pamela C. Cosman |
IEEE Signal Process. Lett. | 3 |
| 2004 | Optimal allocation of bandwidth for source coding, channel coding, and spreading in CDMA systemsabstractThis paper investigates the tradeoffs between source coding, channel coding, and spreading in code-division multiple-access systems, operating under a fixed total bandwidth constraint. We consider two systems, each consisting of a uniform source with a uniform quantizer, a channel coder, an interleaver, and a direct-sequence spreading module. System A is quadrature phase-shift keyed modulated and has a linear block channel coder. A minimum mean-squared error receiver is also employed in this system. System B is binary phase-shift keyed modulated. Rate-compatible punctured convolutional codes and soft-decision Viterbi decoding are used for channel coding in system B. The two systems are analyzed for both an additive white Gaussian noise channel and a flat Rayleigh fading channel. The performances of the systems are evaluated using the end-to-end mean squared error. A tight upper bound for frame-error rate is derived for nonterminated convolutional codes for ease of analysis of system B. We show that, for a given bandwidth, an optimal allocation of that bandwidth can be found using the proposed method. Pamela C. Cosman, Laurence B. Milstein |
IEEE Trans. Commun. | 2 |
| 2004 | Video compression for lossy packet networks with mode switching and a dual-frame bufferabstractVideo codecs that use motion compensation benefit greatly from the development of algorithms for near-optimal intra/inter mode switching within a rate-distortion framework. A separate development has involved the use of multiple-frame prediction, in which more than one past reference frame is available for motion estimation. In this paper, we show that using a dual-frame buffer (one short-term frame and one long-term frame available for prediction) together with intra/inter mode switching improves the compression performance of the coder. We improve the mode-switching algorithm with the use of half-pel motion vectors. In addition, we investigate the effect of feedback in making more informed and effective mode-switching decisions. Feedback information is used to limit drift errors due to packet losses by synchronizing the long-term frame buffers of both the encoder and the decoder. Athanasios Leontaris, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2003 | Video Compression with Intra/Inter Mode Switching and a Dual Frame BufferabstractVideo codecs that use motion compensation have achieved performance improvements from the use of intra/inter mode switching decisions within a rate-distortion framework. A separate development has involved the use of multiple frame prediction, in which more than one past reference frame is available for motion estimation. It is shown that using a dual frame buffer, together with intra/inter mode switching improves the compression performance of the coder. Also, the mode switching algorithm is improved with the use of half-pel motion vectors. Athanasios Leontaris, Pamela C. Cosman |
DCC | 2 |
| 2003 | Error correction by means of arithmetic codes: an application to resilient image transmissionabstractIn this paper, two novel maximum a posteriori (MAP) estimators for the decoding of arithmetic codes in the presence of transmission errors are presented. Trellis search techniques and a forbidden symbol are employed to obtain forward error correction. The proposed system is applied to lossless image compression and transmission across the BSC; the results are compared in terms of both performance and complexity with a traditional separated source and channel coding approach based on convolutional codes. Marco Grangetto, Gabriella Olmo, Pamela C. Cosman |
ICASSP (4) | 3 |
| 2003 | Human Body Model Acquisition and Tracking Using Voxel Data
Ivana Mikic, Mohan M. Trivedi, Edward Hunter, Pamela C. Cosman |
Int. J. Comput. Vis. | 4 |
| 2003 | Preserving step edges in low bit rate progressive image compressionabstractWith the growing importance of low-bandwidth applications, such as wireless access to the Internet, images are often sent or received at low bit rates. At these bit rates, they suffer from significant distortion and artifacts, making it difficult for those viewing the images to understand them. We present two progressive compression algorithms that focus on preserving the clarity of important image features, such as edges, at compression ratios of 80:1 and more. Both algorithms capture and encode the locations of important edges in the images. The first algorithm then transmits a standard SPIHT (set partitioning in hierarchical trees) bit stream, and at the decoder applies a nonlinear edge-enhancement procedure to improve the clarity of the encoded edges. The second approach uses a modified wavelet transform to "remove" the edges, and encodes the remaining texture information using SPIHT. With both approaches, features in the images that may be important for recognition are well preserved, even at low bit rates. Dirck Schilling, Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2003 | Fast and memory efficient text image compression with JBIG2abstractIn this paper, we investigate ways to reduce encoding time, memory consumption and substitution errors for text image compression with JBIG2. We first look at page striping where the encoder splits the input image into horizontal stripes and processes one stripe at a time. We propose dynamic dictionary updating procedures for page striping to reduce the bit rate penalty it incurs. Experiments show that splitting the image into two stripes can save 30% of encoding time and 40% of physical memory with a small coding loss of about 1.5%. Using more stripes brings further savings in time and memory but the return diminishes. We also propose an adaptive way to update the dictionary only when it has become out-of-date. The adaptive updating scheme can resolve the time versus bit rate tradeoff and the memory versus bit rate tradeoff well simultaneously. We then propose three speedup techniques for pattern matching, the most time-consuming encoding activity in JBIG2. When combined together, these speedup techniques can save up to 75% of the total encoding time with at most 1.7% of bit rate penalty. Finally, we look at improving reconstructed image quality for lossy compression. We propose enhanced prescreening and feature monitored shape unifying to significantly reduce substitution errors in the reconstructed images. Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2003 | Decision trees for error concealment in video decodingabstractWhen macro-blocks are lost in a video decoder such as MPEG-2, the decoder can try to conceal the error by estimating or interpolating the missing area. Many different methods for this type of post-processing concealment have been proposed, operating in the spatial, frequency, or temporal domains, or some hybrid combination of them. In this paper, we show how the use of a decision tree that can adaptively choose among several different error concealment methods can outperform each single method. We also propose two promising new methods for temporal error concealment. Song Cen, Pamela C. Cosman |
IEEE Trans. Multim. | 2 |
| 2003 | End-to-end differentiation of congestion and wireless lossesabstractIn this paper, we explore end-to-end loss differentiation algorithms (LDAs) for use with congestion-sensitive video transport protocols for networks with either backbone or last-hop wireless links. As our basic video transport protocol, we use UDP in conjunction with a congestion control mechanism extended with an LDA. For congestion control, we use the TCP-Friendly Rate Control (TFRC) algorithm. We extend TFRC to use an LDA when a connection uses at least one wireless link in the path between the sender and receiver. We then evaluate various LDAs under different wireless network topologies, competing traffic, and fairness scenarios to determine their effectiveness. In addition to evaluating LDAs derived from previous work, we also propose and evaluate a new LDA, ZigZag, and a hybrid LDA, ZBS, that selects among base LDAs depending upon observed network conditions. We evaluate these LDAs via simulation, and find that no single base algorithm performs well across all topologies and competition. However, the hybrid algorithm performs well across topologies and competition, and in some cases exceeds the performance of the best base LDA for a given scenario. All of the LDAs are reasonably fair when competing with TCP, and their fairness among flows using the same LDA depends on the network topology. In general, ZigZag and the hybrid algorithm are the fairest among all LDAs. Song Cen, Pamela C. Cosman, Geoffrey M. Voelker |
IEEE/ACM Trans. Netw. | 2 |
| 2002 | Image quality evaluation based on recognition times for fast image browsing applicationsabstractMean squared error (MSE) and peak signal-to-noise-ratio (PSNR) are the most common methods for measuring the quality of compressed images, despite the fact that their inadequacies have long been recognized. Quality for compressed still images is sometimes evaluated using human observers who provide subjective ratings of the images. Both SNR and subjective quality judgments, however, may be inappropriate for evaluating progressive compression methods which are to be used for fast browsing applications. In this paper, we present a novel experimental and statistical framework for comparing progressive coders. The comparisons use response time studies in which human observers view a series of progressive transmissions, and respond to questions about the images as they become recognizable. We describe the framework and use it to compare several well-known algorithms (JPEG, set partitioning in hierarchical trees (SPIHT), and embedded zerotree wavelet (EZW)), and to show that a multiresolution decoding is recognized faster than a single large-scale decoding. Our experiments also show that, for the particular algorithms used, at the same PSNR, global blurriness slows down recognition more than do localized "splotch" artifacts. Dirck Schilling, Pamela C. Cosman |
IEEE Trans. Multim. | 2 |
| 2001 | Articulated Body Posture Estimation from Multi-Camera Voxel DataabstractWe present a framework for articulated body model acquisition and tracking from voxel data. A 3D voxel reconstruction of the person's body is computed from silhouettes extracted from four cameras. The model acquisition process is fully automated. In the first frame, body parts are located sequentially. The head is located first, since its shape and size are unique and stable. Other parts are found by sequential template growing and fitting. This initial estimate of body part locations, sizes and orientations is then used as a measurement for the extended Kalman filter which ensures a valid articulated body model. The same filter, with a slightly modified state and state transition matrix, is then used for tracking. The performance of the system has been evaluated on several video sequences with promising results. Ivana Mikic, Mohan M. Trivedi, Edward Hunter, Pamela C. Cosman |
CVPR (1) | 4 |
| 2001 | Feature-Preserving Image Coding for Very Low Bit RatesabstractMany progressive wavelet-based image coders are designed for good performance on natural images. They attempt to achieve the greatest reduction in mean squared error (MSE) with each bit sent, an approach that is most effective when the image is composed chiefly of low-frequency content. Many images, however, include sharp-edged objects, text characters or graphics that are not well handled by standard wavelet-based methods. These features, which may contain information important for recognition, become distorted and obscured when highly compressed by standard wavelet-based methods. In this paper, we present a new progressive image coder that treats an image as being composed of three types of information: edges, texture, and edge-associated detail. The locations of important edges are encoded using line graphic techniques. Texture is encoded using a wavelet-based zerotree approach. Detail near edges-that cannot be efficiently encoded as texture-is encoded separately with a bitplane coding technique. With this approach, features in the image that may be important for recognition are well preserved, even at low bit rates. Dirck Schilling, Pamela C. Cosman |
Data Compression Conference | 2 |
| 2001 | Fast and memory efficient JBIG2 encoderabstractWe propose a fast and memory efficient encoding strategy for text image compression with the JBIG2 standard. The encoder splits up the input image into horizontal stripes and encodes one stripe at a time. Construction of the current dictionary is based on updating dictionaries from previous stripes. We describe separate updating processes for the singleton exclusion dictionary and for the modified-class dictionary. Experiments show that, for both dictionaries, splitting the page into two stripes can save 30% of encoding time and 40% of physical memory with a small loss of about 1.5% in compression. Further gains can be obtained by using more stripes but with diminishing returns. The same updating processes are also applied to compressing multi-page document images and shown to improve compression by 8-10% over coding a multi-page document as a collection of single-page documents. Pamela C. Cosman |
ICASSP | 2 |
| 2001 | Dictionary design for text image compression with JBIG2abstractThe JBIG2 standard for lossy and lossless bilevel image coding is a very flexible encoding strategy based on pattern matching techniques. This paper addresses the problem of compressing text images with JBIG2. For text image compression, JBIG2 allows two encoding strategies: SPM and PM&S. We compare in detail the lossless and lossy coding performance using the SPM-based and PM&S-based JBIG2, including their coding efficiency, reconstructed image quality and system complexity. For the SPM-based JBIG2, we discuss the bit rate tradeoff associated with symbol dictionary design. We propose two symbol dictionary design techniques: the class-based and tree-based techniques. Experiments show that the SPM-based JBIG2 is a more efficient lossless system, leading to 8% higher compression ratios on average. It also provides better control over the reconstructed image quality in lossy compression. However, SPM's advantages come at the price of higher encoder complexity. The proposed class-based and tree-based symbol dictionary designs outperform simpler dictionary formation techniques by 8% for lossless and 16-18% for lossy compression. Pamela C. Cosman |
IEEE Trans. Image Process. | 2 |
| 2000 | Symbol Dictionary Design for the JBIG2 StandardabstractThe JBIG2 standard for lossy and lossless bi-level image coding is a very flexible encoding strategy based on pattern matching techniques. The encoder collects a set of symbols in a dictionary and encodes a page by reference to the dictionary symbols. JBIG2 allows the encoder to view all symbols and choose a good set for the dictionary. We propose a two-pass technique for choosing the dictionary entries, and discuss the trade-offs in bit allocation that these choices entail. The proposed dictionary design technique outperforms simpler dictionary formation techniques by up to 7-9% and 10-18% for lossless and lossy compression, respectively. In the lossy case, we further point out one way to significantly speed up the symbol bitmap compression procedure. Dirck Schilling, Pamela C. Cosman, Hyung Hwa Ko |
Data Compression Conference | 3 |
| 2000 | Moving Shadow and Object Detection in Traffic ScenesabstractWe present an algorithm for segmentation of traffic scenes that distinguishes moving objects from their moving cast shadows. A fading memory estimator calculates mean and variance of all three color components for each background pixel. Given the statistics for a background pixel, simple rules for calculating its statistics when covered by a shadow are used. Then, MAP classification decisions are made for each pixel. In addition to the color features, we examine the use of neighborhood information to produce smoother classification. We also propose the use of temporal information by modifying class a priori probabilities based on predictions from the previous frame. Ivana Mikic, Pamela C. Cosman, Greg T. Kogut, Mohan M. Trivedi |
ICPR | 2 |
| 2000 | Combined forward error control and packetized zerotree wavelet encoding for transmission of images over varying channelsabstractOne method of transmitting wavelet based zerotree encoded images over noisy channels is to add channel coding without altering the source coder. A second method is to reorder the embedded zerotree bitstream into packets containing a small set of wavelet coefficient trees. We consider a hybrid mixture of these two approaches and demonstrate situations in which the hybrid image coder can outperform either of the two building block methods, namely on channels that can suffer packet losses as well as statistically varying bit errors. Pamela C. Cosman, Jon K. Rogers, P. Greg Sherwood, Kenneth Zeger |
IEEE Trans. Image Process. | 1 |
| 2000 | Universal lossless compression via multilevel pattern matchingabstractA universal lossless data compression code called the multilevel pattern matching code (MPM code) is introduced. In processing a finite-alphabet data string of length n, the MPM code operates at O(log log n) levels sequentially. At each level, the MPM code detects matching patterns in the input data string (substrings of the data appearing in two or more nonoverlapping positions). The matching patterns detected at each level are of a fixed length which decreases by a constant factor from level to level, until this fixed length becomes one at the final level. The MPM code represents information about the matching patterns at each level as a string of tokens, with each token string encoded by an arithmetic encoder. From the concatenated encoded token strings, the decoder can reconstruct the data string via several rounds of parallel substitutions. A O(1/log n) maximal redundancy/sample upper bound is established for the MPM code with respect to any class of finite state sources of uniformly bounded complexity. We also show that the MPM code is of linear complexity in terms of time and space requirements. The results of some MPM code compression experiments are reported. John C. Kieffer, En-Hui Yang, Gregory J. Nelson, Pamela C. Cosman |
IEEE Trans. Inf. Theory | 4 |
| 1999 | Decision Trees for Error Concealment in Video DecodingabstractWhen macroblocks are lost in an MPEG decoder, the decoder can try to conceal the error by estimating or interpolating the missing area. Many different methods for this type of concealment have been proposed, operating in the spatial, frequency, or temporal domains, or some hybrid combination of them. We show how the use of a decision tree that can adaptively choose among several different error concealment methods can outperform each single method. We also propose two promising new methods for temporal error concealment. Song Cen, Pamela C. Cosman, Faramarz Azadegan |
Data Compression Conference | 2 |
| 1999 | Edge-Enhanced Image Coding for Low Bit RatesabstractMany current progressive wavelet-based image coders attempt to achieve the greatest reduction in mean squared error (MSE) with each bit sent. In so doing, they tend to send information on the lowest-frequency wavelet coefficients first. At very low bit rates, images compressed by these coders are therefore dominated by low frequency information and blotchy artifacts. These effects combine to hamper recognition of objects in the images. In this paper, we present a new progressive image coder which employs edge enhancement with the goal of improving the visual appearance and recognizability of compressed images at very low bit rates. Important edges in the original image are captured and transmitted as side information together with a traditional wavelet coder bit stream. The decoder combines the two complementary information sources in a manner which, for certain image classes, can yield highly recognizable images at very low bit rates. Dirck Schilling, Pamela C. Cosman |
ICIP (3) | 2 |
| 1999 | Comparison of error concealment strategies for MPEG videoabstractWhen macroblocks are lost in an MPEG decoder, the decoder can try to conceal the error by estimating the missing area. Many different methods for this type of concealment have been proposed. In previous work, we showed how the use of a decision tree adaptively choosing among several different error concealment methods can outperform each single method. In this paper, we improve the decision tree approach, and compare it against the use of concealment motion vectors. Song Cen, Pamela C. Cosman |
WCNC | 2 |
| 1998 | Robust Wavelet Zerotree Image Compression with Fixed-Length PacketizationabstractWe present a novel robust image compression algorithm in which the output of a wavelet zerotree-style coder is manipulated into fixed-length segments. The segments are independently decodable, and errors occurring in one segment do not propagate into any other. The method provides both excellent compression performance and graceful degradation under increasing packet losses. We extend the basic scheme to perform region-based compression, in which specified portions of the image are coded to higher quality with little or no side information required by the decoder. Jon K. Rogers, Pamela C. Cosman |
Data Compression Conference | 2 |
| 1998 | Image Compression for Memory Constrained Printers
Pamela C. Cosman, Tamás Frajka, Kenneth Zeger |
ICIP (3) | 1 |
| 1998 | Memory constrained wavelet based image codingabstractWe present a method for ordering the wavelet coefficient information in a compressed bitstream that allows an image to be sequentially decoded, with lower memory requirements than conventional wavelet decompression schemes. We also introduce a hybrid filtering scheme that uses different horizontal and vertical filters, each with different depths of wavelet decomposition. This reduces decoder memory requirements by reducing the instantaneous number of wavelet coefficients needed for inverse filtering. Pamela C. Cosman, Kenneth Zeger |
IEEE Signal Process. Lett. | 1 |
| 1998 | Wavelet zerotree image compression with packetizationabstractWe describe a combined wavelet zerotree coding and packetization method that provides excellent image compression and graceful degradation against packet erasure. For example, using 53-byte packets (48-byte payload), the algorithm compresses the 512/spl times/512 gray-scale Lena image to 0.2 b/pixel with a peak signal-to-noise ratio (PSNR) of 32.2 dB with no packet erasure, and 26.3 dB on average for 10% packets erased. Jon K. Rogers, Pamela C. Cosman |
IEEE Signal Process. Lett. | 2 |
| 1997 | Image quality in lossy compressed digital mammogramsabstractThe substitution of digital representations for analog images provides access to methods for digital storage and transmission and enables the use of a variety of digital image processing techniques, including enhancement and computer assisted screening and diagnosis. Lossy compression can further improve the efficiency of transmission and storage and can facilitate subsequent image processing. Both digitization (or digital acquisition) and lossy compression alter an image from its traditional form, and hence it becomes important that any such alteration be shown to improve or at least not damage the utility of the image in a screening or diagnostic application. One approach to demonstrating in a quantifiable manner that a specific image mode is at least equal to another is by clinical experiment simulating ordinary practice and suitable statistical analysis. In this paper we describe a general protocol for performing such a verification and present preliminary results of a specific experiment designed to show that 12 bpp digital mammograms compressed in a lossy fashion to 0.015 bpp using an embedded wavelet coding scheme result in no significant differences from the analog or digital originals. Die Ersetzung analoger Bilder durch digitale Darstellungen erlaubt eine digitale Speicherung und Übertragung sowie den Einsatz einer Vielzahl von Methoden der digitalen Bildverarbeitung, z.B. zur Verbesserung der Bildqualität und zum computerunterstützten Screening bzw. zur computerunterstützten Diagnose. Eine verlustbehaftete Kompression kann die Effizienz der Übertragung oder Speicherung weiter steigern und eine nachfolgende Bildverarbeitung erleichtern. Sowohl die Digitalisierung (oder digitale Aufnahme) als auch die verlustbehaftete Kompression ändern ein Bild bezüglich seiner ursprünglichen Form. Deswegen ist es wichtig, zu zeigen daβ eine solche Veränderung die Nützlichkeit des Bildes bei Screening- oder diagnostischen Anwendungen steigert oder wenigstens nicht beeinträchtigt. Eine Möglichkeit, auf quantifizierbare Weise zu zeigen, daβ eine bestimmte Bilddarstellung einer anderen zumindest äquivalent ist, ist ein die gewöhnliche Praxis simulierendes klinisches Experiment und eine geeignete statistische Analyse. In diesem Artikel beschreiben wir ein allgemeines Protokoll für die Durchführung einer solchen Verifikation. Wir präsentieren weiters vorläufige Resultate eines spezifischen Experiments, welches zeigt, daβ die verlustbehaftete Kompression digitaler Mammogramme von 12 bpp auf 0.15 bpp mittels einer eingebetteten Wavelet-Codierung zu keinen signifikanten Unterschieden von den analogen oder digitalen Originalen führt. La substitution d'images analogiques par des représentations numériques donne accès à des méthodes de stockage et de transmission numériques, et permet l'utilisation d'une grande variété de techniques de traitement d'images, incluant le rehaussement, les tests de dépistage assisté ordinateur et le diagnostic. La compression avec pertes peut encore améliorer l'efficacité de la transmission et du stockage, et peut faciliter le traitement ultérieur des images. La numérisation et la compression avec pertes altérant toutes deux une image par rapport à sa forme traditionnelle, il devient important de montrer qu'une telle altération améliore, ou du moins ne réduit pas, l'utilité de l'image dans un screening ou une application de diagnostic. Une approche pour démontrer d'une manière quantifiable qu'un mode d'image spécifique est au moins égal à un autre est l'expérimentation clinique simulant la pratique ordinaire jointe à une analyse statistique adaptée. Dans cet article, nous décrivons un protocole général pour effectuer une telle vérification et présentons les résultats préliminaires d'une expérience faite pour montrer que des mamogrammes numérisés à 12 bpp et comprimés avec pertes à 0.15 bpp à l'aide d'une technique de codage par ondelettes incluses ne présentent pas de différences significatives par rapport aux versions originales analogique ou numérique. Sharon M. Perlmutter, Pamela C. Cosman, Robert M. Gray, Richard A. Olshen, D. Ikeda, C. N. Adams, B. J. Betts, Mark B. Williams, Keren Perlmutter, Jia Li 0001, Anuradha K. Aiyer, Laurie Lee Fajardo, R. Birdwell, B. L. Daniel |
Signal Process. | 2 |
| 1997 | Medical image compression with lossless regions of interest
Jacob Ström, Pamela C. Cosman |
Signal Process. | 2 |
| 1996 | Visual computing education at UCSDabstractOur goal at the University of California, San Diego, is to emphasize the increased importance of visual computing and image engineering by expanding and redesigning our curriculum. The planned curriculum development includes a 1-year undergraduate sequence in image processing and machine perception, and a redesign of the current core graduate sequence to put increased emphasis on synthesis of ideas from separate areas into a core sequence of visual computing. We are also introducing many advanced courses at both the undergraduate and graduate levels. Ramesh Jain 0001, Pamela C. Cosman |
ICIP (1) | 2 |
| 1996 | Vector quantization of image subbands: a surveyabstractSubband and wavelet decompositions are powerful tools in image coding because of their decorrelating effects on image pixels, the concentration of energy in a few coefficients, their multirate/multiresolution framework, and their frequency splitting, which allows for efficient coding matched to the statistics of each frequency band and to the characteristics of the human visual system. Vector quantization (VQ) provides a means of converting the decomposed signal into bits in a manner that takes advantage of remaining inter and intraband correlation as well as of the more flexible partitions of higher dimensional vector spaces. Since 1988, a growing body of research has examined the use of VQ for subband/wavelet transform coefficients. We present a survey of these methods. Pamela C. Cosman, Robert M. Gray, Martin Vetterli |
IEEE Trans. Image Process. | 1 |
| 1995 | Tree-Structured Vector Quantization with Significance Map for Wavelet Image CodingabstractVariable-rate tree-structured VQ is applied to the coefficients obtained from an orthogonal wavelet decomposition. After encoding a vector, we examine the spatially corresponding vectors in the higher subbands to see whether or not they are "significant", that is, above some threshold. One bit of side information is sent to the decoder to inform it of the result. When the higher bands are encoded, those vectors which were earlier marked as insignificant are not coded. An improved version of the algorithm makes the decision not to code vectors from the higher bands based on a distortion/rate tradeoff rather than a strict thresholding criterion. Results of this method on the test image "Lena" yielded a PSNR of 30.15 dB at 0.174 bits per pixel. Pamela C. Cosman, Sharon M. Perlmutter, Keren Perlmutter |
Data Compression Conference | 1 |
| 1995 | Evaluating quality and utility in digital mammographyabstractImage quality and utility become crucial issues for engineers, scientists, patients, regulators, administrators, insurance companies, and lawyers whenever there are changes in the technology by which medical images are produced. Examples of such changes include analog-to-digital conversion, lossy compression for efficient transmission and storage, image enhancement, and computer-aided methodology for diagnosis that affects the appearances of images. This paper is a summary of some principles for designing protocols for clinical experiments to quantify the relative qualities and utilities of different images, here analog, digital, and lossy compressed digital mammograms. A talk supplemented this paper with a status report on the specific experiment described which is scheduled to be conducted during summer 1995. Robert M. Gray, Richard A. Olshen, D. Ikeda, Pamela C. Cosman, Sharon M. Perlmutter, Cheryl L. Nash, Keren Perlmutter |
ICIP | 4 |
| 1994 | Measurement Accuracy as a Measure of Image Quality in Compressed MR Chest ScansabstractWe investigated the effects of lossy image compression on measurement accuracy in magnetic resonance images. Thirty chest scans were compressed to five different levels using predictive pruned tree-structured vector quantization (predictive PTSVQ). Three radiologists measured the diameters of the four principal blood vessels on each image. Errors were analyzed relative to both an independent standard and personal performance on uncompressed images. Data were compared with both t and Wilcoxon tests. We conclude that for the purpose of measuring blood vessels in the chest, there is no significant difference in measurement accuracy when images are compressed up to 16:1 with predictive PTSVQ.> Sharon M. Perlmutter, Chien-Wen Tseng, Pamela C. Cosman, King C. P. Li, Richard A. Olshen, Robert M. Gray |
ICIP (1) | 3 |
| 1994 | Evaluating quality of compressed medical images: SNR, subjective rating, and diagnostic accuracyabstractCompressing a digital image can facilitate its transmission, storage, and processing. As radiology departments become increasingly digital, the quantities of their imaging data are forcing consideration of compression in picture archiving and communication systems (PACS) and evolving teleradiology systems. Significant compression is achievable only by lossy algorithms, which do not permit the exact recovery of the original image. This loss of information renders compression and other image processing algorithms controversial because of the potential loss of quality and consequent problems regarding liability, but the technology must be considered because the alternative is delay, damage, and loss in the communication and recall of the images. How does one decide if an image is good enough for a specific application, such as diagnosis, recall, archival, or educational use? The authors describe three approaches to the measurement of medical image quality: signal-to-noise ratio (SNR), subjective rating, and diagnostic accuracy. They compare and contrast these measures in a particular application, consider in some depth recently developed methods for determining diagnostic accuracy of lossy compressed medical images and examine how good the easily obtainable distortion measures like SNR are at predicting the more expensive subjective and diagnostic ratings. The examples are of medical images compressed using predictive pruned tree-structured vector quantization, but the methods can be used for any digital image processing that produces images different from the original for evaluation.> Pamela C. Cosman, Robert M. Gray, Richard A. Olshen |
Proc. IEEE | 1 |
| 1993 | Using vector quantization for image processingabstractA review is presented of vector quantization, the mapping of pixel intensity vectors into binary vectors indexing a limited number of possible reproductions, which is a popular image compression algorithm. Compression has traditionally been done with little regard for image processing operations that may precede or follow the compression step. Recent work has used vector quantization both to simplify image processing tasks, such as enhancement classification, halftoning, and edge detection, and to reduce the computational complexity by performing the tasks simultaneously with the compression. The fundamental ideas of vector quantization are explained, and vector quantization algorithms that perform image processing are surveyed.> Pamela C. Cosman, Karen L. Oehler, Eve A. Riskin, Robert M. Gray |
Proc. IEEE | 1 |
| 1993 | Tree-structured vector quantization of CT chest scans: image quality and diagnostic accuracyabstractThe authors apply a lossy compression algorithm to medical images, and quantify the quality of the images by the diagnostic performance of radiologists, as well as by traditional signal-to-noise ratios and subjective ratings. The authors' study is unlike previous studies of the effects of lossy compression in that they consider nonbinary detection tasks, simulate actual diagnostic practice instead of using paired tests or confidence rankings, use statistical methods that are more appropriate for nonbinary clinical data than are the popular receiver operating characteristic curves, and use low-complexity predictive tree-structured vector quantization for compression rather than DCT-based transform codes combined with entropy coding. The authors' diagnostic tasks are the identification of nodules (tumors) in the lungs and lymphadenopathy in the mediastinum from computerized tomography (CT) chest scans. Radiologists read both uncompressed and lossy compressed versions of images. For the image modality, compression algorithm, and diagnostic tasks the authors consider, the original 12 bit per pixel (bpp) CT image can be compressed to between 1 bpp and 2 bpp with no significant changes in diagnostic accuracy. The techniques presented here for evaluating image quality do not depend on the specific compression algorithm and are useful new methods for evaluating the benefits of any lossy image processing technique. Pamela C. Cosman, Chien-Wen Tseng, Robert M. Gray, Richard A. Olshen, Lincoln E. Moses, H. Christian Davidson, Colleen J. Bergin, Eve A. Riskin |
IEEE Trans. Medical Imaging | 1 |
| 1992 | Combining Vector Quantization and Histogram EqualizationabstractCombined vector quantization and adaptive histogram equalization Pamela C. Cosman Eve A. Riskin Robert M. Gray tDurand Building, Department of Electrical Engineering Stanford University, Stanford, CA, 94305-4055 Department of Electrical Engineering, FT- 10 University of Washington, Seattle, WA 98195 ABSTRACT Adaptive histogram equalization is a contrast enhancement technique in which each pixel is remapped to an intensity proportional to its rank among surrounding pixels in a selected neighborhood. We present work in which adaptive histogram equalization is performed on the codebook of a tree-structured vector quantizer so that encoding with the resulting codebook performs both compression and contrast enhancement. The algorithm was tested on magnetic resonance brain scans from different subjects and the resulting images were significantly contrast enhanced. 1. INTRODUCTION Histogram equalization refers to a set of contrast enhancement techniques which attempt to spread out the intensity levels occurring in an image over the full available range.1 Histogram equalization is a competitor of interactive intensity windowing, which is the established contrast enhancement technique for medical images. In global histogram equalization, one calculates the intensity histogram for the entire image and then remaps each pixel's intensity proportional to its rank among all the pixel intensities. In adaptive histogram equalization (AHE), the histogram is calculated only for pixels in a context region, usually a square, and the remapping is done for the center pixel of the square. This can be called pointwise histogram equalization because, for each point in the image, one calculates the histogram for the square context region centered on that point. Because this is very computationally intensive, the bilinear interpolative version is an alternative that lowers the computational complexity.2 It calculates the histogram for only a set of non-overlapping context regions that cover the image and the reniapping of pixel intensity values is then exact for only the small number of pixels that are at the centers of these context regions. For all other pixels, a bilinear interpolation from the nearest context region centers determines the appropriate remapping function. With the bilinear interpolative version of AHE, the remapping function for a given pixel of intensity i at location (, y) is determined from the nearest 4 context regions as shown in figure 1. Ifm+_ denotes the mapping at the grid pixel (x+, y.) to the upper right of (x, y), and similar subscripts are used for the other surrounding context regions, then the interpolated AHE result is given by2: in(i) = a[bm(i) + (1 — b)m_(i)J + [1 — u]{bm_(i) + (1 — b)m__(i)], b= here y+—y- O-8194-0805-O/92/$4.QO SPIE Vol. 1653 Image Capture, Formatting, and Display (1992) / 213 Downloaded From: http://proceedings.spiedigitallibrary.org/ on 05/20/2014 Terms of Use: http://spiedl.org/terms Pamela C. Cosman, Eve A. Riskin, Robert M. Gray |
Inf. Process. Manag. | 1 |
| 1991 | Combining Vector Quantization and Histogram EqualizationabstractHistogram equalization is performed on the codebook of a tree-structured vector quantizer. Encoding with the resulting codebook performs both compression and contrast enhancement. It is also possible to perform intensity windowing on the codebook, or a combination of intensity windowing and histogram equalization so that these need not be separate post-processing steps.> Pamela C. Cosman, Eve A. Riskin, Robert M. Gray |
Data Compression Conference | 1 |