Amy R. Reibman

dblp:62/5025 · DBLP profile ↗
← Back
97ranked-venue papers
23as first author
9since 2021 · last 2026
0000-0003-1859-1091ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 79 · 22 first-author · 8 since 2021Computer networks · 13 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Improving Animal Pose Estimation through Species Similarity Measures and Rigorous Label Definition
abstract
Effective image-based analysis of animals, their phenotypes, and their behavior requires accurate localization of key bodyparts. Keypoint detection algorithms can be either generalized, for a large set of animal species, or specialized and targeted to a specific small set of species. In this paper, we explore specialized models, and address two critical aspects of the associated data-label quality: selection of training data and definition of keypoints. Using antelope species as an example, we introduce a variety of species-similarity measures that we apply for selecting relevant training samples, and we demonstrate that training with the automatically selected species leads to improved pose estimation performance while reducing the required number of images. Then, we demonstrate that labeling keypoints with more precise locations leads to improved localization performance that would be valuable for downstream tasks.
Medhashree Parhy, Shaan Chanchani, Claire Kim, Josh Mansky, Zian Pan, Parth Thakre, Haoyu Chen 0006, Amy R. Reibman
WACV8
2024 Minimizing Human Labor for In-the-Wild Camera Trap Processing Pipeline
abstract
Camera traps are an important tool in ecological studies for non-intrusive monitoring of various animals. However, annotating camera trap data usually requires large amount of human labor. Therefore, we propose a solution for practitioners with limited human resources: an automated pipeline for curating in-the-wild camera trap data that mimics human annotators and significantly reduces the amount of human labor needed. We also propose evaluation protocols for estimating system performance on unlabeled data, and present experiments that demonstrate our pipeline's strengths, weaknesses, and flexibility to accommodate users' requirements.
Haoyu Chen 0006, Amy R. Reibman
MMSP2
2024 Shadow Augmentation for Handwashing Action Recognition: From Synthetic to Real Datasets
abstract
Video analytics systems designed for deployment in outdoor conditions can be vulnerable to many environmental changes, particularly changes in shadow. Existing works have shown that shadow and its introduced distribution shift can cause system performance to degrade sharply. In this paper, we explore mitigation strategies to shadow-induced breakdown points of an action recognition system, using the specific application of handwashing action recognition for improving food safety. Using synthetic data, we explore the optimal shadow attributes to be included when training an action recognition system in order to improve performance under different shadow conditions. Exper-imental results indicate that heavier and larger shadow is more effective at mitigating the breakdown points. Building upon this observation, we propose a shadow augmentation method to be applied to real-world data. Results demonstrate the effectiveness of the shadow augmentation method for model training and consistency of its effectiveness across different neural network architectures and datasets.
Shengtai Ju, Amy R. Reibman
MMSP2
2024 Image Quality Assessment in End-to-end Face Analytics Systems
abstract
End-to-end face analytics systems perform tasks such as face detection, alignment, and recognition sequentially, and the performance of each task depends on the success of the preceding tasks. Typically, these systems are deployed in resource and bandwidth-constrained environments. In such systems, resource utilization can be improved by linking the quality of face images processed by these systems to the performance of analytics tasks. This approach ensures that the system processes face images that are best suited for analytics. Recently, the development of face image Quality Estimators (QEs) has received significant attention. However, none of these face image QEs have been evaluated in an end-to-end manner to determine how they affect the overall performance of a face analytics system. In this paper, we utilize a carefully curated dataset to evaluate task-specific face image QEs in the context of an end-to-end face analytics system. Through our experiments, we demonstrate that task-specific QEs are effective in rejecting images not suitable for end-to-end face analytics systems and can significantly improve overall system performance.
Praneet Singh, Amy R. Reibman
MMSP2
2023 Gallery-Query Protocol for Evaluating Face Image Quality Metrics
abstract
As more automated face recognition systems are integrated into society, face image quality estimators (QE) become important. These QEs help quantify whether an input face image contains reliable information necessary for face recognition. However, to assess the effectiveness of face QEs, it is essential to have reliable evaluation protocols. Current face QE evaluation protocols require long computation times and do not have explicit real-world implications. In this paper, we propose a novel face QE evaluation protocol named “Gallery-Query (GQ) Protocol”. The GQ protocol is significantly faster in evaluating face QEs when compared to previous approaches. Furthermore, it has a very explicit real-world use case in constructing an optimal gallery set for face recognition tasks. In addition to this, we used these evaluation protocols to investigate the generalizability of face Quality Estimators (QEs) across various face recognition models, which also has implications for real-world use cases.
Haoyu Chen 0006, Praneet Singh, Edward J. Delp, Amy R. Reibman
MMSP4
2022 Video-Analytics Task-Aware Quad-Tree Partitioning and Quantization for HEVC
abstract
Video analytics systems designed for computer vision tasks use deep learning models that rely on high-quality input data to maximize performance. However, in a real-world system, these inputs are often compressed using video codecs such as HEVC. Video compression degrades the quality of the inputs, thereby degrading the performance of these models. Region-of-interest (ROI) coding enables bits to be allocated to improve performance; however, the method to select regions should be computationally simple since it must occur during or before the video is compressed and transmitted for further processing. In this paper, we propose a task-aware quad-tree (TA-QT) partitioning and quantization method to achieve ROI coding for HEVC and other video coding standards. TA-QT uses a lightweight edge-based model to guide task-aware video encoding to improve end-stage video analytics (ESVA) performance while reducing both bit-rate and encoding time. We demonstrate the effectiveness of our approach in terms of (a) the performance of the ESVA on compressed inputs, (b) transmission bit-rates, and (c) encoding time.
Praneet Singh, Edward J. Delp, Amy R. Reibman
ICIP3
2021 Turkey Behavior Identification Using Video Analytics And Object Tracking
abstract
In this paper, we propose a method to identify behavior of experimental turkeys by automatically analyzing video recordings. Monitoring turkey health during production is crucial for improved turkey production. Turkey health can be reflected through their common behavior, and changes in the frequency and duration of their behavior can be used to detect sick turkeys early. Video recordings can be manually annotated to assist identifying turkey behaviors, but this is both time consuming and labor intensive. In this paper, we monitor and detect changes in turkey behavior using video analytics. Behaviors of interest include eating, drinking, preening, and pecking. Identifying these behaviors requires accurate estimates of turkeys’ and turkey heads’ locations. Re-identification of each turkey is crucial after significant shape deformation such as wing flapping and fast walking. Therefore, our system integrates a state-of-the-art turkey tracker and a head tracker with a behavior identification module to identify turkey behavior. Results demonstrate that our system is effective and accurate at estimating the spatial location of turkeys and their heads, and identifying all behaviors of interest with high recall.
Shengtai Ju, Marisa A. Erasmus, Fengqing Zhu 0001, Amy R. Reibman
ICIP4
2021 Estimating Image Quality for Person Re-Identification
abstract
Task-based image quality is an important component in designing a real-life analytics system, and has been studied in many fields such as biometric recognition and pedestrian detection. Person re-identification (re-id) is an application in surveillance where one attempts to identify a person after they leave the surveillance system and then re-enter it. As a newer field in recognition applications, re-id lacks image quality-related research. In this paper, we propose an unsupervised and automated quality measure for the query images used in re-id, which we call "identifiability". Images that are less identifiable are more challenging for any re-identification system, creating output results that would likely be less reliable. Our proposed method measures feature consistency in the presence of perturbations as an indicator of identifiability. We then introduce two evaluation protocols for such a quality measure. We demonstrate our proposed quality measure is effective at ranking an image’s usefulness to a recognition system, and at identifying an image’s robustness against further compression.
Haoyu Chen 0006, Edward J. Delp, Amy R. Reibman
MMSP3
2021 Special issue on Open Media Compression: Overview, Design Criteria, and Outlook on Emerging Standards
abstract
Universal access to and provisioning of multimedia content is now a reality. It is easy to generate, distribute, share, and consume any multimedia content, anywhere, anytime, or any device. Open media standards took a crucial role toward enabling all these use cases leading to a plethora of applications and services that have now become a commodity in our daily life. Interestingly, most of these services adopt a streaming paradigm, are typically deployed over the open, unmanaged Internet, and account for most of today’s Internet traffic. Currently, the global video traffic is greater than 60% of all Internet traffic[1], and it is expected that this share will grow to more than 80% in the near future[2]. In addition, Nielsen’s law of Internet bandwidth states that the users’ bandwidth grows by 50% per year, which roughly fits data from 1983 to 2019[3]. Thus, the users’ bandwidth can be expected to reach approximately 1 Gb/s by 2022. At the same time, network applications will grow and utilize the bandwidth provided, just like programs and their data expand to fill the memory available in a computer system. Most of the available bandwidth today is consumed by video applications, and the amount of data is further increasing due to already established and emerging applications, e.g., ultrahigh definition, high dynamic range, or virtual, augmented, mixed realities, or immersive media applications in general.
Christian Timmerer, Mathias Wien, Lu Yu 0003, Amy R. Reibman
Proc. IEEE4
2020 Robustness Analysis of Face Obscuration
abstract
Face obscuration is needed by law enforcement and mass media outlets to guarantee privacy. Sharing sensitive content where obscuration or redaction techniques have failed to completely remove all identifiable traces can lead to many legal and social issues. Hence, we need to be able to systematically measure the face obscuration performance of a given technique. In this paper we propose to measure the effectiveness of eight obscuration techniques. We do so by attacking the redacted faces in three scenarios: obscured face identification, verification, and reconstruction. Threat modeling is also considered to provide a vulnerability analysis for each studied obscuration technique. Based on our evaluation, we show that the k-same based methods are the most effective.
Hanxiang Hao, David Guera, János Horváth, Amy R. Reibman, Edward J. Delp
FG4
2019 Video Quality Temporal Pooling using a Visibility Measure
abstract
Both mobile and egocentric videos contain much larger motion than broadcast videos. To estimate the quality of videos with huge motion, the temporal pooling strategy should be adaptive to the specific content. Existing methods focus more on the values and variations of frame quality scores and ignore the masking effect of motion. In this paper, a temporal pooling strategy using a visibility measure is proposed to estimate the quality of videos containing large motion, where the imperceivable details during motion are not considered. We then introduce a strategy to measure the influence of measured visibility on pooling and design a subjective test to gather data for the strategy by synthetically creating shaky videos. Our pooling method is demonstrated to be more effective than existing strategies at pooling frame scores estimated by different image quality metrics.
Amy R. Reibman
ICME2
2019 Hand-hygiene activity recognition in egocentric video
abstract
Food safety is affected by the conditions and practices during different manufacturing steps to prevent contamination and food-borne illnesses. In this paper, we focus on detecting hand-hygiene actions in Egocentric videos. We create a two-stage system to localize and recognize all the hand-hygiene actions in each untrimmed video. In the first stage, we apply a low-cost hand mask and motion histogram features to localize the temporal regions of hand-hygiene actions. In the second stage, we use the two-stream network model combined with a search algorithm to recognize all types of hand-hygiene actions that happen in the untrimmed video. The system achieves a detection accuracy close to 80% on our dataset with 100 participants.
Chengzhang Zhong, Amy R. Reibman, Hansel Mina Cordoba, Amanda J. Deering
MMSP2
2018 Controllable Image Illumination Enhancement with an Over-Enhancement Measure
abstract
The quality of images or videos that suffer from exposure distortions can be enhanced using histogram equalization or retinex methods. However, the relationship between the visual quality and the degree of enhancement is an inverted U-shaped function with a peak point, and many existing methods have parameters that have no clear relationship with image quality. We introduce a controllable illumination enhancement system, where the degree of enhancement can be adjusted using a single parameter. We then propose an over-enhancement measure, Lightness Order Measure (LOM), which quantifies the unnaturalness based on a local inversion of lightness order. We explore the relationship between the peak point and LOM in a subjective test. The results indicate that LOM reduces content dependency compared to existing methods. Our subjective test also evaluates the image quality of our enhancement, and demonstrates the effectiveness of our method.
Amy R. Reibman
ICIP2
2018 Video Classification of Farming Activities with Motion-Adaptive Feature Sampling
abstract
Recently, video has been applied in different industrial applications including autonomous driving vehicles. However, to develop autonomous farming vehicles, the video analysis must be targeted for specific farming activities. So an important first step is to classify the videos into their specific farming activity. In this paper, we propose a video classification framework that includes two branches that process videos differently based on their motions. A gradient-based method is proposed for separating videos into two subsets which are then processed by different feature sampling strategies. The result shows that two motion-based feature sampling strategies provide more efficient features; thus better classification performances are achieved. We also discuss how the feature sampling strategy influences the classification accuracy and the computational efficiency. In addition to farming videos, this proposed system can also be applied to classify videos captured from various camera movements, such as hand-held or first-person cameras.
Amy R. Reibman, Aaron Ault, James V. Krogmeier
MMSP2
2018 Image quality assessment in first-person videos
Amy R. Reibman
J. Vis. Commun. Image Represent.2
2017 Mutual reference frame-quality assessment for first-person videos
abstract
First-person videos (FPVs) captured by wearable cameras are explored for applications of sharing experiences, recording daily lives, measuring social interactions and behaviors. These applications can be improved by an accurate quality assessment. To maximally use the information present in a FPV, we introduce a new strategy for image quality assessment, called mutual reference (MR). MR does not fit into the previous categorization of full-reference, reduced-reference and no-reference. It uses the overlapping content between images to provide effective information for quality estimation. We propose a framework of mutual reference frame-quality assessment for FPVs (MRFQAFPV) to implement the MR strategy based on a MR quality estimator (QE), LVI. The effectiveness of MRFQAFPV is demonstrated in a subjective test by comparing with 3 no-reference QEs and frame-to-frame motion.
Amy R. Reibman
ICIP2
2017 Quality-adaptive deep learning for pedestrian detection
abstract
Pedestrian detection is a fundamental task for many applications including autonomous vehicles and surveillance systems. In a mobile or networked environment bandwidth is limited and adaptive datarate streaming is used. Video compression can introduce significant quality degradation that impacts the accuracy of video analytics. In this paper, we examine the problem of a changing video data-rate and examine how it affects the performance of video analytics, in particular pedestrian detection, using a two-stage quality-adaptive convolutional neural network system. Our experimental results demonstrate that when adaptive data-rate streaming is used, our proposed quality-adaptive approach reduces the miss rate by 20% compared to the baseline detector.
Khalid Tahboub, David Guera, Amy R. Reibman, Edward J. Delp
ICIP3
2017 Accuracy prediction for pedestrian detection
abstract
In this paper, we address the problem of predicting accuracy for pedestrian detection. We want to be able to predict the accuracy of a video analytic method without actually executing the method. We propose the use of texture descriptors and random forests to predict the accuracy of various pedestrian detection methods. Our experimental results demonstrate that using the local binary pattern (LBP) or a bank of Schmid and Gabor filters can capture spatial textural information associated with video quality degradation that can be used to predict accuracy. We also demonstrate how predicting the absolute accuracy can save network and computational resources.
Khalid Tahboub, Amy R. Reibman, Edward J. Delp
ICIP2
2017 Enhancing viewability for first-person videos based on a human perception model
abstract
First-person videos (FPVs) captured by wearable cameras have undesired shakiness because of fast changing views. When existing video stabilization techniques are applied, FPVs are transformed into cinematographic videos, losing the First-person motion information (FPMI) such as the recorder's interests and actions. We propose a system that can enhance viewability of FPVs by stabilizing them while preserving their FPMI. The viewability is charaterized based on a human perception model. Objective tests show that our method has competitive stabilization performance relative to existing video stabilization techniques. And subjective tests show that spectators still experience the FPMI from the resulting videos while shakiness is reduced.
Amy R. Reibman
MMSP2
2016 Characterizing distortions in first-person videos
abstract
First-person videos (FPVs) captured by wearable cameras often contain heavy distortions, including motion blur, rolling shutter artifacts and rotation. Existing image and video quality estimators are inefficient for this type of video. We develop a method specifically to measure the distortions present in FPVs, without using a high quality reference video. Our local visual information (LVI) algorithm measures motion blur, and we combine homography estimation with line angle histogram to measure rolling shutter artifacts and rotation. Our experiments demonstrate that captured FPVs have dramatically different distortions compared to traditional source videos. We also show that LVI is responsive to motion blur, but insensitive to rotation and shear.
Amy R. Reibman
ICIP2
2016 DashCam video compression using historical data
abstract
While dashcam videos (DCVs) are used to document unanticipated situations such as accidents, they are often used to create a training set for vehicle detection and vehicle behavior modeling. This requires a substantial volume of stored DCVs. In this paper, we propose a system that effectively compresses these DCVs by taking advantage of existing historical data. First, our video retrieval and alignment preprocessors construct a reference video for a new DCV based on GPS information and ORB (Oriented FAST and Rotated BRIEF) features. The 3D-HEVC encoder jointly compresses these two videos. Our illumination matching algorithm makes the system robust over different illumination conditions. Our system reduces the bit-rate around 30% when tested over 80 sequences and 3 different scenarios: highway, boulevard and downtown.
Amy R. Reibman
PCS2
2016 Software to Stress Test Image Quality Estimators
abstract
An image quality estimator (QE) can be used to improve the performance of a system, but only if its scores are easily interpretable. In this paper, we present software, entitled “Stress Testing Image Quality Estimators (STIQE)” that systematically explores the performance of a QE, with the goal of enabling users to interpret the QE's scores. Our software allows consistent and reproducible benchmarks of new QEs as they are developed, so the most effective QE for an application can be chosen. We demonstrate that results produced by the software provide new insights into hidden aspects of existing QEs.
Amy R. Reibman
QoMEX2
2016 Full-Reference Video Quality Estimation for Videos With Different Spatial Resolutions
abstract
Full-reference (FR) video quality estimators (QEs) resize either the distorted input video or the reference video to compute the quality when these videos have different spatial resolutions. This resizing operation causes several limitations. Multiscale Image Quality Estimator (MIQE) overcomes those limitations for images, but it does not consider the temporal characteristics of video. In this paper, we develop an FR video QE that integrates MIQE with the motion information to estimate the quality of the distorted video without resampling the reference or the test videos. We also perform subjective tests to compare the proposed algorithm with the existing QEs. In these tests, the reference and the input videos are displayed at their native resolutions. The test results show that the proposed algorithm outperforms other QEs when the reference video and the input video have different spatial resolutions. We have also evaluated the performance of the approach using the Scalable Video Database.
Ali Murat Demirtas, Amy R. Reibman, Hamid Jafarkhani
IEEE Trans. Circuits Syst. Video Technol.2
2014 Full reference video quality estimation for videos with different spatial resolutions
abstract
Full reference video quality estimators (QEs) either resize the input video or the reference video to compute the quality when these videos have different spatial resolutions. This resizing operation causes several limitations. Multiscale Image Quality Estimator (MIQE) [1] overcomes those limitations for images but it does not consider the temporal characteristics of video. In this work, we develop a video quality estimator that integrates MIQE with the motion information to estimate the quality. We also perform subjective tests to compare the proposed algorithm with the existing QEs. Test results show that the proposed algorithm outperforms other QEs.
Ali Murat Demirtas, Amy R. Reibman, Hamid Jafarkhani
ICIP2
2014 Full-Reference Quality Estimation for Images With Different Spatial Resolutions
abstract
Multimedia communication is becoming pervasive because of the progress in wireless communications and multimedia coding. Estimating the quality of the visual content accurately is crucial in providing satisfactory service. State of the art visual quality assessment approaches are effective when the input image and reference image have the same resolution. However, finding the quality of an image that has spatial resolution different than that of the reference image is still a challenging problem. To solve this problem, we develop a quality estimator (QE), which computes the quality of the input image without resampling the reference or the input images. In this paper, we begin by identifying the potential weaknesses of previous approaches used to estimate the quality of experience. Next, we design a QE to estimate the quality of a distorted image with a lower resolution compared with the reference image. We also propose a subjective test environment to explore the success of the proposed algorithm in comparison with other QEs. When the input and test images have different resolutions, the subjective tests demonstrate that in most cases the proposed method works better than other approaches. In addition, the proposed algorithm also performs well when the reference image and the test image have the same resolution.
Ali Murat Demirtas, Amy R. Reibman, Hamid Jafarkhani
IEEE Trans. Image Process.2
2013 Image quality estimation for different spatial resolutions
abstract
State of the art visual quality assessment methods are effective when the input image and the reference image have the same resolution. However, estimating the quality of an image that has spatial resolution different than that of the reference image is still a challenging problem. In this work, we design a quality estimator (QE) to estimate the quality of a distorted image with a lower resolution compared to the reference image. We also present a subjective test environment to explore the success of the proposed algorithm in comparison with other QEs. The subjective tests demonstrate that the proposed method works better than other approaches.
Ali Murat Demirtas, Amy R. Reibman, Hamid Jafarkhani
ICIP2
2013 A probabilistic pairwise-preference predictor for image quality
abstract
Current image quality estimators (QEs) compute a single score to estimate the perceived quality of a single input image. When comparing image quality between two images with such a QE, one only knows which image has a higher score; there is no knowledge about the uncertainty of these scores or what fraction of viewers might actually prefer the image with the lower score. In this paper, we present a Probabilistic Pairwise Preference Predictor (P4) that estimates the probability that one image will be preferred by a random viewer relative to a second image. We train a multilevel Bayesian logistic regression model using results from a large-scale subjective test and present the degree to which various factors influence subjective quality. We demonstrate our model provides well-calibrated estimates of pairwise image preferences using a validation set comprising pairs with 60 reference images outside the training set.
Amy R. Reibman, Kenneth Shirley, Chao Tian 0002
ICIP1
2013 Perceptual Visual Signal Compression and Transmission
abstract
One- and two-way communication with digital compressed visual signals is now an integral part of the daily life of millions. Such commonplace use has been realized by decades of advances in visual signal compression. The design of effective, efficient compression and transmission strategies for visual signals may benefit from proper incorporation of human visual system (HVS) characteristics. This paper overviews psychophysics and engineering associated with the communication of visual signals. It presents a short history of advances in perceptual visual signal compression, and describes perceptual models and how they are embedded into systems for compression and transmission, both with and without current compression standards.
Hong Ren Wu, Amy R. Reibman, Weisi Lin, Sheila S. Hemami
Proc. IEEE2
2012 A strategy to jointly test image quality estimators subjectively
abstract
We present an automated algorithm to design subjective tests that have a high likelihood of finding misclassification errors in many image quality estimators (QEs). In our algorithm, a collection of existing QEs collaboratively determine the best pairs of images that will test the accuracy of each individual QE.We demonstrate that the resulting subjective test provides valuable information regarding the accuracy of the cooperating QEs. The proposed strategy is particularly useful for comparing efficacy of QEs across multiple distortion types and multiple reference images.
Amy R. Reibman
ICIP1
2012 An automatic grid corner extraction technique for camera calibration
abstract
Camera calibration is essential for many computer vision and image processing applications. However, this calibration process can be rather time consuming and may require a significant amount of human intervention. Calibration models traditionally employ a calibration grid whose four corner points must be marked by hand on a per-frame basis. The objective of this work is to develop a technique for processing these frames rapidly, with as little human intervention as possible. We propose an algorithm to extract the boundaries of the calibration grid automatically, based on a spectral analysis of HD (high-definition) video frames. The accuracy of the intrinsic parameters estimated using our automatic method is evaluated through comparison with those obtained using a method that requires hand labeling of the corner points.
Lixia Yang, Chao Tian 0002, Vinay A. Vaishampayan, Amy R. Reibman
ICIP4
2011 Systematic stress testing of image quality estimators
abstract
We present a methodology to systematically stress objective image quality estimators (QEs). Using computational results instead of expensive subjective tests, we obtain rigorous information of a QE's performance on a constrained but comprehensive set of degraded images. Our process quantifies many of a QE's potential vulnerabilities. Knowledge of these weaknesses can be used to improve a QE during its design process, to assist in selecting which QE to deploy in a real system, and to interpret the results of a chosen QE once deployed.
Frank M. Ciaramello, Amy R. Reibman
ICIP2
2011 Performance of H.264 with isolated bit error: Packet decode or discard?
abstract
During wireless video transmission, when an error is detected at the receiver, the packet is dropped immediately and error concealment is performed. However, this procedure does not always provide satisfactory results, especially when the motion is complicated or scene change occurs. Moreover, there are often uncorrupted macroblocks in the distorted packet. Hence, it can be important to utilize the information in the distorted packet. In this work, our aim is to explore which produces better quality: discarding or decoding the corrupted packet. First, we examine the effect of an error in each syntax element on the visual quality. Second, we compare the visual degradation caused by an error in each syntax element with that caused by discarding the whole packet. We also evaluate the effects of motion and quantization.
Ali Murat Demirtas, Amy R. Reibman, Hamid Jafarkhani
ICIP2
2010 No-reference image and video quality estimation: Applications and human-motivated design
Sheila S. Hemami, Amy R. Reibman
Signal Process. Image Commun.2
2010 A Versatile Model for Packet Loss Visibility and its Application to Packet Prioritization
abstract
In this paper, we propose a generalized linear model for video packet loss visibility that is applicable to different group-of-picture structures. We develop the model using three subjective experiment data sets that span various encoding standards (H.264 and MPEG-2), group-of-picture structures, and decoder error concealment choices. We consider factors not only within a packet, but also in its vicinity, to account for possible temporal and spatial masking effects. We discover that the factors of scene cuts, camera motion, and reference distance are highly significant to the packet loss visibility. We apply our visibility model to packet prioritization for a video stream; when the network gets congested at an intermediate router, the router is able to decide which packets to drop such that visual quality of the video is minimally impacted. To show the effectiveness of our visibility model and its corresponding packet prioritization method, experiments are done to compare our perceptual-quality-based packet prioritization approach with existing Drop-Tail and Hint-Track-inspired cumulative-MSE-based prioritization methods. The result shows that our prioritization method produces videos of higher perceptual quality for different network conditions and group-of-picture structures. Our model was developed using data from high encoding-rate videos, and designed for high-quality video transported over a mostly reliable network; however, the experiments show the model is applicable to different encoding rates.
Ting-Lan Lin, Sandeep Kanumuri, Yuan Zhi, David Poole 0003, Pamela C. Cosman, Amy R. Reibman
IEEE Trans. Image Process.6
2009 Perceptual quality based packet dropping for generalized video GOP structures
abstract
Our work builds a general visibility model of video packets which is applicable to various types of GOP (group of pictures). The data used for analysis and building the model come from three subjective experiment sets with different encoding and decoding parameters on H.264 and MPEG-2 videos. We consider factors not only within a packet but also across its vicinity to account for possible temporal and spatial masking effects. This model can be useful for an intermediate router in a congested network to drop less visible packets to maintain overall video quality. Experiments are done to compare our perceptual-quality-based packet dropping approach with existing drop-tail and hint-track-inspired cumulative-MSE-based dropping methods. The result shows that our dropping method produces videos of higher perceptual quality for different network conditions and GOP structures.
Ting-Lan Lin, Yuan Zhi, Sandeep Kanumuri, Pamela C. Cosman, Amy R. Reibman
ICASSP5
2009 Video Outage Detection: Algorithm and evaluation
abstract
We present a video outage detection algorithm (VODA) that detects catastrophic failures in video systems by mimicking human behavior. VODA uses a continuity detector for video, audio, and motion, along with blackness, silence, and stillness detection. An outage is declared when all individual features exhibit a sudden drop; an alarm is generated when an outage lasts over two seconds. We analyze the performance of VODA on a large corpus of movie and TV episodes, and we show it has good performance when applied to 160 hours of recorded broadcast TV.
Amy R. Reibman, Allan R. Wilks
PCS1
2008 Perceptual impact of burthy versus isolated packet losses in H.264 compressed video
abstract
When video packets are lost in congested networks, one loss pattern creates a different visual impact than another. We conduct a subjective experiment with H.264 videos and conclude that isolated losses are better than bursty losses in terms of perceptual video quality. A network-implementable video quality model is developed for a router to drop packets so as to achieve good visual quality.
Ting-Lan Lin, Pamela C. Cosman, Amy R. Reibman
ICIP3
2008 A no-reference Spatial Aliasing Measure for digital image resizing
abstract
We present a no-reference spatial aliasing measure (SAM) for images, that assesses the impact of visual jaggy edges in resized digital images. The measure makes use of the fact that visible spatial aliasing is likely to occur in areas with strong directional energy. In these regions, aliasing energy often appears in frequency regions that have little image content. We show the effectiveness of SAM using a subjective test.
Amy R. Reibman, Shan Suthaharan
ICIP1
2007 Characterizing packet-loss impairments in compressed video
abstract
We examine metrics to predict the visibility of packet losses in MPEG-2 and H.264 compressed video. We use subjective data that has a wide range of parameters, including different error concealment strategies and different compression standards. We evaluate SSIM, MSE, and a slice-boundary mismatch (SBM) metric for their effectiveness at characterizing packet-loss impairments.
Amy R. Reibman, David Poole 0003
ICIP (5)1
2007 Capacity analysis of MediaGrid: a P2P IPTV platform for fiber to the node (FTTN) networks
abstract
This paper studies the conditions under which P2P sharing can increase the capacity of IPTV services over FTTN networks. For a typical FTTN network, our study shows a) P2P sharing is not beneficial when the total traffic in a local video office is low; b) P2P sharing increases the load on FTTN switches and routers in local video offices; c) P2P sharing is the most beneficial when the network bottleneck is experienced in the southbound segment of a local video office (equivalently a northbound segment of an FTTN switch); and d) sharing among all FTTN serving communities is not needed when network congestion problems are solved by using some other technologies such as program pre-caching or replication. Based on the analytical results, design for IPTV services which monitors FTTN network conditions and decides when and how to share videos among peers to maximize the service capacity. Simulations and bounds both validate the potential benefits of the MediaGrid IPTV service platform.
Yennun Huang, Yih-Farn Robin Chen, Rittwik Jana, Hongbo Jiang 0001, Michael Rabinovich, Amy R. Reibman, Bin Wei 0003
IEEE J. Sel. Areas Commun.6
2007 Quality Evaluation of Motion-Compensated Edge Artifacts in Compressed Video
abstract
Little attention has been paid to an impairment common in motion-compensated video compression: the addition of high-frequency (HF) energy as motion compensation displaces blocking artifacts off block boundaries. In this paper, we employ an energy-based approach to measure this motion-compensated edge artifact, using both compressed bitstream information and decoded pixels. We evaluate the performance of our proposed metric, along with several blocking and blurring metrics, on compressed video in two ways. First, ordinal scales are evaluated through a series of expectations that a good quality metric should satisfy: the objective evaluation. Then, the best performing metrics are subjectively evaluated. The same subjective data set is finally used to obtain interval scales to gain more insight. Experimental results show that we accurately estimate the percentage of the added HF energy in compressed video.
Athanasios Leontaris, Pamela C. Cosman, Amy R. Reibman
IEEE Trans. Image Process.3
2006 Predicting H.264 Packet Loss Visibility using a Generalized Linear Model
abstract
We consider modeling the visibility of individual and multiple packet losses in H.264 videos. We propose a model for predicting the visibility of multiple packet losses and demonstrate its performance on dual losses (two nearby packet losses). We extract the factors affecting visibility using a reduced-reference method. We predict the probability that a loss is visible using a generalized linear model. We achieve MSE values (between actual and predicted probabilities) of 0.0253 and 0.0398 for individual and dual losses respectively. We also examine the effect of various factors on visibility.
Sandeep Kanumuri, Sitaraman G. Subramanian, Pamela C. Cosman, Amy R. Reibman
ICIP4
2006 Quality assessment for super-resolution image enhancement
abstract
A typical image formation model for super-resolution (SR) introduces blurring, aliasing, and added noise. The enhancement itself may also introduce ringing. In this paper, we use subjective tests to assess the visual quality of SR-enhanced images. We then examine how well some existing objective quality metrics can characterize the observed subjective quality. Even full-reference metrics like MSE and SSIM do not always capture visual quality of SR images with and without residual aliasing.
Amy R. Reibman, Robert M. Bell, Sharon Gray
ICIP1
2006 Modeling packet-loss visibility in MPEG-2 video
abstract
We consider the problem of predicting packet loss visibility in MPEG-2 video. We use two modeling approaches: CART and GLM. The former classifies each packet loss as visible or not; the latter predicts the probability that a packet loss is visible. For each modeling approach, we develop three methods, which differ in the amount of information available to them. A reduced reference method has access to limited information based on the video at the encoder's side and has access to the video at the decoder's side. A no-reference pixel-based method has access to the video at the decoder's side but lacks access to information at the encoder's side. A no-reference bitstream-based method does not have access to the decoded video either; it has access only to the compressed video bitstream, potentially affected by packet losses. We design our models using the results of a subjective test based on 1080 packet losses in 72 minutes of video.
Sandeep Kanumuri, Pamela C. Cosman, Amy R. Reibman, Vinay A. Vaishampayan
IEEE Trans. Multim.3
2005 Comparison of blocking and blurring metrics for video compression
abstract
We evaluate the performance on compressed video of a number of available similarity, blocking, and blurring quality metrics. Using a systematic, objective framework based on simple subjective comparisons, we evaluate the ability of each metric to rank images correctly according to the subjective impact of differences in spatial content, quantization parameters, amounts of filtering, distances from the most recent I-frame, and long-term frame prediction strategies. The indicated weaknesses of available metrics can be used as guides in the development of future quality metrics.
Athanasios Leontaris, Amy R. Reibman
ICASSP (2)2
2005 Measuring the added high frequency energy in compressed video
abstract
A major focus of video quality assessment research has been to quantify the amount of blocking, blurring, and ringing impairments. However, little attention has been paid to another impairment common in motion-compensated video compression systems: the addition of high frequency (HF) energy as motion compensation moves blocking artifacts off block boundaries. In this paper, we employ an energy-based approach to measure this motion-compensated edge artifact (MCEA) impairment, using both compressed bitstream information and decoded pixels. Experimental results show that we can accurately estimate the percentage of this energy in compressed video.
Athanasios Leontaris, Pamela C. Cosman, Amy R. Reibman
ICIP (2)3
2005 Multiple Description Coding for Video Delivery
abstract
Multiple description coding (MDC) is an effective means to combat bursty packet losses in the Internet and wireless networks. MDC is especially promising for video applications where retransmission is unacceptable or infeasible. When combined with multiple path transport (MPT), MDC enables traffic dispersion and hence reduces network congestion. This work describes principles in designing MD video coders employing temporal prediction and presents several predictor structures that differ in their tradeoffs between mismatch-induced distortion and coding efficiency. The paper also discusses example video communication systems integrating MDC and MPT.
Yao Wang 0001, Amy R. Reibman, Shunan Lin
Proc. IEEE2
2004 Visibility of individual packet losses in MPEG-2 video
abstract
The ability of a human to visually detect whether a packet has been lost during the transport of compressed video depends heavily on the location of the packet loss and the content of the video. In this paper, we explore when humans can visually detect the error caused by individual packet losses. Using the results of a subjective test based on 1080 packet losses in 72 minutes of video, we design a classifier that uses objective factors extracted from the video to predict the visibility of each error. Our classifier achieves over 93% accuracy.
Amy R. Reibman, Sandeep Kanumuri, Vinay A. Vaishampayan, Pamela C. Cosman
ICIP1
2004 Quality monitoring of video over a packet network
abstract
We consider monitoring the quality of compressed video transmitted over a packet network from the perspective of a network service provider. Our focus is on no-reference methods, which do not access the original signal, and on evaluating the impact of packet losses on quality. We present three methods to estimate mean squared error (MSE) due to packet losses directly from the video bitstream. NoParse uses only network-level measurements (like packet loss rate), QuickParse extracts the spatio-temporal extent of the impact of the loss, and FullParse extracts sequence-specific information including spatio-temporal activity and the effects of error propagation. Our simulation results with MPEG-2 video subjected to transport packet losses illustrate the performance possible using the three methods.
Amy R. Reibman, Vinay A. Vaishampayan, Yegnaswamy Sermadevi
IEEE Trans. Multim.1
2003 Low complexity quality monitoring of MPEG-2 video in a network
abstract
We consider monitoring the quality of compressed video transmitted over a packet network from the perspective of a network service provider. We estimate the mean-squared error (MSE) caused by packet loss, by examining only the received video bitstream. We merge the best aspects of two of our previously reported methods, with the goal of creating a low-complexity quality monitor that will be able to process many streams simultaneously in the network.
Amy R. Reibman, Vinay A. Vaishampayan
ICIP (3)1
2003 Quality monitoring for compressed video subjected to packet loss
abstract
We consider monitoring the quality of compressed video transmitted over a packet network from the perspective of a network service provider. We estimate the mean-squared error (MSE) caused by packet loss by examining only the received video bitstream. We describe our fullparse method, which extracts sequence-specific information including spatio-temporal activity and the effects of error propagation. We show that fullparse performs well with manageable complexity.
Amy R. Reibman, Vinay A. Vaishampayan
ICME1
2003 Scalable video coding with managed drift
abstract
Traditional scalable video encoders sacrifice coding efficiency to reduce error propagation because they have avoided using enhancement-layer (EL) information to predict the base layer (BL) to prevent the error propagation termed "drift". Drift can produce very poor video quality if left unchecked. We propose a video coder with significantly better compression efficiency because it intentionally allows the drift produced by predicting the BL from the EL. Our drift management system balances the tradeoff between compression efficiency and error propagation. The proposed scalable coder uses a spatially adaptive procedure that optimally selects key encoder parameters: the quantizer and the prediction strategy. Our numerical results indicate the encoder is very powerful, and the selection procedure is effective. The video quality of our coder at low rates is only marginally worse than the drift-free case, while its overall compression efficiency is not much worse than a one-layer nonscalable encoder.
Amy R. Reibman, Léon Bottou, Andrea Basso 0001
IEEE Trans. Circuits Syst. Video Technol.1
2002 Personalized multimedia services using a mobile service platform
abstract
In this paper we address the research issues in providing personalized multimedia services, which enable a mobile user to remotely record video programs, control cameras, and request the delivery of pre-recorded or live video content to his or her own mobile device. We describe a mobile service platform that authenticates users who send service requests from various mobile devices, transcodes video content based on user and device profiles, and authorizes the delivery of content from a media server to the proper client device. The media server adapts automatically to the fluctuations of the wireless channel conditions for reasonable viewing on the client device. The mobile service platform essentially manages the control path, while the media server handles the actual content delivery. We discuss various aspects of the integration and report our successful experiments conducted on wireless LAN and CDPD networks.
Yih-Farn Robin Chen, Huale Huang, Rittwik Jana, Sam John, Serban Jora, Amy R. Reibman, Bin Wei 0003
WCNC6
2002 Introduction to the special issue on wireless communication
Chang Wen Chen, Reginald L. Lagendijk, Amy R. Reibman, Wenwu Zhu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2002 Multiple-description video coding using motion-compensated temporal prediction
abstract
We propose multiple description (MD) video coders which use motion-compensated predictions. Our MD video coders utilize MD transform coding and three separate prediction paths at the encoder to mimic the three possible scenarios at the decoder: both descriptions received or either of the single descriptions received. We provide three different algorithms to control the mismatch between the prediction loops at the encoder and decoder. We present simulation results comparing the three approaches to two standards-based approaches to MD video coding. We show that when the main prediction loop at the encoder uses a two-channel reconstruction, it is important to have side prediction loops and transmit some redundancy information to control mismatch. We also examine the performance of our MD video coder with partial mismatch control in the presence of random packet loss, and demonstrate a significant improvement compared to more traditional approaches.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard, Rohit Puri
IEEE Trans. Circuits Syst. Video Technol.1
2001 Managing Drift in DCT-Based Scalable Video Coding
abstract
When compressed video is transmitted over erasure-prone channels, errors will propagate whenever temporal or spatial prediction is used. Typical tools to combat this error propagation are packetization, re-synchronizing codewords, intra-coding, and scalability. In recent years, the concern over so-called "drift" has sent researchers toward structures for scalability that do not use enhancement-layer information to predict base-layer information and hence have no drift. In this paper, we propose alternative structures for scalability that use previous enhancement-layer information to predict the current base layer, while simultaneously managing the resulting possibility of drift. These structures allow better compression efficiency, while introducing only limited impairments in the quality of the reconstruction.
Amy R. Reibman, Léon Bottou
Data Compression Conference1
2001 DCT-based scalable video coding with drift
abstract
Scalable video coders have traditionally avoided using enhancement-layer (EL) information to predict the base layer (BL), so as to avoid so-called "drift". As a result, they are less efficient than a one-layer coder. Fine granularity scalable (FGS) coders avoid using EL information to predict the EL as well, suffering even further inefficiencies. In this paper, we explore a scalable video coder that allows drift, by predicting the BL from EL information. However, we show that through careful management of the amount of drift introduced, the video quality at low rates is only marginally worse than the drift-free case, while the overall compression efficiency is not much worse than a one-layer encoder.
Amy R. Reibman, Léon Bottou, Andrea Basso 0001
ICIP (2)1
2001 Multiple description video using rate-distortion splitting
abstract
We consider a simple multiple description (MD) video coder, that uses redundancy-rate-distortion criteria to split a one-layer stream generated by a standard video coder into two correlated streams. Our simulation results demonstrate that this MD coder has much better performance for large redundancies than our previous MDTC video coder, although it cannot perform as well at low redundancies. This MD video coder is very simple to implement and is compatible with H.263 to the extent that each description can be decoded by a standard H.263 decoder. This MD coder was used in a previous study on the transport of MD and layered video over an EGPRS wireless network, where the fact that it creates two streams with very balanced rates was a strong advantage.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard
ICIP (1)1
2001 Multiple description coding using pairwise correlating transforms
abstract
The objective of multiple description coding (MDC) is to encode a source into multiple bitstreams supporting multiple quality levels of decoding. In this paper, we only consider the two-description case, where the requirement is that a high-quality reconstruction should be decodable from the two bitstreams together, while lower, but still acceptable, quality reconstructions should be decodable from either of the two individual bitstreams. This paper describes techniques for meeting MDC objectives in the framework of standard transform-based image coding through the design of pairwise correlating transforms. The correlation introduced by the transform helps to reduce the distortion when only a single description is received, but it also increases the bit rate beyond that prescribed by the rate-distortion function of the source. We analyze the relation between the redundancy (i.e., the extra bit rate) and the single description distortion using this transform-based framework. We also describe an image coder that incorporates the pairwise transform and show its redundancy-rate-distortion performance for real images.
Yao Wang 0001, Michael T. Orchard, Vinay A. Vaishampayan, Amy R. Reibman
IEEE Trans. Image Process.4
2000 Transmission of Multiple Description and Layered Video over an EGPRS Wireless Network
abstract
We investigate the ability of multiple descriptions (MD) and layered coding to improve the quality of video transmitted over EGPRS networks. One-layer video sent over a single channel on such a network has a fairly sharp quality transition, depending on a user's location. Either the video can be transmitted reliably (if the video rate is less than or equal to what the channel can sustain), or it is subjected to many lost packets. In this system, MD and layered video may offer two ways to improve the video quality beyond that of the one-layer video. First, each sub-stream can be sent on a separate channel, essentially doubling the assigned bandwidth and increasing the video quality. Second, MD and layered video are more error resilient than one-layer video, potentially improving the video quality seen by users in poor locations. We find that for the system scenarios considered, one and two-layer coding outperform MD coding, depending upon the number of wireless channels used for the video transport.
Amy R. Reibman, Yao Wang 0001, Xiaoxin Qiu, Zhimei Jiang, Kapil K. Chawla
ICIP1
2000 Error-resilient transcoding for video over wireless channels
abstract
We describe a method to maintain quality for video transported over wireless channels. The method is built on three fundamental blocks. First, we use a transcoder that injects spatial and temporal resilience into an encoded bitstream. The amount of resilience is tailored to the content of the video and the prevailing error conditions, as characterized by bit error rate. Second, we derive analytical models that characterize how corruption propagates in a video that is compressed using motion-compensated encoding and subjected to bit errors. Third, we use rate distortion theory to compute the optimal allocation of bit rate among spatial resilience, temporal resilience, and source rate. Furthermore, we use the analytical models to generate the resilience rate distortion functions that are used to compute the optimal resilience. The transcoder then injects this optimal resilience into the bitstream. Simulation results show that using a transcoder to optimally adjust the resilience improves video quality in the presence of errors while maintaining the same input bit rate.
Gustavo de los Reyes, Amy R. Reibman, Shih-Fu Chang, Justin C.-I. Chuang
IEEE J. Sel. Areas Commun.2
1999 Performance of multiple description coders on a real channel
abstract
We explore the ability of multiple description (MD) source coders to achieve good performance on channels other than ideal MD channels. We examine both the overall system design and compare the performance of a system with MD source coder to that of a more traditional system using a layered source coder. For the memoryless channels we consider, MD source coding cannot achieve acceptable performance for a memoryless Gaussian source without appropriate channel coding. Also, in memoryless channels, a system with MD source coding outperforms a layered source coding system only in very poor channels. The introduction of memory in the channel degrades the performance of both systems equally. Using interleaving to reduce the impact of memory in the channel has more influence on performance than the choice of source coder.
Amy R. Reibman, Hamid Jafarkhani, Michael T. Orchard, Yao Wang 0001
ICASSP1
1999 Multiple Description Coding for Video Using Motion Compensated Prediction
abstract
We propose multiple description (MD) video coders which use motion compensated predictions. Our MD video coders utilize MD transform coding and three separate prediction paths at the encoder, to mimic the three possible scenarios at the decoder: both descriptions received or either of the single descriptions received. We provide three different algorithms to control the mismatch between the prediction loops at the encoder and decoder. The results show that when the main prediction loop is the central loop, it is important to have side prediction loops and transmit some redundancy information to control mismatch.
Amy R. Reibman, Hamid Jafarkhani, Yao Wang 0001, Michael T. Orchard, Rohit Puri
ICIP (3)1
1999 Issues of Quality and Multiplexing When Smoothing Rate Adaptive Video
abstract
We have proposed a smoothing and rate adaptation algorithm-SAVE (Smoothed Adaptive Video over Explicit rate networks)-for transport of compressed video over rate-controlled networks. SAVE attempts to preserve quality as much as possible, and exercises control over the source rate only when essential to prevent unacceptable delay. In order to understand the impact on quality of rate adaptation, we have evolved the quality metrics typically used to evaluate the efficacy of mechanisms to transport video. We investigate the dynamic nature of rate reduction: any prolonged impairment is likely to be noticeable. We study the sensitivity of SAVE to its parameters and network characteristics. Finally, the utility of the proposed scheme is measured by its ability to multiplex a large number of streams effectively. Our evaluations are based on experiments with 20 traces of entertainment videos using different compression algorithms.
Nick G. Duffield, K. K. Ramakrishnan, Amy R. Reibman
IEEE Trans. Multim.3
1999 Modeling one- and two-layer variable bit rate video
abstract
This paper presents a source model for VBR video traffic. A finite-state Markov chain is shown to accurately model one- and two-layer video of all activity levels on a per source basis. Our model captures the source dynamics, including the short-term correlations essential for studying network performance. The modeling technique is shown to be applicable for both H.261 and MPEG2 encoded video of a variety of activity levels. The traffic model is shown in a simulation study to be able to accurately characterize both the single-source buffer occupancy over a wide range of buffer sizes and the multiplexing behavior. The VBR video model is also used to model the enhancement layer of two-layer SNR scalable video. We show that two-layer encoding has significantly better statistical multiplexing gains than one-layer video, particularly when the network admits calls based on a leaky-bucket characterization.
Kavitha Chandra, Amy R. Reibman
IEEE/ACM Trans. Netw.2
1998 On combining watermarking with perceptual coding
abstract
A watermark is a data stream inserted into multimedia content. It contains information relevant to the ownership or authorized use of the content. A watermark which could be recovered without a priori knowledge of the identity of the content could be used by Web search mechanisms to flag unauthorized distribution of the content. Since media will be compressed on these sites, a mark detection algorithm that operated in the compressed domain would be useful. We describe a watermark algorithm which operates in the compressed domain and does not require a reference.
Jack Lacy, Schuyler R. Quackenbush, Amy R. Reibman, David H. Shur, James H. Snyder
ICASSP3
1998 Video Transcoding for Resilience in Wireless Channels
abstract
We describe a method to maintain an acceptable quality for video transported over wireless networks under time-varying conditions. We use a transcoder to modify the resilience of the encoded bitstream by using source coding techniques to provide the appropriate level of resilience for the prevailing channel conditions. We develop a statistical model for image loss versus resilience and wireless conditions. Simulation results indicate that using a transcoder to adjust the resilience can improve video quality when errors occur without significantly sacrificing quality when there are no errors. Also, simulation results compare favorably to the analytical model.
Gustavo de los Reyes, Amy R. Reibman, Justin C.-I. Chuang, Shih-Fu Chang
ICIP (1)2
1998 Optimal Pairwise Correlating Transforms for Multiple Description Coding
abstract
Multiple description coding (MDC) addresses the problem of encoding a source into two (or more) bitstreams such that a high-quality reconstruction is decodable from the two bitstreams together, while a lower, but still acceptable, quality reconstruction is decodable if either of the two bitstreams is lost. Recent research has proposed using transforms to introduce a controlled amount of correlation between the two bitstreams in order to achieve MDC objectives. This paper considers several optimality issues related to such transform based MDC methods. Redundancy rate-distortion (RRD) performance of a general class of transforms is derived and used to identify the optimal transform for achieving any given amount of redundancy. Then, the paper introduces a more general transform-based MDC framework incorporating both the transform mode of redundancy and a second mode of redundancy. The optimal allocation of redundancy among these two modes is analyzed.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman
ICIP (1)3
1998 SAVE: An Algorithm for Smoothed Adaptive Video over Explicit Rate Networks
abstract
Supporting compressed video efficiently on networks is a challenge because of its burstiness. Although a large number of applications using compressed video are rate adaptive, it is also important to preserve quality as much as possible. We propose a smoothing and rate adaptation algorithm, called SAVE, that the compressed video source uses in conjunction with explicit rate based control in the network. SAVE smoothes the demand from the source to the network, thus helping achieve good multiplexing gains. SAVE maintains the quality of the video and ensures that the delay at the source buffer does not exceed a bound. We examine the effectiveness of SAVE across 28 different traces (entertainment and teleconferencing videos) using different compression algorithms.
Nick G. Duffield, K. K. Ramakrishnan, Amy R. Reibman
INFOCOM3
1998 VBR video: tradeoffs and potentials
abstract
The authors examine the transport and storage of video compressed with a variable bit rate (VBR). They focus primarily on networked video, although they also briefly consider other applications of VBR video, including satellite transmission (channel sharing), playback of stored video, and wireless transport. Packet video research requires careful integration between the network and the video systems; however, a major stumbling block has resulted because commonly used terms are often interpreted differently by the video and networking communities. The paper then, has two main goals: (i) to clarify the definitions of terms that are often used with different meaning by networking and video-coding researchers and (ii) to explore the tradeoffs entailed by each of the various modalities of VBR transmission (unconstrained, shaped, constrained, and feedback). In particular, they evaluate the tradeoff among the advantages (better video quality, less delay, and more calls) that were identified by early proponents of VBR video transmission. An underlying theme of this paper is that increased interaction between the video and network design has potential for improving overall decoded video quality without changing the network capacity.
T. V. Lakshman, Antonio Ortega, Amy R. Reibman
Proc. IEEE3
1998 SAVE: an algorithm for smoothed adaptive video over explicit rate networks
abstract
Supporting compressed video efficiently on networks is a challenge because of its burstiness. Although a large number of applications using compressed video allow adaptive rates, it is also important to preserve quality as much as possible. We propose a smoothing and rate adaptation algorithm for compressed video, called SAVE, that is used in conjunction, with explicit rate based control in the network. SAVE smooths the demand from the source to the network, thus helping achieve good multiplexing gains. SAVE maintains the quality of the video and ensures that the delay at the source buffer does not exceed a bound. We show that SAVE is effective by demonstrating its performance across 28 different traces (entertainment and teleconferencing videos) that use different compression algorithms.
Nick G. Duffield, K. K. Ramakrishnan, Amy R. Reibman
IEEE/ACM Trans. Netw.3
1997 Redundancy Rate-Distortion Analysis Of Multiple Description Coding Using Pairwise Correlating Transforms
abstract
The objective of multiple description coding (MDC) is to encode a source into two (or more) bitstreams supporting two quality levels of decoding. A high-quality reconstruction should be decodable from the two bitstreams together, while lower, but still acceptable, quality reconstructions should be decodable from either of the two individual bitstreams. This paper describes techniques for meeting MDC objectives in the framework of standard transform-based image coding through the design of pairwise transforms.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman, Vinay A. Vaishampayan
ICIP (1)3
1997 Multiple description image coding for noisy channels by pairing transform coefficients
abstract
Multiple description coding (MDC) is a way of trading off coding gain with robustness to channel errors. This paper presents a new method for MDC using the framework of transform coding. Instead of using the Karhunen-Loeve transform (KLT) that decorrelates all the coefficients, we choose the transform bases so that the coefficients are correlated pair-wise. This is accomplished by rotating every two basis vectors in the KLT. Each pair of correlated coefficients are then split between two descriptions. Only 45/spl deg/ rotation is considered which leads to two balanced streams. In the actual implementation, the DCT is employed in place of the KLT and the rotation of transform bases is accomplished by rotating the DCT coefficients. Experimental results show that this method can lead to satisfactory image reconstruction from any one description with a relatively small (20% for "lena") overhead over a standard JPEG coder.
Yao Wang 0001, Michael T. Orchard, Amy R. Reibman
MMSP3
1997 Joint Selection of Source and Channel Rate for VBR Video Transmission Under ATM Policing Constraints
abstract
Variable bit-rate (VBR) transmission of video over ATM networks has long been said to provide substantial benefits, both in terms of network utilization and video quality, when compared with conventional constant bit-rate (CBR) approaches. However, realistic VBR transmission environments will certainly impose constraints on the rate that each source can submit to the network. We formalize the problem of optimizing the quality of the transmitted video by jointly selecting the source rate (number of bits used for a given frame) and the channel rate (number of bits transmitted during a given frame interval). This selection is subject to two sets of constraints, namely, (1) the end-to-end delay has to be constant to allow for real-time video display and (2) the transmission rate has to be consistent with the traffic parameters negotiated by user and network. For a general class of constraints, including such popular ones as the leaky bucket, we introduce an algorithm to find the optimal solution to this problem. This algorithm allows us to compare VBR and CBR under the same end-to-end delay constraints. Our results indicate that variable-rate transmission can increase the quality of the decoded sequences without increases in the end-to-end delay. Finally, we show that for the leaky-bucket channel, the channel constraints can be combined with the buffer constraints, such that the system is identical to CBR transmission with an additional, infrequently imposed constraint. Therefore, video quality with a leaky-bucket channel can achieve the same quality of a CBR channel with larger physical buffers, without adding to the physical delay in the system.
Chi-Yuan Hsu, Antonio Ortega, Amy R. Reibman
IEEE J. Sel. Areas Commun.3
1996 Forward error control for MPEG-2 video transport in a wireless ATM LAN
abstract
The performance of error control based on forward error correction (FEC) for MPEG-2 video transmission in an indoor wireless ATM LAN is studied. A multipath fading model is used to investigate the effect of errors on video transport. Combined source and channel coding techniques that employ single layer and scalable MPEG-2 coding to combat channel errors are compared. Simulation results indicate that FEC-based error control in combination with 2-layer video coding techniques can lead to acceptable quality for indoor wireless ATM video.
Ender Ayanoglu, Pramod Pancha, Amy R. Reibman, Shilpa Talwar
ICIP (2)3
1996 Forward Error Control for MPEG-2 Video Transport in a Wireless ATM LAN
Ender Ayanoglu, Pramod Pancha, Amy R. Reibman, Shilpa Talwar
Mob. Networks Appl.3
1996 Packet loss resilience of MPEG-2 scalable video coding algorithms
abstract
Transmission of compressed video over packet networks with nonreliable transport benefits when packet loss resilience is incorporated into the coding. One promising approach to packet loss resilience, particularly for transmission over networks offering dual priorities such as ATM networks, is based on layered coding which uses at least two bitstreams to encode video. The base-layer bitstream, which can be decoded independently to produce a lower quality picture, is transmitted over a high priority channel. The enhancement-layer bitstream(s) contain less information, so that packet losses are more easily tolerated. The MPEG-2 standard provides four methods to produce a layered video bitstream: data partitioning, signal-to-noise ratio scalability, spatial scalability, and temporal scalability. Each was included in the standard in part for motivations other than loss resilience. This paper compares the performance of these techniques (excluding temporal scalability) under various loss rates using realistic length material and discusses their relative merits. Nonlayered MPEG-2 coding gives generally unacceptable video quality for packet loss ratios of 10/sup -3/ for small packet sizes. Better performance can be obtained using layered coding and dual-priority transmission. With data partitioning, cell loss ratios of 10/sup -4/ in the low-priority layer are definitely acceptable, while for SNR scalable encoding, cell loss ratios of 10/sup -3/ are generally invisible. Spatial scalable encoding can provide even better visual quality under packet losses; however, it has a high implementation complexity.
Rangarajan Aravind, M. Reha Civanlar, Amy R. Reibman
IEEE Trans. Circuits Syst. Video Technol.3
1995 Post processing transform coded images using edges
abstract
Discrete cosine transform coding, a popular image compression strategy, results in two visible artifacts: blocking and ringing. These are both high frequency artifacts. Since images contain high frequency information the artifacts are removed using a space varying low pass filter as a post processor. Low frequency blocks and flat regions of blocks containing a strong edge are filtered. Low frequency blocks are identified in the transform coefficient domain; edge blocks are identified in the spatial domain. This does not require any alterations in the compressed bit stream. Improvement is demonstrated both subjectively and objectively.
William E. Lynch, Amy R. Reibman, Bede Liu
ICASSP2
1995 Video transport in wireless ATM
abstract
Wireless ATM LANs have the potential to support multi-Mb/s bandwidths to mobile users with guaranteed quality of service. However, the lossy nature of the wireless medium will pose problems for loss-sensitive applications. Techniques to minimize the effect of these losses will therefore be required. In this paper, we examine the use of combined source and channel coding for MPEG video transport in a wireless ATM environment. Although forward error correction (FEC) provides protection against channel bit errors, the bandwidth overhead can become a significant drawback in a fixed bandwidth scenario. Additional protection against losses can be realized by using two-layer video coding. In this work, we compare the performance of a 1-layer main profile and 2-layer data partitioning and SNR scalable MPEG-2 encoders in a system with random channel errors and forward error correction. The results indicate that if the channel bit error rate is known an optimum FEC level can be chosen for the 1-layer case; However, at this fixed FEC level, if a critical bit error rate is exceeded then video quality degrades dramatically. The 2-layer cases appear to lead to more graceful degradation in quality at this critical bit error rates. In particular, SNR scalability may lead to better video quality over a larger range of bit error rates than a 1-layer approach.
Ender Ayanoglu, Pramod Pancha, Amy R. Reibman
ICIP (3)3
1995 An error concealment algorithm for images subject to channel errors
abstract
We present an algorithm to conceal bit errors in still images and image sequences that are coded using the discrete cosine transform (DCT) and variable length codes (VLCs). No modification is necessary to an existing encoder, and no additional bit rate is required. The concealment algorithm is kept simple so that real-time decoding and concealment is possible. A single bit error in these images can cause a block to split into several blocks or several blocks to merge into one. This causes the DCT coefficients of all subsequent blocks to be correctly decoded but stored in the wrong location in the image. Furthermore, the DC coefficient of all subsequent blocks may be incorrect. The error concealment algorithm uses transform domain information to identify the location of the affected blocks and to correct errors. The image quality after error concealment is shown to be significantly improved.
Wai-Man Lam, Amy R. Reibman
IEEE Trans. Image Process.2
1995 An adaptive congestion control scheme for real time packet video transport
abstract
We show that modulating the source rate of a video encoder based on congestion signals from the network has two major benefits: the quality of the video transmission degrades gracefully when the network is congested and the transmission capacity is used efficiently. Source rate modulation techniques have been used in the past in designing fixed rate video encoders used over telephone networks. In such constant bit rate encoders, the source rate modulation is done using feedback information about the occupancy of a local buffer. Thus, the feedback information is available instantaneously to the encoder. In the scheme proposed, the feedback may be delayed by several frames because it comes from an intermediate switching node of a packet switched network. The paper shows the proposed scheme performs quite well despite this delay in feedback. We believe the use of such schemes will simplify the architecture used for supporting real time video services in future nationwide gigabit networks.
Hemant Kanakia, Partho Pratim Mishra, Amy R. Reibman
IEEE/ACM Trans. Netw.3
1995 Traffic descriptors for VBR video teleconferencing over ATM networks
abstract
This paper examines the problem of video transport over ATM networks using knowledge of both video system design and broadband networks. The following issues are addressed: video system delay caused by internal buffering, traffic descriptors (TD) for video, and call admission. We find that while different video sequences require different TD parameters, the following trends hold for all sequences examined. First, increasing the delay in the video system decreases the necessary peak rate and significantly increases the number of calls that can be carried by the network. Second, as an operational traffic descriptor for video, the leaky-bucket algorithm appears to be superior to the sliding-window algorithm. And finally, with a delay in the video system, the statistical multiplexing gain from VBR over CBR video is upper bounded by roughly a factor of four, and to obtain a gain of about 2.0 can require the operational traffic descriptor to have a window or bucket size on the order of a thousand cells. We briefly discuss how increasing the complexity of the video system may enable the size of the bucket or window to be reduced.>
Amy R. Reibman, Arthur W. Berger
IEEE/ACM Trans. Netw.1
1994 Edge Compensated Transform Coding
abstract
Transform based image compression has difficulty with image regions containing edges. Edge compensated transform coding (ECTC) addresses this problem by preprocessing to remove edges. This preprocessing is adapted to transform coding. The edge information is sent in a side channel and the edges are replaced at the receiver. Subjective improvement is demonstrated.>
William E. Lynch, Amy R. Reibman, Bede Liu
ICIP (1)2
1994 Multiplexing of variable rate encoded streams
abstract
Discusses the problem of multiplexing several variable rate encoded streams into a single stream. Moreover, in order to facilitate editing of the resulting stream it is required that data generated during the same time interval is multiplexed together. Particular emphasis is placed on controlling encoder rates and combining data in such a way as to avoid overflow and underflow of buffers at encoder and decoder. Applications include satellite or cable transmission of a fixed number of different video channels, multimedia presentations with multiple video streams, and video on demand.>
Barry G. Haskell, Amy R. Reibman
IEEE Trans. Circuits Syst. Video Technol.2
1994 Methods for performance evaluation of VBR video traffic models
abstract
Models for predicting the performance of multiplexed variable bit rate video sources are important for engineering a network. However, models of a single source are also important for parameter negotiations and call admittance algorithms. In this paper we propose to model a single video source as a Markov renewal process whose states represent different bit rates. We also propose two novel goodness-of-fit metrics which are directly related to the specific performance aspects that we want to predict from the model. The first is a leaky bucket contour plot which can be used to quantify the burstiness of any traffic type. The second measure applies only to video traffic and measures how well the model can predict the compressed video quality.>
David M. Lucantoni, Marcel F. Neuts, Amy R. Reibman
IEEE/ACM Trans. Netw.3
1993 Variable bit-rate video coding for ATM and broadcast applications
Barry G. Haskell, Amy R. Reibman
ICASSP (1)2
1993 Recovery of lost or erroneously received motion vectors
Wai-Man Lam, Amy R. Reibman, Bede Liu
ICASSP (5)2
1993 An Adaptive Congestion Control Scheme for Real-Time Packet Video Transport
abstract
In this paper we show that modulating the source rate of a video encoder based on feedback information from the network results in graceful degradation in picture quality during periods of congestion. Such source rate modulation techniques have been used in the past in designing video encoders used to generate data at a fixed rate. In such constant bit rate encoders, the source rate modulation is done using feedback information about the occupancy of a local buffer. Since the buffer is local, the feedback information is available instantaneously to the encoder. In the proposed scheme, the feedback information is delayed because it comes from within a packet-switching network. The feedback provides information about the traffic at switches along the path of the video connection. We show that the proposed scheme performs well, even though the information is delayed for relatively large intervals of time. We believe that the use of such schemes will simplify the architecture used for supporting real-time services in future nationwide gigabit networks.
Hemant Kanakia, Partho Pratim Mishra, Amy R. Reibman
SIGCOMM3
1993 Design of quantizers for decentralized estimation systems
abstract
The authors consider parameter estimation in decentralized systems with distributed processors. They restrict the local processors to be quantizers and consider the optimal design of the systems to minimize the estimation error. They present necessary conditions for the optimal system based on the Bayes distortion functions and Fisher's information. The numerical results compare the resulting quantizers obtained by different distortion criteria.>
Wai-Man Lam, Amy R. Reibman
IEEE Trans. Commun.2
1992 Self-synchronizing variable-length codes for image transmission
abstract
Variable-length codes are vulnerable to loss of synchronization if they are transmitted through a noisy channel. In order to achieve synchronization regardless of the preceding synchronization slippage, the concept of extended synchronizing codewords is introduced. An iterative algorithm is proposed to design extended synchronizing codewords efficiently. The extended synchronizing codewords are placed in fixed positions of an image to control error propagation due to synchronization slippage. However, the image can still be partially corrupted by channel noises. An error concealment method to further improve the image quality is discussed.>
Wai-Man Lam, Amy R. Reibman
ICASSP2
1992 TES Modeling for Analysis of a Video Multiplexer
Duan-Shin Lee, Benjamin Melamed, Amy R. Reibman, Bhaskar Sengupta
Perform. Evaluation3
1992 Constraints on variable bit-rate video for ATM networks
abstract
Constraints on the encoded bit rate of a video signal that are imposed by a channel and encoder and decoder buffers are considered. Conditions that ensure that the video encoder and decoder buffers do not overflow or underflow when the channel can transmit a variable bit rate are presented. Using these conditions and a commonly proposed network-user contract, the effect of a (BISDN) network policing function on the allowable variability in the encoded video bit rate is examined. It is shown how these ideas might be implemented in a system that controls both the encoded and transmitted bit rates. The performance of video that has been encoded using the derived constraints for the leaky bucket channel is presented.>
Amy R. Reibman, Barry G. Haskell
IEEE Trans. Circuits Syst. Video Technol.1
1991 Asymptotic quantization for signal estimation with noisy channels
abstract
The authors consider signal estimation using a communication system which consists of a quantizer, a noisy channel, and a decision device. By assuming that the lengths of the quantization intervals become arbitrarily small, they derive the asymptotic signal estimation error for the communication system. The asymptotic error consists of an interaction term which depends on both the quantization error and the channel error. The most significant property of the interaction term is that it can be negative. Next, the authors use the companding approach to find the optimal uniform quantizer. They then suggest using a suboptimal piecewise uniform quantizer to approximate the optimal nonuniform quantizer. Finally, they present several analytic and numerical results to justify the validity of the obtained asymptotic expression.>
Wai-Man Lam, Amy R. Reibman
ICASSP2
1991 DCT-based embedded coding for packet video
Amy R. Reibman
Signal Process. Image Commun.1
1990 Hough transform and signal detection theory performance for images with additive noise
Douglas J. Hunt, Loren W. Nolte, Amy R. Reibman, W. Howard Ruedger
Comput. Vis. Graph. Image Process.3
1989 Performance of large distributed detection networks and the sign detector
abstract
A comparison is made of the performance of fusion and serial distributed networks and the centralized system when there are many local sensors. For the fusion network, the author compares its asymptotic efficacy to an approximation of the probability of error that is shown to predict performance more accurately when the signal is not vanishingly small. This result also applies to the sign detector, which is special case of the fusion network. For the serial network, no analytic performance equation is available for general noise. However, for Laplacian noise, the author shows that the performance is bounded as the number of processors increases. She compares the performance of the two networks numerically, using the framework of detecting a known signal in generalized Gaussian noise. The fusion network always performs better than the serial network for many sensors. The fusion network performs best in Laplacian noise and worst in nearly uniform noise. The serial network performs best in nearly uniform noise and worst in Laplacian noise.>
Amy R. Reibman
ICASSP1
1988 Performance of redundant and distributed detection systems with processor faults
abstract
Examines the performance of signal detectors when processors can have faults. The fault tolerance of two systems, the redundant system and the distributed fusion network, are examined. The optimal fault-tolerant design for each system is presented. The performance of the fault-tolerant systems is compared to the faulty simplex processor. In the absence of processor failures, the redundant system is identical to the simplex processor. However, when failures occur, the redundant system performs better than the simplex processor. The fault-free fusion network does not perform as well as the fault-free simplex processor. However, in the cases examined, when failures occur in systems with many (N> 12) channels, the fusion network performs better than the simplex processor.>
Amy R. Reibman, Loren W. Nolte
ICASSP1