John G. Apostolopoulos

dblp:42/979 · DBLP profile ↗
← Back
111ranked-venue papers
17as first author
11since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 69 · 14 first-author · 1 since 2021Computer networks · 17 · 1 first-author · 1 since 2021Theory of computation · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2023 Streaming Erasure Codes over Multicast Relayed Networks
abstract
This paper studies streaming erasure codes in a relayed multicast setting, where a source wishes to transmit a sequence of messages to two different destinations through a common relay. Our construction extends previously proposed works on the single-destination setting studied in Fong et al. and Facenda et al. to the multicast setting, where each destination can recover the source packets with a correspondingly different delay. A key property of our construction is that it does not require prior knowledge of the maximum number of erasures on the relay-destination link. Instead, it enables the recovery of the source stream with a decoding delay that depends on the number of erasures on the relay-destination links. We demonstrate that if some divisibility conditions are satisfied, then the proposed construction can simultaneously achieve the single-destination delay in Fong et al. for both receivers. Finally we also explain how our proposed construction can be applied in the setting of a single destination when the maximum number of erasures on the relay-destination link is not known beforehand.
Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ISIT4
2023 Deep Reinforcement Learning for Latency-Sensitive Communication With Adaptive Redundant Retransmissions
abstract
This paper studies packet repetition strategies over erasure channels with memory and a long feedback delay. The problem is initially formulated as a communications problem where a source wishes to transmit one message packet to a destination while minimizing both the delay and the number of transmissions. At each time instant, the sender is provided a delayed acknowledgement feedback about past attempts, and must decide whether to attempt a new transmission or not. This problem is then re-formulated as an episodic reinforcement learning problem, where an agent attempts to learn the optimal transmission policy, provided delayed feedback about past transmission attempts. The agent is helped by a channel estimator, which attempts to capture the channel memory and use that to predict probabilities of erasures in a future window. This channel estimator is also data-driven and learns the channel model without anya priorichannel knowledge. The paper presents a lower bound on the achievable trade-off between delay and number of transmissions for any channel modeled as a Markov process. Experimental results show that the combination of the proposed channel estimator and the agent can noticeably outperform naive strategies for channels with memory, and achieves results close to the lower bound.
Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Commun.4
2023 Streaming Erasure Codes Over Multi-Access Relayed Networks
abstract
Many emerging multimedia streaming applications involve multiple users communicating under strict latency constraints. In this paper we study streaming codes for a network involving two source nodes, one relay node and a destination node. In this paper’s setting, each source node transmits a stream of messages, through the relay, to a destination, who is required to decode the messages under a strict delay constraint. For the case of a single source node, a class of streaming codes has been proposed by Fong et al., using the concept of delay-spectrum. The current paper presents a novel framework, which constructs streaming codes for a relayed multi-user setting by sequentially constructing the codes for each link. This requires a characterization of the set of all achievable delay spectra for a given rate, blocklength and number of erasures, beyond the specific choice considered by Fong et al. This characterization is presented in the paper for systematic codes. Using this novel framework, the first proposed scheme involves greedily selecting the rate on the link from relay to destination and using properties of the delay-spectrum to find feasible streaming codes that satisfy the required delay constraints. A closed form expression for the achievable rate region is provided, and conditions for when the proposed scheme is optimal are established by a natural outer bound. The second proposed scheme builds upon this approach, but uses a numerical optimization-based approach to improve the achievable rate region over the first scheme. Experimental results show that the proposed schemes achieve significant improvements over baseline schemes based on single-user codes.
Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory5
2023 Adaptive Relaying for Streaming Erasure Codes in a Three Node Relay Network
abstract
This paper investigates adaptive streaming codes over a three-node relayed network. In this setting, a source node transmits a sequence of message packets to a destination with help of a relay. The source-to-relay and relay-to-destination links are unreliable and introduce at most$N_{1}$and$N_{2}$packet erasures, respectively. The destination node must recover each message packet within a strict delay constraint$T$. The paper presents a new construction of streaming codes for all feasible parameters$\{N_{1}, N_{2}, T\}$. Our work improves upon the construction in Fong et al. by adapting the relaying strategy based on the erasure patterns from source to relay. Specifically, the code employs the notion of symbol estimates, which allows the relay to forward information about symbols before it can decode that symbol, and variable-rate encoding, which decreases the rate used to encode a packet as more erasures affect that packet. The codes proposed in this paper achieve rates higher than the ones proposed by Fong et al. whenever$N_{2} > N_{1}$, and achieve the same rate when$N_{2} \leq N_{1}$, in which case the rate is optimal. The paper also presents an upper bound on the achievable rate that takes into account erasures in both links in order to bound the rate in the second link. The upper bound is shown to be tighter than a trivial bound that considers only the erasures in the second link.
Gustavo Kasper Facenda, M. Nikhil Krishnan, Elad Domanovitz, Silas L. Fong, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory7
2022 On State-Dependent Streaming Erasure Codes over the Three-Node Relay Network
abstract
This paper investigates low-latency adaptive streaming codes for a three-node relay network. A source node transmits a sequence of source packets (messages) to the destination through a relay node. We focus on a particular case where the link connecting the source and relay nodes is almost reliable, but the link connecting the relay to the destination is not. The relay node can observe the erasure pattern that has occurred in the transmission between the source node and itself and adapt its relaying strategy based on that observation. Every source packet must be perfectly recovered by the destination with a strict delay T, as long as the number of erasures in the relay-to-destination link lies below some design parameter. We then characterize capacity as a function of such design parameter. The achievability scheme employs two different relaying strategies, based on whether an erasure has or has not occurred in the link from source to relay. The converse is proven by analyzing a periodic erasure pattern and lower bounding the minimum redundancy across channel packets. We show that the achievable rate can be improved compared to non-adaptive schemes previously proposed, indicating that exploiting the knowledge of the erasure pattern by the relay node is essential in achieving capacity.
Gustavo Kasper Facenda, Elad Domanovitz, M. Nikhil Krishnan, Ashish Khisti, Silas L. Fong, Wai-tian Tan, John G. Apostolopoulos
ISIT7
2022 State-Dependent Symbol-Wise Decode and Forward Codes Over Multihop Relay Networks
abstract
This paper studies low-latency streaming codes for the multi-hop network. The source transmits a sequence of messages to a destination through a chain of relays, and requires the destination to reconstruct each message by its deadline. We assume that each communication link is subjected to a certain maximum number of packet erasures. The case of a single relay (a three-node network) was considered in Fong et al. (2020). A coding scheme known as symbol-wise decode and forward was proposed. In the present work, we propose an alternative scheme that is different from Fong et al. (2020) and still achieves the same rate as in Fong et al. (2020) for the one hop case as the field-size goes to infinity. Furthermore, our proposed scheme naturally generalizes to the case of multiple-relay nodes yielding new achievable rates for this setting. The main difference with Fong et al. (2020) is that our proposed scheme exploits the ability of the relay nodes to adapt the transmission based on the erasures on the previous link. Hence, we refer to our scheme as “state-dependent” and contrast it with the scheme in Fong et al. (2020) that is state-independent. Our scheme requires the relay nodes to append a header to the transmitted packets, and we show that the size of the header does not depend on the field-size of the code. We also derive an upper bound on the maximal streaming rate achievable over a network with an arbitrary number of relays. We show that this upper bound matches our achievable rate in the special case when the maximal number of erasures on the first link is greater than or equal to the maximal number of erasures on each of the following links, and the field size goes to infinity.
Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory5
2022 Corrections to "Optimal Streaming Erasure Codes Over the Three-Node Relay Network"
abstract
In the above article[1], an upper bound on the maximum achievable rate is corrected. If we add the restriction that the packets transmitted by the relay must be independent of the erasures introduced by the first-hop channel, then no correction is needed and the converse proof need not be rectified.
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory6
2022 Low-Latency Network-Adaptive Error Control for Interactive Streaming
abstract
We introduce a novel network-adaptive algorithm that is suitable for alleviating network packet losses for low-latency interactive communications between a source and a destination. Our network-adaptive algorithm estimates in real-time the best parameters of a recently proposed streaming code that uses forward error correction (FEC) to correct both arbitrary and burst losses, which cause a crackling noise and undesirable jitters, respectively in audio. In particular, the destination estimates appropriate coding parameters based on its observed packet loss pattern and sends them back to the source for updating the underlying code. Besides, a new explicit construction of practical low-latency streaming codes that achieve the optimal tradeoff between the capability of correcting arbitrary losses and the capability of correcting burst losses is used. Simulation evaluations based on statistical losses and real-world packet loss traces reveal the following: (i) Our proposed network-adaptive algorithm combined with our optimal streaming codes can achieve significantly higher performance compared to uncoded and non-adaptive FEC schemes over UDP (User Datagram Protocol); (ii) Our explicit streaming codes can significantly outperform traditional MDS (maximum-distance separable) streaming schemes when they are used along with our network-adaptive algorithm. In addition, we study different factors that can affect the performance of our network-adaptive algorithm.
Salma Emara, Silas L. Fong, Baochun Li, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Multim.7
2021 Streaming Erasure Codes over Multi-Access Relay Networks
abstract
Applications where multiple users communicate with a common server and desire low latency are common and increasing. This paper studies a network with two source nodes, one relay node and a destination node, where each source nodes wishes to transmit a sequence of messages, through the relay, to the destination, who is required to decode the messages with a strict delay constraint$T$. The network with a single source node has been studied in [1]. We start by introducing two important tools: the delay spectrum, which generalizes delay-constrained point-to-point transmission, and concatenation, which, similar to time sharing, allows combinations of different codes in order to achieve a desired regime of operation. Using these tools, we are able to generalize the two schemes previously presented in [1], and propose a novel scheme which allows us to achieve optimal rates under a set of well-defined conditions. Such novel scheme is further improved in order to achieve higher rates in the scenarios where the conditions for optimality are not met.
Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ISIT5
2021 Guaranteed Rate of Streaming Erasure Codes over Multi-Link Multi-hop Network
abstract
We study the problem of transmitting a sequence of messages (streaming messages) through a multi-link, multi-hop packet erasure network. Each message must be reconstructed in-order and under a strict delay constraint. Special cases of our setting with a single link on each hop have been studied recently - the case of a single relay-node, is studied in Fong et al [1]; the case of multiple relays, is studied in Domanovitz et al [2]. As our main result, we propose an achievable rate expression that reduces to previously known results when specialized to their respective settings. Our proposed scheme is based on the idea of concatenating single-link codes from [2] in a judicious manner to achieve the required delay constraints. We propose a systematic approach based on convex optimization to maximize the achievable rate in our framework.
Elad Domanovitz, Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ITW5
2021 High Rate Streaming Codes Over the Three-Node Relay Network
abstract
In this paper, we investigate streaming codes over a three-node relay network. Source node transmits a sequence of message packets to the destination via a relay. Source-to-relay and relay-to-destination links are unreliable and introduce at most N1and N2packet erasures, respectively. Destination needs to recover each message packet with a strict decoding delay constraint of T time slots. We propose streaming codes under this setting for all feasible parameters $\{N_{1},\ N_{2},\ T\}$. Relay naturally observes erasure patterns occurring in the source-to-relay link. In our code construction, we employ a channel-state-dependent relaying strategy, which rely on these observations. In a recent work, Fong et al. provide streaming codes featuring channel-state-independent relaying strategies, for all feasible parameters $\{N_{1},\ N_{2},\ T\}$. Our schemes offer a strict rate improvement over the schemes proposed by Fong et al., whenever $N_{1}\lt N_{2}$.
M. Nikhil Krishnan, Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ITW6
2020 Streaming Erasure Codes over Multi-hop Relay Network
abstract
A typical path over the internet is composed of multiple hops. When considering the transmission of a sequence of messages (streaming messages) through packet erasure channel over a three-node network, it has been shown that taking into account the erasure pattern of each segment can result in improved performance compared to treating the channel as a point-to-point link. Since rarely there is only a single relay between the sender and the destination, it calls for trying to extend this scheme to more than a single relay. In this paper, we first extend the upper bound on the rate of transmission of a sequence of messages for any number of relays. We further suggest an achievable adaptive scheme that is shown to achieve the upper bound up to the size of an additional header that is required to allow each receiver to meet the delay constraints.
Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ISIT5
2020 Optimal Streaming Erasure Codes Over the Three-Node Relay Network
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory6
2020 Optimal Multiplexed Erasure Codes for Streaming Messages With Different Decoding Delays
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory6
2019 Optimal Streaming Erasure Codes over the Three-Node Relay Network
abstract
This paper investigates low-latency streaming codes for a three-node relay network. The source transmits a sequence of messages (streaming messages) to the destination through the relay between them, where the first-hop channel from the source to the relay and the second-hop channel from the relay to the destination are subject to packet erasures. Every source message generated at a time slot must be recovered perfectly at the destination within the subsequent T time slots. In any sliding window of $ {T}+1$ time slots, we assume no more than $ {N}_{1}$ and ${N}_{2}$ erasures are introduced by the first-hop channel and second-hop channel respectively. We fully characterize the maximum achievable rate in terms of T, $ {N}_{1}$ and $ {N}_{2}$ . The achievability is proved by using a symbol-wise decode-forward strategy where the source symbols within the same message are decoded by the relay with different delays. The converse is proved by analyzing the maximum achievable rate for each channel when the erasures in the other channel are consecutive (bursty). In addition, we show that traditional message-wise decode-forward strategies, which require the source symbols within the same message to be decoded by the relay with the same delay, are sub-optimal in general.
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
ISIT6
2019 Optimal Multiplexed Erasure Codes for Streaming Messages with Different Decoding Delays
abstract
This paper considers multiplexing two sequences of messages with two different decoding delays over a packet erasure channel. In each time slot, the source constructs a packet based on the current and previous messages and transmits the packet, which may be erased when the packet travels from the source to the destination. The destination must perfectly recover every source message in the first sequence subject to a decoding delay Tv, and every source message in the second sequence subject to a shorter decoding delay Tu≤ Tv,. We assume that the channel loss model introduces a burst erasure of a fixed length B on the discrete timeline. Under this channel loss assumption, the capacity region for the case where Tv, ≤ Tu+B was previously solved. In this paper, we fully characterize the capacity region for the remaining case Tv, ) Tu+B. The key step in the achievability proof is achieving the non-trivial corner point of the capacity region through using a multiplexed streaming code constructed by superimposing two single-stream codes. The main idea in the converse proof is obtaining a genie-aided bound when the channel is subject to a periodic erasure pattern where each period consists of a length-B burst erasure followed by a length-Tunoiseless duration.
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
ISIT6
2019 Low-Latency Network-Adaptive Error Control for Interactive Streaming
abstract
We introduce a novel network-adaptive algorithm that is suitable for alleviating network packet losses for low-latency interactive communications between a source and a destination. Network packet losses happen in a bursty manner as well as an arbitrary manner, where the former is usually due to network congestion and the latter can be caused by unreliable wireless links. Our network-adaptive algorithm estimates in real time the best parameters of a recently proposed streaming code that corrects both arbitrary losses (which cause crackling noise in audio) and burst losses (which cause undesirable jitters and pauses in audio) using forward error correction (FEC). The network-adaptive algorithm updates the coding parameters in real time as follows: The destination estimates appropriate coding parameters based on its observed packet loss pattern and then the parameters are fed back to the source for updating the underlying code. In addition, a new explicit construction of practical low-latency streaming codes that achieve the optimal tradeoff between the capability of correcting arbitrary losses and the capability of correcting burst losses is provided. Simulation evaluations based on real-world packet loss traces reveal that our proposed network-adaptive algorithm combined with our optimal streaming codes achieves significantly higher reliability compared to uncoded and non-adaptive FEC schemes over UDP (User Datagram Protocol).
Silas L. Fong, Salma Emara, Baochun Li, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ACM Multimedia7
2019 Optimal Streaming Codes for Channels With Burst and Arbitrary Erasures
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory6
2018 Optimal Streaming Codes for Channels with Burst and Arbitrary Erasures
abstract
This paper considers transmitting a sequence of messages (streaming messages) over a packet erasure channel. In each time slot, the source constructs a packet based on the current and the previous messages and transmits the packet, which may be erased when the packet travels from the source to the destination. Every source message must be recovered perfectly at the destination subject to a fixed decoding delay. We assume that the channel loss model introduces either one burst erasure or multiple arbitrary erasures in any fixed-sized sliding window. Under this channel loss assumption, we fully characterize the maximum achievable rate by constructing streaming codes that achieve the optimal rate. In addition, our construction of optimal streaming codes implies the full characterization of the maximum achievable rate for convolutional codes with any given column distance, column span, and decoding delay. Numerical results demonstrate that the optimal streaming codes outperform existing streaming codes of comparable complexity over some instances of the Gilbert-Elliott channel and the Fritchman channel.
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos
ISIT6
2018 Multiplexed Coding for Multiple Streams With Different Decoding Delays
abstract
We consider a communication setup where two source streams with different decoding deadlines, must be simultaneously transmitted over a single channel subjected to burst erasures. The encoder multiplexes the two source streams into a single stream of channel-packets. The decoder must recover the two source streams sequentially by their corresponding deadlines. One of the streams, the urgent stream, has a smaller delay than the other stream. We study the capacity region for such a setting for a certain range of system parameters. We divide the system into three different cases based on the relative values of the delays. For each case we provide achievability and converse bounds, which match under certain conditions. Our proposed coding scheme involves a careful construction of the parity check packets by jointly coding across the two streams despite different deadlines. Interestingly it is possible to transmit the urgent stream at a certain positive rate even when the sum rate equals the capacity associated with the less urgent stream. A separation based approach where we apply separate single-stream codes to each stream is suboptimal. Although our capacity results assume a simplistic channel model with a single erasure burst, we further demonstrate that our proposed code constructions also provide significant performance gains in simulations over statistical channel models with random bursts.
Ahmed Badr, Devin Lui, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory6
2017 FEC for VoIP using dual-delay streaming codes
abstract
We introduce a new class of forward error correction (FEC) codes for VoIP communications which support different recovery delay depending on the channel conditions. Specifically, our proposed class of Dual-Delay (DD) codes can recover from challenging long bursts of losses with close to theoretical minimum delay so as to meet playback deadlines for recovered packets. They further improve conversational interactivity by achieving lower recovery delay during periods of random isolated losses. These DD codes are shown to achieve lower residual loss rates when compared to existing codes over a wide range of parameters of the Gilbert-Elliott channel. Experiments over real world packet traces further show performance gains of DD codes in terms of perceptually motivated ITU-T G.107 E-model.
Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
INFOCOM5
2017 Multiplexed FEC for multiple streams with different playout deadlines
abstract
We study a setting where two source streams with different decoding deadlines must be transmitted over a burst erasure channel. The source streams are multiplexed into a single stream of channel-packets at the encoder and transmitted over a packet erasure channel. The decoder must recover the source-packets within each stream sequentially, by their corresponding deadlines. We consider the burst-erasure channel model and characterize the capacity region for a certain range of system parameters. We show that the operation of the system can be divided into three different regimes based on the relative values of decoding deadlines. On the achievability side, we show that jointly coding across the two streams, despite their different deadlines, is necessary to achieve capacity. On the converse side we develop information theoretic outer bounds on the capacity region. We find that the capacity region exhibits a “corner point” where we can transmit the urgent stream at a positive rate, yet attain a sum-rate equal to the capacity of the non-urgent stream.
Ahmed Badr, Devin Lui, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ISIT6
2017 Market-based dynamic service mode switching in virtualized wireless networks
abstract
Consider a wireless networking architecture, where multiple infrastructure access points (AP) dynamically offer (bid) service deals (modes) to a mobile over time. Akin to a market offer, each service deal is comprised of service/quality attributes (e.g, AP, rate) for a price (cost). The mobile dynamically selects the desirable deal at each time, so as to efficiently trade long term latency/quality for cumulative price. Offered service modes depend on the randomly fluctuating congestion state of the AP infrastructure. Dynamically switching from one service mode/deal to another, the mobile encounters `friction' (e.g. bandwidth loss, disconnection risk) and, hence, has an incentive to stick with the current deal for as long as this is significantly competitive. A model of this architecture is first developed, which allows for the formulation and computation of the optimal control for the mobile to accept an offered deal amongst many and switch into the corresponding service mode. A suite of low-complexity heuristic controls for mode switching is also discussed. The performance of the optimal and heuristic controls is probed via simulation. Finally, the `switching curve' structure of the optimal control is demonstrated on a simple system where the curves can be plotted.
Maria Dimakopoulou, Nicholas Bambos, Martin Valdez-Vivas, John G. Apostolopoulos
PIMRC4
2017 Layered Constructions for Low-Delay Streaming Codes
abstract
We study error correction codes for multimedia streaming applications where a stream of source packets must be transmitted in real-time, with in-order decoding, and strict delay constraints. In our setup, the encoder observes a stream of source packets in a sequential fashion, and M channel packets must be transmitted between the arrival of successive source packets. Each channel packet can depend on all the source packets observed up to and including that time, but not on any future source packets. The decoder must reconstruct the source stream with a delay of T packets. We consider a class of packet erasure channels with burst and isolated erasures, where the erasure patterns are locally constrained. Our proposed model provides a tractable approximation to statistical models, such as the Gilbert-Elliott channel, for capacity analysis. When M = 1, i.e., when the source-packet arrival and channel-packet transmission rates are equal, we establish upper and lower bounds on the capacity, that are within one unit of the decoding delay T. We also establish necessary and sufficient conditions on the column distance and column span of a convolutional code to be feasible, and in turn establish a fundamental tradeoff between these. Our proposed codes-maximum distance and span codes- achieve a near-optimal tradeoff between the column distance and column span, and involve a layered construction. When M > 1, we establish the capacity for the burst-erasure channel and an achievable rate in the general case. Extensive numerical simulations over Gilbert-Elliott and Fritchman channel models suggest that our codes also achieve significant gains in the residual loss probability over statistical channel models.
Ahmed Badr, Pratik Patil, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
IEEE Trans. Inf. Theory5
2016 Content-independent and loss-pattern-aware distortion evaluation for streaming media
abstract
It is well known that dispersed and burst packet losses introduce significantly different amount of distortions. Since perceptual models are typically content dependent, it is challenging to characterize how losses interact with concealment. This paper presents loss-pattern-aware distortion (LoPAD), a content-independent metric that explicitly models the impact of different loss patterns. LoPAD operates solely on the loss trace without analyzing received media. It is fast, and supports offline and cloud-based monitoring of network impairment. Taking audio conferencing as target application, we show that full-reference PESQ scores for a collection of speech samples can be closely approximated by LoPAD. For various combinations of erasure channel models and forward error correction (FEC) codes, the correlation coefficients between LoPAD and PESQ-DMOS range from 0.90 to 0.97.
Wai-tian Tan, John G. Apostolopoulos, Ahmed Badr, Ashish Khisti
ICIP3
2015 Software defined networking for video: Overview & multicast study
abstract
Software Defined Networking (SDN) is an architectural trend in networking towards the use of centralized controllers to improve global visibility and simplify various network operation tasks. Due to their need for QoS, sizable traffic involved, and dynamic nature, video applications are particularly suitable candidates for dynamic interaction with SDN. This paper provides an overview of different general ways that video applications have interacted with networks, and outline what opportunities SDN provides. We then develop and evaluate several methods that exploit SDN to construct multicast video trees and show that over realistic topology, SDN optimized trees can support 60-90% more traffic.
Wai-tian Tan, Herb Wildfeuer, John G. Apostolopoulos
ICIP3
2015 Introduction to the Special Section on Visual Computing in the Cloud: Fundamentals and Applications
abstract
Cloud computing involves a large number of terminals connected through a real-time high-speed network (such as the Internet). The adoption rates for private and hybrid cloud services increased to 40% in 2013, with computing shifting from on-premise infrastructure to the cloud. To keep pace with the ever-accelerating rate of innovation, companies are moving to the cloud. However, visual computing in the cloud brings great challenges, such as how to measure and then improve the quality of experience in cloud computing. This Special Section provides the image/video community a forum to present new academic research and industrial development in running visual computing services in the cloud. This Special Section aims to address fundamental and practical aspects of visual computing in the cloud, such as how to build cloud platforms that can cope with seemingly unlimited supply of content coming from traditional media sources as well as new media uploaded to the Internet (YouTube, Facebook, etc.); how to leverage cloud technology to build high-quality image/video browsing and delivery experiences for a global audience; how to ingest, encode, process, adapt, as well as protect contents and privacy of users; how to provide both on-demand and live-streaming capabilities; how to tag image/video and allow consumers to access the image/video contents with high availability; how to support image/video services in mobile devices; and how to perform real-time image/video analytics in the cloud, to mention a few among a diverse range of challenges.
Jiangchuan Liu, Wenwu Zhu 0001, Touradj Ebrahimi, John G. Apostolopoulos, Xian-Sheng Hua 0001, Chuan Wu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2014 Resource management in cloud computing with frictions and congestion weather
abstract
Cloud infrastructures with virtualized CPU and memory resources have the potential for providing high quality of service at increased levels of energy efficiency. By dynamically tailoring the capacity of a virtual machine to workload demands, a cloud infrastructure can significantly reduce the number of physical resources it has online, saving on decreased power costs. These resource management techniques, however, have yet to gain widespread appeal among network engineers due to the significant delays and setup costs in activating or reconfiguring cloud resources. A further challenge is these "frictions" fluctuate through time, based on complex system-wide supply and demand "weather" patterns in the cloud as a whole. In this paper, we develop a loss queueing model for capacity provisioning for a virtual machine that draws its computation resources from the cloud under varying friction cost. We solve for the optimal control policy using dynamic programming and discuss its intuitive structural properties. Finally we run simulations to compare the performance of the optimal policy against two benchmarks and a heuristic policy.
Martin Valdez-Vivas, Nicholas Bambos, John G. Apostolopoulos
GLOBECOM3
2014 Dynamic resource management in virtualized data centers with bursty traffic
abstract
Reducing energy consumption in data centers has been a persistent goal for the last decade. Advances in virtualization and dynamic power management have helped to curb excess utilization by sharing resources and running system components in lower energy states during periods of low traffic. However, switching components among virtual machines and energy states is subject to setup times and increased energy consumption to bring them online, and this combined with the innate burstiness of traffic in data centers are significant deterrents to the successful deployment of power management techniques. In this paper, we develop a queueing model with a controllable service rate that accounts for switching frictions within a setting that has a changing arrival rate. We solve for the optimal control policy using dynamic programming, and find it has intuitive structural properties. We compare its performance against benchmark heuristics, and the relative advantages among them in different scenarios are discussed.
Martin Valdez-Vivas, Nicholas Bambos, John G. Apostolopoulos
ICC3
2013 Playout buffer responsive wireless streaming for multiple clients
abstract
We consider a problem faced by wireless base stations in which multiple requests to stream data must be accomodated while minimizing the amount of buffering time spent by clients. In our model, clients request content in discrete data chunks, but the base station is restricted in which clients can simultaneously be served data at certain rates. This limitation occurs frequently due to wireless connectivity and congestion issues in which some clients are difficult to reach from the base station, whereas others are more easily serviced. These constraints are represented by a set of admissible service rate vectors from which a scheduler at the base station must choose. We take a queuing theoretic approach to this decision problem and employ a stress alignment approach to ensure that maximal throughput from the data center to the clients is achieved. As a consequence, whenever it is possible to stabilize the backlog of requests from each client, we are able to do so. Numerical experiments show that among the variations of this scheduling algorithm, we can choose parameters to vary the priority levels of different clients and even induce dependencies between service rates of different clients.
Praveen Bommannavar, John G. Apostolopoulos, Nicholas Bambos
GLOBECOM2
2013 Resource allocation and scheduling for energy efficient tracking
abstract
We examine the problem of tracking the states of a collection of systems over a finite horizon in a power limited scenario. Specifically, each system has a sensor which can track a property of interest and has a fixed budget with which to make measurements and communicate to a fusion center. The state at each system varies independently according to a Markov model and the transitions between different states occur according known transition matrices. Different systems can have vastly different state evolution statistics. At each time step, the fusion center can request an update from any number of the systems, subject to the constraint that the corresponding sensor has not exhausted its budget to do so. These measurement updates are expensive and hence resource limited. After the fusion center receives all updates, it must estimate the state at each system with minimum error. We give an optimal policy for the fusion center to request updates from each sensor and also provide an optimal policy for the fusion center to allocate the measurement budget to each sensor before deployment, given the transition matrices corresponding to each system.
Praveen Bommannavar, John G. Apostolopoulos, Nicholas Bambos
ICC2
2013 Deadline aware packet scheduling in switches for multimedia streaming applications
abstract
We consider the problem of scheduling packets in an input queued switch with a focus on processing streamed multimedia data. In such applications, packets arrive with hard service deadlines; after the deadline for a packet has passed, it is no longer useful and does not get delivered - it is dropped. We seek policies to minimize the number of late packets, which are then dropped. The problem is formulated in a Dynamic Programming framework and shown to be intractable. The formulation is contrasted to the related crossbar switch scheduling problem, with an emphasis on the fact that we have a different objective function. A simplified probabilistic version of the streaming problem is used as motivation for a heuristic solution. Finally, we present results from a simulation in which a simple heuristic based on weighting queues according to the deadline of the leading packet consistently outperforms the well-known maximum weight matching (MWM) algorithm.
Praveen Bommannavar, John G. Apostolopoulos, Nicholas Bambos
ICC2
2013 Streaming codes for channels with burst and isolated erasures
abstract
We study low-delay error correction codes for streaming recovery over a class of packet-erasure channels that introduce both burst-erasures and isolated erasures. We propose a simple, yet effective class of codes whose parameters can be tuned to obtain a tradeoff between the capability to correct burst and isolated erasures. Our construction generalizes previously proposed low-delay codes which are effective only against burst erasures. We establish an information theoretic upper bound on the capability of any code to simultaneously correct burst and isolated erasures and show that our proposed constructions meet the upper bound in some special cases. We discuss the operational significance of column-distance and column-span metrics and establish that the rate 1/2 codes discovered by Martinian and Sundberg [IT Trans. 2004] through a computer search indeed attain the optimal column-distance and column-span tradeoff. Numerical simulations over a Gilbert-Elliott channel model and a Fritchman model show significant performance gains over previously proposed low-delay codes and random linear codes for certain range of channel parameters.
Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
INFOCOM4
2013 Robust streaming erasure codes based on deterministic channel approximations
abstract
We study near optimal error correction codes for real-time communication. In our setup the encoder must operate on an incoming source stream in a sequential manner, and the decoder must reconstruct each source packet within a fixed playback deadline of T packets. The underlying channel is a packet erasure channel that can introduce both burst and isolated losses. We first consider a class of channels that in any window of length T +1 introduce either a single erasure burst of a given maximum length B, or a certain maximum number N of isolated erasures. We demonstrate that for a fixed rate and delay, there exists a tradeoff between the achievable values of B and N, and propose a family of codes that is near optimal with respect to this tradeoff. We also consider another class of channels that introduce both a burst and an isolated loss in each window of interest and develop the associated streaming codes. All our constructions are based on a layered design and provide significant improvements over baseline codes in simulations over the Gilbert-Elliott channel.
Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos
ISIT4
2013 Fusion of Median and Bilateral Filtering for Range Image Upsampling
abstract
We present a new upsampling method to enhance the spatial resolution of depth images. Given a low-resolution depth image from an active depth sensor and a potentially high-resolution color image from a passive RGB camera, we formulate it as an adaptive cost aggregation problem and solve it using the bilateral filter. The formulation synergistically combines the median and bilateral filters thus it better preserves the depth edges and is more robust to noise. Numerical and visual evaluations on a total of 37 Middlebury data sets demonstrate the effectiveness of our method. A real-time high-resolution depth capturing system is also developed using commercial active depth sensor based on the proposed upsampling method.
Qingxiong Yang, Narendra Ahuja, Ruigang Yang, Kar-Han Tan, James Davis 0001, W. Bruce Culbertson, John G. Apostolopoulos, Gang Wang 0012
IEEE Trans. Image Process.7
2012 Cloud-based depth sensing quality feedback for interactive 3D reconstruction
abstract
In this paper we propose a cloud-based approach to improve the 3D reconstruction capability of handheld devices with real-time depth sensors. We attempt to characterize the quality of 3D information captured by real time depth sensing devices, and in particular examine how sensors from Prime Sense and Canesta measure distances, and derive simple analytical models on performance limitations for each. We also study the factors that affect depth sensing quality when these devices are used to incrementally build larger or denser 3D models. Empirical experiments confirm our analysis. Our findings allow us to design a quality metric which can interactively inform users to guide them on how to optimize the quality of their captured 3D content.
Kar-Han Tan, John G. Apostolopoulos
ICASSP2
2012 Power budgeted packet scheduling for wireless multimedia
abstract
In this paper we profile a particular tradeoff between power budget and video quality that emerges in the transmission of multimedia packets over a wireless channel. These packets are due to arrive to a receiver at a particular time, so we consider a finite horizon problem over which multimedia data are transmitted. Due to the lossy nature of the wireless channel, however, not every packet can be successfully sent across the channel. Hence, each packet that is lost leads to distortion in the video that is experienced by the receiver. We suppose that there are M packets that must arrive at the receiver within N time steps, but that power limitations constrain the number of transmissions. At each time step, we may make a measurement of the wireless channel and decide whether or not to transmit a packet over the channel at that time. First we will suppose the times at which the channel state is sampled are spaced far enough apart so that the samples are i.i.d. Then we will continue by supposing that the channel state follows a Markov chain.
Praveen Bommannavar, Nicholas Bambos, John G. Apostolopoulos
ICC3
2012 The Road to Immersive Communication
abstract
Communication has seen enormous advances over the past 100 years including radio, television, mobile phones, video conferencing, and Internet-based voice and video calling. Still, remote communication remains less natural and more fatiguing than face-to-face. The vision of immersive communication is to enable natural experiences and interactions with remote people and environments in ways that suspend disbelief in being there. This paper briefly describes the current state-of-the-art of immersive communication, provides a vision of the future and the associated benefits, and considers the technical challenges in achieving that vision. The attributes of immersive communication are described, together with the frontiers of video and audio for achieving them. We emphasize that the success of these systems must be judged by their impact on the people who use them. Recent high-quality video conferencing systems are beginning to deliver a natural experience-when all participants are in custom-designed studios. Ongoing research aims to extend the experience to a broader range of environments. Augmented reality has the potential to make remote communication even better than being physically present. Future natural and effective immersive experiences will be created by drawing upon intertwined research areas including multimedia signal processing, computer vision, graphics, networking, sensors, displays and sound reproduction systems, haptics, and perceptual modeling and psychophysics.
John G. Apostolopoulos, Philip A. Chou, W. Bruce Culbertson, Ton Kalker, Mitchell D. Trott, Susie J. Wee
Proc. IEEE1
2011 ConnectBoard: Enabling Genuine Eye Contact and Accurate Gaze in Remote Collaboration
abstract
Conventional telepresence systems allow remote users to see one another and interact with shared media, but users cannot make eye contact, and gaze awareness with respect to shared media and documents is lost. In this paper, we describe a remote collaboration system based on a see-through display to create an experience where local and remote users are seemingly separated only by a vertical sheet of glass. Users can see each other and media displayed on the shared surface. Face detectors are applied on the local and remote video streams to introduce an offset in the video display so as to bring the local user's face, the local camera, and the remote user's face image into collinearity. This ensures that, when the local user looks at the remote user's image, the camera behind the see-through display captures an image with the “Mona Lisa effect,” where the eyes of an image appears to follow the viewer. Experiments show that, for one-on-one meetings, our system is capable of capturing and delivering realistic, genuine eye contact as well as accurate gaze awareness with respect to shared media.
Kar-Han Tan, Ian N. Robinson, W. Bruce Culbertson, John G. Apostolopoulos
IEEE Trans. Multim.4
2010 Video cross-talk reduction and synchronization for two-way collaboration
abstract
Recent two-way collaboration prototypes attempt to improve natural interactivity, correct eye contact and gaze direction, and media sharing using novel configurations of projectors, screens, and video cameras. These systems are often afflicted by video cross-talk where the content displayed for viewing by the local participant is unintentionally captured by the camera and delivered to the remote participant. Prior attempts to reduce this cross-talk purely in hardware through various forms of multiplexing (e.g., temporal, wavelength (color), polarization) have performance and cost limitations. In this work, careful system characterization and subsequent signal processing algorithms allow us to reduce video cross-talk. The signals themselves are used to detect temporal synchronization offsets which then allow subsequent reduction of the cross-talk signal. Our software-based approach enables the effective use of simpler hardware and optics than prior methods. Results show substantial cross-talk reduction in a system with unsynchronized projector and camera.
Ramin Samadani, John G. Apostolopoulos, Ian N. Robinson, Kar-Han Tan
ICIP2
2010 Enabling genuine eye contact and accurate gaze in remote collaboration
abstract
Conventional telepresence systems allow remote users to see one another and interact with shared media and documents, but users cannot make eye contact, and gaze awareness with respect to shared media and documents is lost. In this paper we describe a remote collaboration system based on a see-through display to create an experience where local and remote users are seemingly separated only by a vertical sheet of glass. Users can see each other and media displayed on the shared surface. Face detectors on the local and remote video streams are used to introduce an offset in the video display so as to bring the local user's face, the local camera, and the remote user's face image into collinearity. This ensures that when the local user looks at the remote user's image, the camera behind the see-through display captures an image with the `Mona Lisa effect', where the eyes of an image appears to follow the viewer. Experiments show that our system is capable of capturing and delivering realistic, genuine eye contact as well as accurate gaze awareness with respect to shared media.
Kar-Han Tan, Ian N. Robinson, W. Bruce Culbertson, John G. Apostolopoulos
ICME4
2010 Gaze awareness and interaction support in presentations
abstract
Modern digital presentation systems use rich media to bring highly sophisticated information visualization and highly effective storytelling capabilities to classrooms and corporate boardrooms. In this paper we address a number of issues that arise when the ubiquitous computer-projector setup is used in large venues like the cavernous auditoriums and hotel ballrooms often used in large scale academic meetings and industrial conferences. First, when the presenter is addressing a large audience the slide display needs to be very large and placed high enough so that it is clearly visible from all corners of the room. This makes it impossible for a presenter to walk up to the display and interact with the display with gestures, gaze, and other forms of paralanguage. Second, it is hard for the audience to know which part of the slide the presenter is looking at when he/she has to look the opposite way from the audience while interacting with the slide material. It is also hard for the presenter to see the audience in these cases. Even though there may be video captures of the presenter, slides, and even the audience, the above factors add up to make it very difficult for a user viewing either a live feed or a recording to grasp the interaction between all the components and participants of a presentation. We address these problems with a novel presentation system which creates a live video view that seamlessly combines the presenter and the presented material, capturing all graphical, verbal, and nonverbal channels of communication. The system also allows the local and remote audiences to have highly interactive exchanges with the presenter while creating a comprehensive view for recording or remote streaming.
Kar-Han Tan, Dan Gelb, Ramin Samadani, Ian N. Robinson, W. Bruce Culbertson, John G. Apostolopoulos
ACM Multimedia6
2010 Fusion of active and passive sensors for fast 3D capture
abstract
We envision a conference room of the future where depth sensing systems are able to capture the 3D position and pose of users, and enable users to interact with digital media and contents being shown on immersive displays. The key technical barrier is that current depth sensing systems are noisy, inaccurate, and unreliable. It is well understood that passive stereo fails in non-textured, featureless portions of a scene. Active sensors on the other hand are more accurate in these regions and tend to be noisy in highly textured regions. We propose a way to synergistically combine the two to create a state-of-the-art depth sensing system which runs in near real time. In contrast the only known previous method for fusion is slow and fails to take advantage of the complementary nature of the two types of sensors.
Qingxiong Yang, Kar-Han Tan, W. Bruce Culbertson, John G. Apostolopoulos
MMSP4
2010 Channel, deadline, and distortion (CD2) aware scheduling for video streams over wireless
abstract
We study scheduling of multimedia traffic on the downlink of a wireless communication system. We examine a scenario where multimedia packets are associated with strict deadlines and are equivalent to lost packets if they arrive after their associated deadlines. Lost packets result in degradation of playout quality at the receiver, which is quantified in terms of the "distortion cost" associated with each packet. Our goal is to design a scheduler which minimizes the aggregate distortion cost over all receivers. We study the scheduling problem in a dynamic programming (DP) framework. Under well justified modeling reductions, we extensively characterize structural properties of the optimal control associated with the DP problem. We leverage these properties to design a low-complexity Channel, Deadline, and Distortion (CD2) aware heuristic scheduling policy amenable to implementation in real wireless systems. We evaluate the performance of CD2via trace-driven simulations using H.264/MPEG-4 AVC coded video. Our experimental results show that CD2comfortably outperforms benchmark schedulers like earliest deadline first (EDF) and best channel first (BCF). CD2achieves these performance gains by using the knowledge of packet deadlines, wireless channel conditions, and application specific information (per-packet distortion costs) in a systematic and unified way for multimedia scheduling.
Aditya Dua, Carri W. Chan, Nicholas Bambos, John G. Apostolopoulos
IEEE Trans. Wirel. Commun.4
2009 ConnectBoard: A remote collaboration system that supports gaze-aware interaction and sharing
abstract
We present ConnectBoard, a new system for remote collaboration where users experience natural interaction with one another, seemingly separated only by a vertical, transparent sheet of glass. It overcomes two key shortcomings of conventional video communication systems: the inability to seamlessly capture natural user interactions, like using hands to point and gesture at parts of shared documents, and the inability of users to look into the camera lens without taking their eyes off the display. We solve these problems by placing the camera behind the screen, where the remote user is virtually located. The camera sees through the display to capture images of the user. As a result, our setup captures natural, frontal views of users as they point and gesture at shared media displayed on the screen between them. Users also never have to take their eyes off their screens to look into the camera lens. Our novel optical solution based on wavelength multiplexing can be easily built with off-the-shelf components and does not require custom electronics for projector-camera synchronization.
Kar-Han Tan, Ian N. Robinson, Ramin Samadani, Bowon Lee, Dan Gelb, Alex Vorbau, W. Bruce Culbertson, John G. Apostolopoulos
MMSP8
2009 Generalized Butterfly Graph and Its Application to Video Stream Authentication
abstract
This paper presents the generalized butterfly graph (GBG) and its application to video stream authentication. Compared with the original butterfly graph, the proposed GBG provides significantly increased flexibility, which is necessary for streaming applications, including supporting arbitrary bit-rate budget for authentication and arbitrary number of video packets. Within the GBG design, the problem of constructing an authentication graph is defined as follows: given the total number of packets to protect, the expected packet loss rate for the network, and the available overhead budget, how should one design the authentication graph to maximize the probability that the received packets are verifiable? Furthermore, given the fact that media packets are typically of unequal importance, we explore two variants of the GBG authentication, packet sorting and unequal authentication protection, which apply unequal treatment to different packets based on their importance. Lastly, we examine how the proposed GBG authentication can be applied within the context of rate-distortion-authentication (R-D-A) optimized streaming: given a media stream protected by GBG authentication, the R-D-A optimized streaming technique computes an optimized transmission schedule by recognizing and accounting for the authentication dependencies in the GBG authentication graph.
Zhishou Zhang, Qibin Sun, John G. Apostolopoulos, Lawrence Wai-Choong Wong
IEEE Trans. Circuits Syst. Video Technol.3
2009 Scheduling Algorithms for Broadcasting Media with Multiple Distortion Measures
abstract
The growing popularity of multimedia streaming applications brings a growth in diversity of media clients (laptops, PDAs, cellphones). Effectively serving this heterogeneous group of users is highly desirable. Scalable media codecs such as H.264/MPEG-4 SVC help make this adaptation possible. To account for the various capabilities and requests of each user, such as varying spatial or temporal resolutions, multiple distortion measures (MDM) are considered. Rather than consider a homogeneity in users, the MDM framework considers multiple different distortion values for each media packet for each user type. We consider the scenario of simultaneously broadcasting a video stream to multiple users over wireless links. The objective is to design a scheduling algorithm which achieves the highest aggregate quality-of-service, measured by distortion and delay, over all different user types. We cast the problem as a stochastic shortest path problem and use dynamic programming to find the optimal policy. For statistically static channels, the optimal policy is shown to be of threshold type. For time-varying channels, a quasi-static policy is introduced. Experimental results show that our policy reduces distortion by up to a factor of 2 over conventional approaches which do not consider MDM.
Carri W. Chan, Nicholas Bambos, Susie J. Wee, John G. Apostolopoulos
IEEE Trans. Wirel. Commun.4
2008 Wireless Video Broadcasting to Diverse Users
abstract
The growing diversity in media clients calls for content providers to adapt media content to adhere to their various needs. It is desirable to serve these heterogeneous users in a fast and efficient manner. Scalable media, such as H.264/MPEG- 4 SVC, helps make this possible. To account for various viewing capabilities of each user, such as different spatial or temporal resolutions, the Multiple Distortion Measures framework is used [1], [2]. MDM associates multiple distortion values with each packet depending on the user types (low/high resolution viewers, low/high frame rate viewers, etc.) who will consume the media packets. In this paper, we examine how to broadcast media packets with multiple distortion measures to multiple users. The tradeoff between media distortion and delay (in the form of retransmissions) plays an integral role in the scheduling decision. We cast the problem as a stochastic shortest path problem and use Dynamic Programming to find the optimal policy. In the case of statistically static channels, the optimal policy is shown to be a threshold policy where the number of allowable retransmissions is dictated by the importance, in terms of incurred distortion, of each packet. Through experimental results, we show that our policy, which considers multiple distortion measures, achieves up to 8 dB gains over conventional approaches. Finally, a policy based on the theoretical results of statistically static channels is empirically shown to have high performance for time-varying channels modeled by a two-state Markov Chain.
Carri W. Chan, Nicholas Bambos, Susie J. Wee, John G. Apostolopoulos
ICC4
2008 Multiple Distortion Measures for video with temporal scalability
abstract
With the recent growth in mobile device types, an interesting question that arises is how to serve media content to users with a diverse set of capabilities. Defining multiple distortion metrics, one for each type of user, allows for specialized service to each user type, which can result in significantly improved quality of service. H.264/MPEG-4 SVC is a scalable video codec that allows for easy spatial and temporal adaptation of encoded video by selecting or dropping packets. We examine using multiple distortion measures (MDM) for scheduling packets in the context of streaming temporally scalable video to users with different target frame rates. We show that gains in PSNR of multiple dB can be achieved when taking MDM into consideration for making scheduling decisions.
Carri W. Chan, Susie J. Wee, John G. Apostolopoulos
ICIP3
2008 Rate-Distortion-Authentication optimized streaming with Generalized Butterfly Graph authentication
abstract
This paper shows how Rate-Distortion-Authentication (R-D-A) optimized streaming may be performed with the Generalized Butterfly Graph (GBG) stream authentication method. R-D-A streaming is designed to compute an optimized transmission policy by accounting for both coding and authentication dependencies, and GBG is designed to protect the authenticity of a media stream. The GBG graph is better suited for streaming than the original Butterfly graph since it is highly flexible and can support an arbitrary number of packets and arbitrary overhead. However, the dependency chains within the GBG graph are much longer and are tangled with each other — making it very difficult to quantify the dependencies necessary to perform R-D-A streaming. We propose a method to estimate the authentication dependencies, which is then used by the R-D-A technique to compute the optimized transmission policy.
Zhishou Zhang, Qibin Sun, John G. Apostolopoulos, Lawrence Wai-Choong Wong
ICIP3
2008 Quality-Optimized and Secure End-to-End Authentication for Media Delivery
abstract
The need for security services, such as confidentiality and authentication, has become one of the major concerns in multimedia communication applications, such as video on demand and peer-to-peer content delivery. Conventional data authentication cannot be directly applied for streaming media when an unreliable channel is used and packet loss may occur. This paper begins by reviewing existing end-to-end media authentication schemes, which can be classified into stream-based and content-based techniques. We then motivate and describe how to design authentication schemes for multimedia delivery that exploit the unequal importance of different packets. By applying conventional cryptographic hashes and digital signatures to the media packets, the system security is similar to that achievable in conventional data security. However, instead of optimizing packet verification probability, we optimize the quality of the authenticated media, which is determined by the packets that are received and able to be decoded and authenticated. The quality of the authenticated media is optimized by allocating the authentication resources unequally across streamed packets based on their relative importance, thereby providing unequal authenticity protection. The effectiveness of this approach is demonstrated through experimental results on different media types (image and video), different compression standards (JPEG, JPEG2000, and H.264), and different channels (wired with packet erasures and wireless with bit errors).
Qibin Sun, John G. Apostolopoulos, Chang Wen Chen, Shih-Fu Chang
Proc. IEEE2
2008 Analysis of Packet Loss for Compressed Video: Effect of Burst Losses and Correlation Between Error Frames
abstract
Video communication is often afflicted by various forms of losses, such as packet loss over the Internet. This paper examines the question of whether the packet loss pattern, and in particular, the burst length, is important for accurately estimating the expected mean-squared error distortion resulting from packet loss of compressed video. We focus on the challenging case of low-bit-rate video where each P-frame typically fits within a single packet. Specifically, we: 1) verify that the loss pattern does have a significant effect on the resulting distortion; 2) explain why a loss pattern, for example a burst loss, generally produces a larger distortion than an equal number of isolated losses; and 3) propose a model that accurately estimates the expected distortion by explicitly accounting for the loss pattern, inter-frame error propagation, and the correlation between error frames. The accuracy of the proposed model is validated with H.264/AVC coded video and previous frame concealment, where for most sequences the total distortion is predicted to within plusmn0.3 dB for burst loss of length two packets, as compared to prior models which underestimate the distortion by about 1.5 dB. Furthermore, as the burst length increases, our prediction is within plusmn0.7 dB, while prior models degrade and underestimate the distortion by over 3 dB. The proposed model works well for video-telephony-type of sequences with low to medium motion. We also present a simple illustrative example, of how knowledge of the effect of burst loss can be used to adapt the schedule of video streaming to provide improved performance for a burst loss channel, without requiring an increase in bit rate.
Yi J. Liang, John G. Apostolopoulos, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.2
2008 Multiple Distortion Measures for Packetized Scalable Media
abstract
As the diversity in end-user devices and networks grows, it becomes important to be able to efficiently and adaptively serve media content to different types of users. A key question surrounding adaptive media is how to do Rate-Distortion optimized scheduling. Typically, distortion is measured with a single distortion measure, such as the Mean-Squared Error compared to the original high resolution image or video sequence. Due to the growing diversity of users with varying capabilities such as different display sizes and resolutions, we introduceMultipleDistortionMeasures(MDM) to account for a diverse range of users and target devices. MDM gives a clear framework with which to evaluate the performance of media systems which serve a variety of users. Scalable coders, such as JPEG2000 and H.264/MPEG-4 SVC, allow for adaptation to be performed with relatively low computational cost. We show that accounting for MDM can significantly improve system performance; furthermore, by combining this with scalable coding, this can be done efficiently. Given these MDM, we propose an algorithm to generateembeddedschedules, which enables low-complexity, adaptive streaming of scalable media packets to minimize distortion across multiple users. We show that using MDM achieves up to 4 dB gains for spatial scalability applied to images and 12 dB gains for temporal scalability applied to video.
Carri W. Chan, Susie J. Wee, John G. Apostolopoulos
IEEE Trans. Multim.3
2008 Content-Aware Playout and Packet Scheduling for Video Streaming Over Wireless Links
abstract
Media streaming over wireless links is a challenging problem due to both the unreliable, time-varying nature of the wireless channel and the stringent delivery requirements of media traffic. In this paper, we use joint control of packet scheduling at the transmitter and content-aware playout at the receiver, so as to maximize the quality of media streaming over a wireless link. Our contributions are twofold. First, we formulate and study the problem of joint scheduling and playout control in the framework of Markov decision processes. Second, we propose a novel content-aware adaptive playout control, that takes into account the content of a video sequence, and in particular the motion characteristics of different scenes. We find that the joint scheduling and playout control can significantly improve the quality of the received video, at the expense of only a small amount of playout slowdown. Furthermore, the content-aware adaptive playout places the slowdown preferentially in the low-motion scenes, where its perceived effect is lower.
Yan Li 0069, Athina Markopoulou, John G. Apostolopoulos, Nicholas Bambos
IEEE Trans. Multim.3
2008 Real-time monitoring of video quality in IP networks
Shu Tao, John G. Apostolopoulos, Roch Guérin
IEEE/ACM Trans. Netw.2
2007 Multiple Distortion Measures for Scalable Streaming with Jpeg2000
abstract
The increasing diversity of end-user devices and networks allow different users to receive and view images and video at different resolutions and rates. Scalable coding methods allow streaming media systems to easily adapt media by selecting packets according to schedules that reflect their importance. Traditional scheduling algorithms are based on a single distortion measure such as mean-squared error relative to the original high resolution image. In this paper, we present a new approach of using multiple distortion measures to schedule packets in a manner that explicitly accounts for a range of rates and resolutions. We show the effectiveness of this approach and examine a common scenario where multiple distortion measures are helpful. We present an algorithm that generates embedded JPEG2000 schedules and achieves 1 to 4 dB improvement over conventional approaches.
Carri W. Chan, Susie J. Wee, John G. Apostolopoulos
ICASSP (2)3
2007 A Joint Packet Selection/Omission and FEC System for Streaming Video
abstract
Media delivery over packet networks is often plagued by packet losses which limit its utility to end users. Forward error correction (FEC) based techniques are important for overcoming this problem. This paper further develops an FEC-based technique to maximize the expected received media quality by jointly choosing which packets to send and which packets to protect - including discarding packets to make additional room for protection. We describe a straight-forward implementation leveraging existing FEC system components. Comprehensive experiments demonstrate that significant gains in PSNR of several dB are achieved when sending H.264/MPEG-4 AVC coded video over a packet erasure channel.
Ying-zong Huang, John G. Apostolopoulos
ICASSP (1)2
2007 Rate-Distortion-Authentication Optimized Streaming with Multiple Deadlines
abstract
Video streaming with authentication is practically important, where a packet is decoded only when it is both received and authenticated. Recent work examined the problem of rate-distortion-authentication (R-D-A) optimized streaming of authenticated video. The original R-D-A technique assumes that each packet has only one deadline, its display deadline, and that a packet is not considered for transmission after its deadline. However, for video protected with an inter-packet graph-based authentication technique, a video packet can still be useful for verification of other packets even if it misses its own display deadline. We formulate the problem of multiple-deadline R-D-A optimized streaming and also propose ways to reduce the complexity. Simulation results using H.264 and NS-2 demonstrate that multiple-deadline R-D-A optimization achieves performance improvements of up to 4 dB over single-deadline R-D-A optimization.
Zhishou Zhang, Qibin Sun, Lawrence Wai-Choong Wong, John G. Apostolopoulos, Susie J. Wee
ICASSP (2)4
2007 Towards Quality of Service for Peer-to-Peer Video Multicast
abstract
Peer-to-peer streaming is a novel, low-cost, paradigm for large-scale video multicast. Viewers contribute their resources to an overlay network to act as relays for a real-time media stream. Early implementations fall short of the requirements of major content owners in terms of quality, reliability, and latency. In this work we show how adding a limited number of servers to a peer-to-peer streaming network can be used to enhance performance while preserving most of the benefits in terms of bandwidth cost savings. We present a theoretical model which is useful to estimate the number of servers needed to ensure fast connection times and improved error resilience. Experimental results show the proposed approach achieves 10times to 100times bandwidth cost savings compared to a content delivery network, and similar performance in terms of quality and startup latency.
Eric Setton, John G. Apostolopoulos
ICIP (5)2
2007 Lossless FMO and Slice Structure Modification for Compressed H.264 Video
abstract
We introduce a scheme to losslessly modify pre-compressed H.264 video to enable at streaming time (1) modification of slice sizes to fit transport packet size, and (2) introduction of an error resilience feature, namely flexible macroblock ordering. By lossless we mean the reconstructed video from the modified bitstream is identical to that from the original compressed bitstream. We outline how this transcoder operates, and discuss some of its restrictions. The bit rate overhead for operating on the pre-compressed, rather than directly encoding the original video with the desired characteristics, is analyzed. Simulation results show 1 to 2 dB improvement when our scheme is applied to QuickTime generated H.264 videos transported over a lossy packet network.
Wai-tian Tan, Eric Setton, John G. Apostolopoulos
ICIP (4)3
2007 Stream Authentication Based on Generlized Butterfly Graph
abstract
This paper proposes a stream authentication method based on the generalized butterfly graph (GBG) framework. Compared with the original Butterfly graph, the proposed GBG graph supports an arbitrary overhead budget and number of packets. Within the GBG framework, the problem of constructing an authentication graph is considered as a design problem: Given total number of packets, packet loss rate, and overhead budget, we show how to design the graph (number of rows and columns and edge allocation among nodes) to maximize the expected number of verified packets. In addition, we also propose a new evaluation metric called loss-amplification-factor (LAF), which measures the extent to which the authentication method exacerbates the effective packet loss rate. Experimental results demonstrate significant performance improvements over existing authentication methods like EMSS, augmented chain, and the original Butterfly.
Zhishou Zhang, John G. Apostolopoulos, Qibin Sun, Susie J. Wee, Lawrence Wai-Choong Wong
ICIP (6)2
2007 Optimal Scheduling of Media Packets with Multiple Distortion Measures
abstract
Due to the increase in diversity of wireless devices, streaming media systems must be capable of serving multiple types of users. Scalable coding allows for adaptations without re-encoding. To account for various viewing capabilities of each user, such as different spatial resolutions, multiple distortion measures are used. In this paper, we examine the question of how to broadcast media packets with multiple distortion measures to multiple users. We cast the problem as a stochastic shortest path problem and use Dynamic Programming to find the optimal policy. We generate an offline algorithm to generate the optimal transmission policy for the general case. We then show the optimal policy can be done online via a simple threshold policy for the case of independent Bernoulli packet losses. Through experimental results, we show that our policy, which considers multiple distortion measures, achieves up to 2dB gains over conventional approaches.
Carri W. Chan, Nicholas Bambos, Susie J. Wee, John G. Apostolopoulos
ICME4
2007 Rate-Distortion-Authentication Optimized Streaming of Authenticated Video
abstract
We define authenticated video as decoded video that results from those received packets whose authenticities have been verified. Generic data stream authentication methods usually impose overhead and dependency among packets for verification. Therefore, the conventional rate-distortion (R-D) optimized video streaming techniques produce highly sub-optimal R-D performance for authenticated video, since they do not account for the overhead and additional dependencies for authentication. In this paper, we study this practical problem and propose an Rate-Distortion-Authentication (R-D-A) optimized streaming technique for authenticated video. Based on packets' importance in terms of both video quality and authentication dependencies, the proposed technique computes a packet transmission schedule that minimizes the expected end-to-end distortion of the authenticated video at the receiver subject to a constraint on the average transmission rate. Simulation results based on H.264 JM 10.2 and NS-2 demonstrate that our proposed R-D-A optimized streaming technique substantially outperforms both prior (authentication-unaware) R-D optimized streaming techniques and data stream authentication techniques. In particular, when the channel capacity is below the source rate, the PSNR of authenticated video quickly drops to unacceptable levels using conventional R-D optimized streaming techniques, while the proposed R-D-A Optimization technique still maintains optimized video quality. Furthermore, we examine a low-complexity version of the proposed algorithm, and also an enhanced version which accounts for the multiple deadlines associated with each packet, which is introduced by stream authentication
Zhishou Zhang, Qibin Sun, Lawrence Wai-Choong Wong, John G. Apostolopoulos, Susie J. Wee
IEEE Trans. Circuits Syst. Video Technol.4
2007 An Optimized Content-Aware Authentication Scheme for Streaming JPEG-2000 Images Over Lossy Networks
abstract
This paper proposes an optimized content-aware authentication scheme for JPEG-2000 streams over lossy networks, where a received packet is consumed only when it is both decodable and authenticated. In a JPEG-2000 codestream, some packets are more important than others in terms of coding dependency and image quality. This naturally motivates allocating more redundant authentication information for the more important packets in order to maximize their probability of authentication and thereby minimize the distortion at the receiver. Towards this goal, with the awareness of its corresponding image content, we formulate an optimization framework to compute an authentication graph to maximize the expected media quality at the receiver, given specific authentication overhead and knowledge of network loss rate. System analysis and experimental results demonstrate that the proposed scheme achieves our design goal in that the rate-distortion (R-D) curve of the authenticated image is very close to the R-D curve when no authentication is required
Zhishou Zhang, Qibin Sun, Lawrence Wai-Choong Wong, John G. Apostolopoulos, Susie J. Wee
IEEE Trans. Multim.4
2006 Architectural Principles for Secure Streaming & Secure Adaptation in the Developing Scalable Video Coding (SVC) Standard
abstract
Scalable video coding has long been known to provide important functionalities such as low-complexity adaptation for diverse clients with different resources and for delivery over heterogeneous networks with time-varying available bandwidths. Recent advancements in scalable video coding have significantly improved the achievable compression performance, and the scalable video coding (SVC) standard is currently under intense development. An additional capability of scalable coding, which we believe has received less attention than it deserves, is the possibility with careful cross-layer design to support secure adaptive streaming at an untrustworthy sender and secure adaptation at an untrustworthy mid-network node. Building on the secure scalable streaming framework for video, and its realization within the JPEG-2000 security (JPSEC) standard, we describe the architectural principles and design details necessary so that SVC can also enable secure adaptive streaming at a sender and secure R-D optimized adaptation at an untrustworthy mid-network node.
John G. Apostolopoulos
ICIP1
2006 On Optimal Embedded Schedules of JPEG-2000 Packets
abstract
The JPEG-2000 compression standard codes images into data units, referred to as packets, such that images can be successfully decoded, while incurring some distortion, with a subset of these packets. This paper examines how to optimally select subsets of packets (referred to as schedules) that minimize distortion subject to varying rate constraints. We solve for the optimal schedule at a single rate by solving the precedence constraint knapsack problem (PCKP) via dynamic programming, and show that with modifications that consider the specific dependencies of JPEG-2000 packets, we can compute the optimal schedules for all rates through a single execution of our algorithm. We then analyze important properties of the optimal schedule. Using these properties, we look at a fused-greedy algorithm, similar to the recently proposed convex hull algorithm, to generate embedded schedules of JPEG-2000 packets. These embedded schedules have the property that all the JPEG-2000 packets in lower rate schedules are included in higher rate schedules; embedded schedules enable low-complexity, adaptive streaming. We demonstrate the algorithm's near optional performance through comparisons to the optimal performance for JPEG-2000 coded data.
Carri W. Chan, Susie J. Wee, John G. Apostolopoulos
ICIP3
2006 Making Packet Erasures to Improve Quality of FEC-Protected Video
abstract
Media delivery over lossy packet networks is a challenging problem, and forward error correction (FEC) based techniques are an important technique for overcoming packet loss. Conventional FEC-based media delivery techniques protect all packets equally, or protect a subset of the packets, or protect different subsets of packets with different levels of protection, e.g., scalable coding with unequal error protection (UEP). This paper proposes an FEC-based technique to maximize the expected received media quality by explicitly discarding packets, when beneficial, in order to provide additional room for FEC. Given knowledge of the importance of each packet, we show that there is a simple and intuitive criterion for the optimal selection of which packets to discard and which to protect, as well as the level of protection, to minimize the expected distortion experienced at the receiver. The proposed approach provides significant gains over the conventional approaches, and these gains are illustrated for the case of sending H.264 coded video data over a packet erasure channel with known packet loss rate.
Ying-zong Huang, John G. Apostolopoulos
ICIP2
2006 Rate-Distortion Optimized Streaming of Authenticated Video
abstract
Stream authentication methods usually impose overhead and dependency among packets. The straightforward application of state-of-the-art rate-distortion (R-D) optimized streaming techniques produce highly sub-optimal R-D performance for authenticated video, since they do not account for the additional dependencies. This paper proposes an R-D optimized streaming technique for authenticated video, by accounting for authentication dependencies and overhead. It schedules packet transmission based on packets' importance in terms of both video quality and authentication dependencies. The proposed technique works with any stream authentication method as long as the verification probability can be quantitatively computed from packet loss probability. Simulation results based on H.264 JM 10.1 and NS-2 demonstrate that the proposed authentication-aware R-D optimized streaming technique substantially outperforms authentication-unaware R-D optimized streaming techniques. In particular, when the channel capacity is below the source rate, the PSNR of authenticated video quickly drops to unacceptable levels using conventional R-D optimized streaming techniques, while the proposed technique still maintains R-D optimized video quality.
Zhishou Zhang, Qibin Sun, Lawrence Wai-Choong Wong, John G. Apostolopoulos, Susie J. Wee
ICIP4
2006 Receiver-Based Optimization for Video Delivery Over Wireless Links
abstract
We consider transfer of video frames over a time-varying wireless channel. When the channel is good, the transmitter can send frames at a higher rate than the receiver can consume them via playout. In that case, we introduce the idea of admitting new frames even when the receiver buffer is full, by selectively evicting frames already in the buffer; we can also control the playout rate, so as to optimize the tradeoff between video distortion and the time to freeze when the channel turns bad and frames arrive at a lower rate than should be played out. The decision/control problem of whether to admit a new frame, which already stored one to evict to accommodate the new one, and at what rate to play out frames is formulated within a dynamic programming framework, and an interesting connection to the Knapsack problem is made. Application of the idea in a relevant simple system shows significant performance gains, indicating that it is a promising approach for improving video delivery performance over challenging wireless channels
Carri W. Chan, John G. Apostolopoulos, Yan Li 0069, Nicholas Bambos
ICME2
2006 A Content-Aware Stream Authentication Scheme Optimized for Distortion and Overhead
abstract
This paper proposes a content-aware authentication scheme optimized to account for distortion and overhead for media streaming. When authenticated media is streamed over a lossy network, a received packet is consumed only when it is both decodable and authenticated. In most media formats, some packets are more important than others. This naturally motivates allocating more redundant authentication information for the more important packets in order to maximize their probability of authentication and thereby minimize distortion at the receiver. Toward this goal, with awareness of the media content, we formulate an optimization framework to compute an authentication graph to maximize the expected media quality at the receiver, given specific authentication overhead and knowledge of network loss rates. Experimental results with JPEG-2000 coded images demonstrate that the proposed method achieves our design goal in that the R-D curve of the authenticated image is very close to the R-D curve when no authentication is required
Zhishou Zhang, Qibin Sun, Lawrence Wai-Choong Wong, John G. Apostolopoulos, Susie J. Wee
ICME4
2006 The emerging JPEG-2000 security (JPSEC) standard
abstract
The emergence of digital imaging applications is accelerating the need for security of digital imagery. The emerging international standard ISO/IEC JPEG-2000 security (JPSEC) is designed to provide security for digital imagery, and in particular digital imagery coded with the JPEG-2000 image coding standard. This paper provides an overview of the JPSEC standard, including a description of its basic architecture and examples of its use.
John G. Apostolopoulos, Susie J. Wee, Frédéric Dufaux, Touradj Ebrahimi, Qibin Sun, Zhishou Zhang
ISCAS1
2006 Joint Power-Playout Control for Media Streaming Over Wireless Links
abstract
Media streaming applications over wireless links face various challenges, due to both the nature of the wireless channel and the stringent delivery requirements of media traffic. In this paper, we seek to improve the performance of media streaming over an interference-limited wireless link, by using appropriate transmission and playout control. In particular, we choose both the power at the transmitter and the playout scheduling at the receiver, so as to minimize the power consumption and maximize the media playout quality. We formulate the problem using a dynamic programming approach, and study the structural properties of the optimal solution. We further develop a justified, low-complexity heuristic that achieves significant performance gain over benchmark systems. In particular, our joint power-playout heuristic outperforms: 1) the optimal power control policy in the regime where power is most important and 2) the optimal playout control policy in the regime where media (playout) quality is most important; furthermore, this heuristic has only a slight performance loss as compared to the optimal joint power-playout control policy over the entire range of the investigation
Yan Li 0069, Athina Markopoulou, Nicholas Bambos, John G. Apostolopoulos
IEEE Trans. Multim.4
2005 End-to-end rate-distortion optimized mode selection for multiple description video coding
abstract
Multiple description (MD) video coding can be used to reduce the detrimental effects caused by transmission over lossy packet networks. Each approach to MD coding consists of a tradeoff between compression efficiency and error resilience. How effectively each method achieves this tradeoff depends on the network conditions as well as on the characteristics of the video itself. This paper proposes an adaptive MD coding approach which adjusts to these conditions through the use of adaptive MD mode selection. The encoder in this system is able to accurately estimate the expected end-to-end distortion, accounting for both coding and packet-loss-induced distortions, as well as for the bursty nature of channel losses and the effective use of multiple transmission paths. With this model of end-to-end expected distortion, the encoder selects between MD coding modes in a rate-distortion optimized manner to most effectively trade-off compression efficiency for error resilience. We show how this approach adapts to the local characteristics of the video as well as to current network conditions and demonstrate the resulting gains in performance.
Brian A. Heng, John G. Apostolopoulos, Jae S. Lim
ICASSP (5)2
2005 Enterprise Streaming: Different Challenges from Internet Streaming
abstract
Media streaming over the best-effort public Internet has been a focus of research for over a decade. Enterprise or corporate streaming is another area of media streaming that is practically very important and has a different set of challenges and feasible solutions. For example, the quality and reliability requirements for enterprise streaming are much stricter than for typical Internet streaming. Furthermore, in typical enterprise streaming scenarios, a single entity has control over most elements of the system, including the end-points and the infrastructure. This entity has the powerful ability to monitor, adapt, and deploy new infrastructure as necessary. The goal of this paper is to describe enterprise streaming and identify the basic differences between typical Internet streaming and enterprise streaming, and how these differences alter the challenges that must be overcome for enterprise streaming to be successful. Specifically, we examine enterprise streaming media content delivery network design and operation, video conferencing, peer-to-peer networking (P2P), voice over IP (VoIP), and briefly touch upon wireless and security issues
John G. Apostolopoulos, Mitchell D. Trott, Ton Kalker, Wai-tian Tan
ICME1
2005 Examining Memory in Reconstruction Distortion: Dropping Additional Packets to Improve Video Quality
abstract
The source coding process and the packet loss process create certain dependencies between encoded video units in terms of the reconstruction distortion of the video signal at the receiver in case of transmission over packet erasure channels. In this paper, we examine the importance of this "distortion memory" via a specific class of memory-based models denoted Distortion Chains that are used for predicting the distortion of the reconstructed video signal in case of missing multiple packets at the receiver. We show that taking into account even the smallest amount of memory that is possible can yield substantial gains in terms of prediction accuracy and packet selection (packet dropping) performance. An additional and rather surprising result of our study is the fact that in certain situations dropping an additional video packet (which could otherwise be delivered) can actually improve the quality of the reconstructed video
Jacob Chakareski, John G. Apostolopoulos
MMSP2
2005 Joint Packet Scheduling and Content-Aware Playout Control for Video Streaming over Wireless Links
abstract
Media streaming over wireless links is a challenging problem due to both the unreliable, time-varying nature of the wireless channel and the stringent delivery requirements of media traffic. In this paper, we use joint control of packet scheduling at the transmitter and content-aware playout at the receiver, so as to maximize the quality of media streaming over a wireless link. Our contributions are twofold. First, we formulate and study the problem of joint scheduling and playout control within a dynamic programming framework. Second, we propose a novel content-aware playout control, that takes into account the content of a video sequence, and in particular the motion characteristics of different scenes. We find that the joint scheduling and playout control can significantly improve the quality of the received video, at the expense of only a small amount of playout slowdown. Furthermore, thanks to the content-aware playout, the slowdown takes place mainly in the low-motion scenes, where its perceived effect is limited
Yan Li 0069, Athina Markopoulou, John G. Apostolopoulos, Nicholas Bambos
MMSP3
2005 Real-time monitoring of video quality in IP networks
abstract
This paper investigates the problem of assessing the quality of video transmitted over IP networks. Our goal is to develop a methodology that is both reasonably accurate and simple enough to support the large-scale deployments that the increasing use of video over IP are likely to demand. For that purpose, we focus on developing an approach that is capable of mapping network statistics, e.g., packet losses, available from simple measurements, to the quality of video sequences reconstructed by receivers. A first step in that direction is a loss-distortion model that accounts for the impact of network losses on video quality, as a function of application-specific parameters such as the video codec and loss recovery technique, coded bit rate, packetization, video characteristics, etc. The model, although accurate, is poorly suited to large-scale, on-line monitoring, because of its dependency on many parameters that are difficult to estimate in real-time. As a result, we introduce a "relative quality" metric that bypasses this problem by measuring video quality against a quality benchmark that the network is expected to provide. The approach offers a lightweight video quality monitoring solution that is suitable for large-scale deployments. We assess its feasibility and accuracy through extensive simulations and experiments.
Shu Tao, John G. Apostolopoulos, Roch Guérin
NOSSDAV2
2005 Rate-distortion hint tracks for adaptive video streaming
abstract
We present a technique for low-complexity rate-distortion (R-D) optimized adaptive video streaming based on the concept of rate-distortion hint track (RDHT). RDHTs store the precomputed characteristics of a compressed media source that are crucial for high performance online streaming but difficult to compute in real time. This enables low-complexity adaptation to variations in transport conditions such as available data rate or packet loss. An RDHT-based streaming system has three components: 1) information that summarizes the R-D attributes of the media; 2) an algorithm for using the RDHT to predict the distortion for a feasible packet schedule; and 3) a method for determining the best packet schedule to adapt the streaming to the communication channel. A family of distortion models, denoted distortion chains, are presented which accurately predict the distortion produced by arbitrary packet loss patterns. Two distortion chain models are examined which lead to two RDHT-based techniques. We evaluate the proposed techniques for two canonical problems in streaming media, adaptation to available data rate and to packet loss. Experimental results demonstrate that for the difficult case of nonscalably coded H.264 video, the proposed systems provide significant performance gains over conventional low-complexity streaming systems, and achieve this gain with a comparable level of complexity making them suitable for online R-D optimized streaming.
Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.2
2005 Optical flow estimation using temporally oversampled video
abstract
Recent advances in imaging sensor technology make high frame-rate video capture practical. As demonstrated in previous work, this capability can be used to enhance the performance of many image and video processing applications. The idea is to use the high frame-rate capability to temporally oversample the scene and, thus, to obtain more accurate information about scene motion and illumination. This information is then used to improve the performance of image and standard frame-rate video applications. This paper investigates the use of temporal oversampling to improve the accuracy of optical flow estimation (OFE). A method for obtaining high accuracy optical flow estimates at a conventional standard frame rate, e.g., 30 frames/s, by first capturing and processing a high frame-rate version of the video is presented. The method uses the Lucas-Kanade algorithm to obtain optical flow estimates at a high frame rate, which are then accumulated and refined to estimate the optical flow at the desired standard frame rate. The method demonstrates significant improvements in OFE accuracy both on synthetically generated video sequences and on a real video sequence captured using an experimental high-speed imaging system. It is then shown that a key benefit of using temporal oversampling to estimate optical flow is the reduction in motion aliasing. Using sinusoidal input sequences, the reduction in motion aliasing is identified and the desired minimum sampling rate as a function of the velocity and spatial bandwidth of the scene is determined. Using both synthetic and real video sequences, it is shown that temporal oversampling improves OFE accuracy by reducing motion aliasing not only for areas with large displacements but also for areas with small displacements and high spatial frequencies. The use of other OFE algorithms with temporally oversampled video is then discussed. In particular, the Haussecker algorithm is extended to work with high frame-rate sequences. This extension demonstrates yet another important benefit of temporal oversampling, which is improving OFE accuracy when brightness varies with time.
Sukhwan Lim, John G. Apostolopoulos, Abbas El Gamal
IEEE Trans. Image Process.2
2005 Source-channel diversity for parallel channels
abstract
We consider transmitting a source across a pair of independent, nonergodic channels with random states (e.g., slow-fading channels) so as to minimize the average distortion. The general problem is unsolved. Hence, we focus on comparing two commonly used source and channel encoding systems which correspond to exploiting diversity either at the physical layer through parallel channel coding or at the application layer through multiple description (MD) source coding. For on-off channel models, source coding diversity offers better performance. For channels with a continuous range of reception quality, we show the reverse is true. Specifically, we introduce a new figure of merit called the distortion exponent which measures how fast the average distortion decays with signal-to-noise ratio. For continuous-state models such as additive white Gaussian noise (AWGN) channels with multiplicative Rayleigh fading, optimal channel coding diversity at the physical layer is more efficient than source coding diversity at the application layer in that the former achieves a better distortion exponent. Finally, we consider a third decoding architecture: MD encoding with joint source-channel decoding. We show that this architecture achieves the same distortion exponent as systems with optimal channel coding diversity for continuous-state channels, and maintains the advantages of MD systems for on-off channels. Thus, the MD system with joint decoding achieves the best performance from among the three architectures considered, on both continuous-state and on-off channels.
J. Nicholas Laneman, Emin Martinian, Gregory W. Wornell, John G. Apostolopoulos
IEEE Trans. Inf. Theory4
2004 Distortion chains for predicting the video distortion for general packet loss patterns
abstract
When designing a system for video communication over a lossy packet network, it is highly beneficial to have a mechanism for accurately predicting the mean-squared error (MSE) distortion that results from different packet loss patterns. The paper proposes a distortion chains model for accurately predicting the end-to-end distortion for different general packet loss patterns. The performance is examined using JVT/H.264 encoded video sequences and previous frame error concealment. It is shown that, for all tested sequences, the proposed model predicts the total distortion due to a packet loss pattern within a 10% error bound 80% of the time, as compared to the conventional additive approach which achieves the same accuracy less then 40% of the time.
Jacob Chakareski, John G. Apostolopoulos, Wai-tian Tan, Susie J. Wee, Bernd Girod
ICASSP (5)2
2004 Secure media streaming & secure adaptation for non-scalable video
John G. Apostolopoulos
ICIP1
2004 Low-complexity rate-distortion optimized video streaming
Jacob Chakareski, John G. Apostolopoulos, Bernd Girod
ICIP2
2004 Benefits of temporal oversampling in optical flow estimation
abstract
Recently it has been shown that optical flow estimation (OFE) accuracy can benefit from temporal oversampling especially for large displacements between frames. In this paper, we show that temporal oversampling also benefits OFE in the case of complex scenes with small displacements but high spatial bandwidth. Using synthetic test sequences and a high-speed real video sequence, it is shown that temporal oversampling improves the performance of OFE by reducing motion aliasing not only for areas with large displacements but also for areas with small displacements but with high spatial frequencies. We also demonstrate that the minimum frame rate necessary to achieve good OFE performance for the tested sequences is largely determined by the minimum frame rate necessary to prevent motion aliasing.
Sukhwan Lim, John G. Apostolopoulos, Abbas El Gamal
ICIP2
2004 Secure transcoding with JPSEC confidentiality and authentication
abstract
The emerging JPEG-2000 part 8 security standard (JPSEC) is being defined to provide security services for JPEG-2000 images. This paper describes how confidentiality and authentication can be applied in a manner that allows mid-network adaptation of protected JPSEC streams while preserving end-to-end security. We achieved this by designing the JPSEC syntax to support the principles of secure scalable streaming and secure transcoding. Specifically, we designed JPSEC encryption methods and signaling syntax that enable an entity to securely adapt or transcode the resulting JPSEC-protected stream without requiring decryption. We discuss tradeoffs in protection, transcoding flexibility, and complexity for the different encryption methods. Furthermore, we show how authentication can be applied to verify that the secure transcoding operation was performed in a valid and permissible manner.
Susie J. Wee, John G. Apostolopoulos
ICIP2
2004 R-D hint tracks for low-complexity R-D optimized video streaming
abstract
This work presents the concept of rate-distortion hint track (RDHT), and evaluates two specific implementations of streaming systems that employ RDHT. Using RDHT, low-complexity streaming can be realized for systems that adapt to variations in transport conditions such as bandwidth or packet loss. An RDHT-based streaming system has three components: (1) an R-D hint track; (2) an algorithm for using the RDHT to predict the distortion for different packet schedules; and (3) a method for determining the best packet schedule. Two RDHT-based systems are presented which perform R-D optimized scheduling with dramatically reduced complexity as compared to conventional on-line R-D optimized streaming algorithms. Experimental results demonstrate that for the difficult case of R-D optimized scheduling of non-scalably coded video the proposed systems provide 7-12 dB gain when adapting to a bandwidth constraint and 2-4 dB gain when adapting to random packet loss, both relative to a conventional streaming system that does not take into account the different importance of individual packets.
Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod
ICME2
2004 WiSE video: using in-band wireless loss notification to improve rate-controlled video streaming
abstract
Both data and multimedia applications over the Internet are expected to perform some kind of congestion control, typically using packet loss as an indication for congestion. However, when packets are lost due to wireless errors, decreasing the rate unnecessarily harms the application's performance. The paper proposes to use an in-band notification mechanism, called WiSE (wireless signaling via ECN), to distinguish wireless errors from congestion losses and improve the performance of rate-controlled video streamed over wireless links. A WiSE agent on the wireless network identifies wireless errors and piggy-backs this information onto other video packets. The WiSE-aware video source can benefit from this notification (1) by avoiding unnecessary decreases in the sending rate in response to wireless errors, and (2) by accurately adjusting the error resilience for the wireless link. Simulations demonstrate that WiSE provides a significant improvement in video quality over a wide range of conditions.
Athina Markopoulou, Eric Setton, M. Kalman, John G. Apostolopoulos
ICME4
2004 Semantic-enhanced distribution & adaptation networks
abstract
Recent years have witnessed significant efforts in deriving and embedding semantic information in content for improved content retrieval, adaptation, and distribution. Relatively little work considers leveraging this semantic information in the infrastructure to better serve the needs of clients and achieve better cost-effectiveness of infrastructure resources. In this paper, we identify semantics that can be derived and extracted from the various components in a content distribution infrastructure, namely content semantics (from content source), infrastructure semantics and client semantics (from content consumer). We develop a semantic-enhanced distribution and adaptation framework (SEDAN) that achieves superior efficiency and provides content access and adaptation features that were not previously possible
Bo Shen 0003, Zhichen Xu, Susie J. Wee, John G. Apostolopoulos
ICME4
2004 Semantic-enhanced distribution and adaptation networks
abstract
Recent years have witnessed significant efforts in deriving and embedding semantic information in content for improved content retrieval, adaptation, and distribution. Relatively little work considers leveraging this semantic information in the infrastructure to serve the needs of clients better and achieve better cost-effectiveness of infrastructure resources. We identify semantics that can be derived and extracted from the various components in a content distribution infrastructure, namely content semantics (from content source), infrastructure semantics and client semantics (from content consumer). We develop a semantic-enhanced distribution and adaptation framework (SEDAN) that achieves superior efficiency and provides content access and adaptation features that were not previously possible.
Bo Shen 0003, Zhichen Xu, Susie J. Wee, John G. Apostolopoulos
ICME4
2004 Divert: Fine-grained Path Selection for Wireless LANs
abstract
The performance of Wireless Local Area Networks (WLANs)often suffers from link-layer frame losses caused by noise, interference, multipath, attenuation, and user mobility. We observe that frame losses often occur in bursts and that three of the five main causes of frame losses -- multipath, attenuation, mobility--depends on the transmission path traversed between an access point (AP)and a client station.In a typical WLAN deployment, different transmission paths to a client exist in places where overlapping coverage is provided by a set of neighboring APs. Using experimental measurements and analysis on a 802.11b testbed, we show that fine-grained path selection among a set of neighboring APs can significantly reduce path-dependent losses in WLANs.We design and implement a WLAN distribution system called Divert, which supports fine-grained path selection for downlink communications, on an 802.11b testbed. Divert reduces frame losses without consuming any extra bandwidth in the wireless medium. Our experimental results show that Divert can reduce frame loss rates in realistic scenarios by as much as 26% compared to a fixed-path scheme that uses the best available transmitter.
Allen K. L. Miu, Godfrey Tan, Hari Balakrishnan, John G. Apostolopoulos
MobiSys4
2003 Analysis of packet loss for compressed video: does burst-length matter?
abstract
Video communication is often afflicted by various forms of losses, such as packet loss over the Internet. The paper examines the question of whether the packet loss pattern, and in particular the burst length, is important for accurately estimating the expected mean-squared error distortion. Specifically, we (1) verify that the loss pattern does have a significant effect on the resulting distortion, (2) explain why a loss pattern, for example a burst loss, generally produces a larger distortion than an equal number of isolated losses, and (3) propose a model that accurately estimates the expected distortion by explicitly accounting for the loss pattern, inter-frame error propagation, and the correlation between error frames. The accuracy of the proposed model is validated with JVT/H.26L coded video and previous frame concealment, where for most sequences the total distortion is predicted to within /spl plusmn/0.3 dB for burst loss of length two packets, as compared to prior models which underestimate the distortion by about 1.5 dB. Furthermore, as the burst length increases, our prediction is within /spl plusmn/0.7 dB, while prior models degrade and underestimate the distortion by over 3 dB.
Yi J. Liang, John G. Apostolopoulos, Bernd Girod
ICASSP (5)2
2003 Comparing application- and physical-layer approaches to diversity on wireless channels
abstract
Diversity techniques often arise as appealing means for improving the performance of multimedia communication over certain types of channels with independent parallel components (e.g., multiple antennas, frequency bands or time slots). Diversity can be obtained by channel coding across parallel components at the physical layer. Alternatively, the physical layer ca present an interface to the parallel components as separate, independent links thus allowing the application layer to implement diversity in the form of multiple description source coding. We compare these two approaches in terms of average end-to-end distortion as a function of channel signal-to-noise ratio (SNR). When specialized to the case of an independent, identically distributed Gaussian source over Rayleigh fading channels, our results suggest that parallel channel coding at the physical layer is more efficient than independent channel coding combined with multiple description source coding. More generally, we provide intuitive guidelines for allowing system designers to identify which types of systems are preferable under different scenarios of practical interest.
J. Nicholas Laneman, Emin Martinian, Gregory W. Wornell, John G. Apostolopoulos, Susie J. Wee
ICC4
2003 Secure scalable streaming and secure transcoding with JPEG-2000
abstract
Secure scalable streaming (SSS) enables low-complexity, high-quality transcoding at intermediate, possibly untrusted, network nodes without compromising the end-to-end security of the system S. J. Wee, J. G. Apostolopoulos (2001). SSS encodes, encrypts, and packetizes video into secure scalable packets in a manner that allows downstream transcoders to perform transcoding operations such as bitrate reduction and spatial downsampling by simply truncating or discarding packets, and without decrypting the data. Secure scalable packets have unencrypted headers that provide hints such as optimal truncation points to downstream transcoders. Using these hints, downstream transcoders can perform near-optimal secure transcoding. This paper presents a secure scalable streaming system based on motion JPEG-2000 coding with AES or triple-DES encryption. The operational rate-distortion (R-D) performance for transcoding to various resolutions and quality levels is evaluated, and results indicate that end-to-end security and secure transcoding can be achieved with near R-D optimal performance. The average overhead is 4.5% for triple-DES encryption and 7% for AES, as compared to the original media coding rate, and only 2-2.5% overhead as compared to end-to-end encryption which does not allow secure transcoding.
Susie J. Wee, John G. Apostolopoulos
ICIP (1)2
2003 Low-latency wireless video over 802.11 networks using path diversity
abstract
Wireless local area networks, such as 802.11b, are becoming wide-spread as they provide simple wireless connectivity and data delivery. This paper examines low-latency (conversational) video communication over 802.11b networks. The challenges to enable low-latency video include overcoming the highly variable delays, losses, and bandwidth of 802.11b wireless networks. To overcome these challenges we (1) employ the H.264/MPEG-4 advanced video coding (AVC) standard for high video compression efficiency and good resilience to losses, (2) use low-latency best-effort transport mechanisms, and (3) exploit the potential path diversity between each mobile client and multiple access points in the infrastructure, where we use multiple paths simultaneously or switch between multiple paths (site selection) as a function of channel characteristics. Our results indicate that the proposed system can provide significant benefits over conventional single access point (single path) systems.
Allen K. L. Miu, John G. Apostolopoulos, Wai-tian Tan, Mitchell D. Trott
ICME2
2003 Research and design of a mobile streaming media content delivery network
abstract
Delivering media to large numbers of mobile users presents challenges due to the stringent requirements of streaming media, mobility, wireless, and scaling to support large numbers of users. This paper presents a mobile streaming media content delivery network (MSM-CDN) designed to overcome these challenges. The MSM-CDN is a network overlay consisting of overlay servers on top of the existing network; these overlay servers are control points that facilitate end-to-end media delivery and mid-network media services. This paper presents an overview of the MSM-CDN system architecture, and describes the testbed prototype that we built based on these architectural principles. The MSM-CDN provides a new platform for media delivery, and we describe a number of research directions related to the MSM-CDN.
Susie J. Wee, John G. Apostolopoulos, Wai-tian Tan, Sumit Roy 0002
ICME2
2002 Modeling path diversity for multiple description video communication
abstract
The use of multiple description (MD) video coding and path diversity has been proposed to provide improved performance over lossy packet networks [1]. The goal of this work was to develop models to accurately and quickly predict and compare the distortion of MD video coding and path diversity against conventional single description (SD) video delivered over a single path. In the process, we developed (1) a model for the loss process of a two-path path diversity system, and (2) a distortion model that maps the loss model to MD distortion values. Given these models we present a number of comparisons between MD video coding and path diversity and conventional SD video over a single path. The proposed model for path diversity may also be useful in other applications not related to MD coding. Furthermore, other forms of MD coding may be analyzed using similar models for MD distortion.
John G. Apostolopoulos, Wai-tian Tan, Susie J. Wee, Gregory W. Wornell
ICASSP1
2002 Performance of a multiple description streaming media content delivery network
abstract
Content delivery networks (CDN) have been widely used to provide reduced delay and packet loss, fault tolerance, and improved scalability for Web content delivery. Additional benefits are provided for video streaming when one designs a streaming media CDN (SM-CDN) for either conventional single description (SD) or multiple description (MD) coding. Specifically, when precise network conditions and topology are known, simulations show that an MD-SM-CDN can provide 20 to 40% reduction in distortion over a conventional SD-SM-CDN, even when the underlying CDN is not designed with MD streaming in mind. This paper examines the performance of an MD-SM-CDN as a function of different network topologies and loss conditions, and compares it with a conventional SD-SM-CDN. This examination provides insight into an MD-SM-CDN's performance when knowledge of network topology and conditions is imprecise or uncertain. Our simulations indicate that an MD-SM-CDN can provide improved performance over a conventional SD-SM-CDN over a wide range of network topologies and loss conditions.
John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan
ICIP (2)1
2002 Optimized video streaming for networks with varying delay
abstract
This paper presents a method for distortion-optimized streaming of predictively coded video over packet networks with varying delay. In networks with significant delay variations, coded video frames can arrive late at the decoder and miss their respective display deadlines. Furthermore, due to predictive coding, a late frame can also prevent a number of subsequent frames from being displayed properly, where the number of affected frames or degree of distortion depends on the particular coding dependencies of the late frame. In this paper, we present an optimized video streaming strategy based on frame reordering for networks with significant delay variations. This streaming strategy minimizes distortion by exploiting the fact that different late frames result in different degrees of distortion. We model the router-induced delay in a wired network with an analytical PDF and we model the link-layer retransmission delay of a wireless network with the 3GPP specification for W-CDMA radio link control. We compute the distortion for different frame reorderings using the network delay models and a source model that accounts for the prediction dependencies of predictively coded video. Our optimized streaming strategies are shown to reduce the number of late frames by 14 to 23% for the situations examined.
Susie J. Wee, Wai-tian Tan, John G. Apostolopoulos, Minoru Etoh
ICME (2)3
2002 On Multiple Description Streaming with Content Delivery Networks
abstract
We propose a system that improves the performance of streaming media CDN by exploiting the path diversity provided by existing CDN infrastructure. Path diversity is provided by the different network paths that exist between a client and its nearby edge servers; and multiple description (MD) coding is coupled with this path diversity to provide resilience to losses. In our system, MD coding is used to code a media stream into multiple complementary descriptions, which are distributed across the edge servers in the CDN. When a client requests a media stream, it is directed to multiple nearby servers which host complementary descriptions. These servers simultaneously stream these complementary descriptions to the client over different network paths. This paper provides distortion models for MDC video and conventional video. We use these models to select the optimal pair of servers with complementary descriptions for each client while accounting for path lengths and path jointness and disjointness. We also use these models to evaluate the performance of MD streaming over CDN in a number of real and generated network topologies. Our results show that distortion reduction by about 20 to 40% can be realized even when the underlying CDN is not designed with MDC streaming in mind. Also, for certain topologies, MDC requires about 50% fewer CDN servers than conventional streaming techniques to achieve the same distortion at the clients.
John G. Apostolopoulos, Tina Wong, Susie J. Wee
INFOCOM1
2001 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221)
abstract
We present a wireless video streaming system that securely and efficiently streams video to heterogeneous clients over timevarying communication links. Clients may differ in their display, power, communication, and computational capabilities and wireless channels may have time-varying bandwidths and quality levels that depend on channel usage and channel conditions. End-to-end system efficiency is achieved by placing transcoders at intermediate network nodes; these transcoders can easily adapt the video stream for particular client capabilities and network conditions.
Susie J. Wee, John G. Apostolopoulos
ICASSP2
2001 Secure scalable streaming enabling transcoding without decryption
abstract
We present a method of secure scalable streaming (SSS) that enables low-complexity and high-quality transcoding to be performed at intermediate, possibly untrusted, network nodes without compromising the end-to-end security of the system. SSS encodes video into secure scalable packets using jointly designed scalable coding and progressive encryption techniques. This combination allows downstream transcoders to perform transcoding operations such as bitrate reduction and spatial downsampling by simply truncating or discarding packets, and without decrypting the data. Secure scalable packets have unencrypted headers that can provide hints such as optimal truncation points to downstream transcoders. Using these hints, downstream transcoders can perform RD-optimal transcoding for fine-grain bitrate reduction. The SSS transcoding operation has low complexity and is stateless, so SSS transcoders can support many simultaneous transcoding sessions. SSS works with existing scalable image and video compression standards and systems including Motion JPEG-2000, 3D subband coding, and MPEG-4 FGS.
John G. Apostolopoulos, Susie J. Wee
ICIP (1)1
2001 Unbalanced multiple description video communication using path diversity
abstract
Multiple description (MD) coders provide important error resilience properties. Specifically, MD coders are designed to provide good performance when the loss is limited to a single description, but it is not known in advance which description. Apostolopoulos (2001) combined MD video coding with a path diversity transmission system for packet networks such as the Internet, where different descriptions are explicitly transmitted through different network paths, to improve the effectiveness of MD coding over a packet network by increasing the likelihood that the loss probabilities for each description are independent. The available bandwidth in each path may be similar or different, resulting in the requirement of balanced or unbalanced operation, where the bit rate of each description may differ based on the available bandwidth along its path. We design a MD video communication system that is effective in both balanced and unbalanced operation. Specifically, unbalanced MD streams are created by carefully adjusting the frame rate of each description, thereby achieving unbalanced rates of almost 2:1 while preserving MD's effectiveness and error recovery capability.
Susie J. Wee, John G. Apostolopoulos
ICIP (1)2
2001 Reliable video communication over lossy packet networks using multiple state encoding and path diversity
John G. Apostolopoulos
VCIP1
2000 Error-Resilient Video Compression Through the Use of Multiple States
abstract
Video compression enables a number of applications by reducing the required bit rate needed to represent a video sequence, however the compressed video is much more susceptible to errors such as bit errors or packet loss. Conventional video compression standards employ an architecture which we refer to as single-state systems since they have a prediction loop with a single state (e.g. the previous decoded frame) which if lost or corrupted can lead to the loss or severe degradation of all subsequent frames until the state is reinitialized (the prediction is refreshed). We combat this problem of incorrect state and error propagation at the decoder by coding the video into multiple independently decodable streams, each with its own prediction process and state, such that if one stream is lost the other streams can still be used to produce usable video. The correctly received streams provide improved error concealment and, more importantly, enable faster state recovery for the lost stream. This approach is conceptually similar to multiple description coding, however it differs in the representation used for each description as well as its use of state recovery.
John G. Apostolopoulos
ICIP1
1999 Field-To-Frame Transcoding with Spatial and Temporal Downsampling
abstract
We present an algorithm for transcoding high-rate compressed bitstreams containing field-coded interlaced video to lower-rate compressed bitstreams containing frame-coded progressive video. We focus on MPEG-2 to H.263 transcoding, however these results can be extended to other lower-rate video compression standards including MPEG-4 simple profile and MPEG-1. A conventional approach to the transcoding problem involves decoding the input bitstream, spatially and temporally downsampling the decoded frames, and re-encoding the result. The proposed transcoder achieves improved performance by exploiting the details of the MPEG-2 and H.263 compression standards when performing interlaced to progressive (or field to frame) conversion with spatial downsampling and frame-rate reduction. The transcoder reduces the MPEG-2 decoding requirements by temporally downsampling the data at the bitstream level and reduces the H.263 encoding requirements by largely bypassing H.263 motion estimation by reusing the motion vectors and coding modes given in the input bitstream. In software implementations, the proposed approach achieved a 5/spl times/ speedup over the conventional approach with only a 0.3 and 0.5 dB loss in PSNR for the Carousel and Bus sequences.
Susie J. Wee, John G. Apostolopoulos, Nick Feamster
ICIP (4)2
1999 Postprocessing for very low bit-rate video compression
abstract
This paper presents a novel postprocessing algorithm developed specifically for very low bit-rate MC-DCT video coders operating at low spatial resolution, postprocessing is intricate in this situation because the low sampling rate (as compared to the image feature size) makes it very easy to overfilter, producing excessive blurring. The proposed algorithm uses pixel-by-pixel processing to identify and reduce both blocking artifacts and mosquito noise while attempting to preserve the sharpness and naturalness of the reconstructed video signal and minimize the system complexity. Experimental results show that the algorithm successfully reduces artifacts in a 16 kb/s scene-adaptive coder for video signals sampled at 80 x 112 pixels per frame and 5-10 frames/s. Furthermore, the portability of the proposed algorithm to other block-DCT based compression systems is shown by applying it, without modification, to successfully post-process a JPEG-compressed image.
John G. Apostolopoulos, Nikil Jayant
IEEE Trans. Image Process.1
1998 Critically sampled wavelet representations for multidimensional signals with arbitrary regions of support
abstract
Transform/subband representations form an important element of many signal processing algorithms and applications. Until recently, representations have typically been designed for signals with convenient supports, e.g. 2-D signals with rectangular supports. However, a number of applications require representations for signals with arbitrary (non-rectangular) regions of support. We present a novel algorithm for creating critically sampled perfect reconstruction wavelet representations for signals defined over arbitrary supports. The proposed algorithm selects a subset of vectors from a convenient superset basis which under appropriate conditions provides a basis over the given arbitrary support. The algorithm can be interpreted as solving a corresponding sampling problem.
John G. Apostolopoulos, Jae S. Lim
ICASSP1
1997 Transform/subband representations for signals with arbitrarily shaped regions of support
abstract
Transform/subband representations form a basic building block for many signal processing algorithms and applications. Most of the effort has focused on developing representations for infinite-length signals, with simple extensions to finite-length 1-D and rectangular support 2-D signals. However, many signals may have arbitrary length or arbitrarily shaped (AS) regions of support (ROS). We present a novel framework for creating critically sampled perfect reconstruction transform/subband representations for AS signals. Our method selects an appropriate subset of vectors from an (easily obtained) basis for a larger (superset) signal space, in order to form a basis for the AS signal. In particular, we have developed a number of promising wavelet representations for arbitrary-length l-D signals and AS 2-D/M-D signals that provide high performance with low complexity.
John G. Apostolopoulos, Jae S. Lim
ICASSP1
1995 Representing arbitrarily-shaped regions: a case study of overcomplete representations
abstract
Efficiently representing the interior of an arbitrarily-shaped region is an important problem in many applications, including object- or region-based image/video compression. This paper focuses on representing the interior as a linear combination of vectors defined on a superset basis. Two seemingly different, though highly related, problem formulations are given and a number of approaches that result are discussed and analyzed. A geometric interpretation of the problem is given to provide insight into the approaches and examine their properties. The incorporation of quantization is also briefly discussed. In addition, the chosen set of vectors corresponds to a special/structured class of overcomplete representations with O(N log/sub 2/ N)-type processing, low memory requirements, and other important properties.
John G. Apostolopoulos, Jae S. Lim
ICIP1
1994 Position-dependent encoding
abstract
In typical MC-DCT compression algorithms, up to 90% of the available bit rate is used to encode the location and amplitude information of the nonzero quantized DCT coefficients. Therefore, efficient encoding of the location and amplitude information is extremely important for high quality video compression. In this paper we describe the position-dependent encoding (PDE) approach for encoding the DCT coefficients. This method attempts to exploit the inherent differences in statistical properties of the run-lengths and amplitudes as a function of position. This paper includes preliminary comparisons between the PDE scheme and conventional coding approaches. The PDE approach can be applied in any transform/subband filtering scheme for image or video compression.>
John G. Apostolopoulos, Aleksandar Pfajfer, Hae Mook Jung, Jae S. Lim
ICASSP (5)1
1993 Designing a video compression system for high definition television
John G. Apostolopoulos, Peter A. Monta, Julien J. Nicolas, Jae S. Lim
ICASSP (1)1