EDBT 2026 Demo / reviewers in the wild / expert
Wai-tian Tan
dblp:76/1151 · also Wai-Tian Tan
· DBLP profile ↗
88ranked-venue papers
18as first author
15since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 16 first-author · 2 since 2021Computer networks · 12 · 2 first-author · 4 since 2021Theory of computation · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Streaming Erasure Codes over Multicast Relayed NetworksabstractThis paper studies streaming erasure codes in a relayed multicast setting, where a source wishes to transmit a sequence of messages to two different destinations through a common relay. Our construction extends previously proposed works on the single-destination setting studied in Fong et al. and Facenda et al. to the multicast setting, where each destination can recover the source packets with a correspondingly different delay. A key property of our construction is that it does not require prior knowledge of the maximum number of erasures on the relay-destination link. Instead, it enables the recovery of the source stream with a decoding delay that depends on the number of erasures on the relay-destination links. We demonstrate that if some divisibility conditions are satisfied, then the proposed construction can simultaneously achieve the single-destination delay in Fong et al. for both receivers. Finally we also explain how our proposed construction can be applied in the setting of a single destination when the maximum number of erasures on the relay-destination link is not known beforehand. Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 3 |
| 2023 | Deep Reinforcement Learning for Latency-Sensitive Communication With Adaptive Redundant RetransmissionsabstractThis paper studies packet repetition strategies over erasure channels with memory and a long feedback delay. The problem is initially formulated as a communications problem where a source wishes to transmit one message packet to a destination while minimizing both the delay and the number of transmissions. At each time instant, the sender is provided a delayed acknowledgement feedback about past attempts, and must decide whether to attempt a new transmission or not. This problem is then re-formulated as an episodic reinforcement learning problem, where an agent attempts to learn the optimal transmission policy, provided delayed feedback about past transmission attempts. The agent is helped by a channel estimator, which attempts to capture the channel memory and use that to predict probabilities of erasures in a future window. This channel estimator is also data-driven and learns the channel model without anya priorichannel knowledge. The paper presents a lower bound on the achievable trade-off between delay and number of transmissions for any channel modeled as a Markov process. Experimental results show that the combination of the proposed channel estimator and the agent can noticeably outperform naive strategies for channels with memory, and achieves results close to the lower bound. Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Commun. | 3 |
| 2023 | Streaming Erasure Codes Over Multi-Access Relayed NetworksabstractMany emerging multimedia streaming applications involve multiple users communicating under strict latency constraints. In this paper we study streaming codes for a network involving two source nodes, one relay node and a destination node. In this paper’s setting, each source node transmits a stream of messages, through the relay, to a destination, who is required to decode the messages under a strict delay constraint. For the case of a single source node, a class of streaming codes has been proposed by Fong et al., using the concept of delay-spectrum. The current paper presents a novel framework, which constructs streaming codes for a relayed multi-user setting by sequentially constructing the codes for each link. This requires a characterization of the set of all achievable delay spectra for a given rate, blocklength and number of erasures, beyond the specific choice considered by Fong et al. This characterization is presented in the paper for systematic codes. Using this novel framework, the first proposed scheme involves greedily selecting the rate on the link from relay to destination and using properties of the delay-spectrum to find feasible streaming codes that satisfy the required delay constraints. A closed form expression for the achievable rate region is provided, and conditions for when the proposed scheme is optimal are established by a natural outer bound. The second proposed scheme builds upon this approach, but uses a numerical optimization-based approach to improve the achievable rate region over the first scheme. Experimental results show that the proposed schemes achieve significant improvements over baseline schemes based on single-user codes. Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2023 | Adaptive Relaying for Streaming Erasure Codes in a Three Node Relay NetworkabstractThis paper investigates adaptive streaming codes over a three-node relayed network. In this setting, a source node transmits a sequence of message packets to a destination with help of a relay. The source-to-relay and relay-to-destination links are unreliable and introduce at most$N_{1}$and$N_{2}$packet erasures, respectively. The destination node must recover each message packet within a strict delay constraint$T$. The paper presents a new construction of streaming codes for all feasible parameters$\{N_{1}, N_{2}, T\}$. Our work improves upon the construction in Fong et al. by adapting the relaying strategy based on the erasure patterns from source to relay. Specifically, the code employs the notion of symbol estimates, which allows the relay to forward information about symbols before it can decode that symbol, and variable-rate encoding, which decreases the rate used to encode a packet as more erasures affect that packet. The codes proposed in this paper achieve rates higher than the ones proposed by Fong et al. whenever$N_{2} > N_{1}$, and achieve the same rate when$N_{2} \leq N_{1}$, in which case the rate is optimal. The paper also presents an upper bound on the achievable rate that takes into account erasures in both links in order to bound the rate in the second link. The upper bound is shown to be tighter than a trivial bound that considers only the erasures in the second link. Gustavo Kasper Facenda, M. Nikhil Krishnan, Elad Domanovitz, Silas L. Fong, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 6 |
| 2022 | On State-Dependent Streaming Erasure Codes over the Three-Node Relay NetworkabstractThis paper investigates low-latency adaptive streaming codes for a three-node relay network. A source node transmits a sequence of source packets (messages) to the destination through a relay node. We focus on a particular case where the link connecting the source and relay nodes is almost reliable, but the link connecting the relay to the destination is not. The relay node can observe the erasure pattern that has occurred in the transmission between the source node and itself and adapt its relaying strategy based on that observation. Every source packet must be perfectly recovered by the destination with a strict delay T, as long as the number of erasures in the relay-to-destination link lies below some design parameter. We then characterize capacity as a function of such design parameter. The achievability scheme employs two different relaying strategies, based on whether an erasure has or has not occurred in the link from source to relay. The converse is proven by analyzing a periodic erasure pattern and lower bounding the minimum redundancy across channel packets. We show that the achievable rate can be improved compared to non-adaptive schemes previously proposed, indicating that exploiting the knowledge of the erasure pattern by the relay node is essential in achieving capacity. Gustavo Kasper Facenda, Elad Domanovitz, M. Nikhil Krishnan, Ashish Khisti, Silas L. Fong, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 6 |
| 2022 | State-Dependent Symbol-Wise Decode and Forward Codes Over Multihop Relay NetworksabstractThis paper studies low-latency streaming codes for the multi-hop network. The source transmits a sequence of messages to a destination through a chain of relays, and requires the destination to reconstruct each message by its deadline. We assume that each communication link is subjected to a certain maximum number of packet erasures. The case of a single relay (a three-node network) was considered in Fong et al. (2020). A coding scheme known as symbol-wise decode and forward was proposed. In the present work, we propose an alternative scheme that is different from Fong et al. (2020) and still achieves the same rate as in Fong et al. (2020) for the one hop case as the field-size goes to infinity. Furthermore, our proposed scheme naturally generalizes to the case of multiple-relay nodes yielding new achievable rates for this setting. The main difference with Fong et al. (2020) is that our proposed scheme exploits the ability of the relay nodes to adapt the transmission based on the erasures on the previous link. Hence, we refer to our scheme as “state-dependent” and contrast it with the scheme in Fong et al. (2020) that is state-independent. Our scheme requires the relay nodes to append a header to the transmitted packets, and we show that the size of the header does not depend on the field-size of the code. We also derive an upper bound on the maximal streaming rate achievable over a network with an arbitrary number of relays. We show that this upper bound matches our achievable rate in the special case when the maximal number of erasures on the first link is greater than or equal to the maximal number of erasures on each of the following links, and the field size goes to infinity. Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 3 |
| 2022 | Corrections to "Optimal Streaming Erasure Codes Over the Three-Node Relay Network"abstractIn the above article[1], an upper bound on the maximum achievable rate is corrected. If we add the restriction that the packets transmitted by the relay must be independent of the erasures introduced by the first-hop channel, then no correction is needed and the converse proof need not be rectified. Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2022 | Low-Latency Network-Adaptive Error Control for Interactive StreamingabstractWe introduce a novel network-adaptive algorithm that is suitable for alleviating network packet losses for low-latency interactive communications between a source and a destination. Our network-adaptive algorithm estimates in real-time the best parameters of a recently proposed streaming code that uses forward error correction (FEC) to correct both arbitrary and burst losses, which cause a crackling noise and undesirable jitters, respectively in audio. In particular, the destination estimates appropriate coding parameters based on its observed packet loss pattern and sends them back to the source for updating the underlying code. Besides, a new explicit construction of practical low-latency streaming codes that achieve the optimal tradeoff between the capability of correcting arbitrary losses and the capability of correcting burst losses is used. Simulation evaluations based on statistical losses and real-world packet loss traces reveal the following: (i) Our proposed network-adaptive algorithm combined with our optimal streaming codes can achieve significantly higher performance compared to uncoded and non-adaptive FEC schemes over UDP (User Datagram Protocol); (ii) Our explicit streaming codes can significantly outperform traditional MDS (maximum-distance separable) streaming schemes when they are used along with our network-adaptive algorithm. In addition, we study different factors that can affect the performance of our network-adaptive algorithm. Salma Emara, Silas L. Fong, Baochun Li, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Multim. | 5 |
| 2021 | Fast Manifold Landmarking Using Extreme Eigen-PairsabstractManifold landmarking is the problem of selecting a subset of discrete locations on a continuous manifold for label assignment, in order to reduce interpolation error of subsequent semi-supervised learning. In this paper, we select landmarks to minimize the condition number (λmax/λmin) of a submatrix of an alignment matrix Φ, which is equivalent to minimizing an interpolation error bound. Specifically, we design an efficient greedy scheme, where at each iteration t + 1 we choose one landmark i (thus deleting the corresponding row and column i of Φt) so that the resulting submatrix Φt+1has the smallest condition number. Towards fast landmark selection, at iteration t + 1, we first compute the two extreme eignevectors v1and vNcorresponding to λminand λmaxof Φtvia known methods like LOBPCG. We show that λmin(λmax) of submatrix Φt+1, from deleting the chosen row-column pair, can be approximated by an upper (lower) bound that is an easily computable function of eigen-pair {v1, λmin} ({vN, λmax}) of Φt. The error bounds of the obtained approximations can be numerically computed during the greedy step. Leveraging these proofs, we minimize a bound of the condition number for submatrix Φt+1at each greedy step t + 1. Experiments on synthetic and real-world manifold data demonstrate the superiority of our proposed landmarking algorithm compared to several state-of-the-art schemes. Gene Cheung, Yongchao Wang 0002, Wai-tian Tan |
ICASSP | 4 |
| 2021 | Streaming Erasure Codes over Multi-Access Relay NetworksabstractApplications where multiple users communicate with a common server and desire low latency are common and increasing. This paper studies a network with two source nodes, one relay node and a destination node, where each source nodes wishes to transmit a sequence of messages, through the relay, to the destination, who is required to decode the messages with a strict delay constraint$T$. The network with a single source node has been studied in [1]. We start by introducing two important tools: the delay spectrum, which generalizes delay-constrained point-to-point transmission, and concatenation, which, similar to time sharing, allows combinations of different codes in order to achieve a desired regime of operation. Using these tools, we are able to generalize the two schemes previously presented in [1], and propose a novel scheme which allows us to achieve optimal rates under a set of well-defined conditions. Such novel scheme is further improved in order to achieve higher rates in the scenarios where the conditions for optimality are not met. Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 4 |
| 2021 | Guaranteed Rate of Streaming Erasure Codes over Multi-Link Multi-hop NetworkabstractWe study the problem of transmitting a sequence of messages (streaming messages) through a multi-link, multi-hop packet erasure network. Each message must be reconstructed in-order and under a strict delay constraint. Special cases of our setting with a single link on each hop have been studied recently - the case of a single relay-node, is studied in Fong et al [1]; the case of multiple relays, is studied in Domanovitz et al [2]. As our main result, we propose an achievable rate expression that reduces to previously known results when specialized to their respective settings. Our proposed scheme is based on the idea of concatenating single-link codes from [2] in a judicious manner to achieve the required delay constraints. We propose a systematic approach based on convex optimization to maximize the achievable rate in our framework. Elad Domanovitz, Gustavo Kasper Facenda, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ITW | 4 |
| 2021 | High Rate Streaming Codes Over the Three-Node Relay NetworkabstractIn this paper, we investigate streaming codes over a three-node relay network. Source node transmits a sequence of message packets to the destination via a relay. Source-to-relay and relay-to-destination links are unreliable and introduce at most N1and N2packet erasures, respectively. Destination needs to recover each message packet with a strict decoding delay constraint of T time slots. We propose streaming codes under this setting for all feasible parameters $\{N_{1},\ N_{2},\ T\}$. Relay naturally observes erasure patterns occurring in the source-to-relay link. In our code construction, we employ a channel-state-dependent relaying strategy, which rely on these observations. In a recent work, Fong et al. provide streaming codes featuring channel-state-independent relaying strategies, for all feasible parameters $\{N_{1},\ N_{2},\ T\}$. Our schemes offer a strict rate improvement over the schemes proposed by Fong et al., whenever $N_{1}\lt N_{2}$. M. Nikhil Krishnan, Gustavo Kasper Facenda, Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ITW | 5 |
| 2021 | Data-Driven Mode and Group Selection for Downlink MU-MIMO With Implementation in Commodity 802.11ac NetworkabstractMulti-user MIMO (MU-MIMO) is a technique that improves spectral efficiency by allowing concurrent communication between one access point (AP) and multiple clients. In practice, the expected gain is not always achieved and is sometimes even negative. We experimentally demonstrate that the downlink MU-MIMO performance in a practical network not only depends on the client's channel but is also influenced by factors that are not captured by conventional models, such as client motion and device type. We propose a data-driven algorithm with a low computational complexity that determines whether a client should operate in MU mode and the MU-MIMO group for clients in MU mode. Such a mode and group selection algorithm is based on a sequence of channel state information (CSI), SNR, and client device type. The algorithm can automatically adapt to the motion and characteristics of individual clients. Experimental results using implementation on a commodity 802.11ac AP show that the proposed data-driven mode and group selection algorithm can improve network throughput by up to 35% over existing algorithms based on conventional models. We also show that the proposed data-driven algorithm has limited sensitivity to environmental changes and can be deployed into new environments without retraining. Shi Su, Wai-tian Tan, Rob Liston, Behnaam Aazhang |
IEEE Trans. Commun. | 2 |
| 2021 | Motion-Aware Optimizations for Downlink MU-MIMO in 802.11ax NetworksabstractMulti-User Multiple-Input and Multiple-Output (MU-MIMO) is a technique that allows concurrent transmissions between one access point (AP) and multiple clients to improve spectral efficiency. In practice, however, the MU-MIMO is sensitive to client mobility and is sometimes even harmful to the performance in networks with moving clients. In this paper, we identify that it is essential to optimize the MU-MIMO performance with moving clients by jointly selecting the sounding period, the number of spatial streams, and client grouping with the consideration of the client density of the network. We develop a data-driven model that estimates client throughput with the consideration of these parameters, as well as an algorithm that jointly determines the parameters for each client with low computational complexity. Using a commodity 802.11ax network, we experimentally demonstrate the significant impact of the key factors on MU-MIMO performance. Based on experimental data, we develop an emulation model to evaluate network performance with different client densities and mobility. Emulation results show that our proposed algorithm outperforms conventional schemes by over 20% in MU-MIMO networks with moving clients. Shi Su, Wai-tian Tan, Rob Liston, Herb Wildfeuer, Behnaam Aazhang |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2021 | DeepWiPHY: Deep Learning-Based Receiver Design and Dataset for IEEE 802.11ax SystemsabstractIn this work, we develop DeepWiPHY, a deep learning-based architecture to replace the channel estimation, common phase error (CPE) correction, sampling rate offset (SRO) correction, and equalization modules of IEEE 802.11ax based orthogonal frequency division multiplexing (OFDM) receivers. We first train DeepWiPHY with a synthetic dataset, which is generated using representative indoor channel models and includes typical radio frequency (RF) impairments that are the source of nonlinearity in wireless systems. To further train and evaluate DeepWiPHY with real-world data, we develop a passive sniffing-based data collection testbed composed of Universal Software Radio Peripherals (USRPs) and commercially available IEEE 802.11ax products. The comprehensive evaluation of DeepWiPHY with synthetic and real-world datasets (110 million synthetic OFDM symbols and 14 million real-world OFDM symbols) confirms that, even without fine-tuning the neural network's architecture parameters, DeepWiPHY achieves comparable performance to or outperforms the conventional WLAN receivers, in terms of both bit error rate (BER) and packet error rate (PER), under a wide range of channel models, signal-to-noise (SNR) levels, and modulation schemes. Yi Zhang 0021, Akash Doshi, Rob Liston, Wai-tian Tan, Jeffrey G. Andrews, Robert W. Heath Jr. |
IEEE Trans. Wirel. Commun. | 4 |
| 2020 | Streaming Erasure Codes over Multi-hop Relay NetworkabstractA typical path over the internet is composed of multiple hops. When considering the transmission of a sequence of messages (streaming messages) through packet erasure channel over a three-node network, it has been shown that taking into account the erasure pattern of each segment can result in improved performance compared to treating the channel as a point-to-point link. Since rarely there is only a single relay between the sender and the destination, it calls for trying to extend this scheme to more than a single relay. In this paper, we first extend the upper bound on the rate of transmission of a sequence of messages for any number of relays. We further suggest an achievable adaptive scheme that is shown to achieve the upper bound up to the size of an additional header that is required to allow each receiver to meet the delay constraints. Elad Domanovitz, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 3 |
| 2020 | Optimal Streaming Erasure Codes Over the Three-Node Relay Network
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2020 | Optimal Multiplexed Erasure Codes for Streaming Messages With Different Decoding Delays
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2019 | Learning Geographically Distributed Data for Multiple Tasks Using Generative Adversarial NetworksabstractWe present a novel method that supports the learning of multiple classification tasks from geographically distributed data. By combining locally trained generative adversarial networks (GANs) with a small fraction of original data samples, our proposed scheme can train multiple discriminative models at a central location with low communication overhead. Experiments using common image datasets (MNIST, CIFAR10, LSUN-20, Celeb-A) show that our proposed scheme can achieve comparable classification accuracy as the ideal classifier trained using all data from all sites. We further demonstrate that our method can scale to 10 sites without sacrificing classification accuracy for large datasets such as LSUN-20. Mehdi Nikkhah, Wai-tian Tan, Rob Liston |
ICIP | 4 |
| 2019 | Client Pre-Screening for MU-MIMO in Commodity 802.11ac Networks via Online LearningabstractMulti-user MIMO (MU-MIMO) is a technique in 802.11ac and 802.11ax that improves spectral efficiency by allowing concurrent communication between one AP and multiple clients. In practice, the expected gain is not always achieved and is sometimes even negative. Using a commodity 802.11ac AP, we experimentally determine that the inclusion of clients either in motion or with low SNR can cause throughput below that of single-user transmissions. We then propose a pre-screening algorithm using reinforcement learning to predict if a client can benefit from participating in MU-MIMO. Our algorithm is based on a sequence of channel state information (CSI), SNR, and client device type, and can automatically adapt to the motion of individual clients. Experimental results using a commodity AP show that the additional implementation of the pre-screening algorithm alone, without otherwise modifying MU-MIMO client grouping or link parameter selection algorithms, can improve system throughput by up to 40% when half of the clients are moving. Over 20% throughput improvement is maintained when between 25% to 75% of the clients are moving. Shi Su, Wai-tian Tan, Rob Liston |
INFOCOM | 2 |
| 2019 | Optimal Streaming Erasure Codes over the Three-Node Relay NetworkabstractThis paper investigates low-latency streaming codes for a three-node relay network. The source transmits a sequence of messages (streaming messages) to the destination through the relay between them, where the first-hop channel from the source to the relay and the second-hop channel from the relay to the destination are subject to packet erasures. Every source message generated at a time slot must be recovered perfectly at the destination within the subsequent T time slots. In any sliding window of $ {T}+1$ time slots, we assume no more than $ {N}_{1}$ and ${N}_{2}$ erasures are introduced by the first-hop channel and second-hop channel respectively. We fully characterize the maximum achievable rate in terms of T, $ {N}_{1}$ and $ {N}_{2}$ . The achievability is proved by using a symbol-wise decode-forward strategy where the source symbols within the same message are decoded by the relay with different delays. The converse is proved by analyzing the maximum achievable rate for each channel when the erasures in the other channel are consecutive (bursty). In addition, we show that traditional message-wise decode-forward strategies, which require the source symbols within the same message to be decoded by the relay with the same delay, are sub-optimal in general. Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 4 |
| 2019 | Optimal Multiplexed Erasure Codes for Streaming Messages with Different Decoding DelaysabstractThis paper considers multiplexing two sequences of messages with two different decoding delays over a packet erasure channel. In each time slot, the source constructs a packet based on the current and previous messages and transmits the packet, which may be erased when the packet travels from the source to the destination. The destination must perfectly recover every source message in the first sequence subject to a decoding delay Tv, and every source message in the second sequence subject to a shorter decoding delay Tu≤ Tv,. We assume that the channel loss model introduces a burst erasure of a fixed length B on the discrete timeline. Under this channel loss assumption, the capacity region for the case where Tv, ≤ Tu+B was previously solved. In this paper, we fully characterize the capacity region for the remaining case Tv, ) Tu+B. The key step in the achievability proof is achieving the non-trivial corner point of the capacity region through using a multiplexed streaming code constructed by superimposing two single-stream codes. The main idea in the converse proof is obtaining a genie-aided bound when the channel is subject to a periodic erasure pattern where each period consists of a length-B burst erasure followed by a length-Tunoiseless duration. Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 4 |
| 2019 | Low-Latency Network-Adaptive Error Control for Interactive StreamingabstractWe introduce a novel network-adaptive algorithm that is suitable for alleviating network packet losses for low-latency interactive communications between a source and a destination. Network packet losses happen in a bursty manner as well as an arbitrary manner, where the former is usually due to network congestion and the latter can be caused by unreliable wireless links. Our network-adaptive algorithm estimates in real time the best parameters of a recently proposed streaming code that corrects both arbitrary losses (which cause crackling noise in audio) and burst losses (which cause undesirable jitters and pauses in audio) using forward error correction (FEC). The network-adaptive algorithm updates the coding parameters in real time as follows: The destination estimates appropriate coding parameters based on its observed packet loss pattern and then the parameters are fed back to the source for updating the underlying code. In addition, a new explicit construction of practical low-latency streaming codes that achieve the optimal tradeoff between the capability of correcting arbitrary losses and the capability of correcting burst losses is provided. Simulation evaluations based on real-world packet loss traces reveal that our proposed network-adaptive algorithm combined with our optimal streaming codes achieves significantly higher reliability compared to uncoded and non-adaptive FEC schemes over UDP (User Datagram Protocol). Silas L. Fong, Salma Emara, Baochun Li, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ACM Multimedia | 5 |
| 2019 | Optimal Streaming Codes for Channels With Burst and Arbitrary Erasures
Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2018 | Learning Sensitive Images Using Generative ModelsabstractThe sheer amount of personal data being transmitted to cloud services and the ubiquity of cellphones cameras and various sensors, have provoked a privacy concern among many people. On the other hand, the recent phenomenal growth of deep learning that brings advancements in almost every aspect of human life is heavily dependent on the access to data, including sensitive images, medical records, etc. Therefore, there is a need for a mechanism that transforms sensitive data in such a way as to preserves the privacy of individuals, yet still be useful for deep learning algorithms. This paper proposes the use of Generative Adversarial Networks (GANs) as one such mechanism, and through experimental results, shows its efficacy. Sen-Ching S. Cheung, Herb Wildfeuer, Mehdi Nikkhah, Wai-tian Tan |
ICIP | 5 |
| 2018 | Optimal Streaming Codes for Channels with Burst and Arbitrary ErasuresabstractThis paper considers transmitting a sequence of messages (streaming messages) over a packet erasure channel. In each time slot, the source constructs a packet based on the current and the previous messages and transmits the packet, which may be erased when the packet travels from the source to the destination. Every source message must be recovered perfectly at the destination subject to a fixed decoding delay. We assume that the channel loss model introduces either one burst erasure or multiple arbitrary erasures in any fixed-sized sliding window. Under this channel loss assumption, we fully characterize the maximum achievable rate by constructing streaming codes that achieve the optimal rate. In addition, our construction of optimal streaming codes implies the full characterization of the maximum achievable rate for convolutional codes with any given column distance, column span, and decoding delay. Numerical results demonstrate that the optimal streaming codes outperform existing streaming codes of comparable complexity over some instances of the Gilbert-Elliott channel and the Fritchman channel. Silas L. Fong, Ashish Khisti, Baochun Li, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 4 |
| 2018 | Multiplexed Coding for Multiple Streams With Different Decoding DelaysabstractWe consider a communication setup where two source streams with different decoding deadlines, must be simultaneously transmitted over a single channel subjected to burst erasures. The encoder multiplexes the two source streams into a single stream of channel-packets. The decoder must recover the two source streams sequentially by their corresponding deadlines. One of the streams, the urgent stream, has a smaller delay than the other stream. We study the capacity region for such a setting for a certain range of system parameters. We divide the system into three different cases based on the relative values of the delays. For each case we provide achievability and converse bounds, which match under certain conditions. Our proposed coding scheme involves a careful construction of the parity check packets by jointly coding across the two streams despite different deadlines. Interestingly it is possible to transmit the urgent stream at a certain positive rate even when the sum rate equals the capacity associated with the less urgent stream. A separation based approach where we apply separate single-stream codes to each stream is suboptimal. Although our capacity results assume a simplistic channel model with a single erasure burst, we further demonstrate that our proposed code constructions also provide significant performance gains in simulations over statistical channel models with random bursts. Ahmed Badr, Devin Lui, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2018 | SKEPRID: Pose and Illumination Change-Resistant Skeleton-Based Person Re-IdentificationabstractCurrently, the surveillance camera-based person re-identification is still challenging because of diverse factors such as people’s changing poses and various illumination. The various poses make it hard to conduct feature matching across images, and the illumination changes make color-based features unreliable. In this article, we present SKEPRID, 1 a skeleton-based person re-identification method that handles strong pose and illumination changes jointly. To reduce the impacts of pose changes on re-identification, we estimate the joints’ positions of a person based on the deep learning technique and thus make it possible to extract features on specific body parts with high accuracy. Based on the skeleton information, we design a set of local color comparison-based cloth-type features, which are resistant to various lighting conditions. Moreover, to better evaluate SKEPRID, we build the PO8LI 2 dataset, which has large pose and illumination diversity. Our experimental results show that SKEPRID outperforms state-of-the-art approaches in the case of strong pose and illumination variation. Tuo Yu, Haiming Jin, Wai-tian Tan, Klara Nahrstedt |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2017 | Training sample selection for deep learning of distributed dataabstractThe success of deep learning - in the form of multi-layer neural networks - depends critically on the volume and variety of training data. Its potential is greatly compromised when training data originate in a geographically distributed manner and are subject to bandwidth constraints. This paper presents a data sampling approach to deep learning, by carefully discriminating locally available training samples based on their relative importance. Towards this end, we propose two metrics for prioritizing candidate training samples as functions of their test trial outcome: correctness and confidence. Bandwidth-constrained simulations show significant performance gain of our proposed training sample selection schemes over convention uniform sampling: up to 15× bandwidth reduction for the MNIST dataset and 25% reduction in learning time for the CIFAR-10 dataset. Wai-tian Tan, Rob Liston |
ICIP | 3 |
| 2017 | FEC for VoIP using dual-delay streaming codesabstractWe introduce a new class of forward error correction (FEC) codes for VoIP communications which support different recovery delay depending on the channel conditions. Specifically, our proposed class of Dual-Delay (DD) codes can recover from challenging long bursts of losses with close to theoretical minimum delay so as to meet playback deadlines for recovered packets. They further improve conversational interactivity by achieving lower recovery delay during periods of random isolated losses. These DD codes are shown to achieve lower residual loss rates when compared to existing codes over a wide range of parameters of the Gilbert-Elliott channel. Experiments over real world packet traces further show performance gains of DD codes in terms of perceptually motivated ITU-T G.107 E-model. Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
INFOCOM | 3 |
| 2017 | Multiplexed FEC for multiple streams with different playout deadlinesabstractWe study a setting where two source streams with different decoding deadlines must be transmitted over a burst erasure channel. The source streams are multiplexed into a single stream of channel-packets at the encoder and transmitted over a packet erasure channel. The decoder must recover the source-packets within each stream sequentially, by their corresponding deadlines. We consider the burst-erasure channel model and characterize the capacity region for a certain range of system parameters. We show that the operation of the system can be divided into three different regimes based on the relative values of decoding deadlines. On the achievability side, we show that jointly coding across the two streams, despite their different deadlines, is necessary to achieve capacity. On the converse side we develop information theoretic outer bounds on the capacity region. We find that the capacity region exhibits a “corner point” where we can transmit the urgent stream at a positive rate, yet attain a sum-rate equal to the capacity of the non-urgent stream. Ahmed Badr, Devin Lui, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 4 |
| 2017 | Layered Constructions for Low-Delay Streaming CodesabstractWe study error correction codes for multimedia streaming applications where a stream of source packets must be transmitted in real-time, with in-order decoding, and strict delay constraints. In our setup, the encoder observes a stream of source packets in a sequential fashion, and M channel packets must be transmitted between the arrival of successive source packets. Each channel packet can depend on all the source packets observed up to and including that time, but not on any future source packets. The decoder must reconstruct the source stream with a delay of T packets. We consider a class of packet erasure channels with burst and isolated erasures, where the erasure patterns are locally constrained. Our proposed model provides a tractable approximation to statistical models, such as the Gilbert-Elliott channel, for capacity analysis. When M = 1, i.e., when the source-packet arrival and channel-packet transmission rates are equal, we establish upper and lower bounds on the capacity, that are within one unit of the decoding delay T. We also establish necessary and sufficient conditions on the column distance and column span of a convolutional code to be feasible, and in turn establish a fundamental tradeoff between these. Our proposed codes-maximum distance and span codes- achieve a near-optimal tradeoff between the column distance and column span, and involve a layered construction. When M > 1, we establish the capacity for the burst-erasure channel and an achievable rate in the general case. Extensive numerical simulations over Gilbert-Elliott and Fritchman channel models suggest that our codes also achieve significant gains in the residual loss probability over statistical channel models. Ahmed Badr, Pratik Patil, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
IEEE Trans. Inf. Theory | 4 |
| 2016 | Content-independent and loss-pattern-aware distortion evaluation for streaming mediaabstractIt is well known that dispersed and burst packet losses introduce significantly different amount of distortions. Since perceptual models are typically content dependent, it is challenging to characterize how losses interact with concealment. This paper presents loss-pattern-aware distortion (LoPAD), a content-independent metric that explicitly models the impact of different loss patterns. LoPAD operates solely on the loss trace without analyzing received media. It is fast, and supports offline and cloud-based monitoring of network impairment. Taking audio conferencing as target application, we show that full-reference PESQ scores for a collection of speech samples can be closely approximated by LoPAD. For various combinations of erasure channel models and forward error correction (FEC) codes, the correlation coefficients between LoPAD and PESQ-DMOS range from 0.90 to 0.97. Wai-tian Tan, John G. Apostolopoulos, Ahmed Badr, Ashish Khisti |
ICIP | 2 |
| 2015 | Software defined networking for video: Overview & multicast studyabstractSoftware Defined Networking (SDN) is an architectural trend in networking towards the use of centralized controllers to improve global visibility and simplify various network operation tasks. Due to their need for QoS, sizable traffic involved, and dynamic nature, video applications are particularly suitable candidates for dynamic interaction with SDN. This paper provides an overview of different general ways that video applications have interacted with networks, and outline what opportunities SDN provides. We then develop and evaluate several methods that exploit SDN to construct multicast video trees and show that over realistic topology, SDN optimized trees can support 60-90% more traffic. Wai-tian Tan, Herb Wildfeuer, John G. Apostolopoulos |
ICIP | 1 |
| 2014 | Loss-Resilient Coding of Texture and Depth for Free-Viewpoint Video ConferencingabstractFree-viewpoint video conferencing allows a participant to observe the remote 3D scene from any freely chosen viewpoint. An intermediate virtual viewpoint image is typically synthesized using two pairs of transmitted texture and depth maps from two neighboring captured viewpoints via depth-image-based rendering (DIBR). To maintain high quality of synthesized images, it is imperative to contain the adverse effects of network packet losses that may arise during texture and depth video transmission. Towards this goal, we develop an integrated approach that exploits the representation redundancy inherent in the multiple streamed videos-a voxel in the 3D scene visible to two captured views is sampled and coded twice in the two views. In particular, at the receiver we first develop an error concealment strategy that adaptively blends corresponding pixels in the two captured views during DIBR, so that pixels from the more reliable transmitted view are weighted more heavily. We then couple it with a sender-side optimization of reference picture selection (RPS) during real-time video coding, so that blocks containing pixel samples of voxels that are visible in both views are more error-resiliently coded in one view only, given adaptive blending will mitigate errors in the other view. Further, synthesized view distortion sensitivities to texture versus depth errors are analyzed, so that relative importance of texture and depth code blocks can be computed for system-wide RPS optimization. Finally, quantization parameter (QP) is adaptively selected per frame, optimally trading off source distortion due to compression with channel distortion due to potential packet losses. Experimental results show that the proposed scheme can outperform previous work by up to 2.9 dB at 5% packet loss rate. Bruno Macchiavello, Camilo C. Dorea, Edson M. Hung, Gene Cheung, Wai-tian Tan |
IEEE Trans. Multim. | 5 |
| 2013 | Determining co-location using a sequential hypothesis test on patterns of silenceabstractIn everyday meetings, automatic association of co-located mobile devices would ease sharing of web-links, media, and other information. We propose a method that compares patterns of silence from device microphones to detect co-location of those devices. This method works with unsynchronized audio capture, requires only 100bps and preserves privacy. We show how to formulate pattern matching in a sequential hypothesis framework so that changes in co-location status (when people leave or join a meeting) can be determined promptly, and how to compute the likelihood ratio in practice. Using 16 hours of captured audio, we show that our approach can correctly determine device co-location with a low error rate of 0.05%, and can detect co-location changes 10 seconds faster than a similar decision rule based on a constant time window. Compared to a prior audio signature method, we achieve higher accuracy at 1/7 the bit rate. Wai-tian Tan, Ramin Samadani, Bowon Lee, Mary Baker |
ICASSP | 1 |
| 2013 | Saliency-cognizant robust view synthesis in free viewpoint video streamingabstractIn free viewpoint video, texture and depth maps from two camera-captured viewpoints are transmitted, so that at receiver, a novel virtual view chosen by the client can be synthesized via depth-image-based rendering (DIBR). When irrecoverable packet losses occur during transmission-typically affecting less important spatial regions in the video given unequal error protection (UEP) is deployed-appropriate error concealment strategies must be used at decoder to minimize resulting visual degradation in the synthesized view. Towards this goal, we propose a new optimization framework based on visual saliency to combine two different concealment techniques. First, given a pixel in the virtual view is typically constructed as a convex combination of corresponding pixels in the left and right captured views, weighted pixel blending (WPB) readjusts the weights in the linear sum to reflect the expected error in code blocks that contain the corresponding pixels. Second, exemplar-based patch matching (EPM) finds the most similar patches in the known spatial region to complete missing pixels in the unknown region. To choose between candidates constructed using the two techniques when filling a given pixel patch in the synthesized view, we first compute a weighted sum of expected error and visual saliency for each candidate patch. The candidate with the smaller sum (one with small expected error and visual saliency, so that even if errors do occur, they do not stand out visually) is selected for pixel completion. Experimental results show that our scheme can outperform the use of co-located blocks from a previous frame by up to 0.7dB in PSNR and improve subjective visual quality. Bruno Macchiavello, Camilo C. Dorea, Edson M. Hung, Gene Cheung, Wai-tian Tan |
ICIP | 5 |
| 2013 | Image-based indoor place-finder using image to plane matchingabstractDetermining indoor location using image based methods has the promise to be fast and intuitive. Nonetheless, a practical implementation needs to operate without knowledge of true camera pose and focal length, and be discriminating enough to identify location from multiple indoor locations that appear similar.In this paper, we propose an accurate and efficient method based on matching established planes in an environment to a query images. This greatly reduces the necessary computation, and improves accuracy by enforcing geometry on local feature descriptors. Accuracy is further improved by computing matching score based on number of matching pixels rather than descriptors as commonly done. Using a database of over 2000 planes and over 120 query images, we show our algorithm maintains accuracy over 86% even for challenging environments with multiple similar locations, and outperforms a feature based method by 10 - 40%. Ju Shen, Wai-tian Tan |
ICME | 2 |
| 2013 | Streaming codes for channels with burst and isolated erasuresabstractWe study low-delay error correction codes for streaming recovery over a class of packet-erasure channels that introduce both burst-erasures and isolated erasures. We propose a simple, yet effective class of codes whose parameters can be tuned to obtain a tradeoff between the capability to correct burst and isolated erasures. Our construction generalizes previously proposed low-delay codes which are effective only against burst erasures. We establish an information theoretic upper bound on the capability of any code to simultaneously correct burst and isolated erasures and show that our proposed constructions meet the upper bound in some special cases. We discuss the operational significance of column-distance and column-span metrics and establish that the rate 1/2 codes discovered by Martinian and Sundberg [IT Trans. 2004] through a computer search indeed attain the optimal column-distance and column-span tradeoff. Numerical simulations over a Gilbert-Elliott channel model and a Fritchman model show significant performance gains over previously proposed low-delay codes and random linear codes for certain range of channel parameters. Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
INFOCOM | 3 |
| 2013 | Robust streaming erasure codes based on deterministic channel approximationsabstractWe study near optimal error correction codes for real-time communication. In our setup the encoder must operate on an incoming source stream in a sequential manner, and the decoder must reconstruct each source packet within a fixed playback deadline of T packets. The underlying channel is a packet erasure channel that can introduce both burst and isolated losses. We first consider a class of channels that in any window of length T +1 introduce either a single erasure burst of a given maximum length B, or a certain maximum number N of isolated erasures. We demonstrate that for a fixed rate and delay, there exists a tradeoff between the achievable values of B and N, and propose a family of codes that is near optimal with respect to this tradeoff. We also consider another class of channels that introduce both a burst and an isolated loss in each window of interest and develop the associated streaming codes. All our constructions are based on a layered design and provide significant improvements over baseline codes in simulations over the Gilbert-Elliott channel. Ahmed Badr, Ashish Khisti, Wai-tian Tan, John G. Apostolopoulos |
ISIT | 3 |
| 2013 | Sensing device co-location through patterns of silenceabstractThis document describes the technology behind the accompanying video, which gives a demonstration of determining the dynamic group membership of a meeting by matching patterns of relative audio silence, or "silence signatures," sensed by mobile devices. Wai-tian Tan, Mary Baker, Bowon Lee, Ramin Samadani |
MobiSys | 1 |
| 2013 | The sound of silenceabstractA list of the dynamically changing group membership of a meeting supports a variety of meeting-related activities. Effortless content sharing might be the most important application, but we can also use it to provide business card information for attendees, feed information into calendar applications to simplify scheduling of follow-up meetings, populate the membership of collaborative editing applications, mailing lists, and social networks, and perform many other tasks. Wai-tian Tan, Mary Baker, Bowon Lee, Ramin Samadani |
SenSys | 1 |
| 2013 | Low-Cost Eye Gaze Prediction System for Interactive Networked Video StreamingabstractEye gaze is now used as a content adaptation trigger in interactive media applications, such as customized advertisement in video, and bit allocation in streaming video based on region-of-interest (ROI). The reaction time of a gaze-based networked system, however, is lower-bounded by the network round trip time (RTT). Furthermore, only low-sampling-rate gaze data is available when commonly available webcam is employed for gaze tracking. To realize responsive adaptation of media content even under non-negligible RTT and using common low-cost webcams, we propose a Hidden Markov Model (HMM) based gaze-prediction system that utilizes the visual saliency of the content being viewed. Specifically, our HMM has two states corresponding to two of human's intrinsic gaze behavioral movements, and its model parameters are derived offline via analysis of each video's visual saliency maps. Due to the strong prior of likely gaze locations offered by saliency information, accurate runtime gaze prediction is possible even under large RTT and using common webcam. We demonstrate the applicability of our low-cost gaze prediction system by focusing on ROI-based bit allocation for networked video streaming. To reduce transmission rate of a video stream without degrading viewer's perceived visual quality, we allocate more bits to encode the viewer's current spatial ROI, while devoting fewer bits in other spatial regions. The challenge lies in overcoming the delay between the time a viewer's ROI is detected by gaze tracking, to the time the effected video is encoded, delivered and displayed at the viewer's terminal. To this end, we use our proposed low-cost gaze prediction system to predict future eye gaze locations, so that optimized bit allocation can be performed for future frames. Through extensive subjective testing, we show that bit-rate can be reduced by up to 29% without noticeable visual quality degradation when RTT is as high as 200 ms. Yunlong Feng, Gene Cheung, Wai-tian Tan, Patrick Le Callet, Yusheng Ji |
IEEE Trans. Multim. | 3 |
| 2012 | Selective freezing of impaired video frames using low-rate shift-invariant hintabstractIrrecoverable data loss may be unavoidable for real-time video communication over common best-effort networks. Rather than always display impaired pictures or always “freeze” the last good picture, it is preferable to transmit additional hints to support selective freezing of heavily damaged pictures only. In particular, errors in impaired pictures tend to be localized, and often manifest themselves as small spatial shifts that are visually preferable to freezes. Therefore, it is essential that the hints can identify localized error and do not penalize small shifts. We show two ways such “shift-invariant” hints for detecting concealment error can be constructed to less than 1% of video rate. Experiments using 720p sequences achieve recall and precision of 90% and 75%, respectively, with respect to a shift-invariant PSNR measure. We also present an adaptive decision rule to obtain shorter and less frequent freezes. Mina Makar, Wai-tian Tan |
ICASSP | 2 |
| 2012 | Reference frame selection for loss-resilient texture & depth map coding in multiview video conferencingabstractIn a free-viewpoint video conferencing system, the viewer can choose any desired viewpoint of the 3D scene for observation. Rendering of images for arbitrarily chosen viewpoint can be achieved through depth-image-based rendering (DIBR), which typically employs “texture-plus-depth” video format for 3D data exchange. Robust and timely transmission of multiple texture and depth maps over bandwidth-constrained and loss-prone networks is a challenging problem. In this paper, we optimize transmission of multiview video in texture-plus-depth format over a lossy channel for free viewpoint synthesis at decoder. In particular, we construct a recursive model to estimate the distortion in synthesized view due to errors in both texture and depth maps, and formulate a rate-distortion optimization problem to select reference pictures for macroblock encoding in H.264 in a computation-efficient way, in order to provide unequal protection to different macroblocks. Results show that the proposed scheme can outperform random insertion of intra refresh blocks by up to 0.73 dB at 5% loss. Bruno Macchiavello, Camilo C. Dorea, Edson M. Hung, Gene Cheung, Wai-tian Tan |
ICIP | 5 |
| 2012 | Gaze-Driven video streaming with saliency-based dual-stream switchingabstractThe ability of a person to perceive image details falls precipitously with larger angle away from his visual focus. At any given bitrate, perceived visual quality can be improved by employing region-of-interest (ROI) coding, where higher encoding quality is judiciously applied only to regions close to a viewer's focal point. Straight-forward matching of viewer's focal point with ROI coding using a live encoder, however, is computation-intensive. In this paper, we propose a system that supports ROI coding without the need of a live encoder. The system is based on dynamic switching between two pre-encoded streams of the same content: one at high quality (HQ), and the other at mixed quality (MQ), where quality of a spatial region depends on its pre-computed visual saliency values. Distributed source coding (DSC) frames are periodically inserted to facilitate switching. Using a Hidden Markov Model (HMM) to model a viewer's temporal gaze movement, MQ stream is pre-encoded based on ROI coding to minimize the expected streaming rate, while keeping the probability of a viewer observing low quality (LQ) spatial regions below an application-specific ϵ. At stream time, the viewer's gaze locations are collected and transmitted to server for intelligent stream switching. In particular, server employs MQ stream only if: i) viewer's tracked gaze location falls inside the high-saliency regions, and ii) the probability that a viewer's gaze point will soon move outside high-saliency regions, computed using tracked gaze data and updated saliency values, is below ϵ. Experiments showed that video streaming rate can be reduced by up to 44%, and subjective quality is noticeably better than a competing scheme at the same rate where the entire video is encoded using equal quantization. Yunlong Feng, Gene Cheung, Wai-tian Tan, Yusheng Ji |
VCIP | 3 |
| 2011 | Face recovery in conference video streaming using robust principal component analysisabstractIrrecoverable data loss is inevitable for low-delay video conferencing over typical loss-prone networks such as the Internet. A semi-super-resolution (SSR) framework has been previously proposed to supply an additional low-resolution (LR) thumbnail to aid error concealment when the high-resolution (HR) image is lost. Super-resolution is an ill-posed problem, however, and previous block-search based SSR methods tend to produce discontinuities in output images, which can be objectionable, especially in human faces where the focus of a viewer usually lies. In this paper, we propose to recover a human face in a lost frame using the same SSR framework, but by operating on the entire face at a time. We leverage on a recent work called robust principal component analysis (RPCA), where the “salient” features (human face in our scenario) in a sequence of previous HR frames can be recovered despite the presence of gross but sparse errors. We propose and derive various improved methods to solve the SSR problem using RPCA. Beyond robust recovery of the human face, transformations of the face in previous HR frames are also deduced, so that the recovered face can be appropriately transformed in the lost frame for natural viewing. Experimental results show that our face-based approach gives much improved face recovery compared to previous SSR block searches. Wai-tian Tan, Gene Cheung |
ICIP | 1 |
| 2011 | Hidden Markov Model for eye gaze prediction in networked video streamingabstractWith the advent of eye gaze tracking technology, eye gaze is increasingly being used as a media interaction trigger in a variety of applications, such as eye typing, video content customization, and network video streaming based on region-of-interest (ROI). The reaction time of a gaze-based networked system, however, is in practice lower-bounded by the round trip time (RTT) of today's networks, which can be large. To improve the efficacy of gaze-based networked systems, in the paper we propose a Hidden Markov Model (HMM)-based gaze prediction strategy to predict future gaze locations to lower end-to-end reaction delay. We first design an HMM with three states corresponding to human's three major types of intrinsic eye movements. HMM parameters are obtained offline on a per-video basis during training phase. During testing phase, a window of noisy gaze observations are collected in real-time as input to a forward algorithm, which computes the most likely HMM state. Given the deduced HMM state, linear prediction is used to predict gaze location RTT seconds into the future. We demonstrate the applicability of our gaze prediction strategy by focusing on ROI-based bit allocation for network video streaming. To reduce transmission rate of a video stream without degrading viewer's perceived visual quality, we allocate more bits to encode the viewer's current spatial ROI, while devoting fewer bits in other spatial regions. The challenge lies in overcoming the delay between the time a viewer's ROI is detected by gaze tracking, to the time the effected video is encoded, delivered and displayed at the viewer's terminal. To this end, we use our proposed gaze-prediction strategy to predict future eye gaze locations, so that optimized bit allocation can be performed for future frames. Our experiments show that bit rate can be reduced by 21% without noticeable visual quality degradation when end-to-end network delay is as high as 200ms. Yunlong Feng, Gene Cheung, Wai-tian Tan, Yusheng Ji |
ICME | 3 |
| 2011 | Error-resilient live video multicast using low-rate visual quality feedbackabstractEffective adaptive streaming systems need informative feedback that supports selection of appropriate actions. Packet level timing and reception statistics are already widely reported in feedback. In this paper, we introduce a method to produce low bit-rate visual quality feedback and evaluate its effectiveness in controlling errors in live video multicast. The visual quality feedback is a digest of picture content, and allows localized comparison in time and space on a continuous basis. This conveniently allows detection and localization of significant errors that may have originated from earlier irrecoverable losses, a task that is typically challenging with packet level feedback only. Our visual quality feedback has low bit overhead, at about 1% for high-definition video encoded at typical rates. For live video multicast with 10 clients, our experimental results show that the added ability to detect and correct large drift errors significantly reduces the resulting visual quality fluctuations. David P. Varodayan, Wai-tian Tan |
MMSys | 2 |
| 2010 | Redundant representation for network video streaming using reconstructed P-frames and SP-framesabstractFor low-delay streaming of pre-encoded video over lossy networks, fast recovery from decoding errors typically involves use of frequent intra-coded frames, which incurs high bandwidth cost. In this paper, we present a redundant representation of video using a bandwidth-efficient but non-resilient main bitstream, and an additional auxiliary bitstream dedicated to recovery from losses in the main bitstream. In particular, we insert primary SP-frames periodically into the main bitstream, and encode corresponding secondary SP-frames and reconstructed P-frames in the auxiliary stream. When a frame loss occurs, a reconstructed P-frame is first sent to re-synchronize decoder back to normal motion compensation loop, then a secondary SP-frame corresponding to the location of the next pre-inserted primary SP-frame in the main stream is sent thereafter, eliminating coding drift. Results show that proposed method out-performs non-redundant representation of I-frame insertion by up to 11 frames in recovery time, and out-performed redundant representation of only reconstructed P-frames by up to 2.2dB in average PSNR. Gene Cheung, Wai-tian Tan |
ICASSP | 2 |
| 2010 | Multi-resolution redundancy for error-resilient video transmissionabstractIn this paper we advocate a multi-resolution mechanism for redundant information generation and transmission with low overheads in bit-rate, in order to enable reliable video communication over challenging lossy networks with low latency. Our previous work entitled RECAP, transmitted a low resolution redundant version of a parent video stream, in order to achieve effective concealment of isolated and burst losses by clever multi-frame super-resolution processing at the decoder. A reference picture selection mechanism was used to transmit the low-resolution redundant video with high guarantee and stop drift in case of losses. In this work, we extend the framework to incorporate an additional distributed Wyner-Ziv coding layer on the low-resolution information to further correct errors, and consequently improve the error-concealed picture quality. The error-concealed super-resolved frame using the low-resolution information alone, now acts as side-information at the decoder to correct additional errors by decoding the Wyner-Ziv layer. Preliminary results are presented to demonstrate the efficacy of the proposed approach. Debargha Mukherjee, Wai-tian Tan |
ICASSP | 2 |
| 2010 | Improving error resilience of scalable H.264 (SVC) via drift controlabstractCommon error concealment schemes mitigate errors for frames in which losses occur only, even though errors propagate to future frames. Drift control is generally challenging due to lack of reliable basis for determining what needs to be corrected and how. In this paper, we show that for scalable or multi-layer video, an available base-layer can serve as such basis to allow continous error drift checking and correction of higher layers even when the base-layer is of much lower spatial resolution. The associated algorithm is low-complexity, incurs no additional bit cost, and experiments using SVC reference software show PSNR improvement of up to 5 dB over concealment methods without drift control. Wai-tian Tan, Andrew J. Patti |
ICASSP | 1 |
| 2009 | Receiver error concealment using acknowledge preview (RECAP) - An approach to resilient video streamingabstractHigh-quality and low-latency video streaming is essential to providing a natural user experience in video conferencing. This is challenging over lossy networks since compressed video is highly fragile while the low-latency requirement limits the effectiveness of traditional error control approaches such as retransmission and forward error correction. In this paper, we advocate a practical solution for low-latency video communications over best-effort networks that employs an additional low-quality, low-resolution but robustly coded copy of the video. This approach, called RECAP, incurs minimal rate overhead, and can be combined with previously decoded frames to achieve effective concealment of isolated and burst losses even under tight delay constraints. RECAP achieves PSNR gains of 2-6 dB against complete frame loss. Chuohao Yeo, Wai-tian Tan, Debargha Mukherjee |
ICASSP | 2 |
| 2009 | Stationary video camera auto-exposure conditioningabstractVideo conferencing without controlled lighting suffers from the spurious automatic exposure (AE) errors commonly seen in Webcams. These errors cause problems for the subsequent processing and compression. For example, since video encoders do not model intensity changes, these AE errors in turn cause severe blocking artifacts. We develop a pixel-domain AE conditioning algorithm for stationary cameras that: (1) effectively reduces spurious AE changes, resulting in natural and artifact-free video; (2) allows maximum compatibility with third party components (may be transparently inserted between any camera driver and encoder/video processing engine); and (3) is fast and requires little memory. This algorithm allows inexpensive cameras to provide higher quality video conferencing. We describe the algorithm, analyze its performance exactly for a specific video source model and validate its performance experimentally using captured video. Ramin Samadani, Wai-tian Tan |
ICIP | 2 |
| 2009 | Community Streaming With Interactive Visual Overlays: System and OptimizationabstractCommunity streaming is an enhanced form of joint content viewing where a sense of community is reinforced by the addition of interactive visual overlays, controlled in real-time by viewers, on top of a shared video stream. As a concrete example, we describe a community video system called ECHO, where personalized avatars are overlaid on top of a real-time encoded video stream of an Internet game for multicast consumption. Recognizing that only the visual overlays are generated live, we propose schemes that encode and schedule the live and non-live portions of the overlaid video separately in order to exploit the difference in delay sensitivity of the two, leading to video streams that contain two sub-streams with different delay constraints. We show that, in the known channel case, a low complexity ldquoearliest deadline firstrdquo packet scheduling algorithm minimizes receiver buffer delay. We also analyze the case where multiple streams are multiplexed, which allows us to quantify the potential gains of allowing different delay constraints for different sub-streams. We show that a ldquowater fillingrdquo strategy maximizes the total number of streams that can be supported. Simulation results show that the bandwidth necessary to maintain low-latency for visual overlays is reduced by about 40% when our proposed sub-stream approach is used. For multiplexing of multiple streams, our approach can increase the number of supported streams (e.g., a 30% increase when around ten streams are multiplexed). Wai-tian Tan, Gene Cheung, Antonio Ortega, Bo Shen 0003 |
IEEE Trans. Multim. | 1 |
| 2008 | Temporal propagation analysis for small errors in a single-frame in H.264 videoabstractThis paper studies the temporal error propagation of small errors in a single frame for H.264 video. Such small errors can arise due to imperfect recovery from loss, e.g., through error concealment. The key contribution of this paper include demonstrating empirically that small errors tend to amplify over time, and is primarily caused by rounding errors in motion compensation and selective application of deblocking filter based on thresholding. Some methods of reducing error amplification is also presented. Wai-tian Tan, Bo Shen 0003, Andrew J. Patti, Gene Cheung |
ICIP | 1 |
| 2007 | Low-Latency Error Control of H.264 Using SP-Frames and Streaming Agent Over Wireless NetworksabstractA key challenge to low-latency wireless video streaming, where persistent retransmission is impractical, is error control. While SP-frame adaptation of H.264 has potential to mitigate error propagation, streaming server is often either situated too far to react in a timely fashion to client feedbacks, or too computationally constrained to perform the necessarily complex adaptation simultaneously for multiple clients in different sessions. In this paper, we present an innovative error control mechanism using SP-frames of H.264 and performed by a network intermediary for video streaming to a wireless client. Using an intermediary means it is more responsive to client feedbacks due to its close proximity, and it can offload computation complexity from the streaming server. Simulation shows that about 2 dB improvement in PSNR is achievable for video streaming with low latency requirements over traditional schemes using I and P-frames only. Gene Cheung, Wai-tian Tan |
ICC | 2 |
| 2007 | Lossless FMO and Slice Structure Modification for Compressed H.264 VideoabstractWe introduce a scheme to losslessly modify pre-compressed H.264 video to enable at streaming time (1) modification of slice sizes to fit transport packet size, and (2) introduction of an error resilience feature, namely flexible macroblock ordering. By lossless we mean the reconstructed video from the modified bitstream is identical to that from the original compressed bitstream. We outline how this transcoder operates, and discuss some of its restrictions. The bit rate overhead for operating on the pre-compressed, rather than directly encoding the original video with the desired characteristics, is analyzed. Simulation results show 1 to 2 dB improvement when our scheme is applied to QuickTime generated H.264 videos transported over a lossy packet network. Wai-tian Tan, Eric Setton, John G. Apostolopoulos |
ICIP (4) | 1 |
| 2007 | Reference Frame Optimization for Multiple-Path Video Streaming With Complexity ScalingabstractRecent video coding standards such as H.264 offer the flexibility to select reference frames during motion estimation for predicted frames. In this paper, we study the optimization problem of jointly selecting the best set of reference frames and their associated transport QoS levels in a multipath streaming setting. The application of traditional Lagrangian techniques to this optimization problem suffers from either bounded worst case error but high complexity or low complexity but undetermined worst case error. Instead, we present two optimization algorithms that solve the problem globally optimally with high complexity and locally optimally with lower complexity. We then present rounding methods to further reduce computation complexity of the second dynamic programming-based algorithm at the expense of degrading solution quality. Results show that our low-complexity dynamic programming algorithm achieves results comparable to the optimal but high-complexity algorithm, and that gradual tradeoff between complexity and optimization quality can be achieved by our rounding techniques Gene Cheung, Wai-tian Tan, Connie Chan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Using SP-Frames for Error Resilience in Optimized Video StreamingabstractSP-frame is a new picture type of H.264 that can identically reconstruct a picture using any one of several reference frames. In this paper, we discuss how this property can be exploited for controlling error propagation caused by packet losses. We first illustrate the benefits of the scheme through example. We then present results for optimized streaming where PSNR performance of a proposed usage of SP-frames is compared to that of P-frames only. Results show that using SP-frames can noticeably reduce distortion caused by burst packet losses compare to schemes based on P-frames only. Wai-tian Tan, Gene Cheung |
ICIP | 1 |
| 2006 | Methods to Improve Coding Efficiency of SP FramesabstractSP-frame is a new picture type supported by H.264, and supports functions such as rate-switching and random-access. In this paper, we investigate several complementary methods to improve the coding efficiency of SP-frames. We show that by appropriately choosing reference pictures, the size of secondary SP frames can be reduced by up to 40% and 2% for random-access and rate-switching, respectively. We also demonstrate that a simple rule exists that allows the joint selection of the two quantization parameters associated with SP frames to minimize "requantization" error. Results shows 0.1 dB PSNR improvement over comparable choices. Wai-tian Tan, Bo Shen 0003 |
ICIP | 1 |
| 2006 | A Case for Internet Streaming via Web ServersabstractHosting Internet streaming services has its unique challenges. Aiming at making Internet streaming services be widely and easily adopted in practice, in this paper, we have designed and implemented a system, called SProxy that can leverage existing Internet infrastructure to free the streaming content providers so that they only need to host streaming content through a regular Web server. SProxy has been extensively tested and evaluated and it provides high quality streaming delivery in both local area networks and wide area networks (e.g. between Japan and US) Songqing Chen, Bo Shen 0003, Wai-tian Tan, Susie J. Wee, Xiaodong Zhang 0001 |
ICME | 3 |
| 2006 | Packet Scheduling of Streaming Video with Flexible Reference Frame using Dynamic Programming and Integer RoundingabstractVideo coding standards like H.264 offer the flexibility to select reference frames during motion estimation for predicted frames. We investigate the packet scheduling problem of streaming video over lossy networks from a real-time encoder with flexible reference frame. In particular, we consider a multi-path streaming setting where each predicted frame of video, in addition to the flexibility to select a reference frame, can schedule one or multiple transmissions on one or multiple delivery paths for the upcoming optimization period. We present an algorithm based on dynamic programming that provides a locally optimal solution with high complexity. We then present a rounding method to reduce computation complexity at the expense of degrading solution quality. Results show that our algorithm performs noticeably better than a greedy scheme, and graceful tradeoff between complexity and solution quality can be achieved Gene Cheung, Wai-tian Tan |
ICME | 2 |
| 2005 | Design and Implementation of an Extrusion-based Break-In Detector for Personal ComputersabstractAn increasing variety of malware, such as worms, spyware and adware, threatens both personal and business computing. Remotely controlled bot networks of compromised systems are growing quickly. In this paper, we tackle the problem of automated detection of break-ins caused by unknown malware targeting personal computers. We develop a host based system, BINDER (Break-IN DEtectoR), to detect break-ins by capturing user unintended malicious outbound connections (referred to as extrusions). To infer user intent, BINDER correlates outbound connections with user-driven input at the process level under the assumption that user intent is implied by user-driven input. Thus BINDER can detect a large class of unknown malware such as worms, spyware and adware without requiring signatures. We have successfully used BINDER to detect real world spyware on daily used computers and email worms on a controlled testbed with very small false positives. Weidong Cui, Randy H. Katz, Wai-tian Tan |
ACSAC | 3 |
| 2005 | Accurate distortion-driven macroblock level rate control via ρ-domain analysisabstractRate control via /spl rho/-domain rate-distortion analysis (both rate and distortion are considered as functions of /spl rho/, the percentage of zero transform coefficients) relies on a distortion model which can be erroneous in many situations, especially at the macroblock level (Z. He and S. K. Mita, 2002). This paper uses an accurate distortion measurement in the transform domain to drive a better macroblock level bit allocation. Specifically, we propose an algorithm based on an accurate distortion-to-/spl rho/ mapping in order to minimized the distortion while matching the target rate. Experimental results show that our proposed algorithm achieved 0.1 to 0.3 dB better PSNR compared to the pure /spl rho/-domain rate control. Wai-tian Tan, Bo Shen 0003 |
ICIP (3) | 1 |
| 2005 | Enterprise Streaming: Different Challenges from Internet StreamingabstractMedia streaming over the best-effort public Internet has been a focus of research for over a decade. Enterprise or corporate streaming is another area of media streaming that is practically very important and has a different set of challenges and feasible solutions. For example, the quality and reliability requirements for enterprise streaming are much stricter than for typical Internet streaming. Furthermore, in typical enterprise streaming scenarios, a single entity has control over most elements of the system, including the end-points and the infrastructure. This entity has the powerful ability to monitor, adapt, and deploy new infrastructure as necessary. The goal of this paper is to describe enterprise streaming and identify the basic differences between typical Internet streaming and enterprise streaming, and how these differences alter the challenges that must be overcome for enterprise streaming to be successful. Specifically, we examine enterprise streaming media content delivery network design and operation, video conferencing, peer-to-peer networking (P2P), voice over IP (VoIP), and briefly touch upon wireless and security issues John G. Apostolopoulos, Mitchell D. Trott, Ton Kalker, Wai-tian Tan |
ICME | 4 |
| 2005 | Loss-Compensated Reference Frame Optimization for Multi-Path Video StreamingabstractRecent video coding standards such as H. 264 offer the flexibility to select reference frames during motion estimation for predicted frames. In this paper, by tracking loss compensation during distortion minimization, we improve upon an earlier proposal to jointly select reference frame, level of QoS and transmission path for each video frame in a multi-path streaming scenario. An algorithm that efficiently calculates the loss compensation value of an earlier correctly decodeable frame during error concealment is presented. Results show significant streaming quality improvement when loss compensation is used. Gene Cheung, Wai-tian Tan |
ICME | 2 |
| 2005 | SP-Frame Selection for Video Streaming over Burst-loss NetworksabstractSP-frame is a new picture type supported by H.264. The traditional usage of SP-frames is for switching between different compressed bit-streams. In this paper, we proposed and evaluated a scheme that uses SP frames as a mechanism to switch within a single compressed stream for the purpose of achieving error resilience and rate scalability. We have only considered the restricted but practical case in which only one secondary SP frame is allowed for every primary SP frame. Nevertheless, simulation results show that the technique can significantly increase the chance of video frames meeting their deadlines, and also improve overall PSNR. Wai-tian Tan, Gene Cheung |
ISM | 1 |
| 2005 | BINDER: An Extrusion-Based Break-In Detector for Personal Computers
Weidong Cui, Randy H. Katz, Wai-tian Tan |
USENIX ATC, General Track | 3 |
| 2005 | Rate-distortion hint tracks for adaptive video streamingabstractWe present a technique for low-complexity rate-distortion (R-D) optimized adaptive video streaming based on the concept of rate-distortion hint track (RDHT). RDHTs store the precomputed characteristics of a compressed media source that are crucial for high performance online streaming but difficult to compute in real time. This enables low-complexity adaptation to variations in transport conditions such as available data rate or packet loss. An RDHT-based streaming system has three components: 1) information that summarizes the R-D attributes of the media; 2) an algorithm for using the RDHT to predict the distortion for a feasible packet schedule; and 3) a method for determining the best packet schedule to adapt the streaming to the communication channel. A family of distortion models, denoted distortion chains, are presented which accurately predict the distortion produced by arbitrary packet loss patterns. Two distortion chain models are examined which lead to two RDHT-based techniques. We evaluate the proposed techniques for two canonical problems in streaming media, adaptation to available data rate and to packet loss. Experimental results demonstrate that for the difficult case of nonscalably coded H.264 video, the proposed systems provide significant performance gains over conventional low-complexity streaming systems, and achieve this gain with a comparable level of complexity making them suitable for online R-D optimized streaming. Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Real-time video transport optimization using streaming agent over 3G wireless networksabstractFeedback adaptation has been the basis for many media streaming schemes, whereby the media being sent is adapted in real time according to feedback information about the observed network state and application state. Central to the success of such adaptive schemes, the feedback must: 1) arrive in a timely manner and 2) carry enough information to effect useful adaptation. In this paper, we examine the use of feedback adaptation for media streaming in 3G wireless networks, where the media servers are located in wired networks while the clients are wireless. We argue that end-to-end feedback adaptation using only information provided by 3G standards is neither timely nor contain enough information for media adaptation at the server. We first show how the introduction of a streaming agent (SA) at the junction of the wired and wireless network can be used to provide useful information in a timely manner for media adaptation. We then show how optimization algorithms can be designed to take advantage of SA feedbacks to improve performance. The improvement of SA feedbacks in peak signal-to-noise ratio is significant over nonagent-based systems. Gene Cheung, Wai-tian Tan, Takeshi Yoshimura |
IEEE Trans. Multim. | 2 |
| 2004 | Distortion chains for predicting the video distortion for general packet loss patternsabstractWhen designing a system for video communication over a lossy packet network, it is highly beneficial to have a mechanism for accurately predicting the mean-squared error (MSE) distortion that results from different packet loss patterns. The paper proposes a distortion chains model for accurately predicting the end-to-end distortion for different general packet loss patterns. The performance is examined using JVT/H.264 encoded video sequences and previous frame error concealment. It is shown that, for all tested sequences, the proposed model predicts the total distortion due to a packet loss pattern within a 10% error bound 80% of the time, as compared to the conventional additive approach which achieves the same accuracy less then 40% of the time. Jacob Chakareski, John G. Apostolopoulos, Wai-tian Tan, Susie J. Wee, Bernd Girod |
ICASSP (5) | 3 |
| 2004 | Graphics-to-video encoding for 3g mobile game viewer multicast using depth values
Gene Cheung, Takashi Sakamoto, Wai-tian Tan |
ICIP | 3 |
| 2004 | R-D hint tracks for low-complexity R-D optimized video streamingabstractThis work presents the concept of rate-distortion hint track (RDHT), and evaluates two specific implementations of streaming systems that employ RDHT. Using RDHT, low-complexity streaming can be realized for systems that adapt to variations in transport conditions such as bandwidth or packet loss. An RDHT-based streaming system has three components: (1) an R-D hint track; (2) an algorithm for using the RDHT to predict the distortion for different packet schedules; and (3) a method for determining the best packet schedule. Two RDHT-based systems are presented which perform R-D optimized scheduling with dramatically reduced complexity as compared to conventional on-line R-D optimized streaming algorithms. Experimental results demonstrate that for the difficult case of R-D optimized scheduling of non-scalably coded video the proposed systems provide 7-12 dB gain when adapting to a bandwidth constraint and 2-4 dB gain when adapting to random packet loss, both relative to a conventional streaming system that does not take into account the different importance of individual packets. Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod |
ICME | 4 |
| 2004 | Double feedback streaming agent for real-time delivery of media over 3G wireless networksabstractA network agent located at the junction of wired and wireless networks can provide additional feedback information to streaming media servers to supplement feedbacks from clients. Specifically, it has been shown that feedbacks from the network agent have lower latency, and they can be used in conjunction with client feedbacks to effect proper congestion control. In this work, we propose the double-feedback streaming agent (DFSA) which further allows the detection of discrepancies in the transmission constraints of the wired and wireless networks. By working together with the streaming server and client, DFSA reduces overall packet losses by exploiting the excess capacity Of the path with more capacity. We show how DFSA can be used to support three modes of operation tailored for different delay requirements of streaming applications. Simulation results under high wireless latency show significant improvement of media quality using DFSA over non-agent-based and earlier agent-based streaming systems. Gene Cheung, Wai-tian Tan, Takeshi Yoshimura |
IEEE Trans. Multim. | 2 |
| 2003 | Low-latency wireless video over 802.11 networks using path diversityabstractWireless local area networks, such as 802.11b, are becoming wide-spread as they provide simple wireless connectivity and data delivery. This paper examines low-latency (conversational) video communication over 802.11b networks. The challenges to enable low-latency video include overcoming the highly variable delays, losses, and bandwidth of 802.11b wireless networks. To overcome these challenges we (1) employ the H.264/MPEG-4 advanced video coding (AVC) standard for high video compression efficiency and good resilience to losses, (2) use low-latency best-effort transport mechanisms, and (3) exploit the potential path diversity between each mobile client and multiple access points in the infrastructure, where we use multiple paths simultaneously or switch between multiple paths (site selection) as a function of channel characteristics. Our results indicate that the proposed system can provide significant benefits over conventional single access point (single path) systems. Allen K. L. Miu, John G. Apostolopoulos, Wai-tian Tan, Mitchell D. Trott |
ICME | 3 |
| 2003 | Research and design of a mobile streaming media content delivery networkabstractDelivering media to large numbers of mobile users presents challenges due to the stringent requirements of streaming media, mobility, wireless, and scaling to support large numbers of users. This paper presents a mobile streaming media content delivery network (MSM-CDN) designed to overcome these challenges. The MSM-CDN is a network overlay consisting of overlay servers on top of the existing network; these overlay servers are control points that facilitate end-to-end media delivery and mid-network media services. This paper presents an overview of the MSM-CDN system architecture, and describes the testbed prototype that we built based on these architectural principles. The MSM-CDN provides a new platform for media delivery, and we describe a number of research directions related to the MSM-CDN. Susie J. Wee, John G. Apostolopoulos, Wai-tian Tan, Sumit Roy 0002 |
ICME | 3 |
| 2003 | Double feedback streaming agent for real-time delivery of media over 3G wireless networksabstractA network agent located at the junction of wired and wireless networks can provide additional feedback information to streaming servers to supplement feedback from clients. Specifically, it has been shown that feedbacks from the network agent have lower latency, and can be used in conjunction with client feedbacks to effect proper congestion control. In this work, we propose the double feedback streaming agent (DFSA) which further allows the detection of discrepancies in the transmission constraints of the wired and wireless networks. By working together with the streaming server and client, DFSA reduce overall packet losses by exploiting the excess capacity of the path with more capacity. We show how DFSA can be used to support three modes of operation tailored for different delay requirements of streaming applications. Simulation results show noticeable improvement of media quality using DFSA over existing streaming systems. Gene Cheung, Wai-tian Tan, Takeshi Yoshimura |
WCNC | 2 |
| 2002 | Modeling path diversity for multiple description video communicationabstractThe use of multiple description (MD) video coding and path diversity has been proposed to provide improved performance over lossy packet networks [1]. The goal of this work was to develop models to accurately and quickly predict and compare the distortion of MD video coding and path diversity against conventional single description (SD) video delivered over a single path. In the process, we developed (1) a model for the loss process of a two-path path diversity system, and (2) a distortion model that maps the loss model to MD distortion values. Given these models we present a number of comparisons between MD video coding and path diversity and conventional SD video over a single path. The proposed model for path diversity may also be useful in other applications not related to MD coding. Furthermore, other forms of MD coding may be analyzed using similar models for MD distortion. John G. Apostolopoulos, Wai-tian Tan, Susie J. Wee, Gregory W. Wornell |
ICASSP | 2 |
| 2002 | Performance of a multiple description streaming media content delivery networkabstractContent delivery networks (CDN) have been widely used to provide reduced delay and packet loss, fault tolerance, and improved scalability for Web content delivery. Additional benefits are provided for video streaming when one designs a streaming media CDN (SM-CDN) for either conventional single description (SD) or multiple description (MD) coding. Specifically, when precise network conditions and topology are known, simulations show that an MD-SM-CDN can provide 20 to 40% reduction in distortion over a conventional SD-SM-CDN, even when the underlying CDN is not designed with MD streaming in mind. This paper examines the performance of an MD-SM-CDN as a function of different network topologies and loss conditions, and compares it with a conventional SD-SM-CDN. This examination provides insight into an MD-SM-CDN's performance when knowledge of network topology and conditions is imprecise or uncertain. Our simulations indicate that an MD-SM-CDN can provide improved performance over a conventional SD-SM-CDN over a wide range of network topologies and loss conditions. John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan |
ICIP (2) | 3 |
| 2002 | Rate-distortion optimized application-level retransmission using streaming agent for video streaming over 3G wireless networkabstractFeedback adaptation has been the basis for many media streaming schemes whereby the media being sent is adapted according to feedback information about the channel. Central to the success of such adaptive schemes, the feedback must (1) arrive in a timely manner, and (2) carry enough information to effect useful adaptation. We examine the use of feedback adaptation for media streaming in a 3G wireless network, where the media servers are located in wired networks while the clients are wireless. We argue that end-to-end feedback adaptation using only information provided by 3G standards is neither timely, nor contain enough information for media adaptation at the server. We then show how the introduction of an streaming agent (SA) at the junction of the wired and wireless network can be used to provide useful information in a timely manner for media adaptation. Gene Cheung, Wai-tian Tan, Takeshi Yoshimura |
ICIP (1) | 2 |
| 2002 | Directed acyclic graph based source modeling for data unit selection of streaming media over QoS networksabstractRegardless of the network loss process, a central problem in rate-distortion optimized video streaming is the modeling of the resulting distortion associated with the loss of different subsets of data units. When the dependency among data units is modeled by a directed acyclic graph, previous work has modeled the distortion contribution of a data unit based on only two events: whether or not the data unit and all its dependent data units are available. The restriction to only two events allows the use of a single scalar to represent the distortion contribution of a data unit, and at the same time limits its accuracy. We consider the general case when a data unit can assume different distortion contributions when different subsets of its dependent data units are available. Rate-distortion optimized streaming using the general model is performed to demonstrate the possible gains. Gene Cheung, Wai-tian Tan |
ICME (2) | 2 |
| 2002 | Optimized video streaming for networks with varying delayabstractThis paper presents a method for distortion-optimized streaming of predictively coded video over packet networks with varying delay. In networks with significant delay variations, coded video frames can arrive late at the decoder and miss their respective display deadlines. Furthermore, due to predictive coding, a late frame can also prevent a number of subsequent frames from being displayed properly, where the number of affected frames or degree of distortion depends on the particular coding dependencies of the late frame. In this paper, we present an optimized video streaming strategy based on frame reordering for networks with significant delay variations. This streaming strategy minimizes distortion by exploiting the fact that different late frames result in different degrees of distortion. We model the router-induced delay in a wired network with an analytical PDF and we model the link-layer retransmission delay of a wireless network with the 3GPP specification for W-CDMA radio link control. We compute the distortion for different frame reorderings using the network delay models and a source model that accounts for the prediction dependencies of predictively coded video. Our optimized streaming strategies are shown to reduce the number of late frames by 14 to 23% for the situations examined. Susie J. Wee, Wai-tian Tan, John G. Apostolopoulos, Minoru Etoh |
ICME (2) | 2 |
| 2001 | Video multicast using layered FEC and scalable compressionabstractThe use of scalable video with layered multicast has been shown to be an effective method to achieve rate control in heterogeneous networks. We propose the use of layered forward error correction (FEC) as an error-control mechanism in a layered multicast framework. By organizing FEC into multiple layers, receivers can obtain different levels of protection commensurate with their respective channel conditions. Efficient network utilization is achieved as FEC streams are multicast, and only to receivers that need them. Furthermore, FEC is used without overall rate expansion by selectively dropping data layers to make room for FEC layers. Effects of bursty losses are amortized by staggering the FEC streams in time, giving rise to a tradeoff between delay and quality. For rate control at the receivers, we propose an equation-based approach that computes network usage as a function of measured network characteristics. We show that equation-based rate control achieves more fair bandwidth sharing amongst competing sessions as compared to existing multicast rate control schemes such as RLM and RLC. Fairness is achieved since competing sessions sharing a path will measure similar network characteristics. Simulations and actual MBONE experiments are performed using error-resilient, scalable video compression. We find that video quality is significantly improved at the same communication rate when layered FEC is used. Wai-tian Tan, Avideh Zakhor |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Error Control for Video Multicast Using Hierarchical FecabstractBit-rate scalable video compression with layered multicast has been shown to be an effective method to achieve rate control in heterogeneous networks. We further propose the use of hierarchical FEC as an error control mechanism that allows receivers to individually trade-off latency for received video quality. The scheme is efficient since FEC packets are used to protect only the more important data layers and is multicast only to receivers that need them, thereby improving network utilization. Furthermore, there is no loss in error correcting capability by using hierarchical FEC when maximum distance separable codes are used. Actual MBONE experiments are performed to evaluate the performance of the proposed scheme. Wai-tian Tan, Avideh Zakhor |
ICIP (1) | 1 |
| 1999 | Real-Time Internet Video Using Error Resilient Scalable Compression and TCP-Friendly Transport ProtocolabstractWe introduce a point to point real-time video transmission scheme over the Internet combining a low-delay TCP-friendly transport protocol in conjunction with a novel compression method that is error resilient and bandwidth-scalable. Compressed video is packetized into individually decodable packets of equal expected visual importance. Consequently, relatively constant video quality can be achieved at the receiver under lossy conditions. Furthermore, the packets can be truncated to instantaneously meet the time varying bandwidth imposed by a TCP-friendly transport protocol. As a result, adaptive flows that are friendly to other Internet traffic are produced. Actual Internet experiments together with simulations are used to evaluate the performance of the compression, transport, and the combined schemes. Wai-tian Tan, Avideh Zakhor |
IEEE Trans. Multim. | 1 |
| 1998 | Internet Video using Error Resilient Scalable Compression and Cooperative Transport Protocol
Wai-tian Tan, Avideh Zakhor |
ICIP (3) | 1 |
| 1996 | Real time software implementation of scalable video codecabstractScalable video compression is becoming increasingly more important in diverse, heterogeneous networks of today. In a previous work, Taubman and Zakhor (see IEEE Transactions on Image Processing, vol.3, no.5, p.572-88, 1994) developed a scalable codec capable of generating bit rates from tens of kilo bits per second to several mega bits per second with fine granularity of the available bit rates. This codec is based on 3-D subband coding and multi-rate quantization of subband coefficients, followed by arithmetic coding. We replace the arithmetic coding portion of the Taubman codec with block coding, and compare the encode/decode speed of this new coder with MPEG. Unlike MPEG, this codec requires symmetric computational power at the decoder and encoder and as such is useful in software only, real time, interactive video applications. We have found the encoding speed of the new encoder to be one order of magnitude faster than MPEG-1, without significant loss in compression efficiency. Wai-tian Tan, Avideh Zalchor |
ICIP (1) | 1 |