Jacob Chakareski

dblp:29/2042 · DBLP profile ↗
← Back
125ranked-venue papers
61as first author
28since 2021 · last 2026
0000-0003-2428-9518ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 85 · 51 first-author · 16 since 2021Computer networks · 36 · 11 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 ISM: Intelligent Multi-Path Scheduler for Multi-Camera Networked Systems
Alireza Mohammadhosseini, Jacob Chakareski, Mallesham Dasari
MMSys2
2026 SPARC: Proximity-aware Scheduling of AR Mapping and Cloud-based GenAI Upsampling for Efficient Multi-User SLAM
abstract
The scalability of multi-user SLAM is fundamentally limited by the constrained network and computational resources. Existing approaches either focus on SLAM for single-user scenarios or overload networks and servers by streaming dense, uniform camera data and treating all users equally. This results in poor pose estimation accuracy or slow updates to multiple users. Our key insight is that the sparsity and heterogeneity of user activity reveal that not all users or frames contribute equally to the shared map. Building on this, we propose SPARC - Proximity-aware Scheduling of AR Mapping and a blur-aware adaptive Cloud-based GenAI sampling method, which together form a cloud-native framework for efficient multi-user SLAM. On the client side, adaptive, context-aware frame transmission selectively forwards high-value frames. On the server side, generative AI (GenAI)-based upsampling reconstructs dense scene features from sparse inputs, while a proximity-aware scheduler prioritizes updates for users with higher drift or critical interactions. Together, these components reduce redundant transmission, improve resource allocation, and enable fairness without sacrificing accuracy. We show through extensive experimentation that our method reduces the latency by 2× to 4× compared to state-of-the-art while maintaining similar or better tracking accuracy. More broadly, this work reimagines SLAM as a cloud-native service, paving the way for scalable, real-time AR/VR applications where many users seamlessly interact in shared environments.
Shneka Muthu Kumara Swamy, Mallesham Dasari, Nicholas Mastronarde, Jacob Chakareski
MMSys4
2025 ASL360: AI-Enabled Adaptive Streaming of Layered 360° Video over UAV-assisted Wireless Networks
abstract
We propose ASL360, an adaptive deep reinforcement learning-based scheduler for on-demand 360° video streaming to mobile VR users in next generation wireless networks. We aim to maximize the overall Quality of Experience (QoE) of the users served over a UAV-assisted 5G wireless network. Our system model comprises a macro base station (MBS) and a UAV-mounted base station which both deploy mm-Wave transmission to the users. The 360°video is encoded into dependent layers and segmented tiles, allowing a user to schedule downloads of each layer’s segments. Furthermore, each user utilizes multiple buffers to store the corresponding video layer’s segments. We model the scheduling decision as a Constrained Markov Decision Process (CMDP), where the agent selects Base or Enhancement layers to maximize the QoE and use a policy gradient-based method (PPO) to find the optimal policy. Additionally, we implement a dynamic adjustment mechanism for cost components, allowing the system to adaptively balance and prioritize the video quality, buffer occupancy, and quality change based on real-time network and streaming session conditions. We demonstrate that ASL360 significantly improves the QoE, achieving approximately 2 dB higher average video quality, 80% lower average rebuffering time, and 57% lower video quality variation, relative to competitive baseline methods. Our results show the effectiveness of our layered and adaptive approach in enhancing the QoE in immersive video streaming applications, particularly in dynamic and challenging network environments.
Alireza Mohammadhosseini, Jacob Chakareski, Nicholas Mastronarde
GLOBECOM2
2025 Low-Complexity Physics-Informed Reinforcement Learning Using Post-Decision States with Stochastic Sampling
abstract
Delay-sensitive Internet of Things (IoT) applications continue to grow in prevalence as new wireless technologies are adopted. Since these applications often operate in unknown dynamic environments, reinforcement learning (RL) has emerged as an effective method to learn optimal decision policies that improve their overall performance. However, typical data-driven RL techniques that have been adopted to solve these problems do not exploit available knowledge of system dynamics. Consequently, they must “learn” some information about the system that may already be known to the system's designer. Post-decision state (PDS) learning, on the other hand, leverages known system information (i.e., it is “physics-informed”) to simplify the learning task and improve learning performance. However, this comes at the cost of increased computational complexity, and makes it impractical to implement on resource constrained devices. This work introduces stochastic PDS learning, a novel RL algorithm that combines traditional PDS learning with stochastic sampling to produce a physics-informed RL agent that can leverage known system information even with limited computational resources. Performance of stochastic PDS learning is compared against numerous traditional RL algorithms in the context of a delaysensitive energy-efficient scheduling problem simulated as an environment in Gymnasium.
Andrew Corra, Nicholas Mastronarde, Jacob Chakareski
ICC3
2025 Bayes-Split-Edge: Bayesian Optimization for Constrained Collaborative Inference in Wireless Edge Systems
abstract
Mobile edge devices (e.g., AR/VR headsets) typically need to complete timely inference tasks while operating with limited on-board computing and energy resources. In this paper, we investigate the problem of collaborative inference in wireless edge networks, where energy-constrained edge devices aim to complete inference tasks within given deadlines. These tasks are carried out using neural networks, and the edge device seeks to optimize inference performance under energy and delay constraints. The inference process can be split between the edge device and an edge server, thereby achieving collaborative inference over wireless networks. We formulate an inference utility optimization problem subject to energy and delay constraints, and propose a novel solution called Bayes-Split-Edge, which leverages Bayesian optimization for collaborative split inference over wireless edge networks. Our solution jointly optimizes the transmission power and the neural network split point. The Bayes-Split-Edge framework incorporates a novel hybrid acquisition function that balances inference task utility, sample efficiency, and constraint violation penalties. We evaluate our approach using the VGG19 model on the ImageNet-Mini dataset, and Resnet101 on Tiny-ImageNet, and real-world mMobile wireless channel datasets. Numerical results demonstrate that Bayes-Split-Edge achieves up to 2.4× reduction in evaluation cost compared to standard Bayesian optimization and achieves near-linear convergence. It also outperforms several baselines, including CMA-ES, DIRECT, exhaustive search, and Proximal Policy Optimization (PPO), while matching exhaustive search performance under tight constraints. These results confirm that the proposed framework provides a sample-efficient solution requiring maximum 20 function evaluations and constraint-aware optimization for wireless split inference in edge computing systems.
Fatemeh Zahra Safaeipour, Jacob Chakareski, Morteza Hashemi
SEC2
2025 Anywhere Avatar: 3D Telepresence with Just a Phone and a Laptop
abstract
We present Anywhere Avatar, a telepresence system that enables full-body and facial avatar reconstruction using a smartphone and a laptop. Users record short videos to generate personalized avatars, which are animated in real time during teleconferencing using webcam-based tracking. Built on pre-trained FLAME and SMPL models, the avatars are rendered in high fidelity using Gaussian splatting. The system runs at near real-time with minimal bandwidth, making expressive 3D telepresence accessible without specialized hardware.
Ruifan Ji, Mingyuan Wu, Bo Chen 0025, Michael Zink, Ramesh K. Sitaraman, Jacob Chakareski, Klara Nahrstedt
ACM Multimedia6
2025 Reinforcement Learning-Based Dynamic Resource Allocation for Aerial 360° Video VR Streaming
abstract
Efficient power use and accurate viewport information are key factors in enabling effective aerial 360° video delivery to virtual reality (VR) clients for emerging remote immersion societal applications. We explore a learning-based framework for transmission power allocation and robust viewport identification in UAV-based 360° video streaming to a ground user/VR client that aims to maximize the delivered viewport quality and minimize the video playback stall time on the user’s VR headset. We model the problem of interest as a Markov decision process (MDP) encompassing the UAV’s transmit power, the VR client’s video playback stall time, and the full-identification outage of the user’s observed viewport in the MDP reward function. Our framework integrates an effective scalable 360° video tiling representation of the captured content that ensures for the client (i) maximum delivered viewport quality given the available UAV transmission rate and (ii) VR application robustness to partial viewport outages, at the same time. We formulate a novel learning-based method for adaptive transmission power allocation and predicted viewport enlargement, to solve the problem of interest. Relative to multiple reference methods, we demonstrate through experiments that our framework can achieve up to 8dB improvement in viewport PSNR and an 85% reduction in full-viewport identification outage, while using 60% less transmit power and experiencing negligible video stall times.
Jacob Chakareski, Lingdong Wang, Nicholas Mastronarde
MMSP1
2025 FBDT: Sum-Throughput Achieving Transport Layer Solution for Multi-RAT Networks
abstract
Emerging mobile applications give rise to new bandwidth-hungry and latency-sensitive traffic classes that challenge existing wireless systems. Addressing them requires innovative approaches such as simultaneous data transmission across multiple Radio Access Technologies (RATs), e.g., WiFi and WiGig. However, existing transport layer multi-RAT traffic aggregation schemes, e.g., multi-path TCP, suffer from Head-of-Line (HoL) blocking and sub-optimal traffic splitting across the RATs that severely penalize their performance. In this paper, we investigate the design of FBDT, a novel multi-path transport layer solution that for the first time can achieve the sum of the throughput rates across the individual RATs network paths, despite their channel conditions' dynamics. We have implemented FBDT in the Linux kernel and show substantial improvement in throughput relative to state-of-the-art schemes, e.g, 2.5x gain in a dual-RAT scenario (WiFi and WiGig) when the client is mobile. Second, we extend FBDT to more than two radios and demonstrate that its throughput performance scales linearly with the number of RATs, in contrast to multi-path TCP, whose performance degrades with an increase in the number of RATs. We evaluate the performance of FBDT on different traffic classes and demonstrate: (i) 2-3 times shorter file download times, (ii) up to 10 times shorter streaming times and 10 dB higher video quality for progressive download video applications, and (iii) up to 9 dB higher viewport quality for interactive mobile VR applications, when our viewport quality maximization framework is employed along with FBDT.
Suresh Srinivasan, Sam Shippey, Ehsan Aryafar, Jacob Chakareski
IEEE Trans. Mob. Comput.4
2025 Neural-Enhanced Rate Adaptation and Computation Distribution for Emerging mmWave Multi-User 3D Video Streaming Systems
abstract
We investigate multitask edge-user communication-computation resource allocation for$360^\circ$video streaming in an edge-computing enabled millimeter wave (mmWave) multi-user virtual reality system. To balance the communication-computation trade-offs that arise herein, we formulate a video quality maximization problem that integrates interdependent multitask/multi-user action spaces and rebuffering time/quality variation constraints. We formulate a deep reinforcement learning framework formulti-taskrate adaptation andcomputation distribution (MTRC) to solve the problem of interest. Our solution does not rely on a priori knowledge about the environment and uses only prior video streaming statistics (e.g., throughput, decoding time, and transmission delay), and content information, to adjust the assigned video bitrates and computation distribution, as it observes the induced streaming performance online. Moreover, to capture the task interdependence in the environment, we leverage neural network cascades to extend our MTRC method to two novel variants denoted as R1C2 and C1R2. We train all three methods with real-world mmWave network traces and$360^\circ$video datasets to evaluate their performance in terms of expected quality of experience (QoE), viewport peak signal-to-noise ratio (PSNR), rebuffering time, and quality variation. We outperform state-of-the-art rate adaptation algorithms, with C1R2 showing best results and achieving$5.21-6.06$dB PSNR gains,$2.18-2.70$x rebuffering time reduction, and$4.14-4.50$dB quality variation reduction.
Babak Badnava, Jacob Chakareski, Morteza Hashemi
IEEE Trans. Multim.2
2024 Joint Communication and Computation Resource Allocation for Emerging mmWave Multi-User 3D Video Streaming Systems
abstract
We consider a multi-user joint rate adaptation and computation distribution problem in a millimeter wave (mmWave) virtual reality (VR) system. The VR system that we consider comprises an edge computing unit (ECU) that serves 360° videos to VR users. We formulate a multi-user quality of experience (QoE) maximization problem, in which VR users are assisted with the ECU to decode/render 360° videos. The ECU provides additional computational resources that can be used for processing video frames, at the expense of increased data volume and required bandwidth. To balance this trade-off, we leverage deep reinforcement learning (DRL) for joint rate adaptation and computational resource allocation optimization. Our proposed method, dubbed Deep VR, does not rely on any predefined assumption about the environment and relies on video playback statistics (i.e., past throughput, decoding time, transmission time, etc.), video information, and the resulting performance to adjust the video bitrate and computation distribution. We train Deep VR with real-world mmWave network traces and 360° video datasets to obtain evaluation results in terms of the average QoE, peak signal-to-noise ratio (PSNR), rebuffering time, and quality variation. Our results indicate that the Deep VR improves the users’ QoE compared to state-of-the-art rate adaptation algorithm. Specifically, we show a 3.08 dB to 4.49 dB improvement in video quality in terms of PSNR, a 12.5x to 14x reduction in rebuffering time, and a 3.07 dB to 3.96 dB improvement in quality variation.
Babak Badnava, Jacob Chakareski, Morteza Hashemi
GLOBECOM2
2024 Scene Graph Driven Hybrid Interactive VR Teleconferencing
abstract
We propose an interactive and intelligent hybrid teleconferencing system compatible with Virtual Reality devices. Our system understands meeting contexts and leverages user interactions to enhance better system configuration. Employing interactive scene graphs [11], the system extracts and transmits essential meeting context to users while relaying user interactions back to the streaming systems for user-involved adaptive streaming and foveated rendering. We demonstrate the system's real-time performance and compatibility with commercial VR devices such as the Meta Quest 3.
Mingyuan Wu, Ruifan Ji, Haozhen Zheng, Beitong Tian, Bo Chen 0025, Jacob Chakareski, Michael Zink, Ramesh K. Sitaraman, Klara Nahrstedt
ACM Multimedia8
2024 Comparative Analysis and Performance Evaluation of Adaptive 360° Video DASH Streaming Solutions
abstract
We explore two competitive viewport-adaptive 360° video streaming methods: Quality Emphasized Region (QER) representations and 360° panorama Tiling, across multiple performance factors. We carry out our analysis and evaluation within the context of the MPEG Dynamic Adaptive Streaming over HTTP (DASH) standard. The performance criteria we use in our analysis include delivered video quality (PSNR, VMAF), encoding and decoding time/complexity per video frame, and required storage. We also formulate analysis for immersive quality optimization in both streaming systems. Our comprehensive experimental evaluation provides valuable insights for streaming platforms and service operators aiming to optimize the viewer experience while achieving high system efficiency. In particular, we show that Tiled streaming provides 35% lower encoding bitrate and 52% lower encoding time. However, QER streaming can offer significant advantages in reducing the induced decoding complexity/time by approximately 38% while enabling 0.9-1.7 dB higher viewport quality on average and also 2.5-3 dB higher quality in optimal rate allocation scenario, resulting in smoother video playback and higher quality of viewing experience. Additionally, QER streaming demonstrates superior video quality performance in longer video segments by effectively covering larger areas of the video. Specifically, by increasing the video segment duration from 1 second to 8 seconds, the QER method shows an average 0.5 dB improvement in quality performance compared to the tiled method.
Alireza M. Hosseini, Jacob Chakareski
MMSP2
2024 Live 360° Video Streaming to Heterogeneous Clients in 5G Networks
abstract
We investigate rate-distortion-computing optimized live 360◦ video streaming to heterogeneous mobile VR clients in 5G networks. The client population comprises devices that feature single (LTE) or dual (LTE/NR) cellular connectivity. The content is compressed using scalable 360◦ tiling at the origin and sent towards the clients over a single backbone network link. A mobile edge server then adapts the incoming streaming data to the individual clients and their respective down-link transmission rates using formal rate-distortion-computing optimization. Single connectivity clients are served by the edge server a baseline representation/layer of the content adapted to their down-link transmission capacity and device computing capability. A dual connectivity client is served in parallel a baseline content layer on its LTE connectivity and a complementary viewport-specific enhancement layer on its NR connectivity, synergistically adapted to the respective down-links’ transmission capacities and its computing capability. We formulate two optimization problems to conduct the operation of the edge server in each case, taking into account the key system components of the delivery process and induced end-to-end latency, aiming to maximize the immersion fidelity delivered to each client. We explore respective geometric programming optimization strategies that compute the optimal solutions at lower complexity. We rigorously analyze the computational complexity of the two optimization algorithms we formulate. In our evaluation, we demonstrate considerable performance gains over multiple assessment factors relative to two state-of-the-art techniques. We also examine the robustness of our approach to inaccurate user navigation prediction, transient NR link loss, dynamic LTE bandwidth variations, and diverse 360◦ video content. Finally, we contrast our results over five popular video quality metrics and share our evaluation dataset.
Jacob Chakareski, Mahmudur Khan 0002
IEEE Trans. Multim.1
2023 Aerial 360-Degree Video Delivery for Immersive First Person View UAV Navigation
abstract
Adaptive transmission of conventional video from a UAV to the ground has been researched for various applications, but the research topic of 360° video transmission from a UAV for the specific application of first-person view (FPV) based navigation is still nascent. In this work, we present adaptive 360° video compression and streaming methods to optimize the perceptual quality of experience of a pilot, who navigates the UAV in real time by viewing this immersive FPV feed, which is sent wirelessly from the UAV to the pilot. This adaptation of the 360° FPV feed is performed in response to the wireless channel conditions and the pilot’s viewport, wherein each 360° frame is split into two regions of variable size, one meant to be within the pilot’s viewport and the other outside. Each region is encoded using different H. 265 quantization parameters (QP) and modulation orders. We model the scenario realistically by generating probability distributions of the variation in frame size and quality with QP, for aerial 360° videos. These models are expressed using a two-term exponential function, whose parameters are also provided. This model achieves lower prediction errors than the single-term exponential and power law functions. Simulations on a set of aerial 360-degree videos demonstrate that the adaptive approach achieves 9.73 dB (21.77 %) greater QoE than a baseline approach that utilizes throughput-based adaptive bit rate algorithm (ABR) to tune QP per GoP, and a 5G new radio adaptive modulation scheme (AMS) to tune modulation order: Additionally, we present a deep reinforcement learning approach to adapt FPV, which achieves an expected pilot QoE just 2.07 dB lower than the adaptive approach, while being significantly faster and requiring no prior knowledge of the environment.
Simran Singh, Jacob Chakareski
ISM2
2023 Interactive Scene Graph Analysis for Future Intelligent Teleconferencing Systems
abstract
In a real-life meeting environment, individuals often demonstrate a remarkable ability to selectively focus their attention on specific visual information. This ability allows them to naturally concentrate on a specific region of interest while tuning out others. Understanding and exploiting such selective attention remains unexplored in a user-centric teleconferencing system, where there is a potential to customize video streaming and foveated rendering based on the viewer’s attention. This paper proposes a novel user-centric scene analysis module that fully leverages the power of selective attention for online meeting scenarios and recognizes the unequal importance of individual pixels in the videos. The module determines the user’s selective attention through the meeting contexts. The contextual representation of the meeting is modeled as a combination of two primary components: proactive user interaction within the system and passive real-time analysis of high-level visual semantics from the scenes. As the meeting progresses, the interactive scene analysis module dynamically updates its contextual representation, offering a dual advantage: (a) Videos can be selectively and adaptively streamed within a user’s attention, resulting in bandwidth savings of up to 78 percent. (b) The module enhances the overall quality of the user experience by facilitating higher user interactivity, particularly in meeting-related tasks such as screen sharing, privacy-preserving user blocking, background removal, automatic user attention shift detection, etc. Our interactive scene analysis module makes significant progress toward enabling an efficient, immersive, and intelligent teleconferencing system.
Mingyuan Wu, Yuhan Lu, Shiv Trivedi, Bo Chen 0025, Qian Zhou 0008, Lingdong Wang, Simran Singh, Michael Zink, Ramesh K. Sitaraman, Jacob Chakareski, Klara Nahrstedt
ISM10
2023 FBDT: Forward and Backward Data Transmission Across RATs for High Quality Mobile 360-Degree Video VR Streaming
abstract
The metaverse encompasses many virtual universes and relies on streaming high-quality 360° videos to VR/AR headsets. This type of video transmission requires very high data rates to meet the desired Quality of Experience (QoE) for all clients. Simultaneous data transmission across multiple Radio Access Technologies (RATs) such as WiFi and WiGig is a key solution to meet this required capacity demand. However, existing transport layer multi-RAT traffic aggregation schemes suffer from Head-of-Line (HoL) blocking and sub-optimal traffic splitting across the RATs, particularly when there is a high fluctuation in their channel conditions. As a result, state-of-the-art multi-path TCP (MPTCP) solutions can achieve aggregate transmission data rates that are lower than that of using only a single WiFi RAT in many practical settings, e.g., when the client is mobile. We make two key contributions to enable high quality mobile 360° video VR streaming using multiple RATs. First, we propose the design of FBDT, a novel multi-path transport layer solution that can achieve the sum of individual transmission rates across the RATs despite their system dynamics. We implemented FBDT in the Linux kernel and showed substantial improvement in transmission throughput relative to state-of-the-art schemes, e.g, 2.5x gain in a dual-RAT scenario (WiFi and WiGig) when the VR client is mobile. Second, we formulate an optimization problem to maximize a mobile VR client's viewport quality by taking into account statistical models of how clients explore the 360° look-around panorama and the transmission data rate of each RAT. We explore an iterative method to solve this problem and evaluate its performance through measurement-driven simulations leveraging our testbed. We show up to 12 dB increase in viewport quality when our optimization framework is employed.
Suresh Srinivasan, Sam Shippey, Ehsan Aryafar, Jacob Chakareski
MMSys4
2023 Performance Evaluation of 5G Delay-Sensitive Single-Carrier Multi-User Downlink Scheduling
abstract
The coexistence of a wide variety of different applications with diverse Quality of Service (QoS) requirements calls for more sophisticated radio resource scheduling (RRS) in 5G networks compared to previous generations. To address this challenge, a growing body of research formulates the RRS problem as a Markov decision process (MDP) and aims to solve it using deep reinforcement learning (DRL). A key consideration when formulating an MDP is the choice of reward function, which determines the goal of the decision agent. Despite the reward function being a critical component of an MDP, there is currently no systematic study comparing how different reward functions affect network performance. To this end, we carry out a comparative study of the delay and overflow performance using several reward functions that aim to minimize packet delays. Through extensive simulations under different traffic and channel conditions, we identify a reward function that can achieve near optimal delay with up to 55 − 67% fewer packet drops than the other investigated options, and does not require any tuning.
Anjali Omer, Filippo Malandra, Jacob Chakareski, Nicholas Mastronarde
PIMRC3
2023 CLoSER: Video caching in small-cell edge networks with local content sharing
Shadab Mahboob, Koushik Kar, Jacob Chakareski, Md. Ibrahim Ibne Alam
Comput. Networks3
2023 Hierarchical Q-learning-enabled neutrosophic AHP scheme in candidate relay set size adaption in vehicular networks
Mohammad Naderi, Jacob Chakareski, Mohammed Ghanbari 0001
Comput. Networks2
2023 mmWave Networking and Edge Computing for Scalable 360° Video Multi-User Virtual Reality
abstract
We investigate a novel multi-user mobile Virtual Reality (VR) arcade system for streaming scalable 8K 360° video with low interactive latency, while providing high remote scene immersion fidelity and application reliability. This is achieved through the integration of embedded multi-layer 360° tiling, edge computing, and wireless multi-connectivity that comprises sub-6 GHz and mmWave (millimeter wave) links. The sub-6 GHz band is used for broadcast of the base layer of the entire 360° panorama to all users, while the directed mmWave links are used for high-rate transmission of VR-enhancement layers that are specific to the viewports of the individual users. The viewport-specific enhancements can comprise compressed and raw 360° tiles, decoded first at the edge server. We aim to maximize the smallest immersion fidelity for the delivered 360 content across all VR users, given rate, latency and computing constraints. We characterize analytically the rate-distortion trade-offs across the spatiotemporal 360° panorama and the computing power required to decompress 360° tiles. The proposed solution consists of geometric programming algorithms and an intermediate step of graph-theoretic VR user to mmWave access point assignment. The results reveal a significant improvement (8-10 dB) in delivered VR user immersion fidelity and spatial resolution (8K vs. 4K) compared to a state-of-the-art method based on sub-6 GHz transmission only. We also show that an increasing number of raw 360° tiles are sent, as the mmWave network link data rate or the edge server/user computing power increase. Finally, we demonstrate that in order to hypothetically deliver the same immersion fidelity, the reference method would incur a much higher (2.5-4.5x) system latency.
Sabyasachi Gupta, Jacob Chakareski, Petar Popovski
IEEE Trans. Image Process.2
2023 User Navigation Modeling, Rate-Distortion Analysis, and End-to-End Optimization for Viewport-Driven 360$^\circ $ Video Streaming
abstract
The emerging technologies of Virtual Reality (VR) and 360$^\circ $video introduce new challenges for state-of-the-art video communication systems. Enormous data volume and spatial user navigation are unique characteristics of 360$^\circ$videos that necessitate a space-time effective allocation of the available network streaming bandwidth over the 360$^\circ$video content to maximize the Quality of Experience (QoE) delivered to the user. Towards this objective, we investigate a framework for viewport-driven rate-distortion optimized 360$^\circ$video streaming that integrates the user view navigation patterns and the spatiotemporal rate-distortion characteristics of the 360$^\circ$video content to maximize the delivered user viewport video quality, for the given network/system resources. The framework comprises a methodology for assigning dynamic navigation likelihoods over the 360$^\circ$video spatiotemporal panorama, induced by the user navigation patterns, an analysis and characterization of the 360$^\circ$video panorama's spatiotemporal rate-distortion characteristics that leverage preprocessed spatial tilling of the content, and an optimization problem formulation and solution that capture and aim to maximize the delivered expected viewport video quality, given a user's navigation patterns, the 360$^\circ$video encoding/streaming decisions, and the available system/network resources. We formulate a Markov model to capture the navigation patterns of a user over the 360$^\circ$video panorama and simultaneously extend our actual navigation datasets by synthesizing additional realistic navigation data. Moreover, we investigate the impact of using two different tile sizes for equirectangular tiling of the 360$^\circ$video panorama. Our experimental results demonstrate the advantages of our framework over the conventional approach of streaming a monolithic uniformly-encoded 360$^\circ$video and a state-of-the-art navigation-speed based reference method. Considerable average and instantaneous viewport video quality gains of up to 5 dB are demonstrated in the case of five popular 4 K 360$^\circ$videos. In addition, we explore the impact of two different popular 360$^\circ$video quality metrics applied to evaluate the streaming performance of our system framework and the two reference methods. Finally, we demonstrate that by exploiting the unequal rate-distortion characteristics of the different spatial sectors of the 360$^\circ$video panorama, we can enable spatially more uniform and temporally higher 360$^\circ$video viewport quality delivered to the user, relative to monolithic streaming.
Jacob Chakareski, Xavier Corbillon, Gwendal Simon, Viswanathan (Vishy) Swaminathan
IEEE Trans. Multim.1
2023 Millimeter Wave and Free-space-optics for Future Dual-connectivity 6DOF Mobile Multi-user VR Streaming
abstract
Dual-connectivity streaming is a key enabler of next-generation six Degrees Of Freedom (6DOF) Virtual Reality (VR) scene immersion. Indeed, using conventional sub-6 GHz WiFi only allows to reliably stream a low-quality baseline representation of the VR content, while emerging high-frequency communication technologies allow to stream in parallel a high-quality user viewport-specific enhancement representation that synergistically integrates with the baseline representation to deliver high-quality VR immersion. We investigate holistically as part of an entire future VR streaming system two such candidate emerging technologies, Free Space Optics (FSO) and millimeter-Wave (mmWave), that benefit from a large available spectrum to deliver unprecedented data rates. We analytically characterize the key components of the envisioned dual-connectivity 6DOF VR streaming system that integrates in addition edge computing and scalable 360° video tiling, and we formulate an optimization problem to maximize the immersion fidelity delivered by the system, given the WiFi and mmWave/FSO link rates, and the computing capabilities of the edge server and the users’ VR headsets. This optimization problem is mixed integer programming of high complexity and we formulate a geometric programming framework to compute the optimal solution at low complexity. We carry out simulation experiments to assess the performance of the proposed system using actual 6DOF navigation traces from multiple mobile VR users that we collected. Our results demonstrate that our system considerably advances the traditional state of the art and enables streaming of 8K-120 frames-per-second (fps) 6DOF content at high fidelity.
Jacob Chakareski, Mahmudur Khan 0002, Tanguy Ropitault, Steve Blandino
ACM Trans. Multim. Comput. Commun. Appl.1
2022 A review of AI-enabled routing protocols for UAV networks: Trends, challenges, and future outlook
abstract
Unmanned Aerial Vehicles (UAVs), as a recently emerging technology, enabled a new breed of unprecedented applications in different domains. This technology's ongoing trend is departing from large remotely-controlled drones to networks of small autonomous drones to collectively complete intricate tasks time and cost-effectively. An important challenge is developing efficient sensing, communication, and control algorithms that can accommodate the requirements of highly dynamic UAV networks with heterogeneous mobility levels. Recently, the use of Artificial Intelligence (AI) in learning-based networking has gained momentum to harness the learning power of cognizant nodes to make more intelligent networking decisions by integrating computational intelligence into UAV networks. An important example of this trend is developing learning-powered routing protocols, where machine learning methods are used to model and predict topology evolution, channel status, traffic mobility, and environmental factors for enhanced routing. This paper reviews AI-enabled routing protocols designed primarily for aerial networks, including topology-predictive and self-adaptive learning-based routing algorithms, with an emphasis on accommodating highly-dynamic network topology. To this end, we justify the importance and adaptation of AI into UAV network communications. We also address, with an AI emphasis, the closely related topics of mobility and networking models for UAV networks, simulation tools and public datasets, and relations to UAV swarming, which serve to choose the right algorithm for each scenario. We conclude by presenting future trends, and the remaining challenges in AI-based UAV networking, for different aspects of routing, connectivity, topology control, security and privacy, energy efficiency, and spectrum sharing.1
Arnau Rovira-Sugranes, Abolfazl Razi, Fatemeh Afghah, Jacob Chakareski
Ad Hoc Networks4
2022 Hardware Acceleration for Postdecision State Reinforcement Learning in IoT Systems
abstract
Reinforcement learning (RL) is increasingly being used to optimize resource-constrained wireless Internet of Things (IoT) devices. However, existing RL algorithms that are lightweight enough to be implemented on these devices, such as$Q$-learning, converge too slowly to effectively adapt to the experienced information source and channel dynamics, while deep RL algorithms are too complex to be implemented on these devices. By integrating basic models of the IoT system into the learning process, the so-called postdecision state (PDS)-based RL can achieve faster convergence speeds than these alternative approaches at lower complexity than deep RL; however, its complexity may still hinder the real-time and energy-efficient operations on IoT devices. In this article, we develop efficient hardware accelerators for PDS-based RL. We first develop an arithmetic hardware acceleration architecture and then propose a stochastic computing (SC)-based reconfigurable hardware architecture. By using simple bitwise computations enabled by SC, we eliminate costly multiplications involved in PDS learning, which simultaneously reduces the hardware area and power consumption. We show that the computational efficiency can be further improved by using extremely short stochastic representations without sacrificing learning performance. We demonstrate our proposed approach on a simulated wireless IoT sensor that must transmit delay-sensitive data over a fading channel while minimizing its energy consumption. Our experimental results show that our arithmetic accelerator is$5.3\times $faster than$Q$-learning and$2.6\times $faster than a baseline hardware architecture, while the proposed SC-based architecture further reduces the critical path of the arithmetic accelerator by 87.9%.
Jianchi Sun, Nikhilesh Sharma, Jacob Chakareski, Nicholas Mastronarde, Yingjie Lao
IEEE Internet Things J.3
2021 Head Rotation Model for Virtual Reality System Level Simulations
abstract
Virtual Reality (VR) promises immersive experiences in diverse areas such as gaming, entertainment, education, healthcare, and remote monitoring. In VR environments, users can navigate 360-degree content by moving or looking around in all directions, by rotating their heads, as in real life. A rapid head rotation can corrupt the wireless link, degrading the user experience. Due to the lack of proper head rotation models, testbeds are usually required to analyze VR systems. In this paper, we propose an open source code package that generates realistic head rotation traces. The code package is based on a simple, yet flexible, time-correlated mathematical model, which is extrapolated from a publicly available VR head rotation measurement-based dataset. We show that the probability density function of head rotation pitch and roll angles can be modeled as Gaussian distributions, while the probability density function of yaw angles can be modeled as a Gaussian mixture distribution. To introduce temporal correlation, we extrapolate the power spectral density of the angular processes, which are modeled with a bi-exponential decay. Finally, we show how the model can support and accelerate the design of future VR systems by proposing the analysis of a distributed Multiple Input Multiple Output (MIMO) system and the design of a situational awareness Machine Learning (ML) based beamforming training for millimeter wave networks.
Steve Blandino, Tanguy Ropitault, Raied Caromi, Jacob Chakareski, Mahmudur Khan 0002, Nada Golmie
ISM4
2021 Full UHD 360-Degree Video Dataset and Modeling of Rate-Distortion Characteristics and Head Movement Navigation
abstract
We investigate the rate-distortion (R-D) characteristics of full ultra-high definition (UHD) 360° videos and capture corresponding head movement navigation data of virtual reality (VR) headsets. We use the navigation data to analyze how users explore the 360° look-around panorama for such content and formulate related statistical models. The developed R-D characteristics and modeling capture the spatiotemporal encoding efficiency of the content at multiple scales and can be exploited to enable higher operational efficiency in key use cases. The high quality expectations for next generation immersive media necessitate the understanding of these intrinsic navigation and content characteristics of full UHD 360° videos.
Jacob Chakareski, Ridvan Aksu, Viswanathan (Vishy) Swaminathan, Michael Zink
MMSys1
2021 Wifi-VLC dual connectivity streaming system for 6DOF multi-user virtual reality
abstract
We investigate a future WiFi-VLC dual connectivity streaming system for 6DOF multi-user virtual reality that enables reliable high-fidelity remote scene immersion. The system integrates an edge server that uses scalable 360° tiling to adaptively split the present 360° view of a VR user into a panoramic baseline content layer and a viewport-specific enhancement content layer. The user is then served the two content layers over complementary WiFi and VLC wireless links such that the delivered viewport quality is maximized for the given WiFi and VLC transmission resources. We formally characterize the actions of the server using rate-distortion optimization that we solve at low complexity. To account for the users' mobility as they explore different 360° viewpoints of the 6DOF remote scene content and maintain reliable high-quality VLC connectivity, we explore dynamic VLC transmitter steering and assignment in the system as graph bottleneck matching that aims to maximize the received VLC SNR across all users. We formulate an effective low-complexity solution to this discrete combinatorial optimization problem of high complexity. The paper also contributes a first actual 6DOF body and head movement VR navigation dataset that we collected and facilitate to assess the performance of our system via simulation experiments. These demonstrate enhanced VLC transmission performance and an up to 7 dB gain in viewport quality over a state-of-the-art VLC cellular system (LiFi), and an up to 10 dB gain in viewport quality over a state-of-the-art traditional wireless streaming method, for 12K-120fps 360° 6DOF VR content. Moreover, the synergistic WiFi-VLC dual connectivity of the proposed system augments its reliability over the reference method LiFi that comprises only VLC links. These outcomes motivate further exploration and prototype implementation of our system.
Jacob Chakareski, Mahmudur Khan 0002
NOSSDAV1
2021 Decentralized Collaborative Video Caching in 5G Small-Cell Base Station Cellular Networks
abstract
We consider the problem of video caching across a set of 5G small-cell base stations (SBS) connected to each other over a high-capacity short-delay back-haul link, and linked to a remote server over a long-delay connection. Even though the problem of minimizing the overall video delivery delay is NP-hard, the Collaborative Caching Algorithm (CCA) that we present can efficiently compute a solution close to the optimal, where the degree of sub-optimality depends on the worst case video-to-cache size ratio. The algorithm is naturally amenable to distributed implementation that requires no explicit coordination between the SBSs, and runs in O(N + K log K) time, where N is the number of SBSs (caches) and K the maximum number of videos. We extend CCA to an online setting where the video popularities are not known a priori but are estimated over time through a limited amount of periodic information sharing between the SBSs. We demonstrate that our algorithm closely approaches the optimal integral caching solution as the cache size increases. Moreover, via simulations carried out on real video access traces, we show that our algorithm effectively uses the SBS caches to reduce the video delivery delay and conserve the remote server’s bandwidth, and that it outperforms two other reference caching methods adapted to our system setting.
Shadab Mahboob, Koushik Kar, Jacob Chakareski
WiOpt3
2020 Multi-Connectivity and Edge Computing for Ultra-Low-Latency Lifelike Virtual Reality
abstract
We explore a novel multi-user mobile VR system for streaming scalable 8K360^° video at high reliability and immersion fidelity, and low interactive latency, via a synergistic integration of scalable 360^° tiling, dual-band millimeter wave (mmWave) and Wi-Fi transmission, and edge computing. High rate directed mmWave links are studied to send VR viewport-specific high-quality enhancement layers of the 360^° content to the individual users, while Wi-Fi broadcast of the base layer of the entire 360^° panorama is sent to all users, to augment the system's reliability. The viewport-specific enhancement layers can comprise compressed and raw 360^° tiles, decoded first at the edge server. We explore the joint optimization of the mmWave access point to user association, the choice of 360^° tiles to be transmitted decompressed, the allocation of mmWave data rate across the compressed tiles in a viewport-specific enhancement layer, and the allocation of computing resources at the edge server and user devices. Our objective is to maximize the minimum delivered VR immersion fidelity across all users, given transmission, latency, and computing constraints. We demonstrate that our framework can enable a significant improvement in immersion fidelity (8dB to 10 dB) and spatial resolution (8Kvs. 4K), over MPEG-DASH that uses Wi-Fi transmission only. We also show that an increasing number of raw 360^° tiles are sent, as the mmWave link rate or the edge server/user computing power increase, exploring rigorously here the fundamental interplay between computing and communication capabilities, end-to-end system latency, and delivered VR immersion fidelity.
Jacob Chakareski, Sabyasachi Gupta
ICME1
2020 Mobile-Edge Cooperative Multi-User 360° Video Computing and Streaming
abstract
We investigate a novel communications system that integrates scalable multi-layer 360° video tiling, viewport-adaptive rate-distortion optimal resource allocation, and VR-centric edge computing and caching, to enable future high-quality untethered VR streaming. Our system comprises a collection of 5G small cells that can pool their communication, computing, and storage resources to collectively deliver scalable 360° video content to mobile VR clients at much higher quality. Our major contributions are rigorous design of multi-layer 360° tiling and related models of statistical user navigation, and analysis and optimization of edge-based multi-user VR streaming that integrates viewport adaptation and server cooperation. We also explore the possibility of network coded data operation and its implications for the analysis, optimization, and system performance we pursue here. We demonstrate considerable gains in delivered immersion fidelity, featuring much higher 360° viewport peak signal to noise ratio (PSNR) and VR video frame rates and spatial resolutions.
Jacob Chakareski, Nicholas Mastronarde
MMSP1
2020 RF-FSO Dual-Path UAV Network for High Fidelity Multi-Viewpoint Scalable 360° Video Streaming
abstract
We explore a novel RF-FSO dual-path UAV net-work for remote scene aerial scalable 360° video capture and streaming, to enable future virtual human teleportation. One UAV captures the 360° video viewpoint and constructs a scalable tiling representation of the data comprising a base layer and an enhancement layer. The base layer is sent by the UAV to a ground-based remote server using a direct RF link. The enhancement layer is relayed by the UAV to the server over a multi-hop path comprising directed UAV to UAV FSO links. The viewport-specific content from the two layers is then integrated at the server to construct high fidelity content to stream to a remote VR user. The dual-path connectivity ensures both reliability and high fidelity remote immersion. We formulate an optimization problem to maximize the delivered immersion fidelity which depends on the content capture rate, FSO and RF link rates, effective routing path selection, and fast UAV deployment. The problem is mixed integer programming and we formulate an optimization framework that captures the optimal solution at lower complexity. Our experimental results demonstrate an up to 6 dB gain in delivered immersion fidelity over a state-of-the-art method and for the first time enable 12K-120fps 360° video streaming at high fidelity.
Mahmudur Khan 0002, Jacob Chakareski, Sabyasachi Gupta
MMSP2
2020 Deep Reinforcement Learning for Delay-Sensitive LTE Downlink Scheduling
abstract
We consider an LTE downlink scheduling system where a base station allocates resource blocks (RBs) to users running delay-sensitive applications. We aim to find a scheduling policy that minimizes the queuing delay experienced by the users. We formulate this problem as a Markov Decision Process (MDP) that integrates the channel quality indicator (CQI) of each user in each RB, and queue status of each user. To solve this complex problem involving high dimensional state and action spaces, we propose a Deep Reinforcement Learning based scheduling framework that utilizes the Deep Deterministic Policy Gradient (DDPG) algorithm to minimize the queuing delay experienced by the users. Our extensive experiments demonstrate that our approach outperforms state-of-the-art benchmarks in terms of average throughput, queuing delay, and fairness, achieving up to 55% lower queuing delay than the best benchmark.
Nikhilesh Sharma, Someshwar Rao Somayajula Venkata, Filippo Malandra, Nicholas Mastronarde, Jacob Chakareski
PIMRC6
2020 Delay-Sensitive Energy-Harvesting Wireless Sensors: Optimal Scheduling, Structural Properties, and Approximation Analysis
abstract
We consider an energy harvesting sensor transmitting latency-sensitive data over a fading channel. We aim to find the optimal transmission scheduling policy that minimizes the packet queuing delay given the available harvested energy. We formulate the problem as a Markov decision process (MDP) over a state-space spanned by the transmitter's buffer, battery, and channel states, and analyze the structural properties of the resulting optimal value function, which quantifies the long-run performance of the optimal scheduling policy. We show that the optimal value function (i) is non-decreasing and has increasing differences in the queue backlog; (ii) is non-increasing and has increasing differences in the battery state; and (iii) is submodular in the buffer and battery states. Taking advantage of these structural properties, we derive an approximate value iteration algorithm that provides a controllable tradeoff between approximation accuracy, computational complexity, and memory, and we prove that it converges to a near-optimal value function and policy. Our numerical results confirm these properties and demonstrate that the resulting scheduling policies outperform a greedy policy in terms of queuing delay, buffer overflows, energy efficiency, and sensor outages.
Nikhilesh Sharma, Nicholas Mastronarde, Jacob Chakareski
IEEE Trans. Commun.3
2020 Viewport-Adaptive Scalable Multi-User Virtual Reality Mobile-Edge Streaming
abstract
Virtual reality (VR) holds tremendous potential to advance our society, expected to make impact on quality of life, energy conservation, and the economy. To bring us closer to this vision, the present paper investigates a novel communications system that integrates for the first time scalable multi-layer 360° video tiling, viewport-adaptive rate-distortion optimal resource allocation, and VR-centric edge computing and caching, to enable next generation high-quality untethered VR streaming. Our system comprises a collection of 5G small cells that can pool their communication, computing, and storage resources to collectively deliver scalable 360° video content to mobile VR clients at much higher quality. The major contributions of the paper are the rigorous design of multi-layer 360° tiling and related models of statistical user navigation, analysis and optimization of edge-based multi-user VR streaming that integrates viewport adaptation and server cooperation, and base station 360° video packet scheduling. We also explore the possibility of network coded data operation and its implications for the analysis, optimization, and system performance we pursue in this setting. The advances introduced by our framework over the state-of-theart comprise considerable gains in delivered immersion fidelity, featuring much higher 360° viewport peak signal to noise ratio (PSNR) and VR video frame rates and spatial resolutions.
Jacob Chakareski
IEEE Trans. Image Process.1
2020 Collaborative Content Placement Among Wireless Edge Caching Stations With Time-to-Live Cache
abstract
Content caching at the Internet edge using a network of wireless edge caching stations (ECSs) is recently considered as a key solution to alleviating the backhaul traffic burden and improving the quality of experience in 5G networks. This paper studies wireless edge caching systems with the following features: first, content files can be partitioned into many coded packets, which then can be cached in multiple ECSs for collaborative content delivery; second, the service provider (SP) deploys time-to-live cache at ECSs and each cached content file has an occupancy time that needs to be guaranteed; third, the content-to-be-cached arrives at the caching system following a stochastic process as users request new content over time. Unlike existing works that determine which content to cache, this paper focuses on how to distribute the coded packets of content-to-be-cached among the network of ECSs in order to reduce the content downloading time. A novel content placement strategy, called stochastic collaborative content placement is proposed based on Lyapunov techniques. The proposed algorithm makes content placement decisions using only currently available information without foreseeing future content arrivals, takes advantage of the spatial content popularity variation with coded caching, and achieves the provable close-to-optimal long-term caching performance. Simulations are carried out on a real-world YouTube video request trace and the results demonstrate a tremendous caching performance improvement against a variety of benchmark schemes.
Lixing Chen, Linqi Song, Jacob Chakareski, Jie Xu 0001
IEEE Trans. Multim.3
2019 Geometric Programming for Lifetime Maximization in Mobile Edge Computing Networks
abstract
Mobile edge computing has emerged as a promising technology to augment the computational capabilities of mobile devices. For a multi-user network in which its users periodically compute their tasks with the help of an edge cloud, we investigate the network lifetime maximization problem based on present user task information. We pursue this objective via a minimum energy efficiency maximization (MEEM) strategy that jointly optimizes the fraction of user task computations offloaded to the cloud and the respective allocation of edge computing and network communication resources across the users. We also investigate the network lifetime maximization problem for the case when the user task information is available for all future time slots, as well. This setting represents an upper bound for the MEEM strategy. Optimal solutions for both investigated strategies are formulated via feasibility testing and geometric programming. We show that MEEM can achieve a 70% lifetime improvement over the state-of-the-art and 450% lifetime improvement over the case of local user task computation only.
Sabyasachi Gupta, Jacob Chakareski
GLOBECOM2
2019 Neighbor Discovery in a Free-Space-Optical UAV Network
abstract
Neighbor discovery is an essential part of the communication link establishment process for any wireless ad-hoc network. This problem of discovering neighbor nodes becomes even more challenging when the transceivers are highly directional. In this paper, we consider a 3D network of unmanned-aerial-vehicles (UAVs) that uses free-space-optical (FSO) transceivers for establishing high speed highly directional communication links. We consider that each UAV is equipped with a spherical structure on which multiple FSO transceivers are placed. The UAVs can electronically steer their communication beams by switching from one transceiver to another. We provide analysis on how optimally placing the transceivers with the appropriate divergence angles can help establish an FSO link at any direction in the 3D space. We also present a neighbor discovery algorithm that ensures discovery within a limited time. We demonstrate through extensive simulations that a UAV with FSO transceivers can successfully discover its neighbor UAVs even without prior location information about them and without any additional omnidirectional radio frequency (RF) channel.
Mahmudur Khan 0002, Jacob Chakareski
GLOBECOM2
2019 Millimeter Wave meets Edge Computing for Mobile VR with High-Fidelity 8K Scalable 360° Video
abstract
We investigate a novel multiple user scalable 8K 360° video mobile virtual reality arcade streaming system that enables high reliability and immersion fidelity, and low interactive latency, by the synergistic integration of scalable 360° content, expected VR user viewport modeling, millimeter wave (mmWave) communication and network edge computation capability. The high data rate mmWave link is used to transmit the video content of the expected user 360 viewport at enhanced quality. To compensate for the dynamic nature of mmWave links and prospective expected viewport characterization error, we integrate a fall back transmission based on Wi-Fi broadcast of a baseline representation of the 360 panorama to all users. In our proposed transmission strategy, the expected viewport content can be sent as raw or encoded at different qualities, which enhances the end-to-end performance, by exploiting effective trade-offs between communication and computation latency at the receiving user. With the aim of maximizing the minimum VR immersion fidelity across all users, we investigate the joint optimization of the mmWave access point (AP) to user association, the data rate for the encoded portion of the 360 viewport content that is to be transmitted, and computation resource allocation. Our experimental results demonstrate that the proposed system can achieve significant improvement in delivered VR user immersion fidelity and quality of experience relative to a state-of-the-art reference method that leverages Wi-Fi transmission only.
Sabyasachi Gupta, Jacob Chakareski, Petar Popovski
MMSP2
2019 UAV-IoT for Next Generation Virtual Reality
abstract
We investigate UAV-IoT data capture and networking for remote scene virtual reality (VR) immersion. We characterize the delivered immersion fidelity as a function of the assigned UAV-IoT capture/network rates and study the optimization problem of maximizing it, for given system/application constraints. We explore fast reinforcement learning to discover the best dynamic UAV-IoT network placement over the scene of interest to maximize the expected remote immersion fidelity. We design scalable source-channel viewpoint coding to maximize the expected reconstruction fidelity of the data captured at every UAV location at the ground-based aggregation point. Finally, we explore layered directional networking and rate-distortion-power optimized embedded scheduling methods to effectively transmit the encoded data and overcome network transients that lead to packet buffering, which represent the fourth system component of our framework. Experimental results demonstrate considerable performance efficiency gains enabled by each system component over the respective state-of-the-art reference methods, in delivered VR immersion fidelity, application interactivity/play-out latency, and transmission power consumption.
Jacob Chakareski
IEEE Trans. Image Process.1
2018 Viewport-Driven Rate-Distortion Optimized 360º Video Streaming
abstract
The growing popularity of virtual and augmented reality communications and 360° video streaming is moving video communication systems into much more dynamic and resource-limited operating settings. The enormous data volume of 360° videos requires an efficient use of network bandwidth to maintain the desired quality of experience for the end user. To this end, we propose a framework for viewport-driven rate-distortion optimized 360° video streaming that integrates the user view navigation pattern and the spatiotemporal rate-distortion characteristics of the 360° video content to maximize the delivered user quality of experience for the given network/system resources. The framework comprises a methodology for constructing dynamic heat maps that capture the likelihood of navigating different spatial segments of a 360° video over time by the user, an analysis and characterization of its spatiotemporal rate-distortion characteristics that leverage preprocessed spatial tilling of the 360° view sphere, and an optimization problem formulation that characterizes the delivered user quality of experience given the user navigation patterns, 360° video encoding decisions, and the available system/network resources. Our experimental results demonstrate the advantages of our framework over the conventional approach of streaming a monolithic uniformly-encoded 360° video and a state-of-the-art reference method. Considerable video quality gains of 4 - 5 dB are demonstrated in the case of two popular 4K 360° videos.
Jacob Chakareski, Ridvan Aksu, Xavier Corbillon, Gwendal Simon, Viswanathan (Vishy) Swaminathan
ICC1
2018 Energy Efficiency Analysis of UAV-Assisted mmWave HetNets
abstract
We study downlink transmission in a multi-band heterogeneous network comprising unmanned aerial vehicle (UAV) small base stations and ground-based dual mode mmWave small cells within the coverage area of a microwave (μW) macro base station. We formulate a two-layer optimization framework to simultaneously find efficient coverage radius for the UAVs and energy efficient radio resource management for the network, subject to minimum quality-of-service (QoS) and maximum transmission power constraints. The outer layer derives an optimal coverage radius/height for each UAV as a function of the maximum allowed path loss. The inner layer formulates an optimization problem to maximize the system energy efficiency (EE), defined as the ratio between the aggregate user data rate delivered by the system and its aggregate energy consumption (downlink transmission and circuit power). We demonstrate that at certain values of the target SINR τ introducing the UAV base stations doubles the EE. We also show that an increase in τ beyond an optimal EE point decreases the EE.
Syed Naqvi, Jacob Chakareski, Nicholas Mastronarde, Jie Xu 0001, Fatemeh Afghah, Abolfazl Razi
ICC2
2018 Structural Properties of Optimal Transmission Policies for Delay-Sensitive Energy Harvesting Wireless Sensors
abstract
We consider an energy harvesting sensor transmit- ting latency-sensitive data over a fading channel. We aim to find the optimal transmission scheduling policy that minimizes the packet queuing delay given the available harvested energy. We formulate the problem as a Markov decision process (MDP) over a state-space spanned by the transmitter's buffer, battery, and channel states, and analyze the structural properties of the resulting optimal value function, which quantifies the long-run performance of the optimal scheduling policy. We show that the optimal value function (i) is non- decreasing and has increasing differences in the queue backlog; (ii) is non-increasing and has increasing differences in the battery state; and (iii) is submodular in the buffer and battery states. Our numerical results confirm these properties and demonstrate that the optimal scheduling policy outperforms a so-called greedy policy in terms of sensor outages, buffer overflows, energy efficiency, and queuing delay.
Nikhilesh Sharma, Nicholas Mastronarde, Jacob Chakareski
ICC3
2018 Effective Deep Learning for Semantic Segmentation Based Bleeding Zone Detection in Capsule Endoscopy Images
abstract
Capsule endoscopy (CE) is a non-invasive way to detect small intestinal abnormalities such as bleeding. It provides a direct vision of the patients entire gastrointestinal (GI) tract. However, a manual inspection of the huge number of images produced thereby is tedious and lengthy, and thus prone to human errors. This makes automated computer assisted decision-making appealing in this context. This paper introduces a novel deep-learning based semantic segmentation approach for bleeding zone detection in CE images. A bleeding image features three regions labeled as bleeding, non-bleeding, and background. Thus, a convolutional neural network (CNN) is trained using SegNet layers with three classes. A given CE image is segmented using our training network and the detected bleeding zones are marked. The proposed network architecture is tested on different color planes and best performance is achieved using the hue saturation and value (HSV) color space. Experimental performance evaluation is carried out on a publicly available clinical dataset, on which our framework achieves 94.42 % global accuracy and 90.69 % weighted intersection over union (IoU), two state-of-the-art classification metrics. Performance gains are demonstrated over several recent state-of-art competiting methods in terms of all performance measures we examined, including mean accuracy, mean IoU, global accuracy and weighted IoU.
Tonmoy Ghosh, Jacob Chakareski
ICIP3
2017 Viewport-adaptive navigable 360-degree video delivery
abstract
The delivery and display of 360-degree videos on Head-Mounted Displays (HMDs) presents many technical challenges. 360-degree videos are ultra high resolution spherical videos, which contain an omnidirectional view of the scene. However only a portion of this scene is displayed on the HMD. Moreover, HMD need to respond in 10 ms to head movements, which prevents the server to send only the displayed video part based on client feedback. To reduce the bandwidth waste, while still providing an immersive experience, a viewport-adaptive 360-degree video streaming system is proposed. The server prepares multiple video representations, which differ not only by their bit-rate, but also by the qualities of different scene regions. The client chooses a representation for the next segment such that its bit-rate fits the available throughput and a full quality region matches its viewing. We investigate the impact of various spherical-to-plane projections and quality arrangements on the video quality displayed to the user, showing that the cube map layout offers the best quality for the given bit-rate budget. An evaluation with a dataset of users navigating 360-degree videos demonstrates that segments need to be short enough to enable frequent view switches.
Xavier Corbillon, Gwendal Simon, Alisa Devlic, Jacob Chakareski
ICC4
2017 Optimal Set of 360-Degree Videos for Viewport-Adaptive Streaming
abstract
With the decreasing price of Head-Mounted Displays (HMDs), 360-degree videos are becoming popular. The streaming of such videos through the Internet with state of the art streaming architectures requires, to provide high immersion feeling, much more bandwidth than the median user's access bandwidth. To decrease the need for bandwidth consumption while providing high immersion to users, scientists and specialists proposed to prepare and encode 360-degree videos into quality-variable video versions and to implement viewport-adaptive streaming. Quality-variable versions are different versions of the same video with non-uniformly spread quality: there exists some so-called Quality Emphasized Regions (QERs). With viewport-adaptive streaming the client, based on head movement prediction, downloads the video version with the high quality region closer to where the user will watch. In this paper we propose a generic theoretical model to find out the optimal set of quality-variable video versions based on traces of head positions of users watching a 360-degree video. We propose extensions to adapt the model to popular quality-variable version implementations such as tiling and offset projection. We then solve a simplified version of the model with two quality levels and restricted shapes for the QER. With this simplified model, we show that an optimal set of four quality-variable video versions prepared by a streaming server, together with a perfect head movement prediction, allow for 45% bandwidth savings to display video with the same average quality as state of the art solutions or allows an increase of 102% of the displayed quality for the same bandwidth budget.
Xavier Corbillon, Alisa Devlic, Gwendal Simon, Jacob Chakareski
ACM Multimedia4
2017 Convexity characterization of virtual view reconstruction error in multi-view imaging
abstract
Virtual view synthesis is a key component of multi-view imaging systems that enable visual immersion environments for emerging applications, e.g., virtual reality and 360-degree video. Using a small collection of captured reference view-points, this technique reconstructs any view of a remote scene of interest navigated by a user, to enhance the perceived immersion experience. We carry out a convexity characterization analysis of the virtual view reconstruction error that is caused by compression of the captured multi-view content. This error is expressed as a function of the virtual viewpoint coordinate relative to the captured reference viewpoints. We derive fundamental insights about the nature of this dependency and formulate a prediction framework that is able to accurately predict the specific dependency shape, convex or concave, for given reference views, multi-view content and compression settings. We are able to integrate our analysis into a proof-of-concept coding framework and demonstrate considerable benefits over a baseline approach.
Vladan Velisavljevic, Camilo C. Dorea, Jacob Chakareski, Ricardo L. de Queiroz
MMSP3
2017 Guest Editorial Special Issue on Visual Computing in the Cloud: Mobile Computing
abstract
Recent advances in mobile devices (e.g., smartphones and wearables) and wireless technologies are fueling a new wave of user demands for an improved user experience. Indeed, users are not only expecting ubiquitous network connections for traditional services (e.g., messaging and calling), but also demanding extensive access to a wealth of video contents and services. However, this growing demand is seriously hindered by the fact that the onboard resources with mobile devices are inherently limited and their growth rate falls behind that of their desktop counterparts. It follows that new solutions should be in order to resolve this fundamental tussle. Fortunately, the emerging cloud computing offers a natural solution to extend the desktop visual experience to mobile devices. It actually provides both computational and storage support for media-rich applications with both front-end and back-end functionalities.
Yonggang Wen 0001, Jacob Chakareski, Pascal Frossard, Di Wu 0001, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Reinforcement Learning for Energy-Efficient Delay-Sensitive CSMA/CA Scheduling
abstract
We study learning-based energy-efficient multi- user scheduling of delay-sensitive data over fading channels. To tradeoff energy and delay, we combine adaptive rate transmission at the physical layer with a rate-adaptive medium access control (MAC) protocol based on carrier sense multiple access with collision avoidance (CSMA/CA). We formulate the multi-user scheduling problem as a constrained Markov decision process (CMDP). We show that the multi-user problem is intractable and propose to decompose it into multiple (coupled) single-user problems. We design a reinforcement learning algorithm to solve the single-user problems online so that users can achieve energy-efficient operation while meeting their delay constraints, even though the channel, traffic, and multi-user dynamics are unknown a priori. Our proposed MAC protocol enables users to meet significantly tighter delay constraints while also consuming less energy than under the 802.11 Distributed Coordination Function (DCF). Moreover, the proposed learning algorithm converges significantly faster than a state-of-the-art solution.
Nicholas Mastronarde, SayedJalil Modares Najafabadi, Changcan Wu, Jacob Chakareski
GLOBECOM4
2016 Fast and low-complexity reinforcement learning for delay-sensitive energy harvesting wireless visual sensing systems
abstract
In this paper, we consider an energy challenged remote sensor transmitting latency-sensitive imagery data over a time-varying channel. The sensor harvests energy from the environment and hence efficient energy consumption is of great importance. In this paper, we aim to find the optimal transmission scheduling and power management policies that maximize the available energy for future transmissions while meeting a queuing delay constraint. We formulate this problem as a Markov Decision Process (MDP) and propose a reinforcement learning (RL) algorithm to solve it online. Our experiments show that the proposed algorithm achieves comparable performance to a state-of-the-art RL algorithm, but at much lower complexity.
Niloofar Toorchi, Jacob Chakareski, Nicholas Mastronarde
ICIP2
2016 Joint Caching, Routing, and Channel Assignment for Collaborative Small-Cell Cellular Networks
abstract
We consider joint caching, routing, and channel assignment for video delivery over coordinated small-cell cellular systems of the future Internet. We formulate the problem of maximizing the throughput of the system as a linear program, in which the number of variables is very large. To address channel interference, our formulation incorporates the conflict graph that arises when wireless links interfere with each other due to simultaneous transmission. We utilize the column generation method to solve the problem by breaking it into a restricted master subproblem that involves a select subset of variables and a collection of pricing subproblems that select the new variable to be introduced into the restricted master problem, if that leads to a better objective function value. To control the complexity of the column generation optimization further, due to the exponential number of independent sets that arise from the conflict graph, we introduce an approximation algorithm that computes a solution that is within ϵ to optimality, at much lower complexity. Our framework demonstrates considerable gains in average transmission rate at which the video data can be delivered to the users, over the state-of-the-art Femtocaching system, of up to 46%. These operational gains in system performance map to analogous gains in video application quality, thereby enhancing the user experience considerably.
Abdallah Khreishah, Jacob Chakareski, Ammar Gharaibeh
IEEE J. Sel. Areas Commun.2
2015 Virtual Machine Placement via Q-Learning with Function Approximation
abstract
While existing virtual machine technologies provide easy-to-use platforms for distributed computing applications, many are far from efficient and not designed to accommodate diverse objectives, which dramatically penalizes their performance. These shortcomings arise from 1) not having a formal optimization framework that readily leads to algorithmic solutions for diverse objectives; 2) not incorporating the knowledge of the underlying network topologies and the communication/interaction patterns among the virtual machines/services, and 3) not considering the time-varying aspects of real-world environments. This paper formalizes an optimization framework and develops corresponding algorithmic solutions using Markov Decision Process and Q-Learning for virtual machine/service placement and migration for distributed computing in time-varying environments. Importantly, the knowledge of the underlying topologies of the computing infrastructure, the interaction patterns between the virtual machines, and the dynamics of the supported applications will be formally characterized and incorporated into the proposed algorithms in order to improve performance results for small-scale and large-scale networks are provided to verify our solution approach.
Thai Duong 0002, Yu-Jung Chu, Thinh P. Nguyen, Jacob Chakareski
GLOBECOM4
2015 Fine-Grained Scalable Video Caching
abstract
Caching has been shown to enhance network performance. In this paper, we study fine-grain scalable video caching. We start from a single cache scenario by providing a solution to the caching allocation problem that optimizes the average expected video quality for the most popular video clips. Actual trace data is applied to verify the performance of our algorithm and compare its backhaul link bandwidth consumption relative to non-scalable video caching. In addition, we extend our analysis to collaborative caching and integrate network coding for further transmission efficiency. Our experimental results demonstrate considerable performance enhancement.
Qiushi Gong, John W. Woods, Koushik Kar, Jacob Chakareski
ISM4
2015 Dynamic dual-reinforcement-learning routing strategies for quality of experience-aware wireless mesh networking
Nuno Filipe Coutinho, Ricardo Matos, Carlos Marques 0001, Andre B. Reis, Susana Sargento, Jacob Chakareski, Andreas Kassler
Comput. Networks6
2015 Cost and profit driven cloud-P2P interaction
Jacob Chakareski
Peer-to-Peer Netw. Appl.1
2015 Guest Editorial: Special Issue on P2P Cloud Systems
Shueng-Han Gary Chan, Mea Wang, Jacob Chakareski, Bin Wei 0003
Peer-to-Peer Netw. Appl.3
2015 Uplink Scheduling of Visual Sensors: When View Popularity Matters
abstract
Decentralized camera sensors capture a 3D scene of interest from multiple perspectives. The captured video signals need to be transmitted to a central station over a shared wireless channel. A collocated server streams the gathered data to a collection of clients interested in experiencing the scene interactively. We design a constrained optimization framework for sharing the transmission bandwidth of the wireless channel across the sensors such that the average video quality over the client population is maximized. We consider scheduling the uplink resources of the wireless channel at the view or packet level. In the first case, the central station partitions the channel capacity across a select subset of views that are transmitted entirely. In the second case, the station coordinates the packet transmissions of every sensor. We formulate exact and approximate algorithms to solve the optimization of interest. We examine their transmission efficiency via simulation experiments that demonstrate considerable gains over the state-of-the-art and study the impact of view popularity.
Jacob Chakareski
IEEE Trans. Commun.1
2015 Retention in Online Blogging: A Case Study of the Blogster Community
abstract
Community-based blogging platforms can be rich sources of information on a variety of specialized topics, from finance to parenting. The usefulness of such platforms depends heavily on user participation and contribution. However, one potential problem is lower retention: users' fail to contribute in the long run. This paper is an investigation of retention in the popular community blogging platform “Blogster.” We use the points users earn for their activities as a proxy of retention and explore the attributes that are associated with their retention. We find that highly retained users are most central in the network and their blogger friends are mutually less connected. They get more views and comments to their posted blogs. We also examine the homophily of retention in the social network. Based on our empirical observations, we build a classifier that is able to detect top retained users with accuracy as high as 94%. Our work has theoretical implications for the social behavior literature of community bloggers and practical design implications for potential community blogging platform developers.
Md. Imrul Kayes, Jacob Chakareski
IEEE Trans. Comput. Soc. Syst.2
2015 Joint Source-Channel Rate Allocation and Client Clustering for Scalable Multistream IPTV
abstract
We design a system framework for streaming scalable internet protocol television (IPTV) content to heterogenous clients. The backbone bandwidth is optimally allocated between source and parity data layers that are delivered to the client population. The assignment of stream layers to clients is done based on their access link data rate and packet loss characteristics, and is part of the optimization. We design three techniques for jointly computing the optimal number of multicast sessions, their respective source and parity rates, and client membership, either exactly or approximatively, at lower complexity. The latter is achieved via an iterative coordinate descent algorithm that only marginally underperforms relative to the exact analytic solution. Through experiments, we study the advantages of our framework over common IPTV systems that deliver the same source and parity streams to every client. We observe substantial gains in video quality in terms of both its average value and standard deviation over the client population. In addition, for energy efficiency, we propose to move the parity data generation part to the edge of the backbone network, where each client connects to its IPTV stream. We analytically study the conditions under which such an approach delivers energy savings relative to the conventional case of source and parity data generation at the IPTV streaming server. Finally, we demonstrate that our system enables more consistent streaming performance, when the clients' access link packet loss distribution is varied, relative to the two baseline methods used in our investigation, and maintains the same performance as an ideal system that serves each client independently.
Jacob Chakareski
IEEE Trans. Image Process.1
2015 A Poisson Hidden Markov Model for Multiview Video Traffic
abstract
Multiview video has recently emerged as a means to improve user experience in novel multimedia services. We propose a new stochastic model to characterize the traffic generated by a Multiview Video Coding (MVC) variable bit-rate source. To this aim, we resort to a Poisson hidden Markov model (P-HMM), in which the first (hidden) layer represents the evolution of the video activity and the second layer represents the frame sizes of the multiple encoded views. We propose a method for estimating the model parameters in long MVC sequences. We then present extensive numerical simulations assessing the model's ability to produce traffic with realistic characteristics for a general class of MVC sequences. We then extend our framework to network applications where we show that our model is able to accurately describe the sender and receiver buffers behavior in MVC transmission. Finally, we derive a model of user behavior for interactive view selection, which, in conjunction with our traffic model, is able to accurately predict actual network load in interactive multiview services.
Lorenzo Rossi 0002, Jacob Chakareski, Pascal Frossard, Stefania Colonnese
IEEE/ACM Trans. Netw.2
2014 Joint source and channel coding of view and rate scalable multi-view video
abstract
We study multicast of multi-view content in the video plus depth format to heterogeneous clients. We design a joint source-channel coding scheme based on view and rate embedded source coding and rateless channel coding. It comprises an optimization framework for joint view selection and source-channel rate allocation, and includes a fast method for separate optimization of the source and channel coding components, at a negligible performance loss wrt the joint solution. We demonstrate performance gains over a state-of-the-art method based on H.264/SVC, in the case of two client classes.
Jacob Chakareski, Vladan Velisavljevic, Vladimir Stankovic 0001
ICIP1
2014 Know thy neighbor: Community-aware recovery of content selection preferences
Jacob Chakareski
Signal Process.1
2014 Vertex selection via a multi-graph analysis
Jacob Chakareski
Signal Process.1
2014 Social video caching
Jacob Chakareski
Signal Process. Image Commun.1
2014 Wireless Streaming of Interactive Multi-View Video via Network Compression and Path Diversity
abstract
We formulate a system framework for network compression of interactive multi-view streaming video. The setup comprises a media server that delivers the content over two independent network paths to a client. Our system features a proxy-server located at the junction of the wired and wireless portions of each path. The proxy dynamically adapts the content data sent over the wireless links, in response to channel quality feedback from the client, such that video distortion at the client is minimized. We analyze the performance of our system and contrast its characteristics with dynamic content adaptation at the source and conventional streaming architectures, including scalable video. We numerically simulate the operation of all streaming systems under comparison and establish a close agreement between our analysis and the experimental findings. The proposed system delivers superior video quality over the reference competitors, while enabling notable transmission rate savings, at the same time.
Jacob Chakareski
IEEE Trans. Commun.1
2014 Transmission Policy Selection for Multi-View Content Delivery Over Bandwidth Constrained Channels
abstract
I formulate an optimization framework for computing the transmission actions of streaming multi-view video content over bandwidth constrained channels. The optimization finds the schedule for sending the packetized data that maximizes the reconstruction quality of the content, for the given network bandwidth. Two prospective multi-view content representation formats are considered: 1) MVC and 2) video plus depth. In the case of each, I formulate directed graph models that characterize the interdependencies between the data units that comprise the content. For the video plus depth format, I develop a novel space-time error concealment strategy that reconstructs the missing content based on received data units from multiple views. I design multiple techniques to solve the optimization problem of interest, at varying degrees of complexity and accuracy. In conjunction, I derive spatiotemporal models of the reconstruction error for the multi-view content that I employ to reduce the computational requirements of the optimization. I study the performance of my framework via simulation experiments. Significant gains in terms of rate-distortion efficiency are demonstrated over various reference methods.
Jacob Chakareski
IEEE Trans. Image Process.1
2013 Wireless streaming of interactive multi-view video: Network compression meets path diversity
abstract
I formulate a system framework for network compression of interactive multi-view streaming video. The setup comprises a media server that delivers the content over two independent network paths to a client. My system features a proxy-server located at the junction of the wired and wireless portions of each path. The proxy dynamically adapts the content data sent over the wireless links, in response to channel quality feedback from the client, such that video distortion at the client is minimized. I analyze the performance of my system and contrast its characteristics with dynamic content adaptation at the source and conventional streaming architectures, including scalable video. I numerically simulate the operation of all streaming systems under comparison and establish a close agreement between my analysis and the experimental findings. The proposed system delivers superior video quality over the reference competitors, while enabling notable transmission rate savings, at the same time.
Jacob Chakareski
GLOBECOM1
2013 Viewpoint-popularity-driven uplink scheduling of multi-camera sensor arrays
abstract
A decentralized camera array records a 3D scene of interest simultaneously, from different perspectives. The captured video signals need to be transmitted over a shared wireless channel to a central station. The gathered data is then streamed to a collection of clients interested in experiencing the scene interactively. I design an optimization framework that allows for sharing the transmission resources of the wireless medium such that the average video quality over the client population is maximized. I consider scheduling the uplink resources either at the view or packet level. That is, in the first case, the central station partitions the channel capacity across a select subset of views that are transmitted in their entirety. In the second case, the station coordinates the packet transmissions of every sensor such that the aggregate data rate over the wireless channel does not exceed its capacity. I formulate algorithms that solve the optimization either exactly or approximatively, at lower complexity. I examine their transmission efficiency via simulation experiments that show a considerable improvement over reference methods. I also study the impact of viewpoint popularity, as governed by the clients that interact with the scene, on the operation of the optimization.
Jacob Chakareski
GLOBECOM1
2013 View selection policy for multi-view video delivery
abstract
We derive an optimization framework for computing a view selection policy for streaming multi-view content over a bandwidth constrained channel. The optimization allows us to determine the decisions of sending the packetized data such that the end-to-end reconstruction quality of the content is maximized, for the given bandwidth resources. Two prospective multi-view content representation formats are considered: MVC and video plus depth. For each, we formulate directed graph models that characterize the interdependencies between the data units comprising the content. For the video plus depth format, we develop a spatial error concealment strategy that reconstructs missing content at the client based on received data from other views. We design multiple techniques to solve the optimization problem of interest either exactly or approximatively, at lower complexity. In conjunction, we derive a spatial model of the reconstruction error for the multi-view content that we employ to reduce the computational requirements of the optimization. We study the performance of our framework via simulation experiments. Significant gains in terms of rate-distortion efficiency are observed over a content-agnostic reference technique.
Jacob Chakareski
ICASSP1
2013 Scheduling space-time dependent packets in multi-view video streaming
abstract
I formulate an optimization framework for scheduling the packet transmissions in streaming multi-view content over bandwidth constrained channels. The optimization allows me to compute the transmission decisions that maximize the expected reconstruction quality of the content at the client, for the given bandwidth resources. Two prospective multi-view content representation formats are considered: MVC and video plus depth. In the case of each, I formulate directed graph models that characterize the interdependencies between the data units comprising the content. For the video plus depth format, I develop a novel space-time error concealment strategy that reconstructs missing content at the client based on received data units from multiple views. I design two techniques for solving the optimization problem of interest. In conjunction, I derive a spatiotemporal model of the reconstruction error for the multi-view content that I employ to reduce the computational requirements of the optimization. I study the performance of my framework via simulation experiments. Significant gains in rate-distortion efficiency are observed over state-of-the-art methods.
Jacob Chakareski
MMSP1
2013 Spotify Me: Facebook-assisted automatic playlist generation
abstract
We design a novel method for automatically generating a playlist of recommended songs in the popular social music sharing application Spotify that are liked with high probability by a user. Our method employs multiple seed artists as an input that are obtained via the Facebook likes of artists and the listening history of songs of a Spotify user. First, we construct an input vector comprising all the artists that the user likes on Facebook and listens to in Spotify. Then, we search for other artists and bands related to them using EchoNest, an online state-of-the-art machine learning platform. We assign a score to every artist in the thereby obtained collection, based on the frequency of his/her appearance. Finally, we construct a playlist comprising randomly selected popular songs associated with the most frequently cited artists. We examine the recommendation performance of our algorithm by computing its WTF score (fraction of disliked songs) and novelty factor (fraction of new liked songs) on playlists generated for different seed input sizes. We observe that our approach substantially outperforms the built-in Spotify Radio recommender. On 30 song playlists, we are able to improve the WTF score by 49% and the novelty factor by 42%, on average. Due to its general design, our method is broadly applicable to a variety of personal content management scenarios.
Arthur Germain, Jacob Chakareski
MMSP2
2013 Informative State-Based Video Communication
abstract
We study state-based video communication where a client simultaneously informs the server about the presence status of various packets in its buffer. In sender-driven transmission, the client periodically sends to the server a single acknowledgement packet that provides information about all packets that have arrived at the client by the time the acknowledgment is sent. In receiver-driven streaming, the client periodically sends to the server a single request packet that comprises a transmission schedule for sending missing data to the client over a horizon of time. We develop a comprehensive optimization framework that enables computing packet transmission decisions that maximize the end-to-end video quality for the given bandwidth resources, in both prospective scenarios. The core step of the optimization comprises computing the probability that a single packet will be communicated in error as a function of the expected transmission redundancy (or cost) used to communicate the packet. Through comprehensive simulation experiments, we carefully examine the performance advances that our framework enables relative to state-of-the-art scheduling systems that employ regular acknowledgement or request packets. Consistent gains in video quality of up to 2B are demonstrated across a variety of content types. We show that there is a direct analogy between the error-cost efficiency of streaming a single packet and the overall rate-distortion performance of streaming the whole content. In the case of sender-driven transmission, we develop an effective modeling approach that accurately characterizes the end-to-end performance as a function of the packet loss rate on the backward channel and the source encoding characteristics.
Jacob Chakareski
IEEE Trans. Image Process.1
2013 User-Action-Driven View and Rate Scalable Multiview Video Coding
abstract
We derive an optimization framework for joint view and rate scalable coding of multi-view video content represented in the texture plus depth format. The optimization enables the sender to select the subset of coded views and their encoding rates such that the aggregate distortion over a continuum of synthesized views is minimized. We construct the view and rate embedded bitstream such that it delivers optimal performance simultaneously over a discrete set of transmission rates. In conjunction, we develop a user interaction model that characterizes the view selection actions of the client as a Markov chain over a discrete state-space. We exploit the model within the context of our optimization to compute user-action-driven coding strategies that aim at enhancing the client's performance in terms of latency and video quality. Our optimization outperforms the state-of-the-art H.264 SVC codec as well as a multi-view wavelet-based coder equipped with a uniform rate allocation strategy, across all scenarios studied in our experiments. Equally important, we can achieve an arbitrarily fine granularity of encoding bit rates, while providing a novel functionality of view embedded encoding, unlike the other encoding methods that we examined. Finally, we observe that the interactivity-aware coding delivers superior performance over conventional allocation techniques that do not anticipate the client's view selection actions in their operation.
Jacob Chakareski, Vladan Velisavljevic, Vladimir Stankovic 0001
IEEE Trans. Image Process.1
2012 Multiple description coding of free viewpoint video for multi-path network streaming
abstract
By transmitting texture and depth videos from two adjacent captured viewpoints, a client can synthesize via depth-image-based rendering (DIBR) any intermediate virtual view of the scene, determined by the dynamic movement of the client's head. In so doing, depth perception of the 3D scene will be created through motion parallax. Due to the stringent playback deadline of interactive free viewpoint video, burst packet losses in the texture and depth video streams caused by transmission over unreliable channels are difficult to overcome and can severely degrade the synthesized view quality at the client. We propose a multiple description coding (MDC) of free viewpoint video in texture-plus-depth format that will be transmitted on two disjoint network paths. Specifically, we encode even frames of the left view and odd frames of the right view separately as one description and transmit it on path one. Similarly, we encode odd frames of the left view and even frames of the right view as the second description and transmit it on path two. Appropriate quantization parameters (QP) are selected for each description, such that its data rate matches optimally the available transmission bandwidth on each of the two paths. If the receiver receives one description but not the other due to burst loss on one of the paths, it can still partially reconstruct the missing frames in the loss-corrupted description using a computationally efficient DIBR-based recovery scheme that we design. Extensive experimental results show that our MDC streaming system can outperform the traditional single-path single-description transmission scheme by up to 7dB in Peak Signal-to-Noise Ratio (PSNR) of the synthesized intermediate view at the receiving client.
Zhi Liu 0002, Gene Cheung, Jacob Chakareski, Yusheng Ji
GLOBECOM3
2012 Multi-graph sampling of online communities via mean hitting time
abstract
We derive a framework for sampling online communities based on the mean hitting time of its members, considering that there are multiple graphs associated with the same vertex set V representing the social network. First, we formulate random walk models on the multi-graph ensemble and define the essential properties of the mean hitting times associated with the corresponding Markov chains on the vertex set V . Then, we design a branch and bound optimization technique for computing the subset of vertices A that exhibits the shortest mean hitting time across the multi-graph, given a constraint on the size of A. We also design a greedy optimization method that computes an approximation to the optimal subset, at lower complexity, and that lends itself to a decentralized implementation, for further complexity reduction. We examine the performance of the sampling framework through a series of simulation experiments involving synthetic and actual samples of online community graphs. We demonstrate substantial improvements in terms of sampling (network) cost reduction and information dissemination speed relative to the state-of-the-art methods of node degree and eigenvector centrality.
Jacob Chakareski
ICASSP1
2012 Quality of experience-based routing in multi-service wireless mesh networks
abstract
We develop an optimization framework for Quality of Experience (QoE)-based routing in multi-service Wireless Mesh Networks (WMNs). The framework takes into account the heterogeneous requirements of different services delivered over a WMN, such that the overall end-user QoE is maximized under given resource constraints. We propose a novel QoE-aware double reinforcement learning strategy for dynamically computing the most efficient routes to deliver the flows of each service type. Comprehensive NS-2-based simulations demonstrate the substantial performance gains that our approach enables over conventional routing techniques such as AODV, with significant improvement over video quality.
Ricardo Matos, Nuno Filipe Coutinho, Carlos Marques 0001, Susana Sargento, Jacob Chakareski, Andreas Kassler
ICC5
2012 Multi-path content delivery: Efficiency analysis and optimization algorithms
Jacob Chakareski
J. Vis. Commun. Image Represent.1
2011 Multi-graph regularization for efficient delivery of user generated content in online social networks
abstract
We present a methodology for enhancing the delivery of user-generated content in online social networks. To this end, we first regularize the social graph via node capacity and link cost information associated with the underlying data network. We then design a technique for constructing the most efficient delivery tree over the regularized social graph. Finally, we derive an optimization algorithm for allocating the nodes' uplink capacities over the content distribution tree. Our system substantially outperforms the conventional method of flooding data over the social graph, over multiple criteria. In particular, a 100% reduction in terms of network cost and data delivery delay is registered.
Jacob Chakareski
ICASSP1
2011 Content preference estimation in online social networks: Message passing versus sparse reconstruction on graphs
abstract
We design two different strategies for computing the unknown content preferences in an online social network based on a small set of nodes in the corresponding social graph for which this information is available ahead of time. The techniques take advantage of the graph's structure and the additional affinity information between the social contacts, expressed through the graph's edge weights, to optimize the computation of the missing preference data. The first strategy is distributed and comprises a local computation step and a message passing step that are iteratively applied at each node in the graph, until convergence. We carry out a graph Laplacian based analysis of the performance of the algorithm and verify the analytical findings via numerical experiments involving sample social networks. The second strategy is centralized and involves a sparse transform of the content preference data represented as a function over the nodes of the social graph. We solve the related optimization problem of reconstructing the unknown preferences via an iterative algorithm based on variable splitting and alternating direction of multipliers. The algorithm takes into account the specifics of the data to be reconstructed by incorporating multiple regularization terms into the optimization. We investigate the underpinnings of the sparse reconstruction technique via numerical experiments that reveal its characteristics and how they affect its performance.
Jacob Chakareski
ICASSP1
2011 Browsing catalogue graphs: Content caching supercharged!!
abstract
We consider a generic scenario of content browsing where a client is presented with a catalogue of items of interest. Upon the selection of an item from a page of the catalogue, the client can choose the next item to browse from a list of related items presented on the same page. The system has limited resources to have all items available for immediate access by the browsing client. Therefore, it pays a penalty when the client selects an unavailable item. Conversely, there is a reward that the system gains when the client selects an immediately available item. We formulate the optimization problem of selecting the subset of items that the system should have for immediate access such that its profit is maximized, for the given system resources. We design two techniques for solving the optimization problem in linear time, as a function of the catalogue size. We examine their performance via numerical simulations that reveal their core properties. We also study their operation on actual YouTube data and compare their efficiency relative to conventional solutions. Substantial performance gains are demonstrated, due to accounting for the content graph imposed by the catalogue of items.
Jacob Chakareski
ICIP1
2011 Ultrasound-based surgical navigation for percutaneous renal intervention: In vivo measurements and in vitro assessment
abstract
This paper evaluates the feasibility of a proposed ultrasound-based surgical navigation system for percutaneous renal intervention via in vivo measurements and in vitro assessment. The system integrates preoperative computer tomography (CT) planning with intraoperative ultrasonography (US) by means of a proposed semi-automatic US to CT rigid registration. The interventional procedure is performed with a visualized guidance interface. The navigation system is evaluated at two levels. Level I evaluation comprises measurements of the accuracy, precision, and processing time of our registration method on in vivo data provided by volunteers. For Level II, expert urologists are asked to rate the perceptual quality of the system via in vitro tests on a kidney phantom. Both objective and subjective evaluations validate the proposed surgical navigation system.
Zhicheng Li 0001, Jacob Chakareski, Lei Wang 0029
ICIP3
2011 Bit allocation for multiview image compression using cubic synthesized view distortion model
abstract
“Texture-plus-depth” has become a popular coding format for multiview image compression, where a decoder can synthesize images at intermediate viewpoints using encoded texture and depth maps of closest captured view locations via depth-image-based rendering (DIBR). As in other resource-constrained scenarios, limited available bits must be optimally distributed among captured texture and depth maps to minimize the expected signal distortion at the decoder. A specific challenge of multiview image compression for DIBR is that the encoder must allocate bits without the knowledge of how many and which specific virtual views will be synthesized at the decoder for viewing. In this paper, we derive a cubic synthesized view distortion model to describe the visual quality of an interpolated view as a function of the view's location. Given the model, one can easily find the virtual view location between two coded views where the maximum synthesized distortion occurs. Using a multiview image codec based on shape-adaptive wavelet transform, we show how optimal bit allocation can be performed to minimize the maximum view synthesis distortion at any intermediate viewpoint. Our experimental results show that the optimal bit allocation can outperform a common uniform bit allocation scheme by up to 1.0dB in coding efficiency performance, while simultaneously being competitive to a state-of-the-art H.264 codec.
Vladan Velisavljevic, Gene Cheung, Jacob Chakareski
ICME3
2011 In-Network Packet Scheduling and Rate Allocation: A Content Delivery Perspective
abstract
We investigate two important problems in media delivery via active network agents. First, we consider streaming multiple video assets over a shared backbone network through an intermediate proxy-server to a set of receiving clients. The proxy is located at the junction of the backbone network and the last hop to each of the clients and coordinates the delivery of the videos from the origin media server to the clients. We propose an optimization framework that enables the proxy to coordinate the streaming process such that the overall end-to-end performance of the video streams is maximized for the given data rate resources on the backbone and the last hop links. Prospective video quality requirements for the associated media sessions are also taken into consideration in the analysis. Through experiments, we study in detail the operation of the framework and the influence of the various constraints that it considers. Furthermore, we measure its performance gains relative to a sender-driven system where the media server controls the delivery of the data with no assistance from an intervening proxy. We establish an analytical relationship between the relative improvement of the proxy-based system, the network conditions on the backbone and the last hops, and the number of streams served. The gains of the proxy-driven system measured in our experiments closely match their expected values predicted by this relationship. In conjunction with the above scenario, we explore the performance gains due to multi-agent packet scheduling where there are multiple active nodes organizing the packet transmissions along the network path between a server-client pair. To this end, we design an optimization framework that coordinates the multiple scheduling agents such that an end-to-end quality-rate performance metric is maximized. We study experimentally the performance benefits due to multi-agent scheduling relative to single-agent scheduling and conventional server-driven streaming, as a function of the number of intermediate nodes at which the packet scheduling is carried out. We quantify analytically the performance gains and match them with high accuracy to the simulation data.
Jacob Chakareski
IEEE Trans. Multim.1
2011 Prioritized Distributed Video Delivery With Randomized Network Coding
abstract
We address the problem of prioritized video streaming over lossy overlay networks. We propose to exploit network path diversity via a novel randomized network coding (RNC) approach that provides unequal error protection (UEP) to the packets conveying the video content. We design a distributed receiver-driven streaming solution, where a client requests packets from the different priority classes from its neighbors in the overlay. Based on the received requests, a node in turn forwards combinations of the selected packets to the requesting peers. Choosing a network coding strategy at every node can be cast as an optimization problem that determines the rate allocation between the different packet classes such that the average distortion at the requesting peer is minimized. As the optimization problem has log-concavity properties, it can be solved with low complexity by an iterative algorithm. Our simulation results demonstrate that the proposed scheme respects the relative priorities of the different packet classes and achieves a graceful quality adaptation to network resource constraints. Therefore, our scheme substantially outperforms reference schemes such as baseline network coding techniques as well as solutions that employ rateless codes with built-in UEP properties. The performance evaluation provides additional evidence of the substantial robustness of the proposed scheme in a variety of transmission scenarios.
Nikolaos Thomos, Jacob Chakareski, Pascal Frossard
IEEE Trans. Multim.2
2010 Multiple Description Coding with Feedback Based Network Compression
abstract
This paper concerns multi path video streaming using adaptive multiple description coding. The adaptation leverages on the fact that multiple descriptions are correlated. Thus if an intermediate node gets feedback telling that another path is likely to deliver a description, this node can compress its description and forward it. Such a compression can also be done already at the source node; however, the feedback arrives more timely and reliably to intermediate nodes that are closer to the final receiver. In this paper we investigate the performance of such adaptation at the source node and an intermediate node, respectively. A trade-off exists between reducing the delay of the feedback by adapting in the vicinity of the receiver and increasing the gain from compression by adapting close to the source. The analysis shows that adaptation in the network provides a better trade-off than adaptation at the source. Schemes which provide simple solutions to adaptation both at the source and in the network are proposed, analyzed, simulated and compared to non-adaptive reference schemes in scenarios that involve last hop that is wireless. The results reveal that the proposed compression schemes offer significant benefits in streaming scenarios.
Jesper H. Sørensen, Jan Østergaard, Petar Popovski, Jacob Chakareski
GLOBECOM4
2010 Client Clustering and Joint Multistream FEC Rate Allocation in IPTV Systems
abstract
This paper addresses the problem of clustering heterogeneous clients in IPTV services over lossy networks. The delivery of the same stream to clients with different capabilities or access networks is surely suboptimal in terms of average quality for the population of receivers. Instead, we propose that the streaming servers deliver distinct multicast streams to different subsets of clients. We formulate an optimization problem where the receivers are clustered depending on the quality of their connection so that the average video quality in the IPTV system is maximized. Then we propose a novel algorithm for determining optimally the clusters, as well as the source and channel rate allocation in each of the clusters. Simulation results show that the proposed algorithm is able to maximize the average quality in the system when each of the servers transmits information to a distinct cluster. In particular, we show that the proposed solution outperforms baseline schemes that serve all clients with the same multicast stream, as it is commonly the case in practical systems.
Jacob Chakareski, Pascal Frossard
ICC1
2010 Multi-stream partitioning and parity rate allocation for scalable IPTV delivery
abstract
We address the joint problem of clustering heterogenous clients and allocating scalable video source rate and FEC redundancy in IPTV systems. We propose a streaming solution that delivers varying portions of the scalably encoded content to different client subsets, together with suitably selected parity data. We formulate an optimization problem where the receivers are clustered depending on the quality of their connection so that the average video quality in the IPTV system is maximized. Then we propose a novel algorithm for determining optimally the client clusters, the source and parity rate allocation to each cluster, and the set of serving rates at which the source+parity data is delivered to the clients. We implement our system through a novel design based on scalable video coding that allows for much more efficient network utilization relative to the case of source versioning. Through simulations we demonstrate that the proposed solution substantially outperforms baseline IPTV schemes that multicast the same source and FEC streams to the whole client population, as is commonly done in practice today.
Jacob Chakareski, Pascal Frossard
ICIP1
2010 Quality of experience optimized scheduling in multi-service wireless mesh networks
abstract
A growing trend has emerged in network architecture research to switch focus from Quality of Service (QoS) to Quality of Experience (QoE) optimization. In this paper, we first present QoE models that characterize user satisfaction of video, audio, and data services over wireless networks. We then develop a novel packet scheduling algorithm for multi-hop wireless networks that jointly optimizes the delivery of multiple video, audio, and data flows according to the QoE metrics. We formulate a multidimensional optimization problem that minimizes the overall distortion across all flows for the given network resources on wireless links. Fairness constraints over the flows are also considered as part of the optimization. Our experimental results, obtained with the NS-2 IEEE 802.16 MESH-mode simulator, show that distortion-aware scheduling can significantly increase the perceived quality of different wireless services under bandwidth constraints. Additionally, improved fairness across the competing flows is demonstrated relative to conventional scheduling techniques.
Andre B. Reis, Jacob Chakareski, Andreas Kassler, Susana Sargento
ICIP2
2010 A non-stationary Hidden Markov Model of multiview video traffic
abstract
Multiview video is increasingly getting attention due to emerging applications such as 3DTV and immersive teleconferencing. In this paper, we present a non-stationary Hidden Markov Model (HMM) for characterizing the data rate of compressed multiview content. The states of the model correspond to different video activity levels and exhibit a Poisson state duration distribution. We derive a stable maximum likelihood algorithm for estimating the parameters of our multiview traffic model. Synthetic data generated by the model exhibits statistics that closely match those of actual multiview data. In addition, we demonstrate the high accuracy of the model in two multiview streaming applications by evaluating the frame loss rate of a constrained network buffer fed by actual and synthetic data.
Lorenzo Rossi 0002, Jacob Chakareski, Pascal Frossard, Stefania Colonnese
ICIP2
2010 Optimal rate allocation for view synthesis along a continuous viewpoint location in multiview imaging
abstract
We consider the scenario of view synthesis via depth-image based rendering in multi-view imaging. We formulate a resource allocation problem of jointly assigning an optimal number of bits to compressed texture and depth images such that the maximum distortion of a synthesized view over a continuum of viewpoints between two encoded reference views is minimized, for a given bit budget. We construct simple yet accurate image models that characterize the pixel values at similar depths as first-order Gaussian auto-regressive processes. Based on our models, we derive an optimization procedure that numerically solves the formulated min-max problem using Lagrange relaxation. Through simulations we show that, for two captured views scenario, our optimization provides a significant gain (up to 2dB) in quality of the synthesized views for the same overall bit rate over a heuristic quantization that selects only two quantizers — one for the encoded texture images and the other for the depth images.
Vladan Velisavljevic, Gene Cheung, Jacob Chakareski
PCS3
2010 Popularity-aware rate allocation in multiview video
abstract
We propose a framework for popularity-driven rate allocation in H.264/MVC-based multi-view video communications when the overall rate and the rate necessary for decoding each view are constrained in the delivery architecture. We formulate a rate allocation optimization problem that takes into account the popularity of each view among the client population and the rate-distortion characteristics of the multi-view sequence so that the performance of the system is maximized in terms of popularity-weighted average quality. We consider the cases where the global bit budget or the decoding rate of each view is constrained. We devise a simple ratevideo- quality model that accounts for the characteristics of interview prediction schemes typical of multi-view video. The video quality model is used for solving the rate allocation problem with the help of an interior point optimization method. We then show through experiments that the proposed rate allocation scheme clearly outperforms baseline solutions in terms of popularity-weighted video quality. In particular, we demonstrate that the joint knowledge of the rate-distortion characteristics of the video content, its coding dependencies, and the popularity factor of each view is key in achieving good coding performance in multi-view video systems.
Attilio Fiandrotti, Jacob Chakareski, Pascal Frossard
VCIP2
2009 Delay-based overlay construction in P2P video broadcast
abstract
We consider streaming video content over an overlay network of peer nodes. Each of the nodes employs a mesh-pull mechanism to organize the download of data units from its neighbours. We propose a novel algorithm for constructing the distribution overlay, where peers are arranged in neighbourhoods that exhibit similar latency values from the origin media server. Such an organization increases data sharing between neighbours in broadcast applications and reduces the play-out latency at a peer. Each of the nodes in the overlay is further equipped with a packet scheduling procedure that requests data units from neighbours in the order of their importance and their popularity within the neighbourhood. Finally, requesting peers share the upload bandwidth of a sending peer in proportion to their transmission rate to that peer in order to discourage free-riding in the system. Our simulation results show that the proposed mesh construction procedure provides improved performance in terms of frame-freeze and playback latency relative to a conventional approach where peer neighbours are selected at random. Corresponding gains in video quality for the media presentation are also registered due to the improved continuity of the playback experience.
Jacob Chakareski, Pascal Frossard
ICASSP1
2009 Randomized Network Coding for UEP video delivery in overlay networks
abstract
This paper presents a receiver-driven video delivery algorithm that exploits a novel Randomized Network Coding (RNC) scheme for unequal error protection (UEP). The main idea of our approach is to account for the unequal importance of media packets in the network coding algorithm for efficient stream delivery in lossy overlay networks. Based on the requests from their neighbours, the network nodes properly combine packets and forward them to their children nodes. The network coding operations at every node are formulated as a log-concave optimization problem, which is solved with a greedy algorithm in only a few iterations. Our experimental results demonstrate that the proposed scheme permits to respect the priorities between the different packet classes. It further outperforms baseline network coding techniques for video streaming in overlay networks.
Nikolaos Thomos, Jacob Chakareski, Pascal Frossard
ICME2
2009 Efficient proxy-driven multiple user video streaming
abstract
We consider streaming multiple video assets over a shared backbone network through an intermediate proxy-server to a set of receiving clients. The proxy is located at the junction of the backbone network and the last hop to each of the clients and coordinates the delivery of the videos from the origin media server to the clients. We propose an optimization framework that enables the proxy to coordinate the streaming process such that the overall end-to-end performance of the video streams is maximized for the given data rate resources on the backbone and the last hop links. Prospective video quality requirements for the associated media sessions are also taken into consideration in the analysis. We study in detail the operation of the framework and the influence of the various constraints that it considers through experiments. Furthermore, we measure its performance gains relative to a sender-driven system where the media server controls the delivery of the data with no assistance from an intervening proxy. We establish an analytical relationship between the relative improvement of the proxy-based system, the network conditions on the backbone and the last hops, and the number of streams served. The gains of the proxy-driven system measured in our experiments match closely their expected values predicted by this relationship.
Jacob Chakareski
MMSP1
2009 Modeling of distortion caused by Markov-model burst packet losses in video transmission
abstract
This paper addresses the problem of distortion modeling for video transmission over burst-loss channels characterized by a finite state Markov chain. A distortion trellis model is proposed, enabling us to estimate at the frame level the expected mean-square error (MSE) distortion caused by Markov-model bursty packet losses. A sliding window algorithm is developed to perform the MSE estimation with low complexity. Simulation results show that the proposed models are accurate for all tested average loss rates and average burst lengths. Based on the experimental results, the proposed techniques are used to analyze the impact of average burst length on the average decoded video quality. The proposed model is further extended to a more general form, and the modeled distortion is compared with simulation data. These experiments demonstrate that the extended model is also accurate for all tested loss rates.
Zhicheng Li 0001, Jacob Chakareski, Xiaodun Niu, Yongjun Zhang 0008, Wanyi Gu
MMSP2
2009 Telemedicine applications of mobile ultrasound
abstract
This work evaluates the feasibility of real-time wireless video streaming of medical ultrasound signals over 802.11g ad-hoc and 3G cellular broadband networks for disaster relief, rescue, and medical transport applications. Scalable H.264 encoded echocardiographic ultrasound image sequences were transmitted in real-time at specified image resolutions (VGA and QVGA) and frame rates (10, 15, 20, and 30 frames per second). Relevant transmission parameters such as data rate, packet loss, delay jitter, and latency were measured. The 802.11g network permits high frame rate, VGA resolution, has low latency and jitter, but is suitable only for short communication ranges, while the 3G cellular broadband network allows medium to low frame rate streaming at QVGA image resolution with medium latency. However, video streaming can take place from any location with 3G service to any other site with Internet connectivity. Physicians with expertise in medical ultrasound subsequently rated the diagnostic content of the transmitted videos relative to the original ultrasound video signals. The physicians observed that the image quality in the case of both 802.11g and 3G wireless transmission was fully to adequately preserved, while missed frames could momentarily decrease the diagnostic value. This research demonstrates that sufficient wireless bandwidth and efficient video compression make diagnostically valuable wireless streaming of ultrasound video feasible.
Peder C. Pedersen, Brett W. Dickson, Jacob Chakareski
MMSP3
2009 Modeling and Analysis of Distortion Caused by Markov-Model Burst Packet Losses in Video Transmission
abstract
This paper addresses the problem of distortion modeling for video transmission over burst-loss channels characterized by a finite-state Markov chain. Based on a detailed analysis of the error propagation and the bursty losses, a distortion trellis model is proposed, enabling us to estimate at the both the frame level and sequence level the expected mean-square error (MSE) distortion caused by Markov-model burst packet losses. The model takes into account the temporal dependencies induced by both the motion-compensated coding scheme and the Markov-model channel losses. The model is applicable to most block-based motion-compensated encoders, and most Markov-model lossy channels as long as the loss pattern probabilities for that channel is computable. Based on the study of the decaying behavior of the error propagation, a sliding window algorithm is developed to perform the MSE estimation with low complexity. Simulation results show that the proposed models are accurate for all tested average loss rates and average burst lengths. Based on the experimental results, the proposed techniques are used to analyze the impact of factors such as average burst length on the average decoded video quality. The proposed model is further extended to a more general form, and the modeled distortion is compared with the data produced from realistic networks loss traces. The experiment results demonstrate that the proposed model is also accurate in estimating the expected distortion for video transmission in real networks.
Zhicheng Li 0001, Jacob Chakareski, Xiaodun Niu, Yongjun Zhang 0008, Wanyi Gu
IEEE Trans. Circuits Syst. Video Technol.2
2008 Distributed Optimization of Media Flows in Peer-to-Peer Overlay Networks
abstract
We consider the problem of rate-distortion (RD) optimized media streaming in unstructured peer-to-peer (P2P) overlay networks. We formulate the aforementioned problem as a distributed rate allocation problem, and we solve it by applying classical decomposition techniques so that the network-wide utility of the media distortion is minimized. Information exchange between the peers is employed to ensure updates on the price of the locally calculated rate allocation. Media packets are also piggybacked with RD preambles that contain information regarding their impact on the decoder distortion and their size. The benefit of the aforementioned approach is that peers can convert the calculated optimal rate allocation into simple forwarding or dropping actions allowing thus a lightweight implementation. Our simulation results indicate that significant quality benefits can be achieved when the precise RD characteristics of a media description are taken into account by the streaming algorithm.
Antonios Argyriou, Jacob Chakareski
GLOBECOM2
2008 Cooperative media streaming using adaptive network compression
abstract
Media content distribution constitutes a growing share of the services on the Internet. Two distinct distribution approaches used today are layered coding (LC) and multiple description coding (MDC). Current wireless connection technologies, e.g. Wimax, have properties which make them unsuitable for media distribution using traditional approaches. In particular, the asymmetric relationship between the uplink and the downlink bandwidth makes the cooperative distribution difficult. A promising concept, termed MDC with Conditional Compression (MDC-CC), has been proposed [11], which essentially acts as an adaptive hybrid between LC and MDC. In order to facilitate the use of MDC-CC, a new overlay network approach is proposed, using tree of meshes. A control system for managing description distribution and compression in a small mesh is implemented in the discrete event simulator NS-2. The two traditional approaches, MDC and LC, are used as references for the performance evaluation of the proposed scheme. The system is simulated in a heterogeneous network environment, where packet errors are introduced. Moreover, a test is performed at different network loads. Performance gain is shown over both LC and MDC.
Janus Heide, Jesper H. Sørensen, Rasmus Krigslund, Petar Popovski, Torben Larsen, Jacob Chakareski
WOWMOM6
2008 Distributed Collaboration for Enhanced Sender-Driven Video Streaming
abstract
We propose a sender-driven system for adaptive streaming from multiple servers to a single receiver over separate network paths. The servers employ information in receiver feedbacks to estimate the available bandwidth on the paths and then compute appropriate transmission schedules for streaming media packets to the receiver based on the bandwidth estimates. An optimization framework is proposed that enables the senders to compute their transmission schedules in a distributed way, and yet to dynamically coordinate them over time such that the resulting video quality at the receiver is maximized. To reduce the computational complexity of the optimization framework an alternative technique based on packet classification is proposed. The substantial reduction in online complexity due to the resulting packet partitioning makes the technique suitable for practical implementations of adaptive and efficient distributed streaming systems. Simulations with Internet network traces demonstrate that the proposed solution adapts effectively to bandwidth variations and packet loss. They show that the proposed streaming framework provides superior performance over a conventional distortion-agnostic scheme that performs proportional packet scheduling on the network paths according to their respective bandwidth values.
Jacob Chakareski, Pascal Frossard
IEEE Trans. Multim.1
2007 Adaptive P2P video streaming via packet labeling
abstract
We consider the scenario of video streaming in peer-to-peer networks. A single media server delivers the video content to a large number of peer hosts by taking advantage of their forwarding capabilities. We propose a scheme that enables the peers to efficiently distribute the media stream among them. Each of the peers connects to the streaming server via multiple multicast trees that provide for robustness in the event of peer disconnection. Moreover, adaptive forwarding of the media content at each peer is enabled by labeling the packets with their importance for the reconstruction of the media stream. We study the performance of the proposed scheme as a function of system parameters such as the play-out delay of the media application, the peer population size and the number of multicast trees employed by the scheme. We show that by placing priorities on forwarding the individual packets at each peer an improved performance is achieved over conventional peer-to-peer systems where no such prioritization is deployed. The gains in performance are particularly significant for low-delay applications and large peer populations.
Jacob Chakareski, Pascal Frossard
VCIP1
2007 The Virtue of Patience in Low-Complexity Scheduling of Packetized Media With Feedback
abstract
We consider streaming pre-encoded and packetized media over best-effort networks in the presence of acknowledgment feedbacks. We first review a rate-distortion (RD) optimization framework that can be employed in such scenarios. As part of the framework, a scheduling algorithm selects the data to send over the network at any given time, so as to minimize the end-to-end distortion, given an estimate of channel resources and a history of previous transmissions and received acknowledgements. In practice, a greedy scheduling strategy is often considered to limit the solution search space, and reduce the computational complexity associated to the RD optimization framework. Our work observes that popular greedy schedulers are strongly penalized by early retransmissions. Therefore, we propose a scheduling algorithm that avoids premature retransmissions, while preserving the low computational complexity aspect of the greedy paradigm. Such a scheduling strategy maintains close to optimal RD performance when adapting to network bandwidth fluctuations. Our experimental results demonstrate that the proposed patient greedy scheduler provides a reduction of up to 50% in transmission rate relative to conventional greedy approaches, and that it brings up to 2 dB of quality improvement in scheduling classical MPEG-based packet video streams.
Christophe De Vleeschouwer, Jacob Chakareski, Pascal Frossard
IEEE Trans. Multim.2
2006 Distributed Streaming via Packet Partitioning
abstract
We propose a system for adaptive streaming from multiple servers to a single receiver over separate network paths. Based on incoming packets, the receiver estimates the available bandwidth on every path and returns this information to the servers. An optimization algorithm is designed that enables the servers to independently partition the media packets among them according to the bandwidth information and such that the resulting video quality at the receiver is maximized. To this end, the algorithm takes advantage of a source pruning technique that preprocesses the media stream ahead of time. Simulation results demonstrate that the proposed streaming framework provides superior performance over a conventional transmission scheme that performs proportional packet scheduling based only on the available network bandwidth. Due to its low-complexity aspect, the framework is suitable for practical implementations of adaptive and efficient distributed streaming systems
Jacob Chakareski, Pascal Frossard
ICME1
2006 Streaming of Scalable Video from Multiple Servers using Rateless Codes
abstract
This paper presents a framework for efficiently streaming scalable video from multiple servers over heterogeneous network paths. We propose to use rateless codes, or Fountain codes, such that each server acts as an independent source, without the need to coordinate its sending strategy with other servers. In this case, the problem of maximizing the received video quality and minimizing the bandwidth usage, is simply reduced to a rate allocation problem. We provide an optimal solution for an ideal scenario where the loss probability on each server-client path is exactly known. We then present a heuristic-based algorithm, which implements an unequal error protection scheme for the more realistic case of imperfect knowledge of the loss probabilities. Simulation results finally demonstrate the efficiency of the proposed algorithm, in distributed streaming scenarios over lossy channels
Jean-Paul Wagner, Jacob Chakareski, Pascal Frossard
ICME2
2006 Rate-distortion optimized distributed packet scheduling of multiple video streams over shared communication resources
abstract
We consider the problem of distributed packet selection and scheduling for multiple video streams sharing a communication channel. An optimization framework is proposed, which enables the multiple senders to coordinate their packet transmission schedules, such that the average quality over all video clients is maximized. The framework relies on rate-distortion information that is used to characterize a video packet. This information consists of two quantities: the size of the packet in bits, and its importance for the reconstruction quality of the corresponding stream. A distributed streaming strategy then allows for trading off rate and distortion, not only within a single video stream, but also across different streams. Each of the senders allocates to its own video packets a share of the available bandwidth on the channel in proportion to their importance. We evaluate the performance of the distributed packet scheduling algorithm for two canonical problems in streaming media, namely adaptation to available bandwidth and adaptation to packet loss through prioritized packet retransmissions. Simulation results demonstrate that, for the difficult case of scheduling nonscalably encoded video streams, our framework is very efficient in terms of video quality, both over all streams jointly and also over the individual videos. Compared to a conventional streaming system that does not consider the relative importance of the video packets, the gains in performance range up to 6 dB for the scenario of bandwidth adaptation, and even up to 10 dB for the scenario of random packet loss adaptation.
Jacob Chakareski, Pascal Frossard
IEEE Trans. Multim.1
2006 RaDiO edge: rate-distortion optimized proxy-driven streaming from the network edge
Jacob Chakareski, Philip A. Chou
IEEE/ACM Trans. Netw.1
2005 Rate-distortion optimized video streaming over Internet packet traces
abstract
In this paper, we study the performance of rate-distortion optimized video streaming over traces of packet delays and packet losses collected in the Internet. The study provides us with an understanding how important rate-distortion optimization may be for Internet streaming today and when it actually pays off to perform it. We propose a simple technique for channel estimation that can be incorporated within rate-distortion optimized streaming to remove the assumption of known channel statistics. In the experiments, performance of rate-distortion optimized streaming is compared to that of simpler transmission techniques such as ARQ.
Jacob Chakareski, Bernd Girod
ICIP (2)1
2005 Rate-Distortion Optimized Bandwidth Adaptation for Distributed Media Delivery
abstract
We propose a framework for rate-distortion optimized bandwidth adaptation via packet dropping at a network node, when the incoming traffic at the node consists of multiple video streams. The framework enables the node to decide in a rate-distortion optimal sense, which packets, if any, from each stream should be discarded in order to adapt to the available outgoing bandwidth at the node, so that the overall video quality over all streams is maximized. The framework relies on rate-distortion hint track information that is sent along with each video packet. The hint track information consists of two quantities: the size of the video packet in bits, and its importance for the reconstruction quality of the video stream. Experimental results demonstrate that our framework provides significant gains in video quality, both over all streams jointly and also over the individual videos, relative to a conventional system for bandwidth adaptation that does not take into account the different importance of the individual video packets
Jacob Chakareski, Pascal Frossard
ICME1
2005 Rate-Distortion Optimized Packet Scheduling Over Bottleneck Links
abstract
The loss and delay experienced by packets traveling along an Internet network path are mainly governed by the characteristics of a bottleneck link, such as available data rate and queue size. In this work, we propose a framework for rate-distortion optimized packet scheduling with adaptive rate control for media streaming over bandwidth-constrained bottleneck links. The framework computes optimal packet schedules while continuously adapting its instantaneous rate to the following three factors: the available data rate and the current queue size on the bottleneck link, and the congestion that packets transmitted under the schedules will create on the bottleneck link. Experimental results demonstrate that our framework does not lose in rate-distortion performance over rate-distortion optimized packet scheduling without strict rate control, while producing at the same time a much smoother instantaneous rate feeding the bottleneck queue. This in turn contributes to fairness to other flows sharing the bottleneck link and causes less variation in queue size, thereby avoiding queue overflow and unnecessarily long packet delays on the bottleneck link
Jacob Chakareski, Pascal Frossard
ICME1
2005 Examining Memory in Reconstruction Distortion: Dropping Additional Packets to Improve Video Quality
abstract
The source coding process and the packet loss process create certain dependencies between encoded video units in terms of the reconstruction distortion of the video signal at the receiver in case of transmission over packet erasure channels. In this paper, we examine the importance of this "distortion memory" via a specific class of memory-based models denoted Distortion Chains that are used for predicting the distortion of the reconstructed video signal in case of missing multiple packets at the receiver. We show that taking into account even the smallest amount of memory that is possible can yield substantial gains in terms of prediction accuracy and packet selection (packet dropping) performance. An additional and rather surprising result of our study is the fact that in certain situations dropping an additional video packet (which could otherwise be delivered) can actually improve the quality of the reconstructed video
Jacob Chakareski, John G. Apostolopoulos
MMSP1
2005 Distributed Packet Scheduling of Multiple Video Streams over Shared Communication Resources
abstract
We consider the problem of distributed packet selection and scheduling for multiple video streams sharing a communication channel. An optimization framework is proposed to enable the multiple senders to coordinate their packet transmission schedules, such that the overall quality over the video clients is maximized. The framework relies on rate-distortion information that is used to characterize a video packet and that consists of two quantities: the size of the packet in bits, and its importance for the reconstruction quality of the corresponding stream. Using the framework, each of the senders allocates to its own video packets a share of the bandwidth available on the communication channel, that is proportional to the relative importance of these packets. Thereby, a decentralized streaming strategy is provided that allows for trading-off rate and distortion, not only within a single video stream, but also across different streams. Simulation results demonstrate that, for the difficult case of scheduling non-scalably encoded video streams, our framework substantially outperforms a conventional streaming system that does not consider the relative importance of the video packets. The gains in performance reach up to 8 dB in both streaming scenarios under examination, namely adaptation to random packet loss and simultaneous adaptation to packet loss and available bandwidth
Jacob Chakareski, Pascal Frossard
MMSP1
2005 Low-Complexity Adaptive Streaming via Optimized A Priori Media Pruning
abstract
Source pruning is performed whenever the data rate of the compressed source exceeds the available communication or storage resources. In this paper, we propose a framework for rate-distortion optimized pruning of a video source. The framework selects which packets, if any, from the compressed representation of the source should be discarded so that the data rate of the pruned source is adjusted accordingly, while the resulting reconstruction distortion is minimized. The framework relies on a rate-distortion preamble that is created at compression time for the video source and that comprises the video packets' sizes, interdependencies and distortion importance. As one application of the pruning framework, we design a low-complexity rate-distortion optimized ARQ scheme for video streaming. In the experiments, we examine the performance of the pruning framework depending on the employed distortion model that describes the effect of packet interdependencies on the reconstruction quality. In addition, our experimental results show that the enhanced ARQ technique provides a significant performance gain over a conventional system for video streaming that does not take into account the different importance of the individual video packets. These gains are achieved without an increase in packet scheduling complexity, which makes the proposed technique suitable for online R-D optimized streaming
Jacob Chakareski, Pascal Frossard
MMSP1
2005 Layered coding vs. multiple descriptions for video streaming over multiple paths
Jacob Chakareski, Sangeun Han, Bernd Girod
Multim. Syst.1
2005 Rate-distortion hint tracks for adaptive video streaming
abstract
We present a technique for low-complexity rate-distortion (R-D) optimized adaptive video streaming based on the concept of rate-distortion hint track (RDHT). RDHTs store the precomputed characteristics of a compressed media source that are crucial for high performance online streaming but difficult to compute in real time. This enables low-complexity adaptation to variations in transport conditions such as available data rate or packet loss. An RDHT-based streaming system has three components: 1) information that summarizes the R-D attributes of the media; 2) an algorithm for using the RDHT to predict the distortion for a feasible packet schedule; and 3) a method for determining the best packet schedule to adapt the streaming to the communication channel. A family of distortion models, denoted distortion chains, are presented which accurately predict the distortion produced by arbitrary packet loss patterns. Two distortion chain models are examined which lead to two RDHT-based techniques. We evaluate the proposed techniques for two canonical problems in streaming media, adaptation to available data rate and to packet loss. Experimental results demonstrate that for the difficult case of nonscalably coded H.264 video, the proposed systems provide significant performance gains over conventional low-complexity streaming systems, and achieve this gain with a comparable level of complexity making them suitable for online R-D optimized streaming.
Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod
IEEE Trans. Circuits Syst. Video Technol.1
2004 Computing Rate-Distortion Optimized Policies for Streaming Media with Rich Acknowledgments
abstract
We consider the problem of rate-distortion optimized streaming of packetized media over the Internet from a server to a client using rich acknowledgments. Instead of separately acknowledging each media packet as it arrives, the client periodically sends to the server a single acknowledgment packet, denoted rich acknowledgment, that contains information about all media packets that have arrived at the client by the time the rich acknowledgment is sent. Computing the optimal transmission policy for the server involves estimation of the probability that a single packet will be communicated in error as a function of the expected redundancy (or cost) used to communicate the packet. In this paper, we show how to compute this error-cost function, and thereby optimize the server's transmission policy, in this scenario.
Jacob Chakareski, Bernd Girod
Data Compression Conference1
2004 Distortion chains for predicting the video distortion for general packet loss patterns
abstract
When designing a system for video communication over a lossy packet network, it is highly beneficial to have a mechanism for accurately predicting the mean-squared error (MSE) distortion that results from different packet loss patterns. The paper proposes a distortion chains model for accurately predicting the end-to-end distortion for different general packet loss patterns. The performance is examined using JVT/H.264 encoded video sequences and previous frame error concealment. It is shown that, for all tested sequences, the proposed model predicts the total distortion due to a packet loss pattern within a 10% error bound 80% of the time, as compared to the conventional additive approach which achieves the same accuracy less then 40% of the time.
Jacob Chakareski, John G. Apostolopoulos, Wai-tian Tan, Susie J. Wee, Bernd Girod
ICASSP (5)1
2004 Low-complexity rate-distortion optimized video streaming
Jacob Chakareski, John G. Apostolopoulos, Bernd Girod
ICIP1
2004 R-D hint tracks for low-complexity R-D optimized video streaming
abstract
This work presents the concept of rate-distortion hint track (RDHT), and evaluates two specific implementations of streaming systems that employ RDHT. Using RDHT, low-complexity streaming can be realized for systems that adapt to variations in transport conditions such as bandwidth or packet loss. An RDHT-based streaming system has three components: (1) an R-D hint track; (2) an algorithm for using the RDHT to predict the distortion for different packet schedules; and (3) a method for determining the best packet schedule. Two RDHT-based systems are presented which perform R-D optimized scheduling with dramatically reduced complexity as compared to conventional on-line R-D optimized streaming algorithms. Experimental results demonstrate that for the difficult case of R-D optimized scheduling of non-scalably coded video the proposed systems provide 7-12 dB gain when adapting to a bandwidth constraint and 2-4 dB gain when adapting to random packet loss, both relative to a conventional streaming system that does not take into account the different importance of individual packets.
Jacob Chakareski, John G. Apostolopoulos, Susie J. Wee, Wai-tian Tan, Bernd Girod
ICME1
2004 Rate-distortion optimized video streaming with rich acknowledgments
abstract
We consider an unconventional procedure for communicating to the server the receipt of media packets for Internet video streaming. Instead of separately acknowledging each media packet as it arrives, we periodically send to the server a single acknowledgment packet, denoted rich acknowledgment, that contains information about all media packets that have arrived at the client by the time the rich acknowledgment is sent. We investigate rate-distortion optimized sender-driven streaming that employs rich acknowledgments. Performance gains of up to 1.3 dB for streaming packetized video content are observed over rate-distortion optimized sender-driven systems that employ conventional acknowledgments.
Jacob Chakareski, Bernd Girod
VCIP1
2004 Application layer error-correction coding for rate-distortion optimized streaming to wireless clients
abstract
This paper addresses the problem of streaming packetized media over a lossy packet network to a wireless client, in a rate-distortion optimized way. We introduce an incremental redundancy error-correction scheme that combats the effects of both packet loss and bit errors in an end-to-end fashion, without support from the underlying network or from an intermediate base station. The scheme is employed within an optimization framework that enables the sender to compute which packets it should send, out of all the packets it could send at a given transmission opportunity, in order to meet an average transmission-rate constraint while minimizing the average end-to-end distortion. Experimental results show that our system is robust and maintains quality of service over a wide range of channel conditions. Up to 8 dB performance gains are registered over systems that are not rate-distortion optimized, at bit-error rates as large as 10/sup -2/.
Jacob Chakareski, Philip A. Chou
IEEE Trans. Commun.1
2003 Rate-distortion Optimized Packet Scheduling and Routing for Media Streaming with Path Diversity
abstract
The diversity for media streaming in a rate-distortion optimization framework was considered. A sender-driven transmission scenario was also investigated. Diversity was achieved by using multiple transmission paths over the network. The proposed framework enables the sender to decide at every instant which packets, if any, to transmit and over which transmission paths in order to meet a rate constraint while minimizing the end-to-end distortion. Experimental results demonstrate the benefit of exploiting packet diversity in rate-distortion optimized sender-driven streaming of packetized media.
Jacob Chakareski, Bernd Girod
DCC1
2003 Server diversity in rate-distortion optimized media streaming
abstract
We consider diversity for media streaming in a receiver-driven rate-distortion optimization framework. Diversity is achieved by requesting media packets from multiple servers. A framework is proposed that enables the receiver to decide at every instant which packets, if any, to request for transmission and from which servers in order to meet a rate constraint while minimizing the end-to-end distortion. Experimental results demonstrate the benefit of exploiting server diversity in rate-distortion optimized receiver-driven streaming of packetized media.
Jacob Chakareski, Bernd Girod
ICIP (3)1
2003 Video streaming with diversity
abstract
Packet path diversity is one of the recent advances in network-adaptive video streaming. In this paper we first examine two specific techniques for video streaming that exploit diversity to achieve improved performance. The first technique uses a framework for rate-distortion optimized scheduling of the packet transmissions over the available network paths. The second technique exploits feedback and channel probing to determine which network path should be used for transmission and to adapt the source encoding of the video to mitigate error propagation effects. In the final section, we compare the performance of these two techniques by analyzing experimental results and discuss their respective advantages and drawbacks.
Jacob Chakareski, Eric Setton, Yi J. Liang, Bernd Girod
ICME1
2003 Layered coding vs. multiple descriptions for video streaming over multiple paths
abstract
In this paper, we examine the performance of specific implementations of multiple description coding and of layered coding for video streaming over error-prone packet switched networks. We compare their performance using different transmission schemes with and without network path diversity. It is shown that given the specific implementations there is a large variation in relative performance between multiple description coding and layered coding depending on the employed transmission scheme. For scenarios where the packet transmission schedules can be optimized in a rate-distortion sense, layered coding provides a better performance. The converse is true for scenarios where the packet schedules are not rate-distortion optimized.
Jacob Chakareski, Sangeun Han, Bernd Girod
ACM Multimedia1
2002 Computing Rate-Distortion Optimized Policies for Streaming Media to Wireless Clients
abstract
We consider the problem of streaming packetized media over the Internet from a server through a base station to a wireless client, in a rate-distortion optimized way. For error control, we employ an incremental redundancy scheme, in which the server can incrementally transmit parity packets in response to negative acknowledgements fed back from the client. Computing the optimal transmission policy for the server involves estimation of the probability that a single packet will be communicated in error as a function of the expected redundancy (or cost) used to communicate the packet. In this paper, we show how to compute this error-cost function, and thereby optimize the server's transmission policy, in this scenario.
Jacob Chakareski, Behnaam Aazhang, Philip A. Chou
DCC1
2002 Application layer error correction coding for rate-distortion optimized streaming to wireless clients
abstract
This paper addresses the problem of streaming packetized media over a lossy packet network to a wireless client, in a rate-distortion optimized way. We introduce an incremental redundancy scheme that combats the effects of both packet loss and bit errors in an end-to-end fashion, without support from the underlying network or from an intermediate base station. The scheme is combined with an optimization framework that enables the sender to compute which packets it should send, out of all the packets it could send at a given transmission opportunity, in order to meet an average rate constraint while minimizing the average end-to-end distortion. Experimental results show that our system is robust and maintains a quality of service over a wide range of channel conditions. Up to 8 dB performance gains are registered over systems that are not rate-distortion optimized, at bit error rates as large as 10−2.
Jacob Chakareski, Philip A. Chou
ICASSP1