VLDB 2026 Research / reviewers in the wild / expert
Xinggong Zhang
dblp:45/702
· DBLP profile ↗
77ranked-venue papers
6as first author
36since 2021 · last 2026
0000-0003-0484-5951ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 13 since 2021Computer networks · 29 · 2 first-author · 20 since 2021Systems, architecture and hardware · 10 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Promptus: Can Prompt Streaming Replace Video StreamingabstractWith the exponential growth of video traffic, traditional video streaming systems are approaching their limits in communication capacity. To further reduce bitrate while maintaining quality, we propose Promptus, a disruptive semantic communication system that streams prompts instead of videos. Promptus represents the real-world video with a series of "prompts" for delivery and employs Stable Diffusion to generate the same video at the receiver. To ensure that the generated video is pixel-aligned with the original video, a gradient descent-based prompt fitting framework is proposed. Further, a low-rank decomposition-based bitrate control algorithm is introduced to achieve adaptive bitrate. For inter-frame compression, an interpolation-aware fitting algorithm is proposed. Evaluations across various video genres demonstrate that, compared to H.265, Promptus can achieve more than a 4x bandwidth reduction while preserving the same perceptual quality. On the other hand, at extremely low bitrates, Promptus can enhance the perceptual quality by 0.139 and 0.118 (in LPIPS) compared to VAE and H.265, respectively, and decreases the ratio of severely distorted frames by 89.3% and 91.7%. Our work opens up a new paradigm for efficient video communication. Jiangkai Wu, Yunpeng Tan, Junlin Hao, Xinggong Zhang |
AAAI | 6 |
| 2026 | Smaller is Better: Generative Models Can Power Short Video Preloading
Jiangkai Wu, Xinggong Zhang |
ICC | 3 |
| 2026 | MalMoE: Mixture-of-Experts Enhanced Encrypted Malicious Traffic Detection Under Graph Drift
Yunpeng Tan, Qingyang Li 0010, Mingxin Yang, Yannan Hu, Lei Zhang 0157, Xinggong Zhang |
INFOCOM | 6 |
| 2026 | CBDDoSCLIP: A Lightweight Multimodal Framework for Carpet-Bombing DDoS Attack Detection
Haotian Meng, Zongyuan Zhang, Tianyang Duan, Xinggong Zhang |
IWQoS | 6 |
| 2026 | R2-Mesh: Reinforcement Learning Powered Mesh Reconstruction via Geometry and Appearance Refinement
Haoyang Wang 0015, Xinggong Zhang |
MMM (2) | 3 |
| 2026 | NetRadar: Enabling Robust Carpet Bombing DDoS Detection
Junchen Pan, Lei Zhang 0157, Xiaoyong Si, Xinggong Zhang, Yong Cui 0001 |
NDSS | 5 |
| 2026 | HybridPrompt: Bridging Generative Priors and Traditional Codecs for Mobile StreamingabstractIn Video on Demand (VoD) scenarios, traditional codecs are the industry standard due to their high decoding efficiency. However, they suffer from severe quality degradation under low bandwidth conditions. While emerging generative neural codecs offer significantly higher perceptual quality, their reliance on heavy frame-by-frame generation makes real-time playback on mobile devices impractical. We ask: is it possible to combine the blazing-fast speed of traditional standards with the superior visual fidelity of neural approaches? We present HybridPrompt, the first generative-based video system capable of achieving real-time 1080p decoding at over 150 FPS on a commercial smartphone. Specifically, we employ a hybrid architecture that encodes Keyframes using a generative model while relying on traditional codecs for the remaining frames. A major challenge is that the two paradigms have conflicting objectives: the "hallucinated" details from generative models often misalign with the rigid prediction mechanisms of traditional codecs, causing bitrate inefficiency. To address this, we demonstrate that the traditional decoding process is differentiable, enabling an end-to-end optimization loop. This allows us to use subsequent frames as additional supervision, forcing the generative model to synthesize keyframes that are not only perceptually high-fidelity but also mathematically optimal references for the traditional codec. By integrating a two-stage generation strategy, our system outperforms pure neural baselines by orders of magnitude in speed while achieving an average LPIPS gain of 8% over traditional codecs at 200kbps. Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Zongming Guo, Xinggong Zhang |
NOSSDAV | 6 |
| 2026 | Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
Tianyi Gong, Zijian Cao 0007, Zixing Zhang 0009, Jiangkai Wu, Xinggong Zhang, Shuguang Cui, Fangxin Wang 0001 |
NSDI | 5 |
| 2026 | Artic: AI-oriented Real-time Communication for MLLM Video Assistant
Jiangkai Wu, Junquan Zhong, Xinggong Zhang |
SIGCOMM | 5 |
| 2026 | Camel: Frame-Level Bandwidth Estimation for Low-Latency Live Streaming under Video Bitrate UndershootingabstractLow-latency live streaming (LLS) has emerged as a popular web application, with many platforms adopting real-time protocols such as WebRTC to minimize end-to-end latency. However, we observe a counter-intuitive phenomenon: even when the actual encoded bitrate does not fully utilize the available bandwidth, stalling events remain frequent. This insufficient bandwidth utilization arises from the intrinsic temporal variations of real-time video encoding, which cause conventional packet-level congestion control algorithms to misestimate available bandwidth. When a high-bitrate frame is suddenly produced, sending at the wrong rate can either trigger packet loss or increase queueing delay, resulting in playback stalls. To address these issues, we present Camel, a novel frame-level congestion control algorithm (CCA) tailored for LLS. Our insight is to use frame-level network feedback to capture the true network capacity, immune to the irregular sending pattern caused by encoding. Camel comprises three key modules: the Bandwidth and Delay Estimator and the Congestion Detector, which jointly determine the average sending rate, and the Bursting Length Controller, which governs the emission pattern to prevent packet loss. We evaluate Camel on both large-scale real-world deployments and controlled simulations. In the real-world platform with 250M users and 2B sessions across 150+ countries, Camel achieves up to a 70.8% increase in 1080P resolution ratio, a 14.4% increase in media bitrate, and up to a 14.1% reduction in stalling ratio. In simulations under undershooting, shallow buffers, and network jitter, Camel outperforms existing congestion control algorithms, with up to 19.8% higher bitrate, 93.0% lower stalling ratio, and 23.9% improvement in bandwidth estimation accuracy. Zhidong Jia, Li Jiang 0021, Wei Zhang 0074, Lan Xie, Feng Qian 0001, Leju Yan, Zhou Sha, Yixuan Ban, Xinggong Zhang |
WWW | 13 |
| 2025 | PromptMobile: Efficient Promptus for Low Bandwidth Mobile Video StreamingabstractTraditional video compression algorithms exhibit significant quality degradation at extremely low bitrates. Promptus emerges as a new paradigm for video streaming, substantially cutting down the bandwidth essential for video streaming. However, Promptus is computationally intensive and can not run in real-time on mobile devices. This paper presents PromptMobile, an efficient acceleration framework tailored for on-device Promptus. Specifically, we propose (1) a two-stage efficient generation framework to reduce computational cost by 8.1x, (2) a fine-grained inter-frame caching to reduce redundant computations by 16.6%, (3) system-level optimizations to further enhance efficiency. The evaluations demonstrate that compared with the original Promptus, PromptMobile achieves a 13.6x increase in image generation speed. Compared with other streaming methods, PromptMobile achives an average LPIPS improvement of 0.016 (compared with H.265), reducing 60% of severely distorted frames (compared to VQGAN). Jiangkai Wu, Haoyang Wang 0015, Peiheng Wang, Xinggong Zhang, Zongming Guo |
APNet | 5 |
| 2025 | Graph-Based Encrypted Malicious Traffic Detection Under Flow Distribution Drift With Flow Sampling
Yunpeng Tan, Qingyang Li 0010, Mingxin Yang, Xinggong Zhang |
APNet | 4 |
| 2025 | Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AIabstractAI Video Chat emerges as a new paradigm for Real-time Communication (RTC), where one peer is not a human, but a Multimodal Large Language Model (MLLM). This makes interaction between humans and AI more intuitive, as if chatting face-to-face with a real person. However, this poses significant challenges to latency, because the MLLM inference takes up most of the response time, leaving very little time for video streaming. Due to network uncertainty, transmission latency becomes a critical bottleneck preventing AI from being like a real person. To address this, we call for AI-oriented RTC research, exploring the network requirement shift from "humans watching video" to "AI understanding video". We begin by recognizing the main differences between AI Video Chat and traditional RTC. Then, through prototype measurements, we identify that ultra-low bitrate is a key factor for low latency. To reduce bitrate dramatically while maintaining MLLM accuracy, we propose Context-Aware Video Streaming that recognizes the importance of each video region for chat and allocates bitrate almost exclusively to chat-important regions. To evaluate the impact of video streaming quality on MLLM accuracy, we build the first benchmark, named Degraded Video Understanding Benchmark (DeViBench). Finally, we discuss some open questions and ongoing solutions for AI Video Chat. DeViBench is open-sourced at: https://github.com/pku-netvideo/DeViBench. Jiangkai Wu, Xinggong Zhang |
HotNets | 4 |
| 2025 | Modeling Virtual Reality Traffic with Head Movement in Remote RenderingabstractThe proliferation of virtual reality (VR) content, particularly in resource-intensive applications, has been met by remote rendering to overcome local hardware limitations. Along with numerous advantages, remote rendering VR brings about a new traffic type that features huge throughput and burstiness generally, the understanding and modeling of which is critical for performing VR networking optimization to guarantee the Quality of Experience (QoE) of VR traffic transmission, including synthetic traffic generation and Network Slicing orchestrators. However, existing VR traffic modeling studies are limited in that they do not consider the impact of user interactions on VR traffic. In contrast, we carry out extensive traffic measurements in this paper, and discover that head movements actively affect the VR frame sizes generated. We analyze traffic features and further model the relationship between angular velocities and frame sizes quantitatively. A linear regressor is modeled to predict the frame size by considering history frame sizes and angular velocities jointly. We use Air Light VR (ALVR) to stream VR content in various scenarios, construct the datasets, and validate our model on top of them. The result shows that the our model is capable of reducing the 95% square prediction error by 18-30 compared to the state-of-the-art model. To the best of our knowledge, this is the first investigation into the intricate relationship between remote rendering VR traffic and head movement. Our dataset and results will be publicly available and reproducible. Yihang Zhang 0007, Zhidong Jia, Li Jiang 0021, Qingyang Li 0010, Xinggong Zhang, Zongming Guo |
ICC | 5 |
| 2025 | 3DGCoding: Novel Framework for 3D Gaussian Video Incremental Training and CodingabstractFree-viewpoint videos (FVVs) streaming plays a pivotal role in driving the progress of immersive AR/VR applications. To be streamable, FVVs should simultaneously satisfy perframe decoding, compactness, and real-time rendering. Neural radiance based works suffer from low rendering speed, while previous 3D Gaussians (3DGs) based works necessitate large data volumes and pose challenges for compact representation. To support FVVs streaming, we present 3DGCoding, a novel framework for efficient reconstruction and coding of 3DGs for dynamic scenes. We design a learnable 3D mask to partition 3DGs into static and dynamic components, capturing all types of changes and making 3DGs quantitatively and attributively consistent for coding. Then, we propose a 2D-based coding process to further address inter-frame redundancy, achieving independent decoding and smaller data volumes. Experiments confirm that 3DGCoding could stream with 0.49MB per frame, and achieves competitive performance compared with state-of-the-art FVVs streaming methods. Peiheng Wang, Haodan Zhang, Quanlu Jia, Jiangkai Wu, Haoyang Wang 0015, Xinggong Zhang |
ICME | 7 |
| 2025 | Sync5D: Novel View Synthesis from a Single Image with 5D Consistency
Junlin Hao, Yunpeng Tan, Jiangkai Wu, Peiheng Wang, Xinggong Zhang, Zongming Guo |
PRCV (10) | 6 |
| 2024 | BurstRTC: Harnessing Variable Bit-Rate of RTC through Frame-Bursting Congestion ControlabstractThe rapid growth of online interactive video applications reflects the increasing popularity of real-time communication (RTC). Despite advancements in network and video technologies, the worse quality of experience (QoE) such as large delay, rebuffering and low image quality, etc. remains to be complained generally. We argue that this is mainly due to the legacy network-oriented congestion control (CC), which assumes continuous stream of packets are sent. But it is not satisfied for RTC since bit-rate variation is inherent to RTC’s video encoder. Zhidong Jia, Yihang Zhang 0007, Qingyang Li 0010, Xinggong Zhang |
APNet | 4 |
| 2024 | StarTCP: Handover-aware Transport Protocol for StarlinkabstractLegacy transport protocols such as TCP and QUIC suffer from high packet loss and low link utilization in Starlink. From the measurement data, we figure out the ground-satellite link (GSL) handover is mainly to blame. The periodic handovers result in link interruptions and bursty losses with a fixed interval of 15s, which impair TCP’s performance. Based on this finding, we present a handover-aware transport protocol, StarTCP, which proactively stalls transmission during handovers to avoid bursty losses and erroneous congestion signals. Preliminary results indicate that StarTCP can efficiently reduce packet loss and enhance throughput in Starlink. Li Jiang 0021, Yihang Zhang 0007, Yannan Hu, Yong Cui 0001, Xinggong Zhang |
APNet | 5 |
| 2024 | SRFC: Scalable Radiance Fields Streaming with Planar CodecabstractVolumetric videos afford comprehensive and immer-sive viewing experiences with six degrees of freedom (6DoF) for navigation, allowing users to move freely within a three-dimensional space. Radiance fields (RF) is emerging to reproduce photorealistic 3D scenes, which achieves lighting consistency and realistic transfer between the real and virtual worlds. In this paper, we investigate how to deliver photorealistic volumetric video under the network bandwidth constraints and low-power device. We design a novel Scalable Radiance Fields Video Streaming, SRFC, to enable streaming RF video with planar codec. An Orthogonal Layered Depth Image (OLDI) mapping is introduced to map 3D to view-dependent 2D plane images, in order to reduce bitrates and decoding complexity by the commercial 2D codec. Moreover, to address the issue of deviation of the user viewpoint, we propose a dual-tiered structure in radiance fields and optimizes user's perceptual quality by adaptive bitrate and viewpoint decision. The evaluations demonstrate that SRFC is capable of reducing bandwidth requirements by 4 × and increasing decoding frame rate by nearly 3 × while maintaining satisfactory user perception of video quality. Quanlu Jia, Haodan Zhang, Haoyang Wang 0015, Jiangkai Wu, Xinggong Zhang, Zongming Guo |
ICC | 6 |
| 2024 | Tackling Bit-Rate Variation of RTC Through Frame-Bursting Congestion ControlabstractInteractive video applications signal widespread interest in Real-time communication (RTC), yet issues like frame delay, rebuffering, etc. remain a common complaint. We argue that this is mainly due to legacy network-oriented congestion control (CC), which assumes a continuous stream of packets is being sent. But this assumption doesn't hold in RTC since the video encoder exhibits inherent bit-rate variation: (1) Bursty spiking bit-rate leads to packets waiting in the sending buffer, which increases frame delay. (2) Low bit-rate causes insufficient packets available for sending, making current CCs hard to detect available bandwidth. In response, we propose BurstRTC, a novel paradigm for RTC transport protocol. Each frame is emitted as a whole, and the video bit-rate is directly controlled by network congestion feedback. BurstRTC uses frame-bursting to estimate available bandwidth efficiently regardless of bit-rate variation. Considering the impact of bit-rate variation on network congestion, BurstRTC models frame size as a Gaussian distribution instead of a fixed size and further derives its frame delay, preventing suboptimal performance of purely network-oriented designs. An analytic method for determining the target bit-rate replaces the trial-and-error updates of gradient-based methods, ensuring fast convergence to the available bandwidth. We evaluated the performance of BurstRTC and found that, compared with GCC, BurstRTC achieves up to$59.8 \%$higher bit-rate and up to$\mathbf{4 8. 9 \%}$lower frame delay. Further, compared with SQP and Pudica, BurstRTC can also reduce tail frame delay by up to$89.2 \%$, and improve average bit-rate by up to$15.6 \%$. Zhidong Jia, Yihang Zhang 0007, Qingyang Li 0010, Xinggong Zhang |
ICNP | 4 |
| 2024 | NetSentry: Scalable Volumetric DDoS Detection with Programmable SwitchesabstractDistributed Denial of Service (DDoS) attack is a critical and persistent threat to the Internet. Recent DDoS detection schemes based on emerging programmable switches can achieve higher processing throughput and improve detection accuracy. However, with limited data plane memory, such schemes are not suitable for handling a large number of concurrent flows. Prior arts that attempt to increase memory efficiency have failed to do so without the expense of cost and accuracy. In this paper, we propose NetSentry, the first programmable switch based dynamic pooled testing DDoS detector. NetSentry detects DDoS in a pooled testing manner, where multiple flows are grouped to share the same storage unit on the data plane. NetSentry designs an elastic flow aggregation mechanism to dynamically adjust the detection granularity. Further, to achieve accurate DDoS detection for aggregated flows, NetSentry implements frequency domain DDoS detection on programmable switches. Evaluations of NetSentry’s hardware prototype show that NetSentry can achieve better accuracy while saving up to 91% of the data plane memory required to store flow features compared to the state-of-the-art programmable switch-based flow classification scheme. Junchen Pan, Kunpeng He, Lei Zhang 0157, Zhuotao Liu, Xinggong Zhang, Yong Cui 0001 |
IWQoS | 5 |
| 2024 | RTCC: Enable End-to-end Sub-RTT Congestion Control for Next-generation NetworkabstractThe advancement of next-generation networks such as 5G/6G and satellite systems has significantly increased available network bandwidth, while also exacerbating network burstiness. This surge presents a formidable challenge for congestion control (CC), a pivotal mechanism for achieving high bandwidth utilization and low latency by adjusting congestion windows or modifying sending rates. Traditional end-to-end CC algorithms fall short of optimality due to their reliance on congestion signals in acknowledgment packets, which introduce a delay of one round-trip time (RTT). In this paper, to mitigate end-to-end delayed feedback, we introduce a novel Real-Time Congestion Control (RTCC) algorithm that integrates machine learning with conventional model-based CC. RTCC employs a Multi-Layer Perceptron (MLP) to model network conditions and predict current congestion signals accurately. A tailored network model then utilizes these predictions to manage packet accumulation in the network bottleneck. Moreover, an online model-updating mechanism is proposed to adapt to diverse network environments. We integrate RTCC into QUIC and conduct comprehensive experiments in both emulated test-beds and real-world settings, including WiFi/4G/5G and cross-continent networks. The results demonstrate RTCC's efficacy, with up to a 32% increase in average throughput, a reduction in RTT by up to 21% compared to BBR V2, and the maintenance of fair bandwidth allocation. Yihang Zhang 0007, Zhidong Jia, Qingyang Li 0010, Xinggong Zhang, Zongming Guo |
SECON | 4 |
| 2024 | Inferring Video Streaming Quality of Real-Time Communication Inside NetworkabstractReal-time video streaming is getting indispensable in people’s daily life, and poses heavy loads and stringent performance requirements on the network. For Internet Service Providers (ISPs), ensuring high-quality real-time video communication is a widely concerned issue. However, inferring the quality of real-time video streaming based on passively-collected network traffic is a great challenge due to limited information in the User Datagram Protocol (UDP) header and the encryption of the application-level protocol. In this paper, we propose IReaV-T to Infer Real-time Video streaming quality with a generalized Transformer, which understands the intrinsic state of the network and predicts the future real-time video quality. By applying novel embedding methods, IReaV-T could make full use of observed traffic features and distinguish different real-time video applications. Extensive comparative experiments demonstrate the effectiveness of IReaV-T, showing that IReaV-T could predict future real-time video quality with mean squared Video Multimethod Assessment Fusion (VMAF) score error less than 6. Yihang Zhang 0007, Sheng Cheng 0002, Zongming Guo, Xinggong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | LEOTP: An Information-Centric Transport Layer Protocol for LEO Satellite NetworksabstractLow Earth orbit (LEO) satellite networks have attracted extensive research due to their potential to provide high-quality Internet access services. However, the existing TCP variants, which are designed for terrestrial networks, can hardly work in LEO satellite networks with characteristics such as error-prone, bandwidth variations, and link switching. To address these challenges, in this paper we present a new information-centric transport layer protocol LEOTP to guarantee reliable, high-throughput, and low-latency data transmission in LEO satellite networks. It leverages the idea of Information-Centric Networking (ICN) with a Request-Response transmission model and in-network caching. The connectionless transmission paradigm in LEOTP makes it resilient to dynamic topology changes. The caches equipped in intermediate nodes help to recover packet loss while the hop-by-hop congestion control mechanism provides a fast reaction to time-varying network conditions. We evaluate the performance of LEOTP in emulated Starlink constellation, which shows that it increases the throughput by 8%-12% with 40%-60% delay reduction compared with the state-of-the-art TCP variants in the transcontinental data transmission. Li Jiang 0021, Yihang Zhang 0007, Jinyu Yin, Xinggong Zhang, Bin Liu 0001 |
ICDCS | 4 |
| 2023 | Rebuffering but not Suffering: Exploring Continuous-Time Quantitative QoE by User's Exiting Behaviors
Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
INFOCOM | 3 |
| 2023 | QUTY: Towards Better Understanding and Optimization of Short Video QualityabstractShort video applications such as TikTok and Instagram have attracted tremendous attention recently. However, it is very limited for industry and academia to understand the user's Quality of Experience (QoE) on short video, let alone how to improve the QoE in short video streaming. Haodan Zhang, Yixuan Ban, Zongming Guo, Zhimin Xu 0001, Yue Wang 0032, Xinggong Zhang |
MMSys | 7 |
| 2023 | ZGaming: Zero-Latency 3D Cloud Gaming by Image PredictionabstractIn cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction. Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang |
SIGCOMM | 6 |
| 2023 | ActRay: Online Active Ray Sampling for Radiance FieldsabstractThanks to the high-quality reconstruction and photorealistic rendering, the Neural Radiance Field (NeRF) has garnered extensive attention and has been continuously improved. Despite its high visual quality, the prohibitive training time limits its practical application. Although significant acceleration has been achieved, it is still far from real-time training, due to the need for tens of thousands of iterations. In this paper, a feasible solution is to reduce the number of required iterations by always training the rays with the highest loss values, instead of the traditional method of training each ray with a uniform probability. To this end, we propose an online active ray sampling strategy, ActRay. Specifically, to avoid the substantial overhead of calculating the actual loss values for all rays in each iteration, a rendering-gradient-based loss propagation algorithm is presented to efficiently estimate the loss values. To further narrow the gap between the estimated loss and the actual loss, an online learning algorithm based on the Upper Confidence Bound (UCB) is proposed to control the sampling probability of the rays, thereby compensating for the bias in loss estimation. We evaluate ActRay on both real-world and synthetic scenes, and the promising results show that it accelerates radiance field training by 6.5x. Besides, we test ActRay under all kinds of radiance field representations (implicit, explicit, and hybrid), proving that it is general and effective to different representations. We believe this work will contribute to the practical application of radiance fields, because it has taken a step closer to real-time radiance field training. ActRay is open-source at: https://pku-netvideo.github.io/actray/. Jiangkai Wu, Yunpeng Tan, Quanlu Jia, Haodan Zhang, Xinggong Zhang |
SIGGRAPH Asia | 6 |
| 2023 | ABRF: Adaptive BitRate-FEC Joint Control for Real-Time Video StreamingabstractAdaptive Forward Error Correction (AFEC) algorithms are proposed to achieve efficient Forward Error Correction (FEC) in Real-Time Communication (RTC). However, current AFEC approaches suffer two key limitations. 1) They do not consider the squeezing effect of redundancy on source bitrate while the squeezing usually happens in real RTC applications with limited bandwidth. 2) They estimate the future packet loss simply, ignoring some critical features of packet loss like randomness and multi-pattern. These drawbacks stop them from providing better Quality of Experience (QoE) in RTC services. We propose ABRF, a general QoE-oriented Adaptive BitRate-FEC joint control algorithm. ABRF makes predictions on the network loss pattern in the coming time and jointly calculates the optimal bitrate-FEC decision based on a QoE model for real-time video streaming. Moreover, ABRF is equipped with a fast adaptation method which helps it generalize across diverse network environments. In terms of Video Multi-method Assessment Fusion (VMAF), experimental results tell that ABRF decreases VMAF degradation caused by packet loss by 68%-95% compared with other AFEC algorithms in real-time video streaming in real-world Internet. Sheng Cheng 0002, Xinggong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | RAM360: Robust Adaptive Multi-Layer 360$^\circ$ Video Streaming With Lyapunov OptimizationabstractViewport-adaptive streaming approaches are emerging as the most promising way to deliver high-quality 360 videos over mobile networks. However, the viewport prediction is only reliable within a short prediction window, i.e., a short playback buffer, which conicts with maintaining a long buffer to avoid playback rebuffering. To deal with this problem, we present RAM360, a Robust Adaptive Multi-layer 360 video streaming system, to ensure high viewport quality and low stall ratio concurrently. We make three technical contributions. First, we design a QoE-driven robust multi-layer streaming framework, where each chunk is encoded by multiple independent layers with different quality levels. The client dynamically decides which chunk and layer to be downloaded by their QoE contributions. Thus, the base-layer could be prefetched to avoid the risk of stalling while the viewport quality is improved by downloading enhancement layer. Second, we establish a novel QoE model to represent the quality of whole playback session, not that of chunk. It aims to maximize the overall QoE of playback session. Third, we introduce the Lyapunov optimization theory to solve the QoE optimization problem, which is an online algorithm with near-optimality solution. We demonstrate that RAM360 can significantly outperform the existing schemes regarding viewport quality, stall ratio, and QoE through extensive experiments with public datasets. Haodan Zhang, Yixuan Ban, Zongming Guo, Xinggong Zhang |
IEEE Trans. Multim. | 5 |
| 2022 | Bandwidth-Efficient Multi-video Prefetching for Short Video StreamingabstractApplications that allow sharing of user-created short videos exploded in popularity in recent years. A typical short video application allows a user to swipe away the current video being watched and start watching the next video in a video queue. Such user interface causes significant bandwidth waste if users frequently swipe a video away before finishing watching. Solutions to reduce bandwidth waste without impairing the Quality of Experience (QoE) are needed. Solving the problem requires adaptively prefetching of short video chunks, which is challenging as the download strategy needs to match unknown user viewing behavior and network conditions. In our work, we first formulate the problem of adaptive multi-video prefetching in short video streaming. Then, to facilitate the integration and comparison of researchers' algorithms towards solving the problem, we design and implement a discrete-event simulator, which we release as open source. Finally, based on the organization of the Short Video Streaming Grand Challenge at ACM Multimedia 2022, we analyze and summarize the algorithms of the contestants, with the hope of promoting the research community towards addressing this problem. Xutong Zuo, Yishu Li, Mohan Xu, Wei Tsang Ooi, Jiangchuan Liu, Junchen Jiang, Xinggong Zhang, Kai Zheng 0003, Yong Cui 0001 |
ACM Multimedia | 7 |
| 2022 | STC: FoV Tracking Enabled High-Quality 16K VR Video Streaming on Mobile PlatformsabstractThe ultra-high-definition 16K Virtual Reality (VR) video is coming to ages with more ”real” virtual experience and less cybersickness. However, the huge bitrate and decoding overhead would overwhelm today’s network and mobile hardware. The widely-known Field-of-View (FoV) adaptation streaming method still has severe bitrate wastes and decoding overhead as it delivers FoV areas with grid-like static tiles. Inspired by this, we present a novel ShiftTile-traCking (STC) streaming scheme, which crops and delivers tiles by tracking FoV movement. It is equivalent to deliver an FoV planar video instead of VR videos. This would save huge bit-rate and reduce decoding complexity. We mainly entail three contributions. 1) To reduce projection distortions, a novel FoV-centric sphere projection is proposed, which projects VR videos with the center of users’ FoVs. 2) To cover diverse FoV movement trajectories with a limited number of tiles, we propose an optimal tiling algorithm by trajectory clustering. 3) To be resilient to FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which scales FoV areas by the prediction accuracy. The evaluation shows that under the same network conditions, STC improves up to 1.3dB V-PSNR, reduces up to 13.2% buffering ratio, and achieves 60% faster decoding speed (61.5 frames per second) compared with state-of-the-art solutions. Chengyuan Zheng, Jinyu Yin, Fangzhen Wei, Yu Guan 0005, Zongming Guo, Xinggong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | SODA: Similar 3D Object Detection Accelerator at Network Edge for Autonomous DrivingabstractOffloading the 3D object detection from autonomous vehicles to MEC is appealing because of the gains on quality, latency, and energy. However, detection requests lead to repetitive computations since the multitudinous requests share approximate detection results. It is crucial to reduce such fuzzy redundancy by reusing the previous results. A key challenge is that the requests mapping to the reusable result are only similar but not identical. An efficient method for similarity matching is needed to justify the use case. To this end, by taking advantage of TCAM's ap-proximate matching capability and NMC's computing efficiency, we design SODA, a first-of-its-kind hardware accelerator which sits in the mobile base stations between autonomous vehicles and MEC servers. We design efficient feature encoding and partition algorithms for SODA to ensure the quality of the similarity matching and result reuse. Our evaluation shows that SODA significantly improves the system performance and the detection results exceed the accuracy requirements on the subject matter, qualifying SODA as a practical domain-specific solution. Wenquan Xu, Haoyu Song 0001, Linyang Hou, Xinggong Zhang, Chuwen Zhang, Wei Hu 0003, Yi Wang 0004, Bin Liu 0001 |
INFOCOM | 5 |
| 2021 | LightFEC: Network Adaptive FEC with a Lightweight Deep-Learning ApproachabstractNowadays, the interest of real-time video streaming reaches a peak. To deal with the problem of packet loss and optimize users' Quality of Experience (QoE), Forward error correction (FEC) has been studied and applied extensively. The performance of FEC depends on whether the future loss pattern is precisely predicted, while the previous researches have not provided a robust packet loss prediction method. In this work, we propose LightFEC to make accurate and fast prediction of packet loss pattern. By applying long short-term memory (LSTM) networks, clustering algorithms and model compression methods, LightFEC is able to accurately predict packet loss in various network conditions without consuming too much time. According to the results of well-designed experiments, we find out that LightFEC outperforms other schemes on prediction accuracy, which improves the packet recovery ratio while keeping the redundancy ratio at a low level. Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
ACM Multimedia | 3 |
| 2021 | Adapting Named Data Networking (NDN) for Better Consumer Mobility Support in LEO Satellite NetworksabstractLarge low Earth orbit (LEO) satellite constellations provide low-latency and high-bandwidth Internet connectivity at the global scale. One major challenge is to handle frequent satellite handovers. Named Data Networking (NDN) adopts a pull-based communication model, which allows users to retrieve data that fail to come back because of satellite handovers by retransmitting the corresponding requests, hence simplifying mobility management when retrieving data. However, we find that relying on such retransmissions alone can be highly inefficient in typical LEO satellite constellations. Specifically, typical inter-satellite topologies and satellite handover strategies may produce bad cases for retransmissions, generating a significant amount of additional traffic. Motivated by this observation, this paper attempts to consolidate NDN's advantage in mobility management with the Data Recovery Link Service (DRLS), a shim layer service operating between the network and link layer in the NDN protocol stack. DRLS hides recurring satellite handovers from forwarding by recovering data from the previously connected satellite via alternative paths, thus ensuring the bidirectional request-response exchange of NDN without retransmitting requests. A prototype of DRLS is implemented in the reference NDN software forwarder and evaluated through simulations. Results prove the efficacy of the proposed mechanism at reducing the overall traffic volume. Zhongda Xia, Yu Zhang 0036, Teng Liang, Xinggong Zhang, Binxing Fang |
MSWiM | 4 |
| 2021 | PrefCache: Edge Cache Admission With User Preference Learning for Video Content DistributionabstractWith the deployment of video streaming in 4G/5G mobile network, Content Delivery Networks (CDN) are extending to the network edge to provide end-users better Quality of Experience (QoE). However, small cache size and irregular request patterns make it a great challenge for edge caching in video content distribution. Most of the existing cache policies are item-wise, they admit each video object separately, which performs poorly on the network edge due to irregular request patterns. We observe that compared with single video objects, users' preferences for video topics are much more constant, thus are easier to be predicted. So we propose PrefCache, a novel cache admission policy based on preference learning, for video content edge caching. PrefCache enables an edge cache to learn users' preferences for videos in real-time. Once receiving a video object, PrefCache decides whether to admit it to the cache by whether it is under users' preference. We make three contributions in this work. (1) First, we design an information collector, which can proactively collect the preference-related information without any modification of clients and video providers. (2) Second, we propose a tree-structure model to learn and compress users' preferences. (3) Third, to decide which videos should be admitted to the cache in real-time, an explore-and-exploit method is applied. We carried out extensive experiments with 24 hours of trace data from a large commercial video content provider. The experimental results demonstrate that PrefCache can improve hit ratio up to 12%, and save 92% memory / 98% CPU overhead, compared to the state-of-the-art cache policies. Yu Guan 0005, Xinggong Zhang, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | STC: Enabling 16K VR streaming on mobile platforms with FoV trackingabstract16K VR videos are coming to ages. But it could overwhelm mobile hardware for its huge bandwidth consumption and decoding complexity. To enable 16K VR video streaming over mobile platforms, we present a novel ShiftTile-Tracking (STC) streaming system, which crops and transmits video by tracking the Field-of-View (FoV) movement of users. The video chunk is split into ShiftTiles with frame granularity, which always covers FoV areas along the FoV movement trajectory. This transforms a 360-degree VR video into a traditional planar video, which leads to huge bandwidth saving and faster decoding speed. In the system design, we mainly entail two contributions. 1) To accommodate various FoV movement trajectories with a limited number of ShiftTiles, we propose an optimal tiling algorithm by FoV trajectory clustering. 2) To be resilient to the FoV prediction errors, we propose an accuracy-sensitive streaming algorithm, which expands the FoV area if the FoV prediction errors are high. The evaluation shows that under the same real-world 4G network conditions, the proposed STC improves 0. 9dBV-PSNR, reduces 12.4% buffering ratio, and achieves 45% faster decoding speed (64 frames per second) on average compared with the state-of-the-art solutions. This enables 16K VR video streaming on current mobile platforms. Chengyuan Zheng, Jinyu Yin, Yu Guan 0005, Xinggong Zhang, Zongming Guo |
GLOBECOM | 4 |
| 2020 | MA360: Multi-Agent Deep Reinforcement Learning Based Live 360-Degree Video Streaming on EdgeabstractThe mobile edge caching has made video service providers deliver live 360-degree videos worldwide. However, these services still suffer from the huge network traffic on the core network due to the spherical nature and the diverse requests generated from large user populations. It is challenging to optimize the Quality of Experience (QoE) and the bandwidth consumption simultaneously under the significant number of users as well as dynamic network and playback status. In this paper, we propose a Multi-Agent deep reinforcement learning based 360-degree video streaming system, named MA360, to tackle this multi-user live 360-degree video streaming problem in the context of the edge cache network. Specifically, MA360 employs the Mean Field Actor-Critic (MFAC) algorithm to make clients collaboratively and distributively request tiles aiming at maximizing the overall QoE while minimizing the total bandwidth consumption. Experiments over real-world datasets show that MA360 can improve the QoE while significantly reducing the bandwidth consumption compared with several state-of-the-art edge-assisted 360-degree video streaming strategies. Yixuan Ban, Yuanxing Zhang, Haodan Zhang, Xinggong Zhang, Zongming Guo |
ICME | 4 |
| 2020 | DeepRS: Deep-Learning Based Network-Adaptive FEC for Real-Time Video CommunicationsabstractAs real-time multimedia streaming thriving, Forward Error Correction (FEC) methods have been studied and applied extensively these years. Most of researchers paid their attention to the coding algorithms, attempted to balance the trade off between recovery ratio and delay with fewer redundance. However, when packet loss pattern changes dynamically, the redundance waste is too serious to be ignored. In this work, we propose a novel algorithm which adjusts the redundance ratio of FEC encoder according to the prediction of packet loss. Receivers are additionally required to feedback observed packet loss pattern. Streaming sender collects the feedbacked packet loss pattern and predicts the number of packet loss in the incoming short period. As for implementation, we adopt long short-term memory (LSTM) network as our deep learning algorithm, and exquisitely embed it in our adaptive FEC system. With the extensive experiments, our proposed scheme outperforms other FEC methods greatly both in the simulations and evaluations on traces observed from the real world. Sheng Cheng 0002, Xinggong Zhang, Zongming Guo |
ISCAS | 3 |
| 2020 | APL: Adaptive Preloading of Short Video with Lyapunov OptimizationabstractShort video applications, like TikTok, have attracted many users across the world. It can feed short videos based on users' preferences and allow users to slide the boring content anywhere and anytime. To reduce the loading time and keep playback smoothness, most of the short video apps will preload the recommended short videos in advance. However, these apps preload short videos in fixed size and fixed order, which can lead to huge playback stall and huge bandwidth waste. To deal with these problems, we present an Adaptive Preloading mechanism for short videos based on Lyapunov Optimization, also called APL, to achieve near-optimal playback experience, i.e., maximizing playback smoothness and minimizing bandwidth waste considering users' sliding behaviors. Specifically, we make three technical contributions: (1) We design a novel short video streaming framework which can dynamically preload the recommended short videos before the current video is downloaded completely. (2) We formulate the preloading problem into a playback experience optimization problem to maximize the playback smoothness and minimize the bandwidth waste. (3) We transform the playback experience optimization problem during the whole viewing process into a single-step greedy algorithm based on the Lyapunov optimization theory to make the online decisions during playback. Through extensive experiments based on the real datasets that generously provided by TikTok, we demonstrate that APL can reduce the stall ratio by 81%/12% and bandwidth waste by 11%/31% compared with no-preloading/fixed-preloading mechanism. Haodan Zhang, Yixuan Ban, Xinggong Zhang, Zongming Guo, Zhimin Xu 0001, Shengbin Meng, Yue Wang 0032 |
VCIP | 3 |
| 2020 | Statistical Learning Based Congestion Control for Real-Time Video CommunicationabstractThe existing congestion control is hard to simultaneously achieve low latency, high throughput, good adaptability and fair bandwidth allocation, mainly because of the hardwired control strategy and egocentric convergence objective. To address these issues, we propose an end-to-end statistical learning based congestion control, named Iris. By exploring the underlying principles of self-inflicted delay, we find that RTT variation is linearly related to the difference between sending rate and receiving rate, which inspires us to control video bit rate using a statistical-learning congestion control model. The key idea of Iris is to force all flows to converge to the same queue load and adjust bit rate by the model. All flows keep a small and fixed number of packets queuing in the network, thus the fair bandwidth allocation and low latency are both achieved. Besides, the adjustment step size of sending rate is updated by online learning, to better adapt to dynamically changing networks. We carried out extensive experiments to evaluate the performance of Iris, with the implementations over transport layer and application layer respectively. The testing environment includes emulated network, real-world Internet and commercial cellular networks. Compared against Transmission Control Protocol (TCP) flavors and state-of-the-art protocols, Iris is able to achieve high bandwidth utilization, low latency and good fairness concurrently. Especially for HyperText Transfer Protocol (HTTP) video streaming service, Iris is able to increase the video bitrate up to 25% and Peak Signal to Noise Ratio (PSNR) up to 1 dB. Tongyu Dai, Xinggong Zhang, Yihang Zhang 0007, Zongming Guo |
IEEE Trans. Multim. | 2 |
| 2019 | Optimal Viewport-Adaptive 360-Degree Video Streaming Against Random Head MovementabstractRecently, a significant interest in 360-degree virtual reality (VR) video has been formed. However, a key problem is how to design a robust adaptive streaming approach and implement a practical system. The traditional streaming methods which are not sensitive to user's viewport could cause huge bandwidth budget with low video quality, while the viewport-adaptive schemes may be not accurate enough especially under random head movement. In this paper, we have designed an optimal viewport-adaptive 360-degree video streaming scheme, which is to maximize Quality of Experience (QoE) by predicting user's viewport with a probabilistic model, prefetching video segments into the buffer and replacing some unbefitting segments. In this way, continuous and smooth playback, high bandwidth utilization, low viewport prediction error and high peak signal-to-noise ratio in the viewport (V-PSNR) can be obtained. In order to deal with user's head movement, we have reduced prediction error rate by employing a probabilistic viewport prediction model, and a replacement strategy has been applied to update the downloaded segments when the user's viewport suddenly changes. To reduce smoothness loss, the segments' oscillation during playback has been considered. We also developed a prototype system with our method. The well-designed experiments provided numerous results which proved the better performance of our scheme. Zhimin Xu 0001, Xinggong Zhang, Zongming Guo |
ICC | 3 |
| 2019 | UtilCache: Effectively and Practicably Reducing Link Cost in Information-Centric NetworkabstractMinimizing total link cost in Information-Centric Network (ICN) by optimizing content placement is challenging in both effectiveness and practicality. To attain better performance, upstream link cost caused by a cache miss should be considered in addition to content popularity. To make it more practicable, a content placement strategy is supposed to be distributed, adaptive, with low coordination overhead as well as low computational complexity. In this paper, we present such a content placement strategy, UtilCache, that is both effective and practicable. UtilCache is compatible with any cache replacement policy. When the cache replacement policy tends to maintain popular contents, UtilCache attains low link cost. In terms of practicality, UtilCache introduces little coordination overhead because of piggybacked collaborative messages, and its computational complexity depends mainly on content replacement policy, which means it can be O(1) when working with LRU. Evaluations prove the effectiveness of UtilCache, as it saves nearly 40% link cost more than current ICN design. Lemei Huang, Yu Guan 0005, Xinggong Zhang, Zongming Guo |
ICC | 3 |
| 2019 | CACA: Learning-based Content-aware Cache Admission for Video Content in Edge CachingabstractIn the last decades, network caches (Content Distribution Network, CDN) have been widely deployed in video delivery system. As cache has been pushed to network edge as far as possible, small cache size and irregular request pattern make it a great challenge for edge cache to catch popular video contents. Although we can apply cache admission policies to block cold contents out, however, all current admission policies are still based on request pattern (content size, frequency), which perform poorly in edge cache. This paper proposes a novel feature-based cache admission policy, Content-feature Aware Cache Admission(CACA). It admits video objects to cache by video features, not by request pattern anymore. The intuition behind that is, for a group of users, their preferred contents may change at any time, but their preferred content features would maintain for a while. Popularity of video features (such as topic, author), is much more predicable than that of single video object. To mine critical features from huge feature space, this paper proposes a tree-structure reinforcement learning algorithm. Critical features are learned from a feature-partition tree which is spanned and pruned by history popularity. Then, an Exploration-and-Exploitation method is used to select the Top-K critical features. Video contents with these features will be admitted to cache. We carried out extensive experiments with 24-hours data traces from a commercial video content provider. The experimental results demonstrate that the proposed CACA is able to improve hit ratio up to 15%, reduce back-to-origin up to 20% and save 95% memory, compared with state-of-art cache admission policies. Yu Guan 0005, Xinggong Zhang, Zongming Guo |
ACM Multimedia | 2 |
| 2019 | Pano: optimizing 360° video streaming with a better understanding of quality perceptionabstractStreaming 360° videos requires more bandwidth than non-360° videos. This is because current solutions assume that users perceive the quality of 360° videos in the same way they perceive the quality of non-360° videos. This means the bandwidth demand must be proportional to the size of the user's field of view. However, we found several quality-determining factors unique to 360° videos, which can help reduce the bandwidth demand. They include the moving speed of a user's viewpoint (center of the user's field of view), the recent change of video luminance, and the difference in depth-of-fields of visual objects around the viewpoint. Yu Guan 0005, Chengyuan Zheng, Xinggong Zhang, Zongming Guo, Junchen Jiang |
SIGCOMM | 3 |
| 2019 | QoE-Driven Adaptive K-Push for HTTP/2 Live StreamingabstractDynamic adaptive streaming (DAS) over HTTP has been widely deployed over the Internet. However, due to the pull-based nature of HTTP/1.1, there exists intolerable streaming latency and high request overhead in the current DAS systems. With dynamic k-push, HTTP/2 live streaming promises to achieve low live latency with less overhead and small segment duration. In this paper, we propose a quality of experience (QoE) driven adaptive k-push mechanism (QK-Push) for HTTP/2 live streaming. The client just sends one request to set push length ($K$ ) and bitrate (v) parameters and the server would push back $K$ segments in a batch. To determine k-push parameters, a probabilistic buffer model is first designed to avoid buffer underflow/overflow. Also, three QoE objective functions are designed to ensure the high streaming quality (bitrate), playback continuity, and smoothness. QK-Push casts this multi-objective optimization problem as a Pareto optimal problem. To solve it, a Nash bargaining solution is designed to balance the needs for video quality, bitrate smoothness, and request overhead. Finally, the segments in each push cycle are selected by solving the Nash problem with a discrete space Lagrangian method. We implement an HTTP/2 live streaming prototype system, with the QK-Push algorithm over modified dash.js and media presentation description. To evaluate the performances, the extensive live streaming experiments are carried out over a controllable network test bed and real Internet trace. The results demonstrate that the proposed QK-Push algorithm is able to improve the average bitrate up to 13%, reduce the bitrate oscillations up to 81%, decrease the startup delay up to 58%, and increase the estimate the mean opinion score up to 12% compared to the current HTTP/1.1 system. Zhimin Xu 0001, Xinggong Zhang, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | TFDASH: A Fairness, Stability, and Efficiency Aware Rate Control Approach for Multiple Clients Over DASHabstractDynamic adaptive streaming over HTTP (DASH) has recently been widely deployed in the Internet and adopted in the industry. It, however, does not impose any adaptation logic for selecting the quality of video segments requested by clients and suffers from lackluster performance with respect to a number of desirable properties: efficiency, stability, and fairness when multiple players compete for a bottleneck link. In this paper, we propose a throughput-friendly DASH rate control scheme for video streaming with multiple clients over DASH to well balance the tradeoffs among efficiency, stability, and fairness. The core idea behind guaranteeing fairness and high efficiency (bandwidth utilization) is to avoid OFF periods during the downloading process for all clients, i.e., the bandwidth is in perfect-subscription or over-subscription with bandwidth utilization approach to 100%. We also propose a dual-threshold buffer model to solve the instability problem caused by the above idea. As a result, by integrating these novel components, we also propose a probability-driven rate adaption logic taking into account several key factors that most influence visual quality, including buffer occupancy, video playback quality, video bit-rate switching frequency and amplitude, to guarantee high-quality video streaming. Our experiments evidently demonstrate the superior performance of the proposed method. Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Name-Based Routing with On-Path Name Lookup in Information-Centric NetworkabstractName-based routing is one of the core ideas in Information-centric network (ICN). In name-based routing, there is a tradeoff between the cost of name announcement and name lookup. Some ICN architectures introduce an efficient way of name lookup but pay high price in name announcement, others cut off most information exchange in name announcement yet introduce heavy burden in name lookup. In order to solve this problem and balance the cost of name announcement and lookup, we propose Name-based routing with On-Path Name Lookup (OPNL). OPNL looks up name prefixes on the path to name's guaranteed destination. It accomplishes distributed name lookup with lighter burden while maintaining little information exchange in name announcement. Results of simulation experiments show that OPNL makes a tradeoff between the cost of name announcement and lookup to have better scalability, eliminates storage overhead and communication overhead compared with prior works and attains even better performance. Yu Guan 0005, Lemei Huang, Xinggong Zhang, Zongming Guo |
ICC | 3 |
| 2018 | CUB360: Exploiting Cross-Users Behaviors for Viewport Prediction in 360 Video Adaptive StreamingabstractTo ensure 360-degree video's continuous playback and reduce the bandwidth waste, predicting user's future fixation is indispensable. However, existing methods concentrate either on user's motion information or content information. None of them consider users watching behaviors' inconsistency which embodies user's attention distribution more explicitly. So in this paper, we exploit Cross-Users Behaviors for viewport prediction in 360-degree video adaptive streaming, namely CUB360, trying to concurrently consider user's personalized information and cross-users behaviors information to predict future viewport. Besides, we use a QoE-driven framework to optimize existing video streaming approaches and propose a general algorithm aiming at solving the NP problem at a low complexity. Extensive experimental results over real datasets demonstrate that compared with traditional adaptive streaming method, our proposal can significantly boost the prediction accuracy by 20.2% absolutely and 48.1 % relatively. Besides, the mean quality can get 30.28% gain while quality variance can be reduced by 29.89%. Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yue Wang 0032 |
ICME | 4 |
| 2018 | Learning-based Congestion Control for Internet Video Communication over Wireless NetworksabstractWith the deployment of real-time video applications and wireless networks, the real-time congestion control becomes a hot topic. Most existing congestion control algorithms are not designed for low-latency real-time flows, or perform poorly in the face of highly variable channel capacities. In this paper, we proposed a novel Learning-based Congestion Control (LCC) for real-time video communication over wireless networks. The key idea of LCC is employing Kernel Density Estimation for one-way delay and sending rate to capture the underlying information about channel state. Then LCC bases on the estimated probability density and Bayesian theorem to quickly adapt sending rate to the changing channel. We implemented LCC in WebRTC framework and extensive experiments were carried out. Compared with the native WebRTC congestion control (GCC), experimental results show that LCC achieves higher channel utilization, even more than 4.2× throughput in lossy links. LCC is also much better at adapting to the variable channel than GCC. Besides, LCC performs well in delay constraint and intra-protocol fairness. Tongyu Dai, Xinggong Zhang, Zongming Guo |
ISCAS | 2 |
| 2018 | Probabilistic Viewport Adaptive Streaming for 360-degree VideosabstractRecently, there has been a significant interest towards 360-degree virtual reality (VR) video. However, it is a big challenge for them to stream over Internet for huge bit-rates. In this paper, we have designed a novel viewport adaptive streaming scheme for 360-degree videos with probabilistic viewport prediction and optimal segments prefetching by Dynamic Adaptive Streaming over HTTP (DASH). In this way, continuous and smooth video playback, low viewport prediction error and high PSNR are obtained. To avoid head-movement prediction error, a probabilistic viewport prediction model is proposed, which leverages the probability distribution of user's orientation. Further, an optimal segments prefetching method is implemented. Finally, we also implement our method in a real system. The numerous experiment results have demonstrated that the proposed method has achieved significant performance gains compared with the existing methods. Our related work also win the Runner-up in ICME 2017 DASH-IF Grand Challenge: Dynamic Adaptive Streaming over HTTP. Zhimin Xu 0001, Xinggong Zhang, Kai Zhang 0007, Zongming Guo |
ISCAS | 2 |
| 2018 | CLS: A Cross-user Learning based System for Improving QoE in 360-degree Video Adaptive StreamingabstractViewport adaptive streaming is emerging as a promising way to deliver high quality 360-degree video. It is still a critical issue to predict user's viewpoint and deliver partial video within the viewport. Current widely-used motion-based or content-saliency methods have low precision, especially for long-term prediction. In this paper, benefiting from data-driven learning, we propose a Cross-user Learning based System (CLS) to improve the precision of viewport prediction. Since users have similar region-of-interest (ROI) when watching a same video, it is possible to exploit cross-users' ROI behavior to predict viewport. We use a machine learning algorithm to group users according to historical fixations, and predict the viewing probability by the class. Additionally, we present a QoE-driven rate allocation to minimize the expected streaming distortion under bandwidth constraint, and give a Multiple-Choice Knapsack solution. Experiments demonstrate that CLS provides 2dB quality improvement than full-image streaming and 1.5 dB quality improvement than linear regression (LR) method. On average, the precision of viewpoint prediction improve 15% compared with LR. Lan Xie, Xinggong Zhang, Zongming Guo |
ACM Multimedia | 2 |
| 2017 | Dynamic threshold based rate adaptation for HTTP live streamingabstractThe Dynamic Adaptive Streaming over HTTP (DASH) is specified to cope with the changing network conditions and provide an adaptive bit-rate HTTP-based streaming solution. While there have been many researches of rate adaptation algorithms on adaptive HTTP streaming, much of the work is focused on Video on Demand (VoD) service - which is not same as live streaming. It is generally preferred to minimize the end-to-end delay and make full use of the bandwidth for live services. In this paper, we propose a buffer-based rate adaptation algorithm with dynamic threshold which can decrease the rate transitions and provide a seamless playback under a low latency requirement. The rate adaptation metrics not only take into account the momentary value of bandwidth but also consider its fluctuation as the recognition of bandwidth is crucial over small buffer. Experiments demonstrate that our proposed rate adaptation scheme outperforms the methods using fixed threshold or instant throughput. Lan Xie, Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
ISCAS | 3 |
| 2017 | A Caching Miss Ratio Aware Path Selection Algorithm for Information-Centric NetworksabstractIn Information-Centric Networks (ICN), contents are cached on some intermediary routers. This creates thus a new situation which is totally different from the traditional path-selection paradigm: the source/destination paradigm no longer exists; instead, the new paradigm is how to find a path through a selected group of caches, so that the content is delivered via the shortest way. This paper addresses this issue and proposes a path-selection algorithm taking into account both the caching capability of router and the more traditional link cost between routers. We formulated the problem as a convex optimization problem (named ESP) which aims to get expected shortest path (ESP) by minimizing the transportation cost. By applying the Lagrangian dual theorem, we solved the ESP problem and obtained a criterion for request (and reversely, data) routing. Based on this path-selection criterion, we provide a fully distributed distance-based ESP algorithm that enables routers maintain routes to nearest content, without knowing a network topology and the caching miss ratio of content at other routers. Simulations confirm the efficiency of our approach versus the traditional shortest path algorithm. Weihong Lin, Xinggong Zhang, Yu Guan 0005, Zongming Guo |
LCN | 2 |
| 2017 | 360ProbDASH: Improving QoE of 360 Video Streaming Using Tile-based HTTP Adaptive StreamingabstractRecently, there has been a significant interest towards 360-degree panorama video. However, such videos usually require extremely high bitrate which hinders their widely spread over the Internet. Tile-based viewport adaptive streaming is a promising way to deliver 360-degree video due to its on-request portion downloading. But it is not trivial for it to achieve good Quality of Experience (QoE) because Internet request-reply delay is usually much higher than motion-to-photon latency. In this paper, we leverage a probabilistic approach to pre-fetch tiles countering viewport prediction error, and design a QoE-driven viewport adaptation system, 360ProbDASH. It treats user's head movement as probability events, and constructs a probabilistic model to depict the distribution of viewport prediction error. A QoE-driven optimization framework is proposed to minimize total expected distortion of pre-fetched tiles. Besides, to smooth border effects of mixed-rate tiles, the spatial quality variance is also minimized. With the requirement of short-term viewport prediction under a small buffer, it applies a target-buffer-based rate adaptation algorithm to ensure continuous playback. We implement 360ProbDASH prototype and carry out extensive experiments on a simulation test-bed and real-world Internet with real user's head movement traces. The experimental results demonstrate that 360ProbDASH achieves at almost 39% gains on viewport PSNR, and 46% reduction on spatial quality variance against the existed viewport adaptation methods. Lan Xie, Zhimin Xu 0001, Yixuan Ban, Xinggong Zhang, Zongming Guo |
ACM Multimedia | 4 |
| 2017 | An optimal spatial-temporal smoothness approach for tile-based 360-degree video streamingabstractThe world is becoming more and more virtual than we ever thought it would be. Many video service providers have rolled out 360-degree videos which provide immersive experience to users. However, huge bandwidth occupation of 360-degree video hinders its wide spreading over the Internet. Besides, only part of the video is displayed on the screen, transmitting whole video results in waste of bandwidth and computational resources. Tile-based adaptive streaming is regarded as a bandwidth-friendly approach which only delivers specific portion of the whole video. It requires the clients to decide which portion and at which bitrates to deliver. However, due to both space and time partition of 360-degree videos in tile-based adaptive streaming, there still exists a challenge on the quality inconsistence on spatial and temporal domains. In this paper, we propose a optimal spatial-temporal smoothness approach under restricted network for tile-based adaptive streaming. The bitrates of tiles are determined optimally, aiming at maximizing the overall quality while minimizing the spatial and temporal quality variation. By conducting extensive experiments over real bandwidth dataset and user's head movement traces, our approach can get a significant improvement. Specifically, the Viewport-PSNR can be raised by 24.1% compared with traditional delivery of whole 360 video; while the spatial and temporal stability can be improved by 40.5% and 24.6% respectively compared with tile-based streaming. Yixuan Ban, Lan Xie, Zhimin Xu 0001, Xinggong Zhang, Zongming Guo, Yueyu Hu |
VCIP | 4 |
| 2015 | Unequal error protection for real-time video streaming using expanding window reed-solomon codeabstractExpanding Window FEC is an emerging scheme for robust real-time video streaming over wireless networks, with low latency and reduction of error propagation. In this work, we focus on the problem of Expanding Window FEC redundancy allocation which has not been adequately addressed in current works. We first analyse the error probability of the adopted Expanding Window Reed-Solomon code (EW-RS), and introduce an equivalent error probability to simplify them. Then we are able to formulate the optimal redundancy allocation into a constrained nonlinear optimization problem, where by allocating the redundancy unequally considering the unequal importance of different frames and their dependency based on the expanding window, unequal error protection (UEP) is achieved and the overall distortion is minimized. Moreover, to reduce the computation complexity, a high-efficiency hill-climbing algorithm is developed to obtain the suboptimal allocation. At last, the experimental results demonstrate the effectiveness of both the proposed allocation scheme and solution algorithm. Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo |
ICIP | 2 |
| 2015 | A fairness-aware smooth rate adaptation approach for dynamic HTTP streamingabstractRecently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed over the Internet. Under time-varying network conditions, it is, however, still a big challenge to provide smooth video bit-rate with high video quality, especially when multiple clients compete for the network resources where the fairness must be considered. In this paper, a fairness-aware smooth rate adaptation approach is designed for DASH under the scenario that multiple clients are competing for the network resources. To avoid the unfair bandwidth estimated by the client induced by the off-intervals during the downloading process, a probe-based bandwidth estimation method is designed which includes a logarithmic law based increase probing scheme and a conservative back-off based decrease probing scheme. Then, with the probed bandwidth, a dual-threshold based video bit-rate switching scheme is designed that buffer overflow/underflow is avoided, and smooth video bit-rate is also provided. The extensive experiments on our network testbed demonstrate that the proposed approach outperforms the existing schemes significantly. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
ICIP | 3 |
| 2015 | Delay-constrained rate control for real-time video streaming over wireless networksabstractRate control is a big challenge for real-time video streaming on the internet with the needs of low latency, bandwidth-consuming and stable video rate. However, most of the existing Internet congestion control protocols ignore these needs, and some of them use the packet loss event as congestion signal which is deviation especially over error-prone wireless networks. In this paper, we propose a delay-constrained rate control algorithm by locking queueing delay onto a desired objective. The shadow price of video rate is controlled by queueing delay. All flows adapt video rate according to distortion weight and shadow price so as to achieve a distributed bandwidth sharing with low latency, efficient utilization, and distortion fairness. A closed-loop rate control system is designed for the purpose of stable and agile control. The control parameters are analyzed using control-theoretic approach. Additionally, we construct a real-time wireless video streaming test-bed and conduct extensive experiments over it. Compared with the current widely used methods, the experimental results show that the proposed algorithm can achieve 3dB or more gains in PSNR, and better performance on bandwidth utilization, flow stability with well guaranteed multi-flow fairness. Yufeng Geng, Xinggong Zhang, Chao Zhou 0003, Zongming Guo |
VCIP | 2 |
| 2015 | A Novel JSCC Scheme for UEP-Based Scalable Video Transmission Over MIMO SystemsabstractIn this paper, we propose a novel joint source-channel coding (JSCC) scheme for scalable video transmission over multiple-input multiple-output (MIMO) systems. By exploiting the diversity of MIMO antennas and forward error correction (FEC)-based protection, our method aims to provide unequal error protection (UEP) for the video layers, which are mapped to appropriate antennas. Moreover, JSCC is also considered that we extract a proper subset of video layers and allocate suitable FEC redundancy to them. Jointly considering video layer extraction, FEC rate allocation, and video layer scheduling, we are able to achieve UEP so as to minimize end-to-end distortion. We formulate the scheme as a nonlinear integer optimization problem, which is known to be NP-hard. To find a near-optimal solution efficiently, we propose a low-complexity branch-and-bound algorithm, which partitions the original problem into a series of subproblems by a video layer branching technique. In each branch, the upper and lower distortion bounds are derived. In particular, we transform the video layer scheduling subproblem into a 0/1 multiple knapsack problem, which is NP-complete, and employ an evolutionary Lagrangian method to find a solution efficiently. For the FEC allocation subproblem, a Lagrange duality algorithm with fuzzy surrogate subgradient is proposed. The experimental results demonstrate that the proposed method has good efficiency while achieving close performance to the optimal results. Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2014 | Joint multi-CDN and LT-coding for video transport over HTTPabstractVideo transport over HTTP is becoming more and more popular. Many video service providers construct huge content distribution networks(CDN) to support HTTP streaming service, however, they seldom exploit the benefits of multiple servers to achieve higher bandwidth and reliability by parallel streaming. In this paper, we study the problem of jointing multi-CDN and LT-coding for video transport over HTTP. Using LT coding, a client could download the same video segment from multiple servers without considering data segmentation and server scheduling issue. Thus, we are able to treat all CDN servers as a virtual server with higher bandwidth and reliability. To reduce the ACK overhead, a stochastic model is designed to predict the amount of data to be sent from each server while guaranteeing the decoding probability. Compared with the existing schemes, the experimental results show that our proposed scheme obtains less overhead and fewer number of HTTP requests. Besides, we also achieve better video quality and better robustness to fluctuant bandwidth. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
ISCAS | 3 |
| 2014 | Probabilistic chunk scheduling approach in parallel multiple-server DASHabstractRecently parallel Dynamic Adaptive Streaming over HTTP (DASH) has emerged as a promising way to supply higher bandwidth, connection diversity and reliability. However, it is still a big challenge to download chunks sequentially in parallel DASH due to heterogeneous and time-varying bandwidth of multiple servers. In this paper, we propose a novel probabilistic chunk scheduling approach considering time-varying bandwidth. Video chunks are scheduled to the servers which consume the least time while with the highest probability to complete downloading before the deadline. The proposed approach is formulated as a constrained optimization problem with the objective to minimize the total downloading time. Using the probabilistic model of time-varying bandwidth, we first estimate the probability of successful downloading chunks before the playback deadline. Then we estimate the download time of chunks. A near-optimal solution algorithm is designed which schedules chunks to the servers with minimal downloading time while the completion probability is under the constraint. Compared with the existing schemes, the experimental results demonstrate that our proposed scheme greatly increases the number of chunks that are received orderly. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
VCIP | 3 |
| 2014 | A Control-Theoretic Approach to Rate Adaption for DASH Over Multiple Content Distribution ServersabstractRecently, dynamic adaptive streaming over HTTP (DASH) has been widely deployed on the Internet. However, the research about DASH over multiple content distribution servers (MCDS-DASH) is limited. Compared with traditional single-server DASH, MCDS-DASH is able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switching over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. In our method, multiple fragments are grouped into a block and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments download in an orderly way. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel proportional-derivative controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on our network testbed and the Internet demonstrate the good efficiency of the proposed method. Chao Zhou 0003, Chia-Wen Lin, Xinggong Zhang, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Adaptive channel scheduling for Scalable Video broadcasting over MIMO wireless networksabstractVideo broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions of users, it is challenging to minimize all users' video transmission distortion in MIMO broadcasting. In this paper, we investigate the channel scheduling problem, which maps Scalable Video layers to MIMO heterogenous channels to protect video layers unequally, so as to minimize overall received video distortion for all users. We formulate this problem into an integer non-linear optimization problem, which is hard to be solved. An efficient near-optimal algorithm is proposed, which is based on simulated-annealing theory. Experimental results demonstrate the efficiency of our proposed algorithm, and its performance is very close to the optimal results. Compared with the existing MIMO scheduling methods, the proposed scheme significantly improves the overall quality of video broadcasting in MIMO networks. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
ISCAS | 2 |
| 2013 | Topology-aware content-centric networkingabstractMaking data the first class entity, Information-Centric Networking (ICN) replaces conventional host-to-host model with content sharing model. However, the huge amount of content and the volatility of replicas cached across the Internet pose significant challenges for addressing content only by name. In this paper, we propose a topology-aware name-based routing protocol which combines the benefits of location-oriented routing and content-centric routing together. We adopt a URL-like naming scheme, which defines register locations and content identifier. Node with copies sends Register messages towards a register using location-oriented routing protocols. All en-path routers record forwarding entries in forwarding table (FIB) as the "bread crumb" to this content. Following the bread crumb, routers know the "best" topology path to the available copies. An Interest is either forwarded towards a "known" copy by the content identifier, or towards the register nodes where it would find the bread crumb to the "best" copies. Compared with the existing flooding or name resolution methods, Our design shows a good potential in terms of scalability, availability and overhead. Xinggong Zhang, Feng Lao, Zongming Guo |
SIGCOMM | 1 |
| 2013 | A control theory based rate adaption scheme for dash over multiple serversabstractRecently, Dynamic Adaptive Streaming over HTTP (DASH) has been widely deployed in the Internet. However, the research about DASH over Multiple Content Distribution Servers (MCDS) is few. Compared with traditional single-server-DASH, MCDS are able to offer expanded bandwidth, link diversity, and reliability. It is, however, a challenging problem to smooth video bitrate switchings over multiple servers due to their diverse bandwidths. In this paper, we propose a block-based rate adaptation method considering both the diverse bandwidths and feedback buffered video time. Multiple fragments are grouped into a block, and the fragments are downloaded in parallel from multiple servers. We propose to adapt video bitrate at the block level rather than at the fragment level. By dynamically adjusting the block length and scheduling fragment requests to multiple servers, the requested video bitrates from the multiple servers are synchronized, making the fragments downloaded orderly. Then, we propose a control-theoretic approach to select an appropriate bitrate for each block. By modeling and linearizing the rate adaption system, we propose a novel Proportional-Derivative (PD) controller to adapt video bitrate with high responsiveness and stability. Theoretical analysis and extensive experiments on the Internet demonstrate the good efficiency of our DASH designs. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
VCIP | 2 |
| 2013 | Optimal adaptive channel scheduling for scalable video broadcasting over MIMO wireless networksabstractVideo broadcasting is an efficient way to deliver video content to multiple receivers. However, due to heterogeneous channel conditions in MIMO wireless networks, it is challenging for video broadcasting to map scalable video layers to proper MIMO transmit antennas to minimize the average overall video transmission distortion. In this paper, we investigate the channel scheduling problem for broadcasting scalable video content over MIMO wireless networks. An adaptive channel scheduling based unequal error protection (UEP) video broadcasting scheme is proposed. In the scheme, video layers are protected unequally by being mapped to appropriate antennas, and the average overall distortion of all receivers is minimized. We formulate this scheme into a non-linear combinatorial optimization problem. It is not practical to solve the problem by an exhaustive search method with heavy computational complexity. Instead, an efficient branch-and-bound based channel scheduling algorithm, named TBCS, is developed. TBCS finds the global optimal solution with much lower complexity. The complexity is further reduced by relaxing the termination condition of TBCS, which produces a (1 − ε)-optimal solution. Experimental results demonstrate both the effectiveness and efficiency of our proposed scheme and algorithm. As compared with some existing channel scheduling methods, TBCS improves the quality of video broadcasting across all receivers significantly. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
Comput. Networks | 2 |
| 2013 | Modeling and Analysis of Skype Video Calls: Rate Control and Video QualityabstractVideo-conferencing has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video call, and analyze the network impacts in large-scale networks. We first measure the behaviors of Skype video call on a controlled network testbed. By varying packet loss rate, propagation delay and available network bandwidth, we observe how Skype adjusts its sending rate, FEC redundancy, video rate and frame rate. It is found that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that it employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype video calls. Extrapolating from the models, we conduct numerical analysis to study the network impacts. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded. Through a case study of a 4G wireless network, we demonstrate that the proposed models can be used in user-QoE-aware network provisioning. Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2012 | A novel JSCC scheme for scalable video transmission over MIMO systemsabstractMIMO recently emerges as one of promising techniques for wireless video streaming. It is still a challenge to provide un-equal error protections by joint source-channel coding (JSCC) over multiple diverse MIMO sub-channels. In this paper, a joint source-channel coding and antenna mapping scheme for scalable video transmission over MIMO systems is proposed. Bandwidth are elaborately allocated between video source and channel protections by layer extracting and FEC coding. For the extracted layers, we determine i) which antenna will they be transmitted over and ii) how much redundancy bits will be added for error protections. We formulate this scheme into a non-linear integer optimization problem, whose complexity is very high. Instead, a low-complexity branch-and-bound algorithm is presented. Source layers are partitioned into subsets of layers, and the selected layer are mapped to antennas using Min-max scheduling algorithm. By branching and pruning, the computation complexity are reduced significantly. We carry out extensive numerical experiments under various network conditions. The results demonstrate our algorithm's efficiency and the overall transmission quality is improved significantly. Xinggong Zhang, Chao Zhou 0003, Zongming Guo |
ICIP | 1 |
| 2012 | Cross-entropy based antenna selection for scalable video streaming over MIMO wireless networksabstractIn this paper, we investigate the antenna selection (AS) problem for scalable video streaming over MIMO wireless networks. By scheduling scalable video layers over MIMO antennas with different signal strength, the video layers are transmitted with un-equal error protections. Considering layer dependencies and various antenna conditions, it is a non-linear combinatorial problem for AS to minimize the overall end-to-end distortion. To find the optimal solution with low complexity, a cross-entropy based solution, named CEBAS, is proposed. All solutions are indexed by unique binary strings, and the primal problem is reformulated to a binary combination problem. Then, random strings are generated using the probability distribution of solutions, which is updated by the cross-entropy optimization method. The feasibility of solution is guaranteed by our proposed projection strategy. CEBAS is iterative in nature and converges to the global optimum in probability. Simulation results reveal both the effectiveness and efficiency of our proposed algorithm. When comparing CEBAS against other existing algorithms, consistent superior performance has been observed. Chao Zhou 0003, Xinggong Zhang, Zongming Guo |
ICIP | 2 |
| 2012 | Profiling Skype video calls: Rate control and video qualityabstractVideo telephony has recently gained its momentum and is widely adopted by end-consumers. But there have been very few studies on the network impacts of video calls and the user Quality-of-Experience (QoE) under different network conditions. In this paper, we study the rate control and video quality of Skype video calls. We first measure the behaviors of Skype video calls on a controlled network testbed. By varying packet loss rate, propagation delay and bandwidth, we observe how Skype adjusts its rates, FEC redundancy and video quality. We find that Skype is robust against mild packet losses and propagation delays, and can efficiently utilize the available network bandwidth. We also find that Skype employs an overly aggressive FEC protection strategy. Based on the measurement results, we develop rate control model, FEC model, and video quality model for Skype. Extrapolating from the models, we conduct numerical analysis to study the network impacts of Skype. We demonstrate that user back-offs upon quality degradation serve as an effective user-level rate control scheme. We also show that Skype video calls are indeed TCP-friendly and respond to congestion quickly when the network is overloaded. Xinggong Zhang, Yang Xu 0011, Yong Liu 0013, Zongming Guo, Yao Wang 0001 |
INFOCOM | 1 |
| 2012 | Sparsity estimation in image compressive sensingabstractCompressive sensing is an emerging technology which can recover a K-sparse signal vector from M = O(Klog(K=N)) measurements. However, it is a challenge to know exactly how many measurements an image requires to achieve an acceptable recovered visual quality. In this paper, we study the relationship between the image's complexity and its sparsity. We propose a mathematical model to estimate the number of needed measurements by using the image's texture, the edge density and the target reconstruction quality. There exists a linear function between them. The experimental results with a large number of photo pictures show that, quite most reconstructed images using our pre-calculated number of measurements have good enough quality, which confirms our proposed image-complexity-based model well. Shanzhen Lan, Xinggong Zhang, Zongming Guo |
ISCAS | 3 |
| 2012 | Parallelizing video transcoding using Map-Reduce-based cloud computingabstractDue to the complexity of video coding, fast transcoding is still a challenge. Various parallel coding methods have been proposed. In this paper, we present a parallel transcoding system over Map/Reduce cloud computing architecture. Input video sequences are divided into segments, and mapped to multiple computers. The sub-tasks are launched in parallel with processing results concatenated to the final output sequences. For heterogeneous clips, computing capacity, and task-launching overhead, the task scheduling over cloud is an NP-hard problem. We propose a low-complexity heuristic algorithm, Max-MCT, to find out the optimal solutions for task scheduling. By estimating the low-bound of finish time, we transform the problem into a virtual knapsack problem. But it is not an optimal solution for the original problem therefore we use a minimal complete time (MCT) algorithm to minimize the entire finish time. We carry out extensive experiments on numerical simulations. The results verified that our algorithm outperforms the existing algorithms. Feng Lao, Xinggong Zhang, Zongming Guo |
ISCAS | 2 |
| 2012 | A control-theoretic approach to rate adaptation for dynamic HTTP streamingabstractRecently, dynamic adaptive HTTP streaming has been widely used for video content delivery over Internet. However, it is still a challenge how to switch video bitrate under time-varying bandwidth. In this paper, we propose a novel control-theoretic approach to adapt video segments in dynamic HTTP streaming. The rate control is based on a sink-buffer, which has an overflow-threshold and an underflow-threshold. The objective is to maximize the playback quality while keeping the receiver buffer from either overflow or underflow. Using control theory, we formulate this rate control scheme as a proportional (P) control system, which exists oscillations and steady-errors. Furthermore, we design a proportional derivative (PD) controller to improve its adaptation performance. The conditions for stability and settling time of the PD controller are also derived. Numerous experiment results demonstrate the effectiveness of our proposed PD control scheme for dynamic HTTP streaming. Chao Zhou 0003, Xinggong Zhang, Longshe Huo, Zongming Guo |
VCIP | 2 |
| 2010 | Collision-detection based rate-adaptation for video multicasting over IEEE 802.11 wireless networksabstractWireless video multicasting/broadcasting is an efficient method for simultaneous transmission of data to a group of users. But the multicasting rates are fixed in current IEEE 802.11 PHYs standard. In this paper, we propose a novel collision-detection based rate-adaptation scheme (CDRA), which fully exploits the potential of rate adaptation capability of wireless physical layer, to improve service qualities of video multicasting. The received signal strength indication (RSSI) and packet error ratio (PER) are comprehensively used to detect collision. The PER-guided rate adjustment algorithm is performed when no collision happens. Otherwise the collision-avoid mechanism works. By detecting the collision, our scheme could adaptively select the maximum data rates for video multicasting. We construct a practical multicasting test-bed in IEEE 802.11b network and carry out extensive experiments. The results show that CDRA achieves throughput gain up to 166% and PSNR gain to 139% compared with existing methods. Chao Zhou 0003, Xinggong Zhang, Lichuan Lu, Zongming Guo |
ICIP | 2 |
| 2010 | Time-constrained packet scheduling optimization for video streaming in wireless ad-hoc networksabstractPacket schedule is effective to improve the quality of video streaming over time-varying wireless ad-hoc networks. In this paper, we propose a time-constrained packet scheduling algorithm, which minimizes the video distortion by allocating transmission opportunities to packets under the constraint of playback delay. The problem is formulated in a constrained convex optimization framework and solved with Lagrangian method. The packet loss probabilities and transmission time in IEEE802.11 wireless channel are predicted by using a Markov Chain model. The experiments in NS-2 simulator validate that the algorithm achieves a significant improvement on the quality of streaming. Xinggong Zhang, Zongming Guo |
ISCAS | 1 |
| 2009 | Rate-distortion based path selection for video streaming over wireless ad-hoc networksabstractIn wireless ad-hoc networks with low bandwidth, high bit error, and node mobility, the quality of streaming video is highly dependent on the quality of routing paths. Most of existing ad-hoc routing algorithms select the path according to network parameters. However, the selected path may not be the best path for video applications. This paper proposed a rate-distortion based (RD) path selection algorithm for video streaming over wireless networks. The video distortion on application layer is used as path metric. On each link, the rate-distortion due to transmission error and congestion is estimated using queuing theory. The algorithm selects the path with the minimum expected rate-distortion as the routing path. Extensive experiments are carried out over NS-2 simulation environment. Simulation results demonstrate that RD routing algorithm can improve the quality of video streaming significantly as compared to the conventional shortest-path routing algorithm. Xinggong Zhang, Zongming Guo |
ICME | 1 |