VLDB 2026 Research / reviewers in the wild / expert
Mengbai Xiao
dblp:160/4835
· DBLP profile ↗
48ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0002-6305-8125ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 8 since 2021Computer networks · 14 · 2 first-author · 9 since 2021Systems, architecture and hardware · 11 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | cGraph: A Compact and Efficient Graph-Based Index for Approximate Nearest Neighbor Search
Yu Liu 0085, Mengbai Xiao, Jing Qiao, Dongxiao Yu |
ICDCS | 2 |
| 2026 | PARS: Optimizing High-Dimensional Vector Search with Dimensional Reduction and Ray Tracing Core
Yiling Ma, Mengbai Xiao, Zhixiong Xiao, Yuan Yuan 0014, Dongxiao Yu, Feng Li 0002 |
ICDCS | 2 |
| 2026 | PACE: A Multi-Round Cell-Enhanced Prefetching Strategy in Volumetric Video StreamingabstractIn recent years, volumetric video streaming has emerged as a key application in the field of virtual reality (VR) and augmented reality (AR), attracting growing research and industry interest. Despite its potential, the extremely high bandwidth requirements of volumetric content far exceed the capacity of current networks to support full-resolution streaming. To address this limitation, industry practitioners often employ field-of-view (FoV) prediction to downscale the streaming content based on user gaze direction. Although this approach makes the streaming feasible, our empirical measurement reveals that, working with the FoV prediction, the commonly used sequential prefetching mechanism severely constrains streaming efficiency. To overcome this bottleneck, we introduce PACE, a smart prefetching strategy built on multi-round cell-level download scheduling. PACE divides the prefetching process of each group-of-frames (GoF) into several rounds, where distinct video cells are downloaded in each round according to periodically updated FoV predictions. This design decouples FoV prediction from the constraint of prefetch length, enabling more adaptive and responsive streaming. Moreover, PACE integrates a greedy quality selection policy that dynamically adjusts video quality to make full use of available bandwidth and enhance the overall quality of experience (QoE). Comprehensive evaluations show that PACE improves FoV prediction accuracy by up to 23.1% and boosts QoE by as much as 54.8%, while maintaining robust performance across diverse network and user scenarios. Shuquan Liu, Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | MOTA: Mixture of Traffic Agents for Robust Network Traffic ClassificationabstractNetwork traffic classification plays a crucial role in a wide range of applications, e.g., Quality of Service (QoS) enhancement, resource management, and network security. However, the widespread adoption of encryption protocols (e.g., SSL/TLS) and the emergence of anonymous communication systems (e.g., Tor) have introduced significant challenges due to the presence of complex and varied network noise. Although considerable effort has been made to improve the robustness of the traffic classification, the performances of existing state-of-theart methods cannot be guaranteed in the presence of a mixture of noises, and are not stable in different application scenarios. In this paper, we innovate in proposing a network traffic classification method based on MoA (Mixture of Agents), namely MOTA. By leveraging a light-weight MoA architecture, MOTA efficiently fine-tunes mainstream Large Language Models (LLMs) to adapt to different application scenarios of traffic classification, and fully exploits the collaboration of the LLMs to ensure the robustness against mixed noises. Our extensive experiments show that, the classification accuracy is$\geq 99 {\%}$across multiple public datasets injected with mixed noises, significantly outperforming existing SOTA methods. Moreover, despite incorporating multiple LLMs, MOTA maintains millisecond-level inference latency on a server equipped with four NVIDIA GeForce RTX 4090 GPUs, owing to its lightweight design. Shaowei Li, Zhiwen Gan, Mengbai Xiao, Pengfei Hu 0001, Xiuzhen Cheng, Feng Li 0002 |
IWQoS | 3 |
| 2025 | Temporal-Spatial Object Relations Modeling for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) is a challenging task where an agent is required to navigate to a natural language described location via vision observations. The navigation abilities of the agent can be enhanced by the relations between objects, which are usually learned using internal objects or external datasets. The relationships between internal objects are modeled employing graph convolutional network (GCN) in traditional studies. However, GCN tends to be shallow, limiting its modeling ability. To address this issue, we utilize a cross attention mechanism to learn the connections between objects over a trajectory, which takes temporal continuity into account, termed as Temporal Object Relations (TOR). The external datasets have a gap with the navigation environment, leading to inaccurate modeling of relations. To avoid this problem, we construct object connections based on observations from all viewpoints in the navigational environment, which ensures complete spatial coverage and eliminates the gap, called Spatial Object Relations (SOR). Additionally, we observe that agents may repeatedly visit the same location during navigation, significantly hindering their performance. For resolving this matter, we introduce the Turning Back Penalty (TBP) loss function, which penalizes the agent’s repetitive visiting behavior, substantially reducing the navigational distance. Experimental results on the REVERIE, SOON, Touchdown and R2R datasets demonstrate the effectiveness of the proposed method. Yanwei Zheng, Dongchen Sui, Chuanlin Lan, Xinpeng Zhao 0001, Xiao Zhang 0015, Jingke Meng, Mengbai Xiao, Yifei Zou, Dongxiao Yu |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | A Long-Term-Planning Learning Strategy to Coordinate Viewport Prediction and Video Transmission in 360° Video StreamingabstractFueled by Metaverse, 360° video streaming has seen tremendous growth in the past years. However, our measurement reveals that current 360° streaming systems suffer from a dilemma that severely limits QoE. On the one hand, viewport prediction requires the shortest possible prediction distance for high predicting accuracy; On the other hand, video transmission requires more buffered data to compensate for bandwidth fluctuations otherwise substantial playback rebuffering would be incurred. There is so far no existing method that can break this dilemma so the QoE optimization for 360° video streaming was naturally bottlenecked. This work is the first attempt to tackle this challenge by developing QUTA – a novel learning-based streaming system. Specifically, according to our measurement, three kinds of internal streaming parameters have significant impacts on the prediction distance, namely, download pause, data rate threshold, and playback rate. On top of this, we design a new long-term-planning (LTP) continuous control deep reinforcement learning method that tunes the parameters dynamically based on the network condition and the streaming context. Extensive evaluations based on real system prototypes show that QUTA not only improves the prediction accuracy and QoE performance by up to 68.4% but also exhibits strong temporal and spatial robustness. Mengbai Xiao, Dongxiao Yu, Vaneet Aggarwal, Xiuzhen Cheng |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | SLVS: A Self-Learning Approach to Achieve Near-Second Low-Latency Video Streaming Under Highly Variable NetworksabstractFueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and some DASH-based algorithms were specifically proposed for low-latency video delivery. We conducted a measurement study for the state-of-the-art algorithms with large-scale network traces. It reveals that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting in the playback latency that has substantial variations across highly fluctuating networks. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel self-learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Moreover, we further generalize SLVS to complement the streaming platform already in service to make it compatible with any live streaming services. Extensive evaluations based on real system prototypes show that SLVS can control playback latency down to 1 s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly fluctuating networks as well as long periods of video viewing. Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | A Novel Spatial-Temporal Learning Method for Enhancing Generalization in Adaptive Video StreamingabstractAdaptive video streaming has become a fundamental technology for video delivery. With the rise of deep reinforcement learning (DRL), streaming vendors are increasingly adopting DRL-driven adaptive bitrate (ABR) algorithms. In real-world deployments, most ABR approaches are developed with the aim of maintaining good performance across a wide variety of network environments. However, contrary to this expectation, our empirical findings show that even when trained on extensive real-world network trace data, these DRL-based ABR algorithms achieve only 43.1% to 48.9% of Quality-of-Experience (QoE) under highly diverse network conditions, which falls significantly short of the 100% optimum. We termed this problem as “ABR Under-Generalization”. To overcome this problem, we introduce BETA – a novel DRL-based ABR framework that incorporates both spatial and temporal learning mechanisms: 1) Spatially, BETA features a detector that flags the network conditions likely to cause poor performance, then trains specialized ABR models tailored for those conditions; 2) Temporally, BETA enhances its learning by incorporating multi-step decision experiences at each training epoch, enabling the trained model to account for long-term environmental dynamics. Comprehensive evaluations show that BETA outperforms state-of-the-art ABR algorithms, yielding average QoE gains of 19.4% to 50.9%, and achieving improvements of up to 244.1% under severely fluctuating network conditions. Huaren Wei, Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Action-Aware Visual-Textual Alignment for Long-Instruction Vision-and-Language NavigationabstractTraditional Vision-and-Language Navigation (VLN) requires an agent to navigate to a target location solely based on visual observations, guided by natural language instructions. Compared to this task, long-instruction VLN involves longer instructions, extended trajectories, and the need to consider more contextual information for global path planning. As a result, it is more challenging and requires accurately aligning the instructions with the agent’s current visual observations, which is accompanied by two significant issues. Firstly, there is a misalignment between actions. The visual observations of the agent at each step lack explicit action-related details, while the instructions contain action-oriented words. Secondly, there is a misalignment between global instructions and local visual observations. The instructions describe the entire navigation trajectory, whereas the agent’s visual observations only provide localized information about a specific position along the trajectory. To address these issues, this article introduces the Action-Perception Alignment Framework (APAF). In this framework, we first design the Action-Contextual Encoding Module (ACEM), which enriches the agent’s visual perception by encoding potential actions with relative heading and elevation angles. We then propose the Dynamic Instruction Weighting Module (DIWM), which adjusts the importance of instruction words based on the agent’s current visual observations, emphasizing those words most relevant to the agent’s visual observations. Our approach significantly outperforms existing methods, achieving state-of-the-art results with improvements of 8.5% and 4.0% in Success Rate (SR) on the long-instruction R4R and RxR datasets, respectively. Yanwei Zheng, Chuanlin Lan, Dongchen Sui, Xinpeng Zhao 0001, Xiao Zhang 0015, Mengbai Xiao, Dongxiao Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2025 | Enabling the Awareness of Video Perceived Quality for Short-Form Video StreamingabstractIn recent years, fueled by the rapid advances in high-speed mobile networks, streaming short-form videos over mobile devices (e.g., TikTok) has become ubiquitous among mobile users. Despite the widespread application, our investigation based on a real video data source revealed that a large proportion of short videos watched by viewers have suboptimal video quality (e.g., with low VMAF scores), which indicates that the Quality-of-Experience (QoE) is in fact far from optimal. This problem is primarily due to the lack of awareness of video quality optimization based on the features of the video content such as the scene complexity. To tackle this problem, this work develops a novel system called Quality Aware Short Video Streaming (QASVS), which adopts machine learning techniques to learn the video content features and then automatically generate quality-driven bitrate decision models to optimize the perceived video quality and QoE. Extensive evaluations show that QASVS is able to improve the video quality by 11.1%∼27.9% while significantly reducing the playback rebuffering compared to the state-of-the-art streaming algorithms. Therefore, QASVS is able to provide an effective way for streaming vendors to deliver high-performance short-video services. Mengbai Xiao, Hui Yuan 0001, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | TAMO:Fine-Grained Root Cause Analysis via Tool-Assisted LLM Agent With Multi-Modality Observation Data in Cloud-Native SystemsabstractImplementing large language models (LLMs)-driven root cause analysis (RCA) in cloud-native systems has become a key topic of modern software operations and maintenance. However, existing LLM-based approaches face three key challenges: multi-modality input constraint, context window limitation, and dynamic dependence graph. To address these issues, we propose a tool-assisted LLM agent with multi-modality observation data for fine-grained RCA, namely TAMO, including multi-modality alignment tool, root cause localization tool, and fault types classification tool. In detail, TAMO unifies multi-modal observation data into time-aligned representations for cross-modal feature consistency. Based on the unified representations, TAMO then invokes its specialized root cause localization tool and fault types classification tool for further identifying root cause and fault type underlying system context. This approach overcomes the limitations of LLMs in processing real-time raw observational data and dynamic service dependencies, guiding the model to generate repair strategies that align with system context through structured prompt design. Experiments on two benchmark datasets demonstrate that TAMO outperforms state-of-the-art (SOTA) approaches with comparable performance. Xiao Zhang 0015, Yuan Yuan 0040, Mengbai Xiao, Fuzhen Zhuang, Dongxiao Yu |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | patchDPCC: A Patchwise Deep Compression Framework for Dynamic Point CloudsabstractWhen compressing point clouds, point-based deep learning models operate points in a continuous space, which has a chance to minimize the geometric fidelity loss introduced by voxelization in preprocessing. But these methods could hardly scale to inputs with arbitrary points. Furthermore, the point cloud frames are individually compressed, failing the conventional wisdom of leveraging inter-frame similarity. In this work, we propose a patchwise compression framework called patchDPCC, which consists of a patch group generation module and a point-based compression model. Algorithms are developed to generate patches from different frames representing the same object, and more importantly, these patches are regulated to have the same number of points. We also incorporate a feature transfer module in the compression model, which refines the feature quality by exploiting the inter-frame similarity. Our model generates point-wise features for entropy coding, which guarantees the reconstruction speed. The evaluation on the MPEG 8i dataset shows that our method improves the compression ratio by 47.01% and 85.22% when compared to PCGCv2 and V-PCC with the same reconstruction quality, which is 9% and 16% better than that D-DPCC does. Our method also achieves the fastest decoding speed among the learning-based compression models. Zirui Pan, Mengbai Xiao, Xu Han 0016, Dongxiao Yu, Yao Liu 0001 |
AAAI | 2 |
| 2024 | UltraPrecise: A GPU-Based Framework for Arbitrary-Precision Arithmetic in Database SystemsabstractFixed-point decimal operations in databases with arbitrary-precision arithmetic refer to the ability to store and operate decimal fraction numbers with an arbitrary length of digits. This type of operation has become a requirement for many applications, including scientific databases, financial data processing, geometric data processing, and cryptography. However, the state-of-the-art fixed-point decimal technology either provides high performance for low-precision operations or supports arbitrary-precision arithmetic operations at low performance. In this paper, we present a design and implementation of a framework called UltraPrecise which supports arbitrary-precision arithmetic for databases on GPU, aiming to gain high performance for arbitrary-precision arithmetic operations. We build our framework based on the just-in-time compilation technique and optimize its performance via data representation design, PTX acceleration, and expression scheduling. UltraPrecise achieves comparable performance to other high-performance databases for low-precision arithmetic operations. For high-precision, we show that UltraPrecise consistently outperforms existing databases by two orders of magnitude, including workloads of RSA encryption and trigonometric function approximation. Xin Li 0078, Mengbai Xiao, Dongxiao Yu, Rubao Lee, Xiaodong Zhang 0001 |
ICDE | 2 |
| 2024 | RoIRTC: Toward Region-of-Interest Reinforced Real-Time Video CommunicationabstractIn this paper, we propose a region-of-interest (RoI) reinforced real-time communication system, RoIRTC, for improving the quality of videos delivered in real-time communication. RoIRTC uses a novel RoI magnification transformation for spatially adapting the camera-captured video frame. To automatically detect the RoI, it intelligently leverages a deep-learning-based saliency prediction model without affecting the video collector’s processing throughput or the encoder’s efficiency. Evaluation results based on actual remote learning videos show that RoIRTC that performs RoI magnification can improve the median PSNR by 2.6 dB compared to the naive WebRTC implementation. Compared to an approach that mimics the "background blur" scheme used in many real-time communication systems, RoIRTC can also improve the median PSNR by 4.2 dB. Shuoqian Wang, Mengbai Xiao, Yao Liu 0001 |
ICME | 2 |
| 2024 | LiteQUIC: Improving QoE of Video Streams by Reducing CPU Overhead of QUICabstractQUIC is the underlying protocol of the next generation HTTP/3, serving as the major vehicle delivering video data nowadays. As a userspace protocol based on UDP, QUIC features low transmission latency and has been widely deployed by content providers. However, the high computational overhead of QUIC shifts system knobs to CPUs in high-bandwidth scenarios. When CPU resources become the constraint, HTTP/3 exhibits even lower throughput than HTTP/1.1. In this paper, we carefully analyze the performance bottleneck of QUIC and find it results from ACK processing, packet sending, and data encryption. By reducing the ACK frequency, activating UDP generic segmentation offload (GSO), and incorporating PicoTLS, a high-performance encryption library, the CPU overhead of QUIC could be effectively reduced in stable network environments. However, simply reducing the ACK frequency also impairs the transmission throughput of QUIC under poor network conditions. To solve this, we develop LiteQUIC, which involves two mechanisms towards alleviating the overhead of ACK processing in addition to GSO and PicoTLS. We evaluate LiteQUIC in the DASH-based video streaming, and the results show that LiteQUIC achieves 1.2× higher average bitrate and 93.3% lower rebuffering time than an optimized version of QUIC with GSO and PicoTLS. Pengqiang Bi, Yifei Zou, Mengbai Xiao, Dongxiao Yu, Qun Xie |
ACM Multimedia | 3 |
| 2024 | An Intelligent Prefetch Strategy with Multi-round Cell Enhancement in Volumetric Video StreamingabstractOver the past few years, volumetric video streaming, a cutting-edge application in virtual reality (VR) and augmented reality (AR), has gained significant attention. However, its enormous bandwidth demands exceed the capabilities of current networks to support full-size transmission. As a result, streaming vendors in the industry typically use field-of-view (FoV) prediction to reduce the streaming video size. While this approach makes transmission feasible, our measurements have shown that the adopted sequential video prefetch mechanism significantly hinders streaming performance. To address this challenge, we propose PACE, an intelligent video prefetch strategy that leverages multi-round cell-enhanced downloading. PACE divides the prefetching process for each group-of-frame (GoF) into multiple rounds, downloading different cells in each round based on periodically updated FoV prediction results. This method allows the FoV prediction to operate independently of the prefetch length limitations. Additionally, PACE incorporates a greedy-policy-based video quality decision algorithm to fully utilize network bandwidth and maximize the quality of experience (QoE). Extensive evaluations demonstrate that PACE can enhance FoV prediction accuracy by up to 23.1% and improve QoE performance by up to 54.8% while showing strong robustness across a wide range of streaming environments. Shuquan Liu, Mengbai Xiao, Dongxiao Yu, Xiuzhen Cheng |
SECON | 3 |
| 2024 | BETA: A Novel Learning-Based Adaptive Streaming Approach with Spatial and Temporal OptimizationabstractAdaptive video streaming (DASH) has become a key technology in video transmission. Given the advantages of deep re-inforcement learning (DRL), streaming vendors are increasingly focusing on DRL-based adaptive bitrate (ABR) algorithms. In practice, almost all the ABR algorithms are designed with the in-tention to work well in any size/shape of networks. However, far from expectations, our measurements revealed that, even with extensive training on a vast scale of real network trace data, if the network condition varies over a wide range, the achieved Quality-of-Experience (QoE) is only 43.1% ~ 48.9%, far below the optimal 100%. We termed this problem as “ABR Under-Generalization”. To address this problem, we developed BETA in this work, a novel ABR-specialized DRL-based approach that integrates spatial and temporal modules. 1) Spatial: BETA offers a detector that identi-fies the potential network condition that may cause poor performance and then trains complementary ABR algorithms specifically functioning in it; 2) Temporal: At each training epoch, BETA learns from the decision experience of multiple-step segments to cover long-term environmental feedback. Extensive evaluations demonstrate that BETA achieves an average QoE improvement of 19.4% to 50.9% compared to the state-of-the-art ABR algorithms, with gains of up to 244.1% in the challenging networks with dramatically variable throughput. Mengbai Xiao, Dongxiao Yu, Xiuzhen Cheng |
SECON | 3 |
| 2024 | DeepReal: Short-form Video Streaming with Fine-grained Bitrate AdaptationabstractShort-form video streaming has seen explosive growth in recent years. Typically, the short video apps (e.g., TikTok) are installed on mobile devices, with the video content delivered over wireless networks. Our measurement study shows that almost all the state-of-the-art adaptive bitrate (ABR) algorithms for short video streaming fall short of providing satisfactory Quality-of-Experience (QoE). This is due to their reliance on bandwidth-sensitive adaptation logic with over-discretized encoding bitrate levels, which hampers their ability to manage significant wireless bandwidth fluctuations. To tackle this problem, we propose DeepReal, a fine-grained ABR approach for short videos, the core of which is a novel clustering-augmented continuous control learning algorithm. On the one hand, the clustering offers more expertise for the ABR logic training. On the other hand, the continuous-valued action domain enables the trained ABR logic to support any number of discrete bitrate levels. As a result, DeepReal's bitrate decision is able to effectively adapt to the substantial network dynamics. Large-scale evaluations show that DeepReal improves QoE by 23.5% to 33.3% while maintaining strong spatial and temporal robustness. It offers a tailored and innovative solution for current short-form video services. Mengbai Xiao, Dongxiao Yu, Xiuzhen Cheng |
SECON | 3 |
| 2024 | Reordering and Compression for Hypergraph ProcessingabstractHypergraphs are applicable to various domains such as social contagion, online groups, and protein structures due to their effective modeling of multivariate relationships. However, the increasing size of hypergraphs has led to high computation costs, necessitating efficient acceleration strategies. Existing approaches often require consideration of algorithm-specific issues, making them difficult to directly apply to arbitrary hypergraph processing tasks. In this paper, we propose a compression-array acceleration strategy involving hypergraph reordering to improve memory access efficiency, which can be applied to various hypergraph processing tasks without considering the algorithm itself. We introduce a new metric called closeness to optimize the ordering of vertices and hyperedges in the one-dimensional array representation. Moreover, we present an$\frac{1}{2w}$-approximation algorithm to obtain the optimal ordering of vertices and hyperedges. We also develop an efficient update mechanism for dynamic hypergraphs. Our extensive experiments demonstrate significant improvements in hypergraph processing performance, reduced cache misses, and reduced memory footprint. Furthermore, our method can be integrated into existing hypergraph processing frameworks, such as Hygra, to enhance their performance. Yu Liu 0085, Mengbai Xiao, Dongxiao Yu, Huashan Chen, Xiuzhen Cheng |
IEEE Trans. Computers | 3 |
| 2024 | A GPU-Enabled Real-Time Framework for Compressing and Rendering Volumetric VideosabstractNowadays, volumetric videos have emerged as an attractive multimedia application providing highly immersive watching experiences since viewers could adjust their viewports at 6 degrees-of-freedom. However, the point cloud frames composing the video are prohibitively large, and effective compression techniques should be developed. There are two classes of compression methods. One suggests exploiting the conventional video codecs (2D-based methods) and the other proposes to compress the points in 3D space directly (3D-based methods). Though the 3D-based methods feature fast coding speeds, their compression ratios are low since the failure of leveraging inter-frame redundancy. To resolve this problem, we design a patch-wise compression framework working in the 3D space. Specifically, we search rigid moves of patches via the iterative closest point algorithm and construct a common geometric structure, which is followed by color compensation. We implement our decoder on a GPU platform so that real-time decoding and rendering are realized. We compare our method with GROOT, the state-of-the-art 3D-based compression method, and it reduces the bitrate by up to 5.98$\times$. Moreover, by trimming invisible content, our scheme achieves comparable bandwidth demand of V-PCC, the representative 2D-based method, in FoV-adaptive streaming. Dongxiao Yu, Ruopeng Chen, Xin Li 0078, Mengbai Xiao, Yao Liu 0001 |
IEEE Trans. Computers | 4 |
| 2024 | Value of Information: A Comprehensive Metric for Client Selection in Federated Edge LearningabstractFederated edge learning (FEEL) is a novel paradigm that enables privacy-preserving and distributed machine learning on end devices. However, FEEL faces challenges from data/system heterogeneity among the participating clients and resource constraints of edge networks, which affect the efficiency and accuracy of the learning process. In this paper, we propose a comprehensive framework for client selection in FEEL based on the concept of Value-of-Information (VoI), which measures how valuable a client is for the global model aggregation. Our framework consists of two independent components: a VoI estimator that uses reinforcement learning to learn the relationship between VoI and various heterogeneous factors of clients; and a greedy client selector that chooses the most valuable clients under network resource constraints. Compared with most of the previous works that use concrete criteria to evaluate and select heterogeneous clients, our VoI-based approach is more comprehensive. Extensive experiments on different datasets and learning tasks are conducted, which show that our framework outperforms several state-of-the-art methods in terms of accuracy. Yifei Zou, Shikun Shen, Mengbai Xiao, Peng Li 0017, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Computers | 3 |
| 2023 | VQBA: Visual-Quality-Driven Bit Allocation for Low-Latency Point Cloud StreamingabstractVideo-based Point Cloud Compression (V-PCC) is an emerging standard for encoding dynamic point cloud data. With V-PCC, point cloud data is segmented, projected, and packed on to 2D video frames, which can be compressed using existing video coding standards such as H.264, H.265 and AV1. This makes it possible to support point cloud streaming via reliable video transmission systems. On the other hand, despite recent advances, many issues still remain and prevent V-PCC from being used in low-latency point cloud streaming. For instance, point cloud registration and patch generation can take a long time. Shuoqian Wang, Mufeng Zhu, Na Li 0032, Mengbai Xiao, Yao Liu 0001 |
ACM Multimedia | 4 |
| 2023 | An Intelligent Learning Approach to Achieve Near-Second Low-Latency Live Video Streaming under Highly Fluctuating NetworksabstractFueled by the rapid advances in high-speed mobile networks, live video streaming has seen explosive growth in recent years and many DASH-based bitrate adaptive streaming algorithms were specifically proposed for low-latency video delivery. However, our investigations revealed that these algorithms are susceptible to network condition changes due to the use of solo universal adaptation logics, resulting the playback latency that has substantial variations across highly-fluctuating network environments and fails to meet the service quality requirement all the time. To tackle this challenge, this paper proposes Stateful Live Video Streaming (SLVS), which is a novel learning approach that learns the various network features and optimizes the adaptation logic separately for different network conditions, then dynamically tunes the logic at runtime, so that bitrate decision can better match the changing networks. Extensive evaluations show that SLVS can control playback latency down to 1s while improving Quality-of-Experience (QoE) by 17.7% to 31.8%. Moreover, it has strong robustness to maintain near-second latency over highly-fluctuating networks as well as long-period of video viewing. Ke Liu 0004, Mengbai Xiao, Bingshu Wang, Vaneet Aggarwal |
ACM Multimedia | 3 |
| 2023 | patchVVC: A Real-time Compression Framework for Streaming Volumetric VideosabstractNowadays, volumetric video has emerged as an attractive multimedia application, which provides highly immersive watching experiences. However, streaming the volumetric video demands prohibitively high bandwidth. Thus, effectively compressing its underlying point cloud frames is essential to deploying the volumetric videos. The existing compression techniques are either 3D-based or 2D-based, but they still have drawbacks when being deployed in practice. The 2D-based methods compress the videos in an effective but slow manner, while the 3D-based methods feature high coding speeds but low compression ratios. In this paper, we propose patchVVC, a 3D-based compression framework that reaches both a high compression ratio and a real-time decoding speed. More importantly, patchVVC is designed based on point cloud patches, which makes it friendly to an field of view adaptive streaming system that further reduces the bandwidth demands. The evaluation shows patchVCC achieves the real-time decoding speed and the comparable compression ratios as the representative 2D-based scheme, V-PCC, in an FoV-adaptive streaming scenario. Ruopeng Chen, Mengbai Xiao, Dongxiao Yu, Yao Liu 0001 |
MMSys | 2 |
| 2023 | oBBR: Optimize Retransmissions of BBR Flows on the Internet
Pengqiang Bi, Mengbai Xiao, Dongxiao Yu |
USENIX ATC | 2 |
| 2022 | Maze: A Cost-Efficient Video Deduplication System at Web-scaleabstractWith the advancement and dominant service of Internet videos, the content-based video deduplication system becomes an essential and dependent infrastructure for Internet video service. However, the explosively growing video data on the Internet challenges the system design and implementation for its scalability in several ways. (1) Although the quantization-based indexing techniques are effective for searching visual features at a large scale, the costly re-training over the complete dataset must be done periodically. (2) The high-dimensional vectors for visual features demand increasingly large SSD space, degrading I/O performance. (3) Videos crawled from the Internet are diverse, and visually similar videos are not necessarily the duplicates, increasing deduplication complexity. (4) Most videos are edited ones. The duplicate contents are more likely discovered as clips inside the videos, demanding processing techniques with close attention to details. An Qin 0001, Mengbai Xiao, Ben Huang, Xiaodong Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Mind the Gap: Broken Promises of CPU Reservations in Containerized Multi-tenant CloudsabstractContainerization is becoming increasingly popular, but unfortunately, containers often fail to deliver the anticipated performance with the allocated resources. In this paper, we first demonstrate the performance variance and degradation are significant (by up to 5x) in a multi-tenant environment where containers are co-located. We then investigate the root cause of such performance degradation. Contrary to the common belief that such degradation is caused by resource contention and interference, we find that there is a gap between the amount of CPU a container reserves and actually gets. The root cause lies in the design choices of today's Linux scheduling mechanism, which we call Forced Runqueue Sharing and Phantom CPU Time. In fact, there are fundamental conflicts between the need to reserve CPU resources and Completely Fair Scheduler's work-conserving nature, and this contradiction prevents a container from fully utilizing its requested CPU resources. As a proof-of-concept, we implement a new resource configuration mechanism atop the widely used Kubernetes and Linux to demonstrate its potential benefits and shed light on future scheduler redesign. Our proof-of-concept, compared to the existing scheduler, improves the performance of both batch and interactive containerized apps by up to 5.6x and 13.7x. Li Liu 0045, An Wang 0002, Mengbai Xiao, Yue Cheng 0001, Songqing Chen |
SoCC | 4 |
| 2021 | NestGPU: Nested Query Processing on GPUabstractNested queries are commonly used to express complex use-cases by connecting the output of a subquery as an input to the outer query block. However, their execution is highly time-consuming. Researchers have proposed various algorithms and techniques that unnest subqueries to improve performance. Since this is a customized approach that needs high algorithmic and engineering efforts, it is largely not an open feature in most existing database systems.Our approach is general-purpose and GPU-acceleration based, aiming for high performance at a minimum development cost. We look into the major differences between nested and unnested query structures to identify their merits and limits for GPU processing. Furthermore, we focus on the nested approach that is algorithmically simple and rich in parallels, in relatively low space complexity, and generic in program structure. We create a new code generation framework that best fits GPU for the nested method. We also make several critical system optimizations including massive parallel scanning with indexing, effective vectorization to optimize join operations, exploiting cache locality for loops and efficient GPU memory management. We have implemented the proposed solutions in NestGPU, a GPU-based column-store database system that is GPU device independent. We have extensively evaluated and tested the system to show the effectiveness of our proposed methods. Sofoklis Floratos, Mengbai Xiao, Hao Wang 0002, Chengxin Guo, Yuan Yuan 0014, Rubao Lee, Xiaodong Zhang 0001 |
ICDE | 2 |
| 2021 | Mixer: Efficiently Understanding and Retrieving Visual Content at Web-ScaleabstractVisual contents, including images and videos, are dominant on the Internet today. The conventional search engine is mainly designed for textual documents, which must be extended to process and manage increasingly high volumes of visual data objects. In this paper, we present Mixer, an effective system to identify and analyze visual contents and to extract their features for data retrievals, aiming at addressing two critical issues: (1) efficiently and timely understanding visual contents, (2) retrieving them at high precision and recall rates without impairing the performance. In Mixer, the visual objects are categorized into different classes, each of which has representative visual features. Subsystems for model production and model execution are developed. Two retrieval layers are designed and implemented for images and videos, respectively. In this way, we are able to perform aggregation retrievals of the two types in efficient ways. The experiments with Baidu's production workloads and systems show that Mixer halves the model production time and raises the feature production throughput by 9.14x. Mixer also achieves the precision and recall of video retrievals at 95% and 97%, respectively. Mixer has been in its daily operations, which makes the search engine highly scalable for visual contents at a low cost. Having observed productivity improvement of upper-level applications in the search engine, we believe our system framework would generally benefit other data processing applications. Mengbai Xiao, An Qin 0001, Yongwei Wu 0002, Xinjie Huang, Xiaodong Zhang 0001 |
Proc. VLDB Endow. | 1 |
| 2020 | GPU-Accelerated Computation of Vietoris-Rips Persistence BarcodesabstractThe computation of Vietoris-Rips persistence barcodes is both execution-intensive and memory-intensive. In this paper, we study the computational structure of Vietoris-Rips persistence barcodes, and identify several unique mathematical properties and algorithmic opportunities with connections to the GPU. Mathematically and empirically, we look into the properties of apparent pairs, which are independently identifiable persistence pairs comprising up to 99% of persistence pairs. We give theoretical upper and lower bounds of the apparent pair rate and model the average case. We also design massively parallel algorithms to take advantage of the very large number of simplices that can be processed independently of each other. Having identified these opportunities, we develop a GPU-accelerated software for computing Vietoris-Rips persistence barcodes, called Ripser++. The software achieves up to 30x speedup over the total execution time of the original Ripser and also reduces CPU-memory usage by up to 2.0x. We believe our GPU-acceleration based efforts open a new chapter for the advancement of topological data analysis in the post-Moore’s Law era. Simon Zhang, Mengbai Xiao, Hao Wang 0002 |
SoCG | 2 |
| 2020 | SphericRTC: A System for Content-Adaptive Real-Time 360-Degree Video CommunicationabstractWe present the SphericRTC system for real-time 360-degree video communication. 360-degree video allows the viewer to observe the environment in any direction from the camera location. This more-immersive streaming experience allows users to more-efficiently exchange information and can be beneficial in the real-time setting. Our system applies a novel approach to select representations of 360-degree frames to allow efficient, content-adaptive delivery. The system performs joint content and bitrate adaptation in real-time by offloading expensive transformation operations to the GPU via CUDA. The system demonstrates that the multiple sub-components -- viewport feedback, representation selection, and joint content and bitrate adaptation -- can be effectively integrated within a single framework. Compared to a baseline implementation, views in SphericRTC have consistently higher visual quality. The median Viewport-PSNR of such views is 2.25 dB higher than views in the baseline system. Shuoqian Wang, Mengbai Xiao, Kenneth Chiu, Yao Liu 0001 |
ACM Multimedia | 3 |
| 2020 | AdaP-360: User-Adaptive Area-of-Focus Projections for Bandwidth-Efficient 360-Degree Video Streamingabstract360-degree video is an emerging medium that presents an immersive view of the environment to the user. Despite its potential to provide an immersive watching experience, 360-degree video has not achieved widespread popularity. A significant cause of this slow adoption is the high-bandwidth requirements of the format. The primary source of bandwidth inefficiency in 360-degree video streaming, un-addressed in popular transmission methods, is the discrepancy between the pixels sent over the network (typically the full omnidirectional view) and the pixels displayed in the head-mounted display's field of view. At worst, roughly 88% of transmitted pixels remain unviewed. Chao Zhou 0004, Shuoqian Wang, Mengbai Xiao, Sheng Wei 0001, Yao Liu 0001 |
ACM Multimedia | 3 |
| 2019 | Catfish: Adaptive RDMA-enabled R-Tree for Low Latency and High ThroughputabstractR-tree is a foundational data structure used in spatial databases and scientific databases. With the advancement of Internet and computer architectures, in-memory data processing for R-tree in distributed systems has become a common platform. We have observed new performance challenges to process R-tree as the amount of multidimensional datasets become increasingly huge. Specifically, an R-tree server can be heavily overloaded while the network and client CPU are lightly loaded, and vice versa. In this paper, we present the design and implementation of Catfish, an RDMA enabled R-tree for low latency and high throughput by adaptively utilizing the available network bandwidth and computing resources to balance the workloads between clients and servers. We design and implement two basic mechanisms of using RDMA for the client-server R-tree. First, in the fast messaging design, we use RDMA writes to send R-tree requests to the server and let server threads process R-tree requests to achieve low query latency. Second, in the RDMA offloading design, we use RDMA reads to offload tree traversal from the server to the client, which rescues the server as it is overloaded. We further develop an adaptive scheme to effectively switch an R-tree search between fast messaging and RDMA offloading, maximizing the overall performance. Our experiments show that the adaptive solution of Catfish on InfiniBand significantly outperforms R-tree that uses only fast messaging or only RDMA offloading in both latency and throughput. Catfish can also deliver up to one order of magnitude performance over the traditional schemes using TCP/IP on 1 Gbps and 40 Gbps Ethernet. We make a strong case to use RDMA to effectively balance workloads in distributed systems for low latency and high throughput. Mengbai Xiao, Hao Wang 0002, Liang Geng, Rubao Lee, Xiaodong Zhang 0001 |
ICDCS | 1 |
| 2019 | DirectLoad: A Fast Web-Scale Index System Across Large Regional CentersabstractThe freshness of web page indices is the key to improving searching quality of search engines. In Baidu, the major search engine in China, we have developed DirectLoad, an index updating system for efficiently delivering the webscale indices to nationwide data centers. However, the web-scale index updating suffers from increasingly high data volumes during network transmission and inefficient I/O transactions due to slow disk operations. DirectLoad accelerates the index updating streams from two aspects: 1) DirectLoad effectively cuts down the overwhelmingly high volume of indices in transmission by removing the redundant data across versions, and mutates regular operations in a key-value storage system for successful accesses to the deduplicated datasets. 2) DirectLoad significantly improves the I/O efficiency by replacing the LSMTree with a memory-resident table (memtable) and appendingonly-files (AOFs) on disk. Specifically, the write amplification stemming from sorting operations on disk is eliminated, and a lazy garbage collection policy further improves the I/O performance at the software level. In addition, DirectLoad directly manipulates the SSD native interfaces to remove the write amplification at the hardware level. In practice, 63% updating bandwidth has been saved due to the deduplication, and the write throughput to SSDs is increased by 3x. The index updating cycle of our production workloads has been compressed from 15 days to 3 days after deploying DirectLoad. In this paper, we show the effectiveness and efficiency of an in-memory index updating system, which is disruptive to the framework in a conventional memory hierarchy. We hope that this work contributes a strong case study in the system research literature. An Qin 0001, Mengbai Xiao, Dai Tan, Rubao Lee, Xiaodong Zhang 0001 |
ICDE | 2 |
| 2019 | HYPHA: a framework based on separation of parallelisms to accelerate persistent homology matrix reductionabstractPersistent homology (PH) matrix reduction is an important tool for data analytics in many application areas. Due to its highly irregular execution patterns in computation, it is challenging to gain high efficiency in parallel processing for increasingly large data sets. Simon Zhang, Mengbai Xiao, Chengxin Guo, Liang Geng, Hao Wang 0002, Xiaodong Zhang 0001 |
ICS | 2 |
| 2019 | Companion Paper forabstractThis artifact includes source code, scripts and datasets required to reproduce the experimental figures in the evaluation of the MM'18 paper, which is entitled "MiniView Layout for Bandwidth-Efficient 360-Degree Video". The artifact reports the comparison results among the standard cube layout (CUBE), the equi-angular layout (EAC), and the MiniView layout (MVL) in terms of compressed video size, visual quality of views and decoding and rendering time. Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen, Lucile Sassatelli, Gwendal Simon |
ACM Multimedia | 1 |
| 2019 | vCPU as a container: towards accurate CPU allocation for VMsabstractWith our increasing reliance on cloud computing, accurate resource allocation of virtual machines (or domains) in the cloud have become more and more important. However, the current design of hypervisors (or virtual machine monitors) fails to accurately allocate resources to the domains in the virtualized environment. In this paper, we claim the root cause is that the protection scope is erroneously used as the resource scope for a domain in the current virtualization design. Such design flaw prevents the hypervisor from accurately accounting resource consumption of each domain. In this paper, using virtual CPUs as a container we propose to redefine the resource scope of a domain, so that the new resource scope is aligned with all the CPU consumption incurred by this domain. As a demonstration, we implement a novel system, called VASE (vCPU as a container), on top of the Xen hypervisor. Evaluations on our testbed have shown our proposed approach is effective in accounting system-wide CPU consumption incurred by domains, while introducing negligible overhead to the system. Li Liu 0045, An Wang 0002, Mengbai Xiao, Yue Cheng 0001, Songqing Chen |
VEE | 4 |
| 2018 | BAS-360°: Exploring Spatial and Temporal Adaptability in 360-degree Videos over HTTP/2abstractToday, 360-degree video streaming has become a popular Internet service with the rise of affordable virtual reality (VR) technologies. However, streaming 360-degree videos suffers from the prohibitive bandwidth demand. Existing bandwidth-efficient solutions mainly focus on exploiting the inherent spatial adaptability of 360-degree videos, delivering only video content (spatially-cut tiles) in the viewer's region of interest (ROI) with higher quality. Temporal adaptability, which has been widely leveraged in HTTP streaming, has not been well exploited to select proper quality for video segments according to the bandwidth variations. When these two dimensions of adaptability are jointly considered, bitrate selection for the tiles become more complicated and challenging. The importance of a tile with a spatial coordination played at a specific time should be quantified so that we can determine how to allocate bandwidth for improving the viewer's quality of experience. Furthermore, viewer's head orientation prediction is highly variable, which makes the determination of important tiles highly dynamic. In addition, network fluctuations are very common on the Internet. To overcome these challenges, we propose Bi-Adaptive Streaming for 360-degree videos (BAS-360°). In BAS-360°, both spatial and temporal adaptabilities are explored in the bitrate selection for different tiles. The objective is to minimize the bandwidth waste by allocating bandwidth to more important tiles (the tiles that are more likely to be watched). To tackle the high variability of visual region prediction and the unpredictable network fluctuations, we employ two features provided by HTT P /2: stream termination and stream priority, to efficiently organize tile delivery. Evaluation results show that BAS-360° outperforms naive tile-based 360-degree video streaming strategies when network fluctuations or errors in viewport predictions occur. Mengbai Xiao, Chao Zhou 0004, Viswanathan (Vishy) Swaminathan, Yao Liu 0001, Songqing Chen |
INFOCOM | 1 |
| 2018 | ClusTile: Toward Minimizing Bandwidth in 360-degree Video Streamingabstract360-degree video has the potential to transform the video streaming experience by providing a more-immersive environment for users to interact with than standard streaming video. This experience is hampered, however, by high bandwidth requirements resulting from the extra information associated with the 360-degree frames. Because users cannot see this full 360-degree view, but the full view is transmitted in the majority of 360-degree streaming systems, there is much potential to reduce wasted bandwidth in this domain. We propose ClusTile, a tiling approach formulated to select a set of tiles that allows minimal bandwidth needed to be used when streaming 360-degree video over an expected set of views. These tiles are selected by solving a set of integer linear programs (ILPs) independently on clusters of collected user views. The clustering approach reduces computation requirements of the ILPs to practical levels. Tilings computed from ClusTile can save up to 76% bandwidth compared to standard 360-degree streaming and up to 52% bandwidth compared to best-performing fixed tiling schemes. Chao Zhou 0004, Mengbai Xiao, Yao Liu 0001 |
INFOCOM | 2 |
| 2018 | MiniView Layout for Bandwidth-Efficient 360-Degree VideoabstractWith the recent increase in popularity of VR devices, 360-degree video has become increasingly popular. As more users experience this new medium, it will likely see further increases in popularity as users experience its greater immersiveness compared to traditional video streams. 360-degree video streams must encode the omnidirectional view, and, with current encoding techniques, these views require significantly higher bandwidth than traditional video streams. These larger bandwidth requirements comprise the main barrier toward wider adoption by video streaming services. Mengbai Xiao, Shuoqian Wang, Chao Zhou 0004, Li Liu 0045, Zhenhua Li 0001, Yao Liu 0001, Songqing Chen |
ACM Multimedia | 1 |
| 2017 | OpTile: Toward Optimal Tiling in 360-degree Video Streamingabstract360-degree videos are encoded for adaptive streaming by first projecting the spherical surface onto two-dimensional frames, then encoding these as standard video segments. During playback of these 360-degree videos, the video player renders the portion of the spherical surface in the direction of the user's view. These user viewports typically cover only a small portion of the 360 degree surface, causing much of the downloaded bandwidth to be wasted. Tile-based approaches can reduce the wasted bandwidth by cutting video spatially into motion-constrained rectangles. Streaming logic then only needs to download the tiles necessary to render the viewport seen by the user. Existing tile-based approaches cut 360-degree videos into tiles of fixed sizes. These fixed-size tiling approaches, however, suffer from reduced encoding efficiency. Tiling cuts away portions of the video that can be copied by the encoder from adjacent frames or within the current frame that are needed for effective video compression. Mengbai Xiao, Chao Zhou 0004, Yao Liu 0001, Songqing Chen |
ACM Multimedia | 1 |
| 2016 | DASH2M: Exploring HTTP/2 for Internet Streaming to Mobile DevicesabstractToday HTTP/1.1 is the most popular vehicle for delivering Internet content, including streaming video. Standardized in 2015 with a few new features, HTTP/2 is gradually replacing HTTP 1.1 to improve user experience. Yet, how HTTP/2 can help improve the video streaming delivery has not been thoroughly investigated. In this work, we set to investigate how to utilize the new features offered by HTTP/2 for video streaming over the Internet, focusing on the streaming delivery to mobile devices as, today, more and more users watch video on their mobile devices. For this purpose, we design DASH2M, Dynamic Adaptive Streaming over HTTP/2 to Mobile Devices. DASH2M deliberately schedules the streaming content delivery by comprehensively considering the user's Quality of Experience (QoE), the dynamics of the network resources, and the power efficiency on the mobile devices. Experiments based on an implemented prototype show that DASH2M can outperform prior strategies for users' QoE while minimizing the battery power consumption on mobile devices. Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen |
ACM Multimedia | 1 |
| 2016 | Evaluating and improving push based video streaming with HTTP/2abstractThe sever-initiated push mechanism is one of the most prominent features in the next generation HTTP/2 protocol, having shown its capability on saving network traffic and improving the web page retrieval latency. Our prior work has investigated the server push-based mechanism for HTTP video streaming and proposed a k-push scheme, where the server pushes k video segments following the response to a request. In this study, we further conduct an analysis and evaluation of the k-push scheme in HTTP streaming. Our results uncover that the push mechanism can efficiently increase the network utilization (under certain conditions) compared to regular HTTP streaming. However the results also show that the k-push scheme deteriorates network adaptability and leads to the "over-push" problem, in which the pushed video content waste network resources due to user abandonment behaviors. To overcome these limitations, we propose a new " adaptive-push" scheme, which dynamically adjusts the parameter k to adapt to the runtime environment. To evaluate the performance of adaptive-push, we implemented a prototype system. The experimental results show that compared to k-push, adaptive-push can improve the network adaptability. Furthermore, our real-world trace based simulation results show that adaptive-push can effectively alleviate the over-push problem. Mengbai Xiao, Viswanathan (Vishy) Swaminathan, Sheng Wei 0001, Songqing Chen |
NOSSDAV | 1 |
| 2016 | GoCAD: GPU-Assisted Online Content-Adaptive Display Power Saving for Mobile Devices in Internet StreamingabstractDuring Internet streaming, a significant portion of the battery power is always consumed by the display panel on mobile devices. To reduce the display power consumption, backlight scaling, a scheme that intelligently dims the backlight has been proposed. To maintain perceived video appearance in backlight scaling, a computationally intensive luminance compensation process is required. However, this step, if performed by the CPU as existing schemes suggest, could easily offset the power savings gained from backlight scaling. Furthermore, computing the optimal backlight scaling values requires per-frame luminance information, which is typically too energy intensive for mobile devices to compute. Thus, existing schemes require such information to be available in advance. And such an offline approach makes these schemes impractical. To address these challenges, in this paper, we design and implement GoCAD, a GPU-assisted Online Content-Adaptive Display power saving scheme for mobile devices in Internet streaming sessions. In GoCAD, we employ the mobile device's GPU rather than the CPU to reduce power consumption during the luminance compensation phase. Furthermore, we compute the optimal backlight scaling values for small batches of video frames in an online fashion using a dynamic programming algorithm. Lastly, we make novel use of the widely available video storyboard, a pre-computed set of thumbnails associated with a video, to intelligently decide whether or not to apply our backlight scaling scheme for a given video. For example, when the GPU power consumption would offset the savings from dimming the backlight, no backlight scaling is conducted. To evaluate the performance of GoCAD, we implement a prototype within an Android application and use a Monsoon power monitor to measure the real power consumption. Experiments are conducted on more than 460 randomly selected YouTube videos. Results show that GoCAD can effectively produce power savings without affecting rendered video quality. Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen |
WWW | 2 |
| 2016 | Content-Adaptive Display Power Saving for Internet Video Applications on Mobile DevicesabstractBacklight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for mobile video applications, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the Central Processing Unit (CPU), could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices. In this article, we propose Content-Adaptive Display (CAD) for two typical Internet mobile video applications: video streaming and real-time video communication. CAD uses the mobile device’s Graphics Processing Unit (GPU) rather than the CPU to perform luminance compensation at reduced power consumption. For video streaming where video frames are available in advance, we compute the backlight scaling schedule using a more efficient dynamic programming algorithm than existing work. For real-time video communication where video frames are generated on the fly, we propose a greedy algorithm to determine the backlight scaling at runtime. We implement CAD in one video streaming application and one real-time video call application on the Android platform and use a Monsoon power meter to measure the real power consumption. Experiment results show that CAD can save more than 10% overall power consumption for up to 55.7% videos during video streaming and up to 31.0% overall power consumption in real-time video calls. Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Lei Guo 0004, Songqing Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2015 | Reducing display power consumption for real-time video calls on mobile devicesabstractThe display subsystem of a mobile device usually consumes 38%-68% [1] of the total battery power in video streaming. Therefore, a few schemes have been designed to reduce the display power consumption. The basic idea is to dim the backlight level while properly compensating the pixel luminance to maintain image fidelity. The luminance compensation and proper backlight level calculation are computation intensive and demand per-frame luminance information. For these reasons, existing schemes only work for video-on-demand where each frame (and thus the luminance information) is available in advance. In addition, they demand additional computing resource support. Otherwise, if the computation is conducted on the mobile device, the power consumption due to such computation can easily offset the power savings from dimming the backlight. In this work, we set to investigate power saving for real-time video calls on mobile devices. Different from video-on-demand, real-time video calls are highly delay sensitive and the frame luminance information is not known in advance. Moreover, video calls often involve multiple streaming sources from multiple (≥2) participants, making it more difficult. Because there are few background changes and the frame rate is usually small in video calls, we design a Greedy Display Power saving scheme, called LCD-GDP, which utilizes the commonly available GPU on mobile devices without demanding additional support. Our design is implemented on WebRTC, a popular real-time web browser based video call standard. Experiments show that our scheme can save up to 33% power consumption in video calls without affecting the video call quality. Mengbai Xiao, Yao Liu 0001, Lei Guo 0004, Songqing Chen |
ISLPED | 1 |
| 2015 | Power efficient mobile video streaming using HTTP/2 server pushabstractThis paper proposes a power efficient video streaming mechanism on mobile devices over cellular networks. We first develop an analytical model to identify and quantify the power inefficiency in mobile video streaming, due to the mismatch between HTTP request schedule and the radio resource control schedule. Based on the analytical model, we develop a low power video streaming mechanism by employing the server push technology available in the HTTP/2 protocol. We implemented the server push-based low power streaming mechanism in an HTTP DASH video streaming prototype involving mobile devices and the 4G/LTE cellular network. Our experiments show significant battery power savings on mobile devices using our server push strategy. Sheng Wei 0001, Viswanathan (Vishy) Swaminathan, Mengbai Xiao |
MMSP | 3 |
| 2015 | Content-adaptive display power saving in internet mobile streamingabstractBacklight scaling is a technique proposed to reduce the display panel power consumption by strategically dimming the backlight. However, for Internet streaming to mobile devices, a computationally intensive luminance compensation step must be performed in combination with backlight scaling to maintain the perceived appearance of video frames. This step, if done by the CPU, could easily offset the power savings via backlight dimming. Furthermore, computing the backlight scaling values requires per-frame luminance information, which is typically too energy intensive to compute on mobile devices. Yao Liu 0001, Mengbai Xiao, Xin Li 0078, Mian Dong, Zhenhua Li 0001, Songqing Chen |
NOSSDAV | 2 |