Cong Zhang 0002

dblp:18/2908-2 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-9439-6725ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 27 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Semantic-Aware Edge-Cloud Collaboration for Cost-Efficient Video Understanding
Cong Zhang 0002, Danyang Song, Handi Chen, Edith C. H. Ngai, Jiangchuan Liu, Victor C. M. Leung
ICDCS2
2026 ACPGS: Towards Bandwidth-Efficient Delivery of 3D Gaussian Splatting
Cong Zhang 0002, Jianxin Shi 0005, Xiaoyi Fan 0001, Laizhong Cui, Jiangchuan Liu
NOSSDAV2
2026 DRLLMS: Network-Adaptive Reasoning Control for Interactive LLM Streaming
Tao Lyu 0005, Cong Zhang 0002, Haihan Duan, Xiaoyi Fan 0001, Xiping Hu, Laizhong Cui
NOSSDAV2
2026 IVA: Proactive Bitrate Orchestration for Multiparty Video Conferencing via Conversational Intent
abstract
Multimodal large language models (MLLMs) are increasingly integrated into video conferencing, but mostly for speech-centric tasks such as transcription and summarization. Conferencing adaptation remains dominated by signal-driven congestion control that reacts to bandwidth changes without understanding why certain moments or streams will soon become quality-critical. This separation wastes predictive structure in conversation: text often reveals impending role shifts and interaction-mode transitions—for example, taking the floor or initiating screen sharing—seconds before the corresponding media and bandwidth demands materialize.
Cong Zhang 0002, Edith C. H. Ngai, Jiangchuan Liu, Bo Li 0001, Baochun Li
NOSSDAV3
2026 PivotSketch: Control-Ready Semantic Ranking for Adaptive Video Streaming
Sirui Zhang, Cong Zhang 0002, Xiaoyi Fan 0001, Xiping Hu, Haihan Duan
NOSSDAV3
2025 Blockchain-Enabled Market Clearing Mechanism for Peer-to-Peer Energy Storage Sharing
Haihan Duan, Hengming Dai, Xiaoyi Fan 0001, Cong Zhang 0002, Xiping Hu
IEEE Big Data5
2025 Sparse Manifold Retrieval Network for ICESat-2 Photon Point Cloud Denoising
abstract
The photon point clouds acquired by ICESat-2/ATLAS offer unprecedented potential for Earth observation but are heavily contaminated by noise photons, posing a significant challenge for downstream applications. Traditional denoising methods, which often rely on local density statistics, struggle with complex terrains and varying signal-to-noise ratios. While deep learning presents a promising alternative, existing approaches often inefficiently process the inherently sparse data via 2D projections or non-optimized 3D networks. To address these limitations, this paper introduces a novel deep learning framework for ICESat-2 photon denoising, termed Sparse Manifold Retrieval Network (SMRNet). We propose a Manifold-Aware Convolution (MAC) module to capture the continuous manifold structures of signal photons through multi-scale dilated sparse convolutions, and a Cross-Scale Pyramid Enhancement (CSPE) module to effectively refine multi-level features extracted from the encoder. Evaluated on a manually annotated dataset covering southeastern coastal regions of China, SMRNet demonstrates superior performance over traditional denoising method and data-driven baselines across multiple metrics. The results underscore the effectiveness of SMRNet in enhancing denoising accuracy, particularly in challenging environments with sparse signals and rugged topography.
Hengming Dai, Haihan Duan, Cong Zhang 0002, Xiaoyi Fan 0001, Zhifang Zhao
CloudCom3
2025 Enhancing Tail NFT Recommendation via Dependency-Aware Extreme Multi-Label Learning
abstract
With the rise of Web3, Non-Fungible Tokens (NFTs) have become a new class of digital assets, driving demand for large-scale NFT recommendation systems. Each NFT can be associated to a rich set of semantic, stylistic, and thematic labels, forming a highly complex label space. Similar to e-commerce platforms where detailed product labels enable personalized recommendations, such semantic dependencies between labels can potentially enhance NFT recommendation performance. Thus, NFT recommendation can be naturally formulated as an extreme multi-label (XML) classification problem. Many existing probabilistic label tree (PLT)-based approaches address XML problem by recursively partitioning the label space, which greatly alleviates the demands on expensive computer resources. Yet, the highly skewed distribution of labels in datasets in XML makes tail labels more challenging to predict than head labels. In this paper, Our preliminary analysis reveals that inherent label dependencies can be leveraged to improve tail label recommendations for NFTs. We propose ChainTail, a dependency-aware framework that enhances PLT-based NFT label partitioning and prediction re-scoring. It includes: (1) a Dependency-aware partition module that partitions highly dependent NFT labels into subsets. (2) a Dependency-aware ReScore module that re-ranks prediction scores of labels to eliminate the label-priors. Our experimental results show that ChainTail boosts tail label recommendation on widely used item recommendation datasets.
Cong Zhang 0002, Feng Wang 0001, Edith C. H. Ngai
CloudCom2
2025 Blockchain-Enabled Pricing Mechanism in Energy Markets: Survey and Vision
abstract
The growth of distributed energy resources and local energy markets heightens the need for price formation that is transparent, privacy preserving, and compatible with network constraints. Blockchain provides a trust-minimized substrate for auditable clearing and settlement through consensus, tamperevident ledgers, and smart contracts. This survey organizes blockchain-enabled pricing into three families, namely auction-based, game-theoretic, and optimization-based, and links them to enabling techniques such as metering oracles, secure multiparty computation, zero-knowledge proofs, and verifiable optimality certificates. Applications span wholesale electricity, carbon and green certificates, distributed energy trading, ancillary services, and electric vehicles. Evidence indicates gains in auditability, privacy, network awareness, and automated settlement, alongside challenges in scalability, data protection, grid integration, and regulation. The survey distills design patterns and research directions toward verifiable, interoperable, and governable pricing modules that complement system-operator markets.
Xiaoyi Fan 0001, Cong Zhang 0002, Hengming Dai, Haihan Duan
CloudCom3
2025 SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have greatly improved 3D reconstruction. However, its substantial data size poses a significant challenge for transmission and storage. While many compression techniques have been proposed, they fail to efficiently adapt to fluctuating network bandwidth, leading to resource wastage. We address this issue from the perspective of size-aware compression, where we aim to compress 3DGS to a desired size by quickly searching for suitable hyperparameters. Through a measurement study, we identify key hyperparameters that affect the size - namely, the reserve ratio of Gaussians and bit-width settings for Gaussian attributes. Then, we formulate this hyperparameter optimization problem as a mixed-integer nonlinear programming (MINLP) problem, with the goal of maximizing visual quality while respecting the size budget constraint. To solve the MINLP, we decouple this problem into two parts: discretely sampling the reserve ratio and determining the bit-width settings using integer linear programming (ILP). To solve the ILP more quickly and accurately, we design a quality loss estimator and a calibrated size estimator, as well as implement a CUDA kernel. Extensive experiments on multiple 3DGS variants demonstrate that our method achieves state-of-the-art performance in post-training compression. Furthermore, our method can achieve comparable quality to leading training-required methods after fine-tuning.
Shuzhao Xie, Weixiang Zhang, Shijia Ge, Sicheng Pan, Yunpeng Bai, Cong Zhang 0002, Xiaoyi Fan 0001, Zhi Wang 0001
ACM Multimedia8
2025 DD-LIVM: Pioneering Cross-Domain Photovoltaic Defect Detection Using Large Infrared-Visible Model
abstract
Photovoltaic (PV) defect detection is crucial for preventing power efficiency loss and fire hazards. The industry primarily relies on the fusion of infrared and visible images for defect localization and diagnosis. However, current detection methods exhibit poor generalizability in new site environments or with altered imaging setups. While recent infrared and vision foundation models (FM) facilitate domain-invariant feature maps extraction, directly concatenating them and fine-tuning achieves limited generalizability gain to PV defect detection, due to the asymmetric dual-modal semantics of defects. In this paper, we present the first large infrared-visible model DD-LIVM to enable cross-domain defect detection. The key innovation of DD-LIVM lies in its defect-specific three-step fine-tuning strategy, which utilizes alternating modality masking. Prior to feature fusion and joint fine-tuning, the infrared and visible FM encoders are alternately masked and optimized to enhance their individual semantic utility for defect localization visibility and classification granularity, with feature distances among different defect types regulated through contrastive learning. This approach allows for the extraction of generalizable and defect-specific feature maps. Moreover, for practical employment of DD-LIVM, we propose a domain-agnostic spatial alignment algorithm for infrared-visible images before dual-modal fusion, and develop source data augmentation and adaptive detection head selection schemes based on defects' infrared characteristics to further enhance the generalizability. Extensive experiments on 7,078 dual-modal images from 9 real-world scenarios across 4 cities' PV stations demonstrate that DD-LIVM achieves an accuracy of 87.7% for cross-domain defect detection, surpassing state-of-the-art methods by 17.3%.
Yinan Zhu, Meng Xue 0001, Haiyan Hu 0003, Cong Zhang 0002, Xiaoyi Fan 0001, Qian Zhang 0001
MobiCom4
2025 Optimizing Mobile-Friendly Viewport Prediction for Live 360-Degree Video Streaming
abstract
Viewport prediction is the crucial task for adaptive 360-degree video streaming, as the bitrate control algorithms usually require the knowledge of the user's viewing portions of the frames. Various methods are studied and adopted for viewport prediction from less accurate statistic tools to highly calibrated deep neural networks. Conventionally, it is difficult to implement sophisticated deep learning methods on mobile devices, which have limited computation capability. In this work, we propose an advanced learning-based viewport prediction approach and carefully design it to minimize transmission and computation overhead for mobile terminals. To improve viewport prediction accuracy, we utilize both spatial information through a saliency prediction model and temporal information through a modified LSTM model. Different computations introduced by the neural network models are distributed across the network to keep the computation light on mobile devices. To better adapt to the content dynamics in live streaming, we employ the model-agnostic meta-learning (MAML) method for video saliency prediction. The learned saliency prediction model with optimized initialization via offline meta-training can be fast fine-tuned online using a few samples. We further discuss how to integrate this mobile-friendly viewport prediction (MFVP) approach into a typical 360-degree video live streaming system by formulating and solving the bitrate adaptation problem. Extensive experiment results demonstrate that our approach achieves real-time prediction for live video streaming and surpasses existing methods in prediction accuracy on mobile terminals, which, together with our bitrate adaptation algorithm, significantly improves the streaming QoE from various aspects. Compared to baseline methods, MFVP achieves a 4.7–28.7% improvement in accuracy and demonstrates faster adaptability to dynamic content changes, enabling rapid fine-tuning and adjustment. When integrated into a streaming system and paired with our adaptive bitrate allocation algorithm, MFVP enhances overall video quality by 5.6–12.9% and reduces quality fluctuations by 33.3–50.9%.
Lei Zhang 0066, Peng Chen 0041, Cong Zhang 0002, Tao Long 0002, Weizhen Xu, Laizhong Cui, Jiangchuan Liu
IEEE Trans. Mob. Comput.3
2024 Towards Integrated Energy-Communication-Transportation Hub: A Base-Station-Centric Design in 5G and Beyond
abstract
The rise of 5G communication has transformed the telecom industry for critical applications. With the widespread deployment of 5G base stations comes a significant concern about energy consumption. Key industrial players have recently shown strong interest in incorporating energy storage systems to store excess energy during off-peak hours, reducing costs and partic-ipating in demand response. The fast development of batteries opens up new possibilities, such as the transportation area. An effective method is needed to maximize base station battery utilization and reduce operating costs. In this trend towards next-generation smart and integrated energy-communication-transportation (ECT) infrastructure, base stations are believed to play a key role as service hubs. By exploring the overlap between base station distribution and electric vehicle charging infrastructure, we demonstrate the feasibility of efficiently charging EVs using base station batteries and renewable power plants at the Hub. Our model considers various factors, including base station traffic conditions, weather, and EV charging behavior. This paper introduces an incentive mechanism for setting charging prices and employs a deep reinforcement learning-based method for battery scheduling. Experimental results demonstrate the effectiveness of our proposed ECT-Hub in optimizing surplus energy utilization and reducing operating costs, particularly through revenue-generating EV charging.
Linfeng Shen, Guanzhen Wu, Cong Zhang 0002, Xiaoyi Fan 0001, Jiangchuan Liu
ICDCS3
2024 SCALM: Towards Semantic Caching for Automated Chat Services with Large Language Models
abstract
Large Language Models (LLMs) have become increasingly popular, transforming a wide range of applications across various domains. However, the real-world effectiveness of their query cache systems has not been thoroughly investigated. In this work, we for the first time conducted an analysis on real-world human-to-LLM interaction data, identifying key challenges in existing caching solutions for LLM-based chat services. Our findings reveal that current caching methods fail to leverage semantic connections, leading to inefficient cache performance and extra token costs. To address these issues, we propose SCALM, a new cache architecture that emphasizes semantic analysis and identifies significant cache entries and patterns. We also detail the implementations of the corresponding cache storage and eviction strategies. Our evaluations show that SCALM increases cache hit ratios and reduces operational costs for LLMChat services. Compared with other state-of-the-art solutions in GPTCache, SCALM shows, on average, a relative increase of 63% in cache hit ratio and a relative improvement of 77% in tokens savings.
Jiaxing Li 0006, Chi Xu 0004, Feng Wang 0001, Isaac M. von Riedemann, Cong Zhang 0002, Jiangchuan Liu
IWQoS5
2024 Every Little Bit Helps: A Semantic-aware Tail Label Understanding Framework
abstract
The rapid expansion of AI technology, driven by high-speed networks and high-performance mobile devices, enables personalized and cross-content recommendations across diverse applications. Recent methods consider a large number of textual content keywords or topics as labels for recommendations, transforming the problem into Extreme Multi-label Learning (XML). However, addressing the XML problem in AI-driven recommendation systems that handle extensive user-generated data while ensuring Quality of Service (QoS) faces two main challenges: significant computational costs and inferior tail label prediction performance. We propose SAT, a semantic-aware framework with a tree architecture that effectively tackles these challenges and demonstrates improved performance compared to well-established approaches.
Feng Wang 0001, Cong Zhang 0002, Jiaxing Li 0006, Edith C. H. Ngai, Jiangchuan Liu
IWQoS3
2024 You Only Look Once in Panorama: Object Detection for 360° Videos with MLaaS
abstract
360° videos are gaining popularity, but immersive analytics, particularly in object detection, confront challenges from complex scenes and high data volume. This imposes significant burdens on individual users and resource-limited edge devices. Fortunately, Machine Learning as a Service (MLaaS) offers an economical solution for quick deployment without specific hardware or expertise. However, current MLaaS are mostly 2D image-designated and not optimized for the distinctive characteristics of raw 360° video frames. In this paper, we propose a novel MLaaS-based system to address this challenge. Our solution partitions 360° frames into distortion-free 2D regions with dynamic region of interest prediction. We then present an image-stitching algorithm featuring Skyline representation, seamlessly combining all the 2D regions into a unified frame. This frame is then transmitted to the MLaaS platform, with the detected objects being back-projected to yield the final results. Our experiments demonstrate the superiority of this system over baselines, proving its effectiveness in 360° video object detection tasks.
Linfeng Shen, Miao Zhang 0003, Cong Zhang 0002, Jiangchuan Liu
NOSSDAV3
2024 HALO: HVAC Load Forecasting With Industrial IoT and Local-Global-Scale Transformer
abstract
The evolution of Internet-of-Things (IoT) is fostering the use of intelligent controls for energy conservation. Yet, the efficacy of these strategies is largely tied to diverse load forecasting algorithms. Given the significant contribution of heating, ventilation, and air-conditioning (HVAC) systems to global energy consumption, accurate forecasting of HVAC power usage is crucial for improving overall energy efficiency. However, real-world HVAC load forecasting, bolstered by various IoT devices, is complicated by multiple factors: data variability, power load fluctuations, electronic phenomena (e.g., zero drifts), and the increased time complexity and larger model sizes required to manage accumulating historical data. To address these challenges, we first present an in-depth measurement study on the characteristics of HVAC load at a minute scale based on HVAC data collected in six locations. We propose HALO, a transformer-based framework specifically designed for forecasting HVAC load. HALO incorporates an adaptive data pre-processing stage and a local-global-scale transformer-based load forecasting stage, enabling precise forecasting of HVAC load and optimization of energy utilization. Evaluation based on real-world data traces from a prototype application demonstrates that the proposed framework significantly outperforms existing models.
Cong Zhang 0002, Edith C. H. Ngai, Jiangchuan Liu, Bo Li 0001
IEEE Internet Things J.2
2023 A Low Cost Cross-Platform Video/Image Process Framework Empowers Heterogeneous Edge Application
abstract
Recently, video/image intelligent analytics has been widely used in industrial Artificial Intelligence (AI) applications, such as defect detection, face recognition, and security monitoring. To provide better applicability and compatibility in such applications, the embedded AI models must be developed, compiled, and deployed under different development frameworks, such as cuDNN, RKNN, etc. Unfortunately, these frameworks are supported by various Graphic Processing Unit (GPU) hardware vendors, resulting in different model parameter structures and increased development costs. To address these issues, we propose LiGo, a low cost cross-platform video/image process framework, that simplifies and accelerates video intelligent processing in practical heterogeneous hardware systems. LiGo1 provides video processing pipeline, cross-platform development environments, and unified model serving structures. We demonstrate LiGo's efficiency and flexibility in model generation and deployment through its use in supporting multiple real-world commercial industrial systems.
Danyang Song, Cong Zhang 0002, Yifei Zhu 0001, Jiangchuan Liu
NOSSDAV2
2023 Optimal Volumetric Video Streaming With Hybrid Saliency Based Tiling
abstract
Volumetric video enables a six-degree-of-freedom (6DoF) immersive viewing experience and has a wide range of applications in entertainment and education, among others. Most existing approaches to volumetric video streaming are extensions of VR video streaming solutions that do not take into account user behavior and the properties of the video during the tiling process, and the complexity of decoding is high. To this end, we study volumetric video streaming in this paper and address the research questions mentioned above. In particular, we first propose a hybrid visual saliency and hierarchical clustering empowered 3D tiling scheme that better matches the user’s field of view (FoV). Then, we build a quality of experience (QoE) model considering the volumetric video features as the optimization objective. In addition to the usual encoded version, we introduce the reconstructed version (i.e., decoded version, which allows the user to skip the decoding process and thus reduces the decoding overhead) and propose a joint computational and communication resource allocation scheme to achieve a trade-off between communication and computational resources to maximize the QoE. We perform exhaustive simulations and build a prototype system to verify the performance of the proposed tiling and transmission scheme. The results show that the proposed tiling and transmission scheme performs significantly better than the comparison schemes.
Jie Li 0015, Cong Zhang 0002, Zhi Liu 0002, Richang Hong, Han Hu 0003
IEEE Trans. Multim.2
2020 Joint Communication and Computational Resource Allocation for QoE-driven Point Cloud Video Streaming
abstract
Point cloud video is the most popular representation of hologram, which is the medium to precedent natural content in VR/AR/MR and is expected to be the next generation video. Point cloud video system provides users immersive viewing experience with six degrees of freedom (6DoF) and has wide applications in many fields such as online education and entertainment. To further enhance these applications, point cloud video streaming is in critical demand. The inherent challenges lie in the large size by the necessity of recording the three-dimensional coordinates besides color information, and the associated high computation complexity of encoding/decoding. To this end, this paper proposes a communication and computational resource allocation scheme for QoE-driven point cloud video streaming. In particular, with the goal to maximize the defined QoE by selecting proper quality levels (uncompressed tiles at different quality levels are also considered) for each partitioned point cloud video tile, we formulate this into an optimization problem under the limited communication and computational resources constraints and propose a scheme to solve it. Extensive simulations are conducted and the simulation results show the superior performance of the proposed scheme over the existing schemes.
Jie Li 0015, Cong Zhang 0002, Zhi Liu 0002, Wei Sun 0011, Qiyue Li 0001
ICC2
2020 Look Ahead at the First-mile in Livecast with Crowdsourced Highlight Prediction
abstract
Recently, data-driven prediction strategies have shown the potential of shepherding the optimization strategies for end viewer's Quality-of-Experience in practical streaming applications. The current prediction-based designs have largely focused on optimizing the last-mile, i.e., viewer-side, which 1) need the real-time feedback from viewers to improve the prediction accuracy; and 2) need quick responses to guarantee the effectiveness of optimization strategies in the future. Thanks to the emerged crowdsourced livecast services, e.g., Twitch.tv, we for the first time exploit the opportunity to realize the long-term prediction and optimization with the assistance derived from the first-mile, i.e., source broadcasters.In this paper, we propose a novel framework CastFlag, which analyzes the broadcasters' operations and interactions, predicts the key events (i.e., highlights), and optimizes the transcoding stage in the corresponding live streams, even before the encoding stage. Taking the most popular eSports gamecast as an example, we illustrate the effectiveness of this framework in the game highlight prediction and transcoding workload allocation. The trace-driven evaluation shows the superiority of CastFlag as it: (1) improves the prediction accuracy over other learning-based approaches by up to 30%; (2) achieves an average of 10% saving of the transcoding latency at less cost.
Cong Zhang 0002, Jiangchuan Liu, Zhi Wang 0001, Lifeng Sun
INFOCOM1
2020 MPTCP+: Enhancing Adaptive HTTP Video Streaming over Multipath
abstract
This paper presents a systematic study on adaptive streaming over MPTCP. We start from realworld experiments with Dynamic Adaptive Streaming over HTTP (DASH) and analysis on its performance over MPTCP. We show that DASH can greatly benefit from the improved aggregated throughput by MPTCP; yet the inter-path throughput difference and the intra-path throughput fluctuation have noticeable (negative) impact, too. Without a proper design of path selection and adaptation in MPTCP, they can easily confuse the adaptation logic of DASH, resulting in low bitrates or frequent rebuffering even if high-bandwidth paths are available. We present MPTCP+, an extended multipath TCP solution to offer high quality and smooth playback for adaptive HTTP streaming. MPTCP+ incorporates a path use decision algorithm that smartly disables/enables a path to minimize the inter-path difference, and a novel congestion control algorithm that smooths congestion window evolution with multiple paths. We have implemented MPTCP+ in the MPTCP Linux kernel, with minimum change on the server-side MPTCP module only. It is fully compatible with the existing MPTCP clients and requires no change on the upper-layer protocols, too. Our experiments suggest that MPTCP+ increases the quality of experience (QoE) of DASH by up to 50%.
Jia Zhao 0006, Jiangchuan Liu, Cong Zhang 0002, Yong Cui 0001, Yong Jiang 0001, Wei Gong 0001
IWQoS3
2020 DeepCast: Towards Personalized QoE for Edge-Assisted Crowdcast With Deep Reinforcement Learning
abstract
Today’s anywhere and anytime broadband connection and audio/video capture have boosted the deployment of crowdsourced livecast services (orcrowdcast). Bridging a massive amount of geo-distributed broadcasters and their fellow viewers, such representatives as Twitch.tv, Youtube Gaming, and Inke.tv, have greatly changed the generation and distribution landscape of streaming content. They also enable rich online interactions among the crowd, and strive to offer personalized Quality-of-Experience (QoE) for individual viewers. Given the ultra-large scale and the dynamics of the crowd, personalizing QoE however is much more challenging than in early generation streaming services. The rich interactions among the broadcasters, viewers, and the network system, on the other hand, also offer invaluable data that could be utilized towards informed management. This paper presentsDeepCast, an edge-assisted crowdcast framework that explores the sheer amount of viewing data towards intelligent decisions for personalized QoE demands. DeepCast seamlessly integrates cloud, CDN, and edge servers for crowdcast content distribution, and advocates a data-driven design that extracts the hidden information from the complex interactions among the system components. Through deep reinforcement learning (DRL), it automatically identifies the most suitable strategies for viewer assignment and transcoding at edges. We collect multiple real-world datasets and evaluate the performance of DeepCast with trace-driven experiments. The results demonstrate its flexibility and effectiveness towards better personalized QoE and lower cost for crowdcast systems.
Fangxin Wang 0001, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu, Yifei Zhu 0001, Haitian Pang, Lifeng Sun
IEEE/ACM Trans. Netw.2
2019 Towards Low Latency Multi-viewpoint 360° Interactive Video: A Multimodal Deep Reinforcement Learning Approach
abstract
Recently, the fusion of 360° video and multi-viewpoint video, called multi-viewpoint (MVP) 360° interactive video, has emerged and created much more immersive and interactive user experience, but calls for a low latency solution to request the high-definition contents. Such viewing-related features as head movement have been recently studied, but several key issues still need to be addressed. On the viewer side, it is not clear how to effectively integrate different types of viewing-related features. At the session level, questions such as how to optimize the video quality under dynamic networking conditions and how to build an end-to-end mapping between these features and the quality selection remain to be answered. The solutions to these questions are further complicated given the many practical challenges, e.g., incomplete feature extraction and inaccurate prediction.This paper presents an architecture, called iView, to address the aforementioned issues in an MVP 360° interactive video scenario. To fully understand the viewing-related features and provide a one-step solution, we advocate multimodal learning and deep reinforcement learning in the design. iView intelligently determines video quality and reduces the latency without pre-programmed models or assumptions. We have evaluated iView with multiple real-world video and network datasets. The results showed that our solution effectively utilizes the features of video frames, networking throughput, head movements, and viewpoint selections, achieving at least 27.2%, 15.4%, and 2.8% improvements on the three video datasets, respectively, compared with several state-of-the-art methods.
Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Jiangchuan Liu, Lifeng Sun
INFOCOM2
2019 Intelligent Edge-Assisted Crowdcast with Deep Reinforcement Learning for Personalized QoE
abstract
Recent years have seen booming development and great success in interactive crowdsourced livecast (i.e., crowdcast). Different from traditional livecast services, crowdcast is featured with tremendous video contents at the broadcaster side, highly diverse viewer side content watching environments/preferences as well as viewers' personalized quality of experience (QoE) demands (e.g., individual preferences for streaming delays, channel switching latencies and bitrates). This imposes unprecedented key challenges on how to flexibly and cost-effectively accommodate the heterogeneous and personalized QoE demands for the mass of viewers. In this paper, we propose DeepCast, an edge-assisted crowdcast framework, which makes intelligent decisions at edges based on the massive amount of real-time information from the network and viewers to accommodate personalized QoE with minimized system cost. Given the excessive computation complexity in this context, we propose a data-driven deep reinforcement learning (DRL) based solution that can automatically learn the best suitable strategies for viewer scheduling and transcoding selection. To our best knowledge, DeepCast is the first edge-assisted framework that applies the advance of DRL to explicitly accommodate personalized QoE optimization for crowdcast services. We collect multiple real-world datasets and evaluate the performance of DeepCast using trace-driven experiments. The results demonstrate the superiority of our DeepCast framework and its DRL-based solution.
Fangxin Wang 0001, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu, Yifei Zhu 0001, Haitian Pang, Lifeng Sun
INFOCOM2
2019 Video processing with serverless computing: a measurement study
abstract
The growing demand for video processing and the advantages in scalability and cost reduction brought by the emerging serverless computing have attracted significant attention in serverless computing powered video processing. However, how to implement and configure serverless functions to optimize the performance and cost of video processing applications remains unclear. In this paper, we explore the configuration and implementation schemes of typical video processing functions deployed to the serverless platforms and quantify their influence on the execution duration and monetary cost from a developer's perspective. Our measurement reveals that memory configuration is non-trivial. Dynamic profiling of workloads is necessary to find the best memory configuration. Moreover, compared with calling external video processing APIs, implementing these services locally in serverless functions can be competitive. We also find that the performance of video processing applications could be affected by the underlying infrastructure. Our work provides guidelines for further function-level optimization and complements the existing measurement studies for both serverless computing and video processing.
Miao Zhang 0003, Yifei Zhu 0001, Cong Zhang 0002, Jiangchuan Liu
NOSSDAV3
2018 Highlight-Aware Content Placement in Crowdsourced Livecast Services
abstract
Recent years have witnessed an explosion of crowdsourced livecast (i.e., live broadcast) services, in which any Internet users can act as broadcasters to publish livecasts to fellow viewers. To help grow broadcasters' channels, crowdsourced livecast services provide a past-broadcast saving service, allowing viewers to watch the replays they may have missed. Our real-trace measurement and questionnaire survey show that (1) the duration of most of livecasts is extremely long; (2) a much longer duration largely affects the viewers' Quality-of-Experiences (QoE) when watching the replays. To address this issue and improve viewers' QoE, we propose a crowdsourced framework HighCast based on the interactive messages contributed by the viewers in crowdsourced livecast services. According to a highlight-aware detection module, HighCast can exploit the detection results to schedule the content placement by considering the importance of the predicted streaming highlights. The trace-based evaluations illustrate that the proposed framework improves the prediction accuracy and reduces the viewing latency.
Cong Zhang 0002, Jiangchuan Liu, Haitian Pang, Fangxin Wang 0001
IWQoS1
2018 Optimizing Personalized Interaction Experience in Crowd-Interactive Livecast: A Cloud-Edge Approach
abstract
Enabling users to interact with broadcasters and audience, the crowd-interactive livecast greatly improves viewer's quality of experience (QoE) and attracts millions of daily active users recently. In addition to striking the balance between resource utilization and viewers' QoE met in the traditional video streaming service, this novel service needs to take supererogatory efforts to improve the interaction QoE, which reflects the viewer interaction experience. To tackle this issue, we conduct measurement studies over a large-scale dataset crawled from a representative livecast service provider. We observe that the individual's interaction pattern is quite heterogeneous: only 10% viewers proactively participate in the interaction, and the rest viewers usually watch passively. Incorporating the insight into the emerging cloud-edge architecture, we propose a framework PIECE, which optimizes the Personalized Interaction Experience with Cloud-Edge architecture (PIECE) for intelligent user access control and livecast distribution. In particular, we first devise a novel deep neural network based algorithm to predict users' interaction intensity using the historical viewer pattern. We then design an algorithm to maximize the individual's QoE, by strategically matching viewer sessions and transcoding-delivery paths over cloud-edge infrastructure. Finally, we use trace-driven experiments to verify the effectiveness of PIECE. Our results show that our prediction algorithm outperforms the state-of-the-art algorithms with a much smaller mean absolute error (40% reduction). Furthermore, in comparison with the cloud-based video delivery strategy, the proposed framework can simultaneously improve the average viewers QoE (26% improvement) and interaction QoE (21% improvement), while maintaining a high streaming bitrate.
Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Han Hu 0003, Zhi Wang 0001, Jiangchuan Liu, Lifeng Sun
ACM Multimedia2
2018 Dependency- and similarity-aware caching for HTTP adaptive streaming
Cong Zhang 0002, Jiangchuan Liu, Fei Chen 0010, Yong Cui 0001, Edith C. H. Ngai, Yueming Hu 0001
Multim. Tools Appl.1
2017 Beyond the touch: Interaction-aware mobile gamecasting with gazing pattern prediction
abstract
Recent years have witnessed an explosion of gamecasting applications in the market, in which game players (or gamers in short) broadcast their game scenes in real-time. Such pioneer applications as YouTube Gaming, Twitch, and Mobcrush have attracted a massive number of online broadcasters, and each of them can attract hundreds or thousands of fellow viewers. The growing number however has created significant challenges to the network and end-devices, particularly considering bandwidth- and battery-limited smartphones or tablets are becoming dominating for both gamers and viewers. Yet the unique touch operations of the mobile interface offer opportunities, too. In this paper, our crowdsourced measurement reveals that strong associations exist between the gamers' touch interactions and the viewers' gazing patterns. Motivated by this, we present a novel interaction-aware optimization framework to improve the energy utilization and stream quality for mobile gamecasting (MGC). Our framework incorporates a touch-assisted prediction module to extract association rules for gazing pattern prediction and a tile-based optimization module to utilize energy on mobile devices efficiently. Trace-driven simulations illustrate the effectiveness of our framework in terms of energy consumption and streaming quality. Our user study experiments also demonstrate much improved (3%-13%) quality satisfaction than the state-of-the-art solution with similar network resources.
Cong Zhang 0002, Qiyun He, Jiangchuan Liu, Zhi Wang 0001
INFOCOM1
2017 When Cloud Meets Uncertain Crowd: An Auction Approach for Crowdsourced Livecast Transcoding
abstract
In the emerging crowd sourced live cast services, numerous amateur broadcasters live stream their video contents to worldwide viewers and constantly interact with them through chat messages. Live video contents are transcoded into multiple quality versions to better service viewers with different network and device configurations. Cloud computing becomes a natural choice to handle these computational intensive tasks due to its elasticity and the "pay-as-you-go" billing model. However, given the significantly large number of concurrent channel numbers and the diverse viewer geo-distributions in this new crowd sourced live cast service, even the cloud becomes significantly expensive to cover the whole community and inadequate in fulfilling the latency requirement. In this paper, after observing the abundant computational resources residing in end viewers, we propose a Cloud-Crowd collaborative system, C2, which combines end viewers with cloud to perform video transcoding in a cost-efficient way. To quantify the heterogeneity and uncertainty of viewers and pass the asymmetric information barrier, we incorporate statistical descriptions into our bidding language and design truthful auctions to recruit stable viewers with appropriate incentives. We further tailor redundancy strategies for workloads with different Quality of Service requirements to improve the stability of our system. Desirable economic properties, like social efficiency, ex-post incentive compatibility, individual rationality, are proved to be guaranteed in our studied scenarios. Using traces captured from the popular Twitch platform, we show that C2 achieves up to 93% more cost saving than a pure cloud-based solution, and significantly outperforms other baseline approaches in both social welfare and system stability.
Yifei Zhu 0001, Jiangchuan Liu, Zhi Wang 0001, Cong Zhang 0002
ACM Multimedia4
2017 Seeker: Topic-Aware Viewing Pattern Prediction in Crowdsourced Interactive Live Streaming
abstract
Recently, Crowdsourced Interactive Live Streaming (CILS), such as Twitch.tv and Periscope, has emerged as one of the most popular streaming applications over the Internet. In such applications, a large number of geo-distributed users publish live sources to broadcast their game sessions, personal activities, and other events, while fellow viewers not only watch these live streams, but also contribute interactive messages to influence streaming content. Such explosively increasing popularity has posed significant challenges to predict viewing patterns using traditional time-series approaches, which lack the start/end knowledge of live streams and cannot capture the viewing burst very well.
Cong Zhang 0002, Jiangchuan Liu, Lifeng Sun, Bo Li 0001
NOSSDAV1
2017 CrowdTranscoding: Online Video Transcoding With Massive Viewers
abstract
Driven by the advances in personal computing devices and the prevalence of high-speed network accesses, crowdsourced livecast platforms have emerged in recent years, through which numerous broadcasters lively stream their video content to fellow viewers. Compared to professional video producers and broadcasters, these new generation broadcasters are highly heterogeneous in terms of the network/system configurations and, therefore, the generated video quality, which calls for massive encoding and transcoding in order to unify the video sources and serve multiple quality versions to viewers with different configurations. On the other hand, with the rapid evolution in the hardware industry, high-performance processors become mainstream in personal computer market. More end devices can easily transcode high-quality videos in realtime. We witness huge computational resource among the massive fellow viewers that could potentially be used for transcoding. In this paper, we propose CrowdTranscoding, a novel framework for crowdsourced livecast systems that offloads the transcoding assignment to the massive viewers. We identify that the key challenges in CrowdTranscoding are to detect qualified stable viewers and to properly assign them to the source channels. We put forward a viewer crowdsourcing transcode scheduler to smartly schedule the workload assignment. Our solution has been evaluated under diverse viewer/channel conditions as well as different parameter settings. The trace-driven simulation confirms the superiority of CrowdTranscoder, while our PlanetLab-based and real world end-viewer experiments show the practical performance of our approach, which also give hint to the further enhancement.
Qiyun He, Cong Zhang 0002, Jiangchuan Liu
IEEE Trans. Multim.2
2017 Live Broadcast With Community Interactions: Bottlenecks and Optimizations
abstract
Recent years have witnessed the rapid growth of new live broadcast services, represented by Twitch.tv and YouTube live events, where videos are crowdsourced from amateur users (e.g., game players), rather than from commercial and professional TV broadcaster or content providers. The viewers also actively contribute to the content through embedded open-chat channels. Such community interactions among viewers, or even between broadcasters and viewers, make content generation highly diversified and engaging, particularly for the young generation. In this context, cross-viewer synchronization is highly desirable; otherwise the viewers with shorter broadcast latency may act as spoilers, significantly affecting the user experience of other viewers. In this paper, we show that the end-to-end delay has a dramatically amplified impact on the broadcast latency for individual viewers. We suggest smart rate adaptation to achieve cross-viewer synchronization, and develop distributed algorithms based on dual decomposition. We further extend our solution to the cloud environment, and present the concept of ShadowCast, which moves broadcasters to the cloud to provide high-quality streams beyond broadcasters' network bandwidth constraint. Its practicability and effectiveness is demonstrated by our implementation and test bed experiments.
Xiaoqiang Ma, Cong Zhang 0002, Jiangchuan Liu, Ryan Shea, Di Fu
IEEE Trans. Multim.2
2017 Exploring Viewer Gazing Patterns for Touch-Based Mobile Gamecasting
abstract
Recent years have witnessed an explosion of gamecasting applications, in which game players (or gamers in short) broadcast game playthroughs by their personal devices in real time. Such pioneer platforms, such as YouTube Gaming, Twitch, and Mobcrush, have attracted a massive number of online broadcasters, and each of them can have hundreds or thousands of fellow viewers. The growing number, however, has created significant challenges to the network and end-devices, particularly considering that bandwidth- and battery-limited smartphones or tablets are becoming dominating for both gamers and viewers. Yet the unique touch operations of the mobile interface offer opportunities, too. In this paper, our measurements based on the real traces from gamers and viewers reveal that strong associations exist between the gamers' touch interactions and the viewers' gazing patterns. Motivated by this, we present a novel interaction-aware optimization framework to improve the energy utilization and stream quality for mobile gamecasting. Our framework incorporates a touch-assisted prediction module to extract association rules for gazing pattern prediction and a tilebased optimization module to utilize energy on mobile devices efficiently. Trace-driven simulations illustrate the effectiveness of our framework in terms of energy consumption and stream quality. Our user study experiments also demonstrate much improved (3%-13%) quality satisfaction over the state-of-the-art solution with similar network resources.
Cong Zhang 0002, Qiyun He, Jiangchuan Liu, Zhi Wang 0001
IEEE Trans. Multim.1
2017 Cloud-Assisted Crowdsourced Livecast
abstract
The past two years have witnessed an explosion of a new generation of livecast services, represented by Twitch.tv , GamingLive , and Dailymotion , to name but a few. With such a livecast service, geo-distributed Internet users can broadcast any event in real-time, for example, game, cooking, drawing, and so on, to viewers of interest. Its crowdsourced nature enables rich interactions among broadcasters and viewers but also introduces great challenges to accommodate their great scales and dynamics. To fulfill the demands from a large number of heterogeneous broadcasters and geo-distributed viewers, expensive server clusters have been deployed to ingest and transcode live streams. Yet our Twitch-based measurement shows that a significant portion of the unpopular and dynamic broadcasters are consuming considerable system resources; in particular, 25% of bandwidth resources and 30% of computational capacity are used by the broadcasters who do not have any viewers at all. In this article, through the real-world measurement and data analysis, we show that the public cloud has great potentials to address these scalability challenges. We accordingly present the design of Cloud-assisted Crowdsourced Livecast (CACL) and propose a comprehensive set of solutions for broadcaster partitioning. Our trace-driven evaluations show that our CACL design can smartly assign ingesting and transcoding tasks to the elastic cloud virtual machines, providing flexible and cost-effective system deployment.
Cong Zhang 0002, Jiangchuan Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2016 Utilizing Massive Viewers for Video Transcoding in Crowdsourced Live Streaming
abstract
Driven by the advances in personal computing devices and the prevalence of broadband network and wireless mobile network accesses, Crowdsourced Live Streaming (CLS) platforms have emerged in recent years, through which numerous broadcasters lively stream their video content, e.g., live events or online game scenes, to fellow viewers. Compared to professional video producers and broadcasters, these new generation broadcasters are highly heterogenous in terms of the network/system configurations and therefore the generated video quality, which calls for massive encoding and transcoding in order to unify the video sources and serve multiple quality versions to viewers with different configurations. On the other hand, with the rapid evolution in the hardware industry, high performance processors (e.g., Intel Core i7-4790K CPU) become mainstream in personal computer market. More end devices can easily transcode high quality videos in realtime. We witness huge computational resource among the massive fellow viewers that could potentially be used for transcoding. In this paper, inspired by fog computing, we propose Crowd-Transcoding, a novel framework for CLS systems that offloads the transcoding assignment to the massive viewers. We identify that the key challenges in CrowdTranscoding are to detect qualified stable viewers and to properly assign them to the source channels. We put forward Viewer Crowdsourcing Transcode Scheduler (VCTS) to smartly schedule the workload assignment. Our solution has been evaluated under diverse viewer/channel conditions as well as different parameter settings. The trace-driven simulation confirms the superiority of CrowdTranscoder, while our PlanetLab-based and real world end-viewer experiments show the practical performance of our approach, which also give hint to the further enhancement.
Qiyun He, Cong Zhang 0002, Jiangchuan Liu
CLOUD2
2016 Power-Aware Wireless Transmission for Computation Offloading in Mobile Cloud
abstract
In today's mobile devices, the battery reservoir remains severely limited in capacity, making power consumption a key concern in the design and implementation of mobile applications. In this paper, we closely examine one widely adopted approach to improve the energy efficiency of mobile applications-adaptively offloading the computation to the remote cloud. In particular, we measure the power consumption of computation offloading for two representative real-world mobile cloud applications under various wireless network conditions and identify the unique features of data transmission for computation offloading. We then formulate the power-aware scheduling problem for computation offloading and present a scheduling algorithm that makes adaptive offloading decisions according to the dynamic network conditions. Simulation results show that our proposed method can achieve better battery performance, which also reveal that computation-intensive and delay-tolerant tasks are more likely to benefit from offloading.
Lei Zhang 0066, Cong Zhang 0002, Jiangchuan Liu, Xiaowen Chu 0001, Ke Xu 0002, Yong Jiang 0001
ICCCN2
2016 Towards hybrid cloud-assisted crowdsourced live streaming: measurement and analysis
abstract
Crowdsourced Live Streaming (CLS), most notably Twitch.tv, has seen explosive growth in its popularity in the past few years. In such systems, any user can lively broadcast video content of interest to others, e.g., from a game player to many online viewers. To fulfill the demands from both massive and heterogeneous broadcasters and viewers, expensive server clusters have been deployed to provide video ingesting and transcoding services. Despite the existence of highly popular channels, a significant portion of the channels is indeed unpopular. Yet as our measurement shows, these broadcasters are consuming considerable system resources; in particular, 25% (resp. 30%) of bandwidth (resp. computation) resources are used by the broadcasters who do not have any viewers at all. In this paper, we closely examine the challenge of handling unpopular live-broadcasting channels in CLS systems and present a comprehensive solution for service partitioning on hybrid cloud. The trace-driven evaluation shows that our hybrid cloud-assisted design can smartly assign ingesting and transcoding tasks to the elastic cloud virtual machines, providing flexible system deployment cost-effectively.
Cong Zhang 0002, Jiangchuan Liu
NOSSDAV1
2015 Crowdsourced live streaming over the cloud
abstract
Empowered by today's rich tools for media generation and distribution, and the convenient Internet access, crowdsourced streaming generalizes the single-source streaming paradigm by including massive contributors for a video channel. It calls a joint optimization along the path from crowdsourcers, through streaming servers, to the end-users to minimize the overall latency. The dynamics of the video sources, together with the globalized request demands and the high computation demand from each sourcer, make crowdsourced live streaming challenging even with powerful support from modern cloud computing. In this paper, we present a generic framework that facilitates a cost-effective cloud service for crowdsourced live streaming. Through adaptively leasing, the cloud servers can be provisioned in a fine granularity to accommodate geo-distributed video crowdsourcers. We present an optimal solution to deal with service migration among cloud instances of diverse lease prices. It also addresses the location impact to the streaming quality. To understand the performance of the proposed strategies in the realworld, we have built a prototype system running over the planetlab and the Amazon/Microsoft Cloud. Our extensive experiments demonstrate that the effectiveness of our solution in terms of deployment cost and streaming quality.
Fei Chen 0010, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu
INFOCOM2
2015 On crowdsourced interactive live streaming: a Twitch.tv-based measurement study
abstract
Empowered by today's rich tools for media generation and collaborative production, the multimedia service paradigm is shifting from the conventional single source, to multi-source, to many sources, and now toward crowdsource. Such crowdsourced live streaming platforms as Twitch.tv allow general users to broadcast their content to massive viewers, thereby greatly expanding the content and user bases. The resources available for these non-professional broadcasters however are limited and unstable, which potentially impair the streaming quality and viewers' experience. The diverse live interactions among the broadcasters and viewers can further aggravate the problem.
Cong Zhang 0002, Jiangchuan Liu
NOSSDAV1
2015 Cloud-Assisted Live Streaming for Crowdsourced Multimedia Content
abstract
Empowered by today's rich tools for media generation and distribution, and the convenient Internet access , streaming crowdsourced multimedia content (crowdsourced streaming, in brief) generalizes the single-source streaming paradigm by including massive contributors for a video/data channel. It calls a joint optimization along the path from crowdsourcers , through streaming servers, to the end-users to minimize the overall latency. The dynamics of the video sources, together with the globalized request demands and the high computation demand from each sourcer, make crowdsourced live streaming challenging even with powerful support from modern cloud computing. In this paper, we present a generic framework that facilitates a cost-effective cloud service for crowdsourced live streaming. Through adaptively leasing, the cloud servers can be provisioned in a fine granularity to accommodate geo-distributed video crowdsourcers. We present an optimal solution to deal with service migration among cloud instances of diverse lease prices. It also addresses the location impact to the streaming quality. To understand the performance of the proposed strategies in the real world, we have built a prototype system running over the planetlab and the Amazon/Microsoft Cloud. Our extensive experiments demonstrate that the effectiveness of our solution in terms of deployment cost and streaming quality.
Fei Chen 0010, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu, Yuan Liu 0021
IEEE Trans. Multim.2
2014 Insight Data of YouTube from a Partner's View
abstract
YouTube is arguably the most popular online videos sharing site nowadays. To further augment its service with better revenue, it has started working with content owners (known as YouTube partners) whose copyrighted videos and channels have pulled massive audience. By uploading high-quality premium videos, the partners have essentially changed the user-generated content feature of YouTube and further increased YouTube's popularity. Understanding the latest YouTube access pattern is thus crucial to both YouTube and its partners, as well as to other providers of relevant services. In this paper, we for the first time analyze a large-scale YouTube dataset from a partner's view. We make effective use of Insight, a new analytics service of YouTube that offers inside statistics for partners about their content accesses and audience behaviours. From the raw Insight data that are confined to simple scalars and charts, we reveal the inherent relationship among the various metrics that affect the popularity of the videos. Our findings facilitate YouTube partners to adapt their content deployment and user engagement strategies, having great potentials for them to collaborate with YouTube to generate more views and subsequently increasing their revenues.
Xu Cheng 0004, Mehrdad Fatourechi, Xiaoqiang Ma, Cong Zhang 0002, Lei Zhang 0066, Jiangchuan Liu
NOSSDAV4