Fangxin Wang 0001

dblp:142/0351 · DBLP profile ↗
← Back
85ranked-venue papers
10as first author
64since 2021 · last 2026
0000-0003-2559-045XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 60 · 9 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 19 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Maniflat3D: Learning 3D Geometry Through Planar Representations from Multi-Layer Unwrapping
abstract
Point-based geometric representations such as point clouds and Gaussian Splatting are fundamental for 3D understanding. However, the inherent irregularity and high-dimensional nature of point structures present significant challenges for direct 3D learning approaches, which often struggle with scalability and achieve suboptimal performance due to sparse data distributions. In contrast, 2D learning paradigms benefit from well-established architectures with superior optimization stability and efficiency. To bridge this gap, we propose Maniflat3D, a unified framework that systematically transforms volumetric point-based geometries into structured 2D representations through a two-stage process: a multilayer Ball-Pivoting reconstruction with adaptive density control, followed by Scalable Locally Injective Mapping (SLIM) to produce distortion-minimized, bijective UV parameterizations. Our approach explicitly encodes both geometric and attribute information into the flattened domain, enabling conventional 2D neural networks to effectively learn from complex 3D structures such as Gaussian Splatting. Experiments on the ShapeSplat dataset demonstrate that Maniflat3D achieves comparable performance while reducing parameter count by 90% compared to native 3D baselines, and simultaneously attains 21× compression ratio through neural encoding. These results establish a new paradigm for efficient geometric understanding, demonstrating successful transfer of planar learning advantages to challenging 3D manifold problems through dimensional reduction.
Zijian Cao 0007, Dayou Zhang, Zeyuan Liu, Zhicheng Liang, Fangxin Wang 0001
AAAI5
2026 FedBRICK: Structural Bias Aware Heterogeneous Foundation Model Federated Tuning
abstract
Model-heterogeneous federated tuning (MHFT) enables the privacy-preserving fine-tuning of foundation models in heterogeneous systems by allowing clients and the server to adopt different model architectures. Depth partial training—where each client updates only a subset of the model's layers—alleviates system heterogeneity but exacerbates client drift, which stems from clients optimizing different objectives and therefore degrades overall performance. Beyond the well-known statistical bias—where non-IID data leads to client drift—we identify a structural bias arising from clients deploying only partial layers of the global model, which serves as an important cause of drift. We further provide a theoretical analysis showing that the possible range of structural bias expands linearly with the number of missing layers. To counter this effect, we introduce FedBRICK (Federated Bias Recovery via Inserted Calibrative Kernels), which inserts tiny BRICKs into each client’s subnetwork. We employ a dual-end layer-wise distillation scheme to train these blocks using both client-side local data and a small public proxy set on the server. This design effectively mitigates the structural bias caused by layer dropping, reduces client drift, and remains practical for storage-constrained devices. Extensive experiments on federated learning benchmarks confirm that FedBRICK delivers up to a 5% average accuracy gain while requiring no more than 1.44% extra storage per client.
Xianda Wang, Fangxin Wang 0001
AAAI5
2026 RAPIDS: Reliable Adaptive Priority-based Intelligent Delivery Streaming for Enhanced Error Resilience in Real-time Video
Dayou Zhang, Zijian Cao, Fangxin Wang 0001
IWQoS6
2026 Morphe: High-Fidelity Generative Video Streaming with Vision Foundation Model
Tianyi Gong, Zijian Cao 0007, Zixing Zhang 0009, Jiangkai Wu, Xinggong Zhang, Shuguang Cui, Fangxin Wang 0001
NSDI7
2026 MetaKube: An Experience-Aware LLM Framework for Kubernetes Failure Diagnosis
Xinran Tian, Wanshun Lan, Xuhan Feng, Haoyue Li, Fangxin Wang 0001
WWW7
2026 Fast-Tactical Diffusion for On-Board AAV Spectrum-Level Signal Deception
abstract
Spectrum deception has found broad utility across multiple domains, including electronic warfare, tactical counter-measures, and adversarial sensing suppression. However, generating complex time–frequency signatures relies on sophisticated signal processing pipelines, which poses significant challenges for UAV and other IoT platforms with severely constrained onboard computational resources. Moreover, the limited onboard capability further restricts rapid signal synthesis and adaptation, failing to meet the strict rapid-response requirements of the battlefield. To address this dilemma, we propose the Fast-Tactical Signal Deception Framework (FT-SDF), a specialized generative architecture optimized for real-time signal synthesis. We formulate a novel spectro-temporal diffusion dynamics mechanism that innovatively incorporates additional spectral blurring and reverse process variance, jointly optimizing noise prediction and variance, which is necessary to preserve fine-grained spectral structures and key time–frequency signatures across different modulation schemes. Notably, to ensure strict adherence to communication protocols, we introduce a lightweight spectrum-context encoder that employs a dual-domain embedding strategy for physics-aware conditioning. Furthermore, to enable rapid inference, we develop a variance-aware acceleration mechanism that exploits learned spectral uncertainty to guide a dynamic warm-start schedule, thereby drastically compressing the sampling trajectory. Extensive evaluations on a systematically reconstructed RadioML benchmark demonstrate that FT-SDF outperforms state-of-the-art baselines. Specifically, it achieves a 59.3% reduction in sampling iterations (more than 2-fold inference speedup), while maintaining both high statistical fidelity (FID < 15) and industrial-grade precision (EVM ≤ 14%), demonstrating the feasibility of rapid, controllable generative AI in complex electromagnetic environments.
Lingyun Feng, Chao Zhu 0002, Fangxin Wang 0001, Jianping An
IEEE Internet Things J.6
2026 Live High-Fidelity Semantic Communication via Cross-Modal Fusion for Volumetric Video
abstract
Semantic communication (SC) emerges as a breakthrough paradigm for efficient data transmission in next-generation communication networks. However, SC is still in the infant stage with quite a few limitations, such as insufficient semantic representation capacity, high communication latency, and the susceptibility to channel noise. In this paper, we proposeLiFiSC, a cross-modal fusion based generative semantic communication framework with strong semantic compression capacity and high-fidelity semantic restoration. We then extend it toLiFiSCvv, specifically designed to achieve 3D volumetric video transmission and reconstruction with photorealistic visual quality and pixel-level visual consistency, providing end-to-end live watching experience with acceptable latency. We innovatively incorporate unified vision-language encoding into semantic communication, achieving superior semantic understanding and compression.LiFiSCvvcomprises three key components: (1) Information redundancy reduction through lightweight video analysis and structure-from-motion techniques, decreasing reconstruction cost; (2) Cross-modal fusion learning driven codec mechanisms that enable efficient semantic representation and robust transmission against channel impairments; and (3) Feed-forward vision transformer for rapid volumetric video reconstruction and rendering. Comprehensive evaluation results demonstrate thatLiFiSCvvachieves a live high-fidelity and visually consistent video watching experience, with only seconds of end-to-end latency and around 42× semantic compression, significantly outperforming SOTA methods in general image and volumetric video transmission. The project page is available here: https://inml-tygong.github.io/LiFiSCvv/.
Tianyi Gong, Zijian Cao 0007, Zhicheng Liang, Dayou Zhang, Fangxin Wang 0001, Shuguang Cui
IEEE J. Sel. Areas Commun.5
2026 FedShieldLLM: Measurement, Detection and Protection of Privacy Leakage in Federated LLMs
abstract
Large Language Models (LLMs) exhibit remarkable language understanding and problem-solving capabilities but are heavily reliant on vast amounts of high-quality data, which raises significant privacy concerns. Federated Learning (FL) presents a decentralized training paradigm that minimizes raw data sharing but is still vulnerable to privacy attacks that exploit shared model updates. Federated LLMs, especially dialogue models, are particularly prone to query-based attacks, which can lead to the unintentional disclosure of memorized private information. This study performs extensive experiments across multiple model sizes and federated architectures to assess the extent of query-based privacy leakage and evaluate existing protection mechanisms. To enable standardized and reproducible privacy assessments, we propose a multi-stage privacy leakage detection method, integrating structured privacy attribute extraction, controlled injection, realistic prompting, and quantitative evaluation under both generalized and personalized privacy scenarios. Additionally, we introduce FedShieldLLM, a lightweight, adaptable client-side anonymization framework that performs personalized data de-identification prior to federated training, effectively mitigating privacy risks while preserving task utility. Our contributions include: (1) a systematic analysis of query-based privacy leakage in federated LLMs; (2) the design of a reproducible, architecture-agnostic multi-stage detection method; and (3) the development of FedShieldLLM for secure and flexible client participation in privacy-sensitive federated learning.
Xianda Wang, Zhicheng Liang, Wanshun Lan, Yingchun Chen, Fangxin Wang 0001
IEEE Trans. Mob. Comput.7
2026 Efficient and Privacy-Preserving Federated Knowledge Learning for Distributed LLM
abstract
Federated learning (FL) stands out as a promising solution to address the ever-increasing data scarcity problem of large language model (LLM) training through collecting data from distributed sources. The unique challenges however exist in three aspects, the extreme large communication overhead due to frequent model aggregation, the suboptimal performance due to data heterogeneity, and privacy leakage risk due to model parameter exchange. To address these issues, we propose Federated Knowledge Learning (FKL), a novel token-level knowledge sharing framework tailored for LLM to achieve effective task relevant knowledge aggregation. FKL offers the following advantages: (1) Adaptive Flexibility: Clients and servers can employ models customized for their specific tasks while ensuring seamless collaboration. Each node within the system functions independently and can be flexibly adapted as needed. (2) Communication Efficiency: Instead of huge model or parameter exchange, the unique knowledge sharing mechanism of FKL can significantly reduce the communication overhead while delivering superior performance. (3) Privacy Safeguards: The distributed token-level knowledge-sharing mechanism ensures privacy by deidentifying sensitive information and mitigating potential leakage risks. Experiments on LLMs (125M-7B parameters) validate FKL's versatility, achieving 50% higher accuracy on GLUE benchmarks and a notable 0.7x improvement in privacy task response quality compared to existing FL-based LLM approaches.
Xianda Wang, Zhicheng Liang, Tianyi Gong, Wanshun Lan, Yingchun Chen, Haoyue Li, Fangxin Wang 0001
IEEE Trans. Mob. Comput.8
2026 4DGStream: Variable Bitrate Dynamic Gaussian Splatting Streaming
abstract
While 3D Gaussian Splatting (3DGS) has revolutionized static scene representation, the extension to dynamic scene, i.e., 3DGS video (GSV), faces challenges related to reconstruction quality, rendering speed, and storage requirements. The substantial data volume of current GSV poses significant hurdles for streaming applications, particularly in the realm of AR, VR and MR. To tackle these challenges, we introduce 4DGStream, a novel framework that integrates an efficient GSV compression method, Light4D, and a bitrate adaptation streaming strategy, QoSmooth, to ensure smooth playback while maintaining high visual quality. Light4D employs a binarizationassisted spatiotemporal deformation network to model the deformation of Gaussian primitive attributes over time, while a spatiotemporal-aware masking module prunes trivial Gaussians, further enhancing long-term reconstruction quality. To reduce storage, Light4D uses a binary hash grid to model the entropy of attributes for arithmetic coding, with its binary nature allowing efficient entropy modeling via a Bernoulli distribution. These components enable Light4D to improve the FPS/Storage metric by up to 12.4× over SpacetimeGS and 26.4× over 4DGS on the Neu3D dataset, with performance gains exceeding 3× orders of magnitude compared to other NeRF-based state-of-the-art (SOTA) methods. Here, FPS/Storage reflects the balance between rendering speed and data storage. Despite significant model size reductions, Light4D maintains or surpasses the reconstruction quality of 4DGS. Furthermore, QoSmooth provides effective rate control to enhance playback smoothness, reducing bitrate level switches by 61.6% and increasing time-average utility by 26.2%. All these improvements make 4DGStream highly suited for GSV streaming, improving QoE by 36.7% compared to SOTA methods.
Zhicheng Liang, Dayou Zhang, Linfeng Shen, Miao Zhang 0003, Jian Zhang 0054, Bin Ju, Mallesham Dasari, Fangxin Wang 0001, Jiangchuan Liu
IEEE Trans. Multim.8
2026 Bistatic-Enhancement MIMO ISAC: Joint Beamforming Design in Cell-Free Communication and Bistatic Radar Systems
abstract
Multiple-input multiple-output (MIMO) integrated sensing and communication (ISAC) is a promising solution to achieve higher performances of dual functionalities. However, the existing cell-free/bistatic MIMO ISAC networks struggle to meet strict requirements for data-intensive communication and accuracy-sensitive radar positioning. To further achieve joint enhancement, we propose a novel network where two ISAC transmitters cooperatively perform communication and target positioning, fully leveraging the advantages of cell-free/bistatic principles in communication/radar systems, referred to as bistatic-enhancement MIMO ISAC. An optimization for joint beamforming design is established to maximize the sum data rate for communication users and minimize a novel positioning-enhanced Cramér-Rao lower bound (CRB) that evaluates positioning accuracy under their corresponding requirements. The established problem is solved under two schemes: cooperative block-level and symbol-level beamforming. The solution under the former scheme is derived by an iterative behavior. Under the latter one, inter-user interference is eliminated and co-channel interference is exploited for useful signal enhancement. The problem can be converted into a convex semi-definite problem (SDP) based on semi-definite relaxation (SDR). Experimental results substantiate the effectiveness of the proposed algorithms. More importantly, the proposed bistatic-enhancement network improves positioning accuracy by 32.5% ∼ 47.5% over the conventional bistatic-site one under different schemes.
Boxin He, Wencan Mao, Yaxi Liu 0001, Wei Huangfu, Fangxin Wang 0001, Haijun Zhang 0001
IEEE Trans. Wirel. Commun.5
2025 Cluster Based Heterogeneous Federated Foundation Model Adaptation and Fine-Tuning
abstract
In recent years, the distributed training of foundation models (FMs) has seen a surge in popularity. In particular, federated learning enables collaborative model training among edge clients while safeguarding the privacy of their data. However, federated training of FMs across resource-constrained and highly heterogeneous edge devices encounter several challenges. These include the difficulty of deploying FMs on clients with limited computational resources and the high computation and communication costs associated with fine-tuning and collaborative training. To address these challenges, we propose FedCKMS, a Cluster-Aware Framework with Knowledge-Aware Model Search. Specifically, FedCKMS incorporates three key components. The first component is multi-factor heterogeneity-aware clustering, which groups clients based on both data distribution and resource limitations and selects an appropriate model for each cluster. The second component is knowledge-aware model architecture search, which enables each client to identify the optimal sub-model from the cluster model, facilitating adaptive deployment that accommodates highly heterogeneous computational resources across clients. The final component is cluster-aware knowledge transfer, which facilitates knowledge sharing between clusters and the server, addressing model heterogeneity, and reducing communication overhead. Extensive experiments demonstrate that FedCKMS outperforms state-of-the-art baselines by 3-10% in accuracy.
Xianda Wang, Yaqi Qiao, Duo Wu, Chenrui Wu 0002, Fangxin Wang 0001
AAAI5
2025 ExScene: Free-View 3D Scene Reconstruction with Gaussian Splatting from a Single Image
abstract
The increasing demand for augmented and virtual reality applications has highlighted the importance of crafting immersive 3D scenes from a simple single-view image. However, due to the partial priors provided by single-view input, existing methods are often limited to reconstruct low-consistency 3D scenes with narrow fields of view from single-view input. These limitations make them less capable of generalizing to reconstruct immersive scenes. To address this problem, we propose ExScene, a two-stage pipeline to reconstruct an immersive 3D scene from any given single-view image. ExScene designs a novel multimodal diffusion model to generate a high-fidelity and globally consistent panoramic image. We then develop a panoramic depth estimation approach to calculate geometric information from panorama, and we combine geometric information with high-fidelity panoramic image to train an initial 3D Gaussian Splatting (3DGS) model. Following this, we introduce a GS refinement technique with 2D stable video diffusion priors. We add camera trajectory consistency and color-geometric priors into the denoising process of diffusion to improve color and spatial consistency across image sequences. These refined sequences are then used to fine-tune the initial 3DGS model, leading to better reconstruction quality. Experimental results demonstrate that our ExScene achieves consistent and immersive scene reconstruction using only single-view input, significantly surpassing state-of-the-art baselines.
Tianyi Gong, Yifei Zhong, Fangxin Wang 0001
ICME4
2025 SemConf: A System for Multiparty Semantic Video Conferencing
abstract
Multi-party real-time video conferencing has become an indispensable service in industrial production and daily life. However, the current dynamic and limited network resources can no longer meet the growing service demands of users, resulting lagging and low visual quality. The emerging semantic transmission, together with the network-wide redundant computation capacity, provides new opportunities towards a new paradigm of semantic video conferencing. The key challenge of such fusion lies in the interplay of traditional streaming adaptation and the new semantic processing, calling for a holistic mechanism to optimize the service provision with compatibility and efficiency. In this paper, we for the first time address this challenge, and propose SemConf, a novel framework that integrate the semantic transmission into the video conferencing towards optimal user QoE. Our extensive evaluations, against state-of-the-art baselines, reveal that SemConf achieves a substantial improvement in QoE, with up to 33.6% enhancement in bandwidth-constrained environment. Overall, this work highlights the critical role of the coordination algorithm in balancing computational load and network throughput, showcasing SemConf as a transformative approach in the realm of semantic video conferencing.
Xize Duan, Yili Jin 0001, Lei Zhang 0066, Fangxin Wang 0001
NOSSDAV4
2025 SRBF-Gaussian: Streaming-Optimized 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has emerged as a groundbreaking 3D scene representation technique, offering unprecedented visual quality and rendering efficiency. However, the substantial data volume of 3DGS scenes poses significant challenges for streaming applications. Existing research on 3DGS has primarily focused on compression and rendering efficiency, neglecting the specific requirements of streaming transmission. Moreover, the Spherical Harmonics color representation in 3DGS complicates viewport-based transmission partitioning. Achieving hierarchical Gaussian streaming without noticeable quality degradation also remains a significant challenge.To address these challenges, we propose SRBF-Gaussian, a new paradigm that revolutionizes the traditional 3DGS format. Our approach introduces viewport-dependent color encoding based on Spherical Radial Basis Functions (SRBFs) and HSL color space, enabling selective transmission of viewport-relevant color data. This reduces data transmission while maintaining visual quality. We implement adaptive Gaussian pruning and transmission, optimized for current viewports and network conditions. Additionally, we develop coherent multi-level Gaussian representations for smooth transitions between quality levels. Our system incorporates user-behavior-aware streaming strategies to anticipate and pre-fetch relevant data. In cloud VR scenarios, our approach demonstrates substantial improvements, achieving a 5.63% - 14.17% increase in PSNR, a 7.61% - 59.16% reduction in latency, and a 10.45% - 30.12% improvement in overall Quality of Experience (QoE).
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
VR5
2025 LiveVV: Human-Centered Live Volumetric Video Streaming System
abstract
Volumetric video (VV) has emerged as a prominent medium within the realm of extended reality (XR) with advancements in computer graphics and depth capture hardware. Users can fully immersive themselves in VV with the ability to switch their viewport in six degree of freedom (DOF), including three rotational dimensions (yaw, pitch, and roll) and three translational dimensions (X, Y, and Z). Different from traditional 2-D videos that are composed of pixel matrices, VVs employ point clouds, meshes, or voxels to represent a volumetric scene, resulting in significantly larger data sizes. While previous works have successfully achieved VV streaming in video-on-demand scenarios, the live streaming of VV remains an unresolved challenge due to the limited network bandwidth and stringent latency constraints. In this article, we proposeLiveVV, a holistic live VV streaming system that integrates multiview capture, scene segmentation and reuse, adaptive transmission, and real-time rendering.LiveVVfeatures lightweight VV capture modules for easy deployment, processes static and dynamic content separately to reduce bandwidth consumption, and incorporates a VV adaptive bitrate streaming algorithm (VABR) to ensure fluent playback with high-quality experience. Real-world implementation and evaluation demonstrate thatLiveVVachieves live VV streaming at 24 FPS frame rate with less than 350-ms latency on average, meeting the requirements of real-life application.
Kaiyuan Hu, Yongting Chen, Kaiying Han, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001
IEEE Internet Things J.8
2025 3DGStreaming: Spatial-Heterogeneity-Aware 3-D Gaussian Splatting Compression and Streaming
abstract
3-D Gaussian splatting (3DGS) has emerged as a promising technique for high-quality 3-D scene representation. However, streaming 3DGS scenes poses significant challenges due to large data volumes and complex spatial structures, resulting in nonsmooth scene loading, inferior visual quality, and ineffective streaming adaptation, ultimately impacting user experience adversely. To tackle these challenges and enhance user Quality of Experience (QoE), this article introduces a novel adaptive streaming framework called 3DGStreaming. Our framework comprises three key components: 1) smart Spatial Partitioning for efficient scene division, enabling selective streaming and seamless scene merging; 2) two-step Progressive Scene Generation, involving content-aware downsampling and attribute compression to create multibitrate 3DGS scenes; and 3) Field of View (FoV)-based Bitrate Adaptation using a decision transformer for viewport-based bitrate selection. Extensive experiments demonstrate the superiority of 3DGStreaming over existing state-of-the-art solutions. 3DGStreaming achieves a greater rendering quality with a 5.7%–25.5% increase in PSNR, a 27.8%–69.2% reduction in latency, a 54.7% reduction in training time, and a 17.2%–68.3% improvement in overall QoE.
Dayou Zhang, Zhicheng Liang, Zijian Cao 0007, Dan Wang 0002, Fangxin Wang 0001
IEEE Internet Things J.6
2025 Energy-Efficient Joint Beamforming and Trajectory Optimization for UAV-Enabled Integrated Sensing and Communication
abstract
Uncrewed aerial vehicle (UAV)-enabled ISAC systems have received widespread attention due to the high mobility of UAVs with good line-of-sight (LoS) paths to ensure communication and sensing performance. However, the existing works on UAV-enabled ISAC mainly focus on optimizing communication performance (e.g., sum rate) and sensing performance, resulting in excessive energy consumption and reducing the flight endurance of the UAV. Motivated by this, we draw a trade-off between such performance and energy consumption to achieve robust and efficient UAV-enabled ISAC. In this work, we aim to maximize the worst-case energy efficiency in UAV-enabled ISAC by jointly designing the beamforming and the UAV trajectory, while ensuring the UAV energy constraints and the ISAC performance. Nevertheless, solving this problem is non-trivial due to its non-convex nature, and the high coupling of the transmit beamforming vectors and the UAV dynamics adds an additional layer of complexity. To effectively address this non-convex issue, we alternately optimize the transmit communication and sense beamforming, as well as the UAV dynamic variables to obtain a sub-optimal solution, and the algorithm complexity is lower than the existing algorithms. Experimental results show a trade-off between energy efficiency and average sum rate. Furthermore, they indicate the superiority of the proposed algorithm to enhance energy efficiency by significantly reducing energy consumption without causing excessive sum rate loss.
Boxin He, Wencan Mao, Yaxi Liu 0001, Wei Huangfu, Yu Xiao 0001, Fangxin Wang 0001, Yusheng Ji
IEEE Trans. Commun.6
2025 On-Demand Edge Computing Power Networks Assisted by Reconfigurable Intelligent Surface With Multi-Layer Scheme
abstract
On-demand edge computing power networks with both stationary fog nodes co-located with cellular base stations (CFNs) and mobile fog nodes mounted on vehicles (VFNs) provide promising solutions for coping with high spatio-temporal, compute-intensive, and latency-sensitive applications. Joint scheduling and resource allocation in such a network is challenging due to the trade-off between quality of service (QoS) and energy consumption, limited onboard capacity of IoT devices and fog nodes, and urban obstructions that impede line-of-sight links. To address these issues, this work envisions a network assisted by reconfigurable intelligent surface (RIS) with a multi-layer scheme. The computation tasks are offloaded from IoT devices to VFNs and further to CFNs based on the computational demand and latency requirements, and the RIS assists with wireless communication on both links. We jointly optimized the allocation of the subcarriers, the power, the offloading task bits, the time slot, and the RIS beamforming vectors under the constraints of task input bits and computing capability, to minimize the average energy consumption. To address the non-convex issue, we first decompose it into three sub-problems, and then alternately optimize these sub-problems by adopting successive convex approximation (SCA) where a locally optimal solution can be obtained. Simulation results demonstrate the superiority of the proposed offloading strategy where RIS with a multi-layer scheme is introduced in the on-demand edge computing power networks. Also, the effectiveness, feasibility, scalability, and adaptability of the designed algorithm are verified.
Boxin He, Wencan Mao, Yaxi Liu 0001, Fangxin Wang 0001, Wei Huangfu
IEEE Trans. Commun.4
2025 3D Video Conferencing via On-Hand Devices
abstract
Video conferencing has become indispensable in human communication. Researchers are exploring immersive capabilities to enhance video conferencing experiences by delivering realistic interactions. However, existing methods have stringent and extra hardware beyond a typical video conference, including multiple depth cameras, large screens, and headsets, which pose obstacles to the widespread adoption due to high costs and complex setups. Thus, there is an urgent demand for light-weight systems using only on-hand devices including single RGB camera and standard screen, without additional hardware. We propose DVCO, a novel 3D video conferencing system via on-hand devices. With DVCO, users can experience lifelike virtual conferencing that includes natural contact and interactive features. To achieve this, DVCO has two main components. Virtual Camera Transformation (VCT) and New View Generator (NVG). VCT computes a downscaled sender image from tracking to determine viewpoint and gaze vector, enhancing virtual presence on standard screens. NVG takes an input frame and desired view angle to produce an output reflecting the new view from a single RGB camera. Together, these provide an affordable, easy-to-integrate enhancement for current video conferencing systems without expensive upgrades. Through a user study, it has been demonstrated that DVCO offers an exceptional level of immersion when compared to traditional systems. Experiments are conducted to showcase the superior performance of VCT and NVG in comparison to baseline methods.
Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu
IEEE Trans. Circuits Syst. Video Technol.4
2025 Multi-Task Reinforcement Learning-Based Multiple Access for Dynamic Wireless Networks
abstract
With the rapid development of emergent applications, wireless networks require the provision of high throughput. Meanwhile, wireless scenarios exhibit highly dynamic characteristics, involving frequent changes in the network scale and traffic. To satisfy the high demand for new applications in dynamic wireless scenarios, a novel medium access control (MAC) protocol is required to allow stations to access the channel with high efficiency and adaptability. Based on multi-agent reinforcement learning (MARL), we propose a new MAC protocol, Multi-task Transformer-based Multiple Access (MTMA). Multi-task learning is applied to train a single actor to adapt to multiple wireless environments simultaneously. To improve the scalability, we propose a transformer-based critic network, which can scale to different wireless scenarios. Moreover, a novel network called “Generalization for N (Gen-N)” network is proposed to enhance the generalization ability. We conduct simulation experiments to demonstrate that MTMA: 1) achieves over 95% of upper bound of throughput while maximizing the fairness performance; 2) outperforms classic MAC protocol and MARL-based baselines in scenarios with saturated and light traffic; 3) can adapt to environmental changes quickly in dynamic scenarios; 4) can generalize to unseen scenarios during training. Finally, the ablation experiments are conducted to evaluate the effectiveness of components used in MTMA.
Xinghua Sun, Yili Jin 0001, Fangxin Wang 0001
IEEE Trans. Mob. Comput.4
2025 Video Conferencing With Predictive Generation and Collaborative Computation Across Mobile Headsets
abstract
Virtual Reality (VR) has emerged as a transformative platform for remote collaboration, but its adoption for video conferencing is hindered by challenges related to facial expression reconstruction and computational resource constraints, especially on economical mobile VR headsets. This paper introduces a novel system for VR video conferencing that addresses these challenges through two key modules: Predictive Generation and Collaborative Computation. Predictive Generation leverages multimodal inputs, including voice, head motion, and eye blinks, to synthesize realistic facial animations with low latency, eliminating the need for high-precision hardware. Collaborative Computation enhances computational efficiency by employing a game-theoretic framework for resource sharing among users. Experimental evaluations demonstrate that our system delivers immersive and realistic VR video conferencing experiences with superior facial expression reconstruction and efficient resource utilization. Our approach makes VR video conferencing more accessible and practical for a broader audience across mobile headsets.
Yili Jin 0001, Xize Duan, Kaiyuan Hu, Fangxin Wang 0001, Xue (Steve) Liu, Jiangchuan Liu
IEEE Trans. Mob. Comput.4
2025 Toward Universal Personalization in Federated Learning via Collaborative Foundation Generative Models
abstract
Personalized federated learning (PFL) enhances the performance of customized client models through collaborative training without compromising data privacy and ownership. Some previous PFL methods rely on rich prior knowledge about the types of data heterogeneity (such as class imbalance or feature skew), which greatly limits their application ranges. In this paper, we study theUniversal Personalization in Federated Learning (UniPFL), the problem that has no prior knowledge about the types of data heterogeneity. In real-world PFL scenarios, UniPFL is potential because the data distributions of clients are usually heterogeneous and unknown to the server, where quantity imbalance, class imbalance, feature skew, or hybrid heterogeneity are possible contingencies. To address UniPFL, we proposeFedFD, a novel framework with local data augmentation and global concept fusion, which is based on the recent advances inthe foundation generative models(e.g., diffusion models, BLIP-2). On the client side, FedFD utilizes a diffusion model to assist local training by generating augmented data samples, and is then efficiently fine-tuned to be personalized. On the server side, we customize the aggregation strategies based on model similarities to learn both personalized models and diverse feature concepts. Extensive experiments show that FedFD reaches the state-of-the-art on (1) CIFAR-10 and CIFAR-100 for class imbalance; (2) DomainNet and Office-10 for feature skew, and (3) hybrid heterogeneity with both class and feature shifts.
Chenrui Wu 0002, Zexi Li 0001, Fangxin Wang 0001, Hongyang Chen 0001, Jiajun Bu, Haishuai Wang
IEEE Trans. Mob. Comput.3
2025 MANSY: Generalizing Neural Adaptive Immersive Video Streaming With Ensemble and Representation Learning
abstract
The popularity of immersive videos has prompted extensive research into neural adaptive tile-based streaming to optimize video transmission over networks with limited bandwidth. However, the diversity of users’ viewing patterns and Quality of Experience (QoE) preferences has not been fully addressed yet by existing neural adaptive approaches for viewport prediction and bitrate selection. Their performance can significantly deteriorate when users’ actual viewing patterns and QoE preferences differ considerably from those observed during the training phase, resulting in poor generalization. In this paper, we proposeMANSY, a novel streaming system that embraces user diversity to improve generalization. Specifically, to accommodate users’ diverse viewing patterns, we design a Transformer-based viewport prediction model with an efficient multi-viewport trajectory input output architecture based on implicit ensemble learning. Besides, we for the first time combine the advanced representation learning and deep reinforcement learning to train the bitrate selection model to maximize diverse QoE objectives, enabling the model to generalize across users with diverse preferences. Extensive experiments demonstrate thatMANSYoutperforms state-of-the-art approaches in viewport prediction accuracy and QoE improvement on both trained and unseen viewing patterns and QoE preferences, achieving better generalization.
Duo Wu, Panlong Wu, Miao Zhang 0003, Fangxin Wang 0001
IEEE Trans. Mob. Comput.4
2025 Unimodal Training-Multimodal Prediction: Cross-Modal Federated Learning With Hierarchical Aggregation
abstract
Multimodal learning has significantly advanced the extraction of features from varied data sources, enhancing model performance. Federated learning (FL) complements this by enabling collaborative training while maintaining data privacy. The fusion of these two fields, multimodal federated learning, offers considerable promise. Yet, standard methods often incorrectly assume that each node in the FL network has a full complement of multimodal data, which is rare in real-world applications. In our study, we present a novel architecture designed to surmount these challenges, termed the Unimodal Training - Multimodal Prediction (UTMP) framework, positioned within the multimodal federated learning paradigm. Our proposed model, the HA-Fedformer, is a transformer-based model crafted to facilitate unimodal training on the client-side using exclusively unimodal datasets and to execute multimodal inference by synthesizing insights from multiple clients. Our HA-Fedformer model effectively handles non-IID data through a novel uncertainty-aware aggregation technique and layer-wise Markov Chain Monte Carlo sampling in local encoders. It also resolves misaligned language sequences via cross-modal decoder aggregation, capturing correlations between decoders trained on different modalities. Our comprehensive evaluations conducted on widely recognized sentiment analysis benchmarks demonstrate the superiority of the HA-Fedformer. The results show that our model achieves a substantial uplift in performance.
Rongyu Zhang, Xiaowei Chi, Guiliang Liu, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.6
2025 RepCaM++: Exploring Transparent Visual Prompt With Inference-Time Re-Parameterization for Neural Video Delivery
abstract
Recently, content-aware methods have been employed to reduce bandwidth and enhance the quality of Internet video delivery. These methods involve training distinct content-aware super-resolution (SR) models for each video chunk on the server, subsequently streaming the low-resolution (LR) video chunks with the SR models to the client. Prior research has incorporated additional partial parameters to customize the models for individual video chunks. However, this leads to parameter accumulation and can fail to adapt appropriately as video lengths increase, resulting in increased delivery costs and reduced performance. In this paper, we introduce RepCaM++, an innovative framework based on a novel Re- parameterization Content-aware Modulation (RepCaM) module that uniformly modulates video chunks. The RepCaM framework integrates extra parallel-cascade parameters during training to accommodate multiple chunks, subsequently eliminating these additional parameters through re- parameterization during inference. Furthermore, to enhance RepCaM's performance, we propose the Transparent Visual Prompt (TVP), which includes a minimal set of zero-initialized image-level parameters (e.g., less than 0.1%) to capture fine details within video chunks. We conduct extensive experiments on the VSD4K dataset, encompassing six different video scenes, and achieve state-of-the-art results in video restoration quality and delivery bandwidth compression.
Rongyu Zhang, Xize Duan, Jiaming Liu 0003, Yuan Du, Dan Wang 0002, Shanghang Zhang, Fangxin Wang 0001
IEEE Trans. Mob. Comput.8
2024 MMCOUNT: Stationary Crowd Counting System Based on Commodity Millimeter-Wave Radar
abstract
Millimeter wave sensing promises the capability of sensing the surrounding moving people. However, it is still challenging for stationary crowds because objects with few motions (like changing sitting position) are easily treated as a cluster of noise and thus neglected. In this paper, we propose that people’s respiration and natural fidgeting (restless behavior) carry valuable information, which could be captured by millimeter (mmWave) radar. By performing processing on the captured data including signal enhancement and object recognition, we can successfully extract the number of a crowd and the position of each individual. To verify our system, we test it in different locations like hall, classroom, and meeting room to simulate different practical scenarios including watching a movie, having a class, or attending a meeting. The evaluation results show that our proposed approach could reach a high counting accuracy of up to 95.8% even at small separation distances of 0.4m.
Kaiyuan Hu, Hongjie Liao, Fangxin Wang 0001
ICASSP4
2024 HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
abstract
Virtual Reality (VR) has become increasingly popular for remote collaboration, but video conferencing poses challenges when the user's face is covered by the headset. Existing solutions have limitations in terms of accessibility. In this paper, we propose HeadsetOff, a novel system that achieves photorealistic video conferencing on economical VR headsets by leveraging voice-driven face reconstruction. HeadsetOff consists of three main components: a multimodal predictor, a generator, and an adaptive controller. The predictor effectively predicts user future behavior based on different modalities. The generator employs voice, head motion, and eye blink to animate the human face. The adaptive controller dynamically selects the appropriate generator model based on the trade-off between video quality and delay. Experimental results demonstrate the effectiveness of HeadsetOff in achieving high-quality, low-latency video conferencing on economical VR headsets.
Yili Jin 0001, Xize Duan, Fangxin Wang 0001, Xue (Steve) Liu
ACM Multimedia3
2024 NetLLM: Adapting Large Language Models for Networking
abstract
Many networking tasks now employ deep learning (DL) to solve complex prediction and optimization problems. However, current design philosophy of DL-based algorithms entails intensive engineering overhead due to the manual design of deep neural networks (DNNs) for different networking tasks. Besides, DNNs tend to achieve poor generalization performance on unseen data distributions/environments.
Duo Wu, Xianda Wang, Yaqi Qiao, Zhi Wang 0001, Junchen Jiang, Shuguang Cui, Fangxin Wang 0001
SIGCOMM7
2024 FewVV: Few-Shot Adaptive Bitrate Volumetric Video Streaming With Prompted Online Adaptation
abstract
In recent years, volumetric videos have brought immersive experiences to users. Existing viewport-based volumetric video streaming (VVS) systems prune the point cloud according to visibility to reduce bandwidth consumption, leading to a better responsiveness. They also predict bandwidth and allocate bitrate to different parts of the video to enhance Quality-of-Experience (QoE). However, such designs sometimes result in drastic quality fluctuations in real-world deployment, due to limited generalization performance. Our measurement notes that these systems tend to have a significant accuracy loss under an unseen Out-of-Distribution (OoD) environments. On the other hand, open world prediction/adaptation problem have been addressed in the recent reinforcement learning advances, particularly through prompt-based few-shot and zero-shot learning. Inspired by this development, in this work, we first reformulate the volumetric bitrate adaptation (volumetric ABR) into a sequence prediction problem, then we design a volumetric causal transformer algorithm to solve it. We train our model on a large action trajectory data set, then evaluate it on various OoD scenarios. The result show that FewVV consistently outperforms the existing systems on both performance and generalization.
Fangxin Wang 0001, Dan Wang 0002
IEEE Internet Things J.3
2024 VSAS: Decision Transformer-Based On-Demand Volumetric Video Streaming With Passive Frame Dropping
abstract
Volumetric video is becoming a popular application among various multimedia services, which is envisioned as a fundamental technology for VR, AR, and the emerging metaverse. The commodity RGB-D cameras provide an affordable solution for volumetric video capture, and the VR headset allows immersive and interactive display. From the networking perspective, the primary challenge lies in the smooth and high-quality transmission over the Internet, given the enormous data volume and limited bandwidth. MPEG V-PCC standard has stood out recently as a promising compression and streaming solution that can effectively reduce video size while maintaining high-visual quality. Since MPEG V-PCC is largely backward compatible with the 2-D video compression standard like H.264/AVC, it is natural to use DASH, the most widely used 2-D streaming framework, to stream it. We first propose an integrated framework based on DASH to support MPEG V-PCC Internet streaming. During this, we faced three challenges. First, the lack of a rate-distortion model for MPEG V-PCC encoding parameters. Second, the need for a new bitrate adaptation controller that not only considers the rate of the chunks but also chooses the chunks with proper frame rate. We align the decision transformer to this problem, which expands the success of transformer-based models in natural language processing to the decision problems. Finally, to solve the stalling time issue inherited from DASH, we use a frame-dropping mechanism to eliminate the stalling in DASH playback. Our evaluations show that VSAS achieves an average acrlong QoE improvement of$1.67\times $over a range of network conditions.
Fangxin Wang 0001, Dayou Zhang, Dan Wang 0002
IEEE Internet Things J.3
2024 Proffler: Toward Collaborative and Scalable Edge-Assisted Crowdsourced Livecast
abstract
In recent years, crowdsourced livecast has seen remarkable progress due to the interactivity and real-time nature, playing an essential role in multimedia applications in the post-epidemic era. Given the delay sensitivity, large viewing volumes, and heterogeneous viewing patterns, the traditional video streaming methods fail to provide the optimized quality of experience (QoE) for viewers using the minimum system cost over an edge-assisted service architecture. The emerging technology of mobile edge computing (MEC) offers a new perspective of reducing user latency and enhancing the quality of dispatched videos in a promising way. In this paper, we propose Proffler, an integrated framework that addresses this problem through effective stream caching at the network edge server. We first examine the underlying correlations in viewing patterns across different regions and propose a novel transformer-based algorithm, Chili-TF, that achieves accurate viewer request prediction, even for regions with insufficient data. We then design a scalable algorithm, U2VR, that achieves near-optimal video stream allocation as well as viewer scheduling. Extensive real-data-driven experiments further confirm that Proffler can achieve improvements of 20%-55% in average QoE compared to state-of-the-art solutions.
Fangxin Wang 0001, Jiangchuan Liu
IEEE Internet Things J.3
2024 On AoI of Grant-Free Access With HARQ
abstract
For mission-critical URLLC applications, timely status updates are essential. This paper investigates the age of information (AoI) of the three HARQ schemes specified in 5G R16, targeting to provide guidelines for future grant-free access design in 5G-Advanced and beyond. Specifically, we analyze two packet management policies: First-come-first-serve (FCFS) and preemption policy where new packets always preempt the buffer. We also study the AoI in a latency-sensitive scenario where expired packets are discarded. We derive exact expressions of AoI and peak AoI for all schemes and their lower bounds, revealing that the number of the maximum consecutive transmissions is critical for information freshness. Simulation results validate the theoretical analysis and show that Proactive HARQ scheme outperforms K-repetition HARQ scheme unconditionally and Reactive HARQ scheme with moderate system load or above. And discarding expired packets enhances system robustness for overload systems but yields larger AoI.
Jiwen Wang, Ju Ren 0001, Fangxin Wang 0001, Shuai Wang 0013, Jihong Yu
IEEE Trans. Commun.4
2024 Privacy-Preserving Gaze-Assisted Immersive Video Streaming
abstract
Immersive videos, also known as 360$^{\circ }$videos, have gained significant attention in recent years due to their ability to provide an interactive and engaging experience. However, the development of immersive video streaming faces several challenges, including privacy concerns, the need for accurate viewport prediction, and efficient bandwidth allocation. In this paper, we propose a comprehensive system that integrates three specialized modules: the Privacy Protection module, the Viewport Prediction module, and the Bitrate Allocation module. The Privacy Protection module introduces a novel approach to differential privacy tailored for immersive video environments, considering the spatial and temporal correlations in viewport and gaze motion data. The Viewport Prediction module leverages a crossmodal attention mechanism based on the transformer to predict user viewport movements by analyzing the complex interactions between historical data, video content, and gaze patterns. The Bitrate Allocation module employs an adaptive tile-based bitrate allocation strategy using an exponential decay function to optimize video quality and maximize user quality of experience. Experimental results demonstrate that our proposed framework outperforms three state-of-the-art integrated frameworks, achieving an average QoE improvement of 21.61%. This paper offers substantial novelty in addressing privacy concerns, leveraging gaze information for viewport prediction, and utilizing underlying correlations between different features.
Yili Jin 0001, Wenyi Morty Zhang, Fangxin Wang 0001, Xue (Steve) Liu
IEEE Trans. Mob. Comput.4
2024 LMaaS: Exploring Pricing Strategy of Large Model as a Service for Communication
abstract
One of the most important features of next-generation communication is to incorporate intelligence towards semantic communication, where highly condensed semantic information considering both source and channel features will be extracted and transmitted. The recent popular large models such as GPT4 and the boosting learning techniques are envisioned to accelerate its practical implementation in the near future. Given the characteristics of “training once and widely use” of those multimodal large language models, we argue that a pay-as-you-go service mode will be suitable in this context, referred to as Large Model as a Service (LMaaS). However, the trading and pricing problem is quite complex with heterogeneous and dynamic customer environments, making the pricing optimization problem challenging in seeking on-hand solutions. In this paper, we optimize the profit of both the seller and customers. We formulate the LMaaS market trading as a Stackelberg game with two steps. In the first step, we optimize the seller's pricing decision and propose an Iterative Model Pricing (IMP) algorithm that optimizes the prices of large models iteratively by reasoning customers’ future rental decisions, which achieves a near-optimal pricing solution. In the second step, we optimize customers’ selection decisions by designing a robust selecting and renting (RSR) algorithm, which is guaranteed to be optimal with rigorous theoretical proof. Extensive experiments confirm the effectiveness and robustness of our algorithms, outperforming the state-of-the-art solution by 43.96% in profit at the customer side and achieving the near-optimal profit at the seller side.
Panlong Wu, Yanjie Dong 0003, Zhaorui Wang 0001, Fangxin Wang 0001
IEEE Trans. Mob. Comput.5
2024 FedFMSL: Federated Learning of Foundation Models With Sparsely Activated LoRA
abstract
Foundation models (FMs) have shown great success in natural language processing, computer vision, and multimodal tasks. FMs have a large number of model parameters, thus requiring a substantial amount of data to help optimize the model during the training. Federated learning has revolutionized machine learning by enabling collaborative learning from decentralized data while still preserving clients’ data privacy. Despite the great benefits foundation models can have empowered by federated learning, their bulky model parameters cause severe communication challenges for modern networks and computation challenges especially for edge devices. Moreover, the data distribution of different clients can be different thus inducing statistical challenges. In this paper, we propose a novel two-stage federated learning algorithm called FedFMSL. A global expert is trained in the first stage and a local expert is trained in the second stage to provide better personalization. We construct a Mixture of Foundation Models (MoFM) with these two experts and design a gate neural network with an inserted gate adapter that joins the aggregation every communication round in the second stage. To further adapt to edge computing scenarios with limited computational resources, we design a novel Sparsely Activated LoRA (SAL) algorithm that freezes the pre-trained foundation model parameters inserts low-rank adaptation matrices into transformer blocks, and activates them progressively during the training. We employ extensive experiments to verify the effectiveness of FedFMSL, results show that FedFMSL outperforms other SOTA baselines by up to 59.19% in default settings while tuning less than 0.3% parameters of the foundation model.
Panlong Wu, Kangshuo Li, Yanjie Dong 0003, Victor C. M. Leung, Fangxin Wang 0001
IEEE Trans. Mob. Comput.6
2024 ILCAS: Imitation Learning-Based Configuration- Adaptive Streaming for Live Video Analytics With Cross-Camera Collaboration
abstract
The high-accuracy and resource-intensive deep neural networks (DNNs) have been widely adopted by live video analytics (VA), where camera videos are streamed over the network to resource-rich edge/cloud servers for DNN inference. Common video encoding configurations (e.g., resolution and frame rate) have been identified with significant impacts on striking the balance between bandwidth consumption and inference accuracy and therefore their adaption scheme has been a focus of optimization. However, previous profiling-based solutions suffer from high profiling cost, while existing deep reinforcement learning (DRL) based solutions may achieve poor performance due to the usage of fixed reward function for training the agent, which fails to craft the application goals in various scenarios. In this paper, we proposeILCAS, the first imitation learning (IL) based configuration-adaptive VA streaming system. Unlike DRL-based solutions,ILCAStrains the agent with demonstrations collected from the expert which is designed as an offline optimal policy that solves the configuration adaption problem through dynamic programming. To tackle the challenge of video content dynamics,ILCASderives motion feature maps based on motion vectors which allowILCASto visually “perceive” video content changes. Moreover,ILCASincorporates a cross-camera collaboration scheme to exploit the spatio-temporal correlations of cameras for more proper configuration selection. Extensive experiments confirm the superiority ofILCAScompared with state-of-the-art solutions, with 2-20.9% improvement of mean accuracy and 19.9–85.3% reduction of chunk upload lag.
Duo Wu, Dayou Zhang, Miao Zhang 0003, Fangxin Wang 0001, Shuguang Cui
IEEE Trans. Mob. Comput.5
2024 Multi-Level Personalized Federated Learning on Heterogeneous and Long-Tailed Data
abstract
Federated learning (FL) offers a privacy-centric distributed learning framework, enabling model training on individual clients and central aggregation without necessitating data exchange. Nonetheless, FL implementations often suffer from non-i.i.d. and long-tailed class distributions across mobile applications, e.g., autonomous vehicles, which leads models to overfitting as local training may converge to sub-optimal. In our study, we explore the impact of data heterogeneity on model bias and introduce an innovative personalized FL framework, Multi-level Personalized Federated Learning (MuPFL), which leverages the hierarchical architecture of FL to fully harness computational resources at various levels. This framework integrates three pivotal modules: Biased Activation Value Dropout (BAVD) to mitigate overfitting and accelerate training; Adaptive Cluster-based Model Update (ACMU) to refine local models ensuring coherent global aggregation; and Prior Knowledge-assisted Classifier Fine-tuning (PKCF) to bolster classification and personalize models in accord with skewed local data with shared knowledge. Extensive experiments on diverse real-world datasets for image classification and semantic segmentation validate thatMuPFLconsistently outperforms state-of-the-art baselines, even under extreme non-i.i.d. and long-tail conditions, which enhances accuracy by as much as 7.39% and accelerates training by up to 80% at most, marking significant advancements in both efficiency and effectiveness.
Rongyu Zhang, Chenrui Wu 0002, Fangxin Wang 0001, Bo Li 0001
IEEE Trans. Mob. Comput.4
2024 TrimStream: Adaptive Realtime Video Streaming Through Intelligent Frame Retrospection in Adverse Network Conditions
abstract
Realtime video streaming (RVS) services are gaining popularity in various applications such as video conferencing, online education, and mixed reality. However, adverse network conditions can significantly damage video transmission, leading to a decline in users' Quality of Experience (QoE). Existing approaches have made considerable efforts to address these problems, including bitrate adaptation, FEC (forward error correction) encoding, and super-resolution techniques. Nevertheless, these methods either focus solely on adjusting transmission configurations (ABR) or consume additional network and computational resources to enhance QoE (FEC or super-resolution), making them suboptimal for adverse network conditions. In this paper, we analyze the limitations of conventional RVS systems when confronted with adverse network conditions and proposeTrimStream, a novel RVS solution based on intelligent frame retrospection, to effectively handle such scenarios. Our approach leverages the high similarity observed between frames in realtime video streaming. The core idea is to store a subset of correctly received frames and exploit frame similarity to minimize transmission while breaking down frame-level dependencies. We formulate the frame caching problem to maximize QoE in RVS and present an online frame cache algorithm. Furthermore, we design a vision-transformer-based, cost-effective frame matching framework that combines different levels of frame information. Our evaluation results demonstrate thatTrimStreamoutperforms state-of-the-art solutions by$14.8\% \sim 21.1\%$improvement in overall QoE.
Dayou Zhang, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.6
2024 DSJA: Distributed Server-Driven Joint Route Scheduling and Streaming Adaptation for Multi-Party Realtime Video Streaming
abstract
The widespread availability of convenient wireless network connection and video capture have fueled the development of multi-party realtime video streaming (MRVS) services, such as Zoom or Microsoft Teams. These services have transformed the generation and distribution of realtime streaming content and offer a new way of online communication, striving to provide high Quality-of-Experience (QoE) for individuals. However, delivering high QoE in MRVS is more challenging than in traditional video scenarios due to the stringent delay requirements and complex multi-party interactive architectures. In this paper, we propose DSJA, a distributed server-driven multi-party realtime video streaming framework that conquers the challenges. We first design an appropriate QoE model for MRVS services to capture the interplay among perceptual quality, variations, bitrate mismatch, loss damage, and streaming delay. We then model the QoE maximization problem in MRVS as a route scheduling and streaming adaptation problem. Afterward, we design DSJA which seamlessly integrates multiple selective forwarding units (SFU) architecture and server-driven approaches based on a two-step solution of route scheduling and streaming adaptation. DSJA first determines the most suitable SFU and streaming routes for each video session based on SFUs' job queuing delay and path latency. Then, the server conducts joint loss and bitrate adaptation decisions to optimize the streaming configuration of all clients, considering network conditions and QoE preferences. Our evaluations show that our framework outperforms state-of-the-art solutions by$23.1\% \sim 41.7\%$from the perspective of QoE, and reduces the backbone network transmission by$14.0\% \sim 36.6\%$.
Dayou Zhang, Dan Wang 0002, Fangxin Wang 0001
IEEE Trans. Mob. Comput.5
2024 Fumos: Neural Compression and Progressive Refinement for Continuous Point Cloud Video Streaming
abstract
Point cloud video (PCV) offers watching experiences in photorealistic 3D scenes with six-degree-of-freedom (6-DoF), enabling a variety of VR and AR applications. The user's Field of View (FoV) is more fickle with 6-DoF movement than 3-DoF movement in 360-degree video. PCV streaming is extremely bandwidth-intensive. However, current streaming systems require hundreds of Mbps bandwidth, exceeding the bandwidth capabilities of commodity devices. To save bandwidth, FoV-adaptive streaming predicts a user's FoV and only downloads point cloud data falling in the predicted FoV. But it is difficult to accurately predict the user's FoV even 2-3 seconds before playback due to 6-DoF. Misprediction of FoV or network bandwidth dips results in frequent stalls. To avoid rebuffering, existing systems would cause incomplete FoV and degraded experience, deteriorating the user's quality of experience (QoE). In this paper, we describe Fumos, a novel system that preserves interactive experience by avoiding playback stalls while maintaining high perceptual quality and high compression rate. We find a research gap in inter-frame redundant utilization and progressive mechaism. Fumos has three crucial designs, including (1) Neural compression framework with inter-frame coding, namely N-PCC, which achieves both bandwidth efficiency and high fidelity. (2) Progressive refinement streaming framework that enables continuous playback by incrementally upgrading a fetched portion to a higher quality (3) System-level adaptation that employs Lyapunov optimization to jointly optimize the long-term user QoE. Experimental results demonstrate that Fumos significantly outperforms Draco, achieving an average decoding rate acceleration of over 260×. Moreover, the proposed compression framework N-PCC attains remarkable BD-Rate gains, averaging 91.7% and 51.7% against the state-of-the-art point cloud compression methods G-PCC and V-PCC, respectively.
Zhicheng Liang, Junhua Liu 0003, Mallesham Dasari, Fangxin Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2023 Learning Cautiously in Federated Learning with Noisy and Heterogeneous Clients
abstract
Federated learning (FL) is a distributed framework for collaborative training with privacy guarantees. In real-world scenarios, clients may have Non-IID data (local class imbalance) with poor annotation quality (label noise). The co-existence of label noise and class imbalance in FL’s small local datasets renders conventional FL methods and noisy-label learning methods both ineffective. To address the challenges, we propose FEDCNI without using an additional clean proxy dataset. It includes a noise-resilient local solver and a robust global aggregator. For the local solver, we design a more robust prototypical noise detector to distinguish noisy samples. Further to reduce the negative impact brought by the noisy samples, we devise a curriculum pseudo labeling method and a denoise Mixup training strategy. For the global aggregator, we propose a switching re-weighted aggregation method tailored to different learning periods. Extensive experiments demonstrate our method can substantially outperform state-of-the-art solutions in mix-heterogeneous FL environments.
Chenrui Wu 0002, Zexi Li 0001, Fangxin Wang 0001, Chao Wu 0001
ICME3
2023 Cluster-driven GNN-based Federated Recommendation with Biased Message Dropout
abstract
Due to the remarkable ability to model the high-order links within user-item relations, the graph neural network (GNN) is gradually applied to personalized recommendations in many online services. Besides, federated learning (FL) recently emerged as a powerful framework that enables collaborative training while protecting user data privacy. However, the integration of GNN and FL still exists with vital challenges unsolved, e.g., learning from non-IID local sub-graphs with only low-order user-item interactions jointly and overcoming the over-fitting problems with high training efficiency. In this paper, we propose CdFed, a Cluster-driven GNN-based Federated Learning framework, to address the GNN+FL challenges. CdFedhas two major components. First, to learn from non-IID sub-graphs, we design an Adaptive Model Clustering (AMC) strategy that takes advantage of the similarity across the uploaded model weights and updates clusters adaptively in each communication round. Second, we develop a Biased Message Dropout (BMD) strategy to combat the overfitting problem and accelerate the training process of federated learning. Together with AMC and BMD, CdFedcan implicitly complete the missing links between sub-graphs more efficiently and greatly improve the model’s generalization ability in non-IID scenarios. We have conducted extensive evaluations and the results reveal that our proposed approaches can outperform the SOTA solution by 24% in model performance and 8x in training speed.
Rongyu Zhang, Chenrui Wu 0002, Fangxin Wang 0001
ICME4
2023 SJA: Server-driven Joint Adaptation of Loss and Bitrate for Multi-Party Realtime Video Streaming
abstract
The outbreak of COVID-19 has dramatically promoted the explosive proliferation of multi-party realtime video streaming (MRVS) services, represented by Zoom and Microsoft Teams. Different from Video-on-Demand (VoD) or live streaming, MRVS enables all-to-all realtime video communication, bringing significant challenges to service providing. First, unreliable network transmission can cause network loss, resulting in delay increase and visual quality degradation. Second, the transformation from two-party to multi-party communication makes resource scheduling much more difficult. Moreover, optimizing the overall QoE requires a global coordination, which is quite challenging given the various impact factors such as bitrate and loss.In this paper, we propose the SJA framework, which is, to our best knowledge, the first server-driven joint loss and bitrate adaptation framework in multi-party realtime video streaming services towards maximized QoE. We comprehensively design an appropriate QoE model for MRVS services to capture the interplay among perceptual quality, variations, bitrate mismatch, loss damage, and streaming delay. We mathematically formulate the QoE maximization problem in MRVS services. A Lyapunov-based relaxation and the SJA algorithm are further designed to address the optimization problem with close-to-optimal performance. Evaluations show that our framework can outperform the SOTA solutions by 18.4% ∼ 46.5%.
Dayou Zhang, Zi Zhu, Lei Zhang 0066, Fangxin Wang 0001, Dan Wang 0002
INFOCOM5
2023 Collaborative Streaming and Super Resolution Adaptation for Mobile Immersive Videos
abstract
Tile-based streaming and super resolution are two representative technologies adopted to improve bandwidth efficiency of immersive video steaming. The former allows selective download of contents in the user viewport by splitting the video into multiple independently decodable tiles. The latter leverages client-side computation to reconstruct the received video into higher quality using advanced neural network models. In this work, we propose CASE, a collaborated adaptive streaming and enhancement framework for mobile immersive videos, which integrates super resolution with tile-based streaming to optimize user experience with dynamic bandwidth and limited computing capability. To coordinate the video transmission and reconstruction in CASE, we identify and address several key design issues including unified video quality assessment, computation complexity model for super resolution, and buffer analysis considering the interplay between transmission and reconstruction. We further formulate the quality-of-experience (QoE) maximization problem for mobile immersive video streaming and propose a rate adaptation algorithm to make the best decisions for download and for reconstruction based on the Lyapunov optimization theory. Extensive evaluation results validate the superiority of our proposed approach, which presents stable performance with considerable QoE improvement, while enabling trade-off between playback smoothness and video quality.
Lei Zhang 0066, Yanjie Dong 0003, Fangxin Wang 0001, Laizhong Cui, Victor C. M. Leung
INFOCOM4
2023 OmniSense: Towards Edge-Assisted Online Analytics for 360-Degree Videos
abstract
With the reduced hardware costs of omnidirectional cameras and the proliferation of various extended reality applications, more and more 360° videos are being captured. To fully unleash their potential, advanced video analytics is expected to extract actionable insights and situational knowledge without blind spots from the videos. In this paper, we present OmniSense, a novel edge-assisted framework for online immersive video analytics. OmniSense achieves both low latency and high accuracy, combating the significant computation and network resource challenges of analyzing 360° videos. Motivated by our measurement insights into 360° videos, OmniSense introduces a lightweight spherical region of interest (SRoI) prediction algorithm to prune redundant information in 360° frames. Incorporating the video content and network dynamics, it then smartly scales vision models to analyze the predicted SRoIs with optimized resource utilization. We implement a prototype of OmniSense with commodity devices and evaluate it on diverse real-world collected 360° videos. Extensive evaluation results show that compared to resource-agnostic baselines, it improves the accuracy by 19.8% – 114.6% with similar end-to-end latencies. Meanwhile, it hits 2.0× – 2.4× speedups while keeping the accuracy on par with the highest accuracy of baselines.
Miao Zhang 0003, Yifei Zhu 0001, Linfeng Shen, Fangxin Wang 0001, Jiangchuan Liu
INFOCOM4
2023 Understanding User Behavior in Volumetric Video Watching: Dataset, Analysis and Prediction
abstract
Volumetric video emerges as a new attractive video paradigm in recent years since it provides an immersive and interactive 3D viewing experience with six degree-of-freedom (DoF). Unlike traditional 2D or panoramic videos, volumetric videos require dense point clouds, voxels, meshes, or huge neural models to depict volumetric scenes, which results in a prohibitively high bandwidth burden for video delivery. Users' behavior analysis, especially the viewport and gaze analysis, then plays a significant role in prioritizing the content streaming within users' viewport and degrading the remaining content to maximize user QoE with limited bandwidth. Although understanding user behavior is crucial, to the best of our best knowledge, there are no available 3D volumetric video viewing datasets containing fine-grained user interactivity features, not to mention further analysis and behavior prediction.
Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Yongting Chen, Miao Zhang 0003, Fangxin Wang 0001
ACM Multimedia7
2023 FSVVD: A Dataset of Full Scene Volumetric Video
abstract
Recent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand out to provide unprecedented immersive and interactive video watching experience. Despite the tremendous potential, the research towards 3D volumetric video is still in its infancy, relying on sufficient and complete datasets for further exploration. However, existing related volumetric video datasets mostly only include a single object, lacking details about the scene and the interaction between them. In this paper, we focus on the current most widely used data format, point cloud, and for the first time release a full-scene volumetric video dataset that includes multiple people and their daily activities interacting with the external environments. Comprehensive dataset description and analysis are conducted, with potential usage of this dataset. The dataset and additional tools can be accessed via the following website: https://cuhksz-inml.github.io/full_scene_volumetric_video_dataset/.
Kaiyuan Hu, Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001
MMSys5
2023 RepCaM: Re-parameterization Content-aware Modulation for Neural Video Delivery
abstract
Recently, content-aware methods have been utilized to reduce the bandwidth and improve the quality of Internet video delivery. Existing methods train corresponding content-aware super-resolution (SR) models for each video chunk on the server and stream low-resolution (LR) video chunks along with SR models to the client. Previous works introduce additional partial parameters to privatize the models of different video chunks. However, this still leads to the accumulation of parameters and even fails to modulate when the length of video increases, bringing extra delivery costs and performance degradation. In this paper, we introduce a novel Re-parameterization Content-aware Modulation (RepCaM) method to modulate all the video chunks with an end-to-end training strategy. Our method adopts extra parallel-cascade parameters during training to fit multiple chunks while removing the additional parameters through re-parameterization during inference. Therefore, RepCaM increases no extra model size compared with the original SR model. Moreover, in order to improve the training efficiency on servers, we propose an online Video Patch Sampling (VPS) method to speed up the training convergence. We conduct extensive experiments on VSD4K and newly collected dataset (VSD4K-2022), achieving state-of-the-art results in video restoration quality and delivery bandwidth compression. Code is available at: https://github.com/Neural-video-delivery/RepCaM-Pytorch-NOSSDAV2023.
Rongyu Zhang, Lixuan Du, Jiaming Liu 0003, Congcong Song, Fangxin Wang 0001, Xiaoqi Li 0009, Ming Lu 0002, Yandong Guo, Shanghang Zhang
NOSSDAV5
2023 CaV3: Cache-assisted Viewport Adaptive Volumetric Video Streaming
abstract
Volumetric video (VV) recently emerges as a new form of video application providing a photorealistic immersive 3D viewing experience with 6 degree-of-freedom (DoF), which empowers many applications such as VR, AR, and Metaverse. A key problem therein is how to stream the enormous size VV through the network with limited bandwidth. Existing works mostly focused on predicting the viewport for a tiling-based adaptive VV streaming, which however only has quite a limited effect on resource saving. We argue that the content repeatability in the viewport can be further leveraged, and for the first time, propose a client-side cache-assisted strategy that aims to buffer the repeatedly appearing VV tiles in the near future so as to reduce the redundant VV content transmission. The key challenges exist in three aspects, including (1) feature extraction and mining in 6 DoF VV context, (2) accurate long-term viewing pattern estimation and (3) optimal caching scheduling with limited capacity. In this paper, we propose CaV3, an integrated cache-assisted viewport adaptive VV streaming framework to address the challenges. CaV3 employs a Long-short term Sequential prediction model (LSTSP) that achieves accurate short-term, mid-term and long-term viewing pattern prediction with a multi-modal fusion model by capturing the viewer's behavior inertia, current attention, and subjective intention. Besides, CaV3 also contains a contextual MAB-based caching adaptation algorithm (CCA) to fully utilize the viewing pattern and solve the optimal caching problem with a proved upper bound regret. Compared to existing VV datasets only containing single or co-located objects, we for the first time collect a comprehensive dataset with sufficient practical unbounded 360° scenes. The extensive evaluation of the dataset confirms the superiority of CaV3, which outperforms the SOTA algorithm by 15.6%-43% in viewport prediction and 13%-40% in system utility.
Junhua Liu 0003, Boxiang Zhu, Fangxin Wang 0001, Yili Jin 0001, Shuguang Cui
VR3
2023 Ebublio: Edge-Assisted Multiuser 360° Video Streaming
abstract
As one of the most important manifestations of virtual reality (VR), 360° panoramic videos in recent years have experienced booming development due to the desire for immersive and interactive experiences. Compared to traditional videos, 360° videos are featured with uncertain user Field of View (FoV), more sensitive delay tolerance, and much higher bandwidth requirement, bringing unprecedented challenges to 360° video streaming. Meanwhile, the development of 5G and mobile edge computing starts to pave the way for high-bandwidth low-latency video streaming. Some preliminary works focus on either individual FoV prediction or multiuser Quality of Experience (QoE) oriented cache strategy design, while how to design a holistic solution toward optimizing the overall user QoE with considerations over fairness and long-term system cost remains a nontrivial problem. In this article, we proposeEbublio, a novel intelligent edge caching framework to address the aforementioned challenges in 360° video streaming.Ebublioconsists of a collaborative FoV prediction (CFP) module and a long-term tile caching optimization (LTO) module to jointly optimize the long-term user QoE and system cost. The former module integrates the features of video content, user trajectory, and other users’ records for combined prediction. The latter one employs the Lyapunov framework and a subgradient optimization approach toward the optimal caching replacement policy. Our trace-driven evaluation demonstrates the superiority of our framework, with about 42% improvement in FoV prediction, and 36% improvement in QoE at similar traffic consumption.
Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001, Shuguang Cui
IEEE Internet Things J.3
2023 Backup Battery Allocation and Workload Migration Against Electrical Load Shedding at Edge
abstract
In the 5G era (and the upcoming 6G), mobile edge computing (MEC) has been advocated to serve the massive amount of Internet of Things (IoT) devices by base stations (BSs) and edge data centers (EDCs). Geo-distributed EDCs are generally of much smaller scales as compared to mega data centers and hence of much lower costs, but can have fast response to their users so as to satisfy the demands of real-time applications. As their reliability and availability heavily depend on the electrical power supply, most EDCs are equipped with battery groups as backup power in case of power grid load shedding or outage. In a heterogeneous geo-distributed environment, the QoS of heavily loaded EDCs however can be severely impacted by limited backup power while lightly loaded EDCs may simply waste such precious resources. Moreover, a heavily loaded EDC may suffer from deep discharge of its battery group, which will cause a significant reduction of battery capacity and lifetime. This further aggravates the aforementioned situations should load shedding/outage happen again. In this article, we carefully analyze the workloads in EDCs and classify them into interactive workloads and batch workloads, respectively. We then develop a novel battery allocation framework with smart workload migration for EDCs, which simultaneously protects interactive workloads from being interrupted and minimizes the waiting time of batch workloads. Our extensive evaluations show that our strategies can optimize all the objectives within a limited overall cost as compared to state-of-the-art practical allocation.
Linfeng Shen, Fangxin Wang 0001, Feng Wang 0001, Jiangchuan Liu
IEEE Internet Things J.2
2023 FedAB: Truthful Federated Learning With Auction-Based Combinatorial Multi-Armed Bandit
abstract
Federated learning (FL) emerges as a new distributed machine learning (ML) paradigm that enables thousands of mobile devices to collaboratively train ML models using local data without compromising user privacy. However, the FL learning quality highly relies on the data contribution from the distributed mobile devices. Therefore, a well-designed incentive mechanism with effectiveness, fairness, and reciprocity is in urgent need to guarantee the stable participation of users. In this article, we propose federated auction bandit (FedAB), an incentive and client selection strategy based on a novel multiattribute reverse auction mechanism and a combinatorial multi-armed bandit (CMAB) algorithm. First, we develop a local contribution evaluation method based on importance sampling in the FL context. We then design a novel payment mechanism that is able to preserve individual rationality and incentive compatibility (truthfulness). At last, we design a UCB-based winner selection algorithm that is proven to achieve the server’s utility maximization with fairness and reciprocity. We have conducted extensive experiments on real data sets. The results demonstrate the superiority ofFedAB, with a 10%–50% improvement in total reward, final accuracy, and convergence speed compared to state-of-the-art solutions.
Chenrui Wu 0002, Yifei Zhu 0001, Rongyu Zhang, Fangxin Wang 0001, Shuguang Cui
IEEE Internet Things J.5
2022 Towards Joint Loss and Bitrate Adaptation in Realtime Video Streaming
abstract
Recent years have seen booming development of realtime streaming services, highly improving user experience in remote work, online education, and entertainment. Unlike video-on-demand (VoD) or live services, realtime streaming service has extremely stringent delay requirements, rendering the TCP-based transmission no longer applicable. Existing works based on UDP (or its variants) either suffer from the packet loss problem or only focus on improving several QoS metrics, which cannot achieve satisfactory user QoE. Our insight is to slightly sacrifice the bitrate and video quality to trade for the most significant delay to maximize the overall QoE. We propose Oppugno‡‡Oppugno is a spell in Harry Potter that makes magical creatures attack the caster. It is a metaphor that we use an additional mechanism to mitigate the influence of packet loss., an integrated framework that achieves joint loss adaptation and bitrate adaption towards maximized QoE in realtime streaming services. Oppugno leverages existing UDP mechanisms and employs an advanced deep reinforcement learning algorithm Proximal Policy Optimization (PPO), to adaptively select optimal actions based on network conditions. Trace-driven experiments demonstrate the superiority of our framework, which outperforms the SOTA work by 3.9% ∼ 11.6%.
Dayou Zhang, Fangxin Wang 0001, Dan Wang 0002, Jiangchuan Liu
ICME3
2022 Maxim: DRL-Based Cross-Camera Streaming Configuration for Real-Time Video Analytics
abstract
Real-time video analytics (VA) with requirement of high accuracy necessitates intensive bandwidth resource consumption, calling for an adaptive streaming configuration strategy to strike a balance in VA pipelines. Existing works however suffer from a key limitation: the profiling-based strategy not only wastes unnecessary resources with the golden configuration transmission but also is trapped into only coarse-grained adaptation due the contradiction between profiling granularity and the profiling cost. In this paper, we for the first time reveal the correlation between the video dynamics and motion degree and highlight the limitation of traditional profiling-based strategy. We then propose Maxim, the first learning-based framework that solves the fine-grained VA configuration adaptation problem employing a novel deep reinforcement learning-based methodology without using any golden configuration in inference stage. Maxim could optimize the trade-off between resources cost and inference accuracy. Besides, Maxim also employs an enhanced cross-camera collaboration based on spatial and temporal correlation among cameras, which further improves robustness and performance in a large camera network. Extensive experiments confirm the superiority of our work compared with SOTA works, with a 76.6% improvement in comprehensive.
Yutao Zhou, Fangxin Wang 0001, Zhi Wang 0001
ICME3
2022 CASVA: Configuration-Adaptive Streaming for Live Video Analytics
abstract
The advent of high-accuracy and resource-intensive deep neural networks (DNNs) has fulled the development of live video analytics, where camera videos need to be streamed over the network to edge or cloud servers with sufficient computational resources. Although it is promising to strike a balance between available bandwidth and server-side DNN inference accuracy by adjusting video encoding configurations, the influences of fine-grained network and video content dynamics on configuration performance should be addressed. In this paper, we propose CASVA, a Configuration-Adaptive Streaming framework designed for live Video Analytics. The design of CASVA is motivated by our extensive measurements on how video configuration affects its bandwidth requirement and inference accuracy. To handle the complicated dynamics in live video analytics streaming, CASVA trains a deep reinforcement learning model which does not make any assumptions about the environment but learns to make configuration choices through its experiences. A variety of real-world network traces are used to drive the evaluation of CASVA. The results on a multitude of video types and video analytics tasks show the advantages of CASVA over state-of-the-art solutions.
Miao Zhang 0003, Fangxin Wang 0001, Jiangchuan Liu
INFOCOM2
2022 Batch Adaptative Streaming for Video Analytics
abstract
Video streaming plays a critical role in the video analytics pipeline and thus its adaptation scheme has been a focus of optimization. As machine learning algorithms have become main consumers of video contents, the streaming adaptation decision should be made to optimize their inference performance. Existing video streaming adaptation schemes for video analytics are usually designed to adapt to bandwidth and content variations separately, which fail to consider the coordination between transmission and computation. Given the nature of batch transmission in video streaming and batch processing in deep learning-based inference, we observe that the choices of the batch sizes directly affects the bandwidth efficiency, the response delay and the accuracy of the deep learning inference in video analytics. In this work, we investigate the effect of the batch size in transmission and processing, formulate the optimal batch size adaptation problem, and further develop the deep reinforcement learning-based solution. Practical issues are further addressed for Implementation. Extensive simulations are conducted for performance evaluation, whose results demonstrate the superiority of our proposed batch adaptive streaming approach over the baseline streaming approaches.
Lei Zhang 0066, Ximing Wu, Fangxin Wang 0001, Laizhong Cui, Zhi Wang 0001, Jiangchuan Liu
INFOCOM4
2022 Where Are You Looking?: A Large-Scale Dataset of Head and Gaze Behavior for 360-Degree Videos and a Pilot Study
abstract
360° videos in recent years have experienced booming development. Compared to traditional videos, 360° videos are featured with uncertain user behaviors, bringing opportunities as well as challenges. Datasets are necessary for researchers and developers to explore new ideas and conduct reproducible analyses for fair comparisons among different solutions. However, existing related datasets mostly focused on users' field of view (FoV), ignoring the more important eye gaze information, not to mention the integrated extraction and analysis of both FoV and eye gaze. Besides, users' behavior patterns are highly related to videos, yet most existing datasets only contained videos with subjective and qualitative classification from video genres, which lack quantitative analysis and fail to characterize the intrinsic properties of a video scene. To this end, we first propose a quantitative taxonomy for 360° videos that contains three objective technical metrics. Based on this taxonomy, we collect a dataset containing users' head and gaze behaviors simultaneously, which outperforms existing datasets with rich dimensions, large scale, strong diversity, and high frequency. Then we conduct a pilot study on users' behaviors and get some interesting findings such as user's head direction will follow his/her gaze direction with the most possible time interval. A case of application in tile-based 360° video streaming based on our dataset is later conducted, demonstrating a great performance improvement of existing works by leveraging our provided gaze information. Our dataset is available at https://cuhksz-inml.github.io/head_gaze_dataset/
Yili Jin 0001, Junhua Liu 0003, Fangxin Wang 0001, Shuguang Cui
ACM Multimedia3
2022 Computing Cost Optimization for Multi-BS in MEC by Offloading
Wenzao Li, Fangxin Wang 0001, Yuwen Pan, Lei Zhang 0066, Jiangchuan Liu
Mob. Networks Appl.2
2022 CharmSeeker: Automated Pipeline Configuration for Serverless Video Processing
abstract
Video processing plays an essential role in a wide range of cloud-based applications. It typically involves multiple pipelined stages, which well fits the latest fine-grained serverless computing paradigm if properly configured to match the cost and delay constraints of video. Existing configuration tools, however, are primarily developed for traditional virtual machine clusters with general workloads. This paper presents CharmSeeker, an automated configuration tuning tool for serverless video processing pipelines. We first carefully examine the key steps and the performance bottlenecks for video processing over modern serverless platforms. Then, we identify the configuration space for processing pipelines and leverage a carefully designed Sequential Bayesian Optimization search scheme to identify promising configurations. We further address the practical challenges toward integrating our solution into real-world systems and develop a prototype with AWS Lambda. Evaluation results show that CharmSeeker can find out the optimal or near-optimal configurations that improve the relative processing time up to 408.77%. It is also more robust and scalable to various video processing pipelines compared with state-of-the-art solutions.
Miao Zhang 0003, Yifei Zhu 0001, Jiangchuan Liu, Feng Wang 0001, Fangxin Wang 0001
IEEE/ACM Trans. Netw.5
2021 Workload Migration across Distributed Data Centers under Electrical Load Shedding
abstract
Data centers are essential components in the current digital world. The number and scales of data centers have both increased a lot in recent years. The distributed data centers are standing out as a promising solution due to the development of modern applications which need a massive amount of computation resource and strict response requirement. However, compared to centralized data centers, distributed data centers are more fragile when the power supply is unstable. Power constraints or outages because of electrical load shedding or other reasons will significantly affect the service performance of data centers and damage the quality of service (QoS) for customers. Moreover, unlike conventional data centers, distributed data centers are often unattended, so we need a system that can automatically calculate the best workload schedule to maximize profit in such situations. In this paper, we closely investigate the influence of electrical load shedding in distributed data centers and construct a physical model to estimate the relationship among power, heat and workload. We then use queueing theory to approximate the tasks’ response time and aim to minimize the overall response time of tasks by migration. Our extensive evaluations show that our method can improve the response time with more than 9% reduction.
Linfeng Shen, Fangxin Wang 0001, Feng Wang 0001, Jiangchuan Liu
IWQoS2
2021 Demystifying the Relationship Between Network Latency and Mobility on High-Speed Rails: Measurement and Prediction
abstract
Recent years have seen increasing attention on building High-Speed Railways (HSR) in many countries. Trains running on the railways have a top velocity of up to over 300 km/hour. This makes it become a scenario with unstable connection qualities. In this paper, we propose a novel model that can accurately estimate the mobility status on HSR based on the changing patterns of network latency. Though various impact factors make the prediction complex, we however argue that the recent advance of deep learning applies well in our context, and further we design a neural network model that can estimate the moving velocity based on monitoring network latency’s changing patterns in a short period. In this model, we use a new variable called Round Difference Time (RDT) to describe latency’s changing patterns. We also use the Fourier Transform to extract the hidden time-frequency and use the generated spectrum for estimation. Our data-driven evaluations show that with suitable parameters, this model can get an accuracy of up to 94% on all three lines.
Jiangchuan Liu, Fangxin Wang 0001, Ke Xu 0002
IWQoS3
2021 Towards cloud-edge collaborative online video analytics with fine-grained serverless pipelines
abstract
The ever-growing deployment scale of surveillance cameras and the users' increasing appetite for real-time queries have urged online video analytics. Synergizing the virtually unlimited cloud resources with agile edge processing would deliver an ideal online video analytics system; yet, given the complex interaction and dependency within and across video query pipelines, it is easier said than done. This paper starts with a measurement study to acquire a deep understanding of video query pipelines on real-world camera streams. We identify the potentials and practical challenges towards cloud-edge collaborative video analytics. We then argue that the newly emerged serverless computing paradigm is the key to achieve fine-grained resource partitioning with minimum dependency. We accordingly propose CEVAS, a Cloud-Edge collaborative Video Analytics system empowered by fine-grained Serverless pipelines. It builds flexible serverless-based infrastructures to facilitate fine-grained and adaptive partitioning of cloud-edge workloads for multiple concurrent query pipelines. With the optimized design of individual modules and their integration, CEVAS achieves real-time responses to highly dynamic input workloads. We have developed a prototype of CEVAS over Amazon Web Services (AWS) and conducted extensive experiments with real-world video streams and queries. The results show that by judiciously coordinating the fine-grained serverless resources in the cloud and at the edge, CEVAS reduces 86.9% cloud expenditure and 74.4% data transfer overhead of a pure cloud scheme and improves the analysis throughput of a pure edge scheme by up to 20.6%. Thanks to the fine-grained video content-aware forecasting, CEVAS is also more adaptive than the state-of-the-art cloud-edge collaborative scheme.
Miao Zhang 0003, Fangxin Wang 0001, Yifei Zhu 0001, Jiangchuan Liu, Zhi Wang 0001
MMSys2
2021 Multi-Adversarial In-Car Activity Recognition Using RFIDs
abstract
In-car human activity recognition opens a new opportunity toward intelligent driving behavior detection and touchless human-car interaction. Among the many sensing technologies (e.g., using cameras and wearable sensors), radio frequency identification (RFID) exhibits unique advantages given its low cost, easy deployment, and less privacy concerns. Existing RFID-based solutions for activity recognition are mostly confined to working in stable indoor spaces. The inside space of a car however is much more compact and complex, not to mention the fast-changing driving conditions. All these introduce non-negligible noises that pollute the activity-related information, and the existence of various car models in the market further complicates the problem. In this article, we for the first time closely examine the distinct factors that affect the RFID-based in-car activity recognition. We present RF-CAR, a novel RFID-based tag-free solution that well adapts to different in-car environments. RF-CAR smartly filters the domain-specific features in RF signals and retains activity-related features to the maximum extent. It then integrates a deep learning architecture and an advanced multi-adversarial domain adaptation network for training and prediction. With only one-time pre-training, RF-CAR can adapt to new data domains such as new driving conditions, car models, and human subjects for robust activity recognition. We also demonstrate that it is readily deployable in cars with commercial off-the-shelf (COTS) RFID devices. Our extensive experiments suggest that RF-CAR achieves an overall recognition accuracy of around 95 percent, which significantly outperforms the state-of-the-art solutions.
Fangxin Wang 0001, Jiangchuan Liu, Wei Gong 0001
IEEE Trans. Mob. Comput.1
2020 Intelligent Video Caching at Network Edge: A Multi-Agent Deep Reinforcement Learning Approach
abstract
Today's explosively growing Internet video traffics and viewers' ever-increasing quality of experience (QoE) demands for video streaming bring tremendous pressures to the backbone network. As a new network paradigm, mobile edge caching provides a promising alternative by pushing video content closer at the network edge rather than the remote CDN servers so as to reduce both content access latency and redundant network traffic. However, our large-scale trace analysis shows that different from CDN based caching, edge caching environment is much more complicated with massively dynamic and diverse request patterns, which renders that existing rule-based and model-based caching solutions may not well fit such complicated edge environments. Moreover, although cooperative caching has been proposed to better afford limited storage on each individual edge server, our trace analysis also shows that the request similarity among neighboring edges can be highly dynamic and diverse, which is drastically different from CDN based caching environment, and can easily compromise the benefits from traditional cooperative caching mostly designed based on CDN environment. In this paper, we propose MacoCache, an intelligent edge caching framework that is carefully designed to afford the massively diversified and distributed caching environment to minimize both content access latency and traffic cost. Specifically, MacoCache leverages a multi-agent deep reinforcement learning (MADRL) based solution, where each edge is able to adaptively learn its own best policy in conjunction with other edges for intelligent caching. The real trace-driven evaluation further demonstrates that MacoCache is able to reduce an average of 21% latency and 26% cost compared with the state-of-the-art caching solution.
Fangxin Wang 0001, Feng Wang 0001, Jiangchuan Liu, Ryan Shea, Lifeng Sun
INFOCOM1
2020 Follow me Robot-Mind: Cloud brain based personalized robot service with migration
Long Hu, Yinging Jiang, Fangxin Wang 0001, Kai Hwang 0001, M. Shamim Hossain, Muhammad Ghulam
Future Gener. Comput. Syst.3
2020 Car4Pac: Last Mile Parcel Delivery Through Intelligent Car Trip Sharing
abstract
The explosion of online shopping brings great challenges to traditional logistics industry, where the massive parcels and tight delivery deadline impose a large cost on the delivery process, in particular the last mile parcel delivery. On the other hand, modern cities never lack transportation resources such as the private car trips. Motivated by these observations, we propose a novel and effective last mile parcel delivery mechanism through car trip sharing, to leverage the available private car trips to incidentally deliver parcels during their original trips. To achieve this, the major challenges lie in how to accurately estimate the parcel delivery trip cost and assign proper tasks to suitable car trips to maximize the overall performance. To this end, we develop Car4Pac, an intelligent last mile parcel delivery system to address these challenges. Leveraging the real-world massive car trip trajectories, we first build up a 3D (time-dependent, driver-dependent and vehicle-dependent) landmark graph that accurately predicts the travel time and fuel consumption of each road segment. Our prediction method considers not only traffic conditions of different times, but also driving skills of different people and fuel efficiencies of different vehicles. We then develop a two-stage solution towards the parcel delivery task assignment, which is optimal for one-to-one assignment and yields high-quality results for many-to-one assignment. Our extensive real-world trace driven evaluations further demonstrate the superiority of our Car4Pac solution.
Fangxin Wang 0001, Yifei Zhu 0001, Feng Wang 0001, Jiangchuan Liu, Xiaoqiang Ma, Xiaoyi Fan 0001
IEEE Trans. Intell. Transp. Syst.1
2020 DeepCast: Towards Personalized QoE for Edge-Assisted Crowdcast With Deep Reinforcement Learning
abstract
Today’s anywhere and anytime broadband connection and audio/video capture have boosted the deployment of crowdsourced livecast services (orcrowdcast). Bridging a massive amount of geo-distributed broadcasters and their fellow viewers, such representatives as Twitch.tv, Youtube Gaming, and Inke.tv, have greatly changed the generation and distribution landscape of streaming content. They also enable rich online interactions among the crowd, and strive to offer personalized Quality-of-Experience (QoE) for individual viewers. Given the ultra-large scale and the dynamics of the crowd, personalizing QoE however is much more challenging than in early generation streaming services. The rich interactions among the broadcasters, viewers, and the network system, on the other hand, also offer invaluable data that could be utilized towards informed management. This paper presentsDeepCast, an edge-assisted crowdcast framework that explores the sheer amount of viewing data towards intelligent decisions for personalized QoE demands. DeepCast seamlessly integrates cloud, CDN, and edge servers for crowdcast content distribution, and advocates a data-driven design that extracts the hidden information from the complex interactions among the system components. Through deep reinforcement learning (DRL), it automatically identifies the most suitable strategies for viewer assignment and transcoding at edges. We collect multiple real-world datasets and evaluate the performance of DeepCast with trace-driven experiments. The results demonstrate its flexibility and effectiveness towards better personalized QoE and lower cost for crowdcast systems.
Fangxin Wang 0001, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu, Yifei Zhu 0001, Haitian Pang, Lifeng Sun
IEEE/ACM Trans. Netw.1
2019 A Location-Aware Duty Cycle Approach toward Energy-Efficient Mobile Crowdsensing
abstract
This paper aims to explore the problem of energy economy of mobile devices in the Mobile Crowdsensing (MCS) scenario. The neighbor scanning mechanism of mobile devices usually consumes most of the energy in multi-hop message transmission. Traditional mechanisms such as chaotic neighbor detection and continuous neighbor discovery, can easily exhaust the limited energy. Since they are actually unnecessarily considering the strong correlation between the transmission opportunity and the social characteristics. Therefore, diminishing energy consumption is of key importance toward energy-efficient MCS. Sleeping strategy stands out as a promising approach to improve energy efficiency and hold the the network metrics in MCS applications, while challenges still lie in achieving effective scheduling given the changing environment. In this study, we proposed a novel self-adapt sleeping scheduling approach based on the correlation between the pedestrian's own historical trajectory and geographic grid information for energy saving (ESGeo). As the grid-based method can well record the encounter characteristics of nodes, ESGeo is able to pick flexible duty-cycling strategies for mobile devices in each distinguishable grid. This enables, mobile devices to avoid excessive scanning in low-probability encounter areas. Extensive simulation results further demonstrated that the proposed approach can significantly outperform the typical routing approaches in terms of energy-efficiency, without largely affecting the overall networking performance.
Wenzao Li, Fangxin Wang 0001, Jiangchuan Liu
ICPADS4
2019 Towards Low Latency Multi-viewpoint 360° Interactive Video: A Multimodal Deep Reinforcement Learning Approach
abstract
Recently, the fusion of 360° video and multi-viewpoint video, called multi-viewpoint (MVP) 360° interactive video, has emerged and created much more immersive and interactive user experience, but calls for a low latency solution to request the high-definition contents. Such viewing-related features as head movement have been recently studied, but several key issues still need to be addressed. On the viewer side, it is not clear how to effectively integrate different types of viewing-related features. At the session level, questions such as how to optimize the video quality under dynamic networking conditions and how to build an end-to-end mapping between these features and the quality selection remain to be answered. The solutions to these questions are further complicated given the many practical challenges, e.g., incomplete feature extraction and inaccurate prediction.This paper presents an architecture, called iView, to address the aforementioned issues in an MVP 360° interactive video scenario. To fully understand the viewing-related features and provide a one-step solution, we advocate multimodal learning and deep reinforcement learning in the design. iView intelligently determines video quality and reduces the latency without pre-programmed models or assumptions. We have evaluated iView with multiple real-world video and network datasets. The results showed that our solution effectively utilizes the features of video frames, networking throughput, head movements, and viewpoint selections, achieving at least 27.2%, 15.4%, and 2.8% improvements on the three video datasets, respectively, compared with several state-of-the-art methods.
Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Jiangchuan Liu, Lifeng Sun
INFOCOM3
2019 Intelligent Edge-Assisted Crowdcast with Deep Reinforcement Learning for Personalized QoE
abstract
Recent years have seen booming development and great success in interactive crowdsourced livecast (i.e., crowdcast). Different from traditional livecast services, crowdcast is featured with tremendous video contents at the broadcaster side, highly diverse viewer side content watching environments/preferences as well as viewers' personalized quality of experience (QoE) demands (e.g., individual preferences for streaming delays, channel switching latencies and bitrates). This imposes unprecedented key challenges on how to flexibly and cost-effectively accommodate the heterogeneous and personalized QoE demands for the mass of viewers. In this paper, we propose DeepCast, an edge-assisted crowdcast framework, which makes intelligent decisions at edges based on the massive amount of real-time information from the network and viewers to accommodate personalized QoE with minimized system cost. Given the excessive computation complexity in this context, we propose a data-driven deep reinforcement learning (DRL) based solution that can automatically learn the best suitable strategies for viewer scheduling and transcoding selection. To our best knowledge, DeepCast is the first edge-assisted framework that applies the advance of DRL to explicitly accommodate personalized QoE optimization for crowdcast services. We collect multiple real-world datasets and evaluate the performance of DeepCast using trace-driven experiments. The results demonstrate the superiority of our DeepCast framework and its DRL-based solution.
Fangxin Wang 0001, Cong Zhang 0002, Feng Wang 0001, Jiangchuan Liu, Yifei Zhu 0001, Haitian Pang, Lifeng Sun
INFOCOM1
2019 WiCAR: wifi-based in-car activity recognition with multi-adversarial domain adaptation
abstract
In-car human activity recognition is playing a critical role in detecting distracted driving and improving human-car interaction. Among multiple sensing technologies, WiFi-based in-car activity recognition exhibits unique advantages since it does not rely on visible light, avoids privacy leaks and is cost-efficient with integrated WiFi signals in cars. Existing WiFi-based recognition systems mostly focus on the relatively stable indoor space, which only yield reasonably good performance in limited situations. Based on our field studies, the in-car activity recognition, however, is much more complicated suffering from more impact factors. First, the external moving objects and the surrounding WiFi signals can cause various disturbances to the in-car activity sensing. Second, considering the compact in-car space, different car models can also lead to different multipath distortions. Moreover, different people can also perform activities in different shapes. Such extraneous information related to specific driving conditions, car models and human subjects is implicitly contained for training and prediction, inevitably leading to poor recognition performance for new environment and people.
Fangxin Wang 0001, Jiangchuan Liu, Wei Gong 0001
IWQoS1
2019 Toward Optimal Resource Allocation for Task Offloading in Mobile Edge Computing
Wenzao Li, Yuwen Pan, Fangxin Wang 0001, Lei Zhang 0066, Jiangchuan Liu
QSHINE3
2019 On Spatial Diversity in WiFi-Based Human Activity Recognition: A Deep Learning-Based Approach
abstract
The deeply penetrated WiFi signals not only provide fundamental communications for the massive Internet of Things devices but also enable cognitive sensing ability in many other applications, such as human activity recognition. State-of-the-art WiFi-based device-free systems leverage the correlations between signal changes and body movements for human activity recognition. They have demonstrated reasonably good recognition results with a properly placed transceiver pair, or, in other words, when the human body is within a certain sweet zone. Unfortunately, the sweet zone is not ubiquitous. When the person moves out of the area and enters a dead zone, or even just the orientation changes, the recognition accuracy can quickly decay. In this paper, we closely examine such spatial diversity in WiFi-based human activity recognition. We identify the dead zones and their key influential factors, and accordingly present WiSDAR, a WiFi-based spatial diversity-aware device-free activity recognition system. WiSDAR overshadows the dead zones yet with only one physical WiFi sender and receiver. The key innovation is extending the multiple antennas of modern WiFi devices to construct multiple separated antenna pairs for activity observing. Profiling activity features from multiple spatial dimensions can be more complicated and offer much richer information for further recognition. To this end, we propose a deep learning-based framework that integrates the hidden features from both temporal and spatial dimensions, achieving highly accurate and reliable recognition results. WiSDAR is fully compatible with commercial off-the-shelf WiFi devices, and we have implemented it on the commonly available Intel WiFi 5300 cards. Our real-world experiments demonstrate that it recognizes human activities with a stable accuracy of around 96%.
Fangxin Wang 0001, Wei Gong 0001, Jiangchuan Liu
IEEE Internet Things J.1
2019 Backup Battery Analysis and Allocation against Power Outage for Cellular Base Stations
abstract
Base stations have been widely deployed to satisfy the service coverage and explosive demand increase in today's cellular networks. Their reliability and availability heavily depend on the electrical power supply. Battery groups are installed as backup power in most of the base stations in case of power outages due to severe weathers or human-driven accidents, particularly in remote areas. The limited numbers and capacities of batteries, however, can hardly sustain a long power outage without a well-designed allocation strategy. As a result, the service interruption occurs along with an increasing maintenance cost. Meanwhile, a deep discharge of a battery in such case can also accelerate the battery degradation and eventually contribute to a higher battery replacement cost. In this paper, we closely examine the base station features and backup battery features from a 1.5-year dataset of a major cellular service provider, including 4,206 base stations distributed across 8,400 square kilometers and more than 1.5 billion records on base stations and battery statuses. Through exploiting the correlations between the battery working conditions and battery statuses, we build up a deep learning based model to estimate the remaining lifetime of backup batteries. We then develop BatAlloc, a battery allocation framework to address the mismatch between the battery supporting ability and diverse power outage incidents. We present an effective solution that minimizes both the service interruption time and the overall cost. Our real trace-driven experiments show that BatAlloc cuts down the average service interruption time from 4.7 hours to nearly zero with only 85 percent of the overall cost compared to the current practical allocation.
Fangxin Wang 0001, Xiaoyi Fan 0001, Feng Wang 0001, Jiangchuan Liu
IEEE Trans. Mob. Comput.1
2018 Task Scheduling with Optimized Transmission Time in Collaborative Cloud-Edge Learning
abstract
Deep learning has been applied in many recent advanced applications in the field of transportation, finance and medicine. These applications require significant computation resources and large-scale training samples. Cloud becomes a natural choice for conducting these learning tasks due to its abundant resources. However, deeper penetration of deep learning techniques in mission critical applications, like driverless car, calls for stricter time requirement to guarantee its interaction and larger amount of dataset for training to guarantee its accuracy, which cannot be easily satisfied by the cloud and makes the network transmission become the bottleneck. Edge learning emerges to be a promising direction to reduce data transmission time by processing and compressing the raw data at the edge of the network, while brings the concern of accuracy reduction at the meantime. To balance this tradeoff under cloud-edge architecture, we study a task scheduling problem for reducing weighted transmission time which takes learning accuracy into consideration. We also propose efficient scheduling algorithms which are able to achieve up to 50% reduction in makespan with extensive trace-driven simulations.
Yutao Huang, Yifei Zhu 0001, Xiaoyi Fan 0001, Xiaoqiang Ma, Fangxin Wang 0001, Jiangchuan Liu, Ziyi Wang 0002, Yong Cui 0001
ICCCN5
2018 Sensing Power Spectrum Density of True Ultrasounds on Mobile Devices
abstract
Many efforts have been made on sensing ultrasound with commercial-off-the-shelf (COTS) mobile devices in the recent literature. Yet due to the limited sound sample rate, current COTS mobile devices can not directly capture any sound at the frequency over 24 kHz. This issue prevents true ultrasound, of which the frequency is typically over 40 kHz, from benefiting the existing sound sensing applications. In this work, we show that by subtly customizing the sampling process, we can make COTS mobile devices hear the true ultrasound that is typically beyond their capability to fully capture. Particularly, we present a system that enable COTS mobile devices to sense the power spectrum density (PSD) of true ultrasounds, of which the frequency can be as high as 60 kHz.
Yuchi Chen, Wei Gong 0001, Jiangchuan Liu, Fangxin Wang 0001, Haitian Pang
IWQoS4
2018 Edge Computing Empowered Generative Adversarial Networks for Realtime Road Sensing
abstract
Automobiles have become one of the necessities of modern life and deeply penetrated into our daily activities. They unfortunately also introduce numerous social problems, among which traffic accidents are most notoriously threatening automobile drivers and other road users. Advanced driver-assistance systems (ADAS) are under rapid development in recent years, which can necessarily reduce or even eliminate the driver errors, significantly relieving on drivers suffering or stress. These state-of-the-art ADAS mainly rely on built-in cameras, radars and ultrasound sensors to provide road sensing services for object detection, which are further advanced by recent explosion of vision and neural network technologies.
Yiting He, Xiaoyi Fan 0001, Feng Wang 0001, Fangxin Wang 0001, Jiangchuan Liu
IWQoS4
2018 Ridesharing as a Service: Exploring Crowdsourced Connected Vehicle Information for Intelligent Package Delivery
abstract
Nowadays online shopping has become explosively popular and the vast numbers of generated packages have brought great challenges to the traditional logistics industry, especially the last mile package delivery. Traditional delivery approaches rely on dedicated couriers for package dispatch, while the labor cost is quite expensive and the quality is hard to guarantee due to the diverse delivery addresses and tight deadlines. On the other hand, modern cities are full of available transportation resources such as private car trips. The mobile crowdsourcing through 4G/5G and vehicle-related communications enables the vehicle resources to be connected as an intelligent transportation system. As such, we believe ridesharing will be a core service for connected vehicles, which we refer to as Ridesharing as a Service (RaaS). In this paper, we focus on the quality of service (QoS) of RaaS in the last mile package delivery. Mining from real-world car trips, we build up a citywide routing graph and conduct a personalized travel cost prediction considering both the travel time of each driver and the fuel consumption of each vehicle. We then design an online algorithm to assign proper package delivery tasks to the submitted car trips, aiming to maximize the utility of the ridesharing service provider. Our extensive real-world trace-driven evaluations further demonstrate the superiority of our RaaS based package delivery.
Fangxin Wang 0001, Yifei Zhu 0001, Feng Wang 0001, Jiangchuan Liu
IWQoS1
2018 Practical Key Tag Monitoring in RFID Systems
abstract
With rapid development of radio frequency identification (RFID) technology, ever-increasing research effort has been dedicated to devising various RFID-enabled services. The key tag monitoring, which is to detect anomaly of key tags, is one of the most important services in such important Internet-of-Things applications as inventory management. Yet prior work assumes that all tags are armed with hashing functionality and a reader would report channel states in every slot, which is not supported by commercial off-the-shelf (COTS) RFID tags and readers. To bridge this gap, this paper is devoted to enabling key tag monitoring service with COTS devices. In particular, we introduce two anomaly monitoring protocols to detect whether there is any key tag absent from the system. The first protocol employs Q-query that works in an analog frame slotted Aloha paradigm to interrogate tags and collect tag IDs. An anomaly event will be found if at least one key tag ID is not present in the collected ones. To reduce time cost of the first protocol resulted from tag collisions, we present a collision-free method that uses select-query to specify a key tag to reply in each slot. Once there is no response in a slot, the specified key tag is regarded as a missing tag. We conduct experiments to evaluate two protocols.
Jihong Yu, Wei Gong 0001, Jiangchuan Liu, Lin Chen 0002, Fangxin Wang 0001, Haitian Pang
IWQoS5
2018 Highlight-Aware Content Placement in Crowdsourced Livecast Services
abstract
Recent years have witnessed an explosion of crowdsourced livecast (i.e., live broadcast) services, in which any Internet users can act as broadcasters to publish livecasts to fellow viewers. To help grow broadcasters' channels, crowdsourced livecast services provide a past-broadcast saving service, allowing viewers to watch the replays they may have missed. Our real-trace measurement and questionnaire survey show that (1) the duration of most of livecasts is extremely long; (2) a much longer duration largely affects the viewers' Quality-of-Experiences (QoE) when watching the replays. To address this issue and improve viewers' QoE, we propose a crowdsourced framework HighCast based on the interactive messages contributed by the viewers in crowdsourced livecast services. According to a highlight-aware detection module, HighCast can exploit the detection results to schedule the content placement by considering the importance of the predicted streaming highlights. The trace-based evaluations illustrate that the proposed framework improves the prediction accuracy and reduces the viewing latency.
Cong Zhang 0002, Jiangchuan Liu, Haitian Pang, Fangxin Wang 0001
IWQoS4
2018 Optimizing Personalized Interaction Experience in Crowd-Interactive Livecast: A Cloud-Edge Approach
abstract
Enabling users to interact with broadcasters and audience, the crowd-interactive livecast greatly improves viewer's quality of experience (QoE) and attracts millions of daily active users recently. In addition to striking the balance between resource utilization and viewers' QoE met in the traditional video streaming service, this novel service needs to take supererogatory efforts to improve the interaction QoE, which reflects the viewer interaction experience. To tackle this issue, we conduct measurement studies over a large-scale dataset crawled from a representative livecast service provider. We observe that the individual's interaction pattern is quite heterogeneous: only 10% viewers proactively participate in the interaction, and the rest viewers usually watch passively. Incorporating the insight into the emerging cloud-edge architecture, we propose a framework PIECE, which optimizes the Personalized Interaction Experience with Cloud-Edge architecture (PIECE) for intelligent user access control and livecast distribution. In particular, we first devise a novel deep neural network based algorithm to predict users' interaction intensity using the historical viewer pattern. We then design an algorithm to maximize the individual's QoE, by strategically matching viewer sessions and transcoding-delivery paths over cloud-edge infrastructure. Finally, we use trace-driven experiments to verify the effectiveness of PIECE. Our results show that our prediction algorithm outperforms the state-of-the-art algorithms with a much smaller mean absolute error (40% reduction). Furthermore, in comparison with the cloud-based video delivery strategy, the proposed framework can simultaneously improve the average viewers QoE (26% improvement) and interaction QoE (21% improvement), while maintaining a high streaming bitrate.
Haitian Pang, Cong Zhang 0002, Fangxin Wang 0001, Han Hu 0003, Zhi Wang 0001, Jiangchuan Liu, Lifeng Sun
ACM Multimedia3
2016 DVMP: Incremental traffic-aware VM placement on heterogeneous servers in data centers
abstract
As the tremendous momentum cloud computing has grown, the modern data center networks are facing challenge to handle the increasing traffic demand among virtual machines (VMs). Simply adding more switches and links may increase network capacity but at the same time increase the complexity and infrastructure cost. Thus, intelligent VM placement has been proposed to reduce the intra-DC traffic. Prior solutions model the traffic-aware VM placement problem as a Balanced Minimum K-cut Problem (BMKP). However, the assumptions of “once-for-all” VM placement on physical servers with equal VM slots are often not realistic in practical data centers, and thus the naive BMKP model may lead to suboptimal placement solutions. In this work, we revisit the problem by considering the server heterogeneity and propose an incremental traffic-aware VM placement algorithm. Given that the BMKP model cannot be directly applied, we make a number of transformations to re-establish the model. First, by introducing pseudo VM slots on physical servers with less VM slots, we allow the number of available VM slots of each server to be different. Second, pseudo edges with infinite costs are added between existing VMs, and thus previously deployed VMs on the same physical server will still be packed together. Third, a change on the number of pseudo VM slots is applied, so that existing VMs placed on different physical servers will still be separated. In this way, we reduce the problem to a new BMKP problem, which results in a much better solution. The evaluation results show that DVMP can reduce up to 28%, 39% and 55% traffic compared with naive BMKP model, greedy VM placement and random VM placement, respectively.
Dan Li 0001, Syed Shah-e-Mardan Ali Rizvi, Fangxin Wang 0001, Wu He
IWQoS3
2016 CCDN: Content-Centric Data Center Networks
abstract
Data center networks continually seek higher network performance to meet the ever increasing application demand. Recently, researchers are exploring the method to enhance the data center network performance by intelligent caching and increasing the access points for hot data chunks. Motivated by this, we come up with a simple yet useful caching mechanism for generic data centers, i.e., a server caches a data chunk after an application on it reads the chunk from the file system, and then uses the cached chunk to serve subsequent chunk requests from nearby servers. To turn the basic idea above into a practical system and address the challenges behind it, we design content-centric data center networks (CCDNs), which exploits an innovative combination of content-based forwarding and location [Internet Protocol (IP)]-based forwarding in switches, to correctly locate the target server for a data chunk on a fully distributed basis. Furthermore, CCDN enhances traditional content-based forwarding to determine the nearest target server, and enhances traditional location (IP)-based forwarding to make high utilization of the precious memory space in switches. Extensive simulations based on real-world workloads and experiments on a test bed built with NetFPGA prototypes show that, even with a small portion of the server's storage as cache (e.g., 3%) and with a modest content forwarding information base size (e.g., 1000 entries) in switches, CCDN can improve the average throughput to get data chunks by 43% compared with a pure Hadoop File System (HDFS) system in a real data center.
Dan Li 0001, Fangxin Wang 0001, Anke Li, K. K. Ramakrishnan, Ying Liu 0024, Xue (Steve) Liu
IEEE/ACM Trans. Netw.3
2015 Bandwidth guaranteed virtual network function placement and scaling in datacenter networks
abstract
Enterprises deploy their middlebox services in cloud seeking for easy management, flexible scalability and economic savings. However, existing elastic virtual network function(VNF) placement strategy often leads to an unpredictable placing location due to the ever-changing workload, which may waste much precious bandwidth resource and bring a lot of VM operation overhead(e.g. VM launch, termination and migration). A key problem for cloud providers is how to conduct an effective service placement and provide resource provision according to various workload, satisfying the bandwidth requirement of each service while saving as much cloud resource as possible. In this paper we solve both the virtual network function(VNF) placement and scaling problem based on preplanned allocation with bandwidth guarantee. We first propose a concept of VNF instance communication graph to describe the bandwidth demand of each VNF instance and explore the placement requirement for bandwidth savings. Then we design an on-line heuristic algorithm to achieve approximate optimal allocation. At last, we also provide an off-line optimal solution for comparison. Our simulation shows that our heuristic solution saves 20% more bandwidth resource and reduce more VM migration overhead than existing elastic placement solution. Its performance is also very close to the optimal solution.
Fangxin Wang 0001, Ruilin Ling, Jing Zhu 0007, Dan Li 0001
IPCCC1