Xiuquan Qiao

dblp:66/4187 · DBLP profile ↗
← Back
43ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-0140-0650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 21 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ShadowLLM: Resource-Efficient Hot Standby for Heterogeneous Edge LLM Serving
Zhenguo Chen, Pujun Ding, Yuanwei Zhu, Jing Lv, Yakun Huang, Xiuquan Qiao
ICDCS7
2026 Intent-Driven Cognitive XR Networks: Multi-Agent Orchestration for Immersive Communication
abstract
Extended Reality (XR) is emerging as a central use case in the 6G era, requiring intelligent, adaptive, and low-latency communication to support immersive user experiences. However, existing network architectures remain reactive and lack the capability to interpret or act upon users’ fine-grained multimodal intents, limiting their responsiveness and efficiency. To address this gap, this paper presents theCognitive XR Network (CXN)framework, an intent-driven architecture that integrates perception, reasoning, and control through a hierarchy of Artificial Intelligence (AI) agents. CXN comprises three cooperating agents: the User Intent Agent, which infers and predicts user intents from multimodal sensory inputs; the Network Orchestration Agent, which performs global coordination through multi-agent reinforcement learning; and distributed Resource Management Agents which execute localized decisions under strategic guidance. Together, these agents form a closed cognitive loop that enables proactive, intent-aware orchestration across radio, compute, and caching domains. Simulations under realistic XR collaboration scenarios demonstrate that CXN sustains an intent satisfaction rate above 85% and an average interaction latency below 60 ms with 90 concurrent users, outperforming both reactive and centralized learning baselines.
Yakun Huang, Yaru Zhao 0001, Zhenguo Chen, Jing Lv, Xiuquan Qiao
IEEE J. Sel. Areas Commun.6
2026 PortaCap: Portable Volumetric Video Capturing System for Metaverse Interaction
abstract
Portable volumetric video capturing systems present a compelling alternative to traditional, bulky prototype systems used for streaming and interacting with volumetric content. Their key advantages, particularly flexibility and ease of deployment, make them suitable for a wide range of applications. However, despite their potential, there has been limited exploration into the design of such portable systems. This paper addresses this gap by conducting an in-depth analysis of portability and proposing an optimized camera array configuration tailored for high-quality volumetric content generation. Our approach begins with a novel, flexible camera calibration method that leverages geometric priors, enabling accurate alignment of multiple cameras without requiring specialized expertise. Building on this, we introduce a meticulous fusion technique that integrates captured and inferred data to reconstruct complete volumetric representations. This method achieves a fusion latency of less than 100 ms, ensuring real-time performance. We integrate these innovations into a portable capturing system, named PortaCap, which incorporates a carefully designed camera deployment strategy. Through both quantitative and qualitative evaluations, PortaCap demonstrates significant improvements in volumetric content quality, achieving enhancements ranging from 13% to 33.8%. These results underscore the system's potential to advance the state-of-the-art in portable volumetric video capture.
Chongli Zhang, Yakun Huang, Yuanwei Zhu, Dexing Cai, Shibo Fang, Chunsheng Wang, Xiuquan Qiao
IEEE Trans. Mob. Comput.9
2025 WebGS360: Towards web-based visualization of Gaussian Splatting from panoramic images
Chongli Zhang, Jing Lv, Xiuquan Qiao, Yakun Huang
Comput. Graph.5
2025 Two grids are better than one: Hybrid indoor scene reconstruction framework with adaptive priors
abstract
Indoor scene reconstruction from multi-view images is a pivotal technology within the field of robotics and augmented reality . Previous researches have predominantly focused on neural radiance fields aided by geometric monocular priors. However, due to the inductive smoothness bias introduced by deep Multi-Layer Perceptron (MLP) networks, these methods struggle to recover the scene surface with complex and fine geometry details. Additionally, when used as additional supervision signals during optimization, priors in different regions make different contributions. Simply incorporating them in all regions may lead to a decrease in the accuracy. To tackle these issues, we present a generic end-to-end framework named AdaptSurf, which combines Signed Distance Field (SDF) voxel grids and feature voxel grids to enhance the capability of reconstructing accurate geometry details, respectively. Furthermore, we design a policy network to adaptively enable the estimated depth or normal priors to supervise the learning process, which improves the reconstruction accuracy and accelerates neural surface reconstruction. Qualitative and quantitative experiments show that AdaptSurf yields high-quality surfaces, especially for fine-grained details and smooth regions. Furthermore, the policy network exhibits an interpretable behavior that depends on the voxel features, which helps to improve the quality of surface reconstruction.
Boyuan Bai, Xiuquan Qiao, Hongru Zhao, Wenzhe Shi, Hengjia Zhang, Yakun Huang
Neurocomputing2
2025 WebARNav: Mobile Web AR Indoor Navigation With Edge-Assisted Vision Localization
abstract
The gradual maturation of mobile augmented reality (AR) and localization technologies is enabling the development of immersive AR-enabled indoor localization and navigation systems. Existing indoor localization technologies (e.g., WiFi, infrared, Bluetooth) and navigation services do not provide intuitive 3D AR experiences and can be expensive to deploy. This paper introduces WebARNav, a cross-platform indoor localization system that provides user-friendly AR navigation services with low overhead and remarkable accuracy. First, we propose a lightweight location fusion framework for indoor navigation on the mobile web, which leverages accurate edge-supported vision localization to guide and correct lightweight pedestrian dead reckoning localization. Second, we improve the accuracy of localization using an attention-based feature extraction method and a dual-stream retrieval and co-visibility re-ranking technique for initial localization. Third, we significantly improve accuracy and speed up retrieval as users move by generating a topological map for traveling localization. We conducted extensive experiments on various indoor datasets to demonstrate localization accuracy and navigation experience. The study shows that WebARNav achieves a localization frequency of over 30 Hz and reduces the average trajectory error by 76% and 95% for single- and multi-floor office scenes, respectively, compared to the PDR-only method. The proposed traveling localization method also reduces the localization latency by 15.2%, 55.1%, and 98.6% in the baseline datasets, with an accuracy improvement of over 4%.
Yakun Huang, Shengwei Meng, Yuanwei Zhu, Jacky Cao, Xiuquan Qiao, Xiang Su 0001
IEEE Trans. Mob. Comput.6
2025 FPSelector: A Flexible Path Selector for Mobile Augmented Reality Offloading
abstract
Mobile Augmented Reality (MAR) applications pose unique challenges due to computation intensity, constrained device resources, and high interactive rendering requirements. The emergence of 5 G and edge computing offers opportunities to offload computation to the edge and cloud, indirectly enhancing the computing capability and usage duration of MAR devices. However, existing general task offloading and multipath transmission techniques do not address the challenges in offloading path selection with multiple edges, dynamic resource competition awareness, and spatial computation with strong task dependencies. This paper contributes FPSelector, a flexible path selector for MAR offloading. We present a two-tier MAR-specific offloading scheme with multiple edge nodes. In offloading decisions, we design a reinforcement learning model to generate the selection policy for each packet of an AR data stream. This model incorporates an action masking mechanism, a comprehensive reward function, and state features complemented by a resource prediction module, making FPSelector aware of dynamic heterogeneous environments. Moreover, we propose an online learning strategy to facilitate real-time selection. To validate its efficacy, we compare FPSelector's performance against leading schedulers under various scenarios, demonstrating a notable reduction of 9.9% and 9.6% in overall completion time for 4 K and 8 K video-based MAR applications compared to its closest competitor.
Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Xiaoli Liu 0005, Xiang Su 0001, Anna Brunström, Özgü Alay, Sasu Tarkoma
IEEE Trans. Mob. Comput.3
2024 SM3: Self-supervised Multi-task Modeling with Multi-view 2D Images for Articulated Objects
abstract
Reconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on annotated datasets to model articulated objects within limited categories. However, these approaches fall short of effectively addressing the diversity present in the real world. To tackle this issue, we propose a self-supervised interaction perception method, referred to as SM3, which leverages multi-view RGB images captured before and after interaction to model articulated objects, identify the movable parts, and infer the parameters of their rotating joints. By constructing 3D geometries and textures from the captured 2D images, SM3achieves integrated optimization of movable part and joint parameters during the reconstruction process, obviating the need for annotations. Furthermore, we introduce the MMArt dataset, an extension of PartNet-Mobility, encompassing multi-view and multi-modal data of articulated objects spanning diverse categories. Evaluations demonstrate that SM3surpasses existing benchmarks across various categories and objects, and its adaptability in real-world scenarios has been thoroughly validated.
Haowen Wang 0001, Zhengping Che, Yakun Huang, Xiuquan Qiao, Jian Tang 0008
ICRA8
2024 ISCom: Interest-Aware Semantic Communication Scheme for Point Cloud Video Streaming on Metaverse XR Devices
abstract
In the metaverse era, point cloud video (PCV) streaming on mobile XR devices is pivotal. While most current methods focus on PCV compression from traditional 3-DoF video services, emerging AI techniques extract vital semantic information, producing content resembling the original. However, these are early-stage and computationally intensive. To enhance the inference efficacy of AI-based approaches, accommodate dynamic environments, and facilitate applicability to metaverse XR devices, we present ISCom, an interest-aware semantic communication scheme for lightweight PCV streaming. ISCom is featured with a region-of-interest (ROI) selection module, a lightweight encoder-decoder training module, and a learning-based scheduler to achieve real-time PCV decoding and rendering on resource-constrained devices. ISCom’s dual-stage ROI selection provides significantly reduces data volume according to real-time interest. The lightweight PCV encoder-decoder training is tailored to resource-constrained devices and adapts to the heterogeneous computing capabilities of devices. Furthermore, We provide a deep reinforcement learning (DRL)-based scheduler to select optimal encoder-decoder model for various devices adaptivelly, considering the dynamic network environments and device computing capabilities. Our extensive experiments demonstrate that ISCom outperforms baselines on mobile devices, achieving a minimum rendering frame rate improvement of 10 FPS and up to 22 FPS. Furthermore, our method significantly reduces memory usage by 41.7% compared to the state-of-the-art AITransfer method. These results highlight the effectiveness of ISCom in enabling lightweight PCV streaming and its potential to improve immersive experiences for emerging metaverse application.
Yakun Huang, Boyuan Bai, Yuanwei Zhu, Xiuquan Qiao, Xiang Su 0001, Lei Yang 0063, Ping Zhang 0003
IEEE J. Sel. Areas Commun.4
2024 RDFC-GAN: RGB-Depth Fusion CycleGAN for Indoor Depth Completion
abstract
Raw depth images captured in indoor scenarios frequently exhibit extensive missing values due to the inherent limitations of the sensors and environments. For example, transparent materials frequently elude detection by depth sensors; surfaces may introduce measurement inaccuracies due to their polished textures, extended distances, and oblique incidence angles from the sensor. The presence of incomplete depth maps imposes significant challenges for subsequent vision applications, prompting the development of numerous depth completion techniques to mitigate this problem. Numerous methods excel at reconstructing dense depth maps from sparse samples, but they often falter when faced with extensive contiguous regions of missing depth values, a prevalent and critical challenge in indoor environments. To overcome these challenges, we design a novel two-branch end-to-end fusion network named RDFC-GAN, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure, by adhering to the Manhattan world assumption and utilizing normal maps from RGB-D information as guidance, to regress the local dense depth values from the raw depth map. The other branch applies an RGB-depth fusion CycleGAN, adept at translating RGB imagery into detailed, textured depth maps while ensuring high fidelity through cycle consistency. We fuse the two branches via adaptive fusion modules named W-AdaIN and train the model with the help of pseudo depth maps. Comprehensive evaluations on NYU-Depth V2 and SUN RGB-D datasets show that our method significantly enhances depth completion performance particularly in realistic indoor settings.
Haowen Wang 0001, Zhengping Che, Mingyuan Wang 0003, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 HiVAT: Improving QoE for Hybrid Video Streaming Service With Adaptive Transcoding
abstract
Mobile video streaming enables flexible delivery of videos to mobile devices, supporting emerging video formats. The transition from conventional 2D videos to immersive formats, such as virtual reality and holographic videos, significantly increases the demand for computation and network resources. Existing streaming techniques are predominantly developed for specific video types, neglecting fair adaptive transmission and optimal resource utilization in services involving multiple video types. This paper investigates hybrid video streaming, encompassing 2D, 360-degree, and volumetric videos. To accommodate resource-intensive hybrid video streaming on mobile devices, we proposeHiVAT, an adaptive transcoding-based system that ensures Quality of Experience (QoE) for each stream type. We contribute 1) a transcoding-based framework to address the challenges of high bandwidth and decoding overhead on mobile devices; 2) a universal QoE model involving traditional factors, viewport smoothness, degree of immersion, etc., for transcoded video streams; 3) a multi-agent adaptive bitrate controller that collaboratively determines hybrid video quality levels to achieve high and fair QoE across multiple streams; and 4) a learning-based task scheduler to optimize computation resource usage, thereby improving the overall serviceability of the system. We evaluateHiVATagainst state-of-the-art methods, witnessing an average QoE improvement of 5.9% and 9.9% on linear and logarithmic metrics, respectively.
Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Xiang Su 0001
IEEE Trans. Mob. Comput.3
2023 Cloud-Edge-Device Collaborative Image Retrieval and Recognition for Mobile Web
Yakun Huang, Shouyi Wu, Xiuquan Qiao, Hongshun He
CollaborateCom (2)4
2023 DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template Field
abstract
Estimating 6D poses and reconstructing 3D shapes of objects in open-world scenes from RGB-depth image pairs is challenging. Many existing methods rely on learning geometric features that correspond to specific templates while disregarding shape variations and pose differences among objects in the same category. As a result, these methods underperform when handling unseen object instances in complex environments. In contrast, other approaches aim to achieve category-level estimation and reconstruction by leveraging normalized geometric structure priors, but the static prior-based reconstruction struggles with substantial intra-class variations. To solve these problems, we propose the DTF-Net, a novel framework for pose estimation and shape reconstruction based on implicit neural fields of object categories. In DTF-Net, we design a deformable template field to represent the general category-wise shape latent features and intra-category geometric deformation features. The field establishes continuous shape correspondences, deforming the category template into arbitrary observed instances to accomplish shape reconstruction. We introduce a pose regression module that shares the deformation features and template codes from the fields to estimate the accurate 6D pose of each object in the scene. We integrate a multi-modal representation extraction module to extract object features and semantic masks, enabling end-to-end inference. Moreover, during training, we implement a shape-invariant training strategy and a viewpoint sampling method to further enhance the model's capability to extract object pose features. Extensive experiments on the REAL275 and CAMERA25 datasets demonstrate the superiority of DTF-Net in both synthetic and real scenes. Furthermore, we show that DTF-Net effectively supports grasping tasks with a real robot arm.
Haowen Wang 0001, Zhengping Che, Dong Liu 0058, Feifei Feng, Yakun Huang, Xiuquan Qiao, Jian Tang 0008
ACM Multimedia9
2023 Graph attention network-optimized dynamic monocular visual odometry
Hongru Zhao, Xiuquan Qiao
Appl. Intell.2
2023 An Integrated Cloud-Edge-Device Adaptive Deep Learning Service for Cross-Platform Web
abstract
Deep learning shows great promise in providing more intelligence to the cross-platform web. However, insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning with low-performing web browsers. We propose DeepAdapter, an integrated cloud-edge-device framework that ties the edge, the remote cloud, with the device by cross-platform web technology for adaptive deep learning services towards lower latency, lower mobile energy, and higher system throughput. DeepAdapter consists of context-aware pruning, service updating, and online scheduling. First, the offline pruning module provides a context-aware pruning algorithm that incorporates the latency, the network condition, and the device's computing capability to fit various contexts. Second, the service updating module optimizes branch model cache on the edge for massive mobile users and updates the new model pruning requirements. Third, the online scheduling module matches optimal branch models for mobile users. Also, a two-stage DRL-based online scheduling method named DeepScheduler can handle high concurrent requests between edge centers and remote cloud by designing the reward prediction model. Extensive experiments show that DeepAdapter can decrease average latency by 1.33x, reduce average mobile energy by 1.4x, and improve system throughput by 2.1x with considerable accuracy.
Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
IEEE Trans. Mob. Comput.2
2023 A Semantic-Aware Transmission With Adaptive Control Scheme for Volumetric Video Service
abstract
Volumetric video provides a more immersive holographic virtual experience than conventional video services such as 360-degree and virtual reality (VR) videos. However, due to ultra-high bandwidth requirements, existing compression and transmission technology cannot handle the delivery of real-time volumetric video. Unlike traditional compression methods and the approaches that extend 360-degree video streaming, we propose AITransfer, an AI-powered compression and semantic-aware transmission method for point cloud video data (a popular volumetric data format). AITransfer targets the semantic-level communication beyond transmitting raw point cloud video or compressed video with two outstanding contributions: (1) designing an integrated end-to-end architecture with two fundamental contents of feature extraction and reconstruction to reduce the bandwidth consumption and alleviate the computational pressure; and (2) incorporating the dynamic network condition into end-to-end architecture design and employing a deep reinforcement learning-based adaptive control scheme to provide robust transmission. We conduct extensive experiments on the typical datasets and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide extremely efficient point cloud transmission while maintaining considerable user experience with more than 30.72x compression ratio under the existing network environments.
Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Zhijie Tan, Boyuan Bai, Huadong Ma, Schahram Dustdar
IEEE Trans. Multim.3
2023 Distributed Edge System Orchestration for Web-Based Mobile Augmented Reality Services
abstract
The emergence of edge computing and 5G networks has fueled the growth of mobile Web AR. Although efforts have been made to improve the edge system efficiency for Web AR applications, efficient edge-assisted mobile Web AR services remain technically challenging. This paper presents EARNet, a distributed edge system orchestration approach for mobile Web AR in 5G networks. The design of EARNet makes three novel contributions. First, EARNet manages the edge network dynamics with respect to user mobility and their Web AR service requests by employing landmarks and grid index based edge node localization mechanisms. Second, EARNet takes into account both request serving performance and offloading cost in managing workload balance and quality of service and leverages dynamic hash and max heap mechanisms for efficient Web AR service lookup and AR computations. Third, EARNet designs the service migration schemes by optimizing several performance factors, such as message efficiency, scheduling latency, request density and locality of mobile users and edge nodes, and accuracy of Web AR services after migration. Experimental evaluations are conducted using the real base station deployment data in the Melbourne Central Business District (CBD) area. The results shows the effectiveness of the EARNet edge orchestration approach compared to several baseline approaches.
Pei Ren, Ling Liu 0001, Xiuquan Qiao, Junliang Chen 0001
IEEE Trans. Serv. Comput.3
2022 RGB-Depth Fusion GAN for Indoor Depth Completion
abstract
The raw depth image captured by the indoor depth sen-sor usually has an extensive range of missing depth values due to inherent limitations such as the inability to perceive transparent objects and limited distance range. The incomplete depth map burdens many downstream vision tasks, and a rising number of depth completion methods have been proposed to alleviate this issue. While most existing meth-ods can generate accurate dense depth maps from sparse and uniformly sampled depth maps, they are not suitable for complementing the large contiguous regions of missing depth values, which is common and critical. In this paper, we design a novel two-branch end-to-end fusion network, which takes a pair of RGB and incomplete depth images as input to predict a dense and completed depth map. The first branch employs an encoder-decoder structure to regress the local dense depth values from the raw depth map, with the help of local guidance information extracted from the RGB image. In the other branch, we propose an RGB-depth fusion GAN to transfer the RGB image to the fine-grained textured depth map. We adopt adaptive fusion modules named W-AdaIN to propagate the features across the two branches, and we append a confidence fusion head to fuse the two out-puts of the branches for the final depth map. Extensive ex-periments on NYU-Depth V2 and SUN RGB-D demonstrate that our proposed method clearly improves the depth completion performance, especially in a more realistic setting of indoor environments with the help of the pseudo depth map.
Haowen Wang 0001, Mingyuan Wang 0003, Zhengping Che, Xiuquan Qiao, Mengshi Qi, Feifei Feng, Jian Tang 0008
CVPR5
2022 AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile Web
abstract
Employing today’s deep neural network (DNN) into the cross-platform web with an offloading way has been a promising means to alleviate the tension between intensive inference and limited computing resources. However, it is still challenging to directly leverage the distributed DNN execution into web apps with the following limitations, including (1) how special computing tasks such as DNN inference can provide fine-grained and efficient offloading in the inefficient JavaScript-based environment? (2) lacking the ability to balance the latency and mobile energy to partition the inference facing various web applications’ requirements. (3) and ignoring that DNN inference is vulnerable to the operating environment and mobile devices’ computing capability, especially dedicated web apps. This paper designs AoDNN, an automatic offloading framework to orchestrate the DNN inference across the mobile web and the edge server, with three main contributions. First, we design the DNN offloading based on providing a snapshot mechanism and use multi-threads to monitor dynamic contexts, partition decision, trigger offloading, etc. Second, we provide a learning-based latency and mobile energy prediction framework for supporting various web browsers and platforms. Third, we establish a multi-objective optimization to solve the optimal partition by balancing the latency and mobile energy.
Yakun Huang, Xiuquan Qiao, Schahram Dustdar
INFOCOM2
2022 Enabling DNN Acceleration With Data and Model Parallelization Over Ubiquitous End Devices
abstract
Deep neural network (DNN) shows great promise in providing more intelligence to ubiquitous end devices. However, the existing partition-offloading schemes adopt data-parallel or model-parallel collaboration between devices and the cloud, which does not make full use of the resources of end devices for deep-level parallel execution. This article proposes eDDNN (i.e., enabling Distributed DNN), a collaborative inference scheme over heterogeneous end devices using cross-platform Web technology, moving the computation close to ubiquitous end devices, improving resource utilization, and reducing the computing pressure of data centers. eDDNN implements D2D communication and collaborative inference among heterogeneous end devices with WebRTC protocol, divides the data and corresponding DNN model into pieces simultaneously, and then executes inference almost independently by establishing a layer dependency table. Besides, eDDNN provides a dynamic allocation algorithm based on deep reinforcement learning to minimize latency. We conduct experiments on various data sets and DNNs and further employ eDDNN into a mobile Web AR application to illustrate the effectiveness. The results show that eDDNN can achieve the latency decrease by$2.98\times $, reduce mobile energy by$1.8\times $, and relieve the computing pressure of the edge server by$2.57\times $, against a typical partition-offloading approach.
Yakun Huang, Xiuquan Qiao, Wenhai Lai, Schahram Dustdar, Jiulin Li
IEEE Internet Things J.2
2022 A Collaborative Task Offloading Framework for Smart TV Applications in a Household Computing Environment
abstract
Smart TV can perform interactive computing while also providing video content services. However, this leads to a high delay during interactive computing because of the lack of computing capability, thus smart TVs are unable to undertake large scenes and complex interactive computing tasks. This article proposes a collaborative task offloading framework (CTOF) for interactive computing of smart TV video applications. The main contributions of this article are as follows: 1) a computing offloading mechanism is proposed for interactive computing with video content. A part of the interactive computing task is offloaded to the user’s high-computing mobile device in a household video service environment via a Wi-Fi Direct channel and 2) a collaborative computing offloading algorithm is proposed for complex interactive computing tasks. According to the computing complexity and the computing expansion coefficient, a parallel and serial collaborative smart offloading framework is established to minimize the delay. We conduct extensive experiments to indicate that in the existing experimental network environment, the video-based complex interactive service operation efficiency is improved by 20%. With a smart offloading framework, we can further achieve a satisfactory experience in terms of the interactive computing delay. Consequently, the smart TV can quickly respond to the complex video interactive service and improve the business interaction capability of the smart TV.
Liang Li 0023, Xiuquan Qiao, Huabing Zhang, Yakun Huang, Pei Ren
IEEE Internet Things J.2
2022 Edge AR X5: An Edge-Assisted Multi-User Collaborative Framework for Mobile Web Augmented Reality in 5G and Beyond
abstract
Multi-user mobile Augmented Reality (AR) has been successfully used in various fields as a novel visual interaction technology. But current mainstream wearable device-based and app-based solutions are still facing cross-platform, real-time communication, and intensive computing requirements. Mobile Web technology is envisioned to be a promising supporting technology for cross-platform application of mobile AR especially in 5G networks, which provide pervasive communication and computing resources thereby forming a formidable framework for the practical application of multi-user mobile Web AR. However, the problem of how to use these new techniques properly to achieve efficient communication and computing collaboration is obviously paramount in order for multi-user mobile Web AR to be realized in 5G networks. In this article, we propose the first edge-assisted multi-user collaborative framework for mobile Web AR in the 5G era. First, we propose a heuristic mechanism BA-CPP for efficient communication planning, which allows multi-user interaction synchronization to be achieved. Second, we introduce a motion-aware key frame selection mechanism called Mo-KFP to optimize the computational efficiency of the edge system, and simultaneously alleviate the initialization problem by collaborating with nearby mobile devices using the Device-to-Device (D2D) communication technique. Experiments are conducted in a real-world 5G network, and the results demonstrate the superiority of our proposed collaborative framework.
Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001
IEEE Trans. Cloud Comput.2
2022 A Lightweight Collaborative Deep Neural Network for the Mobile Web in Edge Cloud
abstract
Enabling deep learning technology on the mobile web can improve the user’s experience for achieving web artificial intelligence in various fields. However, heavy DNN models and limited computing resources of the mobile web are now unable to support executing computationally intensive DNNs when deploying in a cloud computing platform. With the help of promising edge computing, we propose a lightweight collaborative deep neural network for the mobile web, named LcDNN, which contributes to three aspects: (1) We design a composite collaborative DNN that reduces the model size, accelerates inference, and reduces mobile energy cost by executing a lightweight binary neural network (BNN) branch on the mobile web. (2) We provide a jointly training method for LcDNN and implement an energy-efficient inference library for executing the BNN branch on the mobile web. (3) To further promote the resource utilization of the edge cloud, we develop a DRL-based online scheduling scheme to obtain an optimal allocation for LcDNN. The experimental results show that LcDNN outperforms existing approaches for reducing the model size by about 16x to 29x. It also reduces the end-to-end latency and mobile energy cost with acceptable accuracy and improves the throughput and resource utilization of the edge cloud.
Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001
IEEE Trans. Mob. Comput.2
2022 Fine-Grained Elastic Partitioning for Distributed DNN Towards Mobile Web AR Services in the 5G Era
abstract
Web-based Deep Neural Networks (DNNs) enhance the ability of object recognition and has attracted considerable attention in mobile Web AR and other services. However, neither performing the DNN inference on mobile Web browsers locally nor offloading computations to the cloud can strike a balance between accuracy and efficiency; generally, rude methods are often accompanied by unsatisfactory accuracy. Collaborative approaches seem to fill this gap by coordinating the distributed hierarchical computing resources, especially in the 5G era, but it still faces challenges in the current solutions, such as the lack of (1) full use of 5G resources for the one point DNN computation partitioning schemes; (2) fine-grained branching mechanism; (3) efficient partitioning method; and (4) multi-objective optimization. To this end, we present the fine-grained elastic computation partitioning mechanism for distributed DNN in 5G networks. First, we elaborate two collaborative scenarios. Second, we study the DNN branching mechanism at layer granularity. Next, we propose a DNN computation partitioning algorithm based on deep reinforcement learning. Finally, we develop a mobile Web AR application as a proof of concept. The experiments were conducted in an actually deployed 5G trial network, and the results show the superiority of this collaborative approach. The common theme is, under the premise that Quality of Service (QoS) is satisfied, to balance multiple interests by orchestrating computations across heterogeneous computing platforms.
Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar
IEEE Trans. Serv. Comput.2
2021 Towards Video Streaming Analysis and Sharing for Multi-Device Interaction with Lightweight DNNs
abstract
Multi-device interaction has attracted a growing interest in both mobile communication industry and mobile computing research community as mobile devices enabled social media and social networking continue to blossom. However, due to the stringent low latency requirements and the complexity and intensity of computation, implementing efficient multi-device interaction for real-time video streaming analysis and sharing is still in its infancy. Unlike previous approaches that rely on high network bandwidth and high availability of cloud center with GPUs to support intensive computations for multi-device interaction and for improving the service experience, we propose MIRSA, a novel edge centric multi-device interaction framework with a lightweight end-to-end DNN for on-device visual odometry (VO) streaming analysis by leveraging edge computing optimizations with three main contributions. First, we design MIRSA to migrate computations from the cloud to the device side, reducing the high overhead for large transmission of video streaming while alleviating the server load of the cloud. Second, we design a lightweight VO network by utilizing temporal shift module to support on-device pose estimation. Third, we provide on-device resource-aware scheduling algorithm to optimize the task allocation. Extensive experiments show MIRSA provides real-time high quality pose estimation as an interactive service and outperforms baseline methods.
Yakun Huang, Hongru Zhao, Xiuquan Qiao, Jian Tang 0008, Ling Liu 0001
INFOCOM3
2021 AITransfer: Progressive AI-powered Transmission for Real-Time Point Cloud Video Streaming
abstract
Point cloud video provides a more immersive holographic virtual experience than conventional video services such as 360 degree video and virtual reality (VR) video. However, the existing network bandwidth and transmission technology can not carry real-time point cloud video streaming due to mass data volume, high processing overheads, and extremely bandwidth-consuming. Unlike previous approaches that extend the VR video streaming, we propose AITransfer, an AI-powered bandwidth-aware and adaptive transmission technique driven by extracting and transferring key point cloud features to reduce the bandwidth consumption and alleviate the computational pressure. AITransfer has two outstanding contributions, including (1) incorporating the dynamic network bandwidth into the design of an end-to-end architecture with two fundamental contents of feature extraction and reconstruction, and (2) employing an online adapter to sense the network bandwidth and match the optimal inference model. We conduct extensive experiments on the typical dataset and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide more than 30.72 times compression ratio under the existing network environments.
Yakun Huang, Yuanwei Zhu, Xiuquan Qiao, Zhijie Tan, Boyuan Bai
ACM Multimedia3
2021 EdgeBooster: Edge-Assisted Real-Time Image Segmentation for the Mobile Web in WoT
abstract
Combining image segmentation with Web technology lays a good foundation for lightweight, cross-platform, and pervasive Web artificial intelligence applications, and further improves the capability of Web-of-Things (WoT) applications. However, no matter whether we use a Web real-time communication media server for advanced processing that views camera inputs as a video stream, or transfer continuous camera frames to the remote cloud for processing, we are unable to obtain a satisfactory real-time experience due to high resource consumption and unacceptable latency. In this article, we present EdgeBooster, a computational-efficient architecture that leverages a common edge server to minimize the communication costs, accelerates the camera frame segmentation, and guarantees an acceptable segmentation accuracy with the prior knowledge. EdgeBooster provides real-time segmentation by developing parallel technology that enables segmentation on slices of a camera frame and using presegmentation based on superpixels to accelerate the graph-based segmentation. It also introduces recent DNN-based segmentation results as the prior knowledge to improve the performance of the graph-based segmentation, especially in nonideal scenes, such as dark light and weak contrast. Finally, it creates a pure frontend segmentation that can provide continuous and stable services for mobile users in unstable networks, such as a weak network or with an unstable edge server. The experimental results show that EdgeBooster is able to achieve a considerable accuracy for the mobile Web, running at no less than 30 frames per second in real scenes.
Yakun Huang, Xiuquan Qiao, Pei Ren, Schahram Dustdar, Junliang Chen 0001
IEEE Internet Things J.2
2020 DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network Pruning
abstract
Deep learning shows great promise in providing more intelligence to the mobile web, but insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning in mobile web applications. In this paper, we present DeepAdapter, a collaborative framework that ties the mobile web with an edge server and a remote cloud server to allow executing deep learning on the mobile web with lower processing latency, lower mobile energy, and higher system throughput. DeepAdapter provides a context-aware pruning algorithm that incorporates the latency, the network condition and the computing capability of the mobile device to fit the resource constraints of the mobile web better. It also provides a model cache update mechanism improving the model request hit rate for mobile web users. At runtime, it matches an appropriate model with the mobile web user and provides a collaborative mechanism to ensure accuracy. Our results show that DeepAdapter decreases average latency by 1.33x, reduces average mobile energy consumption by 1.4x, and improves system throughput by 2.1x with a considerable accuracy. Its contextaware pruning algorithm also improves inference accuracy by up to 0.3% with a smaller and faster model.
Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
INFOCOM2
2020 Interest packets scheduling and size-based flow control mechanism for content-centric networking web servers
Xiuquan Qiao, Pei Ren, Yukai Tu, Guoshun Nan, Junliang Chen 0001, M. Brian Blake
Future Gener. Comput. Syst.1
2019 A Lightweight Collaborative Recognition System with Binary Convolutional Neural Network for Mobile Web Augmented Reality
abstract
Lightweight and precise recognition is a key component of web-based augmented reality (Web AR) applications. Although edge-based distributed deep learning approach is now possible to achieve satisfactory recognition for Web AR applications, it puts significant pressure on the computation and energy consumption of the mobile web browser, especially the app-based embedded browser. Thus, reducing the model size and accelerating the inference are regarded as the two fundamental challenges to enable this edge-based collaborative recognition system efficiently. In this paper, we propose a lightweight collaborative recognition system (LCRS) for Web AR applications. LCRS contributes to three aspects: (1) we design a composite deep neural network for reducing the model size and inference latency by introducing binary convolutional neural network; (2) we provide a joint training method to co-train the general branch and the binary branch; (3) we develop a JavaScript library for the mobile web browser to execute and accelerate inference of the binary branch, which also provides a collaborative mechanism between the mobile web browser and the edge server. We have conducted extensive experiments using several well-known networks and datasets. The experimental results have shown that the proposed system outperforms the existing approaches in terms of reducing the model size by about 16x to 29x, and it also reduces end-to-end latency and outpaces the existing state-of-the-art approaches by over 3x to 60x when applying it in practical Web AR cases.
Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001
ICDCS2
2019 PDMR: priority-based dynamic multi-path routing algorithm for a software defined network
abstract
The ever changing traffic and Quality‐of‐Service (QoS) requirements of different traffic classes in multimedia applications impose tremendous challenges to the routing algorithm design. To address these challenges, a novel priority‐based dynamic multi‐path routing algorithm (PDMR) over Software Defined Network (SDN) is proposed in this paper. PDMR comprehensively considers the priorities of multimedia streams and the real‐time status of the link. The goal of PDMR is to determine what resources to be allocated to a set of flows keeping their priority, improving the utilization of network resources for satisfying stringent QoS requirements of the traffic flow. According to different QoS requirements in multimedia applications, traffic is classified into two classes and allocated their respective priorities. In addition, we derive a classification of the link cost function by considering the real‐time parameters such as link bandwidth, link delay, and link load to allocate the resources flexibly for transmission and to reduce the influence of between traffic classes. Furthermore, our proposed communication model between application services and the controller, which can guarantee the reliable transmission of the designated traffic. By conducting extensive simulations, we demonstrate that our proposed PDMR algorithm significantly outperforms other existing non‐priority‐based schemes in improving network performance.
Xiuquan Qiao, Junliang Chen 0001
IET Commun.2
2019 Session persistence for dynamic web applications in Named Data Networking
Xiuquan Qiao, Pei Ren, Junliang Chen 0001, Wei Tan 0001, M. Brian Blake, Wangli Xu
J. Netw. Comput. Appl.1
2019 Web AR: A Promising Future for Mobile Augmented Reality - State of the Art, Challenges, and Insights
abstract
Mobile augmented reality (Mobile AR) is gaining increasing attention from both academia and industry. Hardware-based Mobile AR and App-based Mobile AR are the two dominant platforms for Mobile AR applications. However, hardware-based Mobile AR implementation is known to be costly and lacks flexibility, while the App-based one requires additional downloading and installation in advance and is inconvenient for cross-platform deployment. In comparison, Web-based AR (Web AR) implementation can provide a pervasive Mobile AR experience to users thanks to the many successful deployments of the Web as a lightweight and cross-platform service provisioning platform. Furthermore, the emergence of 5G mobile communication networks has the potential to enhance the communication efficiency of Mobile AR dense computing in the Web-based approach. We conjecture that Web AR will deliver an innovative technology to enrich our ways of interacting with the physical (and cyber) world around us. This paper reviews the state-of-the-art technology and existing implementations of Mobile AR, as well as enabling technologies and challenges when AR meets the Web. Furthermore, we elaborate on the different potential Web AR provisioning approaches, especially the adaptive and scalable collaborative distributed solution which adopts the osmotic computing paradigm to provide Web AR services. We conclude this paper with the discussions of open challenges and research directions under current 3G/4G networks and the future 5G networks. We hope that this paper will help researchers and developers to gain a better understanding of the state of the research and development in Web AR and at the same time stimulate more research interest and effort on delivering life-enriching Web AR experiences to the fast-growing mobile and wireless business and consumer industry of the 21st century.
Xiuquan Qiao, Pei Ren, Schahram Dustdar, Ling Liu 0001, Huadong Ma, Junliang Chen 0001
Proc. IEEE1
2018 The Frame Latency of Personalized Livestreaming Can Be Significantly Slowed Down by WiFi
abstract
The popular personalized livestreaming (PL) in China, arguably the largest PL market in the world, is more monetized than PL in US and hence demands much lower interactive latencies to ensure a good quality of user experience. However, our pilot experiment shows that the video frame latency, dominant component of PL's interactive latency, can be significantly slowed down by WiFi, the primary Internet access method for PL. Understanding and further improving the frame latency over WiFi, however, have difficulties in 1) measuring end-to-end latency; 2) parsing encrypted PL's traffic and 3) modeling complex relationships between WiFi radio factors and the latency. To tackle these challenges, we design and prototype Latency Doctor (LTDr), a practical system which aims to model and optimize PL's video frame latency over WiFi. We deploy LTDr in our campus and obtain several key observations based on 13.9M video frames extracted from 12K individual views on three leading PLs in China. We observe that 40% frame latencies over WiFi hop are more than 30ms, and channel utilization should be less than 64% for low latency. Then we build a predictive model based on the dataset using the machine learning methodologies. Two real cases show that the median frame latencies are decreased by LTDr from 130ms to 22ms, and 50ms to 12ms respectively over WiFi networks.
Guoshun Nan, Xiuquan Qiao, Jiting Wang, Zeyan Li 0001, Jiahao Bu, Changhua Pei, Mengyu Zhou, Dan Pei
IPCCC2
2015 Design and Implementation: the Native Web Browser and Server for Content-Centric Networking
abstract
Content-Centric Networking (CCN) has recently emerged as a clean-slate Future Internet architecture which has a completely different communication pattern compared with exiting IP network. Since the World Wide Web has become one of the most popular and important applications on the Internet, how to effectively support the dominant browser and server based web applications is a key to the success of CCN. However, the existing web browsers and servers are mainly designed for the HTTP protocol over TCP/IP networks and cannot directly support CCN-based web applications. Existing research mainly focuses on plug-in or proxy/gateway approaches at client and server sides, and these schemes seriously impact the service performance due to multiple protocol conversions. To address above problems, we designed and implemented a CCN web browser and a CCN web server to natively support CCN protocol. To facilitate the smooth evolution from IP networks to CCN, CCNBrowser and CCNxTomcat also support the HTTP protocol besides the CCN. Experimental results show that CCNBrowser and CCNxTomcat outperform existing implementations. Finally, a real CCN-based web application is deployed on a CCN experimental testbed, which validates the applicability of CCNBrowser and CCNxTomcat.
Guoshun Nan, Xiuquan Qiao, Yukai Tu, Wei Tan 0001, Junliang Chen 0001
SIGCOMM2
2015 NDNBrowser: An extended web browser for named data networking
Xiuquan Qiao, Guoshun Nan, Yunlei Sun, Junliang Chen 0001
J. Netw. Comput. Appl.1
2015 A low-latency scheduling approach for high-definition video streaming in a heterogeneous wireless network with multihomed clients
Jiyan Wu, Xiuquan Qiao, Yamei Xia, Chau Yuen, Junliang Chen 0001
Multim. Syst.2
2015 Robust bandwidth aggregation for real-time video delivery in integrated heterogeneous wireless networks
Jiyan Wu, Yanlei Shang, Xiuquan Qiao, Bo Cheng 0001, Junliang Chen 0001
Multim. Tools Appl.3
2015 Recommending Nearby Strangers Instantly Based on Similar Check-In Behaviors
abstract
Chatting with nearby interested strangers instantly in location-based mobile social network (LMSN) has become increasingly popular. Currently, friend recommendation relies only on the simple and limited user profiles, and is agnostic to users' offline behaviors in the real world. For the first time, we focus on utilizing the user's check-in behaviors in the real world, instead of the general acquaintance-based social circles, to instantly recommend nearby strangers to make friends. However, bridging nearby strangers with similar check-in behaviors instantly has some new characteristics, such as lack of common friends and interaction histories, temporal, spatial and user three-dimensional correlation, and sparseness of check-ins. Most existing work about friend recommendations mainly focuses on making friends within the acquaintance-based social circles, and has not fully considered these new characteristics mentioned above. Therefore, how to catch the ephemeral opportunity to recommend nearby interested strangers instantly remains a challenge. In this paper, we present to use “Encounter” probability to measure the behavior similarity of two strangers in the real world based on their check-in histories. To address the sparseness challenge of check-in data, a Kernel Density Estimation (KDE)-based user check-in probability estimation method considering the spatiotemporal dimensions is proposed to estimate each user's check-in probability distribution with time at each spot. Finally, we use a large-scale user check-in dataset of Gowalla to validate the effectiveness of this approach. The experimental results show that our approach outperforms other commonly used similarity computation methods.
Xiuquan Qiao, Wei Tan 0001, Jianchong Su, Wangli Xu, Junliang Chen 0001
IEEE Trans Autom. Sci. Eng.1
2014 CCNxTomcat: An extended web server for Content-Centric Networking
Xiuquan Qiao, Guoshun Nan, Wei Tan 0001, Junliang Chen 0001, Yukai Tu
Comput. Networks1
2013 A Low-Delay, Lightweight Publish/Subscribe Architecture for Delay-Sensitive IOT Services
abstract
In order to build a low-latency lightweight publish/subscribe (pub/sub) system for IOT services, we propose an efficient and scalable broker architecture, called Grid Quorum-based pub/sub system (GQPS). As a core component in the event-driven SOA framework for IOT services, this architecture organizes multiple pub/sub brokers into a quorum-based peer-to-peer topology for efficient topic searching. It also leverages a topic searching algorithm and a caching strategy to achieve a small and constant search latency. Lightweight RESTful interfaces make our GQPS more suitable for IOT services. Cost analysis and experiment study demonstrate that GQPS achieves a significant performance gain in search satisfaction without compromising search cost. We applied GQPS in the District Heating Control and Information Service System in Beijing, China, which validates the feasibility and availability of our architecture.
Yunlei Sun, Xiuquan Qiao, Bo Cheng 0001, Junliang Chen 0001
ICWS2
2012 RESTful Web Service Mashup Based Coal Mine Safety Monitoring and Control Automation with Wireless Sensor Network
abstract
Due to complex environment of the coal mine, it's necessary to monitor the information of underground environment, device and miner instantly in order to ensure the safety of coal mine production. However, the exiting coal mine can not meet the requirements of coverage without blind spots as it is developed by the wired network. This paper proposes a RESTful Web services mashup augmented coal mine safety monitoring and control automation using ZigBee wireless sensor network, which can collect the underground temperature, humidity methane values and personal position through sensor nodes in the coal mine, and also collects the personnel position information inside the mine, and then implement a RESTful Application Programming Interface (API) on sensor nodes to provide access to sensors and actuators, allowing for them to be easily combined with other enterprise information resources based on the success of mashup applications. We also illustrated three different of scenarios for RESTful Web service mashups representing for coal mine safety monitoring and control automation. Finally, we give the conclusions.
Bo Cheng 0001, Xiuquan Qiao, Budan Wu, Xiaokun Wu 0002, Ruisheng Shi, Junliang Chen 0001
ICWS2
2008 A Semantic Description Approach for Telecommunications Network Capability Services
Xiuquan Qiao, Tian You
APNOMS1