VLDB 2026 Research / reviewers in the wild / expert
Yakun Huang
dblp:156/1030
· DBLP profile ↗
32ranked-venue papers
13as first author
30since 2021 · last 2026
0000-0003-4051-0200ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 10 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ShadowLLM: Resource-Efficient Hot Standby for Heterogeneous Edge LLM Serving
Zhenguo Chen, Pujun Ding, Yuanwei Zhu, Jing Lv, Yakun Huang, Xiuquan Qiao |
ICDCS | 6 |
| 2026 | Intent-Driven Cognitive XR Networks: Multi-Agent Orchestration for Immersive CommunicationabstractExtended Reality (XR) is emerging as a central use case in the 6G era, requiring intelligent, adaptive, and low-latency communication to support immersive user experiences. However, existing network architectures remain reactive and lack the capability to interpret or act upon users’ fine-grained multimodal intents, limiting their responsiveness and efficiency. To address this gap, this paper presents theCognitive XR Network (CXN)framework, an intent-driven architecture that integrates perception, reasoning, and control through a hierarchy of Artificial Intelligence (AI) agents. CXN comprises three cooperating agents: the User Intent Agent, which infers and predicts user intents from multimodal sensory inputs; the Network Orchestration Agent, which performs global coordination through multi-agent reinforcement learning; and distributed Resource Management Agents which execute localized decisions under strategic guidance. Together, these agents form a closed cognitive loop that enables proactive, intent-aware orchestration across radio, compute, and caching domains. Simulations under realistic XR collaboration scenarios demonstrate that CXN sustains an intent satisfaction rate above 85% and an average interaction latency below 60 ms with 90 concurrent users, outperforming both reactive and centralized learning baselines. Yakun Huang, Yaru Zhao 0001, Zhenguo Chen, Jing Lv, Xiuquan Qiao |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | PortaCap: Portable Volumetric Video Capturing System for Metaverse InteractionabstractPortable volumetric video capturing systems present a compelling alternative to traditional, bulky prototype systems used for streaming and interacting with volumetric content. Their key advantages, particularly flexibility and ease of deployment, make them suitable for a wide range of applications. However, despite their potential, there has been limited exploration into the design of such portable systems. This paper addresses this gap by conducting an in-depth analysis of portability and proposing an optimized camera array configuration tailored for high-quality volumetric content generation. Our approach begins with a novel, flexible camera calibration method that leverages geometric priors, enabling accurate alignment of multiple cameras without requiring specialized expertise. Building on this, we introduce a meticulous fusion technique that integrates captured and inferred data to reconstruct complete volumetric representations. This method achieves a fusion latency of less than 100 ms, ensuring real-time performance. We integrate these innovations into a portable capturing system, named PortaCap, which incorporates a carefully designed camera deployment strategy. Through both quantitative and qualitative evaluations, PortaCap demonstrates significant improvements in volumetric content quality, achieving enhancements ranging from 13% to 33.8%. These results underscore the system's potential to advance the state-of-the-art in portable volumetric video capture. Chongli Zhang, Yakun Huang, Yuanwei Zhu, Dexing Cai, Shibo Fang, Chunsheng Wang, Xiuquan Qiao |
IEEE Trans. Mob. Comput. | 2 |
| 2026 | EcoPath: Energy-Efficient Multi-Path Data Aggregation for Ubiquitous Connectivity ServicesabstractUbiquitous connectivity is a key 6G usage scenario, in which large-scale sensing systems deployed in remote and underserved regions must deliver heterogeneous sensing data under stringent energy budgets and deadline constraints. This paper presents EcoPath, a two-tier data aggregation framework for clustered large-scale sensor networks. EcoPath separates low-power intra-cluster collection from a high-rate multi-interface backhaul operated by cluster heads, where Multipath QUIC (MPQUIC) can be practically deployed to exploit path diversity. At the cluster head, EcoPath jointly integrates (i) a deadline-aware bundling controller that aggregates sensor frames into MTU-bounded bundles to amortize protocol overhead while bounding additional waiting time, and (ii) a robust multi-path scheduler that prioritizes packets using Weighted Earliest- Deadline-First (W-EDF) with fairness protection and selects backhaul paths via a stability-aware quality metric with hysteresis to avoid flapping under time-varying links. We further formulate an explicit energy–timeliness optimization and show how its outputs parameterize the online bundling and scheduling policies. Extensive simulations with realistic wireless effects, together with baselines and ablations, demonstrate that EcoPath improves energy efficiency and deadline satisfaction for large-scale aggregation. Yaru Zhao 0001, Yuan-Ting Yan, Man He, Yuanwei Zhu, Yi Yue 0001, Yakun Huang |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2025 | P2S-XR: Predictive and Scalable Scheduling for Concurrent Multi-User XR Services
Yaru Zhao 0001, Shou-lu Hou, Binyang Li, Yakun Huang |
IEEE Big Data | 4 |
| 2025 | WebGS360: Towards web-based visualization of Gaussian Splatting from panoramic images
Chongli Zhang, Jing Lv, Xiuquan Qiao, Yakun Huang |
Comput. Graph. | 6 |
| 2025 | Two grids are better than one: Hybrid indoor scene reconstruction framework with adaptive priorsabstractIndoor scene reconstruction from multi-view images is a pivotal technology within the field of robotics and augmented reality . Previous researches have predominantly focused on neural radiance fields aided by geometric monocular priors. However, due to the inductive smoothness bias introduced by deep Multi-Layer Perceptron (MLP) networks, these methods struggle to recover the scene surface with complex and fine geometry details. Additionally, when used as additional supervision signals during optimization, priors in different regions make different contributions. Simply incorporating them in all regions may lead to a decrease in the accuracy. To tackle these issues, we present a generic end-to-end framework named AdaptSurf, which combines Signed Distance Field (SDF) voxel grids and feature voxel grids to enhance the capability of reconstructing accurate geometry details, respectively. Furthermore, we design a policy network to adaptively enable the estimated depth or normal priors to supervise the learning process, which improves the reconstruction accuracy and accelerates neural surface reconstruction. Qualitative and quantitative experiments show that AdaptSurf yields high-quality surfaces, especially for fine-grained details and smooth regions. Furthermore, the policy network exhibits an interpretable behavior that depends on the voxel features, which helps to improve the quality of surface reconstruction. Boyuan Bai, Xiuquan Qiao, Hongru Zhao, Wenzhe Shi, Hengjia Zhang, Yakun Huang |
Neurocomputing | 7 |
| 2025 | WebARNav: Mobile Web AR Indoor Navigation With Edge-Assisted Vision LocalizationabstractThe gradual maturation of mobile augmented reality (AR) and localization technologies is enabling the development of immersive AR-enabled indoor localization and navigation systems. Existing indoor localization technologies (e.g., WiFi, infrared, Bluetooth) and navigation services do not provide intuitive 3D AR experiences and can be expensive to deploy. This paper introduces WebARNav, a cross-platform indoor localization system that provides user-friendly AR navigation services with low overhead and remarkable accuracy. First, we propose a lightweight location fusion framework for indoor navigation on the mobile web, which leverages accurate edge-supported vision localization to guide and correct lightweight pedestrian dead reckoning localization. Second, we improve the accuracy of localization using an attention-based feature extraction method and a dual-stream retrieval and co-visibility re-ranking technique for initial localization. Third, we significantly improve accuracy and speed up retrieval as users move by generating a topological map for traveling localization. We conducted extensive experiments on various indoor datasets to demonstrate localization accuracy and navigation experience. The study shows that WebARNav achieves a localization frequency of over 30 Hz and reduces the average trajectory error by 76% and 95% for single- and multi-floor office scenes, respectively, compared to the PDR-only method. The proposed traveling localization method also reduces the localization latency by 15.2%, 55.1%, and 98.6% in the baseline datasets, with an accuracy improvement of over 4%. Yakun Huang, Shengwei Meng, Yuanwei Zhu, Jacky Cao, Xiuquan Qiao, Xiang Su 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | FPSelector: A Flexible Path Selector for Mobile Augmented Reality OffloadingabstractMobile Augmented Reality (MAR) applications pose unique challenges due to computation intensity, constrained device resources, and high interactive rendering requirements. The emergence of 5 G and edge computing offers opportunities to offload computation to the edge and cloud, indirectly enhancing the computing capability and usage duration of MAR devices. However, existing general task offloading and multipath transmission techniques do not address the challenges in offloading path selection with multiple edges, dynamic resource competition awareness, and spatial computation with strong task dependencies. This paper contributes FPSelector, a flexible path selector for MAR offloading. We present a two-tier MAR-specific offloading scheme with multiple edge nodes. In offloading decisions, we design a reinforcement learning model to generate the selection policy for each packet of an AR data stream. This model incorporates an action masking mechanism, a comprehensive reward function, and state features complemented by a resource prediction module, making FPSelector aware of dynamic heterogeneous environments. Moreover, we propose an online learning strategy to facilitate real-time selection. To validate its efficacy, we compare FPSelector's performance against leading schedulers under various scenarios, demonstrating a notable reduction of 9.9% and 9.6% in overall completion time for 4 K and 8 K video-based MAR applications compared to its closest competitor. Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Xiaoli Liu 0005, Xiang Su 0001, Anna Brunström, Özgü Alay, Sasu Tarkoma |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | SemDA: Communication-Efficient Data Aggregation Through Distributed Semantic TransmissionabstractThis paper introduces SemDA, a communication-efficient data aggregation method that uses distributed semantic communication for improved transmission and analysis. SemDA utilizes an end-to-end trainable network structure that reduces data transmission volume and deepens semantic feature aggregation. Key advances include an attention-based aggregation method for holistic semantic feature integration and a dual-attention decoding network that emphasizes viewpoint and content dimensions. Performance evaluations on CIFAR-10 and ImageNet datasets show that SemDA offers significant improvements in accuracy and system overhead compared to traditional and distributed semantic communication methods. Notable contributions include the proposal of a novel decoding structure, the introduction of a dual-attention decoding mechanism, and extensive evaluations against benchmark methods. Yaru Zhao 0001, Yakun Huang |
ICASSP | 2 |
| 2024 | SM3: Self-supervised Multi-task Modeling with Multi-view 2D Images for Articulated ObjectsabstractReconstructing real-world objects and estimating their movable joint structures are pivotal technologies within the field of robotics. Previous research has predominantly focused on supervised approaches, relying on annotated datasets to model articulated objects within limited categories. However, these approaches fall short of effectively addressing the diversity present in the real world. To tackle this issue, we propose a self-supervised interaction perception method, referred to as SM3, which leverages multi-view RGB images captured before and after interaction to model articulated objects, identify the movable parts, and infer the parameters of their rotating joints. By constructing 3D geometries and textures from the captured 2D images, SM3achieves integrated optimization of movable part and joint parameters during the reconstruction process, obviating the need for annotations. Furthermore, we introduce the MMArt dataset, an extension of PartNet-Mobility, encompassing multi-view and multi-modal data of articulated objects spanning diverse categories. Evaluations demonstrate that SM3surpasses existing benchmarks across various categories and objects, and its adaptability in real-world scenarios has been thoroughly validated. Haowen Wang 0001, Zhengping Che, Yakun Huang, Xiuquan Qiao, Jian Tang 0008 |
ICRA | 6 |
| 2024 | KiProL: A Knowledge-Injected Prompt Learning Framework for Language Generation
Yaru Zhao 0001, Yakun Huang, Bo Cheng 0001 |
PAKDD (6) | 2 |
| 2024 | Distributed realtime rendering in decentralized network for mobile web augmented reality
Huabing Zhang, Liang Li 0023, Qiong Lu, Yi Yue 0001, Yakun Huang, Schahram Dustdar |
Future Gener. Comput. Syst. | 5 |
| 2024 | ISCom: Interest-Aware Semantic Communication Scheme for Point Cloud Video Streaming on Metaverse XR DevicesabstractIn the metaverse era, point cloud video (PCV) streaming on mobile XR devices is pivotal. While most current methods focus on PCV compression from traditional 3-DoF video services, emerging AI techniques extract vital semantic information, producing content resembling the original. However, these are early-stage and computationally intensive. To enhance the inference efficacy of AI-based approaches, accommodate dynamic environments, and facilitate applicability to metaverse XR devices, we present ISCom, an interest-aware semantic communication scheme for lightweight PCV streaming. ISCom is featured with a region-of-interest (ROI) selection module, a lightweight encoder-decoder training module, and a learning-based scheduler to achieve real-time PCV decoding and rendering on resource-constrained devices. ISCom’s dual-stage ROI selection provides significantly reduces data volume according to real-time interest. The lightweight PCV encoder-decoder training is tailored to resource-constrained devices and adapts to the heterogeneous computing capabilities of devices. Furthermore, We provide a deep reinforcement learning (DRL)-based scheduler to select optimal encoder-decoder model for various devices adaptivelly, considering the dynamic network environments and device computing capabilities. Our extensive experiments demonstrate that ISCom outperforms baselines on mobile devices, achieving a minimum rendering frame rate improvement of 10 FPS and up to 22 FPS. Furthermore, our method significantly reduces memory usage by 41.7% compared to the state-of-the-art AITransfer method. These results highlight the effectiveness of ISCom in enabling lightweight PCV streaming and its potential to improve immersive experiences for emerging metaverse application. Yakun Huang, Boyuan Bai, Yuanwei Zhu, Xiuquan Qiao, Xiang Su 0001, Lei Yang 0063, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | FluGCF: A Fluent Dialogue Generation Model With Coherent Concept Entity FlowabstractThe integration of external knowledge graphs into dialogue systems effectively mitigates the generation of generic and uninteresting responses. This approach, particularly the explicit modeling of conversation flows from related concept entities, facilitates the generation of semantically rich and informative responses. However, recent models guided by concept entity flows present two primary limitations: (1) a limited semantic understanding of the post message, which complicates the selection of highly relevant 1-hop concept entities, and (2) an inability to extract dynamic and diverse semantic relations between the post message and 2-hop concept entities. To address these issues, we introduce FluGCF, a novel model that fluently generates dialogues with coherent guidance from concept entity flows. FluGCF employs a ternary fusion to explicitly model multi-hop concept entity flows using a post-aware knowledge encoding mechanism. This mechanism learns semantic concept entity features from both word and sentence-level text features. Additionally, we design a corresponding ternary decoding mechanism that dynamically selects concept entities or words from the vocabulary to enhance fluency and diversity in dialogue generation. FluGCF, implemented in PyTorch, was extensively evaluated on a large-scale dataset, revealing that it surpasses baseline models, including the state-of-the-art knowledge-aware model ConceptFlow, by nearly 15% in terms of fluency. Furthermore, it demonstrated notable enhancements in coherence, diversity and informativeness. Yaru Zhao 0001, Bo Cheng 0001, Yakun Huang, Zhiguo Wan |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | HiVAT: Improving QoE for Hybrid Video Streaming Service With Adaptive TranscodingabstractMobile video streaming enables flexible delivery of videos to mobile devices, supporting emerging video formats. The transition from conventional 2D videos to immersive formats, such as virtual reality and holographic videos, significantly increases the demand for computation and network resources. Existing streaming techniques are predominantly developed for specific video types, neglecting fair adaptive transmission and optimal resource utilization in services involving multiple video types. This paper investigates hybrid video streaming, encompassing 2D, 360-degree, and volumetric videos. To accommodate resource-intensive hybrid video streaming on mobile devices, we proposeHiVAT, an adaptive transcoding-based system that ensures Quality of Experience (QoE) for each stream type. We contribute 1) a transcoding-based framework to address the challenges of high bandwidth and decoding overhead on mobile devices; 2) a universal QoE model involving traditional factors, viewport smoothness, degree of immersion, etc., for transcoded video streams; 3) a multi-agent adaptive bitrate controller that collaboratively determines hybrid video quality levels to achieve high and fair QoE across multiple streams; and 4) a learning-based task scheduler to optimize computation resource usage, thereby improving the overall serviceability of the system. We evaluateHiVATagainst state-of-the-art methods, witnessing an average QoE improvement of 5.9% and 9.9% on linear and logarithmic metrics, respectively. Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Xiang Su 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2023 | Cloud-Edge-Device Collaborative Image Retrieval and Recognition for Mobile Web
Yakun Huang, Shouyi Wu, Xiuquan Qiao, Hongshun He |
CollaborateCom (2) | 1 |
| 2023 | DTF-Net: Category-Level Pose Estimation and Shape Reconstruction via Deformable Template FieldabstractEstimating 6D poses and reconstructing 3D shapes of objects in open-world scenes from RGB-depth image pairs is challenging. Many existing methods rely on learning geometric features that correspond to specific templates while disregarding shape variations and pose differences among objects in the same category. As a result, these methods underperform when handling unseen object instances in complex environments. In contrast, other approaches aim to achieve category-level estimation and reconstruction by leveraging normalized geometric structure priors, but the static prior-based reconstruction struggles with substantial intra-class variations. To solve these problems, we propose the DTF-Net, a novel framework for pose estimation and shape reconstruction based on implicit neural fields of object categories. In DTF-Net, we design a deformable template field to represent the general category-wise shape latent features and intra-category geometric deformation features. The field establishes continuous shape correspondences, deforming the category template into arbitrary observed instances to accomplish shape reconstruction. We introduce a pose regression module that shares the deformation features and template codes from the fields to estimate the accurate 6D pose of each object in the scene. We integrate a multi-modal representation extraction module to extract object features and semantic masks, enabling end-to-end inference. Moreover, during training, we implement a shape-invariant training strategy and a viewpoint sampling method to further enhance the model's capability to extract object pose features. Extensive experiments on the REAL275 and CAMERA25 datasets demonstrate the superiority of DTF-Net in both synthetic and real scenes. Furthermore, we show that DTF-Net effectively supports grasping tasks with a real robot arm. Haowen Wang 0001, Zhengping Che, Dong Liu 0058, Feifei Feng, Yakun Huang, Xiuquan Qiao, Jian Tang 0008 |
ACM Multimedia | 8 |
| 2023 | Beyond Words: An Intelligent Human-Machine Dialogue System with Multimodal Generation and Emotional ComprehensionabstractIntelligent service robots have become an indispensable aspect of modern‐day society, playing a crucial role in various domains ranging from healthcare to hospitality. Among these robotic systems, human‐machine dialogue systems are particularly noteworthy as they deliver both auditory and visual services to users, effectively bridging the communication gap between humans and machines. Despite their utility, the majority of existing approaches to these systems primarily concentrate on augmenting the logical coherence of the system’s responses, inadvertently neglecting the significance of user emotions in shaping a comprehensive communication experience. To tackle this shortcoming, we propose the development of an innovative human‐machine dialogue system that is both intelligent and emotionally sensitive, employing multimodal generation techniques. This system is architecturally comprised of three components: (1) data collection and processing, responsible for gathering and preparing relevant information, (2) a dialogue engine, which generates contextually appropriate responses, and (3) an interaction module, responsible for facilitating the communication interface between users and the system. To validate our proposed approach, we have constructed a prototype system and conducted an evaluation of the performance of the core dialogue engine by utilizing an open dataset. The results of our study indicate that our system demonstrates a remarkable level of multimodal generation response, ultimately offering a more human‐like dialogue experience. Yaru Zhao 0001, Bo Cheng 0001, Yakun Huang, Zhiguo Wan |
Int. J. Intell. Syst. | 3 |
| 2023 | An Integrated Cloud-Edge-Device Adaptive Deep Learning Service for Cross-Platform WebabstractDeep learning shows great promise in providing more intelligence to the cross-platform web. However, insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning with low-performing web browsers. We propose DeepAdapter, an integrated cloud-edge-device framework that ties the edge, the remote cloud, with the device by cross-platform web technology for adaptive deep learning services towards lower latency, lower mobile energy, and higher system throughput. DeepAdapter consists of context-aware pruning, service updating, and online scheduling. First, the offline pruning module provides a context-aware pruning algorithm that incorporates the latency, the network condition, and the device's computing capability to fit various contexts. Second, the service updating module optimizes branch model cache on the edge for massive mobile users and updates the new model pruning requirements. Third, the online scheduling module matches optimal branch models for mobile users. Also, a two-stage DRL-based online scheduling method named DeepScheduler can handle high concurrent requests between edge centers and remote cloud by designing the reward prediction model. Extensive experiments show that DeepAdapter can decrease average latency by 1.33x, reduce average mobile energy by 1.4x, and improve system throughput by 2.1x with considerable accuracy. Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | A Semantic-Aware Transmission With Adaptive Control Scheme for Volumetric Video ServiceabstractVolumetric video provides a more immersive holographic virtual experience than conventional video services such as 360-degree and virtual reality (VR) videos. However, due to ultra-high bandwidth requirements, existing compression and transmission technology cannot handle the delivery of real-time volumetric video. Unlike traditional compression methods and the approaches that extend 360-degree video streaming, we propose AITransfer, an AI-powered compression and semantic-aware transmission method for point cloud video data (a popular volumetric data format). AITransfer targets the semantic-level communication beyond transmitting raw point cloud video or compressed video with two outstanding contributions: (1) designing an integrated end-to-end architecture with two fundamental contents of feature extraction and reconstruction to reduce the bandwidth consumption and alleviate the computational pressure; and (2) incorporating the dynamic network condition into end-to-end architecture design and employing a deep reinforcement learning-based adaptive control scheme to provide robust transmission. We conduct extensive experiments on the typical datasets and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide extremely efficient point cloud transmission while maintaining considerable user experience with more than 30.72x compression ratio under the existing network environments. Yuanwei Zhu, Yakun Huang, Xiuquan Qiao, Zhijie Tan, Boyuan Bai, Huadong Ma, Schahram Dustdar |
IEEE Trans. Multim. | 2 |
| 2022 | AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile WebabstractEmploying today’s deep neural network (DNN) into the cross-platform web with an offloading way has been a promising means to alleviate the tension between intensive inference and limited computing resources. However, it is still challenging to directly leverage the distributed DNN execution into web apps with the following limitations, including (1) how special computing tasks such as DNN inference can provide fine-grained and efficient offloading in the inefficient JavaScript-based environment? (2) lacking the ability to balance the latency and mobile energy to partition the inference facing various web applications’ requirements. (3) and ignoring that DNN inference is vulnerable to the operating environment and mobile devices’ computing capability, especially dedicated web apps. This paper designs AoDNN, an automatic offloading framework to orchestrate the DNN inference across the mobile web and the edge server, with three main contributions. First, we design the DNN offloading based on providing a snapshot mechanism and use multi-threads to monitor dynamic contexts, partition decision, trigger offloading, etc. Second, we provide a learning-based latency and mobile energy prediction framework for supporting various web browsers and platforms. Third, we establish a multi-objective optimization to solve the optimal partition by balancing the latency and mobile energy. Yakun Huang, Xiuquan Qiao, Schahram Dustdar |
INFOCOM | 1 |
| 2022 | Enabling DNN Acceleration With Data and Model Parallelization Over Ubiquitous End DevicesabstractDeep neural network (DNN) shows great promise in providing more intelligence to ubiquitous end devices. However, the existing partition-offloading schemes adopt data-parallel or model-parallel collaboration between devices and the cloud, which does not make full use of the resources of end devices for deep-level parallel execution. This article proposes eDDNN (i.e., enabling Distributed DNN), a collaborative inference scheme over heterogeneous end devices using cross-platform Web technology, moving the computation close to ubiquitous end devices, improving resource utilization, and reducing the computing pressure of data centers. eDDNN implements D2D communication and collaborative inference among heterogeneous end devices with WebRTC protocol, divides the data and corresponding DNN model into pieces simultaneously, and then executes inference almost independently by establishing a layer dependency table. Besides, eDDNN provides a dynamic allocation algorithm based on deep reinforcement learning to minimize latency. We conduct experiments on various data sets and DNNs and further employ eDDNN into a mobile Web AR application to illustrate the effectiveness. The results show that eDDNN can achieve the latency decrease by$2.98\times $, reduce mobile energy by$1.8\times $, and relieve the computing pressure of the edge server by$2.57\times $, against a typical partition-offloading approach. Yakun Huang, Xiuquan Qiao, Wenhai Lai, Schahram Dustdar, Jiulin Li |
IEEE Internet Things J. | 1 |
| 2022 | A Collaborative Task Offloading Framework for Smart TV Applications in a Household Computing EnvironmentabstractSmart TV can perform interactive computing while also providing video content services. However, this leads to a high delay during interactive computing because of the lack of computing capability, thus smart TVs are unable to undertake large scenes and complex interactive computing tasks. This article proposes a collaborative task offloading framework (CTOF) for interactive computing of smart TV video applications. The main contributions of this article are as follows: 1) a computing offloading mechanism is proposed for interactive computing with video content. A part of the interactive computing task is offloaded to the user’s high-computing mobile device in a household video service environment via a Wi-Fi Direct channel and 2) a collaborative computing offloading algorithm is proposed for complex interactive computing tasks. According to the computing complexity and the computing expansion coefficient, a parallel and serial collaborative smart offloading framework is established to minimize the delay. We conduct extensive experiments to indicate that in the existing experimental network environment, the video-based complex interactive service operation efficiency is improved by 20%. With a smart offloading framework, we can further achieve a satisfactory experience in terms of the interactive computing delay. Consequently, the smart TV can quickly respond to the complex video interactive service and improve the business interaction capability of the smart TV. Liang Li 0023, Xiuquan Qiao, Huabing Zhang, Yakun Huang, Pei Ren |
IEEE Internet Things J. | 4 |
| 2022 | Edge AR X5: An Edge-Assisted Multi-User Collaborative Framework for Mobile Web Augmented Reality in 5G and BeyondabstractMulti-user mobile Augmented Reality (AR) has been successfully used in various fields as a novel visual interaction technology. But current mainstream wearable device-based and app-based solutions are still facing cross-platform, real-time communication, and intensive computing requirements. Mobile Web technology is envisioned to be a promising supporting technology for cross-platform application of mobile AR especially in 5G networks, which provide pervasive communication and computing resources thereby forming a formidable framework for the practical application of multi-user mobile Web AR. However, the problem of how to use these new techniques properly to achieve efficient communication and computing collaboration is obviously paramount in order for multi-user mobile Web AR to be realized in 5G networks. In this article, we propose the first edge-assisted multi-user collaborative framework for mobile Web AR in the 5G era. First, we propose a heuristic mechanism BA-CPP for efficient communication planning, which allows multi-user interaction synchronization to be achieved. Second, we introduce a motion-aware key frame selection mechanism called Mo-KFP to optimize the computational efficiency of the edge system, and simultaneously alleviate the initialization problem by collaborating with nearby mobile devices using the Device-to-Device (D2D) communication technique. Experiments are conducted in a real-world 5G network, and the results demonstrate the superiority of our proposed collaborative framework. Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2022 | A Lightweight Collaborative Deep Neural Network for the Mobile Web in Edge CloudabstractEnabling deep learning technology on the mobile web can improve the user’s experience for achieving web artificial intelligence in various fields. However, heavy DNN models and limited computing resources of the mobile web are now unable to support executing computationally intensive DNNs when deploying in a cloud computing platform. With the help of promising edge computing, we propose a lightweight collaborative deep neural network for the mobile web, named LcDNN, which contributes to three aspects: (1) We design a composite collaborative DNN that reduces the model size, accelerates inference, and reduces mobile energy cost by executing a lightweight binary neural network (BNN) branch on the mobile web. (2) We provide a jointly training method for LcDNN and implement an energy-efficient inference library for executing the BNN branch on the mobile web. (3) To further promote the resource utilization of the edge cloud, we develop a DRL-based online scheduling scheme to obtain an optimal allocation for LcDNN. The experimental results show that LcDNN outperforms existing approaches for reducing the model size by about 16x to 29x. It also reduces the end-to-end latency and mobile energy cost with acceptable accuracy and improves the throughput and resource utilization of the edge cloud. Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Schahram Dustdar, Junliang Chen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | Fine-Grained Elastic Partitioning for Distributed DNN Towards Mobile Web AR Services in the 5G EraabstractWeb-based Deep Neural Networks (DNNs) enhance the ability of object recognition and has attracted considerable attention in mobile Web AR and other services. However, neither performing the DNN inference on mobile Web browsers locally nor offloading computations to the cloud can strike a balance between accuracy and efficiency; generally, rude methods are often accompanied by unsatisfactory accuracy. Collaborative approaches seem to fill this gap by coordinating the distributed hierarchical computing resources, especially in the 5G era, but it still faces challenges in the current solutions, such as the lack of (1) full use of 5G resources for the one point DNN computation partitioning schemes; (2) fine-grained branching mechanism; (3) efficient partitioning method; and (4) multi-objective optimization. To this end, we present the fine-grained elastic computation partitioning mechanism for distributed DNN in 5G networks. First, we elaborate two collaborative scenarios. Second, we study the DNN branching mechanism at layer granularity. Next, we propose a DNN computation partitioning algorithm based on deep reinforcement learning. Finally, we develop a mobile Web AR application as a proof of concept. The experiments were conducted in an actually deployed 5G trial network, and the results show the superiority of this collaborative approach. The common theme is, under the premise that Quality of Service (QoS) is satisfied, to balance multiple interests by orchestrating computations across heterogeneous computing platforms. Pei Ren, Xiuquan Qiao, Yakun Huang, Ling Liu 0001, Calton Pu, Schahram Dustdar |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Towards Video Streaming Analysis and Sharing for Multi-Device Interaction with Lightweight DNNsabstractMulti-device interaction has attracted a growing interest in both mobile communication industry and mobile computing research community as mobile devices enabled social media and social networking continue to blossom. However, due to the stringent low latency requirements and the complexity and intensity of computation, implementing efficient multi-device interaction for real-time video streaming analysis and sharing is still in its infancy. Unlike previous approaches that rely on high network bandwidth and high availability of cloud center with GPUs to support intensive computations for multi-device interaction and for improving the service experience, we propose MIRSA, a novel edge centric multi-device interaction framework with a lightweight end-to-end DNN for on-device visual odometry (VO) streaming analysis by leveraging edge computing optimizations with three main contributions. First, we design MIRSA to migrate computations from the cloud to the device side, reducing the high overhead for large transmission of video streaming while alleviating the server load of the cloud. Second, we design a lightweight VO network by utilizing temporal shift module to support on-device pose estimation. Third, we provide on-device resource-aware scheduling algorithm to optimize the task allocation. Extensive experiments show MIRSA provides real-time high quality pose estimation as an interactive service and outperforms baseline methods. Yakun Huang, Hongru Zhao, Xiuquan Qiao, Jian Tang 0008, Ling Liu 0001 |
INFOCOM | 1 |
| 2021 | AITransfer: Progressive AI-powered Transmission for Real-Time Point Cloud Video StreamingabstractPoint cloud video provides a more immersive holographic virtual experience than conventional video services such as 360 degree video and virtual reality (VR) video. However, the existing network bandwidth and transmission technology can not carry real-time point cloud video streaming due to mass data volume, high processing overheads, and extremely bandwidth-consuming. Unlike previous approaches that extend the VR video streaming, we propose AITransfer, an AI-powered bandwidth-aware and adaptive transmission technique driven by extracting and transferring key point cloud features to reduce the bandwidth consumption and alleviate the computational pressure. AITransfer has two outstanding contributions, including (1) incorporating the dynamic network bandwidth into the design of an end-to-end architecture with two fundamental contents of feature extraction and reconstruction, and (2) employing an online adapter to sense the network bandwidth and match the optimal inference model. We conduct extensive experiments on the typical dataset and develop a case study to demonstrate the efficiency and effectiveness. The results show that AITransfer can provide more than 30.72 times compression ratio under the existing network environments. Yakun Huang, Yuanwei Zhu, Xiuquan Qiao, Zhijie Tan, Boyuan Bai |
ACM Multimedia | 1 |
| 2021 | EdgeBooster: Edge-Assisted Real-Time Image Segmentation for the Mobile Web in WoTabstractCombining image segmentation with Web technology lays a good foundation for lightweight, cross-platform, and pervasive Web artificial intelligence applications, and further improves the capability of Web-of-Things (WoT) applications. However, no matter whether we use a Web real-time communication media server for advanced processing that views camera inputs as a video stream, or transfer continuous camera frames to the remote cloud for processing, we are unable to obtain a satisfactory real-time experience due to high resource consumption and unacceptable latency. In this article, we present EdgeBooster, a computational-efficient architecture that leverages a common edge server to minimize the communication costs, accelerates the camera frame segmentation, and guarantees an acceptable segmentation accuracy with the prior knowledge. EdgeBooster provides real-time segmentation by developing parallel technology that enables segmentation on slices of a camera frame and using presegmentation based on superpixels to accelerate the graph-based segmentation. It also introduces recent DNN-based segmentation results as the prior knowledge to improve the performance of the graph-based segmentation, especially in nonideal scenes, such as dark light and weak contrast. Finally, it creates a pure frontend segmentation that can provide continuous and stable services for mobile users in unstable networks, such as a weak network or with an unstable edge server. The experimental results show that EdgeBooster is able to achieve a considerable accuracy for the mobile Web, running at no less than 30 frames per second in real scenes. Yakun Huang, Xiuquan Qiao, Pei Ren, Schahram Dustdar, Junliang Chen 0001 |
IEEE Internet Things J. | 1 |
| 2020 | DeepAdapter: A Collaborative Deep Learning Framework for the Mobile Web Using Context-Aware Network PruningabstractDeep learning shows great promise in providing more intelligence to the mobile web, but insufficient infrastructure, heavy models, and intensive computation limit the use of deep learning in mobile web applications. In this paper, we present DeepAdapter, a collaborative framework that ties the mobile web with an edge server and a remote cloud server to allow executing deep learning on the mobile web with lower processing latency, lower mobile energy, and higher system throughput. DeepAdapter provides a context-aware pruning algorithm that incorporates the latency, the network condition and the computing capability of the mobile device to fit the resource constraints of the mobile web better. It also provides a model cache update mechanism improving the model request hit rate for mobile web users. At runtime, it matches an appropriate model with the mobile web user and provides a collaborative mechanism to ensure accuracy. Our results show that DeepAdapter decreases average latency by 1.33x, reduces average mobile energy consumption by 1.4x, and improves system throughput by 2.1x with a considerable accuracy. Its contextaware pruning algorithm also improves inference accuracy by up to 0.3% with a smaller and faster model. Yakun Huang, Xiuquan Qiao, Jian Tang 0008, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001 |
INFOCOM | 1 |
| 2019 | A Lightweight Collaborative Recognition System with Binary Convolutional Neural Network for Mobile Web Augmented RealityabstractLightweight and precise recognition is a key component of web-based augmented reality (Web AR) applications. Although edge-based distributed deep learning approach is now possible to achieve satisfactory recognition for Web AR applications, it puts significant pressure on the computation and energy consumption of the mobile web browser, especially the app-based embedded browser. Thus, reducing the model size and accelerating the inference are regarded as the two fundamental challenges to enable this edge-based collaborative recognition system efficiently. In this paper, we propose a lightweight collaborative recognition system (LCRS) for Web AR applications. LCRS contributes to three aspects: (1) we design a composite deep neural network for reducing the model size and inference latency by introducing binary convolutional neural network; (2) we provide a joint training method to co-train the general branch and the binary branch; (3) we develop a JavaScript library for the mobile web browser to execute and accelerate inference of the binary branch, which also provides a collaborative mechanism between the mobile web browser and the edge server. We have conducted extensive experiments using several well-known networks and datasets. The experimental results have shown that the proposed system outperforms the existing approaches in terms of reducing the model size by about 16x to 29x, and it also reduces end-to-end latency and outpaces the existing state-of-the-art approaches by over 3x to 60x when applying it in practical Web AR cases. Yakun Huang, Xiuquan Qiao, Pei Ren, Ling Liu 0001, Calton Pu, Junliang Chen 0001 |
ICDCS | 1 |