Yongjie Guan

dblp:254/5967 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
14since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 9 · 2 first-author · 8 since 2021Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 SA-MGRU: Gate-aligned fusion of self-attention and multi-gate GRU for medical image segmentation
Yongjie Guan, Zexuan Ji
Expert Syst. Appl.1
2026 Falcon: A Mobile Semantic Visual Perception Framework Enlightened by Human Vision Systems
abstract
In the realm of computer human visual perception, semantic perception means recognizing objects, people, and scenes. This involves not just detecting shapes and colors but understanding what those visual elements represent. Despite its vital role in enhancing various aspects of emerging applications such as safety for autonomous driving and immersion for mixed reality (MR), real-time segmentation on mobile and edge platforms is challenging due to the nature of dense pixel labeling. To address this issue, we propose Falcon, a lightweight focus-aware segmentation framework that effectively integrates multiple innovations to achieve real-time segmentation on resource-constrained mobile and edge devices. We design a novel, low-dimension feature for efficient pixel labeling with shallow neural networks, an agile focus-aware refining scheme to compensate for the coarse nature of holistic segmentation, and a modularized design to accommodate the diversity of mobile and edge platforms and ensure seamless integration with different segmentation models. We build a prototype implementation of Falcon that supports both on-device executions and edge-assisted offloading, and asynchronous segmentation and refinement. We extensively evaluate the performance of Falcon for autonomous driving and MR applications with real setups and standard datasets. Our results demonstrate that Falcon achieves real-time segmentation, with an impressive rate of up to 40 frames per second.
Xueyu Hou, Yongjie Guan, Tao Han 0002
IEEE Trans. Mob. Comput.2
2025 Device-Server Collaborative Speculative Decoding for Real-Time LLM Streaming
Bishakha Rani Biswas, Yongjie Guan, Mingrui Yin, Tao Han 0002, Xueyu Hou
GLOBECOM2
2025 ForestRAG-AR: An AR and RAG Framework for Context-Aware Silviculture Assistance
abstract
Silvicultural decision making in forestry relies heavily on technical manuals and expert knowledge that are difficult to access in the field. Meanwhile, advances in wearable augmented reality (AR) and retrieval-augmented generation (RAG) offer new opportunities to deliver context-aware, data-driven guidance directly within natural environments. This paper presents ForestRAG-AR, an integrated AR and RAG framework that provides site-specific silvicultural recommendations grounded in authoritative forestry documents. The system combines (i) a perception and context extraction module that interprets local forest conditions, including tree species, stand density, and terrain slope, through on-device sensing, (ii) a RAG-based knowledge backend that adapts a large language model to forestry by embedding regional silviculture manuals and best-management-practice guides, and (iii) an AR interface that visualizes thinning and pruning guidance with natural-language explanations and verifiable citations. A domain-adapted cross-encoder reranker fuses environmental context with text passages to improve retrieval precision and factual grounding. Experiments on 50 forestry queries demonstrate that ForestRAG-AR improves nDCG@10 by 7.8 points and citation precision to 0.96 while maintaining sub-500 ms end-to-end latency.
Md Mehedi Hasan, Bishakha Rani Biswas, Xueyu Hou, Yongjie Guan
SEC4
2025 ImmersiveSlicing: An O-RAN Cross-Layer Reinforcement Learning Framework for Low-Latency Immersive Applications
abstract
The proliferation of immersive applications such as Virtual, Augmented, and Mixed Reality (VR/AR/MR) imposes stringent low-latency and reliability requirements that challenge conventional O-RAN slicing mechanisms. Existing frameworks often fail to anticipate rapid XR traffic fluctuations driven by user motion and gaze dynamics, leading to inefficient resource utilization and SLA violations. To overcome these limitations, we propose a cross-layer intelligent control framework that integrates traffic prediction and reinforcement learning-based slice orchestration across the Non-RT and Near-RT RIC. By coupling long-term foresight with short-term adaptability, the proposed design enables proactive, SLA-aware scheduling under highly dynamic conditions. We further develop a trace-driven network emulator to reproduce realistic 5G behaviors and validate system robustness. Extensive experiments demonstrate that our framework consistently achieves over 95% SLA compliance, below 2% latency violations, and up to 30% latency reduction compared with state-of-the-art baselines, confirming its effectiveness and scalability for next-generation immersive networks.
Mingrui Yin, Sohom Sen, Zhihao Ren, Xiaoyu Fang, Yongjie Guan, Tao Han 0002, Nirwan Ansari
SEC5
2025 MRCoach: A Real-Time IoT-Enabled Mixed Reality System With Semantic-Aware Transmission for Smart Sports and Personalized Coaching
abstract
Real-time transmission of large-scale data, high computational demands, and resource limitations on edge devices pose significant challenges for intelligent sports systems. The proliferation of Internet of Things (IoT) technologies has catalyzed the rise of Smart Sport, where wearable sensors, cameras, and intelligent algorithms are integrated to revolutionize athletic training. Despite this transformation, access to professional coaching remains constrained by factors such as time, cost, and scalability. A critical limitation of existing remote coaching approaches is their inability to perform effective spatio-temporal analysis, hindering comprehensive evaluation and refinement of athletic performance. This paper introduces MRCoach, a mixed reality-based, immersive, and interactive sports coaching system that enables data-driven training without requiring in-person supervision. MRCoach reconstructs 3D volumetric avatars of both learners and expert athletes, allowing users to visualize and compare their movements side-by-side in a mixed reality environment for intuitive skill refinement. To ensure responsive and efficient feedback, we propose an adaptive semantic transmission strategy that prioritizes sport-relevant joints, thereby reducing latency and bandwidth requirements without sacrificing accuracy. Furthermore, a 3D sports analysis framework is developed to evaluate motion based on normalized joint positions, velocities, and accelerations. This framework computes real-time similarity scores and delivers actionable guidance to learners. Experimental results across four sports—tennis, soccer, basketball, and baseball—demonstrate MRCoach’s effectiveness in providing personalized, real-time training experiences. Compared to state-of-the-art baselines such as MagicStream, ExPose, and PIXIE, MRCoach achieves significantly lower end-to-end latency (79.9 ms) and higher frame rates (≥ 56 FPS), while maintaining accurate pose tracking and high avatar fidelity.
Mingrui Yin, Sohom Sen, Yongjie Guan, Dhananjay Jagdish Dubey, Xueyu Hou, Tao Han 0002, Nirwan Ansari
IEEE Internet Things J.3
2025 ViEdge: Video Analytics on Distributed Edge
abstract
With the development of edge computing and increasing demand on video analytics, it is attractive to implement distributed video analytics across edge devices. In this article, we propose ViEdge, a distributed video analytics system across edge devices. ViEdge differs from status quo edge-side video analytics systems in: First, ViEdge does not assume the existence of an edge/cloud server. Instead of processing in a cascaded way between edge devices and server, the video analytics in ViEdge is processed across distributed edge devices in a parallel way. Second, ViEdge addresses two practical challenges in video analytics systems. Specifically, ViEdge optimizes the performance of glance-and-focus object detection pipeline and query related processing with multiple query types. The characters of edge devices (computing capabilities and network conditions) and features of query types (computational complexities and input/output sizes) are comprehensively considered in development of components in ViEdge. By modeling the challenges as multiway number partitioning problems, ViEdge provides practical solutions to optimizing the object detection pipeline and allocations of multiple queries of different types across distributed edge devices. Compared to baseline methods in distributed video analytics across edge devices, ViEdge reaches 1.4 × to 5.3 × speedup in different network environments with neglectable overhead.
Xueyu Hou, Yongjie Guan, Tao Han 0002
ACM Trans. Internet Things2
2024 Health-MR: A Mixed Reality-Based Patient Registration and Monitor Medical System
abstract
In medical procedures such as surgery and diagnosis, it is crucial to provide doctors and nurses with up-to-date patient information. In this paper, we propose Health-MR, a portable Mixed-Reality (MR) system that helps medical staff monitor patient conditions. Health-MR consists of three components: (1) Patient Identification Recognition using face detection, (2) Medical Cloud Database for patient information retrieval, and (3) Non-invasive Heart Rate Measurement via image processing and Fast Fourier Transform (FFT). Our evaluation demonstrates that Health-MR significantly reduces the time needed to query patient information and provides remote, accurate, and real-time heart rate monitoring.
Mingrui Yin, Sohom Sen, Yongjie Guan, Xueyu Hou, Tao Han 0002
MobiCom3
2024 BPS: Batching, Pipelining, Surgeon of Continuous Deep Inference on Collaborative Edge Intelligence
abstract
Users on edge generate deep inference requests continuously over time. Mobile/edge devices located near users can undertake the computation of inference locally for users, e.g., the embedded edge device on an autonomous vehicle. Due to limited computing resources on one mobile/edge device, it may be challenging to process the inference requests from users with high throughput. An attractive solution is to (partially) offload the computation to a remote device in the network. In this paper, we examine the existing inference execution solutions across local and remote devices and propose an adaptive scheduler, a BPS scheduler, for continuous deep inference on collaborative edge intelligence. By leveraging data parallel, neurosurgeon, reinforcement learning techniques, BPS can boost the overall inference performance by up to$8.2 \times$over the baseline schedulers. A lightweight compressor, FF, specialized in compressing intermediate output data for neurosurgeon, is proposed and integrated into the BPS scheduler. FF exploits the operating character of convolutional layers and utilizes efficient approximation algorithms. Compared to existing compression methods, FF achieves up to 86.9% lower accuracy loss and up to 83.6% lower latency overhead.
Xueyu Hou, Yongjie Guan, Nakjung Choi, Tao Han 0002
IEEE Trans. Cloud Comput.2
2023 Dystri: A Dynamic Inference based Distributed DNN Service Framework on Edge
abstract
Deep neural network (DNN) inference poses unique challenges in serving computational requests due to high request intensity, concurrent multi-user scenarios, and diverse heterogeneous service types. Simultaneously, mobile and edge devices provide users with enhanced computational capabilities, enabling them to utilize local resources for deep inference processing. Moreover, dynamic inference techniques allow content-based computational cost selection per request. This paper presents Dystri, an innovative framework devised to facilitate dynamic inference on distributed edge infrastructure, thereby accommodating multiple heterogeneous users. Dystri offers a broad applicability in practical environments, encompassing heterogeneous device types, DNN-based applications, and dynamic inference techniques, surpassing the state-of-the-art (SOTA) approaches. With distributed controllers and a global coordinator, Dystri allows per-request, per-user adjustments of quality-of-service, ensuring instantaneous, flexible, and discrete control. The decoupled workflows in Dystri naturally support user heterogeneity and scalability, addressing crucial aspects overlooked by existing SOTA works. Our evaluation involves three multi-user, heterogeneous DNN inference service platforms deployed on distributed edge infrastructure, encompassing seven DNN applications. Results show Dystri achieves near-zero deadline misses and excels in adapting to varying user numbers and request intensities. Dystri outperforms baselines with accuracy improvement up to 95 ×.
Xueyu Hou, Yongjie Guan, Tao Han 0002
ICPP2
2023 MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real Time
abstract
While recent work explored streaming volumetric content on-demand, there is little effort on live volumetric video streaming that bears the potential of bringing more exciting applications than its on-demand counterpart. To fill this critical gap, in this paper, we propose MetaStream, which is, to the best of our knowledge, the first practical live volumetric content capture, creation, delivery, and rendering system for immersive applications such as virtual, augmented, and mixed reality. To address the key challenge of the stringent latency requirement for processing and streaming a huge amount of 3D data, MetaStream integrates several innovations into a holistic system, including dynamic camera calibration, edge-assisted object segmentation, cross-camera redundant point removal, and foveated volumetric content rendering. We implement a prototype of MetaStream using commodity devices and extensively evaluate its performance. Our results demonstrate that MetaStream achieves low-latency live volumetric video streaming at close to 30 frames per second on WiFi networks. Compared to state-of-the-art systems, MetaStream reduces end-to-end latency by up to 31.7% while improving visual quality by up to 12.5%.
Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002
MobiCom1
2022 DistrEdge: Speeding up Convolutional Neural Network Inference on Distributed Edge Devices
abstract
As the number of edge devices with computing resources (e.g., embedded GPUs, mobile phones, and laptops) in-creases, recent studies demonstrate that it can be beneficial to col-laboratively run convolutional neural network (CNN) inference on more than one edge device. However, these studies make strong assumptions on the devices' conditions, and their application is far from practical. In this work, we propose a general method, called DistrEdge, to provide CNN inference distribution strategies in environments with multiple IoT edge devices. By addressing heterogeneity in devices, network conditions, and nonlinear characters of CNN computation, DistrEdge is adaptive to a wide range of cases (e.g., with different network conditions, various device types) using deep reinforcement learning technology. We utilize the latest embedded AI computing devices (e.g., NVIDIA Jetson products) to construct cases of heterogeneous devices' types in the experiment. Based on our evaluations, DistrEdge can properly adjust the distribution strategy according to the devices' computing characters and the network conditions. It achieves 1.1 to 3 x speedup compared to state-of-the-art methods.
Xueyu Hou, Yongjie Guan, Tao Han 0002, Ning Zhang 0007
IPDPS2
2022 NeuLens: spatial-based dynamic acceleration of convolutional neural networks on edge
abstract
Convolutional neural networks (CNNs) play an important role in today's mobile and edge computing systems for vision-based tasks like object classification and detection. However, state-of-the-art methods on CNN acceleration are trapped in either limited practical latency speed-up on general computing platforms or latency speed-up with severe accuracy loss. In this paper, we propose a spatial-based dynamic CNN acceleration framework, NeuLens, for mobile and edge platforms. Specially, we design a novel dynamic inference mechanism, assemble region-aware convolution (ARAC) supernet, that peels off redundant operations inside CNN models as many as possible based on spatial redundancy and channel slicing. In ARAC supernet, the CNN inference flow is split into multiple independent micro-flows, and the computational cost of each can be autonomously adjusted based on its tiled-input content and application requirements. These micro-flows can be loaded into hardware like GPUs as single models. Consequently, its operation reduction can be well translated into latency speed-up and is compatible with hardware-level accelerations. Moreover, the inference accuracy can be well preserved by identifying critical regions on images and processing them in the original resolution with large micro-flow. Based on our evaluation, NeuLens outperforms baseline methods by up to 58% latency reduction with the same accuracy and by up to 67.9% accuracy improvement under the same latency/memory constraints.
Xueyu Hou, Yongjie Guan, Tao Han 0002
MobiCom2
2022 DeepMix: mobility-aware, lightweight, and hybrid 3D object detection for headsets
abstract
Mobile headsets should be capable of understanding 3D physical environments to offer a truly immersive experience for augmented/mixed reality (AR/MR). However, their small form-factor and limited computation resources make it extremely challenging to execute in real-time 3D vision algorithms, which are known to be more compute-intensive than their 2D counterparts. In this paper, we propose DeepMix, a mobility-aware, lightweight, and hybrid 3D object detection framework for improving the user experience of AR/MR on mobile headsets. Motivated by our analysis and evaluation of state-of-the-art 3D object detection models, DeepMix intelligently combines edge-assisted 2D object detection and novel, on-device 3D bounding box estimations that leverage depth data captured by headsets. This leads to low end-to-end latency and significantly boosts detection accuracy in mobile scenarios. A unique feature of DeepMix is that it fully exploits the mobility of headsets to fine-tune detection results and boost detection accuracy. To the best of our knowledge, DeepMix is the first 3D object detection that achieves 30 FPS (i.e., an end-to-end latency much lower than the 100 ms stringent requirement of interactive AR/MR). We implement a prototype of DeepMix on Microsoft HoloLens and evaluate its performance via both extensive controlled experiments and a user study with 30+ participants. DeepMix not only improves detection accuracy by 9.1--37.3% but also reduces end-to-end latency by 2.68--9.15×, compared to the baseline that uses existing 3D object detection models.
Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002
MobiSys1
2020 Distributed Video Analysis for Mobile Live Broadcasting Services
abstract
While webcast platforms on mobile devices are becoming more and more prevalent, inspection for irregularities is getting harder and harder. To solve this problem, the convolution neural network(CNN) has been applied to recognize or detect specified objections in pictures and videos. However, when supervising large platforms, it isn’t very easy to collect mountain piles of video data and send them to the computation center. Other problems like long time delay and the high computational burden will reduce system performance, especially when dealing with data from live streams. This paper presents a method to coordinate mobile devices with remote servers(computers or embedded systems) to achieve real-time monitoring of live streams. The system can make use of computational capacity on mobile devices and reduce the cost of sending data while guaranteeing accuracy for supervision.
Yuanqi Chen, Yongjie Guan, Tao Han 0002
WCNC2