VLDB 2026 Research / reviewers in the wild / expert
Xueyu Hou
dblp:284/4424
· DBLP profile ↗
14ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 4 first-author · 9 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Falcon: A Mobile Semantic Visual Perception Framework Enlightened by Human Vision SystemsabstractIn the realm of computer human visual perception, semantic perception means recognizing objects, people, and scenes. This involves not just detecting shapes and colors but understanding what those visual elements represent. Despite its vital role in enhancing various aspects of emerging applications such as safety for autonomous driving and immersion for mixed reality (MR), real-time segmentation on mobile and edge platforms is challenging due to the nature of dense pixel labeling. To address this issue, we propose Falcon, a lightweight focus-aware segmentation framework that effectively integrates multiple innovations to achieve real-time segmentation on resource-constrained mobile and edge devices. We design a novel, low-dimension feature for efficient pixel labeling with shallow neural networks, an agile focus-aware refining scheme to compensate for the coarse nature of holistic segmentation, and a modularized design to accommodate the diversity of mobile and edge platforms and ensure seamless integration with different segmentation models. We build a prototype implementation of Falcon that supports both on-device executions and edge-assisted offloading, and asynchronous segmentation and refinement. We extensively evaluate the performance of Falcon for autonomous driving and MR applications with real setups and standard datasets. Our results demonstrate that Falcon achieves real-time segmentation, with an impressive rate of up to 40 frames per second. Xueyu Hou, Yongjie Guan, Tao Han 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Device-Server Collaborative Speculative Decoding for Real-Time LLM Streaming
Bishakha Rani Biswas, Yongjie Guan, Mingrui Yin, Tao Han 0002, Xueyu Hou |
GLOBECOM | 5 |
| 2025 | ForestRAG-AR: An AR and RAG Framework for Context-Aware Silviculture AssistanceabstractSilvicultural decision making in forestry relies heavily on technical manuals and expert knowledge that are difficult to access in the field. Meanwhile, advances in wearable augmented reality (AR) and retrieval-augmented generation (RAG) offer new opportunities to deliver context-aware, data-driven guidance directly within natural environments. This paper presents ForestRAG-AR, an integrated AR and RAG framework that provides site-specific silvicultural recommendations grounded in authoritative forestry documents. The system combines (i) a perception and context extraction module that interprets local forest conditions, including tree species, stand density, and terrain slope, through on-device sensing, (ii) a RAG-based knowledge backend that adapts a large language model to forestry by embedding regional silviculture manuals and best-management-practice guides, and (iii) an AR interface that visualizes thinning and pruning guidance with natural-language explanations and verifiable citations. A domain-adapted cross-encoder reranker fuses environmental context with text passages to improve retrieval precision and factual grounding. Experiments on 50 forestry queries demonstrate that ForestRAG-AR improves nDCG@10 by 7.8 points and citation precision to 0.96 while maintaining sub-500 ms end-to-end latency. Md Mehedi Hasan, Bishakha Rani Biswas, Xueyu Hou, Yongjie Guan |
SEC | 3 |
| 2025 | Empirical Analysis of LLMDPP: Advancing Log Parsing in the LLM EraabstractIn the field of software engineering, the automated analysis of log data is crucial for operations and maintenance teams. This study introduces LLMDPP, a novel log parser that leverages Large Language Models (LLMs) and Determinantal Point Process (DPP) sampling techniques to enhance the efficiency and accuracy of online log parsing. LLMDPP transforms raw log messages into structured log templates through a few-shot learning, thus simplifying the processing and analysis of log data. The study explores the accuracy of LLMDPP in log-parsing tasks and compares the effectiveness of different encoding functions (TF-IDF and Flan-T5-small embedding layer) in DPP sampling. Experimental results show that LLMDPP outperforms traditional methods in both Global Accuracy and Parsing Accuracy, with the semantic information encoding function performing better when the number of samples is low. Furthermore, we simulate an online scenario to evaluate the parsing effectiveness of different sampling methods on unseen log datasets. The results indicate that the DPP sampling method has an advantage in maintaining sample diversity and fairness, which can improve parsing accuracy. Siqin Zhang, Haijing Nan, Xueyu Hou, Jiaqi Zou, Zicong Miao |
MobiSys | 4 |
| 2025 | MRCoach: A Real-Time IoT-Enabled Mixed Reality System With Semantic-Aware Transmission for Smart Sports and Personalized CoachingabstractReal-time transmission of large-scale data, high computational demands, and resource limitations on edge devices pose significant challenges for intelligent sports systems. The proliferation of Internet of Things (IoT) technologies has catalyzed the rise of Smart Sport, where wearable sensors, cameras, and intelligent algorithms are integrated to revolutionize athletic training. Despite this transformation, access to professional coaching remains constrained by factors such as time, cost, and scalability. A critical limitation of existing remote coaching approaches is their inability to perform effective spatio-temporal analysis, hindering comprehensive evaluation and refinement of athletic performance. This paper introduces MRCoach, a mixed reality-based, immersive, and interactive sports coaching system that enables data-driven training without requiring in-person supervision. MRCoach reconstructs 3D volumetric avatars of both learners and expert athletes, allowing users to visualize and compare their movements side-by-side in a mixed reality environment for intuitive skill refinement. To ensure responsive and efficient feedback, we propose an adaptive semantic transmission strategy that prioritizes sport-relevant joints, thereby reducing latency and bandwidth requirements without sacrificing accuracy. Furthermore, a 3D sports analysis framework is developed to evaluate motion based on normalized joint positions, velocities, and accelerations. This framework computes real-time similarity scores and delivers actionable guidance to learners. Experimental results across four sports—tennis, soccer, basketball, and baseball—demonstrate MRCoach’s effectiveness in providing personalized, real-time training experiences. Compared to state-of-the-art baselines such as MagicStream, ExPose, and PIXIE, MRCoach achieves significantly lower end-to-end latency (79.9 ms) and higher frame rates (≥ 56 FPS), while maintaining accurate pose tracking and high avatar fidelity. Mingrui Yin, Sohom Sen, Yongjie Guan, Dhananjay Jagdish Dubey, Xueyu Hou, Tao Han 0002, Nirwan Ansari |
IEEE Internet Things J. | 6 |
| 2025 | ViEdge: Video Analytics on Distributed EdgeabstractWith the development of edge computing and increasing demand on video analytics, it is attractive to implement distributed video analytics across edge devices. In this article, we propose ViEdge, a distributed video analytics system across edge devices. ViEdge differs from status quo edge-side video analytics systems in: First, ViEdge does not assume the existence of an edge/cloud server. Instead of processing in a cascaded way between edge devices and server, the video analytics in ViEdge is processed across distributed edge devices in a parallel way. Second, ViEdge addresses two practical challenges in video analytics systems. Specifically, ViEdge optimizes the performance of glance-and-focus object detection pipeline and query related processing with multiple query types. The characters of edge devices (computing capabilities and network conditions) and features of query types (computational complexities and input/output sizes) are comprehensively considered in development of components in ViEdge. By modeling the challenges as multiway number partitioning problems, ViEdge provides practical solutions to optimizing the object detection pipeline and allocations of multiple queries of different types across distributed edge devices. Compared to baseline methods in distributed video analytics across edge devices, ViEdge reaches 1.4 × to 5.3 × speedup in different network environments with neglectable overhead. Xueyu Hou, Yongjie Guan, Tao Han 0002 |
ACM Trans. Internet Things | 1 |
| 2024 | Health-MR: A Mixed Reality-Based Patient Registration and Monitor Medical SystemabstractIn medical procedures such as surgery and diagnosis, it is crucial to provide doctors and nurses with up-to-date patient information. In this paper, we propose Health-MR, a portable Mixed-Reality (MR) system that helps medical staff monitor patient conditions. Health-MR consists of three components: (1) Patient Identification Recognition using face detection, (2) Medical Cloud Database for patient information retrieval, and (3) Non-invasive Heart Rate Measurement via image processing and Fast Fourier Transform (FFT). Our evaluation demonstrates that Health-MR significantly reduces the time needed to query patient information and provides remote, accurate, and real-time heart rate monitoring. Mingrui Yin, Sohom Sen, Yongjie Guan, Xueyu Hou, Tao Han 0002 |
MobiCom | 5 |
| 2024 | BPS: Batching, Pipelining, Surgeon of Continuous Deep Inference on Collaborative Edge IntelligenceabstractUsers on edge generate deep inference requests continuously over time. Mobile/edge devices located near users can undertake the computation of inference locally for users, e.g., the embedded edge device on an autonomous vehicle. Due to limited computing resources on one mobile/edge device, it may be challenging to process the inference requests from users with high throughput. An attractive solution is to (partially) offload the computation to a remote device in the network. In this paper, we examine the existing inference execution solutions across local and remote devices and propose an adaptive scheduler, a BPS scheduler, for continuous deep inference on collaborative edge intelligence. By leveraging data parallel, neurosurgeon, reinforcement learning techniques, BPS can boost the overall inference performance by up to$8.2 \times$over the baseline schedulers. A lightweight compressor, FF, specialized in compressing intermediate output data for neurosurgeon, is proposed and integrated into the BPS scheduler. FF exploits the operating character of convolutional layers and utilizes efficient approximation algorithms. Compared to existing compression methods, FF achieves up to 86.9% lower accuracy loss and up to 83.6% lower latency overhead. Xueyu Hou, Yongjie Guan, Nakjung Choi, Tao Han 0002 |
IEEE Trans. Cloud Comput. | 1 |
| 2023 | Dystri: A Dynamic Inference based Distributed DNN Service Framework on EdgeabstractDeep neural network (DNN) inference poses unique challenges in serving computational requests due to high request intensity, concurrent multi-user scenarios, and diverse heterogeneous service types. Simultaneously, mobile and edge devices provide users with enhanced computational capabilities, enabling them to utilize local resources for deep inference processing. Moreover, dynamic inference techniques allow content-based computational cost selection per request. This paper presents Dystri, an innovative framework devised to facilitate dynamic inference on distributed edge infrastructure, thereby accommodating multiple heterogeneous users. Dystri offers a broad applicability in practical environments, encompassing heterogeneous device types, DNN-based applications, and dynamic inference techniques, surpassing the state-of-the-art (SOTA) approaches. With distributed controllers and a global coordinator, Dystri allows per-request, per-user adjustments of quality-of-service, ensuring instantaneous, flexible, and discrete control. The decoupled workflows in Dystri naturally support user heterogeneity and scalability, addressing crucial aspects overlooked by existing SOTA works. Our evaluation involves three multi-user, heterogeneous DNN inference service platforms deployed on distributed edge infrastructure, encompassing seven DNN applications. Results show Dystri achieves near-zero deadline misses and excels in adapting to varying user numbers and request intensities. Dystri outperforms baselines with accuracy improvement up to 95 ×. Xueyu Hou, Yongjie Guan, Tao Han 0002 |
ICPP | 1 |
| 2023 | MetaStream: Live Volumetric Content Capture, Creation, Delivery, and Rendering in Real TimeabstractWhile recent work explored streaming volumetric content on-demand, there is little effort on live volumetric video streaming that bears the potential of bringing more exciting applications than its on-demand counterpart. To fill this critical gap, in this paper, we propose MetaStream, which is, to the best of our knowledge, the first practical live volumetric content capture, creation, delivery, and rendering system for immersive applications such as virtual, augmented, and mixed reality. To address the key challenge of the stringent latency requirement for processing and streaming a huge amount of 3D data, MetaStream integrates several innovations into a holistic system, including dynamic camera calibration, edge-assisted object segmentation, cross-camera redundant point removal, and foveated volumetric content rendering. We implement a prototype of MetaStream using commodity devices and extensively evaluate its performance. Our results demonstrate that MetaStream achieves low-latency live volumetric video streaming at close to 30 frames per second on WiFi networks. Compared to state-of-the-art systems, MetaStream reduces end-to-end latency by up to 31.7% while improving visual quality by up to 12.5%. Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002 |
MobiCom | 2 |
| 2022 | DistrEdge: Speeding up Convolutional Neural Network Inference on Distributed Edge DevicesabstractAs the number of edge devices with computing resources (e.g., embedded GPUs, mobile phones, and laptops) in-creases, recent studies demonstrate that it can be beneficial to col-laboratively run convolutional neural network (CNN) inference on more than one edge device. However, these studies make strong assumptions on the devices' conditions, and their application is far from practical. In this work, we propose a general method, called DistrEdge, to provide CNN inference distribution strategies in environments with multiple IoT edge devices. By addressing heterogeneity in devices, network conditions, and nonlinear characters of CNN computation, DistrEdge is adaptive to a wide range of cases (e.g., with different network conditions, various device types) using deep reinforcement learning technology. We utilize the latest embedded AI computing devices (e.g., NVIDIA Jetson products) to construct cases of heterogeneous devices' types in the experiment. Based on our evaluations, DistrEdge can properly adjust the distribution strategy according to the devices' computing characters and the network conditions. It achieves 1.1 to 3 x speedup compared to state-of-the-art methods. Xueyu Hou, Yongjie Guan, Tao Han 0002, Ning Zhang 0007 |
IPDPS | 1 |
| 2022 | NeuLens: spatial-based dynamic acceleration of convolutional neural networks on edgeabstractConvolutional neural networks (CNNs) play an important role in today's mobile and edge computing systems for vision-based tasks like object classification and detection. However, state-of-the-art methods on CNN acceleration are trapped in either limited practical latency speed-up on general computing platforms or latency speed-up with severe accuracy loss. In this paper, we propose a spatial-based dynamic CNN acceleration framework, NeuLens, for mobile and edge platforms. Specially, we design a novel dynamic inference mechanism, assemble region-aware convolution (ARAC) supernet, that peels off redundant operations inside CNN models as many as possible based on spatial redundancy and channel slicing. In ARAC supernet, the CNN inference flow is split into multiple independent micro-flows, and the computational cost of each can be autonomously adjusted based on its tiled-input content and application requirements. These micro-flows can be loaded into hardware like GPUs as single models. Consequently, its operation reduction can be well translated into latency speed-up and is compatible with hardware-level accelerations. Moreover, the inference accuracy can be well preserved by identifying critical regions on images and processing them in the original resolution with large micro-flow. Based on our evaluation, NeuLens outperforms baseline methods by up to 58% latency reduction with the same accuracy and by up to 67.9% accuracy improvement under the same latency/memory constraints. Xueyu Hou, Yongjie Guan, Tao Han 0002 |
MobiCom | 1 |
| 2022 | DeepMix: mobility-aware, lightweight, and hybrid 3D object detection for headsetsabstractMobile headsets should be capable of understanding 3D physical environments to offer a truly immersive experience for augmented/mixed reality (AR/MR). However, their small form-factor and limited computation resources make it extremely challenging to execute in real-time 3D vision algorithms, which are known to be more compute-intensive than their 2D counterparts. In this paper, we propose DeepMix, a mobility-aware, lightweight, and hybrid 3D object detection framework for improving the user experience of AR/MR on mobile headsets. Motivated by our analysis and evaluation of state-of-the-art 3D object detection models, DeepMix intelligently combines edge-assisted 2D object detection and novel, on-device 3D bounding box estimations that leverage depth data captured by headsets. This leads to low end-to-end latency and significantly boosts detection accuracy in mobile scenarios. A unique feature of DeepMix is that it fully exploits the mobility of headsets to fine-tune detection results and boost detection accuracy. To the best of our knowledge, DeepMix is the first 3D object detection that achieves 30 FPS (i.e., an end-to-end latency much lower than the 100 ms stringent requirement of interactive AR/MR). We implement a prototype of DeepMix on Microsoft HoloLens and evaluate its performance via both extensive controlled experiments and a user study with 30+ participants. DeepMix not only improves detection accuracy by 9.1--37.3% but also reduces end-to-end latency by 2.68--9.15×, compared to the baseline that uses existing 3D object detection models. Yongjie Guan, Xueyu Hou, Nan Wu 0012, Bo Han 0001, Tao Han 0002 |
MobiSys | 2 |
| 2020 | TrustServing: A Quality Inspection Sampling Approach for Remote DNN ServicesabstractDeep neural networks (DNNs) are being applied to various areas such as computer vision, autonomous vehicles, and healthcare, etc. However, DNNs are notorious for their high computational complexity and cannot be executed efficiently on resource constrained Internet of Things (IoT) devices. Various solutions have been proposed to handle the high computational complexity of DNNs. Offloading computing tasks of DNNs from IoT devices to cloud/edge servers is one of the most popular and promising solutions. While such remote DNN services provided by servers largely reduce computing tasks on IoT devices, it is challenging for IoT devices to inspect whether the quality of the service meets their service level objectives (SLO) or not. In this paper, we address this problem and propose a novel approach named QIS (quality inspection sampling) that can efficiently inspect the quality of the remote DNN services for IoT devices. To realize QIS, we design a new ID-generation method to generate data (IDs) that can identify the serving DNN models on edge servers. QIS inserts the IDs into the input data stream and implements sampling inspection on SLO violations. The experiment results show that the QIS approach can reliably inspect, with a nearly 100% success rate, the service qualtiy of remote DNN services when the SLA level is 99.9% or lower at the cost of only up to 0.5% overhead. Xueyu Hou, Tao Han 0002 |
SECON | 1 |