EDBT 2026 Demo / reviewers in the wild / expert
Xiaolin Guo
dblp:34/4714
· DBLP profile ↗
19ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Popularity-Aware Layer-Wise Caching and Function Scheduling for Dynamic Workflow at the EdgeabstractServerless Edge Computing (SEC) has emerged as a promising paradigm for delivering low-latency, resource-efficient services for edge-native applications, which are implemented as dependent functions, forming Directed Acyclic Graph (DAG) workflows. Unfortunately, the application's performance is hindered by the notorious issue of cold startup, especially in resource-constrained SEC environments. Layer- wise container caching has been proven to be an effective startup acceleration solution in SEC, due to its fine granularity and flexibility. However, due to the dynamic nature of call graphs and the skewness in function popularity in DAG workflows, as well as the heterogeneity of container layer cold start time and edge computing environments, the performance of existing layer- wise caching mechanisms degrades significantly. To solve this problem, we propose an efficient DAG workflow deployment method in SEC to minimize the application completion time (ACT) in the long term. We model the problem as a joint optimization of container layer- wise caching and function scheduling, which is a Time-coupled Integer Nonlinear Programming (TINLP) problem. To solve it, we first convert it to an Integer Linear Programming (ILP) problem and propose an online algorithm with theoretical performance guarantees. Extensive experiments demonstrate that our method achieves up to$2.92\times$speedup in ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise Caching
Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Daheng Yin |
INFOCOM | 3 |
| 2025 | ADPTD: Adaptive Data Partition With Unbiased Task Dispatching for Video Analytics at the EdgeabstractRecently, edge-assisted methods have been proposed as a promising technique to deliver fast and accurate on-device video analytics by partitioning frame data and dispatching them to edge servers for parallel execution. However, the data partition (DP) reduces the detection latency but decreases accuracy since objects may cross the boundaries of adjacent blocks. The effect of DP on the accuracy and latency depends on multiple vital parameters (e.g., target size, density, network, and computing resources) in an unknown and time-varying fashion. Moreover, these parameters are determined by the application scenarios and edge environment, which are uncertain and heterogeneous at the edge. Hence, how to partition frames to strike a balance between accuracy and latency is a nontrivial and intractable problem. To this end, we propose an online learning-based device-edge–cloud collaboration framework, ADPTD, to guide DP at the edge. We propose an optimal task dispatching algorithm (OTD) to minimize detection latency. Then, we propose a multiarmed bandit-based algorithm to pick a DP strategy and invoke OTD to dispatch tasks in each time slot. Theoretical analysis reveals that ADPTD achieves sublinear regret. Extensive experimental results show that ADPTD outperforms the state-of-the-art methods, achieving a latency reduction of up to$2.53\times $and improving accuracy by up to 49.4%. Zhaowu Huang, Fang Dong 0001, Haopeng Zhu, Mengyang Liu, Dian Shen, Ruiting Zhou, Xiaolin Guo, Baijun Chen |
IEEE Internet Things J. | 7 |
| 2025 | Resource-Efficient DNN Inference With Early Exiting in Serverless Edge ComputingabstractServerless Edge Computing (SEC) has gained widespread adoption in improving resource utilization due to its triggered event-driven model. However, deploying deep neural network (DNN) inference services directly in SEC leads to resource inefficiencies, which stem from two key factors. First, existing methods adopt model-wise function encapsulation, which requires the entire DNN model to occupy memory throughout its execution lifecycle. This increases both memory footprint and occupancy time. Second, uniform DNN inference for diversity input leads to redundant computations and additional inference time. To this end, we propose REDI, a novel framework that leverages fine-grained block-wise function encapsulation and progressive inference to provide resource-efficient DNN inference while ensuring latency requirements. REDI enables the release of memory from already inferred shallow networks and allows each request to exit early based on input data complexity, eliminating redundant computations. To fully unleash the potential, REDI jointly considers resource heterogeneity, data diversity, and environment dynamics to investigate the block-wise function placement problem. We introduce an uncertainty-aware online learning-driven algorithm with bounded regret. Finally, we conduct extensive trace-driven experiments to evaluate our methods, demonstrating that REDI achieves a significant speedup of up to$6.52\times$in terms of resource usage cost compared to state-of-the-art methods. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Jinghui Zhang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Rendering Super Resolution Video Streaming Efficiently with in-Network ComputingabstractEmerging live video streaming applications, e.g., Ultra High Definition videos and interactive video streaming, have put forward new demands for ultra-high bandwidth and reduced delay to match the desired quality of experience. Since current on-device Super-Resolution (SR) approaches are hindered by the limited end-device capabilities, we are motivated to take advantage of Mobile Edge Computing (MEC) and Computing in the Network technologies, such that SR videos can be processed on a more powerful infrastructure by integrating the resources from end-devices through the edge and all the way to the cloud. However, the integration of SR and MEC is non-trivial due to the challenges introduced by the features of SR tasks, and the heterogeneous nature of MEC resources. In this paper, we endeavor to explore and solve these challenges by presenting AVSA, which renders SR live video streaming efficiently with in-network computing. AVSA can adaptively allocate SR work-loads under heterogeneous resources and yield a cost-effective workload allocation with theoretical performance guarantees. Simulation results show that, compared with the state-of-the-art methods, our method achieves up to 13.06× speedup in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Baijun Chen, Daheng Yin |
HPCC | 1 |
| 2024 | SkewCache: Skewed Layer-wise Caching for Function Chains in Serverless Edge ComputingabstractIn serverless edge computing (SEC), traditional monolithic applications are encapsulated in multiple dependent functions, in which event-driven service provisioning introduces the cold start problem. Existing works adopt uniform container caching for each function to mitigate the cold start problem, which remains with relatively low efficiency due to ignoring the skewed invocation frequency and the heterogeneity of cold-start behaviors of functions. In this paper, we propose SkewCache, an efficient layer-wise container caching framework for frequency-skewed function chains in SEC. Based on the container’s layer structure, SkewCache enables fine-grained and balanced container caching on multiple edge servers, which takes into account both the invocation frequency skewness and the cold start latency of each function. The problem is modeled as an integer nonlinear programming (INLP) to minimize the application completion time (ACT). To solve the INLP, we first convert it to an equivalent integer linear programming (ILP) form. Then, we propose an approximation algorithm to solve the ILP with a guaranteed approximation ratio. To evaluate the performance of the proposed algorithm, we conduct intensive simulations and the results show that our algorithms outperform baselines, achieving up to 1.94× speedup in terms of ACT reduction. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Haodong Tian |
HPCC | 1 |
| 2024 | Joint Layer-wise Caching and Request Routing for Serverless Inference Acceleration at the EdgeabstractIn recent years, the serverless paradigm has been introduced into edge computing, enhancing resource utilization. However, it has also introduced the issue of cold start, which hinders low-latency responses. This cold start problem is significantly exacerbated when it comes to Artificial Intelligence (AI) applications, due to the necessity of loading bulky deep learning (DL) model files. Existing methods typically address the cold start issue by caching complete containers, which is inefficient in resource-constrained edge servers. This is because massive DL models may not fit in the limited cache space or lead to fragmentation of the cache memory. To solve the above problem, we seek a layer-wise model caching method that regards the DL model as a chain structure of multiple layers and caches a part of the model layers. However, heterogeneity in cold starts and edge computing’s complex environments pose challenges in selecting model layers and servers. In this paper, we investigate the joint layer-wise caching and request routing problem to minimize application completion time (ACT). We propose an approximation algorithm with a provable performance guarantee. Finally, intensive trace-driven simulations show that our algorithm achieves 3.03× speedup in terms of ACT reduction. Zhaowu Huang, Fang Dong 0001, Xiaolin Guo, Haodong Tian |
HPCC | 3 |
| 2023 | Adaptive Overlap Padding and Resolution Selection for Frame Split-based Edge Video AnalyticsabstractFor providing accurate and fast on-device high-resolution video analytics, edge-assisted methods are widely proposed by using lower-resolution frames and splitting them with overlap padding. The use of low-resolution frames can significantly reduce the computational workload of video analytics. Dividing a frame into multiple overlapping parts simultaneously improves parallelism and maintains processing accuracy. However, accuracy and latency serve as a pair of tradeoff metrics, and prioritizing one to optimize overlap padding or resolution selection will compromise the other metric. Fortunately, we have discovered that in real-world video analytics scenarios with varying object sizes, it is not necessary to simultaneously achieve a high overlap padding size and high resolution. Hence, how to set appropriate overlap padding size and resolution to strike a balance between accuracy and latency in practical scenarios is a nontrivial and intractable problem. To this end, we propose an online learning-based method to achieve adaptive overlap padding and resolution selection, called APR. We model the problem as an integer programming and propose a Muli-armed bandit (MAB) theory-based algorithm to solve it. We discretize the continuum overlap padding size into a finite set to narrow explore space and set the frame split strategy as context information to achieve fast convergence. Theoretical analysis reveals APR achieves sub-linear regret. Extensive experimental results show APR outperforms the benchmark methods, achieving up to 2.06 × speedup in terms of latency and 0.18× increase in accuracy. Haopeng Zhu, Zhaowu Huang, Xiaolin Guo, Mengyang Liu, Baijun Chen, Fang Dong 0001 |
ICPADS | 3 |
| 2023 | WAEVSR: Enabling Collaborative Live Video Super-Resolution in Wide-Area MEC EnvironmentabstractLive video streaming is increasingly popular for its rich content and real-time interactions, but its demand for bandwidth has put a heavy burden on backbone networks. To save bandwidth, recent studies have proposed neural-enhanced live video streaming that deploys deep neural networks (DNNs) for video super-resolution (VSR) on end devices or nearby edge devices to enhance video quality by taking low-resolution frames as input and producing high-resolution output frames. In this solution, the high computational demands of high-quality VSR DNNs make them difficult to support on single end or edge device, necessitating the use of distributed resources in edge facilities. However, the distributed deployment of high-quality VSR DNNs for low-latency inference remains challenging due to the inherent data dependencies of VSR DNNs and the heterogeneity and dynamics of edge facilities. In this paper, we present WAEVSR, a novel collaborative neural-enhanced live video super-resolution system that enables effective leverage of distributed resources to maximize the latency-bounded quality in wide-area MEC environments. WAEVSR consists of two key components: 1) It deploys a parallel-friendly video super-resolution DNN among edge devices, 2) with an inference controller based on the variable-size sliding window to balance the latency and quality of distributed inference in the heterogeneous and dynamics MEC environment. Prototype-based evaluation shows that WAEVSR can achieve 2.5 × lower end-to-end latency than traditional super-resolution serving with a 0.01 drop in SSIM score. The case study also demonstrates its higher stability on latency than vanilla distributed MEC deployment. Daheng Yin, Fang Dong 0001, Baijun Chen, Dian Shen, Ruiting Zhou, Xiaolin Guo, Zhaowu Huang |
IWQoS | 6 |
| 2023 | ROIAdaptor: Adaptive Task Offloading of ROI-Encoded Videos for Edge Video AnalyticsabstractReal-time analytics on video data demands intensive computation resources and high bandwidth consumption. Edge computing enables us to offload resource-intensive analytics tasks to nearby edge servers, effectively reducing the extended latency. Numerous studies have applied ROI encoding technology to decrease video data size, thereby decreasing transmission latency. However, ROI encoding will reduce inference accuracy and indirectly affect inference latency. Existing works have ignored these impacts, leading to significant performance degradation. In this paper, we first identified the necessity of considering video QP settings when offloading, and investigated the impacts of changing QP settings on accuracy and latency to demonstrate that adaptive offloading based on QP settings can improve the efficiency of video analytics. We formulate the problem of minimizing average latency to meet real-time requirements under accuracy constraints, and propose an online offloading algorithm called ROIAdaptor, based on a contextual multi-armed bandit method. Our algorithm is developed based on LinUCB, a contextual multi-armed bandit method, and operates online with historical information, achieving a provable performance bound. Simulation results show that ROIAdaptor can reduce the overall latency by an average of 35.7% and meet the requirements for real-time video analysis, with virtually no loss in accuracy. Zhenxuan Xu, Xiaolin Guo, Zhaowu Huang, Shucun Fu, Fang Dong 0001 |
MSN | 2 |
| 2023 | Enabling Distributed and Optimal RDMA Resource Sharing in Large-Scale Data Center Networks: Modeling, Analysis, and ImplementationabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability issues in solving a large-scale RDMA resource scheduling problem. In this paper, we present how to optimally share the RDMA resources in large-scale data center networks with a distributed manner. First, we propose Distributed RDMA NUM (DRUM) to model the RDMA resource scheduling problem as a new variation of the NUM problem. Second, we present distributed algorithms to efficiently solve the large-scale, interdependent RDMA resource sharing problem for different RDMA use cases. Through theoretical analysis, the convergence and parallelism of proposed algorithms are guaranteed. Finally, we implement the algorithms as a kernel-level indirection module in the real-world RDMA environment, so as to provide end-to-end resource sharing and performance guarantee. Through extensive evaluations by large-scale simulations and testbed experiments, we show that our method significantly improves applications’ performance under resource contention, achieving$1.7-3.1\times $higher throughput, and in a dynamic context, the largest performance improvement reaches 98.1% and 64.1% in terms of latency and throughput, respectively. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, Ciyuan Chen, John C. S. Lui |
IEEE/ACM Trans. Netw. | 4 |
| 2023 | Dynamic Linearization and Extended State Observer-Based Data-Driven Adaptive ControlabstractThis article aims at solving the problems of data-driven control design in the presence of strong uncertainties, hard nonlinearities, and model dependency by using a dynamic linearization (DL) method and an extended state observer (ESO). An unknown nonlinear nonaffine system is considered, whose input–output dynamics is then equivalently reformulated into a modified linear data model (mLDM) in which both a linear parametric increment description that is affine to the control input and the unmodeled uncertainties along with disturbances are included without omission or approximation. The uncertain parameter of the mLDM is estimated in real time by designing an adaptive mechanism, and the unmodeled uncertainties and disturbances are considered as a total extended state which is further estimated by developing a linear ESO. Subsequently, a modified DL-and-ESO-based data-driven adaptive control (mDLESO-DDAC) is proposed by using knowledge from previous control input to improve the control performance. The theoretical results are mathematically proved and then verified by simulations. Ronghu Chi, Xiaolin Guo, Na Lin 0002, Biao Huang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Enabling Latency-Sensitive DNN Inference via Joint Optimization of Model Surgery and Resource Allocation in Heterogeneous EdgeabstractNowadays, edge computing is widely adopted to resolve the emerging deep neural networks (DNNs)-driven intelligence scenarios with the requirement of low-latency and high-accuracy, which includes heterogeneous end devices and DNNs. In such scenarios, the influx of data and computation into a shared edge server incurs prohibitive latency. Thus, we exploit the advantage of Multi-exit DNNs (ME-DNNs) that tasks can exit early at appropriate depths to save inference time. However, naively using ME-DNNs in the heterogeneous edge still fails to deliver fast inference due to improper model surgery and resource allocation. Zhaowu Huang, Fang Dong 0001, Dian Shen, Huitian Wang, Xiaolin Guo, Shucun Fu |
ICPP | 5 |
| 2022 | Exploiting the Computational Path Diversity with In-network Computing for MECabstractWith Computing in the Network technologies, Mobile Edge Computing (MEC) has expanded the resource distribution and tightly integrated computing-network capabilities from the end-devices, through the edge, to the cloud infrastructure, including at points in between. Thus, edge computing is able to deliver a more collaborative processing, better service responding to the increasing application needs in low latency processing. In the presence of integrated computing-network resources and their increased capacity, current proximity-to-data methods in edge computing lead to sub-optimal performance in terms of processing latency. Addressing this issue, this paper presents a Low-latency Adaptive Workload Allocation framework (LAWA) to harness the growing in-network computing resources to deliver low latency processing capabilities for emerging latency-constrained applications. LAWA defines an application by its computational source and destination. Considering the diversity of computing and network resources, we try to find an optimal computational path and its workload allocation. We model the problem as a mixed integer programming problem. To solve this problem, we propose the computational pathfinding and workload allocation algorithms with optimality guarantees. Experimental results show that, comparing with the state-of-the-art methods, our method achieves up to 8.04× speedup, in terms of end-to-end latency. Xiaolin Guo, Fang Dong 0001, Dian Shen, Zhaowu Huang, Zhenyang Ni, Yulong Jiang, Daheng Yin |
SECON | 1 |
| 2020 | Distributed and Optimal RDMA Resource Scheduling in Shared Data Center NetworksabstractRemote Direct Memory Access (RDMA) suffers from unfairness issues and performance degradation when multiple applications share RDMA network resources. Hence, an efficient resource scheduling mechanism is urged to optimally allocates RDMA resources among applications. However, traditional Network Utility Maximization (NUM) based solutions are inadequate for RDMA due to three challenges: 1) The standard NUM-oriented algorithm cannot deal with coupling variables introduced by multiple dependent RDMA operations; 2) The stringent constraint of RDMA on-board resources complicates the standard NUM by bringing extra optimization dimensions; 3) Naively applying traditional algorithms for NUM suffers from scalability and convergence issues in solving a large-scale RDMA resource scheduling problem. Dian Shen, Junzhou Luo, Fang Dong 0001, Xiaolin Guo, John C. S. Lui |
INFOCOM | 4 |
| 2020 | A fault diagnosis method of rolling bearing based on VMD Tsallis entropy and FCM clustering
Xing Ting-ting, Zeng Yan, Meng Zong, Xiaolin Guo |
Multim. Tools Appl. | 4 |
| 2019 | Community Detection in Opportunistic Networks Based on Hierarchical MappingabstractIn order to solve the problem that the community partition results are not reusable in the information sensitive opportunity network, a hierarchical model of opportunity network is proposed. First, the physical node set of the opportunity network is mapped to the virtual node set corresponding to the message type, and the virtual opportunity network layer is established on this basis. Then, at the virtual opportunity network layer, social relations are established for the virtual node set. Finally, the social relations of virtual node sets are processed for community division. The experiment shows that the opportunity network hierarchical model under the unsteady topology condition is very effective. Xiaolin Guo |
CSCWD | 4 |
| 2019 | Rendering differential performance preference through intelligent network edge in cloud data centersabstractSummary Sharing the network infrastructure, the performance of emerging distributed applications and services in data centers is directly impacted by the network. As these applications are becoming more and more demanding, it is challenging to satisfy their requirements of low latency, high throughput, and low packet loss rate simultaneously. Prior approaches typically resort to flow control or scheduling mechanisms, prioritizing flows according to their demands. However, none of the methods can solely satisfy the various demands of data center applications. Addressing this challenge, we propose tasch, a preference aware flow scheduling mechanism equipped in the software network edge (ie, end‐host networking). This mechanism utilizes multiple separate queues for flows with different preferences, which guarantees low packet delay for latency‐sensitive flows and provides bandwidth guarantees for throughput‐sensitive flows. A coordinating algorithm is presented to share the network resource among multiple queues with pareto‐optimality. tasch is implemented as a thin and plugable kernel module in Linux based hypervisors, which lies between the complicated physical network and tenants VMs. Subsequently, based on the flow traces of real‐world applications, extensive experiments were conducted to verify the effectiveness of network management mechanism. Dian Shen, Yidan Gao, Xiaolin Guo, Runqun Xiong |
Concurr. Comput. Pract. Exp. | 4 |
| 2009 | WPBench: a benchmark for evaluating the client-side performance of web 2.0 applicationsabstractIn this paper, a benchmark called WPBench is reported to evaluate the responsiveness of Web browsers for modern Web 2.0 applications. In WPBench, variations of servers and networks are removed and the benchmark result is the closest to what Web users would perceive. To achieve these, WPBench records users' interactions with typical Web 2.0 applications, and then replays Web navigations when benchmarking browsers. The replay mechanism can emulate the actual user interactions and the characteristics of the servers and the networks in a consistent way independent of browsers so that any browser compliant to the standards can be benchmarked fairly. In addition to describing the design and generation of WPBench, we also report the WPBench comparison results on the responsiveness performance for three popular Web browsers: Internet Explorer, Firefox and Chrome. Kaimin Zhang, Lu Wang 0002, Xiaolin Guo, Aimin Pan, Bin B. Zhu |
WWW | 3 |