EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyao Huang
dblp:219/2207
· DBLP profile ↗
18ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0003-2571-1979ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 13 · 8 first-author · 10 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Pricing and Scheduling for SLA-aware Serverless GPU Services
Xiaoyao Huang, Jie Wu 0001 |
ICC | 1 |
| 2026 | A Comprehensive Comparison of Two Resource Allocation Schemes in DCNs under the Hose Model
Jie Wu 0001, Shuo Quan, Xiaoyao Huang |
ICC | 3 |
| 2026 | KV Cache Reuse for Elastic LLM Inference on Edge Devices
Peishuo Wang, Zhenzhe Zheng 0001, Xiaoyao Huang, Jie Wu 0001, Fan Wu 0006, Guihai Chen |
ICDCS | 3 |
| 2026 | TAILOR: Token-Aware Partitioning and Routing for Edge-Cloud Transformer Inference
Xiaoyao Huang, Remington R. Liu, Jie Wu 0001 |
IWQoS | 1 |
| 2025 | STPer: Task Scheduling and Traffic Routing with Spatiotemporal Dynamic Perception in Green Computing Power NetworksabstractThe growing demands of compute-intensive applications have led to the emergence of Computing Power Networks (CPNs), which integrate diverse computing resources across cloud, edge, and devices for efficient task scheduling and resource allocation in a network environment. This paper addresses the critical challenge posed by the dual spatiotemporal dynamics of tasks and resources, which complicates the optimization process of task scheduling in CPNs. Our work is the first to systematically investigate the joint optimization of task scheduling and traffic routing in green CPNs involving the dual dynamics. We formulate the profit maximization problem as an integer nonlinear programming problem that is NP-hard. To tackle this issue, we propose a novel deep reinforcement learning based algorithm capable of SpatioTemporal Dynamic Perception (STPer), which employs Graph Neural Networks(GNN) and Long Short-Term Memory(LSTM) networks. The STPer algorithm effectively captures the spatiotemporal dynamics of both resources and tasks, maximizing platform profits. Extensive simulations demonstrate that STPer significantly outperforms existing benchmark algorithms, highlighting its superior performance in optimizing resource utilization and enhancing profitability in the green CPN. Xiaoyao Huang, Remington R. Liu, Jie Wu 0001 |
GLOBECOM | 1 |
| 2025 | Latency-Aware Transformer Partitioning for Heterogeneous Edge InferenceabstractDeploying transformer models in heterogeneous edge environments poses significant challenges due to limited memory, varied compute capabilities, and asymmetric inter-node bandwidth. To enable efficient distributed inference, it is necessary to partition the model into segments that can be collaboratively executed across multiple edge nodes. However, previous works predominantly focus on either single-point partitioning or static scheduling strategies, which fall short in capturing the joint impact of segmentation and assignment on end-to-end latency. In this work, we formulate a latency-aware transformer deployment problem that jointly optimizes model segmentation and segment-to-node assignment, explicitly accounting for memory constraints and heterogeneous communication costs. We propose efficient algorithms to explore the solution space, including a latency-aware balanced partitioning heuristic (LaBP) and a dynamic programming-based optimal strategy (DP-MCP). Extensive experiments demonstrate that our approach consistently achieves significantly lower inference latency compared to existing baselines while maintaining feasibility under stringent resource constraints. Xiaoyao Huang |
ICNP | 1 |
| 2025 | Joint Prediction and Matching for Computing Resource Exchange PlatformsabstractThe rapid growth of deep learning has created unprecedented demand for computing resources, while many small and enterprise-level clusters remain underutilized. Computing resource exchange platforms offer a solution by aggregating these idle resources. However, effective cluster-task matching depends on accurate performance prediction. Existing approaches, which decouple prediction from matching, often lead to suboptimal decisions due to misaligned objectives. We propose a Matching-Focused Cluster Performance Predictor (MFCP), an end-to-end framework that integrates performance prediction with task matching to improve decision accuracy and resource utilization. Unlike existing methods that prioritize prediction accuracy, MFCP minimizes decision regret by aligning the predictor’s loss with optimal matching objectives. To handle non-differentiable matching optimization, we use continuous relaxation and incorporate constraints via an interior-point method, ensuring meaningful gradients for training. For non-convex optimization, we approximate optimal decisions with gradient descent and estimate gradients using zeroth-order perturbation. Experiments show that MFCP consistently outperforms existing methods across different cluster environments and scales, achieving lower matching regret and higher resource utilization. Da Huo 0002, Zhenzhe Zheng 0001, Xiaoyao Huang, Hao Chen 0181, Jianfeng Hu 0003, Zhiyong Yan, Fan Wu 0006, Jie Wu 0001 |
ICPP | 3 |
| 2025 | DRAGON: Enhancing On-Device Model Performance with Distributed Retrieval-Augmented GenerationabstractSmall language models (SLMs) support efficient deployments on resource-constrained edge devices, but their limited capacity compromises inference performance. Retrieval-augmented generation (RAG) is a promising solution to enhance model performance by integrating external databases, without requiring intensive on-device model retraining. However, large-scale public databases and user-specific private contextual documents are typically located on the cloud and the device, respectively, while existing RAG implementations are primarily centralized. To bridge this gap, we propose DRAGON, a distributed RAG framework to enhance on-device SLMs through both general and personal knowledge without the risk of leaking document privacy. Specifically, DRAGON decomposes multi-document RAG into multiple parallel token generation processes performed independently and locally on the cloud and the device, and employs a newly designed Speculative Aggregation, a dual-side speculative algorithm to avoid frequent output synchronization between the cloud and device. A new scheduling algorithm is further introduced to identify the optimal aggregation side based on real-time network conditions. Evaluations on real-world hardware testbed demonstrate a significant performance improvement of DRAGON—up to 1.9X greater gains over standalone SLM compared to the centralized RAG, substantial reduction in per-token latency, and negligible Time to First Token (TTFT) overhead. Shangyu Liu, Zhenzhe Zheng 0001, Xiaoyao Huang, Fan Wu 0006, Guihai Chen, Jie Wu 0001 |
MobiHoc | 3 |
| 2025 | A Hybrid CNN-Transformer Model for Tomato Leaf Disease Classification Incorporating Gray-Aware Attention and RS-Convolutional Block Attention MechanismsabstractAccurate classification of leaf diseases is crucial for plant health and effective crop management. Existing deep learning approaches are predominantly categorized into Convolutional Neural Network (CNN)-based and Vision Transformer (ViT)-based methods. However, inherent limitations in both approaches constrain further performance gains. CNNs excel at capturing local lesion details but struggle with long-range dependencies. In contrast, ViTs leverage self-attention to model global feature relationships but often overlook fine-grained local information. Moreover, most models rely solely on three-channel (RGB) inputs, underutilizing the texture details present in grayscale data. To address these problems, we propose a hybrid CNN-Transformer model enhanced with a novel Gray-Aware (GA) Attention Module and an improved Residual Convolutional Block Attention Module (Rs-CBAM). Specifically, the hybrid architecture effectively balances fine-grained detail preservation and global contextual understanding. GA strengthens texture representation, while Rs-CBAM further enhances attention to critical regions. Comparative experiments conducted on the PlantVillage and AI Challenger 2018 datasets demonstrate that our model outperforms existing models. Ablation studies further confirm the effectiveness of each proposed enhancement. Overall, the proposed approach provides a promising direction for fine-grained plant disease classification, and GA shows potential in broader image processing tasks requiring enhanced texture representation. The code is available on GitHub: https://github.com/Governeson/GA-CCB. Longquan Zou, Xiaoyao Huang, Linqiang Hu |
SMC | 3 |
| 2024 | Profit-Aware Computing Server Clustering and Task Scheduling in the Computing Power NetworkabstractComputing power network has emerged as an attractive technology to tackle the increasing demand for computational resources in cloud networking. In this paper, we study the problem of optimizing the formation of clusters of computing servers/nodes and also the task scheduling considering the games between the platform and computing nodes to maximize the platform profit. We formulate this problem as an integer programming problem. We propose a deep reinforcement-learning based server clustering and auction-based task scheduling algorithm working at different time scales to solve the problem. The deep reinforcement learning-based server clustering algorithm works at a large time scale to optimize the sizes and compositions of different clusters based on temporal and spatial distribution of the tasks and also their characteristics. The auction-based task scheduling algorithm works at a small time scale to match the tasks with the clusters while satisfying the QoS requirements of tasks so as to maximize the profit of the platform. Extensive simulations are conducted to evaluate the performance of proposed algorithm and the results show its high performance. Xiaoyao Huang, Remington R. Liu, Jie Wu 0001, Baoxian Zhang |
HPCC | 1 |
| 2024 | Platform Profit Maximization for Space-Air-Ground Integrated Computing Power Network Supplied by Green EnergyabstractThe rapid expansion of computing needs from emerging applications pushes a large amount of deployment of computing infrastructures and corresponding energy cost and greenhouse gas emissions of computing generate great concern. In this paper, we study how to maximize the platform profit by optimizing task scheduling in the Space-Air-Ground integrated Computing Power Network supplied by green energy while considering both the user requirements and dynamics of green energy. First, we formalize the problem as a binary integer linear programming problem that is NP-hard. The problem is then further modeled as a Markov decision process. Considering the dual dynamics of user requests and the generation of green energy, we propose a task scheduling strategy based on deep reinforcement learning, which can predict power generation based on the current operating status of each hydroelectric power station and also provide a scheduling strategy. Extensive experiments demonstrate that the proposed algorithm performs better than the baseline algorithms. Xiaoyao Huang, Remington R. Liu, Bo Lei 0002, Wenjuan Xing, Xing Zhang 0001 |
ICC | 1 |
| 2024 | Energy-Efficient Video Streaming With Fixed-Wing UAVabstractFixed-wing unmanned aerial vehicle (UAV) communication is a promising paradigm for providing mobile video services to ground users (GUs) without the need of infrastructure. However, the performance of fixed-wing UAV video streaming is severely limited by the onboard energy while energy-efficient fixed-wing UAV video streaming has not been fully explored. In this paper, we study energy-efficient video streaming with a fixed-wing UAV for provisioning mobile video streaming services to multiple GUs. We define UAV's energy efficiency as the ratio of perceived video quality at all GUs to the UAV's energy consumption and formulate an energy efficiency maximization problem that jointly optimizes the communication time allocation among GUs and the UAV's trajectory. Due to the non-convex nature of the formulated problem, we propose a near-optimal iterative algorithm, which utilizes successive convex approximation and quadratic transform techniques to address the problem efficiently. Extensive simulations demonstrate the high efficiency and effectiveness of our proposed algorithm. Guanglun Huang, Minghe Zhang, Xiaoyao Huang, Baoxian Zhang |
WCNC | 3 |
| 2023 | Deep Reinforcement Learning Based Multistage Profit Aware Task Scheduling Algorithm for Computing Power NetworkabstractComputing power network (CPN), which integrates heterogeneous computing resources and communication network, can tackle the challenges brought by the pervasiveness of mobile and Internet of Things applications. In this paper, we study the optimization of task scheduling in a CPN network by considering the unbalancing between task distribution and resource cost. The design objective is to maximize the system profit while satisfying tasks' delay requirements. We formulate this problem as an integer programming problem. To address this NP-hard problem, we propose a Deep Reinforcement Learning (DRL) based multistage profit-aware task scheduling algorithm which first makes coarse grained task allocation using DRL among regions and then determines an optimized intra-region task assignment by using profit-aware balancing algorithm. Extensive simulations are conducted for performance evaluation and the results show the high performance of the proposed algorithm as compared with baseline algorithms. Xiaoyao Huang, Remington R. Liu, Bo Lei 0002, Guanglun Huang, Baoxian Zhang |
GLOBECOM | 1 |
| 2023 | Platform Profit Maximization in D2D Collaboration Based Multi-Access Edge ComputingabstractMulti-access edge computing (MEC) has been an important and promising paradigm for offering computing services to mobile users with computation-intensive and latency-critical tasks. In this paper, we study a D2D collaboration based MEC system, where the service platform purchases resources from resource-rich collaborative D2D devices when the task arrival rate exceeds the platform’s capability for providing satisfactory QoS. The design objective is to maximize the platform profit while maximally satisfying the delay requirements of tasks. We define delay based utility functions for different participants and accordingly formulate the platform profit maximization problem as a Mixed Integer Non-Linear Programming (MINLP) problem. For the online case where future task arrivals are unknown in advance, we propose a reverse auction based task assignment and urgency-value based transmission scheduling algorithm (RAGM). We present the detailed algorithm design and deduce its computation complexity. We prove that RAGM satisfies individual rationality of all participants. We conduct extensive simulations and the results show the high performance of RAGM as compared with benchmark algorithms. Xiaoyao Huang, Guoliang Ji, Baoxian Zhang, Cheng Li 0005 |
IEEE Trans. Wirel. Commun. | 1 |
| 2020 | Task Allocation in Eco-friendly Mobile Crowdsensing: Problems and Algorithms
Wei Gong 0003, Xiaoyao Huang, Baoxian Zhang |
Mob. Networks Appl. | 2 |
| 2019 | Data Offloading for Mobile Crowdsensing in Opportunistic Social NetworksabstractMobile crowdsensing is a novel paradigm by exploiting mobility, sensing, computation, and communication capability of smart devices. In this paper, we study data offloading problem for mobile crowdsensing in opportunistic social networks. In this scenario, mobile users can upload sensing data directly via cellular networks using various data plans. A mobile user can also resort to another user for data offloading by forwarding sensing data to that user using short-range communications (when they encounter). To minimize total data uploading cost while meeting given uploading deadlines, data plan assignment for users and data forwarding strategy when two users encounter should be elaborately designed. In this paper, we use Benders decomposition algorithm to solve offline data plan assignment problem. Then we propose two algorithms including progress- balanced algorithm and social-aware forwarding algorithm to solve online data forwarding problem. Simulation results show that data offloading between users can largely reduce the total data uploading cost. Simulation results also show that the performance of our proposed online algorithms is close to the offline optimal solution. Wei Gong 0003, Xiaoyao Huang, Guanglun Huang, Baoxian Zhang, Cheng Li 0005 |
GLOBECOM | 2 |
| 2018 | Task assignment for Eco-friendly Mobile CrowdsensingabstractMobile crowdsensing is a sensing paradigm such that mobile users need to move to task locations to perform sensing tasks. In this paper, we focus on studying the task assignment problem of eco-friendly mobile crowdsensing which aims to minimize carbon emissions while meeting various resource limits including task deadlines and transportation constraints. We first describe the eco-friendly mobile crowdsensing system model and formulate the task assignment problem. Then we divide the problem into two subproblems including selection of best transportation type and user-task matching. We model the user-task matching as unbalanced minimum-cost bipartite matching, transform the problem into balanced maximum-weight bipartite matching, and use Kuhn-Munkres algorithm to obtain the optimal solution. Extensive simulations are conducted and the results show the efficiency and effectiveness of our proposed solution. Wei Gong 0003, Xiaoyao Huang, Baoxian Zhang |
MobiQuitous | 2 |
| 2018 | An Adaptive Bi-Threshold-Based On-Demand Energy-Efficient Multicast Routing Protocol for Wireless Ad Hoc and Sensor Networks
Xiaoyao Huang, Baoxian Zhang |
Mob. Networks Appl. | 1 |