VLDB 2026 Research / reviewers in the wild / expert
Yunchu Han
dblp:380/5602
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-7319-9348ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust DNN Partitioning and Resource Allocation Under Uncertain Inference TimeabstractIn edge intelligence systems, deep neural network (DNN) partitioning and data offloading can provide real-time task inference for resource-constrained mobile devices. However, the inference time of DNNs is typically uncertain and cannot be precisely determined in advance, presenting significant challenges in ensuring timely task processing within deadlines. To address the uncertain inference time, we propose a robust optimization scheme to minimize the total energy consumption of mobile devices while meeting task probabilistic deadlines. The scheme only requires the mean and variance information of the inference time, without any prediction methods or distribution functions. The problem is formulated as a mixed-integer nonlinear programming (MINLP) that involves jointly optimizing the DNN model partitioning and the allocation of local CPU/GPU frequencies and uplink bandwidth. To tackle the problem, we first decompose the original problem into two subproblems: resource allocation and DNN model partitioning. Subsequently, the two subproblems with probability constraints are equivalently transformed into deterministic optimization problems using the chance-constrained programming (CCP) method. Finally, the convex optimization technique and the penalty convex-concave procedure (PCCP) technique are employed to obtain the optimal solution of the resource allocation subproblem and a stationary point of the DNN model partitioning subproblem, respectively. The proposed algorithm leverages real-world data from popular hardware platforms and is evaluated on widely used DNN models. Extensive simulations show that our proposed algorithm effectively addresses the inference time uncertainty with probabilistic deadline guarantees while minimizing the energy consumption of mobile devices. Zhaojun Nan, Yunchu Han, Sheng Zhou 0001, Zhisheng Niu |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Joint Memory Frequency and Computing Frequency Scaling for Energy-efficient DNN InferenceabstractDeep neural networks (DNNs) have been widely applied in diverse applications, but the problems of high latency and energy overhead are inevitable on resource-constrained devices. To address this challenge, most researchers focus on the dynamic voltage and frequency scaling (DVFS) technique to balance the latency and energy consumption by changing the computing frequency of processors. However, the adjustment of memory frequency is usually ignored and not fully utilized to achieve efficient DNN inference, which also plays a significant role in the inference time and energy consumption. In this paper, we first investigate the impact of joint memory frequency and computing frequency scaling on the inference time and energy consumption with a model-based and data-driven method. Then by combining with the fitting parameters of different DNN models, we give a preliminary analysis for the proposed model to see the effects of adjusting memory frequency and computing frequency simultaneously. Finally, simulation results in local inference and cooperative inference cases further validate the effectiveness of jointly scaling the memory frequency and computing frequency to reduce the energy consumption of devices. Yunchu Han, Zhaojun Nan, Sheng Zhou 0001, Zhisheng Niu |
GLOBECOM | 1 |
| 2025 | DVFS-Aware DNN Inference on GPUs: Latency Modeling and Performance AnalysisabstractThe rapid development of deep neural networks (DNNs) is inherently accompanied by the problem of high computational costs. To tackle this challenge, dynamic voltage frequency scaling (DVFS) is emerging as a promising technology for balancing the latency and energy consumption of DNN inference by adjusting the computing frequency of processors. However, most existing models of DNN inference time are based on the CPU-DVFS technique, and directly applying the CPUDVFS model to DNN inference on GPUs will lead to significant errors in optimizing latency and energy consumption. In this paper, we propose a DVFS-aware latency model to precisely characterize DNN inference time on GPUs. We first formulate the DNN inference time based on extensive experiment results for different devices and analyze the impact of fitting parameters. Then by dividing DNNs into multiple blocks and obtaining the actual inference time, the proposed model is further verified. Finally, we compare our proposed model with the CPU-DVFS model in two specific cases. Evaluation results demonstrate that local inference optimization with our proposed model achieves a reduction of no less than 66% and 69% in inference time and energy consumption respectively. In addition, cooperative inference with our proposed model can improve the partition policy and reduce the energy consumption compared to the CPUDVFS model. Yunchu Han, Zhaojun Nan, Sheng Zhou 0001, Zhisheng Niu |
ICC | 1 |
| 2025 | Robust Task Offloading and Resource Allocation Under Imperfect Computing Capacity Information in Edge Intelligence SystemsabstractIn edge intelligence systems, task offloading and resource allocation policies critically depend on the required computing capacity of the task, which can only be accurately measured after execution, presenting significant design challenges. In this paper, we address the problem of robust task offloading and resource allocation under imperfect computing capacity information, where the exact value as well as distribution knowledge of the required computing capacity cannot be obtained in advance. Specifically, we formulate theenergy-time cost(ETC) minimization problem using min-max robust optimization. To tackle this challenging issue, we propose a decoupling method. This method first assumes the offloading policy is predetermined and derives two independent subproblems: local ETC and edge ETC. Then, we provide a closed-form optimal solution for the local ETC problem. The edge ETC problem is equivalently transformed into a geometric programming (GP) problem, and we introduce an effective iterative algorithm to obtain a stationary point, utilizing successive convex approximation (SCA). Finally, we design a coordinate descent (CD)-based algorithm to optimize the offloading policy effectively. Extensive simulations demonstrate that the proposed policy significantly outperforms other benchmark methods, achieving near-optimal performance even in the presence of high estimation errors in computing capacity. Zhaojun Nan, Yunchu Han, Jintao Yan, Sheng Zhou 0001, Zhisheng Niu |
IEEE Trans. Mob. Comput. | 2 |
| 2024 | CPU-Utilization-Aware Scheduling for In-Vehicle Distributed ComputingabstractWith the rapid advancement of intelligent vehicle technology, the demand for vehicular computing power is increasing. To alleviate computing loads on the on-board computer, this paper proposes an in-vehicle distributed computing system that leverages in-vehicle devices, such as smartphones and tablet computers, for cooperative task execution. Different from previous works that focused on the impact of CPU frequency management for computation offloading, we consider the scenario where the CPU frequency of devices cannot be adjusted, and study a more practical approach to make offloading and scheduling decisions by considering the CPU utilization of devices. Therefore, we first derive an analytical relationship between CPU utilization and computation latency, and then propose a CPU Utilization-Aware Scheduling (CUAS) policy to minimize response latency consisting of computation and communication latency. Simulations conducted on Simgrid show that our proposed CUAS policy can reduce the response latency by up to 20.60% compared with the benchmarks. Additionally, we established a real-world testbed to validate our system's practicality. Experimental results indicate that our proposed policy can reduce response latency by up to 20.75% compared with the benchmarks. Jintao Yan, Yunchu Han, Zhaojun Nan, Sheng Zhou 0001 |
WCNC | 2 |