EDBT 2026 Demo / reviewers in the wild / expert
Hongmin Geng
dblp:216/1543
· DBLP profile ↗
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-0867-175XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lazy but Efficient: Layer-Wise Task Scheduling with Lazy Pulling for Fast Serverless Inference
Zhexiong Li, Hongmin Geng, Yuepeng Li, Lin Gu 0002, Deze Zeng |
INFOCOM | 2 |
| 2025 | PASS: A Priority-based Model Assignment for Minimal Inference Time in Serverless Edge CloudabstractServerless computing is increasingly being adopted to provision various on-demand services at the edge cloud, including inference tasks based on deep neural networks (DNNs) for the Internet of Things (IoT). This approach leverages the advantages of flexible resource allocation and fine-grained resource management. However, the provisioning of on-demand inference typically requires downloading the DNN model at runtime, which can introduce significant delays. In the edge cloud with heterogeneous network connections, the inevitable model downloading time and the inter-model data transmission impose high challenges to the QoS of inference tasks. In this paper, we investigate how to jointly consider both model downloading time and communication time to minimize inference time. We first formulate this problem into a nonlinear optimization form and proved it as NP-hard. We further propose a Priority-based Model Assignment (PASS) algorithm in polynomial time and trace-driven experimental results show that it reduces the average inference time by 23.6% compared to existing state-of-the-art solutions. Fangshuai Zhu, Deze Zeng, Lin Gu 0002, Yuepeng Li, Hongmin Geng |
ICCCN | 5 |
| 2025 | EdgePrios: Joint Scheduling of Initialization and Execution for Serverless Inference Acceleration in Edge CloudabstractThe rapid deployment of intelligent applications on edge cloud calls for efficient and responsive DNN inference, especially under the burst scenarios of inference request. Serverless inference offers a promising solution by enabling rapid and flexible activation of inference tasks to cope with peak request, but its achievable performance is highly influenced by the initialization overhead. Existing studies on inference acceleration mainly focuses on execution optimization, they usually overlook the fact that inference performance also heavily depends on the initialization. In this paper, we propose EdgePrios, a novel priority-based scheduling mechanism that jointly optimizes initialization and execution phases for serverless inference acceleration. EdgePrios dynamically prioritizes tasks by considering workloads, dependency relationships, and the current status of available resources. It enables precise assignment of tasks to computing resources while minimizing overall inference time in edge cloud. Extensive trace-driven evaluations demonstrate the efficiency of EdgePrios as it outperforms state-of-the-art methods, achieving 15.6%-38.8% reduction in inference time under varying resource configurations, network bandwidths, and application topologies. Hongmin Geng, Yuepeng Li, Lin Gu 0002, Deze Zeng |
IEEE Internet Things J. | 1 |
| 2025 | PASS: A Priority-Based Model Assignment for Intelligent Application Acceleration in Edge CloudabstractThanks to the fine-grained resource management capabilities, serverless computing has been extended to edge cloud environments to support diverse Artificial Intelligence of Things (AIoT) applications, particularly those involving complex workflows of interdependent deep neural network (DNN) inference tasks. However, the inherent on-demand provisioning nature of serverless computing imposes the fact that, in serverless inference processes, the DNN models are typically maintained in the remote storage cluster and retrieved as needed. This inevitably incurs substantial latency overhead, particularly in resource-constrained edge cloud. In this paper, we investigate how to accelerate the AI application with joint consideration of both the model downloading time and intermediate data transmission time. We first formulate this problem into a nonlinear optimization form and prove it as NP-hard. We further propose a Priority-Based Model Assignment (PASS) algorithm and theoretically analyze its upper bound. The trace-driven experimental results demonstrate that our proposed algorithm outperforms other sate-of-art solutions and reduces the average application completion time by 23.6%. Yuepeng Li, Deze Zeng, Lin Gu 0002, Fangshuai Zhu, Hongmin Geng |
IEEE Internet Things J. | 5 |
| 2025 | Layer Redundancy Aware DNN Model Repository Planning for Fast Model Download in Edge CloudabstractThe booming development of artificial intelligence (AI) applications has greatly promoted edge intelligence technology. To support latency-sensitive Deep Neural Network (DNN) based applications, the integration of serverless inference paradigm into edge intelligence has become a widely recognized solution. However, the long DNN model downloading time from central clouds to edge servers hinders inference performance, and asks for establishing model repository within the edge cloud. This paper first identifies the inherent layer redundancy in DNN models, which is potentially beneficial to improve the storage efficiency of the model repository in the edge cloud. However, how to exploit the layer redundancy feature and allocate the DNN layers across different edge servers with capacitated storage resources to reduce the model downloading time remains challenging. To address this issue, we first formulate this problem in Quadratic Integer Programming (QIP) form, based on which a randomized rounding layer redundancy aware DNN model storage planning strategy is proposed. Our approach significantly reduces model downloading time by up to 63% compared to state-of-the-art methods, as demonstrated through extensive trace-driven experiments. Hongmin Geng, Yuepeng Li, Lin Gu 0002, Deze Zeng |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | PLAYS: Minimizing DNN Inference Latency in Serverless Edge Cloud for Artificial Intelligence of ThingsabstractThanks to the capability of fine-grained resource allocation and fast task scheduling, serverless computing has been adopted into edge cloud to accommodate various applications, e.g., deep neural network (DNN) inference for Artificial Intelligence of Things (AIoT). In serverless edge cloud, the servers are started up on-demand. However, as a container-based architecture, the inherent sequential startup feature of container imposes high affection on the DNN inference performance in serverless edge clouds. In this article, we investigate the distributed DNN inference problem in serverless edge cloud with the consideration of such characteristics, aiming to eliminate the extra container startup time cost to minimize the DNN inference latency. We formulate this problem into a nonlinear optimization form and then linearize it into an integer programming problem, which is proved as NP-hard. To tackle the computation complexity, we propose a priority-based layer scheduling (PLAYS) algorithm. Extensive experiment results verify the effectiveness and the adaptability of our PLAYS algorithm in comparison with other state-of-art algorithms under several well known DNN models. Hongmin Geng, Deze Zeng, Yuepeng Li, Lin Gu 0002, Quan Chen 0002, Peng Li 0017 |
IEEE Internet Things J. | 1 |
| 2023 | Layered Structure Aware Containerized Task Scheduling and Image Routing in Edge ComputingabstractUsing docker to encapsulate the task has been regarded as a potential way to achieve efficient task orchestration and management in edge computing. Despite the lightweight nature of containers, downloading a larger container image can still be resource-intensive, particularly in resource-constrained edge environments. Fortunately, the unique layered architecture of the container allows multiple containerized tasks to share the same layer, thereby offering an opportunity for reducing the image downloading overhead via sharing the common layers. To explore the potential of layer sharing on image downloading over-head reduction, we investigate a joint task scheduling and image routing problem in edge environment, aiming at minimizing the image downloading overhead. We first formulate the problem into an integer linear programming form, and then propose a layer-aware scheduling and routing (LSR) algorithm to tackle this problem. Finally, to evaluate the effectiveness of our proposed algorithm, we conduct a group of simulation experiments. The experimental results show that our proposed algorithm can reduce the download time by about 20% in comparison with other approaches. Hongmin Geng, Deze Zeng, Wenbing Chen, Yuepeng Li |
GLOBECOM | 1 |
| 2023 | Energy Efficient Partial Distributed Coded Computing in Edge ComputingabstractEdge computing is considered a promising computing paradigm that can mitigate energy consumption and workload of end devices through task offloading to edge servers. Albeit with high potential, edge computing is still challenged by various forms of “system noise”, e.g., node failures, system failures, and poor network conditions. To this end, distributed coded computing has been proposed for alleviating such effects by introducing redundancy into the computation. However, traditional distributed coded computing only focuses on leveraging the unreliable computing resource, and this indeed increases the risk of task non-completion within the acceptable timeframe. To address this problem, in this paper, we propose a partial distributed coded computing framework that can leverage the reliable and unreliable resources in the edge environment. We further investigate the problem of how to offload the coded subtasks for energy reduction while meeting the task tolerant latency. To tackle the computation complexity, we then propose an Iterated Greedy Algorithm. The experimental results verify the efficiency of our proposed algorithm, and it can reduce the energy consumption by 20% compared with other algorithms. Yuepeng Li, Deze Zeng, Hongmin Geng, Zaihang Yang |
GLOBECOM | 3 |
| 2023 | Layered Structure Aware Dependent Microservice Placement Toward Cost Efficient Edge CloudsabstractAlthough the containers are featured by light-weightness, it is still resource-consuming to pull and startup a large container image, especially in relatively resource-constrained edge cloud. Fortunately, Docker, as the most widely used container, provides a unique layered architecture that allows the same layer to be shared between microservices so as to lower the deployment cost. Meanwhile, it is highly desirable to deploy dependent microservices of an application together to lower the operation cost. Therefore, the balancing of microservice deployment cost and the operation cost should be considered comprehensively to achieve minimal overall cost of an on-demand application. In this paper, we first formulate this problem into a Quadratic Integer Programming form (QIP) and prove it as a NP-hard problem. We further propose a Randomized Rounding-based Microservice Deployment and Layer Pulling (RR-MDLP) algorithm with low computation complexity and guaranteed approximation ratio. Through extensive experiments, we verify the high efficiency of our algorithm by the fact that it significantly outperforms existing state-of-the-art microservice deployment strategies. Deze Zeng, Hongmin Geng, Lin Gu 0002, Zhexiong Li |
INFOCOM | 2 |
| 2022 | Performance Efficient Layer-aware DNN Inference Task Scheduling in GPU ClusterabstractGPU has been widely applied to accelerate the DNN based applications. However, single GPU is overwhelmed by the increasing computation requirement of large-scale DNN inference task. Although GPUs cluster alleviates the pressure of massive inference task, it still traps into low efficiency owing to the limited computing power and network bandwidth. Besides, the exclusiveness of GPU device may cause the straggler problem and hence long inference time. Reinforcement Learning (RL) has been widely used in such task scheduling problems. But the delayed reward during the training may slow the convergence speed or even result in non-convergence. To this end, in this paper, we design an improved reinforcement learning based algorithm, called DRM-DQL, to achieve a layer-aware DNN inference task scheduling. We first analyze and model the layer-wise inference task scheduling problem by deep Q-learning. Then, a delayed reward matching strategy is proposed for matching the global reward value to the immediate reward value, which help the algorithm to get the right experience in DNN layer scheduling. The experiment results demonstrate that our algorithm performs better than both heuristic algorithm and the vanilla DQL algorithm, and show the robustness in various network bandwidths, computing power, and DNN model structures. Hongmin Geng, Deze Zeng, Yuepeng Li |
GLOBECOM | 1 |
| 2022 | Collaborative Learning With a Multi-Branch Framework for Feature EnhancementabstractFeature representation is highly important for many computer vision tasks. A broad range of prior studies have been proposed to strengthen representation ability of architectures via built-in blocks. However, during the forward propagation, the reduction in feature map scales still leads to the lack of representation ability. In this paper, we focus on boosting the representational power of a convolutional network by the multi-branch framework that we term the BranchNet. Each branch is directly supervised by label information to enrich the hierarchy features in BranchNet. Based on this framework, we further propose a collaborative learning loss and a soft target loss to transfer knowledge from deeper layers to shallow layers. BranchNet is an efficient training framework without extra parameters introduced in inference and can be integrated in existing networks, e.g., VGG, ResNet, and DenseNet. We evaluate BranchNet on all of these models and find that our method outperforms the baseline models on the widely-used CIFAR and ImageNet datasets. In particular, on the CIFAR-100 dataset, the classification error of ResNet-164 with BranchNet decreases by 4.51 percent. We also conduct experiments on the representative computer vision tasks of instance segmentation and class activation mapping, further verifying the superiority of BranchNet over the baseline models. Models and code are available athttps://github.com/zyyupup/BranchNet/. Xiao Luan, Weihua Ou, Linghui Liu, Weisheng Li 0001, Yucheng Shu, Hongmin Geng |
IEEE Trans. Multim. | 7 |