Zhilan Huang

dblp:13/4665 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0003-5531-0815ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Heterogeneous Device Cluster-Aware Deployment Optimization for LLM Joint Training
Yukai Tan, Mengyu Sun, Zhilan Huang, Yasen Wang, Zeya Zhu
INFOCOM4
2025 Latency-aware scheduling for data-oriented service requests in collaborative IoT-edge-cloud networks
Mengyu Sun, Shuo Quan, Xuliang Wang, Zhilan Huang
Future Gener. Comput. Syst.4
2025 Greening Edge AI: Optimizing Inference Accuracy and Reducing Carbon Emissions With Renewable Energy
abstract
Machine-learning (ML) inference services are vital for artificial intelligence (AI) applications on edge devices. However, there are challenges to balance between maintaining high inference accuracy and minimizing carbon emissions. To foster green AI services, we introduce a novel framework that taps into green energy sources, empowering Internet of Things (IoT) devices to execute inference tasks at the network’s edge. In the context of multitask concurrent scheduling, our methodology seeks to simultaneously optimize both the inference accuracy and carbon footprint of all inference tasks. To support this endeavor, we propose a unique incentive constraint that ensures equitable allocation of computing resources for IoT devices equipped with energy harvesting modules, thereby promoting efficient collaboration. Furthermore, recognizing the system’s dynamic nature and fluctuating availability of green energy, we introduce the carbon-aware online inference task offloading algorithm (OTOA). This algorithm, anchored in the Lyapunov optimization method, leverages a graph-based offloading strategy, amalgamating elements from the two-sided matching method, to derive an approximate optimal solution using real-time data. Our meticulous theoretical evaluations and in-depth experiments with real-world datasets affirm OTOA’s superiority over existing approaches.
Huirong Ma, Zhilan Huang, Deji Fu, Chunyu Shi
IEEE Internet Things J.2
2023 Online Scheduling of CPU-NPU Co-inference for Edge AI Tasks
abstract
Edge AI is an emerging paradigm that leverages edge computing to pave the last mile delivery of artificial intelligence. To satisfy the stringent timeliness and energy-efficiency requirements of emerging edge AI tasks, specialized AI accelerator of Neural Processing Units (NPU) have been widely equipped by edge nodes. Compared to the traditional centralized processing units (CPU), NPU has better performance and energy-efficiency. However, these benefits come at the cost of reduced inference accuracy. As a result, existing coarse-grained scheduling mechanisms that schedule a whole DNN task to either the CPU or NPU are unable to make the best use of NPU. To address this issue, we propose an online NPU-CPU co-inference scheduling mechanism to schedule the DNN task at the fine-grained layer level, and thus to fully utilize the performance, accuracy, and power diversities of the NPU and CPU. By applying Lyapunov optimization to schedule the network layers dynamically, our proposed online scheduling mechanism is able to ensure the real-time inference speed and cap the long-term time-averaged power consumption, while still approximately minimizes the long-term inference accuracy loss. Via rigorous theoretical analysis as well as realistic trace-driven simulations, we demonstrate the effectiveness of our proposed online scheduling mechanism.
Xiancheng Lin, Zhi Zhou 0006, Xu Chen 0004, Zhilan Huang
WCNC7
2023 Optimizing Service Redeployment in Migration-Oriented IoT Networks
abstract
The Internet of Things (IoT) paradigm has established an effective platform to promote the collaboration of resource-limited and duty-cycle IoT nodes, in order to support relative complex service requests that can hardly be achieved by any single IoT node. The functionalities of IoT nodes are typically encapsulated into IoT services, and the satisfaction of service requests is implemented as IoT service composition. Generally, IoT nodes work in turn in terms of their prespecific working cycles, and IoT network topology is constantly varied due to their active/sleeping behavior switching, causing unscheduled response latency. IoT service composition instantiation should be dynamically adjusted and partially redeployed on-demand for supporting request processing efficiently. To remedy this issue, this article proposes a migration-oriented service redeployment (MSrD) mechanism by migrating certain IoT services from their hosted IoT nodes to neighboring ones, in order to support functionally continuous availability. We formulate this problem as a game-theoretic approach, which is reduced to a potential game, where a Nash equilibrium solution is searched for optimizing this service redeployment game. Extensive experiments are conducted, and numerical results show that our MSrD mechanism is promising, compared with the state-of-art techniques, in achieving service redeployment optimization with efficient energy consumption and timely response latency simultaneously.
Mengyu Sun, Zhangbing Zhou, Xuliang Wang, Zhilan Huang
IEEE Internet Things J.5