Haotong Zhang 0003

dblp:316/3205 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-1729-3383ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SegRNN: Segment Recurrent Neural Network for Long-Term Time-Series Forecasting
abstract
With the proliferation of Internet of Things (IoT) applications, advanced time series forecasting techniques have become increasingly critical for managing and responding to complex temporal dynamics. However, traditional RNN-based methods have faced challenges in the Long-term Time Series Forecasting (LTSF) domain when dealing with excessively long look-back windows and forecast horizons. Consequently, the dominance in this domain has shifted towards Transformer, MLP, and CNN approaches. The substantial number of recurrent iterations are the fundamental reasons behind the limitations of RNNs in LTSF. To address these issues, we propose two novel strategies to reduce the number of iterations in RNNs for LTSF tasks: Segment-wise Iterations and Parallel Multi-step Forecasting (PMF). RNNs that combine these strategies, called SegRNN, significantly reduce the required recurrent iterations for LTSF, resulting in notable improvements in forecast accuracy and inference speed. Extensive experiments demonstrate that SegRNN not only outperforms state-of-the-art Transformer-based models but also reduces runtime and memory usage by more than 78%, making it highly suitable for resource-constrained IoT scenarios. These achievements provide strong evidence that RNNs continue to excel in LTSF tasks and encourage further exploration of this domain with more RNN-based approaches. The code is available at: https://github.com/lss-1138/SegRNN.
Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Feiyu Zhao, Ruichao Mo, Haotong Zhang 0003
IEEE Internet Things J.6
2026 WSDBS: Workflow Scheduling With Dynamic Bandwidth Slicing in Resource-Constrained Edge Computing Environment
abstract
In resource-constrained edge computing, the execution efficiency of workflow applications is significantly affected by bandwidth contention, especially during data transmissions between dependent tasks. However, existing workflow scheduling studies often struggle to optimize transmission delay effectively, whereas bandwidth slicing offers promising potential by leveraging the dynamic nature of bandwidth resources. To address this issue, we propose Workflow Scheduling with Dynamic Bandwidth Slicing (WSDBS), a novel scheduling algorithm that integrates bandwidth slicing into the workflow execution process. By introducing a dual-prediction strategy, WSDBS estimates the availability of both computational and bandwidth resources on servers, facilitating efficient task scheduling decisions under transmission uncertainty. Moreover, a novel transmission urgency metric is developed, which is derived from both link load and transmission criticality. This metric guides bandwidth slicing for the dynamic allocation of server-side bandwidth resources, ultimately alleviating contention among concurrent transmissions. Extensive experiments based on real-world Alibaba cluster traces show that WSDBS consistently improves scheduling efficiency, reducing the average makespan by 10.87%-14.44% over state-of-the-art baselines. These results validate its effectiveness in alleviating bandwidth contention and improving scheduling performance in edge computing environments.
Yuebin Huang, Weiwei Lin 0001, Fang Shi, Haotong Zhang 0003, Simon Fong 0001, Bin Wang 0048
IEEE Trans. Mob. Comput.4
2026 PAWSSP: A Two-Stage Parallelism-Aware Algorithm for Joint Workflow Scheduling and Service Placement in Edge Computing
abstract
In edge computing, workflow applications are optimally scheduled onto edge servers that are pre-equipped with the necessary services to satisfy stringent low-latency demands. However, prior research has not fully addressed the joint optimization of service placement and workflow scheduling, particularly the exploitation of task parallelism to reduce overall makespan. To address this shortcoming, we explore the combined workflow scheduling and service placement (WSP-SP) problem with the goal of minimizing the average makespan of applications. Recognizing that WSP-SP is NP-hard, we propose a two-stage, Parallelism Aware Workflow Scheduling and Service Placement strategy (PAWSSP) that minimizes the makespan with low complexity. In the first stage, a Parallelism Aware Service Placement module (PASP) is designed to adjust the service layout by allocating services with high parallelism onto distinct servers to fully leverage task-level concurrency. In the subsequent workflow scheduling stage, PAWSSP determines task priority by resolving inter-task competition and further reduces waiting times by assigning tasks to servers experiencing lower resource contention. We further extend PAWSSP to make it applicable to both offline and online scenarios. Extensive evaluations demonstrate that PAWSSP performs robustly across diverse scenarios, reducing the average makespan by 3.9%-14.7% compared to existing baselines, while maintaining modest runtime overhead.
Weiwei Lin 0001, Fang Shi, Haotong Zhang 0003, Bin Wang 0048
IEEE Trans. Serv. Comput.4
2025 Cacomp: A Cloud-Assisted Collaborative Deep Learning Compiler Framework for DNN Tasks on Edge
abstract
With the development of edge computing, DNN services have been widely deployed on edge devices. The deployment efficiency of deep learning models relies on the optimization of inference and scheduling policy. However, traditional optimization methods on edge devices still suffer from prohibitively long tuning time due to devices’ low computational power. Meanwhile, the widely used scheduling algorithm, the dominant resource fairness algorithm(DRF algorithm), struggles to maximize the efficiency of model execution on edge devices and inevitably increases average waiting time as it is not applicable in the real-time distributed computing environment. In this paper, we propose Cacomp, a distributed cloud-assisted deep learning compiler framework that features accelerating the optimization on edge devices with assistance from the cloud and a novel inference task scheduling algorithm. Our framework utilizes the tuning records from the cloud devices and proposes a two-step distillation strategy to obtain the best tuning record set for the edge device. For the scheduling process, we propose an RD-DRF algorithm to allocate inference tasks to edge devices based on dominant resource matching in real time. Extensive results show that our framework can achieve up to 2.19× improvement in the optimization time compared with other methods on edge devices. Our proposed scheduling algorithm significantly shortens the average waiting time of inference tasks by 30% and improves resource utilization by 20% on edge devices.
Weiwei Lin 0001, Jinhui Lin, Haotong Zhang 0003, Wentai Wu, Weizheng Wu, Zhetao Li, Keqin Li 0001
IEEE Trans. Computers3
2025 Container Scheduling Strategy Based on Image Layer Reuse and Sequential Arrangement in Mobile Edge Computing
abstract
In Mobile Edge Computing (MEC) scenarios, computational tasks are popularly deployed using containerization to isolate the runtime environment. To complete the execution of the task, the edge server first pulls the image, then instantiates and runs the container. Since it takes a lot of time for the edge server to download the image from the cloud, image reuse reduces the pulling latency significantly. However, the limited storage capacity of edge servers hinders image reuse. Recent works have enhanced reuse efficiency by leveraging the hierarchical structure of images and caching high-value layers. However, their efficiency remains limited due to the lack of multi-container collaboration. This paper proposes a novel container scheduling strategy based on image layer reuse and sequence arrangement (ILR-SA) for MEC scenarios, which achieves efficient scheduling by collaborating multiple containers. First, containers are greedily deployed into the edge cluster. Then, the execution sequence of containers is modeled as an optimal Hamiltonian path problem, efficiently solved by our proposed decomposition algorithm. Finally, an efficient image layer update strategy is used to achieve layer reuse. We conduct rigorous experiments to demonstrate that our proposed container scheduling strategy reduces the computational task completion time by up to 91.3% compared to existing approaches.
Haijie Wu, Weiwei Lin 0001, Haotong Zhang 0003, Fang Shi, Wangbo Shen, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Mob. Comput.3
2024 An optimal container update method for edge-cloud collaboration
abstract
Abstract Emerging computing paradigms provide field‐level service responses for users, for example, edge computing, fog computing, and MEC. Edge virtualization technologies represented by Docker can provide a platform‐independent, low‐resource‐consumption operating environment for edge service. The image‐pulling time of Docker is a crucial factor affecting the start‐up speed of edge services. The layer reuse mechanism of native Docker cannot fully utilize the duplicate data of node local images. In this paper, we propose a chunk reuse mechanism (CRM), which effectively targets node‐local duplicate data during container updates and reduces the volume of data transmission required for image building. We orchestrate the CRM process for cloud and remote‐cloud nodes to ensure that the resource overhead from container update data preparation and image reconstruction is within an acceptable range. The experimental results show that the CRM proposed in this paper can effectively utilize the node local duplicate data in the synchronous update of containers in multiple nodes, reduce the volume of data transmission, and significantly improve container update efficiency.
Haotong Zhang 0003, Weiwei Lin 0001, Shenghai Li, Zhiyan Dai, James Zijun Wang
Softw. Pract. Exp.1
2024 Generation of high-order random key matrix for Hill Cipher encryption using the modular multiplicative inverse of triangular matrices
Yuehong Chen, Haotong Zhang 0003, Dongdong Li 0002, Weiwei Lin 0001
Wirel. Networks3
2022 UltraCDC:A Fast and Stable Content-Defined Chunking Algorithm for Deduplication-based Backup Storage Systems
abstract
Content-Defined Chunking(CDC) is the key stage of data deduplication since it has a significant impact on deduplication system’s throughput and deduplication efficiency. However, existing CDC algorithms suffer from high computation overhead, weak stability, and poor ability to handle low-entropy strings. In this paper, we propose UltraCDC, a fast and stable, high-efficiency deal with low-entropy strings, CDC algorithm for deduplication-based storage systems. There are four key techniques behind UltraCDC, namely, rolling compute boundary conditions, skipping sub-minimum chunk size, normalized chunking, and jumping to detect low-entropy strings. Using a sliding window to rolling compute boundary conditions not only accelerates the chunking stage but also makes it more resistant to boundary shift, the two techniques of skipping sub-minimum chunk size and normalized chunking can complement each other to speed up chunking without sacrificing deduplication ratio too much, and the jumping detection can detect more low-entropy strings than AE-opt2 without affecting chunking speed. We implemented UltraCDC in Destor, and the experimental results show that using the above four techniques, chunking speed is 1.5–10× faster than the state-of-the-art CDC approaches, while deduplication ratio is comparable or even higher than the classic Rabin-base CDC. In terms of the capability to detect low-entropy strings, UltraCDC is a CDC approach with the highest ability to detect low-entropy strings, 102× and 2× higher than Rabin-based CDC and AE-opt2, respectively.
Peng Zhou 0005, Zhenyu Wang 0001, Wen Xia, Haotong Zhang 0003
IPCCC4
2022 Exploring the Internet of Things sequence-structure detection and supertask network generation of temporal-spatial-based graph convolutional neural network
Deyu Qi 0001, Haotong Zhang 0003
J. Supercomput.4