Haijie Wu

dblp:60/8831 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MixloadSched: An interference aware and energy efficient scheduling method for mixed workloads in cloud data centers
Jianzhuo Li, Weiwei Lin 0001, Haijie Wu, Duanyang Du, Keqin Li 0001
Future Gener. Comput. Syst.3
2026 Concord: A GPU cluster scheduler with enhanced interference profiling and asymmetric job packing
Weiwei Lin 0001, Haijie Wu, Jidong Zhai, Keqin Li 0001
Future Gener. Comput. Syst.3
2026 CAGO-ECIL: Cloud-Assisted Genetic Optimization for Edge-Class Incremental Learning with training acceleration
Huayue Zeng, Wangbo Shen, Haijie Wu, Weiwei Lin 0001, C. L. Philip Chen
Future Gener. Comput. Syst.3
2026 BlockEdge: A Hybrid Blockchain Framework for Secure and Efficient Collaboration in EEC Environments
abstract
In End-Edge-Cloud (EEC) computing environments, the diversity of devices often requires cloud-trained models to be adapted for end/edge devices, complicating decentralized project management. To address this, end/edge devices are increasingly using local model sharing instead of traditional cloud solutions. Popular platforms like GitHub and DockerHub lack the necessary data authenticity and security for high-stakes applications. While blockchain can ensure secure data sharing, permissioned blockchains struggle with the dynamic nature of EEC devices. To solve this, we propose BlockEdge, a hybrid blockchain architecture combining a permissioned blockchain with Practical Byzantine Fault Tolerance (PBFT) for cloud-based data management and a permissionless blockchain with Proof of Work (PoW) for decentralized model sharing at the end/edge. We enhance the PoW process with a dynamic mining algorithm and a lazy-loading Merkle tree structure, improving energy efficiency and computational performance. Experimental results show that BlockEdge reduces energy consumption by over 50% and cuts data update time by 91.73%, effectively addressing the energy and time inefficiencies of mainstream consensus mechanisms.
Wangbo Shen, Weiwei Lin 0001, Tiansheng Huang, Mian Guo, Haijie Wu
ACM Trans. Internet Techn.6
2025 Automatically Surfacing Opportunities for Improvements In Internet-Scale Applications
abstract
Modern Internet services generate massive volumes of observability data, yet identifying opportunities for business performance improvements remains elusive. In many cases, such insights manifest only within sub-populations defined by derived attributes that cannot be predefined, might evolve over time, and often cannot be exhaustively enumerated ahead of time. Unfortunately, existing commercial and research systems fall short in one or more aspects of generating such improvement opportunities: expressiveness, automation, and scalability. We present a vision for automatically surfacing opportunities for improvements to tackle these seemingly conflicting and intractable requirements. We highlight the early promise from a proof-of-concept system, showing evaluation on three real-world services and discuss open challenges for future work.
Vipul Harsh, Sayan Sinha, Henry Milner, Haijie Wu, B. Aditya Prakash, Vyas Sekar, Hui Zhang 0001
HotNets4
2025 Temporal Query Network for Efficient Multivariate Time Series Forecasting
abstract
Sufficiently modeling the correlations among variables (aka channels) is crucial for achieving accurate multivariate time series forecasting (MTSF). In this paper, we propose a novel technique called Temporal Query (TQ) to more effectively capture multivariate correlations, thereby improving model performance in MTSF tasks. Technically, the TQ technique employs periodically shifted learnable vectors as queries in the attention mechanism to capture global inter-variable patterns, while the keys and values are derived from the raw input data to encode local, sample-level correlations. Building upon the TQ technique, we develop a simple yet efficient model named Temporal Query Network (TQNet), which employs only a single-layer attention mechanism and a lightweight multi-layer perceptron (MLP). Extensive experiments demonstrate that TQNet learns more robust multivariate correlations, achieving state-of-the-art forecasting accuracy across 12 challenging real-world datasets. Furthermore, TQNet achieves high efficiency comparable to linear-based methods even on high-dimensional datasets, balancing performance and computational cost. The code is available at: https://github.com/ACAT-SCUT/TQNet.
Shengsheng Lin, Haojun Chen, Haijie Wu, Chunyun Qiu
ICML3
2025 Adaptive Incremental Broad Learning System Based on Interval Type-2 Fuzzy Set With Automatic Determination of Hyperparameters
abstract
The fuzzy broad learning system (FBLS) has received increasing attention due to its ability to quickly train from broad learning systems (BLS) and interpretability with fuzzy inference. However, the randomness of BLS brings instability to the training performance of the model, so the hyperparameters of the model are crucial for its performance. Currently, many FBLS use grid search to determine hyperparameters. However, grid search brings longer search time and the parameters obtained have randomness, which may not necessarily be the optimal hyperparameters. In response to these challenges, this paper proposes a fuzzy broad learning system with automatic determination of hyperparameters (ADHFBLS). We construct a novel FBLS based on the interval type-2 fuzzy set and design an incremental learning algorithm for rules and enhancement nodes to support rapid model expansion. Meanwhile, a heuristic hyperparameter automatic optimization algorithm is designed to overcome the randomness and long optimization time of grid search. Experiments have shown that ADHFBLS has higher accuracy and shorter model tuning time compared to some state-of-the-art models based on FBLS.
Haijie Wu, Weiwei Lin 0001, Yuehong Chen, Fang Shi, Wangbo Shen, C. L. Philip Chen
IEEE Trans. Fuzzy Syst.1
2025 Container Scheduling Strategy Based on Image Layer Reuse and Sequential Arrangement in Mobile Edge Computing
abstract
In Mobile Edge Computing (MEC) scenarios, computational tasks are popularly deployed using containerization to isolate the runtime environment. To complete the execution of the task, the edge server first pulls the image, then instantiates and runs the container. Since it takes a lot of time for the edge server to download the image from the cloud, image reuse reduces the pulling latency significantly. However, the limited storage capacity of edge servers hinders image reuse. Recent works have enhanced reuse efficiency by leveraging the hierarchical structure of images and caching high-value layers. However, their efficiency remains limited due to the lack of multi-container collaboration. This paper proposes a novel container scheduling strategy based on image layer reuse and sequence arrangement (ILR-SA) for MEC scenarios, which achieves efficient scheduling by collaborating multiple containers. First, containers are greedily deployed into the edge cluster. Then, the execution sequence of containers is modeled as an optimal Hamiltonian path problem, efficiently solved by our proposed decomposition algorithm. Finally, an efficient image layer update strategy is used to achieve layer reuse. We conduct rigorous experiments to demonstrate that our proposed container scheduling strategy reduces the computational task completion time by up to 91.3% compared to existing approaches.
Haijie Wu, Weiwei Lin 0001, Haotong Zhang 0003, Fang Shi, Wangbo Shen, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Mob. Comput.1
2025 End-Edge-Cloud Heterogeneous Resources Scheduling Method Based on RNN and Particle Swarm Optimization
abstract
Task scheduling in cloud computing is a challenging but crucial task for ensuring service quality and load balance. Mainstream scheduling algorithms, such as heuristic algorithms and reinforcement learning, have made progress in this area. However, online task scheduling algorithms, such as reinforcement learning, can pose computational challenges in scenarios with limited computational power and heterogeneous resources. Heuristic algorithms, which are more suitable for offline scheduling where the types and quantities of tasks are known in advance, also require substantial computational resources for online scheduling. In this work, we propose the end-edge-cloud (EEC) heterogeneous resources scheduling method (EHRSM) based on a recurrent neural network (RNN) model and particle swarm optimization (PSO). EHRSM uses an RNN model trained on a dataset generated by dynamic programming to recognize and cache online tasks, efficiently transforming online task scheduling into offline scheduling. Additionally, a PSO algorithm with Cantor expansion (CE) for coding optimization is used to complete the offline scheduling. Experimental results show that the method is effective in converting online scheduling to offline scheduling, reducing the average task completion time and waiting time. Compared with existing online scheduling methods, EHRSM reduces task completion time by up to 48.24%.
Haijie Wu, Wangbo Shen, Weiwei Lin 0001, Wei Li 0058, Keqin Li 0001
IEEE Trans. Netw. Serv. Manag.1
2025 Kairos: Deterministic Scheduling Enhanced by User Collaboration for Deep Learning Workloads
abstract
As deep learning (DL) workloads scale in complexity and volume, ensuring predictable job queuing times has become a critical challenge for data centers. Existing scheduling solutions primarily focus on minimizing tardiness or job completion times (JCT), often neglecting the need for deterministic queuing, particularly in dynamic and preemptive environments. This paper introducesKairos, a preemption-based scheduling framework enhanced by user collaboration to address these gaps.Kairoscombines adivide-and-conquerstrategy—segmenting jobs into sequential units with adaptive priorities—and a user-collaborative mechanism for better duration estimation. By leveraging real-time feedback from resource contention and queuing delays,Kairosminimizes a novel metric, theQueue inStability Index(QSI), achieving significant improvements in queuing predictability while maintaining competitive JCT. Experimental results demonstrate thatKairosreduces QSI by over 99.8% compared to state-of-the-art deadline-aware baselines, offering robust performance for diverse DL workloads.
Weiwei Lin 0001, Ruichao Mo, Guozhi Liu, Haijie Wu, Shengjun Tang
IEEE Trans. Parallel Distributed Syst.5
2025 MCG-Sched: Multi-Cluster GPU Scheduling for Resource Fragmentation Reduction and Load Balancing
abstract
Since the rapid development of deep learning (DL) technology, large-scale GPU clusters receive a large number of DL workloads daily. To speed up the completion time, the workloads usually occupy several GPUs on a server. However, workload scheduling inevitably generates resource fragmentation, which results in many scattered GPU resources being unavailable. Existing works address improving resource utilization by reducing GPU resource fragmentation, while they focus on resource scheduling for a single cluster and ignore multiple clusters. Multi-cluster scenarios, such as virtual clusters and geo-distributed clusters, require load balancing to avoid some clusters exhausting resources while some clusters are idle while improving resource utilization, which is not well addressed by existing works. In this paper, we propose MCG-Sched, a scheduling strategy to reduce resource fragmentation in multiple GPU clusters while maintaining load balancing among clusters. MCG-Sched measures the fragmented resources with the distribution of workload demands and uses a scheme that minimizes fragmentation in workload scheduling. Meanwhile, MCG-Sched achieves balanced load scheduling across clusters through the load balancing index. MCG-Sched senses the workload requests in the waiting queue, and prioritizes the workloads by combining fragmentation measurement and load balancing index to maximize resource utilization and load balancing during load peak. Our experiments show that MCG-Sched reduces unallocated GPUs up to 1.45× and workload waiting time by more than 40% compared to existing fragmentation-aware methods and achieves effective load balancing.
Haijie Wu, Xiaoxuan Luo, Wangbo Shen, Weiwei Lin 0001
IEEE Trans. Parallel Distributed Syst.1
2025 Prediction of Heterogeneous Device Task Runtime Based on Edge Server-Oriented Deep Neuro-Fuzzy System
abstract
Predicting the runtime of tasks is of great significance as it can help users better understand the future runtime consumption of the tasks and make decisions for their heterogeneous devices, or be applied to task scheduling. Learning features from user task history data for predicting task runtime is a mainstream method. However, this method faces many challenges when applied to edge intelligence. In the Big Data era, user devices and data features are constantly evolving, necessitating frequent model retrains. Meanwhile, the noisy data from these devices requires robust methods for valuable insight extraction. In this paper, we propose an edge server-oriented deep neuro-fuzzy system (ESODNFS) that can be trained and inferred on edge servers, for providing users with task runtime prediction services. We divided the dataset and trained it on multiple improved adaptive-network-based fuzzy inference system units (ANFISU), and finally conducted joint training on a deep neural network (DNN). By partitioning the dataset, we reduced the number of parameters for each ANFISU, and at the same time, multiple units can be trained in parallel, supporting fast training and iteration. Additionally, the application of fuzzy inference can effectively learn the features in noisy data and make accurate predictions. The experimental results show that ESODNFS can accurately predict the runtime of real tasks. Compared with other DNN and DNFS, it can achieve good prediction results while reducing training time by over 35%.
Haijie Wu, Weiwei Lin 0001, Wangbo Shen, Xiumin Wang 0005, C. L. Philip Chen, Keqin Li 0001
IEEE Trans. Serv. Comput.1
2013 Energy-efficient data redistribution in sensor networks
abstract
We address the energy-efficient data redistribution problem in data-intensive sensor networks (DISNs). In a DISN, a large volume of data gets generated, which is first stored in the network and is later collected for further analysis when the next uploading opportunity arises. The key concern in DISNs is to be able to redistribute the data from data-generating nodes into the network under limited storage and energy constraints at the sensor nodes. We formulate the data redistribution problem where the objective is to minimize the total energy consumption during this process while guaranteeing full utilization of the distributed storage capacity in the DISNs. We show that the problem is APX-hard for arbitrary data sizes; therefore, a polynomial time approximation algorithm is unlikely. For unit data sizes, we show that the problem is equivalent to the minimum cost flow problem, which can be solved optimally. However, the optimal solution's centralized nature makes it unsuitable for large-scale distributed sensor networks. Thus, we design a distributed algorithm for the data redistribution problem which performs very close to the optimal, and compare its performance with various intuitive heuristics. The distributed algorithm relies on potential function-based computations, incurs limited message and computational overhead at both the sensor nodes and data generator nodes, and is easily implementable in a distributed manner. We analytically study the convergence and performance of the proposed algorithm and demonstrate its near-optimal performance and scalability under various network scenarios. In addition, we implement the distributed algorithm in TinyOS, evaluate it using TOSSIM simulator, and show that it outperforms EnviroStore, the only existing scheme for data redistribution in sensor networks, in both solution quality and message overhead. Finally, we extend the proposed algorithm to avoid disproportionate energy consumption at different sensor nodes without compromising the solution quality.
Bin Tang 0004, Neeraj Jaggi, Haijie Wu, Rohini Kurkal
ACM Trans. Sens. Networks3
2012 Secure neighbor discovery and wormhole localization in mobile ad hoc networks
Radu Stoleru, Haijie Wu, Harsha Chenji
Ad Hoc Networks2
2011 Secure Neighbor Discovery in Mobile Ad Hoc Networks
abstract
Neighbor discovery is an important part of many protocols for wireless adhoc networks, including localization and routing. When neighbor discovery fails, communications and protocols performance deteriorate. In networks affected by relay attacks, also known as wormholes, the failure may be more subtle. The wormhole may selectively deny or degrade communications. IIn this paper we present Mobile Secure Neighbor Discovery (MSND), which offers a measure of protection against wormholes by allowing participating mobile nodes to securely determine if they are neighbors. To the best of our knowledge, this work is the first to secure neighbor discovery in mobile adhoc networks. MSND leverages concepts of graph rigidity for wormhole detection.We prove security properties of our protocol, and demonstrate its effectiveness through extensive simulations and a real system evaluation employing Epic motes and iRobot robots.
Radu Stoleru, Haijie Wu, Harsha Chenji
MASS2
2011 Geographic routing with constant stretch in large scale sensor networks with holes
abstract
Geographic routing is well suited for large scale sensor networks deployments, because the per node state it maintains is independent of the network size. However, due to the “local minimum” caused by holes/obstacles, the path stretch of geographic routing can grow as O(c2), where c is the length of the optimal path. Recently, VIGOR, a geographic routing protocol based on the visibility graph, shows that a constant path stretch can be achieved. This, however, is possible with increased overhead. To address this issue, we propose GOAL (Geometric Routing using Abstracted Holes), a routing protocol that provably achieves a constant path stretch, with lower message, space and computational overhead. We develop a novel distributed convex hull construction (DCC) algorithm that compactly describes holes. This compact representation of a hole is leveraged by nodes to make locally optimal routing decisions. Our theoretical analysis proves the constant stretch property and average stretch of GOAL. Through extensive simulations and a hardware implementation, we demonstrate the effectiveness of GOAL and its feasibility for large-scale sensor networks. In our network settings, GOAL reduces the energy consumption by up to 32%, routing table size by an order of magnitude, when compared with VIGOR.
Myounggyu Won, Radu Stoleru, Haijie Wu
WiMob3
2010 Energy-efficient data redistribution in sensor networks
abstract
We address the energy-efficient data redistribution problem in data intensive sensor networks (DISNs). The key question in sensor networks with large volumes of sensory data is how to redistribute the data efficiently under limited storage and energy constraints at the sensor nodes. The goal of the redistribution scheme is to minimize the energy consumption during the process, while guaranteeing full utilization of the distributed storage capacity in the DISNs. We formulate this problem as a minimum cost flow problem, which can be solved optimally. However, the optimal solution's centralized nature makes it unsuitable for large-scale distributed sensor networks. We thus design a distributed algorithm for the data redistribution problem which performs very close to the optimal, and compare its performance with various intuitive heuristics. Our proposed algorithm relies on potential function based computations, incurs limited message and computational overhead at both the sensor nodes and data generator nodes, and is easily implementable in a distributed manner. We analytically show the convergence of our algorithm, and demonstrate its near-optimal performance and scalability under various network scenarios considered. Finally, we implement our distributed algorithm in TinyOS and evaluate it using TOSSIM simulator, and show that it outperforms EnviroStore, the only existing scheme for data redistribution in sensor networks, in both solution quality and overhead messages.
Bin Tang 0004, Neeraj Jaggi, Haijie Wu, Rohini Kurkal
MASS3