Minghui Wu 0003

dblp:97/6864-3 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-8636-0212ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 A fine-grained task scheduling strategy for resource auto-scaling over fluctuating data streams
Yinuo Fan, Dawei Sun 0001, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.3
2026 A Popularity-Aware Discriminative Grouping Strategy in Distributed Stream Computing Systems
abstract
Stream grouping strategy plays an important role in stateful stream computing environments. Many existing grouping strategies overlook various cost factors associated with grouping while balancing stream load. To overcome this limitation, we propose Pd-Stream, a popularity-aware discriminative grouping strategy that identifies the hot keys in dynamic real-time streams and assigns them to instances with high balance and low cost. Our solution includes: (1) A stream application model is constructed, along with a skewed data stream model and a data stream grouping model. Data stream grouping optimization problems are formalized. (2) A hot key probability estimation algorithm is designed, which estimates real-time probabilities of hot keys based on their popularity within the sampling window. (3) An instance assignment algorithm is designed using dynamic routing. This algorithm determines the minimal number of candidate instances based on the probabilities of hot keys, and selects the target instance with the lowest load through a dynamic routing table. Experimental results show that Pd-Stream provides near-optimal load balancing with low memory, achieving load imbalance as low as$10^{-5}$and replication factor as low as 1.74. It outperforms state-of-the-art works, reducing latency by 27%–46% and improving throughput by 23%–52%.
Dawei Sun 0001, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
IEEE Trans. Mob. Comput.2
2025 A Prediction-Driven Collaborative Scheduling Strategy for Distributed Stream Computing Systems
abstract
Multi-objective collaborative optimization is essential for improving performance in stream computing systems. However, existing approaches often neglect the interdependencies among communication overhead, load balancing, and energy consumption, and lack predictive capabilities, resulting in delayed scheduling decisions that degrade system latency and throughput. To overcome these limitations, we propose a prediction-driven collaborative framework, named Pc-Stream, which proactively identifies overloaded compute nodes and triggers task migrations in advance. This paper presents this strategy through two key components: (1) A temperature-driven neighborhood adjustment method for task topology partitioning. This method dynamically adjusts the number of migrated tasks based on a predefined temperature. Tasks with high communication volume are batchmigrated to nodes with lower utilization rates during the hightemperature phase, and migrated individually during the lowtemperature phase. (2) A sliding window mechanism that generates multiple sub-sequences for training multiple predictive models. These models enable the system to monitor load trends and proactively migrate tasks from overloaded nodes to those with sufficient resources, thereby reducing communication costs and improving load balance. Experimental results demonstrate that, under dynamic and fluctuating data stream conditions, Pc-Stream significantly enhances overall system performance: reducing average system latency by 49.9 %, and increasing average throughput by 16.9 %.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
ICPADS1
2025 Straggler mitigation via hierarchical scheduling in elastic stream computing systems
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.1
2025 Ls-Stream: Lightening Stragglers in Join Operators for Skewed Data Stream Processing
abstract
Load imbalance can lead to the emergence of stragglers, i.e., join instances that significantly lag behind others in processing data streams. Currently, state-of-the-art solutions are capable of balancing the load between join instances to mitigate stragglers by managing hot keys and random partitioning. However, these solutions rely on either complicated routing strategies or resource-inefficient processing structures, making them susceptible to frequent changes in load between instances. Therefore, we present Ls-Stream, a data stream scheduler that aims to support dynamic workload assignment for join instances to lighten stragglers. This paper outlines our solution from the following aspects: (1) The models for partitioning, communication, matrix, and resource are developed, formalizing problems like imbalanced load between join instances and state migration costs. (2) Ls-Stream employs a two-level routing strategy for workload allocation by combining hash-based and key-based data partitioning, specifying the destination join instances for data tuples. (3) Ls-Stream also constructs a fine-grained model for minimizing the state migration cost. This allows us to make tradeoffs between data transfer overhead and migration benefits. (4) Experimental results demonstrate significant improvements made by Ls-Stream: reducing maximum system latency by 49.3% and increasing maximum throughput by more than 2x compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Keqin Li 0001, Rajkumar Buyya
IEEE Trans. Computers1
2025 A Hierarchical Near-Source Grouping Strategy for Elastic Stream Computing Systems
abstract
Effective task scheduling in stream computing systems can reduce the latency by minimizing inter-node communication. However, this approach often requires restarting tasks to change their deployment locations, resulting in significant system overhead and making it inadequate especially in dynamically changing data stream environments. To address this issue, we propose Ns-Stream, a hierarchical data scheduler that dynamically adjusts data distribution weights between near-source and off-source tasks. Our solution includes: (1) We observe that communication overhead from off-source data processing significantly impacts system latency when tasks' resources are ample. However, as the resources become limited, the computational power required by tasks becomes the key constraint on system performance. (2) During initialization scheduling, we deploy tasks with potential communication to the same node using the graph convolutional network, thus avoiding the need for runtime task scheduling. (3) We dynamically adjust data distribution weights between near-source and off-source tasks based on their computing capabilities, prioritizing local processing of data tuples (within the same worker and node) to optimize resource utilization and reduce data transmission overhead. (4) Experimental results demonstrate significant improvements made by Ns-Stream: reducing maximum system latency by 40% and increasing maximum throughput by 55% compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Rajkumar Buyya
IEEE Trans. Serv. Comput.1
2024 Elastic Scaling of Stateful Operators Over Fluctuating Data Streams
abstract
Elastic scaling of parallel operators has emerged as a powerful approach to reduce response time in stream applications with fluctuating inputs. Many state-of-the-art works focus on stateless operators and change the operator parallelism from one aspect. They often lack efficient management of operator states and overlook the costs associated with resource over-provisioning. To overcome these limitations, we introduce Es-Stream for elastic scaling of stateful operators over fluctuating data streams, which includes: 1) We observe that under-provisioning of operator parallelism leads to data pile-up, resulting in longer system latency, while over-provisioning of operator parallelism causes idle instances and additional resource consumption. 2) The Es-Stream system scales in two dimensions: the parallelism of operators and the number of resources. It dynamically adjusts operators to an optimal parallelism while scaling the resources used by the stream application. 3) When the parallelism of stateful operators changes, upstream operators backup downstream operators’ state and cache the emitted data tuples at dynamic time intervals, ensuring the operator parallelism is adjusted in a low-overhead way. 4) Experimental results demonstrate that Es-Stream provides promising performance improvements, reducing the maximum system latency by 3x and saving the maximum state recovery time by 2x, compared to existing state-of-the-art works.
Minghui Wu 0003, Dawei Sun 0001, Shang Gao 0003, Keqin Li 0001, Rajkumar Buyya
IEEE Trans. Serv. Comput.1
2023 A dynamic resource-aware endorsement strategy for improving throughput in blockchain systems
Minghui Wu 0003, Jianguo Yu, Zhangbing Zhou
Expert Syst. Appl.1
2023 A two-tier coordinated load balancing strategy over skewed data streams
Dawei Sun 0001, Minghui Wu 0003, Zhihong Yang, Atul Sajjanhar, Rajkumar Buyya
J. Supercomput.2
2022 An energy efficient and runtime-aware framework for distributed stream computing systems
Dawei Sun 0001, Yijing Cui, Minghui Wu 0003, Shang Gao 0003, Rajkumar Buyya
Future Gener. Comput. Syst.3
2022 A state lossless scheduling strategy in distributed stream computing systems
Minghui Wu 0003, Dawei Sun 0001, Yijing Cui, Shang Gao 0003, Xunyun Liu, Rajkumar Buyya
J. Netw. Comput. Appl.1