Zhongming Fu

dblp:206/6963 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-3041-6990ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GraphDelta: A distributed incremental framework for efficient dynamic graph computing in edge intelligence
Mengsi He, Zhongming Fu, Zhuo Tang
J. Syst. Archit.2
2025 Optimization of fault tolerance for iterative graph algorithm in spark GraphX based on high performance computing cluster
Mengsi He, Zhongming Fu
CCF Trans. High Perform. Comput.2
2025 UM-Mamba: An efficient U-network with medical visual state space for medical image segmentation
Hejian Chen, Zhongming Fu
Comput. Vis. Image Underst.3
2025 Optimizing Data Locality by Integrating Intermediate Data Partitioning and Reduce Task Scheduling in Spark Framework
abstract
Data locality is crucial for distributed computing systems (e.g., Spark and Hadoop), which is the main factor considered in the task scheduling. Simultaneously, the effects of data locality on reduce tasks are determined by the intermediate data partitioning. While suffering from the problem of data skew, the existing intermediate data partitioning methods only achieves load balancing for reduce tasks. To address the problem, this paper optimizes the data locality for reduce tasks by integrating intermediate data partitioning and task scheduling in Spark framework. First, it presents a distribution skew model to divide the key clusters into skewed and non-skewed distribution. Then, a data locality and load balancing-aware intermediate data partitioning method is proposed, where a priority allocation strategy for the key clusters with skewed distribution is presented, and a balanced allocation strategy for the key clusters with non-skewed distribution is presented. Finally, it proposes a data locality-aware reduce task scheduling algorithm, where an online self-adaptive NARX (nonlinear autoregressive with external input) model is developed to predict the idle time of node. It can ensure that the delayed scheduling decision made can complete the data transmission of reduce tasks earlier. We implement our proposals in Spark-3.5.1 and evaluate the performance using several representative benchmarks. Experimental results indicate that the proposed method and algorithm can reduce the job/application running time by approximately 4% to 46% and decrease the total volume of data transmission by approximately 8% to 54%.
Mengsi He, Zhongming Fu, Zhuo Tang
IEEE Trans. Parallel Distributed Syst.2
2024 Improving Data Locality of Tasks by Executor Allocation in Spark Computing Environment
abstract
The concept of data locality is crucial for distributed systems (e.g., Spark and Hadoop) to process Big Data. Most of the existing research optimized the data locality from the aspect of task scheduling. However, as the execution container of Spark's tasks, the executor launched on different nodes can directly affect the data locality achieved by the tasks. This paper tries to improve the data locality of tasks by executor allocation in Spark framework. First, because of different communication modes at stages, we separately model the communication cost of tasks for transferring input data to the executors. Then formalize an optimal executor allocation problem to minimize the total communication cost of transferring all input data. This problem is proven to be NP-hard. Finally, we present a greed dropping heuristic algorithm to provide solution to the executor allocation problem. Our proposals are implemented in Spark-3.4.0 and its performance is evaluated through representative micro-benchmarks (i.e.,WordCount,Join,Sort) and macro-benchmarks (i.e.,PageRankandLDA). Extensive experiments show that the proposed executor allocation strategy can decrease the network traffic and data access time by improving the data locality during the task scheduling. Its performance benefits are particularly significant for iterative applications.
Zhongming Fu, Mengsi He, Zhuo Tang
IEEE Trans. Cloud Comput.1
2023 PCU-Net: An enhanced U-network by combining PPM and CBAM for Medical Image Segmentation
abstract
U-network is a kind of full convolutional neural network, which is widely used in the field of medical image segmentation. However, it still has the problem of small targets being unable to be segmented and resulting in unsatisfactory segmentation effect. In order to solve this problem, this paper proposes an enhanced U-network by combining pyramid pooling module (PPM) and convolutional block attention module (CBAM). It’s whole network is U-Net architecture, where PPM with bin sizes of 1×1, 2×2 and 3×3 are used in the downsampling part of the network, which can extract input image features of various dimensions. And CBAM is used in the downsampling part of the network, which combines convolution and attention mechanism, and can pay attention to the image from two aspects of space and channel to improve the segmentation ability of the network. Experimental results show that our network outperforms traditional Ushaped segmentation networks by 30% to 40% in metrics IoU, MAE, and Dice, respectively.
Hejian Chen, Zhongming Fu, Mengsi He, Jiayi Deng, Zhuo Tang
ICPADS2
2022 Cross-domain Resemblance Detection based on Meta-learning for Cloud Storage
abstract
Recently, cloud storage has been widely used in our daily life. And there are lots of redundancy among these outsourced data. Conventional deduplication technology efficiently splits these data at the chunk level and removes the duplicate chunks to save the network bandwidth and improve the cloud storage utility. But it ignores the redundancy among similar chunks. Resemblance detection has recently become a hot issue with detecting these redundant parts among similar data. CARD, the state-of-the-art work, can efficiently and effectively remove these redundancies by introducing the neural network with resemblance detection. However, the source domain of the CARD model may have an explicitly different input distribution. The cloud cannot deal with the possible future domain data based on CARD design. This cross-domain setting may serials degrades the performance of CARD. To overcome this problem, we propose a cross-domain resemblance detection scheme called MetaContext. Integrating the chunk-context aware model and the learn-to-learn idea can produce a more robust chunk feature than CARD. As a byproduct, it also outperforms the CARD in speed. Finally, we implement the MetaContext and conduct serial experiments on real workloads. The results show that our method can efficiently and effectively detect and remove the redundancy among similar data.
Baisong Li, Ruixuan Li 0001, Weijun Xiao, Zhongming Fu, Xuming Ye, Renjiao Duan, Zhiyong Xu 0003
IPCCC5
2022 Optimizing Speculative Execution in Spark Heterogeneous Environments
abstract
The execution time of a stage is extended by a few slow running tasks in Spark computing environments. To tackle this so-called straggler problem, Spark adopts speculative execution mechanism under which the scheduler speculatively launch additional backup for the straggler with the hope to complete early. However, due to the characteristics of tasks and the complexity of runtime environments, the Spark original speculative execution strategy and its improved versions cannot deal with this problem effectively. In this paper, we propose a novel strategy called ETWR to improve the efficiency of speculative execution in Spark. We consider the heterogeneous environment when we around to tackle the three key points of speculative execution: straggler identification, backup node selection and effectiveness guarantee. Based on the task type classification, first, we divide the task into sub-phases and use both the process speed and progress rate within a phase to find the straggler promptly. Second, we use the Locally Weighted Regression model to estimate the execution time of the task, which will be used to calculate the task's remaining time and backup time. Third, we present iMCP model to guarantee the effectiveness of speculative tasks, which can additionally keep load balancing for nodes. Finally, the factors of fast node and better location are considered when choosing proper backup nodes. Extensive experiments show that ETWR can reduce the job execution time by 23.8 percent, and improve the cluster throughput by 33.2 percent compared with Spark-2.2.0.
Zhongming Fu, Zhuo Tang
IEEE Trans. Cloud Comput.1
2022 IncGraph: An Improved Distributed Incremental Graph Computing Model and Framework Based on Spark GraphX
abstract
The excavated information will become obsolete when the data changes in dynamic graphs. To compute the up-to-date results, the graph algorithm has to re-compute the entire data from scratch, which will consume huge computation time and resources. To reduce the cost of such calculations, this paper proposes a model called IncGraph to support incremental iterative computation over dynamic graphs. Different from the way of traditional iteration, IncGraph executes the graph algorithm through reusing the results of the previous graph and performs computation on the part of the graph that has changed. IncGraph has two critical components: (1) an incremental iterative computation model that consists of two steps: an incremental step to calculate the results on the changed vertices of the graph, and a merge step to calculate the results on the entire graph by using the results of the previous graph and the incremental step; and (2) an incremental update method to accelerate the iterative process within the iterative graph algorithm. We implement IncGraph model on GraphX and evaluate its performance by using several representative iterative graph algorithms: PageRank, Connected components, and Single Source Shortest Path. The results show that compared with the traditional iteration, when adding the 100k of vertices in different size data sets, the performance optimization ratio of IncGraph is 31.79 percent averagely, and 50.2 percent maximum; and when the percentage of added vertices varied from 0.01 to 10 percent in different data sets, the performance optimization ratio of IncGraph varied from 19.9 to 66.1 percent. Moreover, the result errors of IncGraph is small and can be neglected.
Zhuo Tang, Mengsi He, Zhongming Fu, Li Yang 0012
IEEE Trans. Knowl. Data Eng.3
2021 Optimizing Data Locality by Executor Allocation in Reduce Stage for Spark Framework
Zhongming Fu, Mengsi He, Zhuo Tang, Yang Zhang 0026
PDCAT1
2020 ImRP: A Predictive Partition Method for Data Skew Alleviation in Spark Streaming Environment
Zhongming Fu, Zhuo Tang, Li Yang 0012, Kenli Li 0001, Keqin Li 0001
Parallel Comput.1
2020 An Optimal Locality-Aware Task Scheduling Algorithm Based on Bipartite Graph Modelling for Spark Applications
abstract
In the distributed computing framework of Spark, cross-node/rack data transfer produced by map tasks and reduce tasks are common problems resulting in performance degradation, such as prolonging of entire execution time and network congestion. To address these problems, this article utilizes the bipartite graph modelling to propose an optimal locality-aware task scheduling algorithm. By considering global optimality, the algorithm can generate the optimal scheduling solution for both the map tasks and the reduce tasks for data locality. Because of the different communication modes, this article uses a unified graph to model the map task scheduling and the reduce task scheduling respectively. Then, by calculating the communication cost matrix of tasks, we formulate an optimal task scheduling scheme to minimize overall communication cost and transform the problem as the well-known graph problem: minimum weighted bipartite matching (MWBM), which can be resolved by Kuhn-Munkres algorithm. In addition, this article proposes a locality-aware executor allocation strategy to improve the data locality further. We implement our algorithm and strategy in Spark-2.4.1 and evaluate its performance using several representative micro-benchmarks, macro-benchmarks, and HiBench benchmark suite. The experimental results verify that by reducing the network traffic and access latency, the proposed algorithm can improve the job performance substantially compared to some other task scheduling algorithms.
Zhongming Fu, Zhuo Tang, Li Yang 0012, Chubo Liu
IEEE Trans. Parallel Distributed Syst.1
2017 A Parallel Conditional Random Fields Model Based on Spark Computing Environment
Zhuo Tang, Zhongming Fu, Zherong Gong, Kenli Li 0001, Keqin Li 0001
J. Grid Comput.2