EDBT 2026 Demo / reviewers in the wild / expert
Mengsi He
dblp:262/1294
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-4985-2832ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphDelta: A distributed incremental framework for efficient dynamic graph computing in edge intelligence
Mengsi He, Zhongming Fu, Zhuo Tang |
J. Syst. Archit. | 1 |
| 2025 | Optimization of fault tolerance for iterative graph algorithm in spark GraphX based on high performance computing cluster
Mengsi He, Zhongming Fu |
CCF Trans. High Perform. Comput. | 1 |
| 2025 | Optimizing Data Locality by Integrating Intermediate Data Partitioning and Reduce Task Scheduling in Spark FrameworkabstractData locality is crucial for distributed computing systems (e.g., Spark and Hadoop), which is the main factor considered in the task scheduling. Simultaneously, the effects of data locality on reduce tasks are determined by the intermediate data partitioning. While suffering from the problem of data skew, the existing intermediate data partitioning methods only achieves load balancing for reduce tasks. To address the problem, this paper optimizes the data locality for reduce tasks by integrating intermediate data partitioning and task scheduling in Spark framework. First, it presents a distribution skew model to divide the key clusters into skewed and non-skewed distribution. Then, a data locality and load balancing-aware intermediate data partitioning method is proposed, where a priority allocation strategy for the key clusters with skewed distribution is presented, and a balanced allocation strategy for the key clusters with non-skewed distribution is presented. Finally, it proposes a data locality-aware reduce task scheduling algorithm, where an online self-adaptive NARX (nonlinear autoregressive with external input) model is developed to predict the idle time of node. It can ensure that the delayed scheduling decision made can complete the data transmission of reduce tasks earlier. We implement our proposals in Spark-3.5.1 and evaluate the performance using several representative benchmarks. Experimental results indicate that the proposed method and algorithm can reduce the job/application running time by approximately 4% to 46% and decrease the total volume of data transmission by approximately 8% to 54%. Mengsi He, Zhongming Fu, Zhuo Tang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | Improving Data Locality of Tasks by Executor Allocation in Spark Computing EnvironmentabstractThe concept of data locality is crucial for distributed systems (e.g., Spark and Hadoop) to process Big Data. Most of the existing research optimized the data locality from the aspect of task scheduling. However, as the execution container of Spark's tasks, the executor launched on different nodes can directly affect the data locality achieved by the tasks. This paper tries to improve the data locality of tasks by executor allocation in Spark framework. First, because of different communication modes at stages, we separately model the communication cost of tasks for transferring input data to the executors. Then formalize an optimal executor allocation problem to minimize the total communication cost of transferring all input data. This problem is proven to be NP-hard. Finally, we present a greed dropping heuristic algorithm to provide solution to the executor allocation problem. Our proposals are implemented in Spark-3.4.0 and its performance is evaluated through representative micro-benchmarks (i.e.,WordCount,Join,Sort) and macro-benchmarks (i.e.,PageRankandLDA). Extensive experiments show that the proposed executor allocation strategy can decrease the network traffic and data access time by improving the data locality during the task scheduling. Its performance benefits are particularly significant for iterative applications. Zhongming Fu, Mengsi He, Zhuo Tang |
IEEE Trans. Cloud Comput. | 2 |
| 2023 | PCU-Net: An enhanced U-network by combining PPM and CBAM for Medical Image SegmentationabstractU-network is a kind of full convolutional neural network, which is widely used in the field of medical image segmentation. However, it still has the problem of small targets being unable to be segmented and resulting in unsatisfactory segmentation effect. In order to solve this problem, this paper proposes an enhanced U-network by combining pyramid pooling module (PPM) and convolutional block attention module (CBAM). It’s whole network is U-Net architecture, where PPM with bin sizes of 1×1, 2×2 and 3×3 are used in the downsampling part of the network, which can extract input image features of various dimensions. And CBAM is used in the downsampling part of the network, which combines convolution and attention mechanism, and can pay attention to the image from two aspects of space and channel to improve the segmentation ability of the network. Experimental results show that our network outperforms traditional Ushaped segmentation networks by 30% to 40% in metrics IoU, MAE, and Dice, respectively. Hejian Chen, Zhongming Fu, Mengsi He, Jiayi Deng, Zhuo Tang |
ICPADS | 4 |
| 2022 | IncGraph: An Improved Distributed Incremental Graph Computing Model and Framework Based on Spark GraphXabstractThe excavated information will become obsolete when the data changes in dynamic graphs. To compute the up-to-date results, the graph algorithm has to re-compute the entire data from scratch, which will consume huge computation time and resources. To reduce the cost of such calculations, this paper proposes a model called IncGraph to support incremental iterative computation over dynamic graphs. Different from the way of traditional iteration, IncGraph executes the graph algorithm through reusing the results of the previous graph and performs computation on the part of the graph that has changed. IncGraph has two critical components: (1) an incremental iterative computation model that consists of two steps: an incremental step to calculate the results on the changed vertices of the graph, and a merge step to calculate the results on the entire graph by using the results of the previous graph and the incremental step; and (2) an incremental update method to accelerate the iterative process within the iterative graph algorithm. We implement IncGraph model on GraphX and evaluate its performance by using several representative iterative graph algorithms: PageRank, Connected components, and Single Source Shortest Path. The results show that compared with the traditional iteration, when adding the 100k of vertices in different size data sets, the performance optimization ratio of IncGraph is 31.79 percent averagely, and 50.2 percent maximum; and when the percentage of added vertices varied from 0.01 to 10 percent in different data sets, the performance optimization ratio of IncGraph varied from 19.9 to 66.1 percent. Moreover, the result errors of IncGraph is small and can be neglected. Zhuo Tang, Mengsi He, Zhongming Fu, Li Yang 0012 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Optimizing Data Locality by Executor Allocation in Reduce Stage for Spark Framework
Zhongming Fu, Mengsi He, Zhuo Tang, Yang Zhang 0026 |
PDCAT | 2 |