EDBT 2026 Demo / reviewers in the wild / expert
Ming Dun
dblp:250/0531
· DBLP profile ↗
18ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-0664-9543ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing Streaming Tensor Decomposition on GPUabstractTensors represent multidimensional data and cover various areas of scientific computing. The Canonical Polyadic Decomposition (CPD) emerges to extract latent patterns from large but highly sparse tensors. In real-world scenarios, tensor slices often arrive dynamically over time in streaming form, making traditional CPD algorithms inefficient in processing the entire tensor at each time step. Streaming CPD processes tensor slices incrementally, exploiting a forgetting factor to adjust the weight of historical information to capture dynamics. Current optimizations mainly focus on CPU platforms, failing to meet the real-time processing requirements of modern applications. Efficiently deploying streaming CPD on GPU remains challenging due to frequent data transfers and memory operations throughout the complex workflow, as well as the intricate computational patterns of bottleneck operators. Wenqing Lin, Jianuo Sheng, Shuqin Feng, Ming Dun, Huawei Cao, Qingxiao Sun |
ICS | 4 |
| 2026 | Toward Resource-Efficient Billion-Scale SpGEMM on CPU-GPU Heterogeneous Server
Ming Dun, Shuhan Song, Huawei Cao, Xuejun An, Xiaochun Ye |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2025 | A Co-Design Framework for Graph Processing on CPU-GPU Heterogeneous PlatformsabstractRecently, large-scale graph processing on CPU-GPU heterogeneous platforms has attracted considerable attention. However, disparities in memory bandwidth and parallel computational capabilities between CPUs and GPUs, coupled with the irregular structure of graphs and the inherent unpredictability of graph algorithms, often lead to inefficient utilization of CPU-GPU hardware resources, ultimately degrading graph processing performance. To address this, we propose and implement CoDgraph, a co-design framework for high-performance graph processing on CPU-GPU heterogeneous platforms. Specifically, we introduce a fine-grained partitioning strategy to balance workloads, minimize communication overhead, and enhance data locality. Next, we develop an adaptive co-scheduling computing scheme, leveraging a cost model that accounts for CPU and GPU hardware resources to improve system utilization. Finally, to further optimize largescale graph processing, we design and implement an efficient overlapping pipeline execution mode that employs asynchronous parallel execution. Extensive evaluations demonstrate that CoDgraph outperforms state-of-the-art CPU and CPU-GPU graph processing systems, including Ligra (CoDgraph is$14.59 \times$faster on average) and Subway (CoDgraph is$4.17 \times$faster on average). In addition, CoDgraph also has comparable performance to the advanced in-memory graph processing Tigr on GPU and shows good scalability for different graph scales and CPU-GPU heterogeneous platforms. Yuan Zhang 0031, Huawei Cao, Ming Dun, Jie Zhang 0130, Xiaochun Ye |
ICCD | 4 |
| 2025 | GPromptShield: Elevating Resilience in Graph Prompt Tuning Against Adversarial AttacksabstractThe paradigm of ``pre-training and prompt-tuning", with its effectiveness and lightweight characteristics, has rapidly spread from the language field to the graph field. Several pioneering studies have designed specialized prompt functions for diverse downstream graph tasks based on various graph pre-training strategies. These prompts concentrate on the compatibility between the pre-training pretext and downstream graph tasks, aiming to bridge the gap between them. However, designing prompts blindly to adapt to downstream tasks based on this concept neglects crucial security issues. By conducting covert attacks on downstream graph data, we find that even when the downstream task data closely matches that of the pre-training tasks, it is still feasible to generate highly misleading prompts using simple deceptive techniques. In this paper, we shift the primary focus of graph prompts from compatibility to vulnerability issues in adversarial attack scenarios. We design a highly extensible shield defense system for the prompts, which enhances their robustness from two perspectives:Direct Handling and Indirect Amplification. When downstream graph data contains unreliable biases, the former directly combats invalid information by incorporating hybrid multi-defense prompts to the input graph's feature space, while the latter adopts a training strategy to bypass the invalid components and amplifies valid part. We provide a theoretical derivation that proves their feasibility, indicating that unbiased prompts exist under certain conditions on unreliable data. Extensive experiments across various scenarios of adversarial attacks (including adaptive and non-adaptive attacks) indicate that the prompts within our defense system exhibit enhanced resilience and superiority. This paper explores a new perspective in graph prompt learning, offering a novel option for robust prompt tuning in downstream tasks. Shuhan Song, Ming Dun, Maolei Huang, Huawei Cao, Xiaochun Ye |
ICLR | 3 |
| 2025 | Equipping Graph Autoencoders: Revisiting Masking Strategies from a Robustness PerspectiveabstractMasked Graph Autoencoders (MGAEs), represented by GraphMAE and GraphMAE2, which utilize masked feature (or structure) reconstruction strategies, have demonstrated the potential to surpass contrastive learning. However, current masked reconstruction strategies primarily rely on random strategies, only prove effective on reliable graph data. Therefore, these popular methods face immediate robustness deficiencies issues. Firstly, when the graph is unreliable or under adversarial attacks, the selection of nodes for masked reconstruction has a significant impact on downstream tasks. Secondly, the reconstructed features contains redundant components. In this paper, to overcome the non-robustness caused by randomness, we provide a theoretical analysis and evaluation of the robustness of state-of-the-art MGAEs. Additionally, we design two lightweight plug-and-play tools: Box-Based Weighted Reliability Ranking Masking Strategy and Decoupled Feature Reconstruction. Without incurring additional time overhead, these tools provide a defense armor against adversarial attacks for MGAEs, significantly boosting the robustness performance of downstream tasks. Extensive experiments on real-world graphs attacked by various attacks demonstrate our designs have a considerable robust expressive ability. Especially on datasets with large perturbations, the defense performance could even be improved by up to 20%. Shuhan Song, Ming Dun, Yuan Zhang 0031, Huawei Cao, Xiaochun Ye |
SDM | 3 |
| 2025 | SPMGAE: Self-purified masked graph autoencoders release robust expression power
Shuhan Song, Ming Dun, Yuan Zhang 0031, Huawei Cao, Xiaochun Ye |
Neurocomputing | 3 |
| 2024 | QAAS: quick accurate auto-scaling for streaming processing
Yunchun Li, Ming Dun, Chen Chen 0094, Huaitao Zhang |
Frontiers Comput. Sci. | 4 |
| 2022 | CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUsabstractGraph neural networks (GNNs) suffer from low GPU utilization due to frequent memory accesses. Existing concurrent training mechanisms cannot be directly adapted to GNNs because they fail to consider the impact of input irregularity. This requires pre-profiling the memory footprint of concurrent tasks based on input dimensions to ensure successful co-location on GPU. Moreover, massive training tasks generated from scenarios such as hyper-parameter tuning require flexible scheduling strategies. To address these problems, we propose CoGNN that enables efficient management of GNN training tasks on GPUs. Specifically, the CoGNN organizes the tasks in a queue and estimates the memory consumption of each task based on cost functions at operator basis. In addition, the CoGNN implements scheduling policies to generate task groups, which are iteratively submitted for execution. The experiment results show that the CoGNN can achieve shorter completion and queuing time for training tasks from diverse GNN models. Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Ruizhe Zhang 0012, Ming Dun, Mingzhen Li 0001, Wencong Xiao, Yong Li 0020, Zhongzhi Luan, Depei Qian 0001 |
SC | 5 |
| 2022 | Input-Aware Sparse Tensor Storage Format Selection for Optimizing MTTKRPabstractThe major bottleneck of Canonical polyadic decomposition (CPD) is matricized tensor times Khatri-Rao product (MTTKRP). To optimize the performance of MTTKRP, various sparse tensor formats have been proposed such as CSF and HiCOO. However, due to the spatial complexity of the tensors, no single format fits all tensors. To address this problem, we propose SpTFS, a framework that automatically predicts the optimal storage format for an input sparse tensor. Specifically, SpTFS leverages a set of sampling methods to lower the sparse tensor to fix-sized matrices and sparsity features. In addition, SpTFS adopts both supervised learning based and unsupervised learning based methods to predict the optimal sparse tensor storage formats. For supervised learning, we propose TnsNet that combines convolution neural network (CNN) and the feature layer, which effectively captures the sparsity patterns of the input tensors. Whereas for unsupervised learning, we propose TnsClustering that consists of a feature encoder using convolutional layers and fully connected layers, and a K-means++ model to cluster sparse tensors for optimal tensor format prediction, without massively profiling on the hardware platform. The experimental results show that both TnsNet and TnsClustering can achieve higher prediction accuracy and performance speedup compared to the state-of-the-art works. Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Ming Dun, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
IEEE Trans. Computers | 4 |
| 2022 | Accelerating approximate matrix multiplication for near-sparse matrices on GPUs
Yi Liu 0013, Hailong Yang 0002, Ming Dun, Bohong Yin, Zhongzhi Luan, Depei Qian 0001 |
J. Supercomput. | 4 |
| 2021 | PriPro: Towards Effective Privacy Protection on Edge-Cloud System running DNN InferenceabstractThe huge computation demand for deep learning models and limited computation resources on the edge devices calls for the cooperation between the edge device and cloud service. On a typical edge-cloud system accommodating DNN inference, a deep model is split into two partial models running on the edge device and the cloud service, respectively. The two partial models collaborate closely to satisfy the DNN inference requested by the user. However, user's privacy is vulnerable when transferring the intermediate results generated by the partial model at edge device to cloud service. Existing research works rely on metrics that are either impractical or insufficient to measure the effectiveness of privacy protection methods in the above scenario, especially from a single input aspect. In this paper, we first thoroughly analyze the state-of-the-art methods and drawbacks of existing methods from the aspects of both evaluation metrics and proposed techniques. Then, we propose a new metric system, including privacy accuracy (PA) and privacy index (PI), that can accurately measure the effectiveness of privacy protection methods. Furthermore, we propose PriPro, a privacy protection method that can dynamically inject noise to the intermediate results at various layers regarding the input features through the self-attention mechanism. The experiment results demonstrate our method outperforms existing methods for protecting user privacy on deep models such as AlexNet, VGG, and ResNet. Ruiyuan Gao 0001, Hailong Yang 0002, Shaohan Huang, Ming Dun, Mingzhen Li 0001, Zerong Luan, Zhongzhi Luan, Depei Qian 0001 |
CCGRID | 4 |
| 2021 | csTuner: Scalable Auto-tuning Framework for Complex Stencil Computation on GPUsabstractThe computational patterns of stencil operations are commonly used in HPC applications. Many HPC platforms utilize the computation capability of GPUs to accelerate stencil operations. In recent years, stencils have become more complex in terms of stencil order, memory accesses, and operator patterns. To adapt complex stencils to GPUs, various optimization techniques have been proposed such as blocking and unrolling. However, due to the complexity of GPU architecture, no single parameter setting of the optimization techniques fits all stencils. To address this problem, we propose csTuner, a scalable auto-tuning framework that quickly determines the optimal parameter setting for a given combination of optimization techniques. Specifically, csTuner leverages a set of statistics and machine learning methods to generate parameter groups and sampled parameter settings from the search space. In addition, csTuner adopts the genetic algorithm with approximation to reduce the cost of evolutionary search. The experimental results show that csTuner can find better performing settings with higher auto-tuning speed compared to the state-of-the-art works. Qingxiao Sun, Yi Liu 0013, Hailong Yang 0002, Zhonghui Jiang, Ming Dun, Zhongzhi Luan, Depei Qian 0001 |
CLUSTER | 6 |
| 2021 | An optimized tensor completion library for multiple GPUsabstractTensor computations are gaining wide adoption in big data analysis and artificial intelligence. Among them, tensor completion is used to predict the missing or unobserved value in tensors. The decomposition-based tensor completion algorithms have attracted significant research attention since they exhibit better parallelization and scalability. However, existing optimization techniques for tensor completion cannot sustain the increasing demand for applying tensor completion on ever larger tensor data. To address the above limitations, we develop the first tensor completion library cuTC on multiple Graphics Processing Units (GPUs) with three widely used optimization algorithms such as alternating least squares (ALS), stochastic gradient descent (SGD) and coordinate descent (CCD+). We propose a novel TB-COO format that leverages warp shuffle and shared memory on GPU to enable efficient reduction. In addition, we adopt the auto-tuning method to determine the optimal parameters for better convergence and performance. We compare cuTC with state-of-the-art tensor completion libraries on real-world datasets, and the results show cuTC achieves significant speedup with similar or even better accuracy. Ming Dun, Yunchun Li, Hailong Yang 0002, Qingxiao Sun, Zhongzhi Luan, Depei Qian 0001 |
ICS | 1 |
| 2021 | Towards efficient canonical polyadic decomposition on sunway many-core processor
Ming Dun, Yunchun Li, Qingxiao Sun, Hailong Yang 0002, Wei Li 0125, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
Inf. Sci. | 1 |
| 2021 | Towards efficient tile low-rank GEMM computation on sunway many-core processors
Qingchang Han, Hailong Yang 0002, Ming Dun, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
J. Supercomput. | 3 |
| 2020 | Accelerating De Novo Assembler WTDBG2 on Commodity Servers
Ming Dun, Yunchun Li, Xin You 0001, Qingxiao Sun, Zerong Luan, Hailong Yang 0002 |
ICA3PP (1) | 1 |
| 2020 | SpTFS: sparse tensor format selection for MTTKRP via deep learningabstractCanonical polyadic decomposition (CPD) is one of the most common tensor computations adopted in many scientific applications. The major bottleneck of CPD is matricized tensor times Khatri-Rao product (MTTKRP). To optimize the performance of MTTKRP, various sparse tensor formats have been proposed such as CSF and HiCOO. However, due to the spatial complexity of the tensors, no single format fits all tensors. To address this problem, we propose SpTFS, a framework that automatically predicts the optimal storage format for an input sparse tensor. Specifically, SpTFS leverages a set of sampling methods to lower the sparse tensor to fix-sized matrices and specific features. Then, TnsNet combines CNN and the feature layer to accurately predict the optimal format. The experimental results show that SpTFS achieves prediction accuracy of 92.7% and 96% on CPU and GPU respectively. Qingxiao Sun, Yi Liu 0013, Ming Dun, Hailong Yang 0002, Zhongzhi Luan, Lin Gan 0001, Guangwen Yang 0002, Depei Qian 0001 |
SC | 3 |
| 2019 | Improving the Parallelism of CESM on GPU
Zehui Jin, Ming Dun, Xin You 0001, Hailong Yang 0002, Yunchun Li, Yingchun Lin, Zhongzhi Luan, Depei Qian 0001 |
ICA3PP (2) | 2 |