EDBT 2026 Demo / reviewers in the wild / expert
Zeyao Mo
dblp:14/704
· DBLP profile ↗
24ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0003-3280-5682ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HAGC: A Hardware-Aware Gradient Compression framework for distributed deep learning
Aiqiang Yang, Jie Liu 0002, Bo Yang 0021, Xiang Zhang 0008, Zeyao Mo, Keqin Li 0001 |
J. Syst. Archit. | 6 |
| 2025 | Balancing communication overhead and accuracy in compression integration: a survey
Aiqiang Yang, Zeyao Mo |
J. Supercomput. | 4 |
| 2024 | A More Context-Aware Approach for Textual Adversarial Attacks Using Probability Difference-Guided Beam SearchabstractTextual adversarial attacks expose the vulnerabilities of text classifiers and can be used to improve their robustness. Previous context-aware attack models suffer from several limitations. They generally rely on out-of-date substitutes, solely consider the gold label probability, and use the greedy search when generating adversarial examples, often limiting the attack efficiency. To tackle these issues, we proposeMC-PDBS, aMoreContext-aware textual adversarial attack model usingProbabilityDifference-guidedBeamSearch. MC-PDBS generates substitutes using the newest perturbed text sequences in each attack iteration, enabling the generation of more context-aware adversarial examples. The probability difference is an overall consideration of the probabilities of all class labels, which is more effective than the gold label probability in guiding the selection of attack paths. In addition, the beam search enables MC-PDBS to search attack paths from multiple search channels, thereby avoiding the limited search space problem. Extensive experiments and human evaluation demonstrate that MC-PDBS outperforms previous best models in a series of evaluation metrics, particularly bringing up to a +19.5% attack success rate. Extensive analyses further confirm the effectiveness of MC-PDBS. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Zibo Yi, Mengxue Du, Miaomiao Li 0001, Jie Liu 0002, Zeyao Mo |
IEEE Trans. Knowl. Data Eng. | 10 |
| 2023 | JSweep: A Patch-centric Data-driven Approach for Parallel Sweeps on Large-scale MeshesabstractIn mesh-based numerical simulations, sweep is an important computation pattern. During sweep on meshes, computations on cells are strictly ordered by data dependencies in given directions. Due to this order constraint, parallelizing sweep is challenging, especially for unstructured and deforming meshes. Meanwhile, recent high-fidelity multi-physics simulations of particle transport, including nuclear reactor and inertial confinement fusion, require sweeps on large scale meshes with billions of cells and hundreds of directions. In this paper, we present JSweep, a parallel data-driven framework integrated in the JAxMIN infrastructures. The essential of JSweep is a general patch-centric data-driven abstraction, coupled with a high performance runtime system leveraging hybrid parallelism of MPI+threads and achieving dynamic communication on contemporary multi-core clusters. Built on JSweep, we implement a representative data-driven algorithm, Sn transport, featuring optimizations of vertex clustering, multi-level priority strategy and patch-angle parallelism. Experimental evaluation with two real-world applications on structured and unstructured meshes respectively, demonstrates that JSweep can scale to tens of thousands of processor cores with reasonable parallel efficiency. Aiqing Zhang, Zeyao Mo |
ICPP | 4 |
| 2023 | JXPAMG: a parallel algebraic multigrid solver for extreme-scale numerical simulations
Xiaoqiang Yue, Runzhang Mao, Yuntong Deng, Silu Huang, Haifeng Zou, Shaoliang Hu, Chunsheng Feng, Shi Shu, Zeyao Mo |
CCF Trans. High Perform. Comput. | 11 |
| 2021 | JCOGIN: a programming framework for particle transport on combinatorial geometry
Baoyin Zhang, Zeyao Mo, Xin Wang 0078, Wei Wang 0229, Aiqing Zhang, Xiaolin Cao |
J. Supercomput. | 2 |
| 2020 | αSetup-AMG: an adaptive-setup-based parallel AMG solver for sequence of sparse linear systems
Zeyao Mo, Xiaoqiang Yue, Hengbin An, Shi Shu |
CCF Trans. High Perform. Comput. | 2 |
| 2019 | JAUMIN: a programming framework for large-scale numerical simulation on unstructured meshes
Qingkai Liu, Zeyao Mo, Aiqing Zhang |
CCF Trans. High Perform. Comput. | 2 |
| 2018 | Extreme-scale parallel computing: bottlenecks and strategiesabstractExtreme-scale numerical simulations seriously demand extreme parallel computing capabilities. To address the challenges of these capabilities toward exascale, we systematically analyze the major bottlenecks of parallel computing research from three perspectives: computational scale, computing efficiency, and programming productivity. For these bottlenecks, we propose a series of urgent key issues and coping strategies. This study will be useful in synchronizing development between the numerical computing capability and supercomputer peak performance. Zeyao Mo |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2016 | Graphine: Programming Graph-Parallel Computation of Large Natural Graphs for Multicore ClustersabstractGraph-parallel computation has become a crucial component in emerging applications of web search, data analytics and machine learning. In practice, most graphs derived from real-world phenomena are very large and scale-free. Unfortunately, distributed graph-parallel computation of these natural graphs still suffers strong scalability issues on contemporary multicore clusters. To embrace the multicore architecture in distributed graph-parallel computation, we propose the framework Graphine, which features (i) A Scatter-Combine computation abstraction that is evolved from the traditional vertex-centric approach by fusing the paired scatter and gather operations, executed separately on two edge sides, into a one-sided scatter. Further coupled with active message mechanism, it potentially reduces intermediate message cost and enables fine-grained parallelism on multicore architecture. (ii) An Agent-Graph data model, which leverages an idea similar to vertex-cut but conceptually splits the remote replica into two agent types of scatter and combiner, resulting in less communication. We implement the Graphine framework and evaluate it using several representative algorithms on six large real-world graphs and a series of synthetic graphs with power-law degree distributions. We show that Graphine achieves sublinear scalability with the number of cores per node, number of nodes, and graph sizes (up to one billion vertices), and is 2~15 times faster than the state-of-the-art PowerGraph on a cluster of 16 multicore nodes. Guangming Tan, Zeyao Mo, Ninghui Sun |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | FAST: A Fast Stencil Autotuning Framework Based On An Optimal-solution Space ModelabstractStencil computations comprise an important class of kernels in many scientific computing applications. As the diversity of both architectures and programming models grow, autotuning is emerging as a critical strategy for achieving portable performance across a broad range of execution contexts for stencil computations. However, costly tuning overhead is a major obstacle to its popularity. In this work, we propose a fast stencil autotuning framework FAST based on an Optimal-Solution Space (OSS) model to significantly improve tuning speed. It leverages a feature extractor that comprehensively characterizes stencil computation. Using the extracted features, FAST constructs an OSS database to train an off-line model which provides an on-line prediction. We evaluate FAST with five important stencil computation applications on both an Intel Xeon multicore CPU and an NVIDIA Tesla K20c GPU. Compared with state-of-the-art stencil autotuners like Patus and SDSL, FAST improves autotuning speed by 10-2697 times without any user annotation, while achieving comparable performance. Yulong Luo, Guangming Tan, Zeyao Mo, Ninghui Sun |
ICS | 3 |
| 2015 | Performance Optimization Using Partitioned SpMV on GPUs and Multicore CPUsabstractThis paper presents a sparse matrix partitioning strategy to improve the performance of SpMV on GPUs and multicore CPUs. This method has wide adaptability for different types of sparse matrices, and is different from existing methods which only adapt to some particular sparse matrices. In addition, our partitioning method can obtain dense blocks by analyzing the probability distribution of non-zero elements in a sparse matrix, and result in very low proportion of zero padded. We make the following significant contributions. (1) We present a partitioning strategy of sparse matrices based on probabilistic modeling of non-zero elements in a row. (2) We prove that our method has the highest mean density compared with other strategies according to certain given ratios of partition obtained from the computing powers of heterogeneous processors. (3) We develop a CPU-GPU hybrid parallel computing model for SpMV on GPUs and multicore CPUs in a heterogeneous computing platform. Our partitioning strategy has balanced load distribution and the performance of SpMV is significantly improved when a sparse matrix is partitioned into dense blocks using our method. The average performance improvement of our solution for SpMV is about 15.75 percent on multicore CPUs, compared to that of the other solutions. By considering the rows of a matrix in a unique order based on the probability mass function of the number of non-zeros in a row, the average performance improvement of our solution for SpMV is about 33.52 percent on GPUs and multicore CPUs of a heterogeneous computing platform, compared to that of the partitioning methods based on the original row order of a matrix. Wangdong Yang, Kenli Li 0001, Zeyao Mo, Keqin Li 0001 |
IEEE Trans. Computers | 3 |
| 2014 | A new parallel algorithm for vertex priorities of data flow acyclic digraphs
Zeyao Mo, Aiqing Zhang |
J. Supercomput. | 1 |
| 2013 | Component-based Parallel Programming for Peta-scale Particle Simulations
Xiaolin Cao, Zeyao Mo, Aiqing Zhang |
ICSOFT | 2 |
| 2011 | Parallel implementation of fast multipole method based on JASMIN
Xiaolin Cao, Zeyao Mo, Aiqing Zhang |
Sci. China Inf. Sci. | 2 |
| 2010 | JASMIN: a parallel software infrastructure for scientific computing
Zeyao Mo, Aiqing Zhang, Xiaolin Cao, Qingkai Liu, Hengbin An, Wenbing Pei, Shaoping Zhu |
Frontiers Comput. Sci. China | 1 |
| 2007 | Dynamic Load-Balancing and High Performance Communication in JclusterabstractThis paper describes the dynamic load-balancing and high performance communication provided in Jcluster, an efficient Java parallel environment. For the efficient load-balancing, we implement a task scheduler based on a transitive random stealing algorithm, which improves the random stealing, a well-known load-balancing algorithm. The experiment results show that the scheduler performs efficiently, especially for a large-scale cluster. With the method of asynchronously multithreaded transmission, a high performance PVM-like and MPI-like message passing interface is implemented in pure Java. The evaluation of the communication performance is conducted among Jcluster, LAM-MPI and mpiJava on LAM-MPI based on the Java Grande Forum's pingpong benchmark. Bao-Yin Zhang, Zeyao Mo |
IPDPS | 2 |
| 2007 | Relaxed RS0 or CLJP coarsening strategy for parallel AMG
Zeyao Mo |
Parallel Comput. | 1 |
| 2006 | Towards a parallel framework of grid-based numerical algorithms on DAGsabstractThis paper presents a parallel framework of grid-based numerical algorithms where data dependencies between grid zones can be modeled by a directed acyclic graph (DAG). It consists of three parts on how to partition, order and calculate the vertices of digraph. Numerical results using hundreds of processors on two parallel machines show the efficiencies and moderate scalability of this framework Zeyao Mo, Aiqing Zhang, Xiaolin Cao |
IPDPS | 1 |
| 2006 | Study on Parallel Computing
Guoliang Chen 0001, Guangzhong Sun, Yunquan Zhang, Zeyao Mo |
J. Comput. Sci. Technol. | 4 |
| 2005 | An Efficient Dynamic Load-Balancing Algorithm in a Large-Scale Cluster
Bao-Yin Zhang, Zeyao Mo |
ICA3PP | 2 |
| 2005 | Scheduling Efficiently for Irregular Load Distributions in a Large-scale Cluster
Bao-Yin Zhang, Zeyao Mo |
ISPA | 2 |
| 2004 | A New Scalable Parallel Method for Molecular Dynamics Based on Cell-Block Data Structure
Xiaolin Cao, Zeyao Mo |
ISPA | 2 |
| 2004 | Parallel Flux Sweep Algorithm for Neutron Transport on Unstructured Grid
Zeyao Mo, Lianxiang Fu |
J. Supercomput. | 1 |