EDBT 2026 Demo / reviewers in the wild / expert
Menghan Jia
dblp:216/9294
· DBLP profile ↗
18ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DOA: Dataflow Optimization for Attention on Multi-core DSPs with Three-Level Memory Hierarchy
Zhiquan Lai, Shun Ouyang, Zhaoning Zhang 0001, Menghan Jia, Huayou Su, Dongsheng Li 0001 |
APPT | 7 |
| 2026 | Hierarchy-Aware Neural Subgraph Matching with Enhanced Similarity Measure (Extended Abstract)
Zhouyang Liu, Ning Liu 0015, Yixin Chen 0004, Jiezhong He, Menghan Jia, Dongsheng Li 0001 |
ICDE | 5 |
| 2026 | MeCache: Communication-Efficient Multi-GPU Heterogeneous Graph Neural Network Training
Gongqingjian Jiang, Menghan Jia, Zhiquan Lai, Dongsheng Li 0001 |
IPDPS | 3 |
| 2026 | FastCC: A System-Algorithm Co-Design for Connected Components Computation on Large Power-Law GraphsabstractConnected Components (CC) computation is a fundamental graph analytics kernel. While BFS-sampling has emerged as the state-of-the-art approach for power-law graphs, its performance in existing implementations is severely limited by inheriting unnecessary BFS semantics. The core issue is a fundamental mismatch: BFS requires strict level-synchronization to compute shortest paths, while CC only needs eventual label consistency without ordering constraints. This semantic mismatch manifests as three critical bottlenecks: (i) severe load imbalance from vertex-centric task allocation, which fails to distribute the massive workload of high-degree hub vertices; (ii) redundant writes from dynamic push/pull mode switching, which necessitates costly frontier reconstruction; and (iii) redundant synchronization and computation from enforcing BFS’s strict ordering guarantees, which are superfluous for CC computation. We introduce FastCC , a lightweight multiprocess system-algorithm co-designed solution that breaks this semantic mismatch. The cornerstone of our approach is the strategic decision to fix the highest-degree vertex as the BFS root, creating a predictable computation topology. This enables three synergistic innovations: (1) hybrid task partitioning that employs edge-centric allocation in the critical first iteration to eliminate load imbalance at its source; (2) predictable mode switching that leverages the deterministic computation graph to bypass expensive frontier reconstruction; and (3) lightweight, custom synchronization primitives that relax BFS’s strict ordering to match CC’s eventual consistency requirements, while allowing earlier label propagation within the same iteration to reduce the overall computational workload. Extensive evaluation on large-scale real-world and synthetic power-law graphs demonstrates that FastCC achieves significant performance improvements, with speedups of 10.6-55.5× faster (average: 37.8×) over state-of-the-art CC implementations including ConnectIt and vGraph. FastCC also reduces peak memory footprint by up to 2.87× and exhibits superior, more predictable scalability. The practical efficacy of our approach is validated by its deployment as the core engine in a top-ranked GreenGraph500 solution. Menghan Jia, Yongquan Fu, Yiming Zhang 0003, Xinhai Chen 0001, Dongsheng Li 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | AccuGraph: Memory-Efficient Full-Graph GNN Training on a Single GPU via Subgraph Accumulation
Liyang Wu, Menghan Jia, Yahui Wu, Gongqingjian Jiang, Jiezhong He, Chunye Gong, Yinghui Gao |
ICA3PP (3) | 2 |
| 2025 | HMGraph: Boosting GNN Training on Hierarchical Memory via Coordinated CacheabstractThe GPU-CPU-SSD hierarchical memory systems are commonly employed for large-scale GNN training. However, existing solutions inefficiently utilize high-bandwidth memory due to coarse-grained memory management and poor data placement that ignores graph access patterns. This paper presents HMGraph, a GNN training system unleashing the full potential of hierarchical memory architectures. The core design of HMGraph is Coordinated Cache, integrating GPU memory and CPU memory as a cache layer for the hierarchical memory system and improving GNN efficiency through fine-grained data and memory management. For this goal, three main designs are proposed. First, we design an automatic cache management mechanism that optimizes cache allocation based on the graph data access pattern to enhance the overall cache hit rate. Second, we propose a dynamic data space management strategy to improve the efficiency of dynamic cache. Third, we develop a hierarchical memory-aware data partitioning strategy that further improves the utilization of high-bandwidth memory. Our evaluation of various large-scale graphs reveals that HMGraph significantly outperforms other state-of-the-art systems by 1.4-36.7 ×. Menghan Jia, Zhiquan Lai, Qiao Li 0001, Yiming Zhang 0003, Dongsheng Li 0001 |
ICPP | 2 |
| 2025 | UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform LossabstractPartial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-varying regions, enhancing both simulation accuracy and computational efficiency. However, traditional approaches suffer from high computational complexity and geometric inflexibility, limiting their applicability, and existing supervised learning-based approaches face challenges in zero-shot generalization across diverse PDEs and mesh topologies.
In this paper, we present an $\textbf{U}$nsupervised and $\textbf{G}$eneralizable $\textbf{M}$esh $\textbf{M}$ovement $\textbf{N}$etwork (UGM2N). We first introduce unsupervised mesh adaptation through localized geometric feature learning, eliminating the dependency on pre-adapted meshes. We then develop a physics-constrained loss function, M-Uniform loss, that enforces mesh equidistribution at the nodal level. Experimental results demonstrate that the proposed network exhibits equation-agnostic generalization and geometric independence in efficient mesh adaptation. It demonstrates consistent superiority over existing methods, including robust performance across diverse PDEs and mesh geometries, scalability to multi-scale resolutions and guaranteed error reduction without mesh tangling. Xinhai Chen 0001, Xiang Gao 0020, Qingyang Zhang 0009, Menghan Jia, Xiang Zhang 0008, Jie Liu 0002 |
NeurIPS | 6 |
| 2025 | TriFMatch: a flash subgraph matching algorithm with effective filtering techniques
Jiezhong He, Yixin Chen 0004, Menghan Jia, Zhouyang Liu, Dongsheng Li 0001, Kian-Lee Tan |
Knowl. Inf. Syst. | 3 |
| 2025 | Hierarchy-Aware Neural Subgraph Matching With Enhanced Similarity MeasureabstractSubgraph matching is challenging as it necessitates time-consuming combinatorial searches. Recent Graph Neural Network (GNN)-based approaches address this issue by employing GNN encoders to extract graph information and hinge distance measures to ensure containment constraints in the embedding space. These methods significantly shorten the response time, making them promising solutions for subgraph retrieval. However, they suffer from scale differences between graph pairs during encoding, as they focus on feature counts but overlook the relative positions of features within node-rooted subtrees, leading to disturbed containment constraints and false predictions. Additionally, their hinge distance measures lack discriminative power for matched graph pairs, hindering ranking applications. We propose NC-Iso, a novel GNN architecture for neural subgraph matching. NC-Iso preserves the relative positions of features by building the hierarchical dependencies between adjacent echelons within node-rooted subtrees, ensuring matched graph pairs maintain consistent hierarchies while complying with containment constraints in feature counts. To enhance the ranking ability for matched pairs, we introduce a novel similarity dominance ratio-enhanced measure, which quantifies the dominance of similarity over dissimilarity between graph pairs. Empirical results on nine datasets validate the effectiveness, generalization ability, scalability, and transferability of NC-Iso while maintaining time efficiency, offering a more discriminative neural subgraph matching solution for subgraph retrieval. Zhouyang Liu, Ning Liu 0015, Yixin Chen 0004, Jiezhong He, Menghan Jia, Dongsheng Li 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | vGraph: Memory-Efficient Multicore Graph Processing for Traversal-Centric AlgorithmsabstractTo lower the monetary/energy cost, single-machine multicore graph processing is gaining increasing attention for a wide range of traversal-centric graph algorithms such as BFS, SSSP, CC, and PageRank, of which the processing is relatively simple and the topology data (vertices and edges) dominates the memory footprint. This paper presents$v$Graph, a NUMA-aware, memory-efficient multicore graph processing system for traversal-centric algorithms.$v$Graph proposes an ultralight NUMA-aware graph preprocessing scheme which eliminates almost all complex preprocessing steps and pipelines per-NUMA graph loading and compressing, to effectively reduce inter-NUMA memory accesses while keeping both preprocessing cost and peak memory footprint low. We further optimize$v$Graph with effective HPC techniques including prefetching and work-stealing. Evaluation on a 384GB-memory, four-NUMA machine shows that compared to the state-of-the-art NUMA-aware/-unaware systems,$v$Graph can process much larger real-world and synthetic graphs with various traversal-centric algorithms, achieving significantly higher memory efficiency and lower processing time. Menghan Jia, Yiming Zhang 0003, Xinbiao Gan, Dongsheng Li 0001, Erci Xu, Ruibo Wang, Kai Lu 0001 |
SC | 1 |
| 2021 | Modeling and Analysis of Medical Resource Sharing and Scheduling for Public Health Emergencies based on Petri NetsabstractMedical information systems (MIS) play a vital role in managing and scheduling medical resources to underpin healthcare services, which has become more critically important during major public health emergencies. During the Covid-19 pandemic, MIS is facing significant challenges to cope with the surge in demands of medical resources, resulting in more deaths and wider spreading of the disease. Our research examines how to allocate and utilize the medical resources across hospitals in a more accurate, and effective way to mitigate medical resource shortages and sustain the resource provisions. This paper mainly investigated the hospital’s supply-and-demand problems for medical resources under major public health emergencies by analyzing the allocation of medical staff resources. Furthermore, a formal method based on the Colored Petri Nets (CPN) has been proposed to model and characterize the medical business process and resource scheduling tasks. The experiments demonstrate that our approach can correctly and efficiently complete the dynamical scheduling process for surging requests. Wangyang Yu 0001, Menghan Jia, Bo Yuan 0004 |
MSN | 2 |
| 2021 | XSP: Fast SSSP Based on Communication-Computation Collaboration
Xinbiao Gan, Menghan Jia, Jie Liu 0002, Yiming Zhang 0003 |
NPC | 3 |
| 2021 | VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003 |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | Correction to: VPC: Pruning connected components using vector-based path compression for Graph500
Xinbiao Gan, Tianjing Xu, Menghan Jia, Juan Chen 0001, Yiming Zhang 0003 |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | QRDF: An efficient RDF graph processing system for fast queryabstractAbstract With the rapid growth of RDF (resource description framework) data volume, it is challenging to answer queries over dynamic RDF graphs of which the data frequently changes (by insertion/deletion) over time. This article presents QRDF, an efficient RDF graph processing system that supports not only efficient storage but also fast query of dynamic RDF graphs. QRDF (i) adopts a red‐black (RB) tree for quick RDF data insertion/deletion and (ii) asynchronously moves data from the RB tree into a vector for efficient RDF query. QRDF performs early pruning on its index structure for various RDF query patterns. Experimental results show that the query performance of QRDF over RDF graphs is about 2x higher than that of the state‐of‐the‐art RDF processing systems. Menghan Jia, Yiming Zhang 0003, Dongsheng Li 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Modeling and analysis of medical resource allocation based on Timed Colored Petri net
Wangyang Yu 0001, Menghan Jia, Xianwen Fang, Yao Lu 0021, Jianchun Xu |
Future Gener. Comput. Syst. | 2 |
| 2020 | TopoX: Topology Refactorization for Minimizing Network Communication in Graph ComputationsabstractEfficient graph partitioning is vital for high-performance graph-parallel systems. Traditional graph partitioning methods attempt to both minimize communication cost and guarantee load balancing in computation. However, the skewed degree distribution of natural graphs makes it difficult to simultaneously achieve the two objectives. This article proposes topology refactorization (TR), a topology-aware method allowing graph-parallel systems to separately handle the two objectives: refactorization is mainly focused on reducing communication cost, and partitioning is mainly targeted for balancing the load. TR transforms a skewed graph into a more communication-efficient topology through fusion and fission, where the fusion operation organizes a set of neighboring low-degree vertices into a super-vertex, and the fission operation splits a high-degree vertex into a set of sibling sub-vertices. Based on TR, we design an efficient graph-parallel system (TopoX) which pipelines refactorization with partitioning to both reduce communication cost and balance computation load. Prototype evaluation shows that TopoX outperforms PowerLyra by up to 78.5% (from 37.2%) on real-world graphs and is faster than most other graph-parallel systems, while only introducing small overhead and memory consumption. Yiming Zhang 0003, Menghan Jia, Dongsheng Li 0001, Guangtao Xue, Kian-Lee Tan |
IEEE/ACM Trans. Netw. | 3 |
| 2018 | PruX: Communication Pruning of Parallel BFS in the Graph 500 Benchmark
Menghan Jia, Yiming Zhang 0003, Dongsheng Li 0001, Songzhu Mei |
ICA3PP (1) | 1 |