EDBT 2026 Demo / reviewers in the wild / expert
Jinhui Wei
dblp:294/6332
· DBLP profile ↗
13ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-9850-8384ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell SimulationsabstractParticle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model. Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen 0001, Yutong Lu |
EuroSys | 6 |
| 2026 | POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and CommunicationabstractParticle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a \(99.1\%\) overlap ratio. In cross-platform comparisons, POLAR-PIC achieves \(13.2\%\) of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches \(9.6\%\) on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains \(67.5\%\) weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems. Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Languang Gao, Zhiguang Chen 0001, Yutong Lu |
HPDC | 7 |
| 2026 | ASM-SpMM: Unleashing the Potential of Arm SME for Sparse Matrix Multiplication AccelerationabstractSparse Matrix–Matrix Multiplication (SpMM) is a core kernel in scientific computing, data analytics, and artificial intelligence, supporting applications such as linear solvers and Graph Neural Networks (GNNs). The Scalable Matrix Extension (SME) in Armv9 introduces dedicated matrix acceleration for ARM CPUs, but exploiting its full potential for SpMM requires architecture-aware optimizations to address irregular sparsity and hardware constraints. Jiazhi Jiang, Xijia Yao, Jinhui Wei, Dan Huang 0001, Yutong Lu |
PPoPP | 4 |
| 2025 | Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core ParallelismabstractIn-situ LLM inference on end-user devices has gained significant interest due to its privacy benefits and reduced dependency on external infrastructure. However, as the decoding process is memory-bandwidth-bound, the diverse processing units in modern end-user devices cannot be fully exploited, resulting in slow LLM inference. This paper presents Ghidorah, an LLM inference system for end-user devices with the unified memory architecture. The key idea of Ghidorah can be summarized in two steps: 1) leveraging speculative decoding approaches to enhance parallelism, and 2) ingeniously distributing workloads across multiple heterogeneous processing units to maximize computing power utilization. Ghidorah includes the hetero-core model parallelism (HCMP) architecture and the architecture-aware profiling (ARCA) approach. The HCMP architecture guides partitioning by leveraging the unified memory design of end-user devices and adapting to the hybrid computational demands of speculative decoding. The ARCA approach is used to determine the optimal speculative strategy and partitioning strategy, balancing acceptance rate with parallel capability to maximize the speedup. Additionally, we optimize sparse computation on ARM CPUs. Experimental results show that Ghidorah can achieve up to$7.6 \times$speedup in the dominant LLM decoding phase compared to the sequential decoding approach on NVIDIA Jetson NX. Jinhui Wei, Yuhui Zhou, Jiazhi Jiang, Jiangsu Du |
ICCD | 1 |
| 2025 | IasRT: Interference-Aware and SLO-Driven GPU Scheduling for Real-Time DNN InferenceabstractDeep Neural Network (DNN) inference has become a cornerstone of latency-sensitive applications such as autonomous driving and augmented reality. While GPUs offer high throughput for DNN inference, they often suffer from underutilization due to coarse-grained scheduling and limited concurrency. Existing GPU-sharing methods either lack awareness of fine-grained kernel interference or fail to meet service-level objectives (SLOs) under multi-tenant, multi-priority workloads. In this paper, we propose IasRT, a runtime framework that enables interferenceaware and SLO-driven GPU sharing for real-time DNN inference. IasRT profiles kernel-level resource usage and interference sensitivity, and dynamically partitions GPU streaming multiprocessors (SMs) to collocate jobs with minimal performance degradation. Furthermore, it introduces a dynamic SLO controller to maintain latency targets for multiple latency-sensitive (LS) jobs simultaneously. Evaluations on real-world DNN workloads show that IasRT reduces the 99th percentile latency of LS jobs by up to 38% compared to the state-of-the-art GPU sharing methods, while maintaining similar overall throughput from multiple collocated workloads, demonstrating its effectiveness in high-concurrency environments. Heming Zhong, Jinhui Wei, Yujia Fu, Dan Huang 0001, Yutong Lu |
ICCD | 2 |
| 2025 | Deep reinforcement learning for dynamic strategy interchange in financial markets
Xingyu Zhong, Jinhui Wei, Qingzhen Xu |
Appl. Intell. | 2 |
| 2024 | Enhancing Embedding and Hierarchical Reward Shaping for Multi-Hop Reasoning with Reinforcement Learning
Jinhui Wei, Jicheng Yu |
ADMA (2) | 2 |
| 2024 | Efficient Coupling Streaming AI and Ensemble Simulations on HPC Clusters
Jiazhi Jiang, Hongbin Zhang 0006, Deyin Liu, Jiangsu Du, Xiaojiao Yao, Jinhui Wei, Pin Chen, Dan Huang 0001, Yutong Lu |
Euro-Par (1) | 6 |
| 2024 | GCCR: GAT-Based Category-Aware Course Recommendation
Xiaohuan Xu, Wenjun Ma, Jinhui Wei, Suqin Tang |
KSEM (4) | 3 |
| 2024 | Liger: Interleaving Intra- and Inter-Operator Parallelism for Distributed Large Model InferenceabstractDistributed large model inference is still in a dilemma where balancing cost and effect. The online scenarios demand intraoperator parallelism to achieve low latency and intensive communications makes it costly. Conversely, the inter-operator parallelism can achieve high throughput with much fewer communications, but it fails to enhance the effectiveness. Jiangsu Du, Jinhui Wei, Jiazhi Jiang, Shenggan Cheng, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
PPoPP | 2 |
| 2024 | Community detection in attributed networks via adaptive deep nonnegative matrix factorization
Junwei Cheng, Yong Tang 0001, Chaobo He, Kunlin Han, Ying Li 0081, Jinhui Wei |
Neural Comput. Appl. | 6 |
| 2022 | Attention-based bi-directional refinement network for salient object detection
JunBin Yuan, Jinhui Wei, Kanoksak Wattanachote, Qingzhen Xu, Yongyi Gong |
Appl. Intell. | 2 |
| 2020 | Dynamic GMMU Bypass for Address Translation in Multi-GPU Systems
Jinhui Wei, Jianzhuang Lu, Yunping Zhao |
NPC | 1 |