EDBT 2026 Demo / reviewers in the wild / expert
Jiabin Xie
dblp:90/5279
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-0770-3086ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell SimulationsabstractParticle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model. Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen 0001, Yutong Lu |
EuroSys | 3 |
| 2026 | POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and CommunicationabstractParticle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a \(99.1\%\) overlap ratio. In cross-platform comparisons, POLAR-PIC achieves \(13.2\%\) of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches \(9.6\%\) on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains \(67.5\%\) weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems. Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Languang Gao, Zhiguang Chen 0001, Yutong Lu |
HPDC | 4 |
| 2025 | HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAabstractStencil computations are fundamental to various HPC and intelligent computing applications, often consuming significant execution time. The emergence of specialized matrix units presents new opportunities to accelerate stencil computations. While scalable matrix compute units provide substantial computing horsepower, prior efforts fail to fully utilize the computing capabilities for stencils due to suboptimal matrix-unit utilization, limited instruction-level parallelism, and low cache hit rates. This paper introduces HStencil, a novel stencil computing framework utilizing matrix and vector units. HStencil addresses these challenges through three contributions: 1) microkernels that jointly leverage matrix and vector units to enhance hardware utilization; 2) fine-grained instruction scheduling with interleaved execution to enhance instruction-level parallelism; and 3) spatial prefetch to sustain high performance when working sets exceed cache capacity. Evaluations on representative benchmarks demonstrate that HStencil achieves maximum speedups of 1.81x – 5.76x over auto-vectorization across different CPU platforms, delivers 31% - 91% higher performance versus state-of-the-art methods. Jiabin Xie, Guangnan Feng, Xianwei Zhang 0001, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
SC | 2 |
| 2024 | Extreme-scale Direct Numerical Simulation of Incompressible Turbulence on the Heterogeneous Many-core SystemabstractDirect numerical simulation (DNS) is a technique that directly solves the fluid Navier-Stokes equations with high spatial and temporal resolutions, which has driven much research regarding the nature of turbulence. For high-Reynolds number (Re) incompressible turbulence of particular interest, where the nondimensional Re characterizes the flow regime, the application of DNS is hindered by the fact that the numerical grid size (i.e., the memory requirement) scales with Re3, while the overall computational cost scales with Re4. Recent studies have shown that developing efficient parallel methods for heterogeneous many-core systems is promising to solve this computational challenge. Jiabin Xie, Guangnan Feng, Junxuan Feng, Zhiguang Chen 0001, Yutong Lu |
PPoPP | 1 |
| 2024 | UNR: Unified Notifiable RMA Library for HPCabstractRemote Memory Access (RMA) enables direct access to remote memory to achieve high performance for HPC applications. However, most modern parallel programming models lack schemes for the remote process to detect the completion of RMA operations. Many previous works have proposed programming models and extensions to notify the communication peer, but they did not solve the multi-NIC aggregation, portability, hardware-software co-design, and usability problems. In this work, we proposed a Unified Notifiable RMA (UNR) library for HPC to address these challenges. In addition, we demonstrate the best practice of utilizing UNR within a real-world scientific application, PowerLLEL. We deployed UNR across four HPC systems, each with a different interconnect. The results show that PowerLLEL powered by UNR achieves up to a 36% acceleration on 1728 nodes of the Tianhe-Xingyi supercomputing system. Guangnan Feng, Jiabin Xie, Dezun Dong, Yutong Lu |
SC | 2 |
| 2000 | An ICU Protocol Development and Management SystemabstractPatient care is commonly managed by clinicians using general guidelines or specific protocols. Acute, severe illness may require care in an intensive care unit (ICU), with many aspects of care managed concurrently. Variation of aspects of care among patients, clinicians, ICUs and institutions is well-known, and is due to a lack of standardization in clinical decision making and the inability to implement detailed, patient-responsive protocols. Automation of protocols for specific aspects of care using computing technology and medical informatics/knowledge engineering principles offers a consistent, systematic way to implement a protocol by prompting bedside clinicians in a timely manner monitoring a patient's responses to therapeutic interventions, and monitoring the decision logic used. This paper describes a new generic system for the design and implementation of computerized protocols. Jiabin Xie, Adwait Nerlikar, John R. Glover, Bruce A. McKinley |
CBMS | 1 |