VLDB 2026 Research / reviewers in the wild / expert
Qingcai Jiang
dblp:262/1511
· DBLP profile ↗
15ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0002-9729-8821ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 6 first-author · 14 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State DrivesabstractNear-data processing (NDP) mitigates the data movement bottleneck in modern computing systems by performing computation close to where the data resides. Solid-state drives (SSDs) are well suited for NDP because they: (1) store large application datasets that exceed main memory capacity, and (2) contain multiple heterogeneous computation resources, e.g., general-purpose embedded cores in the SSD controller, DRAM chips, and NAND flash chips, which enable three NDP paradigms: in-storage processing (ISP), processing using DRAM in the SSD (PuD-SSD), and in-flash processing (IFP). These resources offer massive internal parallelism and enable in-place computation, which reduces data movement across the memory hierarchy. A large body of prior SSD-based NDP techniques operate in isolation, mapping computations to only one or two NDP paradigms (i.e., ISP, PuD-SSD, or IFP) within the SSD. These techniques (1) are tailored to specific workloads or kernels, (2) do not offload computations across all three NDP paradigms in the SSD and thus fail to exploit the full computational potential of an SSD, and (3) lack programmer-transparency, often requiring significant manual effort to identify offloadable code regions and map them to the SSD computation resources, which limits their general applicability and ease of deployment. While several prior works propose techniques to partition computation between the host and near-memory accelerators, adapting these techniques to SSDs offers limited benefits because they (1) ignore the heterogeneity of the SSD computation resources, and (2) make offloading decisions based on limited factors such as bandwidth utilization, data movement cost, or memory intensity, while ignoring key factors such as resource utilization. We propose Conduit, a general-purpose, programmertransparent NDP framework for SSDs that accelerates a broad range of workloads by leveraging available SSD computation resources. Conduit operates in two stages. At compile time, Conduit executes a custom compiler (e.g., LLVM) pass that (i) vectorizes suitable application code segments into single-instruction multiple-data (SIMD) operations that align with the SSD's page layout, and (ii) embeds metadata (e.g., operation type, operand sizes) into the vectorized instructions to guide runtime offloading decisions. At runtime, within the SSD, Conduit performs instruction-granularity offloading by evaluating six key application and system features (e.g., operation type, computation resource utilization, data dependence delay), and uses a cost function to select the most suitable SSD computation resource to execute each vectorized instruction. We evaluate Conduit and two prior NDP offloading techniques using an in-house event-driven SSD simulator on six data-intensive applications (e.g., large language model inference and training, encryption). Conduit outperforms the best-performing prior offloading policy by$1.8 \times$and reduces energy consumption by 46 %, with small latency and storage overheads, and no additional hardware cost. Rakesh Nadig, Vamanan Arulchelvan, Mayank Kabra, Harshita Gupta, Rahul Bera, Nika Mansouri-Ghiasi, Nanditha Rao, Qingcai Jiang, Andreas Kosmas Kakolyris, Yu Liang 0004, Mohammad Sadrosadati, Onur Mutlu |
HPCA | 8 |
| 2025 | NDFT: Accelerating Density Functional Theory Calculations via Hardware/Software Co-Design on Near-Data Computing SystemabstractLinear-response time-dependent Density Functional Theory (LR-TDDFT) is a widely used method for accurately predicting the excited-state properties of physical systems. Previous works have attempted to accelerate LR-TDDFT using heterogeneous systems such as GPUs, FPGAs, and the Sunway architecture. However, a major drawback of these approaches is the constant data movement between host memory and the memory of the heterogeneous systems, which results in substantial data movement overhead. Moreover, these works focus primarily on optimizing the compute-intensive portions of LR-TDDFT, despite the fact that the calculation steps are fundamentally memory-bound.To address these challenges, we propose NDFT, a Near-Data Density Functional Theory framework. Specifically, we design a novel task partitioning and scheduling mechanism to offload each part of LR-TDDFT to the most suitable computing units within a CPU-NDP system. Additionally, we implement a hardware/software co-optimization of a critical kernel in LR-TDDFT to further enhance performance on the CPU-NDP system. Our results show that NDFT achieves performance improvements of 5.2x and 2.5x over CPU and GPU baselines, respectively, on a large physical system. Qingcai Jiang, Buxin Tu, Junshi Chen 0003, Hong An |
DAC | 1 |
| 2025 | NDPage: Efficient Address Translation for Near-Data Processing Architectures via Tailored Page TableabstractNear-Data Processing (NDP) has been a promising architectural paradigm to address the memory wall problem for data-intensive applications. Practical implementation of NDP architectures calls for system support for better programmability, where having virtual memory (VM) is critical. Modern computing systems incorporate a 4-level page table design to support address translation in VM. However, simply adopting an existing 4-level page table in NDP systems causes significant address translation overhead because (1) NDP applications generate a lot of address translations, and (2) the limited L1 cache in NDP systems cannot cover the accesses to page table entries (PTEs). We extensively analyze the 4-level page table design in the NDP scenario and observe that (1) the memory access to page table entries is highly irregular, thus cannot benefit from the L1 cache, and (2) the last two levels of page tables are nearly fully occupied. Based on our observations, we propose NDPage, an efficient page table design tailored for NDP systems. The key mechanisms of NDPage are (1) an L1 cache bypass mechanism for PTEs that not only accelerates the memory accesses of PTEs but also prevents the pollution of PTEs in the cache system, and (2) a flattened page table design that merges the last two levels of page tables, allowing the page table to enjoy the flexibility of a 4KB page while reducing the number of PTE accesses. We evaluate NDPage using a variety of data-intensive work-loads. Our evaluation shows that in a single-core NDP system, NDPage improves the end-to-end performance over the state-of-the-art address translation mechanism of 14.3 %; in 4-core and 8-core NDP systems, NDPage enhances the performance of 9.8% and 30.5 %, respectively. Qingcai Jiang, Buxin Tu, Hong An |
DATE | 1 |
| 2025 | Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile DevicesabstractAs the memory demands of individual mobile applications continue to grow and the number of concurrently running applications increases, available memory on mobile devices is becoming increasingly scarce. When memory pressure is high, current mobile systems use a RAM-based compressed swap scheme (called ZRAM) to compress unused execution-related data (called anonymous data in Linux) in main memory. This approach avoids swapping data to secondary storage (NAND flash memory) or terminating applications, thereby achieving shorter application relaunch latency.In this paper, we observe that the state-of-the-art ZRAM scheme prolongs relaunch latency and wastes CPU time because it does not differentiate between hot and cold data or leverage different compression chunk sizes and data locality. We make three new observations. First, anonymous data has different levels of hotness. Hot data, used during application relaunch, is usually similar between consecutive relaunches. Second, when compressing the same amount of anonymous data, small-size compression is very fast, while large-size compression achieves a better compression ratio. Third, there is locality in data access during application relaunch.Based on these observations, we propose a hotness-aware and size-adaptive compressed swap scheme, Ariadne, for mobile devices to mitigate relaunch latency and reduce CPU usage. Ariadne incorporates three key techniques. First, a low-overhead hotness-aware data organization scheme aims to quickly identify the hotness of anonymous data without significant overhead. Second, a size-adaptive compression scheme uses different compression chunk sizes based on the data’s hotness level to ensure fast decompression of hot and warm data. Third, a proactive decompression scheme predicts the next set of data to be used and decompresses it in advance, reducing the impact of data swapping back into main memory during application relaunch.We implement and evaluate Ariadne on a commercial smartphone, Google Pixel 7 with the latest Android 14. Our experimental evaluation results show that, on average, Ariadne reduces application relaunch latency by 50% and decreases the CPU usage of compression and decompression procedures by 15% compared to the state-of-the-art compressed swap scheme for mobile devices. Yu Liang 0004, Aofeng Shen, Chun Jason Xue, Riwei Pan, Haiyu Mao, Nika Mansouri-Ghiasi, Qingcai Jiang, Rakesh Nadig, Lei Li 0067, Rachata Ausavarungnirun, Mohammad Sadrosadati, Onur Mutlu |
HPCA | 7 |
| 2025 | CIExplorer: Microarchitecture-Aware Exploration for Tightly Integrated Custom InstructionabstractExtending existing architectures with customized instruction extensions is emerging to achieve high performance and energy efficiency for specific applications.Automated discovery of custom instructions (CIs) is well-studied nowadays, which requires exploring combinations of different types and quantities of operations, resulting in a vast search space.However, previous works typically use microarchitectureagnostic cost models, leading to suboptimal CIs that may degrade performance.They leverage graph isomorphism to reduce area overhead, but few of them consider its potential to benefit performance-oriented exploration.To this end, we present CIExplorer, a framework for adaptive CI exploration. Qingcai Jiang, Jun Shi 0007, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan |
ICS | 4 |
| 2025 | PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway SupercomputerabstractFirst-principles density functional theory (DFT) with plane wave (PW) basis set is the most widely used method in quantum mechanical material simulations due to its advantages in accuracy and universality. However, a perceived drawback of PW-based DFT calculations is their substantial computational cost and memory usage, which currently limits their ability to simulate large-scale complex systems containing thousands of atoms. This situation is exacerbated in the new Sunway supercomputer, where each process is limited to a mere 16 GB of memory. Herein, we present a novel parallel implementation of plane wave density functional theory on the new Sunway supercomputer (PWDFT-SW). PWDFT-SW fully extracts the benefits of Sunway supercomputer by extensively refactoring and calibrating our algorithms to align with the system characteristics of the Sunway system. Through extensive numerical experiments, we demonstrate that our methods can substantially decrease both computational costs and memory usage. Our optimizations translate to a speedup of 64.8x for a physical system containing 4,096 silicon atoms, enabling us to push the limit of PW-based DFT calculations to large-scale systems containing 16,384 carbon atoms. Qingcai Jiang, Zhenwei Cao, Junshi Chen 0003, Xinming Qin, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2024 | A3PIM: An Automated, Analytic and Accurate Processing-in-Memory OffloaderabstractThe performance gap between memory and processor has grown rapidly. Consequently, the energy and wall-clock time costs associated with moving data between the CPU and main memory predominate the overall computational cost. The Processing-in-Memory (PIM) paradigm emerges as a promising architecture that mitigates the need for extensive data movements by strategically positioning computing units proximate to the memory. Despite the abundant efforts devoted to building a robust and highly-available PIM system, identifying PIM-friendly segments of applications poses significant challenges due to the lack of a comprehensive tool to evaluate the intrinsic memory access pattern of the segment. To tackle this challenge, we propose A3PIM11The code and benchmarks of this work are opened-sourced in: https://github.com/ACSA-PIM/A3PIM: an Automated, Analytic and Accurate Processing-in-Memory offloader. We sys-tematically consider the cross-segment data movement and the intrinsic memory access pattern of each code segment via static code analyzer. We evaluate A 3PIM across a wide range of real-world workloads including GAP and PrIM benchmarks and achieve an average speedup of 2.63x and 4.45x (up to 7.14x and lO.64x) when compared to CPU-only and PIM-only executions, respectively. Qingcai Jiang, Shaojie Tan, Junshi Chen 0003, Hong An |
DATE | 1 |
| 2024 | Enabling 13K-Atom Excited-State GW Calculations via Low-Rank Approximations and HPC on the New Sunway SupercomputerabstractGW approximation is a powerful approach to accurately describe the excited-state of semiconductors. However, GW incurs high computational cost $\mathcal{O}\left(N^{4}\right)$ and large memory usage $\mathcal{O}\left(N^{3}\right)$, limiting its applications to thousands of (2,742) atoms even on leadership supercomputers. Herein we present a massively parallel implementation of accurate and efficient cubic-scaling plane-wave GW calculations by using low-rank approximations and high-performance computing on leadership supercomputers. By using a series of low rank approximations, we can reduce the expensive GW calculations to the cubic-scaling computational cost $\mathcal{O}\left(N^{3}\right)$ and quadratic memory usage $\mathcal{O}\left(N^{2}\right)$. With the help of parallel and communication optimization, the plane-wave GW calculations gain an overall speedup of over 70x and efficiently scale up to 13,824 atoms within a few minutes using 449,280 cores on new Sunway supercomputer. This accomplishment paves the way for excited-state quantum mechanical material simulations at mesoscopic scale (10K atoms) and for the design of next-generation semiconductor devices. Wentiao Wu, Zhengbang Zhou, Qingcai Jiang, Junwei Feng, Xinming Qin, Huanhuan Ma, Zhenwei Cao, Junshi Chen 0003, Xinyong Meng, Bingkun Hou, Yuanfan Xiong, Linhao Wang, Yixuan Sun, Hong An, Jinlong Yang 0003, Wei Hu 0006 |
SC | 3 |
| 2024 | Uncovering the performance bottleneck of modern HPC processor with static code analyzer: a case study on Kunpeng 920
Shaojie Tan, Qingcai Jiang, Zhenwei Cao, Junshi Chen 0003, Hong An |
CCF Trans. High Perform. Comput. | 2 |
| 2024 | Extending the limit of LR-TDDFT on two different approaches: Numerical algorithms and new Sunway heterogeneous supercomputerabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics , computational chemistry, and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost O ( N 5 ∼ N 6 ) and large memory usage O ( N 4 ) especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and accelerate LR-TDDFT in two different aspects: (1) numerical algorithms on the X86 supercomputer and (2) optimizations on the heterogeneous architecture of the new Sunway supercomputer. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate the physical nature of different computation steps. By utilizing these two different methods, our implementation can gain an overall speedup of 10x and 80x and efficiently scales to large systems up to 4096 and 2744 atoms within dozens of seconds. Qingcai Jiang, Zhenwei Cao, Xinhui Cui, Lingyun Wan, Xinming Qin, Huanqi Cao, Hong An, Junshi Chen 0003, Jie Liu 0069, Wei Hu 0006, Jinlong Yang 0003 |
Parallel Comput. | 1 |
| 2023 | High performance computing for first-principles Kohn-Sham density functional theory towards exascale supercomputers
Xinming Qin, Junshi Chen 0003, Zhaolong Luo, Lingyun Wan, Jielan Li, Shizhe Jiao, Qingcai Jiang, Wei Hu 0006, Hong An, Jinlong Yang 0003 |
CCF Trans. High Perform. Comput. | 8 |
| 2022 | Accelerating Parallel First-Principles Excited-State Calculation by Low-Rank Approximation with K-Means ClusteringabstractFirst-principles time-dependent density functional theory (TDDFT) is a powerful tool to accurately describe the excited-state properties of molecules and solids in condensed matter physics, computational chemistry and materials science. However, a perceived drawback in TDDFT calculations is its ultrahigh computational cost and large memory usage especially for plane-wave basis set, confining its applications to large systems containing thousands of atoms. Here, we present a massively parallel implementation of linear-response TDDFT (LR-TDDFT) and reduce the complexity to by combining K-Means clustering based low-rank approximation with iterative eigensolve algorithm. Furthermore, we carefully design the parallel data and task distribution schemes to accommodate with the physical nature in different steps of the computation, also, several optimization methods are employed to effectively handle the matrix operations and data communications of constructing and diagonalizing the LR-TDDFT Hamiltonian. In particular, our method can significantly reduce the cost of computation and memory by nearly 2 orders of magnitude compared to conventional LR-TDDFT calculations. Numerical results demonstrate that our implementation can gain an overall speedup of 10x and efficiently scale up to 12,288 CPU cores for large systems up to 4,096 atoms within dozens of seconds. Qingcai Jiang, Jielan Li, Junshi Chen 0003, Xinming Qin, Lingyun Wan, Jinlong Yang 0003, Jie Liu 0069, Wei Hu 0006, Hong An |
ICPP | 1 |
| 2022 | 2.5 Million-Atom Ab Initio Electronic-Structure Simulation of Complex Metallic Heterostructures with DGDFTabstractOver the past three decades, ab initio electronic structure calculations of large, complex and metallic systems are limited to tens of thousands of atoms in computational accuracy and efficiency on leadership supercomputers. We present a massively parallel discontinuous Galerkin density functional theory (DGDFT) implementation, which adopts adaptive local basis functions to discretize the Kohn-Sham equation, resulting in a block-sparse Hamiltonian matrix. A highly efficient pole expansion and selected inversion (PEXSI) sparse direct solver is implemented in DGDFT to achieve O(N1.5) scaling for quasi two-dimensional systems. DGDFT allows us to compute the electronic structures of complex metallic heterostructures with 2.5 million atoms (17.2 million electrons) using 35.9 million cores on the new Sunway supercomputer. The peak performance of PEXSI can achieve 64 PFLOPS (~5% of theoretical peak), which is un-precedented for sparse direct solvers. This accomplishment paves the way for quantum mechanical simulations into mesoscopic scale for designing next-generation electronic devices. Wei Hu 0006, Hong An, Zhuoqiang Guo, Qingcai Jiang, Xinming Qin, Junshi Chen 0003, Weile Jia, Chao Yang 0001, Zhaolong Luo, Jielan Li, Wentiao Wu, Guangming Tan, Dongning Jia, Qinglin Lu, Yeqi Huang, Liyi Wang, Jinlong Yang 0003 |
SC | 4 |
| 2022 | Bridging the Gap between Deep Learning and Frustrated Quantum Spin System for Extreme-Scale Simulations on New Generation of Sunway SupercomputerabstractEfficient numerical methods are promising tools for delivering unique insights into the fascinating properties of physics, such as the highly frustrated quantum many-body systems. However, the computational complexity of obtaining the wave functions for accurately describing the quantum states increases exponentially with respect to particle number. Here we present a novel convolutional neural network (CNN) for simulating the two-dimensional highly frustrated spin-$1/2$$J_1-J_2$Heisenberg model, meanwhile the simulation is performed at an extreme scale system with low cost and high scalability. By ingenious employment of transfer learning and CNN’s translational invariance, we successfully investigate the quantum system with the lattice size up to$24\times 24$, within 30 million cores of the new generation of sunway supercomputer. The final achievement demonstrates the effectiveness of CNN-based representation of quantum-state and brings the state-of-the-art record up to a brand-new level from both aspects of remarkable accuracy and unprecedented scales. Mingfan Li, Junshi Chen 0003, Qingcai Jiang, Xuncheng Zhao, Rongfen Lin, Hong An, Lixin He |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | RDMA-Based Apache Storm for High-Performance Stream Data Processing
Ziyu Zhang 0003, Zitan Liu, Qingcai Jiang, Junshi Chen 0003, Hong An |
NPC | 3 |