VLDB 2026 Research / reviewers in the wild / expert
Runxi He
dblp:406/2196
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 33% Hardware accelerators and domain-specific architectures · 33% Processor architecture and microarchitecture · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture
chip multiprocessor |
0.9 | 1 | 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs · IEEE Trans. Computers 2025 |
High-performance computing › tensor computation
convolution |
0.9 | 1 | 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs · IEEE Trans. Computers 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
direct convolution |
0.9 | 1 | 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs · IEEE Trans. Computers 2025 |
Methods — techniques the papers use, named apart from their topics
parallelization · 0.9operator fusion · 0.9data packing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUsabstractConvolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007 |
IEEE Trans. Computers | 6 |