VLDB 2026 Research / reviewers in the wild / expert
Ximeng Fu
dblp:430/7450
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware accelerators and domain-specific architectures · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator |
1.0 | 1 | 2026 | AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
convolution optimization |
1.0 | 1 | 2026 | AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › convolution optimization
winograd convolution |
1.0 | 1 | 2026 | AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.3 | 1 | 2026 | AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 Processors · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
winograd convolution · 2.0micro-kernel optimization · 2.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AirWino: Optimized Winograd Convolution for Accelerating CNN Inference on ARMv8 ProcessorsabstractAs Convolutional Neural Networks (CNNs) continue to gain traction in deep learning, Winograd convolution has emerged as a key algorithm to enhance computational efficiency. Although ARM-based CPUs are increasingly prevalent in mobile devices, embedded systems and HPC servers, existing 2D Winograd convolution implementations for ARM often leave room for improvement in transformation efficiency, computational throughput, and overall versatility. Furthermore, the lack of tailored 3D Winograd convolution implementations for ARM architectures stems from the additional complexity of supporting higher-dimensional kernels. AirWino introduces a set of novel optimizations covering transformations, data layouts, micro-kernel computations, and parallelization strategies for both 2D and 3D Winograd convolution. It supports FP32 and FP16 precisions with filter sizes of 3 and 5, targeting a broad range of applications. Evaluations on four distinct ARM platforms show that AirWino consistently outperforms state-of-the-art libraries across various experimental scenarios and hardware configurations, highlighting its efficiency and portability. Haoyuan Gui, Ximeng Fu, Leisheng Li |
AAAI | 4 |