Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Qiukun Han

dblp:373/6998 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
0009-0003-7829-2963ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Processor architecture and microarchitecture · 62% High-performance computing · 38%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
vectorization
0.812024
Boost Linear Algebra Computation Performance via Efficient VNNI Utilization · ASPLOS (3) 2024
Processor architecture and microarchitecture › SIMD
SIMD instructions
0.812024
Boost Linear Algebra Computation Performance via Efficient VNNI Utilization · ASPLOS (3) 2024
High-performance computing › numerical linear algebra
dense linear algebra
0.212024
Boost Linear Algebra Computation Performance via Efficient VNNI Utilization · ASPLOS (3) 2024
High-performance computing
numerical linear algebra
0.212024
Boost Linear Algebra Computation Performance via Efficient VNNI Utilization · ASPLOS (3) 2024

Methods — techniques the papers use, named apart from their topics

peephole optimization · 1.5pattern matching · 1.5auto-vectorization · 1.5
YearPublicationVenuePosition
2024 Boost Linear Algebra Computation Performance via Efficient VNNI Utilization
abstract
Intel's Vector Neural Network Instruction (VNNI) provides higher efficiency on calculating dense linear algebra (DLA) computations than conventional SIMD instructions. However, existing auto-vectorizers frequently deliver suboptimal utilization of VNNI by either failing to recognize VNNI's unique computation pattern at the innermost loops/basic blocks, or producing inferior code through constrained and rudimentary peephole optimizations/pattern matching techniques. Auto-tuning frameworks might generate proficient code but are hampered by the necessity for sophisticated pattern templates and extensive search processes.
Hao Zhou 0009, Qiukun Han, Heng Shi 0005, Yalin Zhang 0004, Jianguo Yao 0002
ASPLOS (3)2