Yanjie Zhen

dblp:238/1870 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Intelligent Hybrid Memory Scheduling Based on Page Pattern Recognition
abstract
Hybrid memory systems exhibit disparities in their heterogeneous memory components' access speeds. Dynamic page scheduling to ensure memory access predominantly occurs in the faster memory components is essential for optimizing the performance of hybrid memory systems. Recent works attempt to optimize page scheduling by predicting their hotness using neural network models. However, they face two crucial challenges: the page explosion problem and the new pages problem. We propose an intelligent hybrid memory scheduler driven by page pattern recognition to address these two challenges. Experimental results demonstrate that our approach outperforms state-of-the-art intelligent schedulers regarding effectiveness and cost.
Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004
DATE1
2024 PatternS: An intelligent hybrid memory scheduler driven by page pattern recognition
Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004
J. Syst. Archit.1
2023 Automatic Deep Learning Operator Fusion on Sunway SW26010 Many-Core Processor
abstract
Deep learning networks (DNNs) have been growing rapidly in recent years, with increasing demands on computing power. Therefore, accelerating the execution of DNN models has become a research hotspot. Operator fusion is a critical optimization strategy to enhance DNN performance in Deep Learning (DL) frameworks, such as TensorFlow, Pytorch, TVM and Halide. However, these frameworks are designed for general optimization and cannot fully harness the specific features of emerging hardware. Moreover, they primarily implement operator fusion at the operator level, missing out on many fusion opportunities and heavily relying on extensive manual optimizations for fused operators. Targeting the Sunway SW26010 Many-Core processor, the basic building block of Sunway TaihuLight supercomputer, we introduce swAutoFuser, an end-to-end automatic operator fusion and code generation framework. swAutoFuser proposes a set of low-level primitives to leverage hardware features and employs an autofuser to achieve primitive level fusion, which breaks operator boundaries and enables more fusion opportunities. In addition, swAutoFuser can automatically generate high-performance fused operator implementations based on a static cost model, significantly reducing the overhead of manually optimizing fused operators. Our experiments demonstrate that swAutoFuser can improve operator performance by 10% to 56%.
Wenxiang Zhang, Wenzhao Wu, Yanjie Zhen, Wenlai Zhao, Guangwen Yang 0002
ICPADS4
2023 A Low-Cost and Pages-Interrelation-Aware Attention Model for Hybrid Memory Scheduling
abstract
Hybrid memory architecture has become an important solution to address the increasing demand for the main memory capacity of big data applications. Due to the varying properties of different components in hybrid memory, accurately predicting the hotness of pages and timely scheduling hot pages to fast memory becomes crucial for optimal performance. However, existing hybrid memory schedulers using non-intelligent policy exhibit low performance. Although schedulers employing neural models can improve performance, they suffer limitations such as long inference time and loss of interrelation between pages. This paper presents PI-Attention, a low-cost and pages-interrelation-aware attention model for hybrid memory scheduling. It addresses the limitations above by utilizing two attention modules in the page and time sequence dimensions. Our experiments show that PI-Attention brings 11.14% performance improvement and a 3.75x reduction in inference time.
Yanjie Zhen, Yu Chen 0004
SMC1