Dean You

dblp:387/8608 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0002-4588-3794ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Hardware accelerators and domain-specific architectures · 41% Embedded and real-time systems · 26% Hardware reliability and fault tolerance · 15%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
error detection
0.912025
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural processing unit
0.912025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.812024
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024
Embedded and real-time systems › real-time scheduling › mixed-criticality scheduling
mixed-criticality systems
0.812024
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024
Embedded and real-time systems
real-time scheduling
0.812024
MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs · RTSS 2024
Memory systems
cache
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Memory systems › cache
cache miss reduction
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Memory systems › cache
prefetching
0.312025
NVR: Vector Runahead on NPUs for Sparse Memory Access · DAC 2025
Processor architecture and microarchitecture › superscalar processor
superscalar out-of-order processor
0.312025
MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors · DAC 2025

Methods — techniques the papers use, named apart from their topics

speculative execution · 0.9runahead execution · 0.9hardware-software co-design · 0.9RTL design · 0.9timing analysis · 0.8context switching · 0.8
YearPublicationVenuePosition
2025 MEEK: Re-thinking Heterogeneous Parallel Error Detection Architecture for Real-World OoO Superscalar Processors
abstract
Heterogeneous parallel error detection is an approach to achieving fault-tolerant processors, leveraging multiple power-efficient cores to re-execute software originally run on a high-performance core. Yet, its complex components, gathering data cross-chip from many parts of the core, raise questions of how to build it into commodity cores without heavy design invasion and extensive re-engineering. We build the first full-RTL design, MEEK, into an open-source SoC, from microarchitecture and ISA to the OS and programming model. We identify and solve bottlenecks and bugs overlooked in previous work, and demonstrate that MEEK offers microsecond-level detection capacity with affordable overheads. By trading off architectural functionalities across codesigned hardware-software layers, MEEK features only light changes to a mature out-of-order superscalar core, simple coordinating software layers, and a few lines of operating-system code. The Repo. of MEEK’s source code: https://github.com/SEU-ACAL/reproduce-MEEK-DAC-25
Zhe Jiang 0004, Minli Julie Liao, Sam Ainsworth 0001, Dean You, Timothy M. Jones 0001
DAC4
2025 NVR: Vector Runahead on NPUs for Sparse Memory Access
abstract
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4 x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16 KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the $\mathbf{L 2}$ cache size by the same amount.
Hui Wang 0166, Zhengpeng Zhao, Jing Wang 0113, Yushu Du, Chenhao Ma 0006, Xiaomeng Han, Dean You, Jiapeng Guan, Zhe Jiang 0004
DAC10
2025 MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
abstract
Runahead execution is a technique to mask memory latency caused by irregular memory accesses. By pre-executing the application code during occurrences of long-latency operations and prefetching anticipated cache-missed data into the cache hierarchy, runahead effectively masks memory latency for subsequent cache misses and achieves high prefetching accuracy; however, this technique has been limited to superscalar out-of-order and superscalar in-order cores. For implementation in scalar in-order cores, the challenges of area-/energy-constraint and severe cache contention remain. Here, we build the first full-stack system featuring runahead, MERE , from SoC and a dedicated ISA to the OS and programming model. Through this deployment, we show that enabling runahead in scalar in-order cores is possible, with minimal area and power overheads, while still achieving high performance. By re-constructing the sequential runahead employing a hardware/software co-design approach, the system can be implemented on a mature processor and SoC. Building on this, an adaptive runahead mechanism is proposed to mitigate the severe cache contention in scalar in-order cores. Combining this, we provide a comprehensive solution for embedded processors managing irregular workloads. Our evaluation demonstrates that the proposed MERE attains 93.5% of a 2-wide out-of-order core’s performance while constraining area and power overheads below 5%, with the adaptive runahead mechanism delivering an additional 20.1% performance gain through mitigating the severe cache contention issues.
Dean You, Jieyu Jiang, Yushu Du, Zhihang Tan, Hui Wang 0166, Jiapeng Guan, Shuai Zhao 0004, Zhe Jiang 0004
ACM Trans. Embed. Comput. Syst.1
2024 MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
abstract
Modern Mixed-Criticality Systems (MCSs) rely on hardware heterogeneity to satisfy ever-increasing computational demands. However, most of the heterogeneous co-processors are designed to achieve high throughput, with their micro-architectures executing the workloads in a streaming manner. This streaming execution is often non-preemptive or limited-preemptive, preventing tasks’ prioritisation based on their importance and resulting in frequent occurrences of algorithmic priority and/or criticality inversions. Such problems present a significant barrier to guaranteeing the systems’ real-time predictability, especially when co-processors dominate the execution of the workloads (e.g., DNNs and transformers).In contrast to existing works that typically enable coarse-grained context switch by splitting the workloads/algorithms, we demonstrate a method that provides fine-grained context switch on a widely used open-source DNN accelerator by enabling instruction-level preemption without any workloads/algorithms modifications. As a systematic solution, we build a real system, i.e., Make Each Switch Count (MESC), from the SoC and ISA to the OS kernel. A theoretical model and analysis are also provided for timing guarantees. Experimental results reveal that, compared to conventional MCSs using non-preemptive DNN accelerators, MESC achieved a 250 x and 300 x speedup in resolving algorithmic priority and criticality inversions, with less than 5% overhead. To our knowledge, this is the first work investigating algorithmic priority and criticality inversions for MCSs at the instruction level.
Jiapeng Guan, Dean You, Yingquan Wang, Ruizhe Yang, Hui Wang 0166, Zhe Jiang 0004
RTSS3