Jiapeng Guan

dblp:375/4909 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
0009-0006-6850-8785ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 FlexStep: Enabling Flexible Error Detection in Multi/Many-core Real-time Systems
abstract
Reliability and real-time responsiveness in safety-critical systems have traditionally been achieved using error detection mechanisms, such as LockStep, which require pre-configured checker cores, strict synchronisation, static error detection regions, or limited preemptions. However, these core-bound hardware mechanisms often lead to significant resource over-provisioning and diminished real-time performance in modern systems where tasks with varying reliability requirements are consolidated on shared processors for efficiency and cost reduction. To address these challenges, this work presents FlexStep, a systematic solution that integrates hardware and software across the SoC, ISA, and OS scheduling layers. FlexStep features a novel microarchitecture that supports dynamic core configuration and asynchronous, preemptive error detection. The FlexStep architecture naturally allows for flexible task scheduling and error detection, enabling new scheduling algorithms that enhance both resource efficiency and real-time schedulability.
Tinglue Wang, Jiapeng Guan, Zhenghui Guo, Renshuang Jiang, Jing Li 0025, Zhe Jiang 0004
DAC4
2025 NVR: Vector Runahead on NPUs for Sparse Memory Access
abstract
Deep Neural Networks are increasingly leveraging sparsity to reduce the scaling up of model parameter size. However, reducing wall-clock time through sparsity and pruning remains challenging due to irregular memory access patterns, leading to frequent cache misses. In this paper, we present NPU Vector Runahead (NVR), a prefetching mechanism tailored for NPUs to address cache miss problems in sparse DNN workloads. Rather than optimising memory patterns with high overhead and poor portability, NVR adapts runahead execution to the unique architecture of NPUs. NVR provides a general micro-architectural solution for sparse DNN workloads without requiring compiler or algorithmic support, operating as a decoupled, speculative, lightweight hardware sub-thread alongside the NPU, with minimal hardware overhead (under 5%). NVR achieves an average 90% reduction in cache misses compared to SOTA prefetching in general-purpose processors, delivering 4 x average speedup on sparse workloads versus NPUs without prefetching. Moreover, we investigate the advantages of incorporating a small cache (16 KB) into the NPU combined with NVR. Our evaluation shows that expanding this modest cache delivers 5x higher performance benefits than increasing the $\mathbf{L 2}$ cache size by the same amount.
Hui Wang 0166, Zhengpeng Zhao, Jing Wang 0113, Yushu Du, Chenhao Ma 0006, Xiaomeng Han, Dean You, Jiapeng Guan, Zhe Jiang 0004
DAC11
2025 MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
abstract
Runahead execution is a technique to mask memory latency caused by irregular memory accesses. By pre-executing the application code during occurrences of long-latency operations and prefetching anticipated cache-missed data into the cache hierarchy, runahead effectively masks memory latency for subsequent cache misses and achieves high prefetching accuracy; however, this technique has been limited to superscalar out-of-order and superscalar in-order cores. For implementation in scalar in-order cores, the challenges of area-/energy-constraint and severe cache contention remain. Here, we build the first full-stack system featuring runahead, MERE , from SoC and a dedicated ISA to the OS and programming model. Through this deployment, we show that enabling runahead in scalar in-order cores is possible, with minimal area and power overheads, while still achieving high performance. By re-constructing the sequential runahead employing a hardware/software co-design approach, the system can be implemented on a mature processor and SoC. Building on this, an adaptive runahead mechanism is proposed to mitigate the severe cache contention in scalar in-order cores. Combining this, we provide a comprehensive solution for embedded processors managing irregular workloads. Our evaluation demonstrates that the proposed MERE attains 93.5% of a 2-wide out-of-order core’s performance while constraining area and power overheads below 5%, with the adaptive runahead mechanism delivering an additional 20.1% performance gain through mitigating the severe cache contention issues.
Dean You, Jieyu Jiang, Yushu Du, Zhihang Tan, Hui Wang 0166, Jiapeng Guan, Shuai Zhao 0004, Zhe Jiang 0004
ACM Trans. Embed. Comput. Syst.8
2024 MESC: Re-thinking Algorithmic Priority and/or Criticality Inversions for Heterogeneous MCSs
abstract
Modern Mixed-Criticality Systems (MCSs) rely on hardware heterogeneity to satisfy ever-increasing computational demands. However, most of the heterogeneous co-processors are designed to achieve high throughput, with their micro-architectures executing the workloads in a streaming manner. This streaming execution is often non-preemptive or limited-preemptive, preventing tasks’ prioritisation based on their importance and resulting in frequent occurrences of algorithmic priority and/or criticality inversions. Such problems present a significant barrier to guaranteeing the systems’ real-time predictability, especially when co-processors dominate the execution of the workloads (e.g., DNNs and transformers).In contrast to existing works that typically enable coarse-grained context switch by splitting the workloads/algorithms, we demonstrate a method that provides fine-grained context switch on a widely used open-source DNN accelerator by enabling instruction-level preemption without any workloads/algorithms modifications. As a systematic solution, we build a real system, i.e., Make Each Switch Count (MESC), from the SoC and ISA to the OS kernel. A theoretical model and analysis are also provided for timing guarantees. Experimental results reveal that, compared to conventional MCSs using non-preemptive DNN accelerators, MESC achieved a 250 x and 300 x speedup in resolving algorithmic priority and criticality inversions, with less than 5% overhead. To our knowledge, this is the first work investigating algorithmic priority and criticality inversions for MCSs at the instruction level.
Jiapeng Guan, Dean You, Yingquan Wang, Ruizhe Yang, Hui Wang 0166, Zhe Jiang 0004
RTSS1
2024 An efficient multi-task learning CNN for driver attention monitoring
abstract
Driver Monitoring System (DMS), usually equipped with a camera, is an emerging vehicle safety system that can monitor driver attentiveness and trigger timely alarms when signs of inattention are detected. Since a single indicator (e.g., eye blink rate) is insufficient and unreliable to analyze driver attentiveness, almost all existing solutions train several independent models to identify driver facial states, such as face landmark, head pose, yawning, eye state, etc. However, apart from neglecting the inherent correlations between these related tasks, multiple models also raise challenges for vehicle safety-critical systems (e.g., hardware resources, software compatibility, and real-time response). In this paper, we propose a multi-task learning CNN framework (DANet) to unify the relevant tasks into one model and simultaneously output various driver facial states. By sharing the common features and parameters of highly related tasks, DANet avoids repetitive computations and mitigates single task overfitting. More importantly, the model provides a comprehensive overview of facial states while maintaining low complexity. We also propose two novel designs: (1) Dual-loss Block, which decomposes the pose estimation task into pose classification and coarse-to-fine regression; (2) Head Pose Penalization, which constrains the network to predict gaze direction based on predicted head pose. Our method achieves compelling results in both speed and accuracy on a vehicle computing platform, marking a momentous step in this field.
Jiapeng Guan, Zhe Jiang 0004
J. Syst. Archit.4