Yuan Yao 0009

dblp:25/4120-9 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
11since 2021 · last 2026
0000-0001-9448-5595ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 NoCWalk: In-Network Page Walks for Efficient Pointer-Chasing Workloads on Multicores
Yuan Yao 0009, Rashid Aligholipour, Stefanos Kaxiras
CF1
2026 Efficient CNN Inference on Ultra-Low-Power MCUs via Saturation-Aware Convolution
abstract
Quantized CNN inference on ultra-low-power MCUs incurs unnecessary computations in neurons that produce saturated output values. These values are too extreme and are eventually clamped to the boundaries allowed by the neuron. Often times, the neuron can save time by only producing a value that is extreme enough to lead to the clamped result, instead of completing the computation, yet without introducing any error. Based on this, we present saturation-aware convolution: an inference technique whereby we alter the order of computations in convolution kernels to induce earlier saturation, and value checks are inserted to omit unnecessary computations when the intermediate result is sufficiently extreme. Our experimental results display up to 24% inference time saving on a Cortex-M0+ MCU, with zero impact on accuracy.
Luca Mottola, Yuan Yao 0009, Stefanos Kaxiras
DATE3
2026 DICE: Detailed Inter-Chiplet End-to-End Phy Modeling for Accurate Chiplet Simulation
Rashid Aligholipour, Stefanos Kaxiras, Yuan Yao 0009
ISCA3
2026 Understanding Simulated Architecture via gem5 Call-Stack Profiling
abstract
Understanding the behavior of simulated architectures in gem5 is critical for studying complex, deeply integrated computing systems. However, conventional analysis methods, which rely heavily on simulation statistics, provide only an indirect view of the simulated system internals. In this work, we show that call-stack profiling of gem5 itself offers a powerful yet underutilized perspective: the simulator’s own call-stack directly reflects the activity of the simulated system, exposing insights that conventional statistics may overlook.Profiling gem5’s call-stacks, however, is challenging due to its highly layered and complex software design patterns. To address this, we introduce a specialized, lightweight profiling framework built on Linux’s perf_event interface which samples and analyzes gem5’s runtime call-stacks throughout the simulation, resolves symbols on the fly, and merges samples into a hierarchical call-tree representation supporting both high-level structural views and focused, user-defined, component-specific analysis. Moreover, all profiling is performed in a dedicated helper process running alongside the main gem5 process, avoiding intrusive changes and overheads to the simulation itself.We apply our framework to gem5’s three major CPU modelsAtomicSimpleCPU, TimingSimpleCPU, and O3CPU-together with the Ruby memory system, and uncover behaviors that are not easily observable in conventional gem5 statistics. Our case studies reveal, for example, that TimingSimpleCPU is inefficient due to its use of a lockup-cache model and, despite its conceptual simplicity, does not simulate faster than a full out-of-order core. In addition, our tool makes it straightforward to detect cache coherence protocol deadlock and livelock-issues that are otherwise difficult to identify, since the simulation either appears to run normally or terminates abruptly, making it hard to pinpoint when these conditions occur.
Johan Söderström, Rashid Aligholipour, Yuan Yao 0009
ISPASS3
2026 ConvReflex: Efficient Ultra-Low-Power CNN Inference via Clamping Prediction
Luca Mottola, Yuan Yao 0009, Stefanos Kaxiras
SenSys3
2025 RXT: RefleXive Address Translation for Pointer-Chasing Workloads
abstract
With increasingly irregular memory access patterns in Indirect Memory Access (IMA) and Graph Processing (GP) applications, Virtual Address Translation (VAT) not only has become a major performance bottleneck, but also a key contributor to high power consumption in out-of-order (OoO) multi-core processor pipeline. We find that this problem can be attributed to a large extent to an inefficient utilization of TLBs in pointer-chasing instructions (i.e. load-to-load), which constitute a significant portion of workloads in IMA/GP. Conventionally, VATs for pointer-chasing are performed via a series of TLB look-ups: one for each load in the pointer chain. However, this approach is sub-optimal as it overlooks the correlations between pointers, missing opportunities to perform VATs through more costeffective calculations instead of relying on the more expensive TLB look-ups. This work draws on the insight that in pointer-chasing workloads, the physical page holding an upstream pointer is often at a fixed distance to the physical page where a downstream pointer resides. Building on this insight, we introduce RefleXive address Translation (RXT), which encodes the physical page distance (termed PageDist) between an upstream and downstream pointer into the unused upper 16 bits of the upstream pointer's virtual address. Consequently, RXT can compute the translation of a pointer by directly adding the PageDist to its residing address (the physical address where the pointer is stored). This process can be recursively applied throughout the pointer-chasing sequence, transforming address translation from a sequence of TLB accesses to computes. Through gem5 full-system simulation, RXT reduces TLB energy consumption for pointer-chasing applications by an average of 41.48 % (up to 62.50 %) and decreases overall core power density by 4.65 % (up to 13.09 %), without compromising application execution time. In fact, RXT even improves program runtime by an average of 0.71 % (up to 2.09 %).
Rashid Aligholipour, Pavlos Aimoniotis, Stefanos Kaxiras, Yuan Yao 0009
IPDPS4
2025 The Fake-Busy and True-Idle Problems of Running Graph Applications on Chiplet-Based Multi-Cores
abstract
We introduce the fake-busy and true-idle problems encountered when running large graph workloads on chipletbased Out-of-Order (OoO) multi-cores. Caused by high interchiplet communication latency and irregular memory access patterns, these issues lead to inefficient use of core pipeline resources such as the Reorder Buffer (RoB) and Load Queue (LQ). Our evaluation shows that reducing RoB and LQ sizes has minimal impact on application performance, revealing new performance optimization opportunities for large graph workloads.
Rashid Aligholipour, Yuan Yao 0009
ISPASS2
2024 TangramFP: Energy-Efficient, Bit-Parallel, Multiply-Accumulate for Deep Neural Networks
abstract
As energy consumption becomes a primary concern for deep learning acceleration, the need to optimize not only data movement but also compute is becoming important. The basic element of compute, the Multiply-Accumulate (MAC) unit, performs the operation X · Y+Z, comprises the compute cores of systolic arrays such as Google’s TPU or Nvidia’s Tensor Cores, and it is found in practically every deep neural network (DNN) accelerator.In this work, we aim to reduce the energy needs of bit-parallel MACs, without perceptible impact on precision, and without affecting the structure of the overall accelerator architecture—in other words, we aim for an energy-efficient drop-in MAC replacement.Although there is a significant body of work on efficient approximate multipliers and MACs, in this work, we propose a novel approach: a tunable floating-point MAC design, TANGRAMFP, that can deliver the full precision of a standard implementation, yet dynamically adjusts to eliminate ineffectual computation. Different from state-of-the-art approaches that are based on truncated multiplication, TANGRAMFP introduces a new class of multipliers where input operands are split and partial products are selectively generated (by enabling or disabling different areas of the logical multiplier array) and added together. In a hardware implementation, this is achieved by decomposing a large multiplier into four smaller ones, at the same overall hardware cost.We demonstrate that TANGRAMFP precision can adhere to the same bounds (measured as Unit-in-Last-Place—ULP—error) as standard IEEE FP16 arithmetic and delivers better precision than a state-of-the-art approach based on bit-serial truncated multiplication aimed at eliminating ineffectual computation in DNNs. At the same time, TANGRAMFP is a drop-in replacement for the standard MAC design, having approximately the same mean ULP error, the same area and latency, while achieving up to 36.57% dynamic power savings (27.44% with a mean error close to standard).
Yuan Yao 0009, Hannah Atmer, Stefanos Kaxiras
SBAC-PAD1
2024 TaDA: Task Decoupling Architecture for the Battery-less Internet of Things
abstract
We present TaDA, a system architecture enabling efficient execution of Internet of Things (IoT) applications across multiple computing units, powered by ambient energy harvesting. Low-power microcontroller units (MCUs) are increasingly specialized; for example, custom designs feature hardware acceleration of neural network inference, next to designs providing energy-efficient input/output. As application requirements are growingly diverse, we argue that no single MCU can efficiently fulfill them. TaDA allows programmers to assign the execution of different slices of the application logic to the most efficient MCU for the job. We achieve this by decoupling task executions in time and space, using a special-purpose hardware interconnect we design, while providing persistent storage to cross periods of energy unavailability. We compare our prototype performance against the single most efficient computing unit for a given workload. We show that our prototype saves up to 96.7% energy per application round. Given the same energy budget, this yields up to a 68.7x throughput improvement.
Weining Song, Stefanos Kaxiras, Thiemo Voigt, Yuan Yao 0009, Luca Mottola
SenSys4
2023 Silent Stores in the Battery-less Internet of Things: A Good Idea?
Weining Song, Stefanos Kaxiras, Luca Mottola, Thiemo Voigt, Yuan Yao 0009
EWSN5
2021 TSOPER: Efficient Coherence-Based Strict Persistency
abstract
We propose a novel approach for hardware-based strict TSO persistency, called TSOPER. We allow a TSO persistency model to freely coalesce values in the caches, by forming atomic groups of cachelines to be persisted. A group persist is initiated for an atomic group if any of its newly written values are exposed to the outside world. A key difference with prior work is that our architecture is based on the concept of a TSO persist buffer, that sits in parallel to the shared LLC, and persists atomic groups directly from private caches to NVM, bypassing the coherence serialization of the LLC. To impose dependencies among atomic groups that are persisted from the private caches to the TSO persist buffer, we introduce a sharing-list coherence protocol that naturally captures the order of coherence operations in its sharing lists, and thus can reconstruct the dependencies among different atomic groups entirely at the private cache level without involving the shared LLC. The combination of the sharing-list coherence and the TSO persist buffer allows persist operations and writes to non-volatile memory to happen in the background and trail the coherence operations. Coherence runs ahead at full speed; persistency follows belatedly. Our evaluation shows that TSOPER provides the same level of reordering as a program-driven relaxed model, hence, approximately the same level of performance, albeit without needing the programmer or compiler to be concerned about false sharing, data-race-free semantics, etc., and guaranteeing all software that can run on top of TSO, automatically persists in TSO.
Per Ekemark, Yuan Yao 0009, Alberto Ros 0001, Konstantinos Sagonas, Stefanos Kaxiras
HPCA2
2020 Pursuing Extreme Power Efficiency With PPCC Guided NoC DVFS
abstract
In sharp contrast to conventional performance indicative based Network-on-Chip (NoC) DVFS, where the direct relation between application performance and NoC power consumption is missing, we exploit the concept of Performance-Power Characteristic Curve (PPCC) newly proposed in the literature to approach maximum NoC power efficiency. PPCC, which defines the direct relation between application performance and NoC power consumption, consists of three distinct regions: an inertial region due to power under-provisioning, a linear region for proportional performance gain, and a saturation region due to power over-provisioning. With PPCC as a guidance, we propose Δ-DVFS, which employs a “profile-then-select” strategy to step-by-step approach maximum NoC power efficiency. Δ-DVFS is built on two observations. First, in multi-threaded applications, maximum NoC power efficiency is achieved at the boundary between the linear region and the saturation region on the PPCC. Second, PPCC stabilizes when threads repeat workloads of the same loop. This is intuitively meaningful because loop repetition stresses NoC with similar workload. Based on the observations, Δ-DVFS uses the first several loop iterations for PPCC profiling. After the profiling is done, Δ-DVFS selects and applies the optimal V/F that achieves maximum NoC power efficiency to the remaining loop iterations. To accurately and timely follow PPCC when threads proceed to different loops, Δ-DVFS utilizes an H-tree loop monitor to detect loop change among distributive threads.
Yuan Yao 0009, Zhonghai Lu
IEEE Trans. Computers1
2018 iNPG: Accelerating Critical Section Access with In-network Packet Generation for NoC Based Many-Cores
abstract
As recently studied, serialized competition overhead for entering critical section is more dominant than critical section execution itself in limiting performance of multi-threaded shared variable applications on NoC-based many-cores. We illustrate that the invalidation-acknowledgement delay for cache coherency between the home node storing the critical section lock and the cores running competing threads is the leading factor to high competition overhead in lock spinning, which is realized in various spin-lock primitives (such as the ticket lock, ABQL, MCS lock, etc.) and the spinning phase of queue spin-lock (QSL) in advanced operating systems. To reduce such high lock coherence overhead, we propose in-network packet generation (iNPG) to turn passive "normal" NoC routers which only transmit packets into active "big" ones that can generate packets. Instead of performing all coherence maintenance at the home node, big routers which are deployed nearer to competing threads can generate packets to perform early invalidation-acknowledgement for failing threads before their requests reach the home node, shortening the protocol round-trip delay and thus significantly reducing competition overhead in various locking primitives. We evaluate iNPG in Gem5 using PARSEC and SPEC OMP2012 programs with five different locking primitives. Compared to a state-of-the-art technique accelerating critical section access, experimental results show that iNPG can effectively reduce lock coherence overhead, expediting critical section access by 1.35x on average and 2.03x at maximum and consequently improving the program Region-of-Interest (ROI) runtime by 7.8% on average and 14.7% at maximum.
Yuan Yao 0009, Zhonghai Lu
HPCA1
2018 Thread Voting DVFS for Manycore NoCs
abstract
We present a thread-voting DVFS technique for manycore networks-on-chip (NoCs). This technique has two remarkable features which differentiate from conventional NoC DVFS schemes. (1) Not only network-level but also thread-level runtime performance indicatives are used to guide DVFS decisions. (2) To resolve multiple perhaps conflicting performance indicatives from many cores, it allows each thread to “vote” for a V/F level in its own performance interest, and a region-based V/F controller makes dynamic per-region V/F decision according to the major vote. We evaluate our technique on a 64-core CMP in full-system simulation environment GEM5 with both PARSEC and SPEC OMP2012 benchmarks. Compared to a network metric (router buffer occupancy) based approach, it can improve the network energy efficacy measured in MPPJ (million packets per joule) by up to 22 percent for PARSEC and 20 percent for SPEC OMP2012, and the system energy efficacy measured in MIPJ (million instructions per joule) by up to 35 percent for PARSEC and 33 percent for SPEC OMP2012.
Zhonghai Lu, Yuan Yao 0009
IEEE Trans. Computers2
2017 Prediction based convolution neural network acceleration: work-in-progress
abstract
Although intra-layer parallelism is commonly used to expedite CNN execution, it is difficult to achieve inter-layer parallelism because of data dependence between layers. In the paper, we propose a two-phase prediction and correction mechanism to break the data dependence between CNN layers so as to enable inter-layer parallelism. Our technique achieves one more order of magnitude (from the order of 10 to the order of 100) CNN acceleration compared to other three state-of-the-art GPU based CNN acceleration mechanisms.
Yuan Yao 0009, Zhonghai Lu
CASES1
2017 Marginal Performance: Formalizing and Quantifying Power Over/Under Provisioning in NoC DVFS
abstract
In network-on-chip (NoC) based CMPs, DVFS is commonly used to co-optimize performance and power. To achieve optimal efficiency, it is important to gain proportional performance growth with power. However, power over/under provisioning often exists. To properly evaluate and guide NoC DVFS techniques, it is highly desirable to formalize and quantify power over/under provisioning. In this paper, we first show that application performance does not grow linearly with network power in an NoC-based CMP. Instead, their relationship is non-linear and can be captured using performance-power characteristics curve (PPCC) with three distinct regions: an inertial region, a linear region, and a saturation region. We note that conventional DVFS metrics such as Performance Per Watt (PPW) cannot accurately evaluate such non-linear relationship. Based on PPCC, we propose a new figure of merit called Marginal Performance (MP) which evaluates the incremental performance per power increment after the inertial region. The MP concept enables to formally define power overand under-provisioning with reference to the linear region in which an efficient NoC DVFS should operate. Applying the PPCC and MP concepts in full-system simulations with PARSEC and SPEC OMP2012 benchmarks, we are able to identify power over/under provisioning occurrences, measure and compare their statistics in two latest NoC DVFS techniques. Moreover, we show evidences that MP can accurately and consistently evaluate the NoC DVFS techniques, avoiding the misjudgement and inconsistency of PPW-based evaluations.
Zhonghai Lu, Yuan Yao 0009
IEEE Trans. Computers2
2017 Dynamic Traffic Regulation in NoC-Based Systems
abstract
In network-on-chip (NoC)-based systems, performance enhancement has primarily focused on the network itself, with little attention paid on controlling traffic injection at the network boundary. This is unsatisfactory because traffic may be over injected, aggravating congestion, and lowering performance. Recently, traffic regulation is proposed as an orthogonal means for performance improvement. Rather than as soon as possible admission, traffic regulation may hold back packet injection by admitting packets into the network only when the accumulated traffic volume at any time interval does not exceed a threshold. These regulation techniques are, however, often static, likely causing overregulation and underregulation. We propose dynamic traffic regulation to improve the system performance for NoC-based multi/many-processor systems-on-chip (MPSoC) and chip multi/many-core processor (CMP) designs. It can be applied to MPSoCs for intellectual property integration in an open-loop fashion by injecting traffic according to its run-time profiled characteristics. It can also be applied to CMPs in a closed-loop fashion by admitting traffic fully adaptive to the traffic and network states. Through extensive experiments and results, we show that both the open-loop and closed-loop dynamic regulation techniques can significantly improve the network and system performance.
Zhonghai Lu, Yuan Yao 0009
IEEE Trans. Very Large Scale Integr. Syst.2
2016 Memory-access aware DVFS for network-on-chip in CMPs
Yuan Yao 0009, Zhonghai Lu
DATE1
2016 DVFS for NoCs in CMPs: A thread voting approach
abstract
As the core count grows rapidly, dynamic voltage/frequency scaling (DVFS) in networks-on-chip (NoCs) becomes critical in optimizing energy efficacy in chip multiprocessors (CMPs). Previously proposed techniques often exploit inherent network-level metrics to do so. However, such network metrics may contradictorily reflect application's performance need, leading to power over/under provisioning. We propose a novel on-chip DVFS technique for NoCs that is able to adjust per-region V/F level according to voted V/F levels of communicating threads. Each region is composed of a few adjacent routers sharing the same V/F level. With a voting-based approach, threads seek to influence the DVFS decisions independently by voting for a preferred V/F level that best suits their own performance interest according to their runtime profiled message generation rate and data sharing characteristics. The vote expressed in a few bits is then carried in the packet header and spread to the routers on the packet route. The final DVFS decision is made democratically by a region DVFS controller based on the majority election result of collected votes from all active threads. To achieve scalable V/F adjustment, each region works independently, and the voting-based V/F tuning forms a distributed decision making process. We evaluate our technique with detailed simulations of a 64-core CMP running a variety of multi-threaded PARSEC benchmarks. Compared with a network without DVFS and a network metric (router buffer occupancy) based approach, experimental results show that our voting based DVFS mechanism improves the network energy efficacy measured in MPPJ (million packets per joule) by about 17.9% and 9.7% on average, respectively, and the system energy efficacy measured in MIPJ (million instructions per joule) by about 26.3% and 17.1% on average, respectively.
Yuan Yao 0009, Zhonghai Lu
HPCA1
2016 Opportunistic Competition Overhead Reduction for Expediting Critical Section in NoC Based CMPs
abstract
With the degree of parallelism increasing, performance of multi-threaded shared variable applications is not only limited by serialized critical section execution, but also by the serialized competition overhead for threads to get access to critical section. As the number of concurrent threads grows, such competition overhead may exceed the time spent in critical section itself, and become the dominating factor limiting the performance of parallel applications. In modern operating systems, queue spinlock, which comprises a low-overhead spinning phase and a high-overhead sleeping phase, is often used to lock critical sections. In the paper, we show that this advanced locking solution may create very high competition overhead for multithreaded applications executing in NoC-based CMPs. Then we propose a software-hardware cooperative mechanism that can opportunistically maximize the chance that a thread wins the critical section access in the low-overhead spinning phase, thereby reducing the competition overhead. At the OS primitives level, we monitor the remaining times of retry (RTR) in a thread's spinning phase, which reflects in how long the thread must enter into the high-overhead sleep mode. At the hardware level, we integrate the RTR information into the packets of locking requests, and let the NoC prioritize locking request packets according to the RTR information. The principle is that the smaller RTR a locking request packet carries, the higher priority it gets and thus quicker delivery. We evaluate our opportunistic competition overhead reduction technique with cycle-accurate full-system simulations in GEM5 using PARSEC (11 programs) and SPEC OMP2012 (14 programs) benchmarks. Compared to the original queue spinlock implementation, experimental results show that our method can effectively increase the opportunity of threads entering the critical section in low-overhead spinning phase, reducing the competition overhead averagely by 39.9% (maximally by 61.8%) and accelerating the execution of the Region-of-Interest averagely by 14.4% (maximally by 24.5%) across all 25 benchmark programs.
Yuan Yao 0009, Zhonghai Lu
ISCA1
2016 Aggregate Flow-Based Performance Fairness in CMPs
abstract
In CMPs, multiple co-executing applications create mutual interference when sharing the underlying network-on-chip architecture. Such interference causes different performance slowdowns to different applications. To mitigate the unfairness problem, we treat traffic initiated from the same thread as an aggregate flow such that causal request/reply packet sequences can be allocated to resources consistently and fairly according to online profiled traffic injection rates. Our solution comprises three coherent mechanisms from rate profiling, rate inheritance, and rate-proportional channel scheduling to facilitate and realize unbiased workload-adaptive resource allocation. Full-system evaluations in GEM5 demonstrate that, compared to classic packet-centric and latest application-prioritization approaches, our approach significantly improves weighted speed-up for all multi-application mixtures and achieves nearly ideal performance fairness.
Zhonghai Lu, Yuan Yao 0009
ACM Trans. Archit. Code Optim.2
2014 Fuzzy flow regulation for Network-on-Chip based chip multiprocessors systems
abstract
Flow regulation is a traffic shaping technique, which can be used to improve communication performance with better utilization of network resources in chip multi-processors (CMPs). This paper presents fuzzy flow regulation. Being different from the static flow regulation policy, our system makes regulation decisions fully dynamically according to traffic dynamism and the state of interconnection network. The central idea is to use fuzzy logic to mimic the behavior of an expert that can recognize the network status and then intelligently control the admission of input flows. As the experiment results show, the maximum improvement in average delay reaches 53.0% against static regulation and 37.4% against no regulation. The maximum improvement in average throughput reaches 37.5% against static regulation and 23.8% against no regulation.
Yuan Yao 0009, Zhonghai Lu
ASP-DAC1
2014 Towards stochastic delay bound analysis for Network-on-Chip
abstract
We propose stochastic performance analysis in order to provide probabilistic quality-of-service guarantees in on-chip packet-switching networks. In contrast to deterministic analysis which gives per-flow absolute delay bound, stochastic analysis derives per-flow probabilistic delay bounding function, which can be used to avoid over-dimensioning network resources. Based on stochastic network calculus, we build a basic analytic model for an on-chip router, propose and exemplify a stochastic performance analysis flow. In experiments, we show the correctness and accuracy of our analysis, and exhibit its potential in enhancing network utilization with a relaxed delay requirement. Moreover, the benefits of such relaxation is demonstrated through a video playback application.
Zhonghai Lu, Yuan Yao 0009, Yuming Jiang 0001
NOCS2