Henry Kao

dblp:249/4465 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0002-6897-4218ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Scalar Interpolation: A Better Balance between Vector and Scalar Execution for SuperScalar Architectures
abstract
Most compilers convert all iterations of a vectorizable loop into vector operations to decrease processing time. This paper proposes Scalar Interpolation, a technique that inserts scalar operations into vectorized loops to increase the utilization of execution units in processors with distinct pipelines for scalar and vector processing. Scalar interpolation inserts scalar operations for an entire iteration of the sequential loop to avoid data movements between vector and scalar registers. A challenge to introducing scalar interpolation is creating a static cost model to guide the compiler’s decision to interpolate scalar operations in a loop. An alternative to a static cost model is to perform auto-tuning in a loop to dynamically discover a sweet spot for the scalar interpolation factor. A performance study on an LLVM-based prototype reveals speedups of up to 30% on Intel Xeon (x86) with a static analysis of the cost model, and 43% on Kunpeng-920 (AArch64) with auto-tuning.
Reza Ghanbari, Henry Kao, João P. L. de Carvalho, Ehsan Amiri, José Nelson Amaral
CGO2
2025 A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
abstract
Modern mobile CPU software pose challenges for conventional instruction cache replacement policies due to their complex runtime behavior causing high reuse distance between executions of the same instruction.Mobile code commonly suffers from large amounts of stalls in the CPU frontend and thus starvation of the rest of the CPU resources.Complexity of these applications and their code footprint are projected to grow at a rate faster than available on-chip memory due to power and area constraints, making conventional hardware-centric methods for managing instruction caches to be inadequate.We present a novel software-hardware co-design approach called TRRIP (Temperature-based Re-Reference Interval Prediction) that enables the compiler to analyze, classify, and transform code based on "temperature" (hot/cold), and to provide the hardware with a summary of code temperature information through a well-defined OS interface based on using code page attributes.TRRIP's lightweight hardware extension employs code temperature attributes to optimize the instruction cache replacement policy resulting in the eviction rate reduction of hot code.TRRIP is designed to be practical and adoptable in real mobile systems that have strict feature requirements on both the software and hardware components.TRRIP can reduce the L2 MPKI for instructions by 26.5% resulting in geomean speedup of 3.9%, on top of RRIP cache replacement running mobile code already optimized using PGO.
Henry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati, Fang Su, Maziar Goudarzi
MICRO1
2022 ESUM: an efficient UTXO schedule model
abstract
The current size of UTXO set in Bitcoin is large and continues to grow, resulting in high memory usage. This paper proposes an efficient UTXO schedule model to lower memory usage based on the observation of UTXO lifespan. Full nodes with the UTXO schedule model store the full UTXO set in disk database and schedule a subset of full UTXO set in memory, converting UTXO querying and accessing into multi-layer activities. Machine learning methods are applied to train a well-performed model based on historical transaction data and market trade data. Experiment results indicate that the UTXO schedule model could reduce nearly 90% memory usage compared to original Bitcoin nodes, allowing more nodes with different memory capability to participate in the blockchain.
Meilin Lv, Kangjian Wei, Henry Kao, Huiping Sun, Zhong Chen 0001
ICBC3
2021 Pitstop: Enabling a Virtual Network Free Network-on-Chip
abstract
Maintaining correctness is of paramount importance in the design of a computer system. Within a multiprocessor interconnection network, correctness is guaranteed by having deadlock-free communication at both the protocol and network levels. Modern network-on-chip (NoC) designs use multiple virtual networks to maintain protocol-level deadlock freedom, at the expense of high power and area overheads. Other techniques involve complex detection and recovery mechanisms, or use misrouting which incurs additional packet latency. Considering that the probability of deadlocks occurring is low, the additional resources needed to avoid/resolve deadlocks should also be low. To this end, we propose Pitstop, a low-cost technique that guarantees correctness by resolving both protocol and network-level deadlocks without the use of virtual networks, complex hardware, or misrouting. Pitstop transfers blocked packets to the network interface (NI) creating a bubble (empty buffer slot) which breaks deadlock. The blocked packet can make forward progress through NI to NI traversals using low complexity bypassing mechanisms. This scheme performs better due to higher utilization of virtual channels and works on arbitrary irregular topologies without any virtual networks. Compared to state-of-the-art solutions, Pitstop can improve performance up to 11% and reduce power and area up to 41% and 40%.
Hossein Farrokhbakht, Henry Kao, Kamran Hasan, Paul Gratz, Tushar Krishna, Joshua San Miguel, Natalie D. Enright Jerger
HPCA2
2019 UBERNoC: unified buffer power-efficient router for network-on-chip
abstract
Networks-on-Chip (NoCs) address many shortcomings of traditional interconnects. However, they consume a considerable portion of a chip's total power - particularly when the utilization is low. As transistor size continues to shrink, we expect NoCs to contribute even more, especially static power. A wide range of prior-art focuses on reducing the contribution of NoC power consumption. These can be categorized into two main groups: (1) power-gating, and (2) simplified router microarchitectures. Maintaining the performance and the flexibility of the network are key challenges that have not yet been addressed by these two groups of low-power architectures. In this paper, we propose UBERNoC, a simplified router microarchitecture, which reduces underutilized buffer space by leveraging an observation that for most switch traversals, only a single packet is present. We use a unified buffer with multiple virtual channels shared amongst the input ports to reduce both power and area. The empirical results demonstrate that compared to a conventional router, UBERNoC achieves 58% and 69% reduction in power and area respectively, with negligible latency overhead.
Hossein Farrokhbakht, Henry Kao, Natalie D. Enright Jerger
NOCS2