VLDB 2026 Research / reviewers in the wild / expert
Hans Kasan
dblp:213/4782
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-8828-8058ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine Learning
Hans Kasan, Dennis Abts, Jungwook Choi, John Kim 0001 |
MICRO | 1 |
| 2024 | Ghost Arbitration: Mitigating Interconnect Side-Channel Timing Attacks in GPUabstractNetwork-on-chip (NoC) is a critical shared resource in scalable multicore processors; however, it is well-known that shared resources can lead to side-channel attacks. In this work, we demonstrate how contention for on-chip bandwidth in GPUs can lead to fine-grain information leakage and enable side-channel attacks. As a case study, we demonstrate how RSA key bit information can be leaked on a real GPU. We also describe how interconnect characteristics from the side-channel or an interconnect-gram can be used to fingerprint kernels executing on the GPU. To defend against such fine-grain side-channel attack, we propose secure arbitration that prevents information leakage while minimizing performance impact during normal execution. In particular, we present a novel ghost arbitration that prevents interconnect contention from being leveraged to leak information by keeping track of “ghost” requests or requests when other nodes receive free arbitration to enable least-recently-used priority. However, if the attacker reverse engineers the arbitration, a naive implementation of ghost arbitration can still lead to information leakage. Thus, we propose a weighted ghost arbitration that exploits “malicious” communication patterns to prevent information leakage with minimal loss in performance. Compared to previously proposed arbitration that is secure (e.g., strict time-division multiplexing), ghost arbitration is able to improve performance by up to$4\times$• Zhixian Jin, Jaeguk Ahn, Hans Kasan, Jina Song, Wonjun Song, John Kim 0001 |
MICRO | 4 |
| 2024 | Uncovering Real GPU NoC Characteristics: Implications on Interconnect ArchitectureabstractA critical component of high-throughput processors such as GPUs is the network-on-chip (NoC) that interconnects the large number of cores and the memory partitions together. In this work, we provide a detailed analysis, in terms of latency and bandwidth, of real GPU NoC across several generations of modern NVIDIA GPUs. Our analysis identifies how non-uniform latency exists between the cores and the memory partitions based on their physical location in the GPU. The non-uniformity can result in up to approximately 70 % difference in on-chip latency. In comparison, the bandwidth provided from the cores to the memory partitions is approximately uniform. However, recent GPUs that consist of multiple GPU “partitions” present different on-chip latency and bandwidth characteristics when communicating between the partitions. Based on our analysis of real GPU interconnect, we discuss potential implications including its impact on timing used in side-channel attacks as well as NoC microarchitectures. We show how the non-uniform latency can be exploited in a timing side-channel attack within a GPU as the core location impacts performance (or timing). In addition, proper understanding (and proper assumptions) of GPU NoC is critical to ensure a network that does not bottleneck the overall system performance. Zhixian Jin, Christopher Rocca, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M. Aamodt, John Kim 0001 |
MICRO | 4 |
| 2023 | VVQ: Virtualizing Virtual Channel for Cost-Efficient Protocol Deadlock AvoidanceabstractDeadlock freedom is a critical component of interconnection networks in large-scale systems. In particular, protocol or high-level deadlock can occur from dependency based on network endpoints. Virtual channels (VCs) are commonly used to avoid such protocol deadlocks in large-scale systems; however, the cost of VCs is high in large-scale networks because of the deep input buffers and VCs need to be replicated. In this work, we propose to virtualize virtual channels to create Virtualizing Virtual Queues (VVQ) buffer architecture. VVQ is based on the observation that a FIFO buffer organization is sufficient when protocol deadlocks do not occur. However, when potential protocol deadlock can occur through blocking, VVQ takes a proactive approach to allow packets from different traffic classes to "jump" the queue and ensure protocol deadlock does not occur. Our proposal shows how a ghost pointer can be leveraged to enable such buffer organization without introducing complex buffer management such as dynamic buffer organizations. Our evaluations show VVQ can match the performance of the baseline VC buffer organization with only half of the total buffer storage. Hans Kasan, John Kim 0001 |
HPCA | 1 |
| 2022 | Dynamic global adaptive routing in high-radix networksabstractGlobal adaptive routing is a critical component of high-radix networks in large-scale systems and is necessary to fully exploit the path diversity of a high-radix topology. The routing decision in global adaptive routing is made between minimal and non-minimal paths, often based on local information (e.g., queue occupancy) and rely on "approximate" congestion information through backpressure. Different heuristic-based adaptive routing algorithms have been proposed for high-radix topologies; however, heuristic-based routing has performance trade-off for different traffic patterns and leads to inefficient routing decisions. In addition, previously proposed global adaptive routing algorithms are static as the same routing decision algorithm is used, even if the congestion information changes. In this work, we propose a novel global adaptive routing that we refer to as dynamic global adaptive routing that adjusts the routing decision algorithm through a dynamic bias based on the network traffic and congestion to maximize performance. In particular, we propose DGB - Decoupled, Gradient descent-based Bias global adaptive routing algorithm. DGB introduces a dynamic bias to the global adaptive routing decision by leveraging gradient descent to dynamically adjust the adaptive routing bias based on the network congestion. In addition, both the local and global congestion information are decoupled in the routing decision - global information is used for the dynamic bias while local information is used in the routing decision to more accurately estimate the network congestion. Our evaluations show that DGB consistently outperforms previously proposed routing algorithms across diverse range of traffic patterns and workloads. For asymmetric traffic pattern, DGB improves throughput by 65% compared to the state-of-the-art global adaptive routing algorithm while matching the performance for symmetric traffic patterns. For trace workloads, DGB provides average performance improvement of 26%. Hans Kasan, Gwangsun Kim, Yung Yi, John Kim 0001 |
ISCA | 1 |
| 2021 | BoomGate: Deadlock Avoidance in Non-Minimal Routing for High-Radix NetworksabstractAvoiding routing deadlock is an important component of an interconnection network. For large-scale systems with high-radix topologies that leverage non-minimal adaptive routing, virtual channels (VCs) are commonly used to prevent routing deadlock. However, VCs in large-scale networks can be costly because of deep buffers and restrict VC usage. In this work, we propose BoomGATE for deadlock avoidance in large-scale networks. In particular, BoomGATE consists of two components - Restricted Intermediate-node Non-minimal Routing (RINR) algorithm and opportunistic flow control (OFC) which both exploit the low-diameter characteristics of high-radix networks while maximizing path diversity within the topology. We identify how routing deadlock in fully-connected topologies are caused by non-minimal routes and propose to restrict the non-minimal routing to ensure deadlock freedom without additional virtual channels. We also propose an algorithm that ensures path diversity is load-balanced across all nodes in the system. However, since path diversity is restricted with the RINR algorithm, complement RINR algorithm with opportunistic flow control (OFC) where “illegal routes” are allowed if and only if sufficient buffer can be guaranteed to ensure cyclical dependency does not occur. We propose both a static and dynamic OFC implementation. We evaluate the performance of BoomGATE and demonstrate there is minimal performance loss compared to global adaptive routing, while reducing the amount of buffers required by 50%. Gyuyoung Kwauk, Seungkwan Kang, Hans Kasan, Hyojun Son, John Kim 0001 |
HPCA | 3 |
| 2021 | Network-on-Chip Microarchitecture-based Covert Channel in GPUsabstractAs GPUs are becoming widely deployed in the cloud infrastructure to support different application domains, the security concerns of GPUs are becoming increasingly important. In particular, the support for multiprogramming in modern GPUs has led to new vulnerabilities since multiple kernels in a GPU can be executed at the same time. In this work, we propose a new microarchitectural timing covert channel for GPUs that can be established based on the shared, on-chip interconnect channels. We first reverse-engineer the organization of the on-chip networks in modern GPUs to understand the core placements throughout the GPU. The hierarchical organization of the GPU results in the sharing of interconnect bandwidth between neighboring cores. Based on this understanding, we identify how contention for the interconnect bandwidth can be exploited for a novel covert channel attack. We propose two types of interconnect-based covert channels that exploit the on-chip network hierarchy. Unlike cache-based covert channels, no states of the on-chip network need to be modified for communication in our interconnect-based covert channel and the impact of contention is very predictable. By exploiting the parallelism of GPUs, our proposed covert channel results in very high bandwidth – achieving approximately 24 Mbps of bandwidth on NVIDIA Volta GPUs and results in one of the highest known microarchitectural covert channel bandwidth. Jaeguk Ahn, Hans Kasan, Zhixian Jin, Leila Delshadtehrani, Wonjun Song, Ajay Joshi, John Kim 0001 |
MICRO | 3 |