VLDB 2026 Research / reviewers in the wild / expert
Hanjoon Kim
dblp:00/9135
· DBLP profile ↗
9ranked-venue papers
6as first author
1since 2021 · last 2024
0000-0003-4510-5685ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Interconnection networks and networks-on-chip · 40% Processor architecture and microarchitecture · 28% Hardware accelerators and domain-specific architectures · 23% |
Topics — the 8 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 1 | 2024 | TCP: A Tensor Contraction Processor for AI Workloads Industrial Product · ISCA 2024 |
Interconnection networks and networks-on-chip
flow control |
0.4 | 2 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 Transportation-network-inspired network-on-chip · HPCA 2014 |
Interconnection networks and networks-on-chip › ring network
hierarchical ring |
0.4 | 2 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 Transportation-network-inspired network-on-chip · HPCA 2014 |
Interconnection networks and networks-on-chip › network topology › network topology design
network-on-chip topology |
0.2 | 1 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 |
Memory systems › cache › prefetching
hardware prefetching |
0.2 | 1 | 2014 | Mutually Aware Prefetcher and On-Chip Network Designs for Multi-Cores · IEEE Trans. Computers 2014 |
Processor architecture and microarchitecture
chip multiprocessor |
0.1 | 2 | 2014 | Mutually Aware Prefetcher and On-Chip Network Designs for Multi-Cores · IEEE Trans. Computers 2014 Transportation-network-inspired network-on-chip · HPCA 2014 |
Interconnection networks and networks-on-chip
congestion control |
0.1 | 1 | 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-Chip · IEEE Trans. Computers 2016 |
Processor architecture and microarchitecture
many-core architecture |
0.1 | 1 | 2014 | Transportation-network-inspired network-on-chip · HPCA 2014 |
Methods — techniques the papers use, named apart from their topics
design space exploration · 0.8data reuse optimization · 0.8circuit-switched fetch network · 0.8virtual channels · 0.4simulation-based evaluation · 0.2credit network design · 0.2priority-based router · 0.2congestion-sensitive throttling · 0.2congestion management · 0.2closed-loop simulation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TCP: A Tensor Contraction Processor for AI Workloads Industrial ProductabstractWe introduce a novel tensor contraction processor (TCP) architecture that offers a paradigm shift from traditional architectures that rely on fixed-size matrix multiplications. TCP aims at exploiting the rich parallelism and data locality inherent in tensor contractions, thereby enhancing both efficiency and performance of AI workloads.TCP is composed of coarse-grained processing elements (PEs) to simplify software development. In order to efficiently process operations with diverse tensor shapes, the PEs are designed to be flexible enough to be utilized as a large-scale single unit or a set of small independent compute units.We aim at maximizing data reuse on both levels of inter and intra compute units. To do that, we propose a circuit switch-based fetch network to flexibly connect compute units to enable inter-compute unit data reuse. We also exploit input broadcast to multiple contraction engines and input buffer based reuse to further exploit reuse behavior in tensor contraction. Our compiler explores the design space of tensor contractions considering tensor shapes and the order of their associated loop operations as well as the underlying accelerator architecture.A TCP chip was designed and fabricated in 5nm technology as the second-generation product of Furiosa AI, offering 256/512/1024 TOPS (BF16/FP8 or INT8/INT4) with 256 MB SRAM and 1.5 TB/s 48 GB HBM3 under 150 W TDP. Commercialization will start in August 2024.We performed an extensive case study of running the LLaMA-2 7B model and evaluated its performance and power efficiency on various configurations of sequence length and batch size. For this model, TCP is 2.7 × and 4.1 × better than H100 and L40s, respectively, in terms of performance per watt. Hanjoon Kim, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee 0014, Jeseung Yeon, Changjae Park, Boncheol Gu, Changman Lee, Jaeick Bae, SungGyeong Bae, Yojung Cha, Wooyoung Choe, Jonguk Choi, Juho Ha, Hyuck Han, Namoh Hwang, Seokha Hwang, Kiseok Jang, Haechan Je, Hojin Jeon, Jaewoo Jeon, Hyunjun Jeong, Yeonsu Jung, Dongok Kang, Hyewon Kim, Muhwan Kim, Sewon Kim, Suhyung Kim, Yong Kim, Youngsik Kim, Younki Ku, Jeong Ki Lee, Juyun Lee, Seokho Lee, Minwoo Noh, Hyuntaek Oh, Gyunghee Park, Jimin Seo, Jungyoung Seong, June Paik, Nuno P. Lopes, Sungjoo Yoo |
ISCA | 1 |
| 2019 | Ghost routers: energy-efficient asymmetric multicore processors with symmetric NoCsabstractAsymmetric multicore architectures have been proposed to exploit the benefits of heterogeneous cores. However, asymmetric cores present challenge to network-on-chip (NoC) designers since the floorplan is not necessarily regular with "nodes" being different size. In contrast, most of the previously proposed NoC topologies commonly assume a regular or symmetric floorplan with equal size nodes. In this work, we first describe how asymmetric floorplan leads to asymmetric topology and can limit overall performance. To overcome the asymmetric floorplan, we present Ghost Routers - extra "dummy" routers that are added to the NoC to create a symmetric NoC architecture for asymmetric multicore architectures. Ghost router provides higher network path diversity and provides higher network performance that leads to higher system performance. Ghost routers also enable simpler routing algorithms because of the symmetric NoC architecture. While ghost routers is a simplistic modification to the NoC architecture, it does increase NoC cost. However, ghost routers exploit the observations that in realistic systems, the cost of NoC is not a significant fraction of overall system cost. Our evaluations show that ghost routers can improve performance by up to 21% while improving overall energy-efficiency of the system by up to 26%. Hyojun Son, Hanjoon Kim, Hao Wang 0011, Nam Sung Kim, John Kim 0001 |
NOCS | 2 |
| 2016 | Design and Analysis of Hybrid Flow Control for Hierarchical Ring Network-on-ChipabstractA cost-efficient network-on-chip is needed in a scalable many-core systems. Recent multicore processors have leveraged a ring topology and hierarchical ring can increase scalability but presents different challenges, including higher hop count and global ring bottleneck. In this work, we describe a hierarchical ring topology that we refer to as a transportation-network-inspired network-on-chip (tNoC) that leverages principles from transportation network systems. In particular, we propose a novel hybridflow control for hierarchical ring topology to scale the topology efficiently. The flow control is hybrid in that the channels are allocated on flit granularity while the buffers are allocated on packet granularity. The hybrid flow control enables a simplified router microarchitecture (to minimize per-hop latency) as router input buffers are minimized and buffers are pushed to the edges, either at the output ports or at the hub routers that interconnect the local rings to the global ring-while still supporting virtual channels to avoid protocol deadlock. We describe a packet-quota-system (PQS) and a separate credit network that provide congestion management, support prioritized arbitration in the network, and provide support for multiflit packets. We also provide alternative designs for the credit network and PQS architectures. A detailed evaluation of a 64-core CMP shows that the tNoC improves performance by up to 21 percent compared with a baseline, buffered hierarchical ring topology while reducing NoC energy by 51 percent. Hanjoon Kim, Gwangsun Kim, Hwasoo Yeo, John Kim 0001, Seung Ryoul Maeng |
IEEE Trans. Computers | 1 |
| 2014 | Transportation-network-inspired network-on-chipabstractA cost-efficient network-on-chip is needed in a scalable many-core systems. Recent multicore processors have leveraged a ring topology and hierarchical ring can increase scalability but presents different challenges, including higher hop count and global ring bottleneck. In this work, we describe a hierarchical ring topology that we refer to as a transportation-network-inspired network-on-chip (tNoC) that leverages principles from transportation network systems. In particular, we propose a novel hybrid flow control for hierarchical ring topology to scale the topology efficiently. The flow control is hybrid in that the channels are allocated on flit granularity while the buffers are allocated on packet granularity. The hybrid flow control enables a simplified router microarchitecture (to minimize per-hop latency) as router input buffers are minimized and buffers are pushed to the edges, either at the output ports or at the hub routers that interconnect the local rings to the global ring - while still supporting virtual channels to avoid protocol deadlock. We also describe a packet-quota-system (PQS) and a separate credit network that provide congestion management, support prioritized arbitration in the network, and provide support for multiflit packets. A detailed evaluation of a 64-core CMP shows that the tNoC improves performance by up to 21% compared with a baseline, buffered hierarchical ring topology while reducing NoC energy by 51%. Hanjoon Kim, Gwangsun Kim, Seung Ryoul Maeng, Hwasoo Yeo, John Kim 0001 |
HPCA | 1 |
| 2014 | Extending bufferless on-chip networks to high-throughput workloadsabstractBufferless networks-on-chip (NoC) has been proposed to reduce network cost by removing router input buffers and improve energy-efficiency. However, bufferless NoC has some limitations that include lower network throughput caused by deflection routing at high load. In addition, the longer router critical path impacts the router frequency, which reduces the amount of bandwidth provided by the network router. These limitations reduce any benefit of bufferless NoC - especially for high-throughput workloads such as GPGPU. In this work, we first provide a simple analysis into how the benefit of bufferless NoC can be reduced for high throughput workloads, especially in terms of energy-efficiency. We then propose clumsy flow control (CFC) - a congestion control mechanism that can reduce the amount of deflection and improve the efficiency of bufferless NoC. The clumsy flow control enables the allocation to be simplified and we propose a novel switch allocation (randomized-deterministic allocation) which significantly reduces the router critical path. The combination of these two techniques result in our bufferless NoC to exceed the system performance of buffered network by approximately 7% (up to 22%) while reducing network area by 53% and energy by 52%. Hanjoon Kim, Miri Kim, Kanghee Won, John Kim 0001 |
NOCS | 1 |
| 2014 | Mutually Aware Prefetcher and On-Chip Network Designs for Multi-CoresabstractHardware prefetching has become an essential technique in high performance processors to hide long external memory latencies. In multi-core architectures with cores communicating through a shared on-chip network, traffic generated by the prefetchers can account for up to 60% of the total on-chip network traffic. However, the distinct characteristics of prefetch traffic have not been considered in on-chip network design. In addition, prefetchers have been oblivious to the network congestion. In this work, we investigate the interactions between prefetchers and on-chip networks, exploiting the synergy of these two components in multi-cores. Firstly, we explore the design space of prefetch-aware on-chip networks. Considering the difference between prefetch and non-prefetch packets, we propose a priority-based router design, which selects non-prefetch packets first over prefetch packets. Secondly, we investigate network-aware prefetcher designs. We propose a prefetch control mechanism sensitive to network congestion—throttling prefetch requests based on the current network congestion. Our evaluation with full system simulations shows that the combination of the proposed prefetch-aware router and congestion-sensitive prefetch control improves the performance of benchmark applications by 11–12% with out-of-order cores, and 21–22% with SMT cores on average, up to 37% on some of the workloads. Hanjoon Kim, Minjeong Shin, John Kim 0001, Jaehyuk Huh 0001 |
IEEE Trans. Computers | 2 |
| 2012 | Providing cost-effective on-chip network bandwidth in GPGPUsabstractNetwork-on-chip (NoC) bandwidth has a significant impact on overall performance in throughput-oriented processors such as GPG-PUs. Although it has been commonly assumed that high NoC bandwidth can be provided through abundant on-chip wires, we show that increasing NoC router frequency results in a more cost-effective NoC. However, router arbitration critical path can limit the NoC router frequency. Thus, we propose a direct all-to-all network overlaid on mesh (DA2mesh) NoC architecture that exploits the traffic characteristics of GPGPU and removes arbitration from the router pipeline. DA2mesh simplifies the router pipeline with 36% improvement of performance while reducing NoC energy by 15%. Hanjoon Kim, John Kim 0001, Woong Seo, Yeongon Cho, Soojung Ryu |
ICCD | 1 |
| 2011 | Exploiting Mutual Awareness between Prefetchers and On-chip Networks in Multi-coresabstractThe unique characteristics of prefetch traffic have not been considered in on-chip network design for multicore architectures. Most prefetchers are often oblivious to the network congestion when generating prefetech requests. In this work, we investigate the interaction between prefetchers and on-chip networks and exploit the synergy of these two components in multi-core architectures. We explore prefetchaware on-chip networks that differentiates between prefetch and demand traffic by prioritizing demand traffic. In addition, we propose prefetch control mechanism based on network congestion. Our evaluations show that the combination of the proposed prefetch-aware router architecture and congestion sensitive prefetch control improves the performance of benchmarks by 11-13% on average, up to 30% on some of the workloads. Minjeong Shin, Hanjoon Kim, John Kim 0001, Jaehyuk Huh 0001 |
PACT | 3 |
| 2010 | On-Chip Network Evaluation FrameworkabstractWith the number of cores on a chip continuing to increase, proper evaluation of on-chip network is critical for not only network performance but also overall system performance. In this paper, we show how a network-only simulation can be limited as it does not provide an accurate representation of system performance. We evaluate traditionally used open loop simulations and compare them to closed-loop simulations. Although they use different methodologies, measurements, and metrics, we identify how they can provide very similar results. However, we show how the results of closed-loop simulations do not correlate well with execution-driven simulations. We then add simple extensions to the closed-loop simulation to model the impact of the processor and the memory system and show how the correlation with execution-driven simulations can be improved. The proposed framework/methodology provides a fast simulation time while providing better insights into the impact of network parameters on overall system performance. Hanjoon Kim, Seulki Heo, Jaehyuk Huh 0001, John Kim 0001 |
SC | 1 |