EDBT 2026 Demo / reviewers in the wild / expert
Tae Hee Han
dblp:48/7551
· DBLP profile ↗
10ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0001-8508-7536ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ScaleMoE: A Fast and Scalable Distributed Training Framework for Large-Scale Mixture-of-Experts ModelsabstractThe size of pre-trained models has continuously increased to support growing demands for solving more complex problems. Especially, mixture-of-experts (MoE) model has become the most popular approach, enabling systems to easily train extremely large-scale models with relatively lower computational requirements. However, the current distributed training frameworks cannot achieve scalable performance for these large-scale MoE models due to substantial communication overheads. In this paper, we propose ScaleMoE, a fast and scalable distributed training framework for large-scale MoE models. We first identify three problems in state-of-the-art distributed training frameworks: high all-to-all communication overheads, severe load imbalance in expert selection, and insufficient consideration of heterogeneous networks. We propose three novel optimizations to resolve these problems. First, to reduce communication volumes, we propose adaptive all-to-all communication that eliminates unnecessary zeros caused by zero padding. Second, to address the load imbalance in expert selection, we propose dynamic expert clustering that rebalances experts using a novel clustering methodology. Lastly, to further minimize communication overheads, we propose topology-aware expert remapping that carefully maps experts to GPU devices while considering heterogeneous network bandwidths. Our evaluations show that ScaleMoE achieves scalable performance, reducing all-to-all communication overheads by up to $\mathbf{8 1 \%}$. In general, ScaleMoE significantly improves system performance, achieving a speedup of up to $3.3 \times$ compared to the state-of-the-art framework. Seohong Choi, Huize Hong, Tae Hee Han, Joonsung Kim 0001 |
PACT | 3 |
| 2025 | LNBN: layer-flexible non-blocking bypass network-on-chip for accelerating DNN inferenceabstractOn-device artificial intelligence has increased the importance of energy-efficient inference in resource-constrained environments. Lightweight deep neural networks (DNNs) reduce computational complexity by decreasing the data dimensionality of layers, leading to reduced data reuse, causing global buffer bottlenecks and inadequate routing flexibility in accelerators. We propose the layer-flexible non-blocking bypass network-on-chip (LNBN) architecture, integrating (1) a configurable non-blocking bypass router that adapts to multicast in large-scale DNNs, and parallel transmission in lightweight DNNs; (2) a flexible conflict-free routing algorithm that minimizes congestion and distinguishes concurrently executable traffic through path allocation based on layer dimensionality; (3) a block-based versatile mapping scheme that enables systematic routing with irregular layer structures and increases data reuse. These techniques significantly improve the performance and energy efficiency of DNN accelerators during inference. LNBN enhances network throughput by 23.35%, leading to an18.36% reduction in inference time and a 22.08% improvement in energy efficiency compared with dataflow-flexible DNN accelerator. Suk Bong Kang, Won Hyeok Kim, Tae Hee Han |
J. Supercomput. | 3 |
| 2023 | MRCN: Throughput-Oriented Multicast Routing for Customized Network-on-ChipsabstractThe relentless proliferation of Big Data and artificial intelligence has compelled computing platform architectures to evolve into heterogeneous multicores for greater energy efficiency. A customized network-on-chip (NoC) supporting interconnection diversity is pivotal for the asymmetric data-access traffic requirements of modern heterogeneous multicore system-on-chip (SoC). A significant portion of on-chip data access comprises single-source multi-destination (SSMD) traffic, which supports barrier synchronization, multi-threading, cache coherency protocols, and deep neural network (DNN) acceleration. By amortizing SSMD traffic, multicast routing is essential for effectively utilizing communication bandwidth. One of the primary concerns in supporting multicast routing in NoCs is to circumvent the additional deadlock conditions caused by branch operations among the active routers. However, it is challenging to implement the throughput-optimized multicast routing in irregular topology-based NoCs because the deadlock conditions become highly complicated, and the Hamiltonian path required to apply the labeling rule may not exist. Two important observations were identified regarding multicast routing in customized NoCs: 1) Even if the NoC lacks a Hamiltonian path, deadlock-freedom can be guaranteed by restricting branch operations to a specific destination. 2) A variable path diversity in a custom topology can be leveraged in routing path allocation and branch. Based on these properties, this study proposes a deadlock-free and throughput-enhanced multicast routing for customized NoC (MRCN). MRCN ensures deadlock freedom by utilizing extended routing and router labeling rules. Furthermore, destination router partitioning and traffic-aware adaptive branching are incorporated to reduce packet routing hops and disperse channel traffic. The effectiveness of MRCN was verified using Noxim, a well-known cycle-accurate NoC simulator, under various topologies and traffic patterns. The simulation revealed that MRCN improved the average latency by 13.98 % and the throughput by 12.16 % under the saturated traffic conditions over the previous multicast routings in customized NoCs. Young Sik Lee, Yong Wook Kim, Tae Hee Han |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | System-Level Signal Analysis Methodology for Optical Network-on-Chip Using Linear Model-Based CharacterizationabstractState-of-the-art silicon photonics technology has demonstrated its potential use in all required building blocks for ultrahigh bandwidth on-chip optical links. However, a robust system-level abstraction model reflecting the properties of optical devices has not been well established. We propose a linear optical device model (LODM) for silicon photonic devices and an associated computation method of optical signal propagation (CMOP) in an optical network-on-chip (ONoC). The CMOP manipulates the optical signal routing paths according to the topology, router configuration, and routing algorithm of the given ONoC architecture; thus, it allows the transformed information to be adaptable in an LODM to facilitate simplified analysis. Furthermore, we construct a linear system model of a microring resonator (MR) to reduce the computation complexities caused by its resonance structure. By using the CMOP, we accelerate the system-level analysis of optical signal propagation in ONoCs, reflecting the propagation loss, interference, and phase shift with close accuracy to analog and mixed-signal extensions (AMS) environments. The evaluation results show that the computation speeds up by three orders of magnitude with 1.57% error in accuracy when compared to the AMS environments. Min Su Kim, Yong Wook Kim, Tae Hee Han |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Insertion Loss-Aware Routing Analysis and Optimization for a Fat-Tree-Based Optical Network-on-ChipabstractFat-tree-based optical network-on-chip (FONoC) is an emerging architecture that enables next-generation computing platforms to achieve ultimate performance and energy efficiency. However, the architecture suffers from high insertion loss, which degrades energy efficiency and signal reliability severely. Focusing primarily on microring resonator (MR) drops, we analyze the relationship between the insertion loss caused by MR drops and the routing paths in the FONoCs. Our approach involves developing a simplified graph model named a drop-characterized fat-tree graph with vertex indexing. We propose three types of routing algorithms: 1) insertion loss-minimized deterministic routing; 2) minimized loss path-prioritized adaptive routing; and 3) insertion loss-constrained adaptive routing. Furthermore, we present the associated optical router architectures and additional insertion loss optimization by minimizing the number of waveguide crossings. Based on our simulation results for the latency, throughput, energy efficiency, and MR activation power, we discuss the tradeoffs and suggest appropriate optimization techniques to be adopted according to the priorities of the design goals. Jae Hoon Lee, Min Soo Kim 0008, Tae Hee Han |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Non-linear library characterization method for FinFET logic cells by L1-minimizationabstractState-of-the-art process technology offers ultra-low power devices operating at ultra-low voltages. However, they show a considerable level of non-linear characteristics. Hence, the accuracy of cell delay and variation modeling for logic cells is expected to be very low with a linear interpolation. In this paper, we propose a compressive sensing based high non-linear cell delay and variation modeling. This paper introduces accuracy optimization methods to fit the delay and variation modeling by pre-processing. Pre-processing is a hybrid approach combining a linear interpolation and compressive sensing for accurate restoration with using less samples. With FinFET cell delay and variation modeling, the experimental results show that the proposed method can obtain a similar or better accuracy with a half of measurement samples than a conventional linear interpolation based modeling. Byung-Su Kim, Hyo-Sig Won, Tae Hee Han, Joon-Sung Yang |
ISCAS | 3 |
| 2016 | AFSEM: Advanced frequent subcircuit extraction method by graph mining approach for optimized cell library developmentsabstractThe optimization of cells and cell combinations used in design is critical to enhance the performance. If frequently used cell combinations are known in advance, a new cell development can be significantly optimized using the cell combinations for chip design. However, extracting frequent cell combinations is an NP hard problem. We propose a new framework, referring as AFSEM, to extract frequent cell combinations for design optimization. To solve this problem, we use a frequent subgraph mining method which is a process of discovering subgraphs. We present an advanced graph modeling and optimized frequent subgraph mining platform for a practical use. The experimental results with various designs demonstrate that the proposed method can discover various types of subcircuits for design optimization with various runtime optimization methods. Byung-Su Kim, Hyo-Sig Won, Tae Hee Han, Joon-Sung Yang |
ISCAS | 3 |
| 2013 | A shortest path adaptive routing technique for minimizing path collisions in hybrid optical network-on-chip
Jae Hoon Lee, Young Seok Kim, Chang Lin Li, Tae Hee Han |
J. Syst. Archit. | 4 |
| 2009 | Frequency and yield optimization using power gates in power-constrained designsabstractManufactured dies exhibit a large spread of maximum frequency and leakage power due to process variations, which have been increasing with technology scaling. Reducing the spread is very important for maximizing the frequency and the yield of power-constrained designs, because otherwise many dies that do not satisfy frequency or power constraints would be discarded. In this paper, we propose two optimization methods to improve the maximum operating frequency and the yield using power gates that already exist in many power-constrained designs. In the first method, we consider the designs of multiple cores, where each of them can be independently power-gated. When each core shows different frequencies due to within-die variations, the strength of a power gate in each core is adjusted to make their maximum operating frequencies even. This allows faster cores to consume less active leakage power, reducing the total power consumption well below a power constraint in a globally-clocked design. We subsequently increase global supply voltage for higher overall frequency until the power constraint is satisfied. In our experiments assuming multicore processors with 2--16 cores, the maximum operating frequency was improved by 4-23%. In the second method, we take leaky-but-fast dies (which otherwise would be discarded) and adjust the strength of the power gates such that they can operate in an acceptable power and frequency region. The problem is extended to designs employing a frequency binning strategy, where we have an additional objective of maximizing the number of dies for higher frequency bins. In our experiments with ISCAS benchmark circuits, most discarded fast-but leaky dies were recovered using the second method. Nam Sung Kim, Jun Seomun, Abhishek A. Sinkar, Jungseob Lee, Tae Hee Han, Ken Choi, Youngsoo Shin |
ISLPED | 5 |
| 2009 | A New Demapper for BICM system with HARQabstractIn this paper, a new demapper generating log-likelihood ratio (LLR) is proposed by using a linear approximation for the nonlinear factor of an exact LLR. The exact LLR is sufficient for the probability based decoding, but its calculation requires a large number of computations in the case of high order modulations such as 16 quadrature amplitude modulation (QAM) and 64QAM. To alleviate the computational complexity a piecewise-linear demapper was proposed by neglecting the nonlinear factor of the exact LLR. However, it is observed that the throughput performance of the piecewise-linear demapper in hybrid automatic repeat request (HARQ) system is deficient in low signal-to-noise ratio region. Therefore, the new demapper is proposed which has an equivalent throughput performance to the exact LLR in HARQ system, and at the same time, only needs a few elementary operations for calculating the simplified LLRs. Experimental evidences which show the viability of the proposed demapper are also provided. Jin Whan Kang, Sang-Hyo Kim, Young Seok Jung, Seokho Yoon, Tae Hee Han |
VTC Fall | 5 |