EDBT 2026 Demo / reviewers in the wild / expert
Chanho Park 0004
dblp:98/3338-4
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0005-7524-539XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NPUWattch: ML-Based Power, Area, and Timing Modeling for Neural AcceleratorsabstractPre-silicon modeling tools for characterizing power, area, and timing (PAT) have enabled numerous architectural studies, but traditional analytical and table-based models begin to exhibit limitations in their applicability as architectural design complexity increases and process technology scales below 5 nm with the emergence of advanced transistors. Previous modeling techniques typically assume static scaling factors across different designs and technology nodes, derived from small circuit design benchmarks using old processes. Consequently, they do not reflect complex design variability and nonlinear projection to advanced technology nodes. Moreover, reference logic and SRAM implementations serving as the baseline for design and technology scaling are often created using different technologies and design rules, leading to significant estimation inaccuracies that distort the relative contributions of individual components. To address these challenges, this paper introduces NPUWattch, a machine learning-based PAT modeling framework for neural accelerators. It leverages neural network regression models to learn complex nonlinear relationships in technology and design scaling based on diverse post-layout logic and SRAM design datasets formulated using unified technology libraries. To this end, we developed technology libraries from 65 nm to 2 nm, constructed and validated diverse logic and SRAM datasets, and trained neural network models using an adaptive loss function to reinforce underrepresented regions of the design space. NPUWattch is validated against the post-layout results of numerous open-source neural accelerators, and evaluation results demonstrate that NPUWattch outperforms existing tools with an average estimation error of 2.7%, offering reliable and accurate PAT estimation. Minkwan Kim, Chanho Park 0004, Hanmok Park, Taigon Song, William J. Song |
HPCA | 3 |
| 2024 | Nona: Accurate Power Prediction Model Using Neural NetworksabstractThis paper proposes a neural-network-based power model, Nona, that accurately predicts the power consumption of heterogeneous CPUs on a commercial mobile device. With aggressive on-device power management in action, it becomes increasingly challenging to make accurate power predictions for diverse applications. To overcome the limitations of the existing power models based on linear regression, Nona uses a lightweight neural network with a small number of performance monitoring counters (PMCs) chosen from a system analysis and a loss function designed for power prediction. Experiments on Google Pixel 6 show that Nona has a 3.4% average prediction error, improving on prior work by 2.6x. HoSun Choi, Chanho Park 0004, Euijun Kim, William J. Song |
DAC | 2 |
| 2023 | NeuroSpector: Systematic Optimization of Dataflow Scheduling in DNN AcceleratorsabstractThis paper presents an optimization framework namedNeuroSpectorthat systematically analyzes the dataflow of deep neural network (DNN) accelerators and rapidly identifies optimal execution methods. The proposed methodology is demonstrated to work effectively with a variety of accelerator architectures and DNN workloads. It has been a baffling challenge to devise scheduling schemes for neural accelerators to maximize energy efficiency and performance. The challenge lies in that hardware specifications associated with multi-dimensional DNN data create an enormous number of possible scheduling options that can be exerted on accelerators. Related work suggested various techniques to solve the challenge encompassing brute-force search of massive solution spaces pruned by user constraints, solving the objective functions of system models, learning-based optimization, etc. However, each suggested technique was devised only for a specific accelerator model. Therefore, we find that they are not adaptively applicable to different accelerators and DNN workloads in that they produce hit-or-miss results with 100.1% greater energy and cycles on average compared to optimal scheduling schemes obtained from fully comprehensive brute-force searches. In contrast, NeuroSpector identifies efficient execution methods for various accelerators and workloads with only 1.5% differences on average to the optimal scheduling solutions. The optimization strategy of NeuroSpector is based on an observation that optimal executions are strongly correlated with minimizing data movements to the lower-level memory hierarchy of accelerators rather than maximizing the utilization of processing elements. Thus, NeuroSpector prioritizes optimizing lower-level components in the accelerator hierarchy, which is proven highly effective for various accelerators and DNN workloads. Chanho Park 0004, Bogil Kim, Sungmin Ryu, William J. Song |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | The Nebula Benchmark Suite: Implications of Lightweight Neural NetworksabstractThis article presents a benchmark suite namedNebulathat implements lightweight neural network benchmarks. Recent neural networks tend to form deeper and sizable networks to enhance accuracy and applicability. However, the massive volume of heavy networks makes them highly challenging to use in conventional research environments such as microarchitecture simulators. We notice that neural network computations are mainly comprised of matrix and vector calculations that repeat on multi-dimensional data encompassing batches, channels, layers, etc. This observation motivates us to develop a variable-sized neural network benchmark suite that provides users with options to select appropriate size of benchmarks for different research purposes or experiment conditions. Inspired by the implementations of well-known benchmarks such as PARSEC and SPLASH suites, Nebula offers various size options from large to small datasets for diverse types of neural networks. The Nebula benchmark suite is comprised of seven representative neural networks built on a C++ framework. The variable-sized benchmarks can be executed i) with acceleration libraries (e.g., BLAS, cuDNN) for faster and realistic application runs or ii) without the external libraries if execution environments do not support them, e.g., microarchitecture simulators. This article presents a methodology to develop the variable-sized neural network benchmarks, and their performance and characteristics are evaluated based on hardware measurements. The results demonstrate that the Nebula benchmarks reduce execution time as much as 25x while preserving similar architectural behaviors as the full-fledged neural networks. Bogil Kim, Chanho Park 0004, William J. Song |
IEEE Trans. Computers | 3 |