Hyemi Min

dblp:207/7230 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-3261-626XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SENNA: Unified Hardware/Software Space Exploration for Parametrizable Neural Network Accelerators
abstract
Parametrizable neural network accelerators enable the deployment of targeted hardware for specialized environments. Finding the best architecture configuration for a given specification, however, is challenging. A large number of hardware configurations have to be considered, and for each hardware instance, an efficient software execution plan needs to be found, leading to a vast search space. Prior work has tackled this problem by dividing the search into subproblems for individual layers of a network. There is no guarantee, however, that the overall best hardware configuration that delivers the desired end-to-end performance across the entire network is among the best individual layer configurations. This work presents SENNA, a unified hardware/software space exploration framework for parametrizable neural network accelerators. To guide the exploration toward the overall best configuration, SENNA employs a multi-objective genetic algorithm with a novel design space representation that encodes the configuration of hardware and software parameters in a single chromosome. Using the Parallel Island Model (PIM), each layer is represented by one or more individual islands each containing a separate population to simultaneously search for the best configuration across the entire network. A tailored gene migration technique enables the exchange of genes between the populations of different islands. SENNA is evaluated with three parametrizable architectures and four neural networks. The evaluation result demonstrates that SENNA achieves upto 1.92x EDP improvement compared to the State-of-the-Art. With equivalent evaluation budgets, SENNA shows 2.5x–9.3x speedup compared to an Oracle scheme and the State-of-the-Art.
Jungyoon Kwon, Hyemi Min, Bernhard Egger 0002
ACM Trans. Embed. Comput. Syst.2
2023 Flexer: Out-of-Order Scheduling for Multi-NPUs
abstract
Recent neural accelerators often comprise multiple neural processing units (NPUs) with shared cache and memory. The regular schedules of state-of-the-art scheduling techniques miss important opportunities for memory reuse. This paper presents Flexer, an out-of-order (OoO) scheduler that maximizes instruction-level parallelism and data reuse on such multi-NPU systems. Flexer employs a list scheduling algorithm to dynamically schedule the tiled workload to all NPUs. To cope with the irregular data access patterns of OoO schedules, several heuristics help maximize data reuse by considering the availability of data tiles at different levels in the memory hierarchy. Evaluated with several neural networks on 2 to 4-core multi-NPUs, Flexer achieves a speedup of up to 2.2x and a 1.2-fold reduction in data transfers for individual layers compared to the best static execution order.
Hyemi Min, Jungyoon Kwon, Bernhard Egger 0002
CGO1
2021 Fast generation of optimized execution plans for parameterizable CNN accelerators: work-in-progress
abstract
Generating an optimal execution plan for a given convolutional neural network (CNN) and a parameterizable hardware accelerator is a challenge.We present a framework that finds an execution plan that maximizes throughput for a given network and a specific configuration of our parameterizable accelerator. The framework first generates tiled dataflows for each layer, then maps the dataflows to the different independent hardware units using techniques borrowed from traditional list scheduling. Evaluated with a number of different networks and different hardware configurations, the presented framework clearly outperforms existing approaches in terms of speedup or schedule generation time.
Hyemi Min, Jungyoon Kwon, Bernhard Egger 0002
CASES1
2018 Architectures and algorithms for user customization of CNNs
abstract
In this paper we present a convolutional neural network architecture that supports user customization through incremental transfer learning. The architecture consists of a large basic inference engine and a small augmenting engine. After training the basic inference engine and augmenting engine on a large general dataset, the basic inference engine is fixed. For user customization, only the augmenting engine is re-trained on-device using a small user specific dataset provided by the user. To accelerate the training of the augmenting engine we map this to a coarsegrained reconfigurable array processor. The complete network architecture is evaluated using the Caffe framework, and a C-code equivalent network is implemented and tested on a CGRA processor. Experiments with NIST'19 and our user-specific datasets show an increase in accuracy of the system from 76.3% to 93.2% after user customization. Mapping this code to a CGRA gives us a speed up of 45x and a 49-and 3-fold reduced energy consumption over an ARMv7 processor and a 3-way VLIW processor, respectively, showing the potential of CGRAs as DNN processors.
Barend Harris, Mansureh S. Moghaddam, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi
ASP-DAC6
2018 Auto-Tuning CNNs for Coarse-Grained Reconfigurable Array-Based Accelerators
abstract
As more and more deep learning tasks are pushed to mobile devices, accelerators for running these networks efficiently gain in importance. We show a that an existing class of general purpose accelerators, modulo-scheduled coarse-grained reconfigurable array (CGRA) processors typically used to accelerate multimedia workloads, can be a viable alternative to dedicated deep neural network processing hardware. To this end, an auto-tuning compiler is presented that maps convolutional neural networks (CNNs) efficiently on such architectures. The auto-tuner analyzes the structure of the CNN and the features of the CGRA, then explores the large optimization space to generate code that allows for an efficient mapping of the network. Evaluated with various CNNs, the auto-tuned code achieves an 11-fold speedup over the initial mapping. Comparing the energy per interference, the CGRA outperforms other general-purpose accelerators and an ARMv8 processor by a significant margin.
Inpyo Bae, Barend Harris, Hyemi Min, Bernhard Egger 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Incremental training of CNNs for user customization: work-in-progress
abstract
This paper presents a convolutional neural network architecture that supports transfer learning for user customization. The architecture consists of a large basic inference engine and a small augmenting engine. Initially, both engines are trained using a large dataset. Only the augmenting engine is tuned to the user-specific dataset. To preserve the accuracy for the original dataset, the novel concept of quality factor is proposed. The final network is evaluated with the Caffe framework, and our own implementation on a coarse-grained reconfigurable array (CGRA) processor. Experiments with MNIST, NIST'19, and our user-specific datasets show the effectiveness of the proposed approach and the potential of CGRAs as DNN processors.
Mansureh S. Moghaddam, Barend Harris, Duseok Kang, Inpyo Bae, Euiseok Kim, Hyemi Min, Hansu Cho, Sukjin Kim, Bernhard Egger 0002, Soonhoi Ha, Kiyoung Choi
CASES6