VLDB 2026 Research / reviewers in the wild / expert
Hyunmin Jeong
dblp:286/1937
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 8 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TCP: A Tensor Contraction Processor for AI Workloads Industrial ProductabstractWe introduce a novel tensor contraction processor (TCP) architecture that offers a paradigm shift from traditional architectures that rely on fixed-size matrix multiplications. TCP aims at exploiting the rich parallelism and data locality inherent in tensor contractions, thereby enhancing both efficiency and performance of AI workloads.TCP is composed of coarse-grained processing elements (PEs) to simplify software development. In order to efficiently process operations with diverse tensor shapes, the PEs are designed to be flexible enough to be utilized as a large-scale single unit or a set of small independent compute units.We aim at maximizing data reuse on both levels of inter and intra compute units. To do that, we propose a circuit switch-based fetch network to flexibly connect compute units to enable inter-compute unit data reuse. We also exploit input broadcast to multiple contraction engines and input buffer based reuse to further exploit reuse behavior in tensor contraction. Our compiler explores the design space of tensor contractions considering tensor shapes and the order of their associated loop operations as well as the underlying accelerator architecture.A TCP chip was designed and fabricated in 5nm technology as the second-generation product of Furiosa AI, offering 256/512/1024 TOPS (BF16/FP8 or INT8/INT4) with 256 MB SRAM and 1.5 TB/s 48 GB HBM3 under 150 W TDP. Commercialization will start in August 2024.We performed an extensive case study of running the LLaMA-2 7B model and evaluated its performance and power efficiency on various configurations of sequence length and batch size. For this model, TCP is 2.7 × and 4.1 × better than H100 and L40s, respectively, in terms of performance per watt. Hanjoon Kim, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee 0014, Jeseung Yeon, Changjae Park, Boncheol Gu, Changman Lee, Jaeick Bae, SungGyeong Bae, Yojung Cha, Wooyoung Choe, Jonguk Choi, Juho Ha, Hyuck Han, Namoh Hwang, Seokha Hwang, Kiseok Jang, Haechan Je, Hojin Jeon, Jaewoo Jeon, Hyunjun Jeong, Yeonsu Jung, Dongok Kang, Hyewon Kim, Muhwan Kim, Sewon Kim, Suhyung Kim, Yong Kim, Youngsik Kim, Younki Ku, Jeong Ki Lee, Juyun Lee, Seokho Lee, Minwoo Noh, Hyuntaek Oh, Gyunghee Park, Jimin Seo, Jungyoung Seong, June Paik, Nuno P. Lopes, Sungjoo Yoo |
ISCA | 5 |
| 2024 | Polarization Diversity and Transfer Learning-Based Modulation Optimization for High-Speed Dual Channel MIMO Backscatter CommunicationabstractThis paper presents a novel approach to address the challenges in backscatter communication for the Internet of Things. The traditional use of I/Q load modulator based on parametric transistor models often suffers from performance degradation due to discrepancies with actual transistor models caused by thermal and environmental noise sources. To overcome this issue, the paper proposes an active circuit modeling technique based on artificial neural network (ANN)-based transfer learning, which utilizes actual measurement data to model the I/Q load modulator accurately. Furthermore, an optimization algorithm is applied to achieve an optimal high-order modulation scheme, leading to improved energy efficiency by 40%. By leveraging machine learning-based modeled I/Q modulators, the proposed approach enables high-speed wireless data communication in a dual-channel configuration. The paper also conducts theoretical analysis to define the required performance of a dual-polarized Vivaldi antenna for implementing polarization diversity in backscatter communication. This analysis provides guidelines for achieving optimal performance in terms of spectral efficiency and error vector magnitude (EVM). The experimental results demonstrate that the proposed approach achieves a spectral efficiency of 2.0 bps/Hz based on 4-QAM modulation within a 150 MHz bandwidth. The measured EVM is 9.35%, indicating the effectiveness of the proposed technique in achieving reliable and efficient wireless data communication in backscatter systems. This paper presents a comprehensive approach combining accurate circuit modeling, optimization algorithms, and theoretical analysis to enable high-speed, ultra-low-power wireless data communication in backscatter communication systems. Hyunmin Jeong, Nohgyeom Ha, Apostolos Georgiadis, Manos M. Tentzeris, Sangkil Kim |
IEEE Internet Things J. | 1 |
| 2023 | AutoScaleDSE: A Scalable Design Space Exploration Engine for High-Level SynthesisabstractHigh-Level Synthesis (HLS) has enabled users to rapidly develop designs targeted for FPGAs from the behavioral description of the design. However, to synthesize an optimal design capable of taking better advantage of the target FPGA, a considerable amount of effort is needed to transform the initial behavioral description into a form that can capture the desired level of parallelism. Thus, a design space exploration (DSE) engine capable of optimizing large complex designs is needed to achieve this goal. We present a new DSE engine capable of considering code transformation, compiler directives (pragmas), and the compatibility of these optimizations. To accomplish this, we initially express the structure of the input code as a graph to guide the exploration process. To appropriately transform the code, we take advantage of ScaleHLS based on the multi-level compiler infrastructure (MLIR). Finally, we identify problems that limit the scalability of existing DSEs, which we name the “design space merging problem.” We address this issue by employing a Random Forest classifier that can successfully decrease the number of invalid design points without invoking the HLS compiler as a validation tool. We evaluated our DSE engine against the ScaleHLS DSE, outperforming it by a maximum of 59×. We additionally demonstrate the scalability of our design by applying our DSE to large-scale HLS designs, achieving a maximum speedup of 12× for the benchmarks in the MachSuite and Rodinia set. Hyegang Jun, Hanchen Ye, Hyunmin Jeong, Deming Chen |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | ScaleHLS: a scalable high-level synthesis framework with multi-level transformations and optimizations: invitedabstractThis paper presents an enhanced version of a scalable HLS (High-Level Synthesis) framework named ScaleHLS, which can compile HLS C/C++ programs and PyTorch models to highly-efficient and synthesizable C++ designs. The original version of ScaleHLS achieved significant speedup on both C/C++ kernels and PyTorch models [14]. In this paper, we first highlight the key features of ScaleHLS on tackling the challenges present in the representation, optimization, and exploration of large-scale HLS designs. To further improve the scalability of ScaleHLS, we then propose an enhanced HLS transform and analysis library supported in both C++ and Python, and a new design space exploration algorithm to handle HLS designs with hierarchical structures more effectively. Comparing to the original ScaleHLS, our enhanced version improves the speedup by up to 60.9× on FPGAs. ScaleHLS is fully open-sourced at https://github.com/hanchenye/scalehls. Hanchen Ye, Hyegang Jun, Hyunmin Jeong, Stephen Neuendorffer, Deming Chen |
DAC | 3 |
| 2022 | ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate RepresentationabstractHigh-level synthesis (HLS) has been widely adopted as it significantly improves the hardware design productivity and enables efficient design space exploration (DSE). Existing HLS tools are built using compiler infrastructures largely based on a single-level abstraction, such as LLVM. How-ever, as HLS designs typically come with intrinsic structural or functional hierarchies, different HLS optimization problems are often better solved with different levels of abstractions. This paper proposes ScaleHLS1, a new scalable and customizable HLS framework, on top of a multi-level compiler infrastructure called MLIR. ScaleHLS represents HLS designs at multiple representation levels and provides an HLS-dedicated analysis and transform library to solve the optimization problems at the suitable levels. Using this library, we provide a DSE engine to generate optimized HLS designs automatically. In addition, we develop an HLS C front-end and a C/C++ emission back-end to translate HLS designs into/from MLIR for enabling an end-to-end compilation flow. Experimental results show that, comparing to the baseline designs without manual directives insertion and code-rewriting, that are only optimized by Xilinx Vivado HLS, ScaleHLS improves the performances with amazing quality-of-results – up to 768.1× better on computation kernel level programs and up to 3825.0× better on neural network models. Hanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong, Jack Huang, Stephen Neuendorffer, Deming Chen |
HPCA | 4 |
| 2021 | TwinDNN: A Tale of Two Deep Neural NetworksabstractCompression technologies for deep neural networks (DNNs), such as weight quantization, have been widely investigated to reduce the model size so that they can be implemented on hardware with strict resource restrictions. However, one major downside of model compression is accuracy degradation. To deal with this problem effectively, we propose a new compressed network inference scheme with a high accuracy but slower DNN coupled with its highly compressed DNN version that typically delivers much faster inference speed but with a lower accuracy. During the inference, we determine the confidence of the prediction of the compressed DNN, and infer the original neural network for the inputs that are considered not confident by the compressed DNN. The proposed design uses a balanced number of resources available on the hardware and can deliver overall accuracy close to the high accuracy model, but with the inference speed closer to the compressed DNN. We demonstrate our design on two image classification tasks: CIFAR-10 and ImageNet. Our experiments show that our design can recover up to 94% of accuracy drop caused by extreme network compression, with more than 90% speedup compared to just using the original DNN. This is more than 17% extra accuracy recovery and 36% extra speedup compared to the previous work with a similar concept on VGG-16. This is the first work that considers using a highly compressed DNN along with the original DNN in parallel to achieve high accuracy and speed at the same time, while maintaining the resource balance by using two different main computation sources efficiently on an FPGA. Hyunmin Jeong, Deming Chen |
ASAP | 1 |
| 2021 | TwinDNN: A Tale of Two Deep Neural NetworksabstractMachine learning is one of the most popular fields in the current era. It is used in various areas, such as speech recognition, face recognition, medical diagnosis, etc. However, the problem is that the neural networks for machine learning applications are becoming too large and slow as they get more complicated and powerful. This problem gets especially serious when neural networks are used for edge devices with a small chip. As a result, researchers have proposed two major solutions to solve this problem. Hyunmin Jeong, Deming Chen |
FCCM | 1 |
| 2021 | PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis FlowabstractThe exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, in high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap between hardware design flows and software applications, and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric Python-based programming and synthesis flow for FPGA. PyLog is powered by a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details and allows designers to focus on algorithm specification. PyLog captures more high-level computation patterns for better optimization than traditional HLS systems. PyLog also has a runtime for running PyLog code directly on FPGA platform without any extra code development. Evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA version by 3.17× and 1.24× on average. Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu |
FPGA | 3 |
| 2021 | PyLog: An Algorithm-Centric Python-Based FPGA Programming and Synthesis FlowabstractThe exploding complexity and computation efficiency requirements of applications are stimulating a strong demand for hardware acceleration with heterogeneous platforms such as FPGAs. However, a high-quality FPGA design is very hard to create as it requires FPGA expertise and a long design iteration time. In contrast, software applications are typically developed in a short development cycle, with high-level languages like Python, which is at a much higher level of abstraction than all existing hardware design flows. To close this gap and simplify FPGA programming, we create PyLog, a high-level, algorithm-centric programming and synthesis flow for FPGA. PyLog features a set of compiler optimization passes and a type inference system to generate high-quality design. It abstracts away the implementation details, and allows designers to focus on algorithm specification. PyLog takes in Python functions and generates complete optimized FPGA system design. PyLog also has a runtime that allows users to run the PyLog code directly on the target FPGA platform without any extra code development. The whole design flow is automated. The evaluation shows that PyLog significantly improves FPGA design productivity and generates highly efficient FPGA designs that outperform highly optimized CPU and FPGA versions by 3.17x and 1.24x on average. Sitao Huang, Kun Wu 0002, Hyunmin Jeong, Chengyue Wang 0002, Deming Chen, Wen-Mei W. Hwu |
IEEE Trans. Computers | 3 |