VLDB 2026 Research / reviewers in the wild / expert
Youngsik Kim
dblp:89/652
· DBLP profile ↗
11ranked-venue papers
5as first author
3since 2021 · last 2024
0000-0003-2842-4190ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | TCP: A Tensor Contraction Processor for AI Workloads Industrial ProductabstractWe introduce a novel tensor contraction processor (TCP) architecture that offers a paradigm shift from traditional architectures that rely on fixed-size matrix multiplications. TCP aims at exploiting the rich parallelism and data locality inherent in tensor contractions, thereby enhancing both efficiency and performance of AI workloads.TCP is composed of coarse-grained processing elements (PEs) to simplify software development. In order to efficiently process operations with diverse tensor shapes, the PEs are designed to be flexible enough to be utilized as a large-scale single unit or a set of small independent compute units.We aim at maximizing data reuse on both levels of inter and intra compute units. To do that, we propose a circuit switch-based fetch network to flexibly connect compute units to enable inter-compute unit data reuse. We also exploit input broadcast to multiple contraction engines and input buffer based reuse to further exploit reuse behavior in tensor contraction. Our compiler explores the design space of tensor contractions considering tensor shapes and the order of their associated loop operations as well as the underlying accelerator architecture.A TCP chip was designed and fabricated in 5nm technology as the second-generation product of Furiosa AI, offering 256/512/1024 TOPS (BF16/FP8 or INT8/INT4) with 256 MB SRAM and 1.5 TB/s 48 GB HBM3 under 150 W TDP. Commercialization will start in August 2024.We performed an extensive case study of running the LLaMA-2 7B model and evaluated its performance and power efficiency on various configurations of sequence length and batch size. For this model, TCP is 2.7 × and 4.1 × better than H100 and L40s, respectively, in terms of performance per watt. Hanjoon Kim, Byeongwook Bae, Hyunmin Jeong, Sang Min Lee 0014, Jeseung Yeon, Changjae Park, Boncheol Gu, Changman Lee, Jaeick Bae, SungGyeong Bae, Yojung Cha, Wooyoung Choe, Jonguk Choi, Juho Ha, Hyuck Han, Namoh Hwang, Seokha Hwang, Kiseok Jang, Haechan Je, Hojin Jeon, Jaewoo Jeon, Hyunjun Jeong, Yeonsu Jung, Dongok Kang, Hyewon Kim, Muhwan Kim, Sewon Kim, Suhyung Kim, Yong Kim, Youngsik Kim, Younki Ku, Jeong Ki Lee, Juyun Lee, Seokho Lee, Minwoo Noh, Hyuntaek Oh, Gyunghee Park, Jimin Seo, Jungyoung Seong, June Paik, Nuno P. Lopes, Sungjoo Yoo |
ISCA | 35 |
| 2022 | Effective Algorithm to Control Depth Level for Performance Improvement of Sound TracingabstractSound tracing, a 3D sound rendering technology based on ray tracing, is a very costly method for calculating sound propagation. To reduce its expense, we propose an algorithm for adjusting the depth based on frame coherence and spatial characteristics. The results of the experiment indicate that when the sound source and listener were indoors, the reflection path loss rate was 3%, the diffraction path loss rate was 15.4%, and the total frame rate increased by 6.25%. When the listener was outdoors and the sound source was indoors, the reflection path and diffraction path loss rate were 0%, and the total frame rate was increased by 33.33 compared to the conventional method. Thus, the proposed algorithm can improve rendering performance while minimizing path loss rate. Eunjae Kim, Juwon Yun, Woo-Nam Chung, Jae-Ho Nah, Youngsik Kim, Cheoung Ghil Kim, Woo-Chan Park |
J. Web Eng. | 5 |
| 2021 | Lossless Compression Algorithm and Architecture for Reduced Memory Bandwidth Requirement with Improved Prediction Based on the Multiple DPCM Golomb-Rice AlgorithmabstractIn a computing environment, higher resolutions generally require more memory bandwidth, which inevitably leads to the consumption more power. This may become critical for the overall performance of mobile devices and graphic processor units with increased amounts of memory access and memory bandwidth. This paper proposes a lossless compression algorithm with a multiple differential pulse-code modulation variable sign code Golomb-Rice to reduce the memory bandwidth requirement. The efficiency of the proposed multiple differential pulse-code modulation is enhanced by selecting the optimal differential pulse code modulation mode. The experimental results show compression ratio of 1.99 for high-efficiency video coding image sequences and that the proposed lossless compression hardware can reduce the bus bandwidth requirement. Imjae Hwang, Juwon Yun, Woo-Nam Chung, Jaeshin Lee, Cheong-Ghil Kim, Youngsik Kim, Woo-Chan Park |
J. Web Eng. | 6 |
| 2016 | Improving Write Performance by Controlling Target Resistance Distributions in MLC PRAMabstractMulti-level cell (MLC) phase change RAM (PRAM) is expected to offer lower cost main memory than DRAM. However, poor write performance is one of the most critical problems for practical applications of MLC PRAM. In this article, we present two schemes to improve write performance by controlling the target resistance distribution of MLC PRAM cells. First, we propose multiple RESET/SET operations that relax the target resistance bands of intermediate logic levels with additional RESET/SET operations, which reduces the program time of intermediate logic levels, thereby improving write performance. Second, we propose a two-step write scheme consisting of lightweight write and idle-time completion write that exploits the fact that hot dirty data tend to be overwritten in a short time period and the MLC PRAM often has long idle times. Experimental results show that the multiple RESET/SET and two-step write schemes result in an average IPC improvement of 15.7% and 10.4%, respectively, on a hybrid DRAM/PRAM main memory subsystem. Furthermore, their integrated solution results in an average IPC improvement of 23.2% (up to 46.4%). Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2015 | Entity Linking Korean Text: An Unsupervised Learning Approach using Semantic RelationsabstractAlthough entity linking is a widely researched topic, the same cannot be said for entity linking geared for languages other than English.Several limitations including syntactic features and the relative lack of resources prevent typical approaches to entity linking to be used as effectively for other languages in general.We describe an entity linking system that leverage semantic relations between entities within an existing knowledge base to learn and perform entity linking using a minimal environment consisting of a part-of-speech tagger.We measure the performance of our system against Korean Wikipedia abstract snippets, using the Korean DBpedia knowledge base for training.Based on these results, we argue both the feasibility of our system and the possibility of extending to other domains and languages in general. Youngsik Kim, Key-Sun Choi |
CoNLL | 1 |
| 2014 | Named Entity Corpus Construction using Wikipedia and DBpedia Ontology
YoungGyun Hahm, Jungyeul Park, Youngsik Kim, Dosam Hwang, Key-Sun Choi |
LREC | 4 |
| 2013 | The No-Prop algorithm: A new learning algorithm for multilayer neural networks
Bernard Widrow, Aaron Greenblatt, Youngsik Kim, Dookun Park |
Neural Networks | 3 |
| 2012 | Write performance improvement by hiding R drift latency in phase-change RAMabstractPhase-change RAM (PRAM) is considered to be one of the most promising candidates to complement or replace DRAM in the near future. However, it is imperative to overcome the limitations of PRAM, especially, long write latency for its widespread applications. R drift latency occupies a significant portion in PRAM write latency thereby adversely affecting system performance. In this paper, we propose a novel method called write status holding register (WSHR) to reduce the write latency due to R drift latency. The WSHR allows for non-blocking accesses to PRAM during R drift latency thereby improving system performance. Our experiments with SPEC benchmarks show that the proposed WSHR gives 53.6%~0% performance improvements in the hybrid DRAM/PRAM main memory (256MB DRAM and 14nm PRAM). Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
DAC | 1 |
| 2012 | A case study on the application of real phase-change RAM to main memory subsystemabstractPhase-change RAM (PCM) has the advantages of better scaling and non-volatility compared with the DRAM which is expected to face its scaling limit in the near future. There have been many studies on applying the PCM to main memory in order to complement or replace the DRAM. One common limitation of these studies is that they are based on synthetic PCM models. In our study, we investigate the feasibility and issues of applying a real PCM to main memory. In this paper, we report our case study of characterizing the PCM and evaluating its usefulness in the main memory. Our results show that the PCM/DRAM hybrid main memory with a modest DRAM size can give comparable performance to that of the DRAM only main memory. However, the hybrid memory with small DRAMs or large footprint programs can suffer from performance degradation due to the long latency of both PCM writes and write preemption penalty, which requires architectural innovations for exploiting the full potential of PCM write performance. Suknam Kwon, Dongki Kim, Youngsik Kim, Sungjoo Yoo, Sunggu Lee |
DATE | 3 |
| 2008 | Automated formal verification of scheduling with speculative code motionsabstractWe present a methodology for formal verification of scheduling phase of High-Level Synthesis (HLS) when speculative code motions are performed during this process. Verification relies on establishing functional equivalence between the result of scheduling and the behavioral specification of the design, using their FSMD models. We propose and formally define a relation between the two FSMDs that is less constrained than the strong equivalence, but stronger than weak equivalence. For verification of scheduling involving speculative code motions, we propose the notion of FSMD recomposition, a transformation that alters the state sequence and/or the operation of each state, while maintaining the functionality. The equivalence conditions are formulated in higher-order logic, and their correctness is verified in the theorem proving environment PVS. The entire verification flow, including formal model extraction and proof generation is fully automated. Youngsik Kim, Nazanin Mansouri |
ACM Great Lakes Symposium on VLSI | 1 |
| 2005 | Exploiting PSL standard assertions in a theorem-proving-based verification environmentabstractAssertion-based design is becoming more widely used in industry. However, little has been done to take advantage of existing design assertions in the theorem-proving verification environments. In this paper, we present our work on development of the semi-automated theorem-proving based verification system ROVERIFIC that makes use of existing design assertions. We have defined generic predicate templates that capture the semantics of PSL, and a subset of Verilog. ROVERIFIC uses these templates, and automatically compiles a design under verification (Verilog) and its assertions (PSL) into the higher-order predicates of the PVS [11] theorem proving system. Design verification can be subsequently conducted by proving the correctness properties. Youngsik Kim, Parija Sule, Nazanin Mansouri |
ACM Great Lakes Symposium on VLSI | 1 |