EDBT 2026 Demo / reviewers in the wild / expert
Younghyun Lee
dblp:119/8219
· DBLP profile ↗
11ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing reservoir connectivity: A path to high-performance liquid state machines
Seungmin Oh, Unhyeon Kang, Jingyeong Hwang, Jiin Bang, Kyungmin Lee, Younghyun Lee, Jongkil Park 0001, Hyun Jae Jang, Changyoung Kim, Suyoun Lee |
Neurocomputing | 7 |
| 2026 | Pangaea v2: CXL-Based Disaggregated Memory System Architecture for Cloud-Native OrchestrationabstractToday’s data centers suffer from CPU and memory resource stranding because they often over-provision resources when deploying servers for worst-case scenarios. This problem gives rise to a disaggregated system architecture allowing each type of resource to be allocated, utilized and freed separately as required. In particular, research on disaggregated memory systems over the past few years has focused primarily on achieving low remote memory access latency over Ethernet, which is known as the RDMA optimization approach.In this paper, we introduce a dynamic rack-scale disaggregated memory system architecture, so called Pangaea v2 using ASIC-CXL H/W and memory orchestration S/W designed to increase the memory utilization of worker nodes between containerized applications execution in a Kubernetes, a major process container platform in the data center. In our evaluation with in-memory database application, disaggregated CXL memory system shows significantly better throughput improved by up to 10.2x/6.7x and 99th tail latency reduced to 96%/93% compared to RDMA with RoCEv2/InfiniBand. Han Deok Lee, Jehoon Park, Younghyun Lee, Junhyeok Im, Jin Jung, Jinin So, Siamak Tavallaei, Woo Taek Shim, Chin-Hua Chang, Sungwook Ryu, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001 |
IEEE Trans. Computers | 3 |
| 2024 | Discovering Efficient Fused Layer Configurations for Executing Multi-Workloads on Multi-Core NPUsabstractAs the AI industry grows rapidly, Neural Processing Units (NPUs) have been developed to deliver AI services more efficiently. One of the most important challenges for NPUs is task scheduling to minimize off-chip memory accesses, which may occur significant performance overhead. To reduce memory accesses, multiple convolution layers can be fused into a fused layer group, which offers numerous optimization opportunities. However, in most Convolutional Neural Networks (CNNs), when multiple layers are fused, the on-chip memory utilization of the fused layers gradually decreases, resulting in non-flat memory usage. In this paper, we propose a scheduling search algorithm to optimize the fusion of multiple convolution layers while reducing the peak on-chip memory usage. The proposed algorithm aims to find a schedule that simultaneously optimizes execution time and peak on-chip memory usage, despite a slight increase in off-chip memory accesses. It organizes the search space into a graph of possible partial schedules and then finds the optimal path. As a result of the improved on-chip memory usage, multiple workloads can be executed on multi-core NPUs with increased throughput. Experimental results show that the fusion schedule explored by the proposed method reduced on-chip memory usage by 39%, while increasing latency by 13%. When the freed on-chip memory was allocated to other workloads and the two workloads were executed concurrently in a multi-core NPU, a 32% performance improvement could be achieved. Younghyun Lee, Hyejun Kim, Yongseung Yu, Myeongjin Cho, Jiwon Seo 0002, Yongjun Park 0001 |
DATE | 1 |
| 2024 | An LPDDR-based CXL-PNM Platform for TCO-efficient Inference of Transformer-based Large Language ModelsabstractTransformer-based large language models (LLMs) such as Generative Pre-trained Transformer (GPT) have become popular due to their remarkable performance across diverse applications, including text generation and translation. For LLM training and inference, the GPU has been the predominant accelerator with its pervasive software development ecosystem and powerful computing capability. However, as the size of LLMs keeps increasing for higher performance and/or more complex applications, a single GPU cannot efficiently accelerate LLM training and inference due to its limited memory capacity, which demands frequent transfers of the model parameters needed by the GPU to compute the current layer(s) from the host CPU memory/storage. A GPU appliance may provide enough aggregated memory capacity with multiple GPUs, but it suffers from frequent transfers of intermediate values among GPU devices, each accelerating specific layers of a given LLM. As the frequent transfers of these model parameters and intermediate values are performed over relatively slow device-to-device interconnects such as PCIe or NVLink, they become the key bottleneck for efficient acceleration of LLMs. Focusing on accelerating LLM inference, which is essential for many commercial services, we develop CXL-PNM, a processing near memory (PNM) platform based on the emerging interconnect technology, Compute eXpress Link (CXL). Specifically, we first devise an LPDDR5X-based CXL memory architecture with 512GB of capacity and 1.1TB/s of bandwidth, which boasts 16× larger capacity and 10× higher bandwidth than GDDR6and DDR5-based CXL memory architectures, respectively, under a module form-factor constraint. Second, we design a CXLPNM controller architecture integrated with an LLM inference accelerator, exploiting the unique capabilities of such CXL memory to overcome the disadvantages of competing technologies such as HBM-PIM and AxDIMM. Lastly, we implement a CXLPNM software stack that supports seamless and transparent use of CXL-PNM for Python-based LLM programs. Our evaluation shows that a CXL-PNM appliance with 8 CXL-PNM devices offers 23% lower latency, 31% higher throughput, and 2.8× higher energy efficiency at 30% lower hardware cost than a GPU appliance with 8 GPU devices for an LLM inference service. Sangsoo Park, Kyungsoo Kim 0003, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim 0006, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, Jinhyun Kim, Yeongon Cho, Yongmin Tai, Jeonghyeon Cho, Hoyoung Song, Jung Ho Ahn, Nam Sung Kim |
HPCA | 8 |
| 2024 | Visually Dehallucinative Instruction GenerationabstractIn recent years, synthetic visual instructions by generative language model have demonstrated plausible text generation performance on the visual question-answering tasks. However, challenges persist in the hallucination of generative language models, i.e., the generated image-text data contains unintended contents. This paper presents a novel and scalable method for generating visually dehallucinative instructions, dubbed CAP2QA, that constrains the scope to only image contents. Our key contributions lie in introducing imagealigned instructive QA dataset CAP2QA-COCO and its scalable recipe. In our experiments, we compare synthetic visual instruction datasets that share the same source data by visual instruction tuning and conduct general visual recognition tasks. It shows that our proposed method significantly reduces visual hallucination while consistently improving visual recognition ability and expressiveness. Sungguk Cha, Jusung Lee, Younghyun Lee, Cheoljong Yang |
ICASSP | 3 |
| 2024 | Robust Face Recognition Based on an Angle-Aware Loss and Masked Autoencoder Pre-TrainingabstractDespite the advances in deep learning techniques, accurate identification using face recognition (FR) systems remains challenging owing to changes in face angles, bad lighting, and occlusions. To address these problems, we propose an optimized approach to improve the robustness of feature extraction models that are used in FR systems. The proposed method leverages an angle-aware loss function, inspired by ArcFace, that provides a large margin for significantly rotated faces. Additionally, a pre-trained weight initialization was derived from a masked autoencoder to enhance the ability of the model to cope with various poor conditions. The experimental results indicate that the proposed method outperforms existing face recognition methods in both normal and adverse environments. Jaehyeop Choi, Youngbaek Kim, Younghyun Lee |
ICASSP | 3 |
| 2023 | Block Group Scheduling: A General Precision-scalable NPU Scheduling Technique with Capacity-aware Memory AllocationabstractPrecision-scalable neural processing units (PSNPUs) efficiently provide native support for quantized neural networks. However, with the recent advancements of deep neural networks, PSNPUs are affected by a severe memory bottleneck owing to the need to perform an extreme number of simple computations simultaneously. In this study, we first analyze whether the memory bottleneck issue can be solved using conventional neural processing unit scheduling techniques. Subsequently, we introduce new capacity-aware memory allocation and block-level scheduling techniques to minimize the memory bottleneck. Compared with the baseline, the new method achieves up to 2.26× performance improvements by substantially relieving the memory pressure of low-precision computations without hardware overhead. Seokho Lee, Younghyun Lee, Hyejun Kim, Taehoon Kim 0001, Yongjun Park 0001 |
DATE | 2 |
| 2023 | Tailoring CUTLASS GEMM using Supervised LearningabstractGeneral matrix multiplication (GEMM) is a core computation kernel for deep neural networks. CUTLASS, a state-of-the-art open-source CUDA-based linear-algebra template library, provides a highly optimized tiling-based GEMM. However, CUTLASS GEMM often cannot achieve the optimal performance when its tiling configuration is not appropriately chosen because the performance varies significantly depending on some factors such as the tile size and shape, as well as the target graphics processing unit (GPU) architecture. Thus, determining the optimal tiling configuration is a major challenge in achieving the best performance of a tiling-based GEMM.To address this problem, we propose CUTLASS-tailor, a novel end-to-end framework that predicts the best tile parameters for target CUTLASS GEMM operations and underlying GPUs using a neural network model. We trained the prediction model using a suitable synthetic dataset that includes various input matrix combinations with different sizes and structures. Furthermore, to cover the various GPUs with a universal model, we also included the number of GPU cores and the amount of shared memory as GPU hardware features for the input of the CUTLASS-tailor network. On a test dataset from several real-world GEMMs, CUTLASS-tailor-based GEMM operations outperformed the GEMM operations using cuBLAS by up to 1.94× on an NVIDIA TitanXp GPU, and also showed that CUTLASS-tailor can find better tile parameters than well-known search algorithms. Yongseung Yu, Donghyun Son, Younghyun Lee, Sunghyun Park 0004, Giha Ryu, Myeongjin Cho, Jiwon Seo 0002, Yongjun Park 0001 |
ICCD | 3 |
| 2021 | Rotated Box Is Back: An Accurate Box Proposal Network for Scene Text Detection
Jusung Lee, Jaemyung Lee 0003, Cheoljong Yang, Younghyun Lee, Joonsoo Lee |
ICDAR (4) | 4 |
| 2012 | Crowd Density Estimation Using Multi-class AdaboostabstractIn this paper, we propose a crowd density estimation algorithm based on multi-class Adaboost using spectral texture features. Conventional methods based on self-organizing maps have shown unsatisfactory performance in practical scenarios, and in particular, they have exhibited abrupt degradation in performance under special conditions of crowd densities. In order to address these problems, we have developed a new training strategy by incorporating multi-class Adaboost with spectral texture features that represent a global texture pattern. According to the representative experimental results, the proposed method shows an average improvement of about 30% in the correct recognition rate, as compared to existing conventional methods. Daehun Kim, Younghyun Lee, Bonhwa Ku, Hanseok Ko |
AVSS | 2 |
| 2010 | License Plate Detection Using Local Structure PatternsabstractWe address the problem of license plate detection in video surveillance systems. The Adaboost based approach, known for relative ease of implementation, makes use of discriminative features such as edges or Haar-like features. In this paper, we propose a novel detection algorithm based on local structure patterns for license plate detection. The proposed algorithm includes post-processing methods to reduce false positive rate using positional and color information of license plates. Experimental results demonstrate effectiveness of the proposed method compared to both the edge and Haar-like feature based methods. Younghyun Lee, Taeyup Song, Bonhwa Ku, Seoungseon Jeon, David K. Han, Hanseok Ko |
AVSS | 1 |