Younghyun Cho

dblp:185/0226 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
abstract
While many advanced LLMs are designed to handle long sequence data, we can still observe notable quality degradation even within the sequence limit.In this work, we introduce a novel approach called Scaling to Emphasize Attention for Long-context retrieval (SEAL), which enhances the retrieval performance of large language models (LLMs) over long contexts.We observe that specific attention heads are closely tied to long-context retrieval, showing positive or negative correlation with retrieval scores, and adjusting the strength of these heads boosts the quality of LLMs in long context by a large margin.Built on this insight, we propose a learning-based mechanism that leverages generated data to emphasize these heads.By applying SEAL, we achieve significant improvements in long-context retrieval performance across various tasks and models.Additionally, when combined with existing training-free context extension techniques, SEAL extends the contextual limits of LLMs while maintaining highly reliable outputs.
Changhun Lee, Minsang Seok, Jungyu Jin, Younghyun Cho, Eunhyeok Park
ACL (1)4
2025 EnQode: Fast Amplitude Embedding for Quantum Machine Learning Using Classical Data
abstract
Amplitude embedding (AE) is essential in quantum machine learning (QML) for encoding classical data onto quantum circuits. However, conventional AE methods suffer from deep, variable-length circuits that introduce high output error due to extensive gate usage and variable error rates across samples, resulting in noise-driven inconsistencies that degrade model accuracy. We introduce EnQode, a fast AE technique based on symbolic representation that addresses these limitations by clustering dataset samples and solving for cluster mean states through a low-depth, machine-specific ansatz. Optimized to reduce physical gates and SWAP operations, EnQode ensures all samples face consistent, low noise levels by standardizing circuit depth and composition. With over 94% fidelity in data mapping, EnQode enables robust, high-performance QML on noisy intermediate-scale quantum (NISQ) devices. Our opensource solution provides a scalable and efficient alternative for integrating classical data with quantum models.
Jason Han, Nicholas S. DiBrita, Younghyun Cho, Hengrui Luo, Tirthak Patel
DAC3
2025 PTQ4VM: Post-Training Quantization for Visual Mamba
abstract
Visual Mamba is an approach that extends the selective space state model, Mamba, to vision tasks. It processes image tokens sequentially in a fixed order, accumulating information to generate outputs. Despite its growing popularity for delivering high-quality outputs at a low computational cost across various tasks, Visual Mamba is highly susceptible to quantization, which makes further performance improvements challenging. Our analysis reveals that the fixed token access order in Visual Mamba introduces unique quantization challenges, which we categorize into three main issues: 1) token-wise variance, 2) channel-wise outliers, and 3) a long tail of activations. To address these challenges, we propose Post-Training Quantization for Visual Mamba (PTQ4VM), which introduces two key strategies: Per-Token Static (PTS) quantization and Joint Learning of Smoothing Scale and Step Size (JLSS). To the our best knowledge, this is the first quantization study on Visual Mamba. PTQ4VM can be applied to various Visual Mamba backbones, converting the pre-trained model to a quantized format in under 15 minutes without notable quality degradation. Extensive experiments on large-scale classification and regression tasks demonstrate its effectiveness, achieving up to 1.83x speedup on GPUs with negligible accuracy loss compared to FP16. Our code is available at https://github.com/YoungHyun197/ptq4vm.
Younghyun Cho, Changhun Lee, Seonggon Kim, Eunhyeok Park
WACV1
2023 Harnessing the Crowd for Autotuning High-Performance Computing Applications
abstract
This paper presents GPTuneCrowd, a crowd-based autotuning framework for tuning high-performance computing applications. GPTuneCrowd collects performance data from various users using a user-friendly tuner interface. GPTuneCrowd then presents novel autotuning techniques, based on transfer learning and parameter sensitivity analysis, to maximize tuning quality using collected data from the crowd. This paper shows several real-world case studies of GPTuneCrowd. Our evaluation shows that GPTuneCrowd’s transfer learning improves the tuned performance of ScaLAPACK’s PDGEQRF by 1.57x and a plasma fusion code NIMROD by 2.97x, over a non-transfer learning autotuner. We use GPTuneCrowd’s sensitivity analysis to reduce the search space of SuperLU_DIST and Hypre. Tuning on the reduced search space achieves 1.17x and 1.35x better tuned performance of SuperLU_DIST and Hypre, respectively, compared to the original search space.
Younghyun Cho, James Demmel, Jacob King, Xiaoye S. Li, Yang Liu 0179, Hengrui Luo
IPDPS1
2022 Dopia: online parallelism management for integrated CPU/GPU architectures
abstract
Recent desktop and mobile processors often integrate CPU and GPU onto the same die. The limited memory bandwidth of these integrated architectures can negatively affect the performance of data-parallel workloads when all computational resources are active. The combination of active CPU and GPU cores achieving the maximum performance depends on a workload's characteristics, making manual tuning a time-consuming task. Dopia is a fully automated framework that improves the performance of data-parallel workloads by adjusting the Degree Of Parallelism on Integrated Architectures. Dopia transparently analyzes and rewrites OpenCL kernels before executing them with the number of CPU and GPU cores expected to yield the best performance. Evaluated on AMD and Intel integrated processors, Dopia achieves 84% of the maximum performance attainable by an oracle.
Younghyun Cho, Jiyeon Park, Florian Negele, Changyeon Jo, Thomas R. Gross, Bernhard Egger 0002
PPoPP1
2020 Evaluation of memory performance in NUMA architectures using Stochastic Reward Nets
Reza Entezari-Maleki, Younghyun Cho, Bernhard Egger 0002
J. Parallel Distributed Comput.2
2020 Performance Modeling of Parallel Loops on Multi-Socket Platforms Using Queueing Systems
abstract
Predicting the performance of parallel loops on modern shared-memory multi-socket multi-core systems in dependence of the allocated resources is an important means to achieve better system utilization. Previous prediction techniques are tied to specific architectures and do not allow for purely online performance predictions without requiring an offline analysis of the parallel program. This paper presents a practical approach based on queueing theory to model the performance of parallel programs in dependence of the number of allocated core resources. Based on the key insight that scalability of scientific parallel loops is limited by memory performance, a hierarchically constructed M/M/1/N/N queue system is used to analytically compute the response time at the different congestion points in the memory system of modern NUMA architectures. After automatically tuning the model to a specific architecture by executing a number of micro-benchmarks, the required parameter values are obtained at runtime from hardware performance counters present in modern commodity AMD and Intel processors. Evaluated with 24 OpenMP parallel loops on a 64-core AMD and a 72-core Intel multi-socket platform, the presented queueing system is able to accurately predict the speedup of parallel loops with a mean absolute percentage error of 8.3 percent on the AMD system and 6.7 percent on the Intel platform.
Younghyun Cho, Surim Oh, Bernhard Egger 0002
IEEE Trans. Parallel Distributed Syst.1
2018 Maximizing system utilization via parallelism management for co-located parallel applications
abstract
With an increasing number of cores and memory controllers in multiprocessor platforms, co-location of parallel applications is gaining on importance. Key to achieve good performance is allocating the proper number of threads to co-located applications. This paper presents NuPoCo, a framework for automatically managing parallelism of co-located parallel applications on NUMA multi-socket multi-core systems. NuPoCo maximizes the utilization of CPU cores and memory controllers by dynamically adjusting the number of threads for co-located parallel applications. Evaluated with various scenarios of co-located OpenMP applications on a 64-core AMD and a 72-core Intel machine, NuPoCo achieves a reduction of the total turnaround time by 10-20% compared to the default Linux scheduler and an existing parallelism management policy focusing on CPU utilization only.
Younghyun Cho, Camilo A. Celis Guzman, Bernhard Egger 0002
PACT1
2018 On-the-fly workload partitioning for integrated CPU/GPU architectures
abstract
Integrating CPUs and GPUs on the same die provides new opportunities for optimization, especially for irregular data-parallel workloads that fail to fully exploit the computational power of the GPU. Such workloads benefit from a proper partitioning between the CPU and the GPU. This paper presents an on-the-fly workload partitioning technique for irregular workloads on integrated architectures. Unlike existing work, no prior analysis of the workload is required. GPU kernels and input data are analyzed and optimized at runtime. The technique executes work chunks of similar load on the GPU and assigns irregular chunks to the CPU. Evaluated with various irregular workloads, the method achieves a 1.4x--7.1x speedup over GPU execution on AMD and Intel processors.
Younghyun Cho, Florian Negele, Seohong Park, Bernhard Egger 0002, Thomas R. Gross
PACT1
2017 POSTER: Improving NUMA System Efficiency with a Utilization-Based Co-scheduling
abstract
This work proposes a co-scheduling technique for co-located parallel applications on Non-Uniform Memory Access (NUMA) multi-socket multi-core platforms. The technique allocates core resources for running parallel applications such that both the utilization of the memory controllers and the CPU cores are maximized. Utilization is predicted using an online performance prediction model based on queuing systems. At runtime, the core allocation is periodically re-evaluated and cores are re-assigned to executing applications. Experimental results show that the proposed co-scheduling technique is able to execute co-located parallel applications in significantly less total execution time than the default Linux scheduler and a conventional scalability-based scheduler.
Younghyun Cho, Camilo A. Celis Guzman, Bernhard Egger 0002
PACT1
2016 Online Scalability Characterization of Data-Parallel Programs on Many Cores
abstract
We present an accurate online scalability prediction model for data-parallel programs on NUMA many-core systems. Memory contention is considered to be the major limiting factor of program scalability as data parallelism limits the amount of synchronization or data dependencies between parallel work units. Reflecting the architecture of NUMA systems, contention is modeled at the last-level caches of the compute nodes and the memory nodes using a two-level queuing model to estimate the mean service time of the individual memory nodes. Scalability predictions for individual or co-located parallel applications are based solely on data obtained during a short sampling period at runtime; this allows the presented model to be employed in a variety of scenarios. The proposed model has been implemented into an open-source OpenCL and the GNU OpenMP runtime and evaluated on a 64-core AMD system. For a wide variety of parallel workloads and configurations, the evaluations show that the model is able to predict the scalability of data-parallel kernels with high accuracy.
Younghyun Cho, Surim Oh, Bernhard Egger 0002
PACT1
2016 Adaptive Space-Shared Scheduling for Shared-Memory Parallel Programs
Younghyun Cho, Surim Oh, Bernhard Egger 0002
JSSPP1
2016 Efficient Checkpointing of Live Virtual Machines
abstract
The ability to save the state of a running virtual machine (VM) for later restoration is an important tool for home, server, and virtual desktop cloud (VDC) environments in order to achieve optimal and balanced hardware utilization. With guest memory sizes of four to eight gigabytes being the norm the time- and space-overhead of storing VM checkpoints still prevents an effective use of the technique. This work presents a method for fast and space-efficient checkpointing of VMs. Based on the observation that operating systems cache disk blocks in memory, the proposed technique transparently intercepts I/O operations and maintains an up-to-date mapping of memory pages and disk blocks containing identical data. At a checkpoint, those memory pages are excluded from the checkpoint image leading to a significant reduction of both the time and space required to take a checkpoint of a running VM. The broad applicability and good performance of the proposed method is demonstrated by an extensive set of experiments. We have implemented the technique for para-virtualized (PV), PVHVM, and fully-virtualized (HVM) guests in the Xen hypervisor. In comparison with an unmodified Xen hypervisor, experiments with Linux and Windows guests, on average, achieve a 86, 76, 53, and 47 percent reduction in the stored data and a 73, 62, 47, and 38 percent shorter time required to take a checkpoint for PV, PVHVM, HVM Linux, and HVM Windows guests, respectively.
Bernhard Egger 0002, Younghyun Cho, Changyeon Jo, Eunbyung Park, Jaejin Lee
IEEE Trans. Computers2