Hyunsuk Ko

dblp:97/9144 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-7015-8351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-tuning
abstract
We introduce LOREN, a curvature-aware zeroth-order (ZO) optimization method for fine-tuning large language models (LLMs). Existing ZO methods, which estimate gradients via finite differences using random perturbations, often suffer from high variance and suboptimal search directions. Our approach addresses these challenges by: (i) reformulating the problem of gradient preconditioning as that of adaptively estimating an anisotropic perturbation distribution for gradient estimation, (ii) capturing curvature through a low-rank block diagonal preconditioner using the framework of natural evolution strategies, and (iii) applying a REINFORCE leave-one-out (RLOO) gradient estimator to reduce variance. Experiments on standard LLM benchmarks show that our method outperforms state-of-the-art ZO methods by achieving higher accuracy and faster convergence, while cutting peak memory usage by up to 27.3% compared with MeZO-Adam.
Hyunseok Seung, Hyunsuk Ko
AAAI3
2026 Mean activation curvature for scalable second-order optimization in deep networks
abstract
Abstract Second-order methods can accelerate deep neural network training, but their adoption is limited by the cost and instability of estimating and inverting curvature matrices. We revisit Kronecker-factored Fisher approximations via an empirical structural analysis of activation and gradient statistics in modern architectures. Across a range of models, we find that activation statistics capture most of the effective curvature directions, while gradient statistics mainly act as a global or diagonal rescaling. Based on this observation, we propose mean activation curvature , a scalable curvature surrogate that yields two optimizers, MAC and SMAC , offering different trade-offs between expressiveness and efficiency. We further extend the construction to self-attention layers with a structured approximation that retains the role of attention scores, and provide a convergence analysis under standard assumptions. Experiments on vision and language models show that MAC / SMAC matches or improves the accuracy of existing second-order baselines while reducing training time and memory usage. Code: github.com/hseung88/mac .
Hyunseok Seung, Hyunsuk Ko
Knowl. Inf. Syst.3
2026 Adaptive Neural In-Loop Filtering via Boundary-Aware Skipping and Time-Distortion Optimization
abstract
Neural network-based in-loop filters (NNILFs) have recently been integrated into emerging video coding frameworks to enhance reconstruction quality beyond the capabilities of traditional signal processing filters. However, the high computational complexity of these filters significantly increases encoding and decoding time, hindering their real-time deployment. In this article, we propose an efficient neural in-loop filter control framework that improves decoding efficiency while preserving coding performance. Our approach introduces boundary strength-aware control to selectively skip NNILF operations on visually less critical regions. In addition, a time-distortion optimization (TDO) strategy is presented to adaptively manage the tradeoff between visual distortion and inference time during encoding. Experimental results on the NNVC reference software show that the proposed methods reduce decoding time by up to 23% and 21% for the random access (RA) and low-delay B (LDB) configurations, respectively. Apart from a minimal luma loss only in RA, chroma gains were also achieved, with Y: 0.10% U: −0.01% V: −0.14% for RA and Y: 0.00% U: −1.34% V: −0.94% for LDB. These results show that the proposed control framework offers a practical and effective solution for integrating NN-based tools into future video coding systems.
Hyukmin Kwon, Juyeon Seo, Donghyun Kim 0017, Sung-Chang Lim, Hyunsuk Ko
ACM Trans. Multim. Comput. Commun. Appl.5
2025 MAC: An Efficient Gradient Preconditioning Using Mean Activation Approximated Curvature
abstract
Second-order optimization methods for training neural networks, such as KFAC, exhibit superior convergence by utilizing curvature information of loss landscape. However, it comes at the expense of high computational burden. In this work, we analyze the two components that constitute the layer-wise Fisher information matrix (FIM) used in KFAC: the Kronecker factors related to activations and pre-activation gradients. Based on empirical observations on their eigenspectra, we propose efficient approximations for them, resulting in a computationally efficient optimization method called MAC. To the best of our knowledge, MAC is the first algorithm to apply the Kronecker factorization to the FIM of attention layers used in transformers and explicitly integrate attention scores into the preconditioning. We also study the convergence property of MAC on nonlinear neural networks and provide two conditions under which it converges to global minima. Our extensive evaluations on various network architectures and datasets show that the proposed method outperforms KFAC and other state-of-the-art methods in terms of accuracy, end-to-end training time, and memory usage.
Hyunseok Seung, Hyunsuk Ko
ICDM3
2025 Phase Distribution Matters: On the Importance of Phase Distribution Alignment (PDA) in Holographic Applications
Seungmi Choi, TaeHwa Lee, Jun Yeong Cha, Suhyun Jo, Hyunmin Ban, Kwan-Jung Oh, Hyunsuk Ko, Hui Yong Kim
ACM Multimedia7
2024 NysAct: A Scalable Preconditioned Gradient Descent using Nyström Approximation
abstract
Adaptive gradient methods are computationally efficient and converge quickly, but they often suffer from poor generalization. In contrast, second-order methods enhance convergence and generalization but typically incur high computational and memory costs. In this work, we introduce NYSACT, a scalable first-order gradient preconditioning method that strikes a balance between state-of-the-art first-order and second-order optimization methods. NYSACT leverages an eigenvalue-shifted Nyström method to approximate the activation covariance matrix, which is used as a preconditioning matrix, significantly reducing time and memory complexities with minimal impact on test accuracy. Our experiments show that NYSACT not only achieves improved test accuracy compared to both first-order and second-order methods but also demands considerably less computational resources than existing second-order methods.
Hyunseok Seung, Hyunsuk Ko
IEEE Big Data3
2024 Facial image deblurring network for robust illuminance adaptation and key structure restoration
Yongrok Kim, Hyukmin Kwon, Hyunsuk Ko
Eng. Appl. Artif. Intell.3
2022 Video Quality Model of Compression, Resolution and Frame Rate Adaptation Based on Space-Time Regularities
abstract
Being able to accurately predict the visual quality of videos subjected to various combinations of dimension reduction protocols is of high interest to the streaming video industry, given rapid increases in frame resolutions and frame rates. In this direction, we have developed a video quality predictor that is sensitive to spatial, temporal, or space-time subsampling combined with compression. Our predictor is based on new models of space-time natural video statistics (NVS). Specifically, we model the statistics of divisively normalized difference between neighboring frames that are relatively displaced. In an extensive empirical study, we found that those paths of space-time displaced frame differences that provide maximal regularity against our NVS model generally align best with motion trajectories. Motivated by this, we built a new video quality prediction engine that extracts NVS features that represent how space-time directional regularities are disturbed by space-time distortions. Based on parametric models of these regularities, we compute features that are used to train a regressor that can accurately predict perceptual quality. As a stringent test of the new model, we apply it to the difficult problem of predicting the quality of videos subjected not only to compression, but also to downsampling in space and/or time. We show that the new quality model achieves state-of-the-art (SOTA) prediction performance on the new ETRI-LIVE Space-Time Subsampled Video Quality (STSVQ) and also on the AVT-VQDB-UHD-1 database.
Dae Yeol Lee, Hyunsuk Ko, Alan C. Bovik
IEEE Trans. Image Process.3
2022 A Subjective and Objective Study of Space-Time Subsampled Video Quality
abstract
Video dimensions are continuously increasing to provide more realistic and immersive experiences to global streaming and social media viewers. However, increments in video parameters such as spatial resolution and frame rate are inevitably associated with larger data volumes. Transmitting increasingly voluminous videos through limited bandwidth networks in a perceptually optimal way is a current challenge affecting billions of viewers. One recent practice adopted by video service providers is space-time resolution adaptation in conjunction with video compression. Consequently, it is important to understand how different levels of space-time subsampling and compression affect the perceptual quality of videos. Towards making progress in this direction, we constructed a large new resource, called the ETRI-LIVE Space-Time Subsampled Video Quality (ETRI-LIVE STSVQ) database, containing 437 videos generated by applying various levels of combined space-time subsampling and video compression on 15 diverse video contents. We also conducted a large-scale human study on the new dataset, collecting about 15,000 subjective judgments of video quality. We provide a rate-distortion analysis of the collected subjective scores, enabling us to investigate the perceptual impact of space-time subsampling at different bit rates. We also evaluated and compare the performance of leading video quality models on the new database. The new ETRI-LIVE STSVQ database is being made freely available at (https://live.ece.utexas.edu/research/ETRI-LIVE_STSVQ/index.html).
Dae Yeol Lee, Somdyuti Paul, Christos G. Bampis, Hyunsuk Ko, Seyoon Jeong, Blake Homan, Alan C. Bovik
IEEE Trans. Image Process.4
2020 Video Quality Model for Space-Time Resolution Adaptation
abstract
Delivering voluminous amounts of video data through limited bandwidth channels is a challenge affecting billions of viewers. Accordingly, it is becoming more important to understand the perceptual effects that arise from various dimension reduction methodologies. Towards this direction, we propose a new video quality model that predicts the perceptual quality of videos undergoing varying levels of spatio-temporal subsampling and compression. The new model is established upon the natural statistics principle of videos, which leverage the fact that pristine videos obey statistical regularities that are disturbed by distortions. We found that there exist space-time paths between video frames that best preserve the statistical regularity inherent in the spatial structure of the video frames. The distribution features extracted from frame differences displaced in the direction of these paths correlate more highly with human subjective quality opinions than those from non-displaced frame differences. Given that non-displaced frame differences are widely utilized in video quality models, the improved efficiency of spatially and/or temporally displaced (possibly by more than one frame) frame differences, is an important finding that may significantly elevate the success of studies on temporal features and video quality.
Dae Yeol Lee, Hyunsuk Ko, Alan C. Bovik
IPAS2
2020 Edge-Preserving Reference Sample Filtering and Mode-Dependent Interpolation for Intra-Prediction
abstract
High Efficiency Video Coding is the latest video compression standard, which achieves the best coding performance up until now. Specifically, intra prediction is a tool that removes spatial redundancy in a single frame and then a predictor is generated from its neighboring reference samples based on a specific interpolation scheme. In this paper, we propose an edge-preserving intra reference sample filtering method using a bilateral filter, which is implemented as hardware-friendly. Two parameters of the bilateral filter are modeled by block size and mean amplitude of the pixel intensity. In addition, a mode-dependent interpolation scheme is proposed, which takes the directionality of angular predictions into account. The experimental results show that a BD rate-reduction of 0.63% can be achieved for all intra configurations by combining the two methods. We also demonstrate that the subjective quality of the reconstructed frames can be improved.
Hyunsuk Ko, Jungwon Kang, Hui Yong Kim
IEEE Trans. Circuits Syst. Video Technol.1
2020 Quality Prediction on Deep Generative Images
abstract
In recent years, deep neural networks have been utilized in a wide variety of applications including image generation. In particular, generative adversarial networks (GANs) are able to produce highly realistic pictures as part of tasks such as image compression. As with standard compression, it is desirable to be able to automatically assess the perceptual quality of generative images to monitor and control the encode process. However, existing image quality algorithms are ineffective on GAN generated content, especially on textured regions and at high compressions. Here we propose a new "naturalness"-based image quality predictor for generative images. Our new GAN picture quality predictor is built using a multi-stage parallel boosting system based on structural similarity features and measurements of statistical similarity. To enable model development and testing, we also constructed a subjective GAN image quality database containing (distorted) GAN images and collected human opinions of them. Our experimental results indicate that our proposed GAN IQA model delivers superior quality predictions on the generative image datasets, as well as on traditional image quality datasets.
Hyunsuk Ko, Dae Yeol Lee, Seunghyun Cho, Alan C. Bovik
IEEE Trans. Image Process.1
2018 Learning-Based Just-Noticeable-Quantization- Distortion Modeling for Perceptual Video Coding
abstract
Conventional predictive video coding-based approaches are reaching the limit of their potential coding efficiency improvements, because of severely increasing computation complexity. As an alternative approach, perceptual video coding (PVC) has attempted to achieve high coding efficiency by eliminating perceptual redundancy, using just-noticeable-distortion (JND) directed PVC. The previous JNDs were modeled by adding white Gaussian noise or specific signal patterns into the original images, which were not appropriate in finding JND thresholds due to distortion with energy reduction. In this paper, we present a novel discrete cosine transform-based energy-reduced JND model, called ERJND, that is more suitable for JND-based PVC schemes. Then, the proposed ERJND model is extended to two learning-based just-noticeable-quantization-distortion (JNQD) models as preprocessing that can be applied for perceptual video coding. The two JNQD models can automatically adjust JND levels based on given quantization step sizes. One of the two JNQD models, called LR-JNQD, is based on linear regression and determines the model parameter for JNQD based on extracted handcraft features. The other JNQD model is based on a convolution neural network (CNN), called CNN-JNQD. To our best knowledge, our paper is the first approach to automatically adjust JND levels according to quantization step sizes for preprocessing the input to video encoders. In experiments, both the LR-JNQD and CNN-JNQD models were applied to high efficiency video coding (HEVC) and yielded maximum (average) bitrate reductions of 38.51% (10.38%) and 67.88% (24.91%), respectively, with little subjective video quality degradation, compared with the input without preprocessing applied.
Sehwan Ki, Sung-Ho Bae, Munchurl Kim, Hyunsuk Ko
IEEE Trans. Image Process.4
2017 Just-noticeable-quantization-distortion based preprocessing for perceptual video coding
abstract
Conventional predictive video coding may no longer become capable of effectively accommodating the demand of high quality video services with continuously increasing spatiotemporal resolutions as before since it is reaching the limit of its coding efficiency improvement. As an alternative, perceptual video coding (PVC) is being exploited by effectively removing perceptual redundancy for coding efficiency improvement, one of which is just-noticeable-distortion (JND) directed PVC. Unfortunately, the previous JND modeling is not often suitable for JND-directed PVC approaches because quantization effects are not considered. Thus, we presents a new DCT-domain JND model that considers the quantization operation in video compression into JND modeling for PVC, which is calledjust noticeable quantization distortion (JNQD) model. Our proposed JNQD model can be applied as preprocessing prior to any video compression scheme by adding a parameter to adapt the model to quantization step sizes. For experiments, our JNQD models have been applied to High Efficiency Video Coding (HEVC) and yielded the maximum and average bitrate reductions of 37.35% and 13.05%, respectively with little subjective video quality degradation, compared to the input without preprocessing applied. Moreover, it can be applicable for any encoder as preprocessing, which can have a large flexibility compared to previous encoder-dependent schemes.
Sehwan Ki, Munchurl Kim, Hyunsuk Ko
VCIP3
2017 Robust uncalibrated stereo rectification with constrained geometric distortions (USR-CGD)
Hyunsuk Ko, Han Suk Shim, Ouk Choi, C.-C. Jay Kuo
Image Vis. Comput.1
2017 A ParaBoost stereoscopic image quality assessment (PBSIQA) system
Hyunsuk Ko, Rui Song 0003, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.1
2009 Fast mode-decision for H.264/AVC based on inter-frame correlations
Hyunsuk Ko, Kiwon Yoo, Kwanghoon Sohn
Signal Process. Image Commun.1