Sanghyo Park 0001

dblp:158/0642 · also Sang-Hyo Park 0001, Sang-hyo Park 0001 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-7282-7686ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Extended Scene Description using Relationship Information for Manipulating 3D Reconstructed Data
abstract
In this paper, we identify a critical limitation in the MPEG-I Scene Description (SD) format: the absence of a native mechanism for representing semantic relationships between objects. To address this, we propose a relationship extension framework that enables both static and dynamic inter-object semantics within 3D scenes. The proposed solution includes two integration strategies: (1) embedding relationship information within existing extensions such as scene interactivity and scene dynamic, and (2) introducing external relationship extensions that preserve the integrity of the core SD specification. We provide a concrete syntax and demonstrate the utility of our approach through applications in scene reconstruction and ISOBMFF-based integration. The results show that both options enable context-aware manipulation and efficient restoration of inter-object relations, supporting advanced interactive and media synchronization scenarios. We propose further evaluation and adoption of these extensions to enhance interoperability and backward compatibility in future MPEG standards.
Chae-yeong Song, Chaewon Moon, Aro Kim, Sanghyo Park 0001, Suhyeon Lee 0002, Sungjei Kim
DCC5
2026 Compression Framework for Light 3D Scene Graph Generation via Pruning-as-Search and Distillation
abstract
3D scene graph generation (3DSGG), which involves classifying objects and predicates, is an emerging topic in 3D scene understanding. Recent studies leveraging graph neural networks (GNNs) have introduced sophisticated architectures that enhance classification performance. However, since GNNs serve as the core and constitute the majority of parameters in 3DSGG models, their computational demands substantially increase overall complexity, which makes it difficult to determine the optimal model capacity. In this paper, we propose the first compression framework for lightweight 3DSGG models, based on pruning-as-search and knowledge distillation. This framework integrates multiple strategies and modules. In phase 1, the framework identifies the optimal compression ratio through pruning-as-search. In phase 2, to mitigate the accuracy loss incurred during compression, we employ structured pruning and a novel knowledge distillation strategy that effectively transfers precise information from the teacher to the compressed model. Experimental results show that our approach reduces model size by more than half while improving classification accuracy. Code is available athttps://github.com/hojunking/3DSGG-compression.
Hojun Song, Chae-yeong Song, Heejung Choi, Sungjei Kim, Sanghyo Park 0001
IEEE Trans. Multim.7
2025 Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
abstract
In recent text-video retrieval, the use of additional captions from vision-language models has shown promising effects on the performance. However, existing models using additional captions often have struggled to capture the rich semantics, including temporal changes, inherent in the video. In addition, incorrect information caused by generative models can lead to inaccurate retrieval. To address these issues, we propose a new framework, Narrating the Video (NarVid), which strategically leverages the comprehensive information available from frame-level captions, the narration. The proposed NarVid exploits narration in multiple ways: 1) feature enhancement through cross-modal interactions between narration and video, 2) query-aware adaptive filtering to suppress irrelevant or incorrect information, 3) dual-modal matching score by adding query-video similarity and query-narration similarity, and 4) hard-negative loss to learn discriminative features from multiple perspectives using the two similarities from different views. Experimental results demonstrate that NarVid achieves state-of-the-art performance on various benchmark datasets.
Chan Hur, Jeong-Hun Hong, Dabin Kang, Semin Myeong, Sanghyo Park 0001, Hyeyoung Park
CVPR6
2025 Human-oriented video retargeting via object detection and patch decision
Jaehyun Bae, Sukee Cho, Byungjun Bae, Sanghyo Park 0001
Multim. Tools Appl.7
2025 3DEKD: 3D Explanation-Based Knowledge Distillation for Pillar-Based 3D Object Detection
abstract
LiDAR-based 3D object detection has been widely utilized in fields such as autonomous driving and robotics. However, the black-box nature of 3D models limits their interpretability, making it difficult to understand their predictions and evaluate significant feature contributions, which are essential for safety-critical applications. To address this, we propose a novel knowledge distillation method for 3D object detection that integrates explanation-based and aggregation techniques to achieve effective knowledge transfer and, as a result, enhance the model’s interpretability. Our method generates attribution maps that highlight the importance of 3D points in the teacher model and aggregates them into a single map. This map is aligned with the student model’s pillar features using corresponding coordinates, allowing pillar-wise feature mapping. Building on this feature mapping, to the best of our knowledge, this is the first study to propose a distillation process that effectively transfers the teacher model’s explanations of critical regions to the student model. Experimental results demonstrate that the proposed method increases 3D and BEV mAP by up to 2.09% and 0.84%, respectively, compared to the existing models.
Heejung Choi, Dabin Kang, Jaehyup Lee, Sanghyo Park 0001
IEEE Signal Process. Lett.5
2024 Pruning-guided feature distillation for an efficient transformer-based pose estimation model
abstract
Abstract The authors propose a compression strategy for a 3D human pose estimation model based on a transformer which yields high accuracy but increases the model size. This approach involves a pruning‐guided determination of the search range to achieve lightweight pose estimation under limited training time and to identify the optimal model size. In addition, the authors propose a transformer‐based feature distillation (TFD) method, which efficiently exploits the pose estimation model in terms of both model size and accuracy by leveraging transformer architecture characteristics. Pruning‐guided TFD is the first approach for 3D human pose estimation that employs transformer architecture. The proposed approach was tested on various extensive data sets, and the results show that it can reduce the model size by 30% compared to the state‐of‐the‐art while ensuring high accuracy.
Aro Kim, Jong Taek Lee, Sungjei Kim, Sanghyo Park 0001
IET Comput. Vis.7
2024 Fast depth intra mode decision using intra prediction cost and probability in 3D-HEVC
Sanghyo Park 0001
Multim. Tools Appl.2
2021 Adaptive fractional motion and disparity estimation skipping in MV-HEVC
Sanghyo Park 0001
J. Vis. Commun. Image Represent.2
2021 Fast Multi-Type Tree Partitioning for Versatile Video Coding Using a Lightweight Neural Network
abstract
In this paper, we propose a fast decision scheme using a lightweight neural network (LNN) to avoid redundant block partitioning in versatile video coding (VVC). A more versatile block structure, named the multi-type tree (MTT) structure, which includes binary trees (BTs) and ternary trees (TTs), is adopted by VCC, in addition to the traditional quadtree structure. The MTT improved the coding efficiency compared with previous video coding standards. However, the new tree structures, mainly TT, significantly increased the complexity of the VVC encoder. Although widespread application of VVC has been inhibited, this problem has not yet been investigated thoroughly in the literature. In this study, we first determine the statistical characteristics of coded parameters that exhibit correlation with the TT and develop two useful types of features—explicit VVC features(EVFs) andderived VVC features(DVFs)—to facilitate the intra coding of VVC. These features can be obtained efficiently during the intra prediction before the determination of the best block partitioning during rate-distortion optimization in VVC encoding. Our LNN model decides whether to terminate the nested TT block structures subsequent to a quadtree based on the features. The experimental results confirm that the proposed method substantially decreases the encoding complexity of VVC with a slight coding loss under the All Intra configuration. Our code, models, and dataset are available athttps://github.com/foriamweak/MTTPartitioning_LNN.
Sanghyo Park 0001, Je-Won Kang
IEEE Trans. Multim.1
2017 An Efficient Motion Estimation Method for QTBT Structure in JVET Future Video Coding
abstract
In this paper, an efficient motion estimation method for quadtree plus binary tree (QTBT) structure is presented for JVET future video coding (FVC). To exploit the possibility to design a new coding standard better than HEVC, an activity called future video coding has been active in MPEG and VCEG. One of the most influential new technologies proposed for FVC is QTBT structure, which can give more flexibility on the prediction block than quadtree-based HEVC. To advance the compression efficiency of QTBT, we propose a method that, for accurate motion estimation, sets a new point and adjusts search range accordingly based on the motion information of parent node. Performance evaluation was conducted on top of joint exploration model (JEM 3.0), which resulted in a 0.41% gain on average in terms of BD-rate under random access scenario without any burden on encoder complexity.
Sanghyo Park 0001, Euee S. Jang
DCC1
2016 Zero coefficient-aware fast butterfly-based inverse discrete cosine transform algorithm
abstract
The latest video coding standards, including Moving Picture Experts Group‐4 (MPEG‐4) advanced video coding (AVC)/H.264 and high‐efficiency video coding (HEVC), use a discrete cosine transform (DCT) process as the core for compression efficiency, sacrificing the computational complexity at decoder. There have been a number of attempts to reduce the complexity of inverse DCT (IDCT). Butterfly‐based factorisation remains the most commonly used method for such a reduction. In this study, the authors propose a zero (Z) coefficient‐aware fast butterfly‐based IDCT algorithm for video decoding. They focus on a reduction in the computational complexity of the butterfly‐based 8 × 8 IDCT by removing the unnecessary computations of one‐dimensional (1D) IDCT kernels, and adaptively applying IDCT kernels based on the number of non‐Z DCT coefficients to speed‐up 1D data. Their experimental results show that the average operation numbers using the proposed IDCT is approximately half that for the 8 × 8 IDCT implemented in the MPEG‐4 AVC/H.264 and HEVC reference software. The improved computational complexity of the proposed method is demonstrated by measuring the running time, which requires only one‐half of the IDCT time using the reference software.
Sanghyo Park 0001, Kiho Choi, Euee S. Jang
IET Image Process.1