Jinsun Park

dblp:169/4840 · DBLP profile ↗
← Back
26ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-2296-819XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 9 since 2021Systems, architecture and hardware · 2 · 1 first-author
YearPublicationVenuePosition
2026 OceanSplat: Object-aware Gaussian Splatting with Trinocular View Consistency for Underwater Scene Reconstruction
abstract
We introduce OceanSplat, a novel 3D Gaussian Splatting-based approach for high-fidelity underwater scene reconstruction. To overcome multi-view inconsistencies caused by scattering media, we design a trinocular setup for each camera pose by rendering from horizontally and vertically translated virtual viewpoints, enforcing view consistency to facilitate spatial optimization of 3D Gaussians. Furthermore, we derive synthetic epipolar depth priors from the virtual viewpoints, which serve as self-supervised depth regularizers to compensate for the limited geometric cues in degraded underwater scenes. We also propose a depth-aware alpha adjustment that modulates the opacity of 3D Gaussians during early training based on their depth along the viewing direction, deterring the formation of medium-induced primitives. Our approach promotes the disentanglement of 3D Gaussians from the scattering medium through effective geometric constraints, enabling accurate representation of scene structure and significantly reducing floating artifacts. Experiments on real-world underwater and simulated scenes demonstrate that OceanSplat substantially outperforms existing methods for both scene reconstruction and restoration in scattering media.
Minseong Kweon, Jinsun Park
AAAI2
2026 AT-adapter: Leveraging attribute knowledge of CLIP for few-shot classification
Yonghyeon Jo, Janghyun Kim, Chanill Park, Seonghoon Choi, Jinsun Park
Neurocomputing6
2025 Interactive Sign Language Question Answering Framework and Dataset for Barrier-Free Exhibition Service
abstract
Recent advancements in deep learning have shown significant potential in developing barrier-free environments for the hearing-impaired. However, the lack of diverse and specialized datasets limits the development of high-level services such as interactive museum exhibitions. In this paper, we introduce KSL-Ex, a Korean Sign Language (KSL) dataset designed to support real-world interactive exhibition services. Our dataset comprises 29,574 sign language video samples, including isolated words, continuous KSL sentences, questionanswer pairs, and detailed annotations. In addition, we propose an interactive sign language question answering framework that leverages state-of-the-art large language models which have demonstrated outstanding performance in question answering tasks. The proposed framework consists of four key components: The sign language recognition (SLR) model, the answering module, the refinement module, and the 3 D sign language animation module. These components effectively process sign language queries, generate proper answers, and produce sign language animations to provide interactive communication interfaces for the hearing-impaired. Experimental results show the effectiveness of both KSL-Ex and the proposed framework, demonstrating their potential in providing real-world interactive sign language services.
Seokyong Heo, Minsu Jeon, Jinwan Kim, Seongheon Ha, Suhyeon Cho, Namyeong Heo, Jinsun Park
FG8
2025 Federated Domain Generalization with Data-free On-server Matching Gradient
abstract
Domain Generalization (DG) aims to learn from multiple known source domains a model that can generalize well to unknown target domains. One of the key approaches in DG is training an encoder which generates domain-invariant representations. However, this approach is not applicable in Federated Domain Generalization (FDG), where data from various domains are distributed across different clients. In this paper, we introduce a novel approach, dubbed Federated Learning via On-server Matching Gradient (FedOMG), which can efficiently leverage domain information from distributed domains. Specifically, we utilize the local gradients as information about the distributed models to find an invariant gradient direction across all domains through gradient inner product maximization. The advantages are two-fold: 1) FedOMG can aggregate the characteristics of distributed models on the centralized server without incurring any additional communication cost, and 2) FedOMG is orthogonal to many existing FL/FDG methods, allowing for additional performance improvements by being seamlessly integrated with them. Extensive experimental evaluations on various settings demonstrate the robustness of FedOMG compared to other FL/FDG baselines. Our method outperforms recent SOTA baselines on four FL benchmark datasets (MNIST, EMNIST, CIFAR-10, and CIFAR-100), and three FDG benchmark datasets (PACS, VLCS, and OfficeHome). The reproducible code is publicly available~\footnote[1]{\url{https://github.com/skydvn/fedomg}}.
Trong-Binh Nguyen, Duong Minh Nguyen, Jinsun Park, Viet Quoc Pham, Won-Joo Hwang
ICLR3
2025 BAC-GCN: Background-Aware CLIP-GCN Framework for Unsupervised Multi-Label Classification
Yonghyeon Jo, Janghyun Kim, Jinsun Park
ACM Multimedia3
2025 Dual interaction network with cross-image attention for medical image segmentation
Jeonghyun Noh, Wangsu Jeon, Jinsun Park
Pattern Recognit. Lett.3
2024 CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
abstract
Recent LiDAR-based 3D Object Detection (3DOD) methods show promising results, but they often do not generalize well to target domains outside the source (or training) data distribution. To reduce such domain gaps and thus to make 3DOD models more generalizable, we introduce a novel unsupervised domain adaptation (UDA) method, called CMDA, which (i) leverages visual semantic cues from an image modality (i.e., camera images) as an effective semantic bridge to close the domain gap in the cross-modal Bird's Eye View (BEV) representations. Further, (ii) we also introduce a self-training-based learning strategy, wherein a model is adversarially trained to generate domain-invariant features, which disrupt the discrimination of whether a feature instance comes from a source or an unseen target domain. Overall, our CMDA framework guides the 3DOD model to generate highly informative and domain-adaptive features for novel data distributions. In our extensive experiments with large-scale benchmarks, such as nuScenes, Waymo, and KITTI, those mentioned above provide significant performance gains for UDA tasks, achieving state-of-the-art performance.
Gyusam Chang, Wonseok Roh, Sujin Jang, Daehyun Ji, Gyeongrok Oh, Jinsun Park, Jinkyu Kim 0001, Sangpil Kim
AAAI7
2024 Exploiting Cross-Modal Cost Volume for Multi-sensor Depth Estimation
Janghyun Kim, Ukcheol Shin, Seokyong Heo, Jinsun Park
ACCV (9)4
2024 ULTRON: Unifying Local Transformer and Convolution for Large-Scale Image Retrieval
Minseong Kweon, Jinsun Park
ACCV (1)2
2024 CrossFormer: Cross-guided attention for multi-modal object detection
Seungik Lee, Jinsun Park
Pattern Recognit. Lett.3
2023 Deep Depth Estimation from Thermal Image
abstract
Robust and accurate geometric understanding against adverse weather conditions is one top prioritized conditions to achieve a high-level autonomy of self-driving cars. However, autonomous driving algorithms relying on the visible spectrum band are easily impacted by weather and lighting conditions. A long-wave infrared camera, also known as a thermal imaging camera, is a potential rescue to achieve high-level robustness. However, the missing necessities are the well-established large-scale dataset and public benchmark results. To this end, in this paper, we first built a large-scale Multi-Spectral Stereo (MS2) dataset, including stereo RGB, stereo NIR, stereo thermal, and stereo LiDAR data along with GNSS/IMU information. The collected dataset provides about 195K synchronized data pairs taken from city, residential, road, campus, and suburban areas in the morning, daytime, and nighttime under clear-sky, cloudy, and rainy conditions. Secondly, we conduct an exhaustive validation process of monocular and stereo depth estimation algorithms designed on visible spectrum bands to benchmark their performance in the thermal image domain. Lastly, we propose a unified depth network that effectively bridges monocular depth and stereo depth tasks from a conditional random field approach perspective. Our dataset and source code are available at https://github.com/UkcheolShin/MS2-MultiSpectralStereoDataset.
Ukcheol Shin, Jinsun Park, In-So Kweon
CVPR2
2022 Lightweight Alpha Matting Network Using Distillation-Based Channel Pruning
Donggeun Yoon, Jinsun Park, Donghyeon Cho
ACCV (3)2
2022 MC-Calib: A generic and robust calibration toolbox for multi-camera systems
François Rameau, Jinsun Park, Oleksandr Bailo, In-So Kweon
Comput. Vis. Image Underst.2
2022 Real-Time Multi-Car Localization and See-Through System
François Rameau, Oleksandr Bailo, Jinsun Park, Kyungdon Joo, In-So Kweon
Int. J. Comput. Vis.3
2022 Learning Lightweight Low-Light Enhancement Network Using Pseudo Well-Exposed Images
abstract
Recently, there has been growing attention on deep learning-based low-light image enhancement algorithms. With this interest, various synthetic low-light image datasets have been released publicly. However, real-world low-light and well-exposed image pair datasets are still lacking. In this paper, we propose a real-world low-light image dataset and a practical lightweight low-light image enhancement network. In order to construct a large-scale real-world low-light dataset, we have not only captured under-exposed images by ourselves but also collected under-exposed images from the Internet. Then, we produce pseudo well-exposed images for each low-light image. Using pairs of a real-world low-light image and a pseudo well-exposed image, we present a lightweight deep CNN model through knowledge distillation. Experimental results demonstrate the effectiveness and practicality of the proposed method on various datasets.
Seonggwan Ko, Jinsun Park, Byungjoo Chae, Donghyeon Cho
IEEE Signal Process. Lett.2
2020 Robust Reference-Based Super-Resolution With Similarity-Aware Deformable Convolution
abstract
In this paper, we propose a novel and efficient reference feature extraction module referred to as the Similarity Search and Extraction Network (SSEN) for reference-based super-resolution (RefSR) tasks. The proposed module extracts aligned relevant features from a reference image to increase the performance over single image super-resolution (SISR) methods. In contrast to conventional algorithms which utilize brute-force searches or optical flow estimations, the proposed algorithm is end-to-end trainable without any additional supervision or heavy computation, predicting the best match with a single network forward operation. Moreover, the proposed module is aware of not only the best matching position but also the relevancy of the best match. This makes our algorithm substantially robust when irrelevant reference images are given, overcoming the major cause of the performance degradation when using existing RefSR methods. Furthermore, our module can be utilized for self-similarity SR if no reference image is available. Experimental results demonstrate the superior performance of the proposed algorithm compared to previous works both quantitatively and qualitatively.
Gyumin Shim, Jinsun Park, In-So Kweon
CVPR2
2020 Non-local Spatial Propagation Network for Depth Completion
Jinsun Park, Kyungdon Joo, Chi-Kuei Liu, In-So Kweon
ECCV (13)1
2020 Propose-and-Attend Single Shot Detector
abstract
We present a simple yet effective prediction module for a one-stage detector. The main process is conducted in a coarse-to-fine manner. First, the module roughly adjusts the default boxes to well capture the extent of target objects in an image. Second, given the adjusted boxes, the module aligns the receptive field of the convolution filters accordingly, not requiring any embedding layers. Both steps build a propose-and-attend mechanism, mimicking two-stage detectors in a highly efficient manner. To verify its effectiveness, we apply the proposed module to a basic one-stage detector SSD. We empirically show that our module significantly lifts the detection accuracy with marginal parameter overhead. Our final model achieves an accuracy comparable to that of state-of-the-art detectors while using a fraction of their model parameter and computational overheads. Moreover, we found that the proposed module has two strong applications. 1) The module can be successfully integrated into a lightweight backbone, further pushing the efficiency of the one-stage detector. 2) The module also allows train-from-scratch without relying on any sophisticated base networks as previous methods do.
Ho-Deok Jang, Sanghyun Woo, Philipp Benz, Jinsun Park, In-So Kweon
WACV4
2020 Lightweight Deep CNN for Natural Image Matting via Similarity-Preserving Knowledge Distillation
abstract
Recently, alpha matting has witnessed remarkable growth by wide and deep convolutional neural networks. However, previous deep learning-based alpha matting methods require a high computational cost to be used in real environments including mobile devices. In this letter, a lightweight natural image matting network with a similarity-preserving knowledge distillation is developed. The similarity-preserving knowledge distillation makes pairwise similarities from a compact student network similar to those from a teacher network. The pairwise similarity measured on spatial, channel, and batch units enables to transfer knowledge of the teacher to the student. Based on the similarity-preserving knowledge distillation, we not only design a lighter and smaller student network than the teacher one but also achieve superior performance compared to that of the student network without the knowledge distillation. In addition, the proposed algorithm can be seamlessly applied to various deep image matting algorithms. Therefore, our algorithm is effective for mobile applications (e.g., human portrait matting), which are in growing demand. The effectiveness of the proposed algorithm is verified on two public benchmark datasets.
Donggeun Yoon, Jinsun Park, Donghyeon Cho
IEEE Signal Process. Lett.2
2019 Vehicular Multi-Camera Sensor System for Automated Visual Inspection of Electric Power Distribution Equipment
abstract
In this paper, we present a multi-camera sensor system along with its control algorithm for automated visual inspection from a moving vehicle. To accomplish this task, we propose a unique hardware configuration consisting of a frontal stereo vision system, six lateral cameras motorized to tilt, and a GPS/IMU sensor mounted on the roof of a car. From the frontal stereo system, we detect electric poles and estimate their corresponding 3D positions. Based on this 3D estimation, the tilt angles of the motorized lateral cameras are controlled in real-time to capture high resolution images of the equipment - typically installed a few meters above the road surface. In addition, inertial odometry information from the GPS/IMU module is utilized for pose estimation, object localization, and re-identification among cameras. Experimental results demonstrate the efficiency and robustness of our system for automated electric equipment maintenance, which can reduce human effort significantly.
Jinsun Park, Ukcheol Shin, Gyumin Shim, Kyungdon Joo, François Rameau, Junhyeok Kim 0004, Dong-Geol Choi, In-So Kweon
IROS1
2019 Camera Exposure Control for Robust Robot Vision with Noise-Aware Image Quality Assessment
abstract
In this paper, we propose a noise-aware exposure control algorithm for robust robot vision. Our method aims to capture best-exposed images, which can boost the performance of various computer vision and robotics tasks. For this purpose, we carefully design an image quality metric that captures complementary quality attributes and ensures light-weight computation. Specifically, our metric consists of a combination of image gradient, entropy, and noise metrics. The synergy of these measures allows the preservation of sharp edges and rich texture in the image while maintaining a low noise level. Using this novel metric, we propose a real-time and fully automatic exposure and gain control technique based on the Nelder-Mead method. To illustrate the effectiveness of our technique, a large set of experimental results demonstrates the higher qualitative and quantitative performance compared with conventional approaches.
Ukcheol Shin, Jinsun Park, Gyumin Shim, François Rameau, In-So Kweon
IROS2
2019 Depth from a Light Field Image with Learning-Based Matching Costs
abstract
One of the core applications of light field imaging is depth estimation. To acquire a depth map, existing approaches apply a single photo-consistency measure to an entire light field. However, this is not an optimal choice because of the non-uniform light field degradations produced by limitations in the hardware design. In this paper, we introduce a pipeline that automatically determines the best configuration for photo-consistency measure, which leads to the most reliable depth label from the light field. We analyzed the practical factors affecting degradation in lenslet light field cameras, and designed a learning based framework that can retrieve the best cost measure and optimal depth label. To enhance the reliability of our method, we augmented an existing light field benchmark to simulate realistic source dependent noise, aberrations, and vignetting artifacts. The augmented dataset was used for the training and validation of the proposed approach. Our method was competitive with several state-of-the-art methods for the benchmark and real-world light field datasets.
Hae-Gon Jeon, Jaesik Park, Gyeongmin Choe, Jinsun Park, Yunsu Bok, Yu-Wing Tai, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Efficient adaptive non-maximal suppression algorithms for homogeneous spatial keypoint distribution
Oleksandr Bailo, François Rameau, Kyungdon Joo, Jinsun Park, Oleksandr Bogdan, In-So Kweon
Pattern Recognit. Lett.4
2017 A Unified Approach of Multi-scale Deep and Hand-Crafted Features for Defocus Estimation
abstract
In this paper, we introduce robust and synergetic hand-crafted features and a simple but efficient deep feature from a convolutional neural network (CNN) architecture for defocus estimation. This paper systematically analyzes the effectiveness of different features, and shows how each feature can compensate for the weaknesses of other features when they are concatenated. For a full defocus map estimation, we extract image patches on strong edges sparsely, after which we use them for deep and hand-crafted feature extraction. In order to reduce the degree of patch-scale dependency, we also propose a multi-scale patch extraction strategy. A sparse defocus map is generated using a neural network classifier followed by a probability-joint bilateral filter. The final defocus map is obtained from the sparse defocus map with guidance from an edge-preserving filtered input image. Experimental results show that our algorithm is superior to state-of-the-art algorithms in terms of defocus estimation. Our work can be used for applications such as segmentation, blur magnification, all-in-focus image generation, and 3-D estimation.
Jinsun Park, Yu-Wing Tai, Donghyeon Cho, In-So Kweon
CVPR1
2017 Weakly- and Self-Supervised Learning for Content-Aware Deep Image Retargeting
abstract
This paper proposes a weakly- and self-supervised deep convolutional neural network (WSSDCNN) for content-aware image retargeting. Our network takes a source image and a target aspect ratio, and then directly outputs a retargeted image. Retargeting is performed through a shift reap, which is a pixel-wise mapping from the source to the target grid. Our method implicitly learns an attention map, which leads to r content-aware shift map for image retargeting. As a result, discriminative parts in an image are preserved, while background regions are adjusted seamlessly. In the training phase, pairs of an image and its image-level annotation are used to compute content and structure tosses. We demonstrate the effectiveness of our proposed method for a retargeting application with insightful analyses.
Donghyeon Cho, Jinsun Park, Tae-Hyun Oh, Yu-Wing Tai, In-So Kweon
ICCV2
2015 Accurate depth map estimation from a lenslet light field camera
abstract
This paper introduces an algorithm that accurately estimates depth maps using a lenslet light field camera. The proposed algorithm estimates the multi-view stereo correspondences with sub-pixel accuracy using the cost volume. The foundation for constructing accurate costs is threefold. First, the sub-aperture images are displaced using the phase shift theorem. Second, the gradient costs are adaptively aggregated using the angular coordinates of the light field. Third, the feature correspondences between the sub-aperture images are used as additional constraints. With the cost volume, the multi-label optimization propagates and corrects the depth map in the weak texture regions. Finally, the local depth map is iteratively refined through fitting the local quadratic function to estimate a non-discrete depth map. Because micro-lens images contain unexpected distortions, a method is also proposed that corrects this error. The effectiveness of the proposed algorithm is demonstrated through challenging real world examples and including comparisons with the performance of advanced depth estimation algorithms.
Hae-Gon Jeon, Jaesik Park, Gyeongmin Choe, Jinsun Park, Yunsu Bok, Yu-Wing Tai, In-So Kweon
CVPR4