VLDB 2026 Research / reviewers in the wild / expert
Min-Cheol Sagong
dblp:234/4281
· DBLP profile ↗
11ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0001-9191-4147ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-View Pedestrian Occupancy Prediction with a Novel Synthetic DatasetabstractWe address an advanced challenge of predicting pedestrian occupancy as an extension of multi-view pedestrian detection in urban traffic. To support this, we have created a new synthetic dataset called MVP-Occ, designed for dense pedestrian scenarios in large-scale scenes. Our dataset provides detailed representations of pedestrians using voxel structures, accompanied by rich semantic scene understanding labels, facilitating visual navigation and insights into pedestrian spatial information. Furthermore, we present a robust baseline model, termed OmniOcc, capable of predicting both the voxel occupancy state and panoptic labels for the entire scene from multi-view images. Through in-depth analysis, we identify and evaluate the key elements of our proposed model, highlighting their specific contributions and importance. Sithu Aung, Min-Cheol Sagong, Junghyun Cho |
AAAI | 2 |
| 2025 | Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor ScenesabstractWe propose a diffusion-based inverse rendering framework that decomposes a single RGB image into geometry, material, and lighting. Inverse rendering is inherently ill-posed, making it difficult to predict a single accurate solution. To address this challenge, recent generative model-based methods aim to present a range of possible solutions. However, finding a single accurate solution and generating diverse solutions can be conflicting. In this paper, we propose a channel-wise noise scheduling approach that allows a single diffusion model architecture to achieve two conflicting objectives. The resulting two diffusion models, trained with different channel-wise noise schedules, can predict a single highly accurate solution and present multiple possible solutions. The experimental results demonstrate the superiority of our two models in terms of both diversity and accuracy, which translates to enhanced performance in downstream applications such as object insertion and material editing. Junyong Choi, Min-Cheol Sagong, SeokYeong Lee, Seung-Won Jung, Ig-Jae Kim, Junghyun Cho |
CVPR | 2 |
| 2025 | VIGFace: Virtual Identity Generation for Privacy-Free Face Recognition Dataset
Min-Cheol Sagong, Gi Pyo Nam, Junghyun Cho, Ig-Jae Kim |
ICCV | 2 |
| 2024 | Conditional Convolution Projecting Latent Vectors on Condition-Specific SpaceabstractDespite rapid advancements over the past several years, the conditional generative adversarial networks (cGANs) are still far from being perfect. Although one of the major concerns of the cGANs is how to provide the conditional information to the generator, there are not only no ways considered as the optimal solution but also a lack of related research. This brief presents a novel convolution layer, called the conditional convolution (cConv) layer, which incorporates the conditional information into the generator of the generative adversarial networks (GANs). Unlike the most general framework of the cGANs using the conditional batch normalization (cBN) that transforms the normalized feature maps after convolution, the proposed method directly produces conditional features by adjusting the convolutional kernels depending on the conditions. More specifically, in each cConv layer, the weights are conditioned in a simple but effective way through filter-wise scaling and channel-wise shifting operations. In contrast to the conventional methods, the proposed method with a single generator can effectively handle condition-specific characteristics. The experimental results on CIFAR, LSUN, and ImageNet datasets show that the generator with the proposed cConv layer achieves a higher quality of conditional image generation than that with the standard convolution layer. Min-Cheol Sagong, Yoon-Jae Yeo, Yong-Goo Shin, Sung-Jea Ko |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Semantic and Instance-Aware Pixel-Adaptive Convolution for Panoptic SegmentationabstractAlthough the weight-sharing property of convolution is one of the major reasons for the success of convolution neural networks, the content-agnostic operation is insufficient for several tasks requiring content-adaptive processing, including panoptic segmentation. Inspired by several recent works on content-adaptive convolutions, we introduce the GuidedPAKA, the first content-adaptive convolution method specialized for panoptic segmentation. Specifically, GuidedPAKA learns the pixel-adaptive kernel attention consisting of the channel and spatial kernel attentions. Instead of commonly used self-attention operation, we guide the channel and spatial kernel attentions using their respective supervision signals, i.e., semantic segmentation maps and local instance affinities. Consequently, these kernel attentions extract features helpful for panoptic segmentation. Experimental results show that the proposed GuidedPAKA improves the performance of panoptic segmentation when integrated into the baseline model. Sumin Song, Min-Cheol Sagong, Seung-Won Jung, Sung-Jea Ko |
ICIP | 2 |
| 2023 | Image generation with self pixel-wise normalization
Yoon-Jae Yeo, Min-Cheol Sagong, Seung Park, Sung-Jea Ko, Yong-Goo Shin |
Appl. Intell. | 2 |
| 2022 | RORD: A Real-world Object Removal Dataset
Min-Cheol Sagong, Yoon-Jae Yeo, Seung-Won Jung, Sung-Jea Ko |
BMVC | 1 |
| 2021 | PEPSI++: Fast and Lightweight Network for Image InpaintingabstractAmong the various generative adversarial network (GAN)-based image inpainting methods, a coarse-to-fine network with a contextual attention module (CAM) has shown remarkable performance. However, due to two stacked generative networks, the coarse-to-fine network needs numerous computational resources, such as convolution operations and network parameters, which result in low speed. To address this problem, we propose a novel network architecture called parallel extended-decoder path for semantic inpainting (PEPSI) network, which aims at reducing the hardware costs and improving the inpainting performance. PEPSI consists of a single shared encoding network and parallel decoding networks called coarse and inpainting paths. The coarse path produces a preliminary inpainting result to train the encoding network for the prediction of features for the CAM. Simultaneously, the inpainting path generates higher inpainting quality using the refined features reconstructed via the CAM. In addition, we propose Diet-PEPSI that significantly reduces the network parameters while maintaining the performance. In Diet-PEPSI, to capture the global contextual information with low hardware costs, we propose novel rate-adaptive dilated convolutional layers that employ the common weights but produce dynamic features depending on the given dilation rates. Extensive experiments comparing the performance with state-of-the-art image inpainting methods demonstrate that both PEPSI and Diet-PEPSI improve the qualitative scores, i.e., the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM), as well as significantly reduce hardware costs, such as computational time and the number of network parameters. Yong-Goo Shin, Min-Cheol Sagong, Yoon-Jae Yeo, Seung-Wook Kim 0002, Sung-Jea Ko |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Simple Yet Effective Way for Improving the Performance of Depth Map Super-ResolutionabstractIn depth map super-resolution (SR), a high-resolution color image plays an important role as guidance for preventing blurry depth boundaries. However, excessive/deficient use of the color image features often causes performance degradation such as texture-copying/edge-smoothing in flat/boundary areas. To alleviate these problems, this letter presents a simple yet effective method for enhancing the performance of the SR without requiring significant modifications to the original SR network. To this end, we present a self-selective concatenation (SSC), which is a substitute for the conventional feature concatenation. In the upsampling layers of the SR network, the SSC extracts spatial and channel attention from both color and depth features such that color features can be selectively used for depth SR. Specifically, the SSC learns to use sufficient color features for rendering sharp depth boundaries, whereas their effects are reduced in smooth regions to prevent texture-copying. The proposed SSC can be included in any existing SR networks that have the encoder-decoder structure. The experimental results show that the proposed method can further improve the performances of existing SR networks in terms of the root mean squared error and peak signal-to-noise ratio. Yoon-Jae Yeo, Min-Cheol Sagong, Yong-Goo Shin, Seung-Won Jung, Sung-Jea Ko |
IEEE Signal Process. Lett. | 2 |
| 2020 | Simple Yet Effective Way for Improving the Performance of Lossy Image CompressionabstractLossy image compression methods with deep neural network (DNN) include a quantization process between encoder and decoder networks as an essential part to increase the compression rate. However, the quantization operation impedes the flow of gradient and often disturbs the optimal learning of the encoder, which results in distortion in the reconstructed images. To alleviate this problem, this paper presents a simple yet effective way that enhances the performance of lossy image compression without imposing training overhead or modifying the original network architectures. In the proposed method, we utilize an auxiliary branch called a shortcut which directly connects the encoder and decoder. Since the shortcut does not include the quantization process, it supports the optimal learning of the encoder by flowing the accurate gradient. Furthermore, to assist the decoder which should handle additional feature maps obtained via the shortcut, we also propose a residual refinement unit (RRU) following the quantizer. The experimental results show that the image compression network trained with the proposed method remarkably improves the performance in terms of peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and multi-scale structural similarity (MS-SSIM). Yoon-Jae Yeo, Yong-Goo Shin, Min-Cheol Sagong, Seung-Wook Kim 0002, Sung-Jea Ko |
IEEE Signal Process. Lett. | 3 |
| 2019 | PEPSI : Fast Image Inpainting With Parallel Decoding NetworkabstractRecently, a generative adversarial network (GAN)-based method employing the coarse-to-fine network with the contextual attention module (CAM) has shown outstanding results in image inpainting. However, this method requires numerous computational resources due to its two-stage process for feature encoding. To solve this problem, in this paper, we present a novel network structure, called PEPSI: parallel extended-decoder path for semantic inpainting. PEPSI can reduce the number of convolution operations by adopting a structure consisting of a single shared encoding network and a parallel decoding network with coarse and inpainting paths. The coarse path produces a preliminary inpainting result with which the encoding network is trained to predict features for the CAM. At the same time, the inpainting path creates a higher-quality inpainting result using refined features reconstructed by the CAM. PEPSI not only reduces the number of convolution operation almost by half as compared to the conventional coarse-to-fine networks but also exhibits superior performance to other models in terms of testing time and qualitative scores. Min-Cheol Sagong, Yong-Goo Shin, Seung-Wook Kim 0002, Seung Park, Sung-Jea Ko |
CVPR | 1 |