Sangil Jung

dblp:216/3570 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map Construction
abstract
Vectorized High-Definition (HD) map construction requires predictions of the category and point coordinates of map elements (e.g. road boundary, lane divider, pedestrian crossing, etc.). State-of-the-art methods are mainly based on point-level representation learning for regressing accurate point coordinates. However, this pipeline has limitations in obtaining element-level information and handling element-level failures, e.g. erroneous element shape or entanglement between elements. To tackle the above issues, we propose a simple yet effective HybrId framework named HIMap to sufficiently learn and interact both point-level and element-level information. Concretely, we introduce a hybrid representation called HIQuery to represent all map elements, and propose a point-element interactor to interactively extract and encode the hybrid information of elements, e.g. point position and element shape, into the HIQuery. Additionally, we present a point-element con-sistency constraint to enhance the consistency between the point-level and element-level information. Finally, the output point-element integrated HIQuery can be directly converted into map elements' class, point coordinates, and mask. We conduct extensive experiments and consistently outperform previous methods on both nuScenes and Argo-verse2 datasets. Notably, our method achieves 77.8 mAP on the nuScenes dataset, remarkably superior to previous SOTAs by 8.3 mAP at least.
Yi Zhou 0020, Hui Zhang 0093, Jiaqian Yu, Yifan Yang 0007, Sangil Jung, Seung In Park, ByungIn Yoo
CVPR5
2024 MapDistill: Boosting Efficient Camera-Based HD Map Construction via Camera-LiDAR Fusion Model Distillation
Xiaoshuai Hao, Ruikai Li, Hui Zhang 0093, Dingzhe Li, Rong Yin 0001, Sangil Jung, Seung In Park, ByungIn Yoo, Haimei Zhao, Jing Zhang 0037
ECCV (3)6
2024 Gradtrans: Transformer-Based Gradient Guidance for Image Generation
abstract
Image generation has been attracting widespread attention in recent years along with the development of generative models. Existing works mostly focus on pursuing high-quality generated samples as a priority. In this work, we introduce a lightweight transformer-based module, called GradTrans, that provides a novel balance on the speed-performance trade-off with generative adversarial networks for image generation. GradTrans effectively leverages the instructive information in the discriminator network to guide the generator network for a higher generation quality at the inference stage without overburdening the cost. Extensive experiments are conducted for unconditional image generation task and style transfer task on diverse datasets, including CIFAR10, STL10 and Horse2Zebra, demonstrating that our proposed GradTrans can surpass different related methods with significantly superior performance, as well as being generalizable with large compatibility to different base models.
Jiaqian Yu, Siyang Pan, Sangil Jung, Wu Bi, Seung In Park, Qiang Wang 0023, ByungIn Yoo
ICIP4
2024 MBFusion: A New Multi-modal BEV Feature Fusion Method for HD Map Construction
abstract
HD map construction is a fundamental and challenging task in autonomous driving to understand the surrounding environment. Recently, Camera-LiDAR BEV feature fusion methods have attracted increasing attention in HD map construction task, which can significantly boost the benchmark. However, existing fusion methods ignore modal interaction and utilize very simple fusion strategy, which suffers from the problems of misalignment and information loss. To tackle this, we propose a novel Multi-modal BEV feature fusion method named MBFusion. Specifically, to solve the semantic misalignment problem between Camera and LiDAR features, we design Cross-modal Interaction Transform (CIT) module to make these two feature spaces interact knowledge with each other to enhance the feature representation by the cross-attention mechanism. Then, we propose a Dual Dynamic Fusion (DDF) module to automatically select valuable information from different modalities for better feature fusion. Moreover, MBFusion is simple, and can be plug-and-played into existing pipelines. We evaluate MBFusion on three architectures, including HDMapNet, VectorMapNet, and MapTR, to show its versatility and effectiveness. Compared with the state-of-the-art methods, MBFusion achieves 3.6% and 4.1% absolute improvements on mAP on the nuScenes and the Argoverse2 datasets, respectively, demonstrating the superiority of our method.
Xiaoshuai Hao, Hui Zhang 0093, Yifan Yang 0007, Yi Zhou 0020, Sangil Jung, Seung In Park, ByungIn Yoo
ICRA5
2023 Object-Centric Multi-Task Learning for Human Instances
Hyeongseok Son, Sangil Jung, Solae Lee, Seongeun Kim, Seung In Park, ByungIn Yoo
BMVC2
2021 RaScaNet: Learning Tiny Models by Raster-Scanning Images
abstract
Deploying deep convolutional neural networks on ultra-low power systems is challenging due to the extremely limited resources. Especially, the memory becomes a bottleneck as the systems put a hard limit on the size of on-chip memory. Because peak memory explosion in the lower layers is critical even in tiny models, the size of an input image should be reduced with sacrifice in accuracy. To overcome this drawback, we propose a novel Raster-Scanning Network, named RaScaNet, inspired by raster-scanning in image sensors. RaScaNet reads only a few rows of pixels at a time using a convolutional neural network and then sequentially learns the representation of the whole image using a recurrent neural network. The proposed method operates on an ultra-low power system without input size reduction; it requires 15.9–24.3× smaller peak memory and 5.3–12.9× smaller weight memory than the state-of-the-art tiny models. Moreover, RaScaNet fully exploits on-chip SRAM and cache memory of the system as the sum of the peak memory and the weight memory does not exceed 60 KB, improving the power efficiency of the system. In our experiments, we demonstrate the binary classification performance of RaScaNet on Visual Wake Words and Pascal VOC datasets.
Jaehyoung Yoo, Changyong Son, Sangil Jung, ByungIn Yoo, Changkyu Choi, Jae-Joon Han, Bohyung Han
CVPR4
2019 Learning to Quantize Deep Networks by Optimizing Quantization Intervals With Task Loss
abstract
Reducing bit-widths of activations and weights of deep networks makes it efficient to compute and store them in memory, which is crucial in their deployments to resource-limited devices, such as mobile phones. However, decreasing bit-widths with quantization generally yields drastically degraded accuracy. To tackle this problem, we propose to learn to quantize activations and weights via a trainable quantizer that transforms and discretizes them. Specifically, we parameterize the quantization intervals and obtain their optimal values by directly minimizing the task loss of the network. This quantization-interval-learning (QIL) allows the quantized networks to maintain the accuracy of the full-precision (32-bit) networks with bit-width as low as 4-bit and minimize the accuracy degeneration with further bit-width reduction (i.e., 3 and 2-bit). Moreover, our quantizer can be trained on a heterogeneous dataset, and thus can be used to quantize pretrained networks without access to their training data. We demonstrate the effectiveness of our trainable quantizer on ImageNet dataset with various network architectures such as ResNet-18, -34 and AlexNet, on which it outperforms existing methods to achieve the state-of-the-art accuracy.
Sangil Jung, Changyong Son, Seohyung Lee, JinWoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, Changkyu Choi
CVPR1