ByungIn Yoo

dblp:26/1322 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
13since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 HIMap: HybrId Representation Learning for End-to-end Vectorized HD Map Construction
abstract
Vectorized High-Definition (HD) map construction requires predictions of the category and point coordinates of map elements (e.g. road boundary, lane divider, pedestrian crossing, etc.). State-of-the-art methods are mainly based on point-level representation learning for regressing accurate point coordinates. However, this pipeline has limitations in obtaining element-level information and handling element-level failures, e.g. erroneous element shape or entanglement between elements. To tackle the above issues, we propose a simple yet effective HybrId framework named HIMap to sufficiently learn and interact both point-level and element-level information. Concretely, we introduce a hybrid representation called HIQuery to represent all map elements, and propose a point-element interactor to interactively extract and encode the hybrid information of elements, e.g. point position and element shape, into the HIQuery. Additionally, we present a point-element con-sistency constraint to enhance the consistency between the point-level and element-level information. Finally, the output point-element integrated HIQuery can be directly converted into map elements' class, point coordinates, and mask. We conduct extensive experiments and consistently outperform previous methods on both nuScenes and Argo-verse2 datasets. Notably, our method achieves 77.8 mAP on the nuScenes dataset, remarkably superior to previous SOTAs by 8.3 mAP at least.
Yi Zhou 0020, Hui Zhang 0093, Jiaqian Yu, Yifan Yang 0007, Sangil Jung, Seung In Park, ByungIn Yoo
CVPR7
2024 MapDistill: Boosting Efficient Camera-Based HD Map Construction via Camera-LiDAR Fusion Model Distillation
Xiaoshuai Hao, Ruikai Li, Hui Zhang 0093, Dingzhe Li, Rong Yin 0001, Sangil Jung, Seung In Park, ByungIn Yoo, Haimei Zhao, Jing Zhang 0037
ECCV (3)8
2024 Efficient Learning on Successive Test Time Augmentation
abstract
Test time augmentation (TTA) has been a promising tool for improving the robustness against out-of-distribution data at inference time. Recent TTA methods try to learn predictive transformations which are supposed to provide the best performance gain on each test sample. However, existing methods are either restricted to predicting one single transformation for each sample or require multiple forward passes of the transformation predictor, leading to a sub-optimal solution regarding efficiency. In this paper, we propose a novel method to predict successive test time augmentations. For the first time, it only requires a single forward pass of the transformation predictor, while can output multiple desired transformations iteratively. The experimental results show that our method provides a significant and consistent improvement in model robustness against various corruptions while significantly surpassing state-of-the-arts in runtime.
Siyang Pan, Jiaqian Yu, Qiang Wang 0023, ByungIn Yoo
ICASSP7
2024 Gradtrans: Transformer-Based Gradient Guidance for Image Generation
abstract
Image generation has been attracting widespread attention in recent years along with the development of generative models. Existing works mostly focus on pursuing high-quality generated samples as a priority. In this work, we introduce a lightweight transformer-based module, called GradTrans, that provides a novel balance on the speed-performance trade-off with generative adversarial networks for image generation. GradTrans effectively leverages the instructive information in the discriminator network to guide the generator network for a higher generation quality at the inference stage without overburdening the cost. Extensive experiments are conducted for unconditional image generation task and style transfer task on diverse datasets, including CIFAR10, STL10 and Horse2Zebra, demonstrating that our proposed GradTrans can surpass different related methods with significantly superior performance, as well as being generalizable with large compatibility to different base models.
Jiaqian Yu, Siyang Pan, Sangil Jung, Wu Bi, Seung In Park, Qiang Wang 0023, ByungIn Yoo
ICIP8
2024 MBFusion: A New Multi-modal BEV Feature Fusion Method for HD Map Construction
abstract
HD map construction is a fundamental and challenging task in autonomous driving to understand the surrounding environment. Recently, Camera-LiDAR BEV feature fusion methods have attracted increasing attention in HD map construction task, which can significantly boost the benchmark. However, existing fusion methods ignore modal interaction and utilize very simple fusion strategy, which suffers from the problems of misalignment and information loss. To tackle this, we propose a novel Multi-modal BEV feature fusion method named MBFusion. Specifically, to solve the semantic misalignment problem between Camera and LiDAR features, we design Cross-modal Interaction Transform (CIT) module to make these two feature spaces interact knowledge with each other to enhance the feature representation by the cross-attention mechanism. Then, we propose a Dual Dynamic Fusion (DDF) module to automatically select valuable information from different modalities for better feature fusion. Moreover, MBFusion is simple, and can be plug-and-played into existing pipelines. We evaluate MBFusion on three architectures, including HDMapNet, VectorMapNet, and MapTR, to show its versatility and effectiveness. Compared with the state-of-the-art methods, MBFusion achieves 3.6% and 4.1% absolute improvements on mAP on the nuScenes and the Argoverse2 datasets, respectively, demonstrating the superiority of our method.
Xiaoshuai Hao, Hui Zhang 0093, Yifan Yang 0007, Yi Zhou 0020, Sangil Jung, Seung In Park, ByungIn Yoo
ICRA7
2024 Object-Centric Representation Learning for Video Scene Understanding
abstract
Depth-aware Video Panoptic Segmentation (DVPS) is a challenging task that requires predicting the semantic class and 3D depth of each pixel in a video, while also segmenting and consistently tracking objects across frames. Predominant methodologies treat this as a multi-task learning problem, tackling each constituent task independently, thus restricting their capacity to leverage interrelationships amongst tasks and requiring parameter tuning for each task. To surmount these constraints, we present Slot-IVPS, a new approach employing an object-centric model to acquire unified object representations, thereby facilitating the model's ability to simultaneously capture semantic and depth information. Specifically, we introduce a novel representation, Integrated Panoptic Slots (IPS), to capture both semantic and depth information for all panoptic objects within a video, encompassing background semantics and foreground instances. Subsequently, we propose an integrated feature generator and enhancer to extract depth-aware features, alongside the Integrated Video Panoptic Retriever (IVPR), which iteratively retrieves spatial-temporal coherent object features and encodes them into IPS. The resulting IPS can be effortlessly decoded into an array of video outputs, including depth maps, classifications, masks, and object instance IDs. We undertake comprehensive analyses across four datasets, attaining state-of-the-art performance in both Depth-aware Video Panoptic Segmentation and Video Panoptic Segmentation tasks.
Yi Zhou 0020, Hui Zhang 0093, Seung In Park, ByungIn Yoo, Xiaojuan Qi 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Object-Centric Multi-Task Learning for Human Instances
Hyeongseok Son, Sangil Jung, Solae Lee, Seongeun Kim, Seung In Park, ByungIn Yoo
BMVC6
2022 Slot-VPS: Object-centric Representation Learning for Video Panoptic Segmentation
abstract
Video Panoptic Segmentation (VPS) aims at assigning a class label to each pixel, uniquely segmenting and identifying all object instances consistently across all frames. Classic solutions usually decompose the VPS task into several subtasks and utilize multiple surrogates (e.g. boxes and masks, centers and offsets) to represent objects. However, this divide-and-conquer strategy requires complex post-processing in both spatial and temporal domains and is vulnerable to failures from surrogate tasks. In this paper, inspired by object-centric learning which learns compact and robust object representations, we present Slot- VPS, the first end-to-end framework for this task. We encode all panoptic entities in a video, including both foreground instances and background semantics, with a unified representation called panoptic slots. The coherent spatio-temporal object's information is retrieved and encoded into the panoptic slots by the proposed Video Panoptic Retriever, enabling to localize, segment, differentiate, and associate objects in a unified manner. Finally, the output panoptic slots can be directly converted into the class, mask, and object ID of panoptic objects in the video. We conduct extensive ablation studies and demonstrate the effectiveness of our approach on two benchmark datasets, Cityscapes- VP S (val and test sets) and VIPER (val set), achieving new state-of-the-art performance of 63.7, 63.3 and 56.2 VPQ, respectively.
Yi Zhou 0020, Hui Zhang 0093, Hana Lee, Shuyang Sun, Pingjun Li, Yangguang Zhu, ByungIn Yoo, Xiaojuan Qi 0001, Jae-Joon Han
CVPR7
2021 Order Regularization on Ordinal Loss for Head Pose, Age and Gaze Estimation
abstract
Ordinal loss is widely used in solving regression problems with deep learning technologies. Its basic idea is to convert regression to classification while preserving the natural order. However, the order constraint is enforced only by ordinal label implicitly, leading to the real output values not strictly in order. It causes the network to learn separable feature rather than discriminative feature, and possibly overfit on training set. In this paper, we propose order regularization on ordinal loss, which makes the outputs in order by explicitly constraining the ordinal classifiers in order. The proposed method contains two parts, i.e. similar-weights constraint, which reduces the ineffective space between classifiers, and differential-bias constraint, which enforces the decision planes in order and enhances the discrimination power of the classifiers. Experimental results show that our proposed method boosts the performance of original ordinal loss on various regression problems such as head pose, age, and gaze estimation, with significant error reduction of around 5%. Furthermore, our method outperforms the state of the art on all these tasks, with the performance gain of 14.4%, 2.2% and 6.5% on head pose, age and gaze estimation respectively.
Tianchu Guo, ByungIn Yoo, Youngjun Kwak, Jae-Joon Han
AAAI3
2021 Controllable Image Restoration for Under-Display Camera in Smartphones
abstract
Under-display camera (UDC) technology is essential for full-screen display in smartphones and is achieved by removing the concept of drilling holes on display. However, this causes inevitable image degradation in the form of spatially variant blur and noise because of the opaque display in front of the camera. To address spatially variant blur and noise in UDC images, we propose a novel controllable image restoration algorithm utilizing pixel-wise UDC-specific kernel representation and a noise estimator. The kernel representation is derived from an elaborate optical model that reflects the effect of both normal and oblique light incidence. Also, noise-adaptive learning is introduced to control noise levels, which can be utilized to provide optimal results depending on the user preferences. The experiments showed that the proposed method achieved superior quantitative performance as well as higher perceptual quality on both a real-world dataset and a monitor-based aligned dataset compared to conventional image restoration algorithms.
Kinam Kwon, Eunhee Kang, Su-Jin Lee, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han
CVPR6
2021 RaScaNet: Learning Tiny Models by Raster-Scanning Images
abstract
Deploying deep convolutional neural networks on ultra-low power systems is challenging due to the extremely limited resources. Especially, the memory becomes a bottleneck as the systems put a hard limit on the size of on-chip memory. Because peak memory explosion in the lower layers is critical even in tiny models, the size of an input image should be reduced with sacrifice in accuracy. To overcome this drawback, we propose a novel Raster-Scanning Network, named RaScaNet, inspired by raster-scanning in image sensors. RaScaNet reads only a few rows of pixels at a time using a convolutional neural network and then sequentially learns the representation of the whole image using a recurrent neural network. The proposed method operates on an ultra-low power system without input size reduction; it requires 15.9–24.3× smaller peak memory and 5.3–12.9× smaller weight memory than the state-of-the-art tiny models. Moreover, RaScaNet fully exploits on-chip SRAM and cache memory of the system as the sum of the peak memory and the weight memory does not exceed 60 KB, improving the power efficiency of the system. In our experiments, we demonstrate the binary classification performance of RaScaNet on Visual Wake Words and Pascal VOC datasets.
Jaehyoung Yoo, Changyong Son, Sangil Jung, ByungIn Yoo, Changkyu Choi, Jae-Joon Han, Bohyung Han
CVPR5
2021 Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination
abstract
We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the chromaticity of each illuminant in the scene and the mixture ratio of illuminants per pixel. Images in our dataset are mostly captured with illuminants existing in the scene, and the ground truth illumination is computed by taking the difference between the images with different illumination combination. Therefore, our dataset captures natural composition in the real-world setting with wide field-of-view, providing more extensive dataset compared to existing datasets for multi-illumination white balance. As conventional single illuminant white balance algorithms cannot be directly applied, we also apply per-pixel DNN-based white balance algorithm and show its effectiveness against using patch-wise white balancing. We validate the benefits of our dataset through extensive analysis including a user-study, and expect the dataset to make meaningful contribution for future work in white balancing.
Dongyoung Kim, Jinwoo Kim 0007, Seonghyeon Nam, Yeonkyung Lee, Nahyup Kang, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han, Seon Joo Kim
ICCV8
2021 Learning Generalized Intersection Over Union for Dense Pixelwise Prediction
abstract
Intersection over union (IoU) score, also named Jaccard Index, is one of the most fundamental evaluation methods in machine learning. The original IoU computation cannot provide non-zero gradients and thus cannot be directly optimized by nowadays deep learning methods. Several recent works generalized IoU for bounding box regression, but they are not straightforward to adapt for pixelwise prediction. In particular, the original IoU fails to provide effective gradients for the non-overlapping and location-deviation cases, which results in performance plateau. In this paper, we propose PixIoU, a generalized IoU for pixelwise prediction that is sensitive to the distance for non-overlapping cases and the locations in prediction. We provide proofs that PixIoU holds many nice properties as the original IoU. To optimize the PixIoU, we also propose a loss function that is proved to be submodular, hence we can apply the Lovász functions, the efficient surrogates for submodular functions for learning this loss. Experimental results show consistent performance improvements by learning PixIoU over the original IoU for several different pixelwise prediction tasks on Pascal VOC, VOT-2020 and Cityscapes.
Jiaqian Yu, Jingtao Xu, Qiang Wang 0023, ByungIn Yoo, Jae-Joon Han
ICML6
2018 Residual Encoder Decoder Network and Adaptive Prior for Face Parsing
abstract
Face Parsing assigns every pixel in a facial image with a semantic label, which could be applied in various applications including face recognition, facial beautification, affective computing and animation. While lots of progress have been made in this field, current state-of-the-art methods still fail to extract real effective feature and restore accurate score map, especially for those facial parts which have large variations of deformation and fairly similar appearance, e.g. mouth, eyes and thin eyebrows. In this paper, we propose a novel pixel-wise face parsing method called Residual Encoder Decoder Network (RED-Net), which combines a feature-rich encoder-decoder framework with adaptive prior mechanism. Our encoder-decoder framework extracts feature with ResNet and decodes the feature by elaborately fusing the residual architectures in to deconvolution. This framework learns more effective feature comparing to that learnt by decoding with interpolation or classic deconvolution operations. To overcome the appearance ambiguity between facial parts, an adaptive prior mechanism is proposed in term of the decoder prediction confidence, allowing refining the final result. The experimental results on two public datasets demonstrate that our method outperforms the state-of-the-arts significantly, achieving improvements of F-measure from 0.854 to 0.905 on Helen dataset, and pixel accuracy from 95.12% to 97.59% on the LFW dataset. In particular, convincing qualitative examples show that our method parses eye, eyebrow, and lip regins more accurately.
Tianchu Guo, Youngsung Kim, Deheng Qian, ByungIn Yoo, Jingtao Xu, Dongqing Zou, Jae-Joon Han, Changkyu Choi
AAAI5
2018 Deep Facial Age Estimation Using Conditional Multitask Learning With Weak Label Expansion
abstract
Accurate age estimation from a facial image is quite challenging, since physical age and apparent age can be quite different, and this difference is dependent on gender, ethnicity, and many other factors. Multitask deep learning is one of the approach to improve age estimation by employing auxiliary tasks, such as gender recognition, that are related to the primary task. However, in traditional multitask learning for age estimation, the relationship between the primary and auxiliary tasks is difficult to describe; how the auxiliary tasks enhance the model for the primary objective is ambiguous. In this letter, we propose a conditional multitask learning method that architecturally factorizes an age variable into gender-conditioned age probabilities in a deep neural network. The lack of accurate training labels with discrete age values is another critical limitation to training age estimation models. Therefore, we propose a label expansion method that increases the number of accurate labels from weakly supervised categorical labels. To verify the generality of the proposed method, we perform intensive experiments on the publicly available MORPH-II and FG-NET datasets. The proposed methods outperform state-of-the art methods in both age estimation and gender recognition accuracy. These performance gains are verified on well-known deep network architectures-VGG-16, CASIA-WebFace, and Alexnet-to confirm the proposed methods generality.
ByungIn Yoo, Youngjun Kwak, Youngsung Kim, Changkyu Choi, Junmo Kim 0002
IEEE Signal Process. Lett.1
2015 Rotating your face using multi-task deep neural network
abstract
Face recognition under viewpoint and illumination changes is a difficult problem, so many researchers have tried to solve this problem by producing the pose- and illumination- invariant feature. Zhu et al. [26] changed all arbitrary pose and illumination images to the frontal view image to use for the invariant feature. In this scheme, preserving identity while rotating pose image is a crucial issue. This paper proposes a new deep architecture based on a novel type of multitask learning, which can achieve superior performance in rotating to a target-pose face image from an arbitrary pose and illumination image while preserving identity. The target pose can be controlled by the user's intention. This novel type of multi-task model significantly improves identity preservation over the single task model. By using all the synthesized controlled pose images, called Controlled Pose Image (CPI), for the pose-illumination-invariant feature and voting among the multiple face recognition results, we clearly outperform the state-of-the-art algorithms by more than 4~6% on the MultiPIE dataset.
Junho Yim, Heechul Jung, ByungIn Yoo, Changkyu Choi, Du-Sik Park, Junmo Kim 0002
CVPR3
2014 HDO: A novel local image descriptor
abstract
This paper presents a simple, yet powerful local image descriptor, called the histograms of dominant orientations (HDO). The HDO consists of two components, namely the dominant orientation and its coherence, which represents how intensively gradients in the local region are distributed along the dominant orientation. For a given image patch, we incorporate these two components into a 1-D histogram and define it as our HDO descriptor. Compared to previous approaches suffering from the presence of clutters and significant distortions, our HDO descriptor has a great ability to preserve the underlying image structure, and it can thus be successfully applied to various applications (e.g., object detection). The proposed method has been extensively tested on several challenging data sets and results show that our HDO descriptor is effective for object detection in images.
ByungIn Yoo, Jae-Joon Han
ICIP2
2014 Randomized decision bush: Combining global shape parameters and local scalable descriptors for human body parts recognition
abstract
This paper presents a novel method which combines global shape parameters and scalable local descriptors for accurate body parts recognition from a single depth image in real-time. Human poses are of extremely large variation in aspects of visual shapes, because human can take poses from daily activities to gymnastic actions. In order to cover wide-range of the human poses, the proposed algorithm employs a unified structure which combines pose clustering and body parts classification. We name the proposed method Randomized Decision Bush (RDB). Specifically, global shape parameters which can discriminate coarse level shapes are utilized for pose clustering while scalable local shape descriptors are employed for accurate classification. RDB splits the various human poses into multiple clusters which contain similar shapes of the poses. As a result, it provides robust clustering which enables fine level classification within the cluster. The experimental results show improvements on recognizing body parts due to the pose clustering and classification with scalable local descriptors. Additionally, we significantly reduce the complexity of training a large number of human shapes.
ByungIn Yoo, Jae-Joon Han, Changkyu Choi, Du-Sik Park, Junmo Kim 0002
ICIP1
2008 The seamless browser: enhancing the speed of web browsing by zooming and preview thumbnails
abstract
In this paper, we present a new web browsing system, Seamless Browser, for fast link traversal on a large screen like TV In navigating web, users mainly suffer from cognitive overhead of determining whether or not to follow links. This overhead can be reduced by providing preview information of the destination of links, and also by providing semantic cues on the nearest location in relation to the anchor. In order to reduce disorientation and annoyance from the preview information, we propose that users will focus on the small area nearside around a pointer, and a small number of hyperlink previews in that focused area will appear differently depending on the distances between the pointer and the hyperlinks: the nearer the distance is, the richer the content of the information scent is. We also propose that users can navigate the link paths by controlling the pointer and the zooming interface, so that users may go backward and forward seamlessly along several possible link paths. We found that combining the pointer and a zoom significantly improved performance for navigational tasks.
ByungIn Yoo, JongHo Lea, YeunBae Kim
WWW1