VLDB 2026 Research / reviewers in the wild / expert
Byeongkeun Kang
dblp:164/6143
· DBLP profile ↗
20ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0003-2537-7720ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning from multimodal pseudo-labels for robust open-vocabulary instance and panoptic segmentation
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 3 |
| 2026 | Generative compositional zero-Shot learning using learnable primitive disparity
Byeongkeun Kang, Yeejin Lee |
Knowl. Based Syst. | 2 |
| 2025 | Generalized Class Discovery in Instance SegmentationabstractThis work addresses the task of generalized class discovery (GCD) in instance segmentation. The goal is to discover novel classes and obtain a model capable of segmenting instances of both known and novel categories, given labeled and unlabeled data. Since the real world contains numerous objects with long-tailed distributions, the instance distribution for each class is inherently imbalanced. To address the imbalanced distributions, we propose an instance-wise temperature assignment (ITA) method for contrastive learning and class-wise reliability criteria for pseudo-labels. The ITA method relaxes instance discrimination for samples belonging to head classes to enhance GCD. The reliability criteria are to avoid excluding most pseudo-labels for tail classes when training an instance segmentation network using pseudo-labels from GCD. Additionally, we propose dynamically adjusting the criteria to leverage diverse samples in the early stages while relying only on reliable pseudo-labels in the later stages. We also introduce an efficient soft attention module to encode object-specific representations for GCD. Finally, we evaluate our proposed method by conducting experiments on two settings: COCO$_{half}$ + LVIS and LVIS + Visual Genome. The experimental results demonstrate that the proposed method outperforms previous state-of-the-art methods. Cuong Manh Hoang, Yeejin Lee, Byeongkeun Kang |
AAAI | 3 |
| 2025 | Generalized Zero-Shot Learning for Point Cloud Segmentation with Evidence-Based Dynamic CalibrationabstractGeneralized zero-shot semantic segmentation of 3D point clouds aims to classify each point into both seen and unseen classes. A significant challenge with these models is their tendency to make biased predictions, often favoring the classes encountered during training. This problem is more pronounced in 3D applications, where the scale of the training data is typically smaller than in image-based tasks. To address this problem, we propose a novel method called E3DPC-GZSL, which reduces overconfident predictions towards seen classes without relying on separate classifiers for seen and unseen data. E3DPC-GZSL tackles the overconfidence problem by integrating an evidence-based uncertainty estimator into a classifier. This estimator is then used to adjust prediction probabilities using a dynamic calibrated stacking factor that accounts for pointwise prediction uncertainty. In addition, E3DPC-GZSL introduces a novel training strategy that improves uncertainty estimation by refining the semantic space. This is achieved by merging learnable parameters with text-derived features, thereby improving model optimization for unseen data. Extensive experiments demonstrate that the proposed approach achieves state-of-the-art performance on generalized zero-shot semantic segmentation datasets, including ScanNet v2 and S3DIS. Hyeonseok Kim, Byeongkeun Kang, Yeejin Lee |
AAAI | 2 |
| 2025 | Unsupervised contrastive learning using out-of-distribution data for long-tailed dataset
Cuong Manh Hoang, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 3 |
| 2025 | Content-aware preserving image generation
Giang H. Le, Anh Q. Nguyen, Byeongkeun Kang, Yeejin Lee |
Neurocomputing | 3 |
| 2025 | Completely weakly supervised class-incremental learning for semantic segmentation
David Minkwan Kim, Soeun Lee, Byeongkeun Kang |
Pattern Recognit. Lett. | 3 |
| 2024 | MSTA3D: Multi-scale Twin-attention for 3D Instance SegmentationabstractRecently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods. Duc Dang Trung Tran, Byeongkeun Kang, Yeejin Lee |
ACM Multimedia | 2 |
| 2024 | Pixel-level clustering network for unsupervised image segmentation
Cuong Manh Hoang, Byeongkeun Kang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Improving weakly-supervised object localization using adversarial erasing and pseudo label
Byeongkeun Kang, Sinhae Cha, Yeejin Lee |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Enhancing long-term person re-identification using global, local body part, and head streams
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 3 |
| 2023 | FDCNet: Feature Drift Compensation Network for Class-Incremental Weakly Supervised Object LocalizationabstractThis work addresses the task of class-incremental weakly supervised object localization (CI-WSOL). The goal is to incrementally learn object localization for novel classes using only image-level annotations while retaining the ability to localize previously learned classes. This task is important because annotating bounding boxes for every new incoming data is expensive, although object localization is crucial in various applications. To the best of our knowledge, we are the first to address this task. Thus, we first present a strong baseline method for CI-WSOL by adapting the strategies of class-incremental classifiers to mitigate catastrophic forgetting. These strategies include applying knowledge distillation, maintaining a small data set from previous tasks, and using cosine normalization. We then propose the feature drift compensation network to compensate for the effects of feature drifts on class scores and localization maps. Since updating network parameters to learn new tasks causes feature drifts, compensating for the final outputs is necessary. Finally, we evaluate our proposed method by conducting experiments on two publicly available datasets (ImageNet-100 and CUB-200). The experimental results demonstrate that the proposed method outperforms other baseline methods. Sejin Park 0002, Taehyung Lee 0003, Yeejin Lee, Byeongkeun Kang |
ACM Multimedia | 4 |
| 2023 | Multiscale Vision Transformer With Deep Clustering-Guided Refinement for Weakly Supervised Object LocalizationabstractThis work addresses the task of weakly-supervised object localization. The goal is to learn object localization using only image-level class labels, which are much easier to obtain compared to bounding box annotations. This task is important because it reduces the need for labor-intensive ground-truth annotations. However, methods for object localization trained using weak supervision often suffer from limited accuracy in localization. To address this challenge and enhance localization accuracy, we propose a multiscale object localization transformer (MOLT). It comprises multiple object localization transformers that extract patch embeddings across various scales. Moreover, we introduce a deep clustering-guided refinement method that further enhances localization accuracy by utilizing separately extracted image segments. These segments are obtained by clustering pixels using convolutional neural networks. Finally, we demonstrate the effectiveness of our proposed method by conducting experiments on the publicly available ILSVRC-2012 dataset. David Minkwan Kim, Sinhae Cha, Byeongkeun Kang |
VCIP | 3 |
| 2022 | Sampling Agnostic Feature Representation for Long-Term Person Re-IdentificationabstractPerson re-identification is a problem of identifying individuals across non-overlapping cameras. Although remarkable progress has been made in the re-identification problem, it is still a challenging problem due to appearance variations of the same person as well as other people of similar appearance. Some prior works solved the issues by separating features of positive samples from features of negative ones. However, the performances of existing models considerably depend on the characteristics and statistics of the samples used for training. Thus, we propose a novel framework named sampling independent robust feature representation network (SirNet) that learns disentangled feature embedding from randomly chosen samples. A carefully designed sampling independent maximum discrepancy loss is introduced to model samples of the same person as a cluster. As a result, the proposed framework can generate additional hard negatives/positives using the learned features, which results in better discriminability from other identities. Extensive experimental results on large-scale benchmark datasets verify that the proposed model is more effective than prior state-of-the-art models. Seongyeop Yang, Byeongkeun Kang, Yeejin Lee |
IEEE Trans. Image Process. | 2 |
| 2019 | Incremental Class Discovery for Semantic Segmentation With RGBD Sensing
Yoshikatsu Nakajima, Byeongkeun Kang, Hideo Saito 0001, Kris Makoto Kitani |
ICCV | 2 |
| 2019 | Random Forest With Learned Representations for Semantic SegmentationabstractWe present a random forest framework that learns the weights, shapes, and sparsities of feature representations for real-time semantic segmentation. Typical filters (kernels) have predetermined shapes and sparsities and learn only weights. A few feature extraction methods fix weights and learn only shapes and sparsities. These predetermined constraints restrict learning and extracting optimal features. To overcome this limitation, we propose an unconstrained representation that is able to extract optimal features by learning weights, shapes, and sparsities. We, then, present the random forest framework that learns the flexible filters using an iterative optimization algorithm and segments input images using the learned representations. We demonstrate the effectiveness of the proposed method using a hand segmentation dataset for hand-object interaction and using two semantic segmentation datasets. The results show that the proposed method achieves real-time semantic segmentation using limited computational and memory resources. Byeongkeun Kang, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2018 | Accurate and Efficient Video De-Fencing Using Convolutional Neural Networks and Temporal InformationabstractDe-fencing is to eliminate the captured fence on an image or a video, providing a clear view of the scene. It has been applied for many purposes including assisting photographers and improving the performance of computer vision algorithms such as object detection and recognition. However, the state-of-the-art de-fencing methods have limited performance caused by the difficulty of fence segmentation and also suffer from the motion of the camera or objects. To overcome these problems, we propose a novel method consisting of segmentation using convolutional neural networks and a fast/robust recovery algorithm. The segmentation algorithm using convolutional neural network achieves significant improvement in the accuracy of fence segmentation. The recovery algorithm using optical flow produces plausible de-fenced images and videos. The proposed method is experimented on both our diverse and complex dataset and publicly available datasets. The experimental results demonstrate that the proposed method achieves the state-of-the-art performance for both segmentation and content recovery. Byeongkeun Kang, Ji Dai, Truong Q. Nguyen |
ICME | 2 |
| 2018 | Depth-Adaptive Deep Neural Network for Semantic SegmentationabstractIn this paper, we present the depth-adaptive deep neural network using a depth map for semantic segmentation. Typical deep neural networks receive inputs at the predetermined locations regardless of the distance from the camera. This fixed receptive field presents a challenge to generalize the features of objects at various distances in neural networks. Specifically, the predetermined receptive fields are too small at a short distance, and vice versa. To overcome this challenge, we develop a neural network that is able to adapt the receptive field not only for each layer but also for each neuron at the spatial location. To adjust the receptive field, we propose the depth-adaptive multiscale (DaM) convolution layer consisting of the adaptive perception neuron and the in-layer multiscale neuron. The adaptive perception neuron is to adjust the receptive field at each spatial location using the corresponding depth information. The in-layer multiscale neuron is to apply the different size of the receptive field at each feature space to learn features at multiple scales. The proposed DaM convolution is applied to two fully convolutional neural networks. We demonstrate the effectiveness of the proposed neural networks on the publicly available RGB-D dataset for semantic segmentation and the novel hand segmentation dataset for hand-object interaction. The experimental results show that the proposed method outperforms the state-of-the-art methods without any additional layers or preprocessing/postprocessing. Byeongkeun Kang, Yeejin Lee, Truong Q. Nguyen |
IEEE Trans. Multim. | 1 |
| 2017 | A computational framework for driver's visual attention using a fully convolutional architectureabstractIt is a challenging and important task to perceive and interact with other traffic participants in a complex driving environment. The human vision system plays one of the crucial roles to achieve this task. Particularly, visual attention mechanisms allow a human driver to cleverly attend to the salient and relevant regions of the scene to further make necessary decisions for the safe driving. Thus, it is significant to investigate human vision systems with great potential to improve assistive, and even autonomous, vehicular technologies. In this paper, we investigate driver's gaze behavior to understand visual attention. We, first, present a Bayesian framework to model visual attention of a human driver. Further, based on the framework, we develop a fully convolutional neural network to estimate the salient region in a novel driving scene. We systematically evaluate the proposed method using on-road driving data and compare it with other state-of-the-art saliency estimation approaches. Our analyses show promising results. Ashish Tawari, Byeongkeun Kang |
Intelligent Vehicles Symposium | 2 |
| 2016 | Long-Range Motion Trajectories Extraction of Articulated Human Using Mesh EvolutionabstractThis letter presents a novel approach to extract reliable dense and long-range motion trajectories of articulated human in a video sequence. Compared with existing approaches that emphasize temporal consistency of each tracked point, we also consider the spatial structure of tracked points on the articulated human. We treat points as a set of vertices, and build a triangle mesh to join them in image space. The problem of extracting long-range motion trajectories is changed to the issue of consistency of mesh evolution over time. First, self-occlusion is detected by a novel mesh-based method and an adaptive motion estimation method is proposed to initialize mesh between successive frames. Furthermore, we propose an iterative algorithm to efficiently adjust vertices of mesh for a physically plausible deformation, which can meet the local rigidity of mesh and silhouette constraints. Finally, we compare the proposed method with the state-of-the-art methods on a set of challenging sequences. Evaluations demonstrate that our method achieves favorable performance in terms of both accuracy and integrity of extracted trajectories. Yuanyuan Wu 0001, Xiaohai He, Byeongkeun Kang, Haiying Song, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 3 |