VLDB 2026 Research / reviewers in the wild / expert
Yeejin Lee
dblp:84/208
· DBLP profile ↗
21ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-3439-5042ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning from multimodal pseudo-labels for robust open-vocabulary instance and panoptic segmentation
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 2 |
| 2026 | Generative compositional zero-Shot learning using learnable primitive disparity
Byeongkeun Kang, Yeejin Lee |
Knowl. Based Syst. | 3 |
| 2025 | Generalized Class Discovery in Instance SegmentationabstractThis work addresses the task of generalized class discovery (GCD) in instance segmentation. The goal is to discover novel classes and obtain a model capable of segmenting instances of both known and novel categories, given labeled and unlabeled data. Since the real world contains numerous objects with long-tailed distributions, the instance distribution for each class is inherently imbalanced. To address the imbalanced distributions, we propose an instance-wise temperature assignment (ITA) method for contrastive learning and class-wise reliability criteria for pseudo-labels. The ITA method relaxes instance discrimination for samples belonging to head classes to enhance GCD. The reliability criteria are to avoid excluding most pseudo-labels for tail classes when training an instance segmentation network using pseudo-labels from GCD. Additionally, we propose dynamically adjusting the criteria to leverage diverse samples in the early stages while relying only on reliable pseudo-labels in the later stages. We also introduce an efficient soft attention module to encode object-specific representations for GCD. Finally, we evaluate our proposed method by conducting experiments on two settings: COCO$_{half}$ + LVIS and LVIS + Visual Genome. The experimental results demonstrate that the proposed method outperforms previous state-of-the-art methods. Cuong Manh Hoang, Yeejin Lee, Byeongkeun Kang |
AAAI | 2 |
| 2025 | Generalized Zero-Shot Learning for Point Cloud Segmentation with Evidence-Based Dynamic CalibrationabstractGeneralized zero-shot semantic segmentation of 3D point clouds aims to classify each point into both seen and unseen classes. A significant challenge with these models is their tendency to make biased predictions, often favoring the classes encountered during training. This problem is more pronounced in 3D applications, where the scale of the training data is typically smaller than in image-based tasks. To address this problem, we propose a novel method called E3DPC-GZSL, which reduces overconfident predictions towards seen classes without relying on separate classifiers for seen and unseen data. E3DPC-GZSL tackles the overconfidence problem by integrating an evidence-based uncertainty estimator into a classifier. This estimator is then used to adjust prediction probabilities using a dynamic calibrated stacking factor that accounts for pointwise prediction uncertainty. In addition, E3DPC-GZSL introduces a novel training strategy that improves uncertainty estimation by refining the semantic space. This is achieved by merging learnable parameters with text-derived features, thereby improving model optimization for unseen data. Extensive experiments demonstrate that the proposed approach achieves state-of-the-art performance on generalized zero-shot semantic segmentation datasets, including ScanNet v2 and S3DIS. Hyeonseok Kim, Byeongkeun Kang, Yeejin Lee |
AAAI | 3 |
| 2025 | Unsupervised contrastive learning using out-of-distribution data for long-tailed dataset
Cuong Manh Hoang, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 2 |
| 2025 | Content-aware preserving image generation
Giang H. Le, Anh Q. Nguyen, Byeongkeun Kang, Yeejin Lee |
Neurocomputing | 4 |
| 2024 | MSTA3D: Multi-scale Twin-attention for 3D Instance SegmentationabstractRecently, transformer-based techniques incorporating superpoints have become prevalent in 3D instance segmentation. However, they often encounter an over-segmentation problem, especially noticeable with large objects. Additionally, unreliable mask predictions stemming from superpoint mask prediction further compound this issue. To address these challenges, we propose a novel framework called MSTA3D. It leverages multi-scale feature representation and introduces a twin-attention mechanism to effectively capture them. Furthermore, MSTA3D integrates a box query with a box regularizer, offering a complementary spatial constraint alongside semantic queries. Experimental evaluations on ScanNetV2, ScanNet200 and S3DIS datasets demonstrate that our approach surpasses state-of-the-art 3D instance segmentation methods. Duc Dang Trung Tran, Byeongkeun Kang, Yeejin Lee |
ACM Multimedia | 3 |
| 2024 | Improving weakly-supervised object localization using adversarial erasing and pseudo label
Byeongkeun Kang, Sinhae Cha, Yeejin Lee |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Enhancing long-term person re-identification using global, local body part, and head streams
Duy Tran Thanh, Yeejin Lee, Byeongkeun Kang |
Neurocomputing | 2 |
| 2023 | FDCNet: Feature Drift Compensation Network for Class-Incremental Weakly Supervised Object LocalizationabstractThis work addresses the task of class-incremental weakly supervised object localization (CI-WSOL). The goal is to incrementally learn object localization for novel classes using only image-level annotations while retaining the ability to localize previously learned classes. This task is important because annotating bounding boxes for every new incoming data is expensive, although object localization is crucial in various applications. To the best of our knowledge, we are the first to address this task. Thus, we first present a strong baseline method for CI-WSOL by adapting the strategies of class-incremental classifiers to mitigate catastrophic forgetting. These strategies include applying knowledge distillation, maintaining a small data set from previous tasks, and using cosine normalization. We then propose the feature drift compensation network to compensate for the effects of feature drifts on class scores and localization maps. Since updating network parameters to learn new tasks causes feature drifts, compensating for the final outputs is necessary. Finally, we evaluate our proposed method by conducting experiments on two publicly available datasets (ImageNet-100 and CUB-200). The experimental results demonstrate that the proposed method outperforms other baseline methods. Sejin Park 0002, Taehyung Lee 0003, Yeejin Lee, Byeongkeun Kang |
ACM Multimedia | 3 |
| 2022 | Dense depth estimation from multiple 360-degree images using virtual depth
Seongyeop Yang, Kunhee Kim, Yeejin Lee |
Appl. Intell. | 3 |
| 2022 | Lossless White Balance for Improved Lossless CFA Image and Video CompressionabstractColor filter array is a spatial multiplexing of pixel-sized filters fabricated over pixel sensors in most color image sensors. The state-of-the-art lossless coding techniques of raw sensor data captured by such sensors leverage spatial or cross-color correlation using lifting schemes. In this paper, we propose a lifting-based lossless white balance algorithm. When applied to the raw sensor data, the spatial bandwidth of the implied chrominance signals decreases. We propose to use this white balance as a pre-processing step to lossless CFA subsampled image/video compression, improving the overall coding efficiency of the raw sensor data. Yeejin Lee, Keigo Hirakawa |
IEEE Trans. Image Process. | 1 |
| 2022 | Sampling Agnostic Feature Representation for Long-Term Person Re-IdentificationabstractPerson re-identification is a problem of identifying individuals across non-overlapping cameras. Although remarkable progress has been made in the re-identification problem, it is still a challenging problem due to appearance variations of the same person as well as other people of similar appearance. Some prior works solved the issues by separating features of positive samples from features of negative ones. However, the performances of existing models considerably depend on the characteristics and statistics of the samples used for training. Thus, we propose a novel framework named sampling independent robust feature representation network (SirNet) that learns disentangled feature embedding from randomly chosen samples. A carefully designed sampling independent maximum discrepancy loss is introduced to model samples of the same person as a cluster. As a result, the proposed framework can generate additional hard negatives/positives using the learned features, which results in better discriminability from other identities. Extensive experimental results on large-scale benchmark datasets verify that the proposed model is more effective than prior state-of-the-art models. Seongyeop Yang, Byeongkeun Kang, Yeejin Lee |
IEEE Trans. Image Process. | 3 |
| 2020 | Shift-And-Decorrelate Lifting: CAMRA for Lossless Intra Frame CFA Video CompressionabstractIn this letter, we improve lossless intra-frame compression of color filter array (CFA) video based on Camera-Aware Multi-Resolution Analysis (CAMRA). CAMRA-based compression leverages the correlation between LH and HL subbands of wavelet-transformed CFA video frames. The key contribution of this letter is an analysis showing that the chrominance components shared by LH and HL subbands of LeGall 5/3 wavelet transform are misaligned, negatively impacting the coding efficiency. To decorrelate the subbands more effectively with minimal dynamic range growth, we developed a new lifting scheme that corrects for this misalignment. We validated our theoretical analysis and the performance of the proposed compression scheme using videos of natural scenes captured in a raw format. The experimental results verify that our proposed transform improves the coding efficiency of CFA intra-frame lossless compression. Yeejin Lee, Keigo Hirakawa |
IEEE Signal Process. Lett. | 1 |
| 2018 | Camera-Aware Multi-Resolution Analysis for Raw Image Sensor Data Compressionabstractnalysis, or CAMRA. Specifically, by CAMRA we refer to modifications that we make to wavelet transform of CFA sampled images in order to achieve a very high degree of decorrelation at the finest scale wavelet coefficients; and a series of color processing steps applied to the coarse scale wavelet coefficients, aimed at limiting the propagation of lossy compression errors through the subsequent camera processing pipeline. We validated our theoretical analysis and the performance of the proposed compression schemes using the images of natural scenes captured in a raw format. The experimental results verify that our proposed methods improve coding efficiency relative to the standard and the state-of-the-art compression schemes for CFA sampled images. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2018 | Depth-Adaptive Deep Neural Network for Semantic SegmentationabstractIn this paper, we present the depth-adaptive deep neural network using a depth map for semantic segmentation. Typical deep neural networks receive inputs at the predetermined locations regardless of the distance from the camera. This fixed receptive field presents a challenge to generalize the features of objects at various distances in neural networks. Specifically, the predetermined receptive fields are too small at a short distance, and vice versa. To overcome this challenge, we develop a neural network that is able to adapt the receptive field not only for each layer but also for each neuron at the spatial location. To adjust the receptive field, we propose the depth-adaptive multiscale (DaM) convolution layer consisting of the adaptive perception neuron and the in-layer multiscale neuron. The adaptive perception neuron is to adjust the receptive field at each spatial location using the corresponding depth information. The in-layer multiscale neuron is to apply the different size of the receptive field at each feature space to learn features at multiple scales. The proposed DaM convolution is applied to two fully convolutional neural networks. We demonstrate the effectiveness of the proposed neural networks on the publicly available RGB-D dataset for semantic segmentation and the novel hand segmentation dataset for hand-object interaction. The experimental results show that the proposed method outperforms the state-of-the-art methods without any additional layers or preprocessing/postprocessing. Byeongkeun Kang, Yeejin Lee, Truong Q. Nguyen |
IEEE Trans. Multim. | 2 |
| 2017 | Lossless compression of CFA sampled image using decorrelated Mallat wavelet packet decompositionabstractThis paper presents a rigorous analysis of wavelet transform on color filter array (CFA) sampled images. The presented analysis suggests that the wavelet coefficients of HL and LH subbands are highly correlated. Hence, we propose a novel lossless compression scheme for CFA sampled images using the decorrelated Mallat wavelet packet decomposition. We validated our theoretical analysis and the performance of the proposed compression scheme using images of natural scenes captured in a raw format. The experimental results verify that our proposed method improves coding efficiency relative to the standard and the state-of-the-art lossless compression schemes CFA sampled images. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
ICIP | 1 |
| 2017 | Joint Defogging and DemosaickingabstractImage defogging is a technique used extensively for enhancing visual quality of images in bad weather conditions. Even though defogging algorithms have been well studied, defogging performance is degraded by demosaicking artifacts and sensor noise amplification in distant scenes. In order to improve the visual quality of restored images, we propose a novel approach to perform defogging and demosaicking simultaneously. We conclude that better defogging performance with fewer artifacts can be achieved when a defogging algorithm is combined with a demosaicking algorithm simultaneously. We also demonstrate that the proposed joint algorithm has the benefit of suppressing noise amplification in distant scenes. In addition, we validate our theoretical analysis and observations for both synthesized data sets with ground truth fog-free images and natural scene data sets captured in a raw format. Yeejin Lee, Keigo Hirakawa, Truong Q. Nguyen |
IEEE Trans. Image Process. | 1 |
| 2014 | Stereo image defoggingabstractThis paper presents a new approach to estimate fog-free images from stereo foggy images. We investigate a new way to estimate transmission by computing the scattering coefficient and depth information of a scene. However, most existing visibility restoration algorithms estimate transmission independently on scattering coefficient and object distance. In the proposed method, the natural color of a foggy image is recovered using depth information from a stereo image pair even though prior knowledge or multiple images taken at different times are not required. Furthermore, we explore a new way to measure the scattering coefficient by using a stereo image pair from an image processing perspective. Experimental results verify that the proposed method outperforms the conventional defogging methods. Yeejin Lee, Kristofor B. Gibson, Zucheul Lee, Truong Q. Nguyen |
ICIP | 1 |
| 2014 | Depth-Assisted Frame Rate Up-Conversion for Stereoscopic VideoabstractIn this letter, we propose a depth-assisted frame rate up-conversion (DA-FRUC) scheme which finds applications in 3D video processing. By considering the depth cue in video plus depth representation, we categorize the blocks of the interpolated frame as depth-continuous and depth-discontinuous groups. The motion vector (MV) outliers of the depth continuous blocks are then detected and corrected by layer-constrained MV refinement method. Moreover, a depth-based adaptive interpolation and block segmentation method is proposed to deal with the disocclusion and occlusion at the boundary of the foreground area. Experimental results show that the proposed method achieves better subjective and objective performance for the interpolated color frames. Additionally, the interpolated depth frames obtained by the proposed method is more accurate, which benefit the video quality in view synthesis. Jiande Sun 0001, Yeejin Lee, Truong Q. Nguyen |
IEEE Signal Process. Lett. | 4 |
| 2007 | Improving generalization capability of neural networks based on simulated annealingabstractThis paper presents a single-objective and a multiobjective stochastic optimization algorithms for global training of neural networks based on simulated annealing. The algorithms overcome the limitation of local optimization by the conventional gradient-based training methods and perform global optimization of the weights of the neural networks. Especially, the multiobjective training algorithm is designed to enhance generalization capability of the trained networks by minimizing the training error and the dynamic range of the network weights simulataneously. For fast convergence and good solution quality of the algorithms, we suggest the hybrid simulated annealing algorithm with the gradient-based local optimization method. Experimental results show that the performance of the trained networks by the proposed methods is better than that by the gradient-based local training algorithm and, moreover, the generalization capability of the networks is significantly improved by preventing overfitting phenomena. Yeejin Lee, Jong-Seok Lee, Sun-Young Lee, Cheol Hoon Park |
IEEE Congress on Evolutionary Computation | 1 |