VLDB 2026 Research / reviewers in the wild / expert
Minseung Kim
dblp:161/4023
· DBLP profile ↗
10ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Waste Detection on Low-Light EnvironmentabstractOn waste-sorting conveyor lines, illumination often fluctuates sharply or drops to very low levels, causing conventional object detectors to fail. To achieve lighting-robust performance without extra lamps, vision rooms, or site-specific retraining, we propose an Illumination-Invariant Convolution (IIC) block that can be inserted into any backbone network. Working in log-intensity space under a Lambertian model, IIC applies learnable zero-mean cross-channel filters that mute lighting artefacts and boost material cues, and then merges the resulting maps with base features to produce lighting-robust representations. We integrate IIC into the “nano” versions of YOLOv5/8/10/11 and lightweight RT-DETR, training on roughly 200 k conveyor-belt images (29 classes) from the AI-Hub waste dataset. IIC raises YOLO mAP50 by 2.5-4.5 pp and mAP50-95 by up to 4.6 pp; even the data-hungry, transformer-based RT-DETR gains up to 1.5 pp. In low-light video tests, the IIC-augmented models successfully detected objects the original networks missed, raising recall. This demonstrates that inserting the IIC module at a network's input provides a straightforward path to illumination-robust, field-ready waste-sorting systems. Jehwan Choi, Minseung Kim, Kang-Hyun Jo |
HSI | 2 |
| 2025 | Waste Object Detection Using Bright to Dark Feature AlignmentabstractFactory environments vary significantly in lighting and camera conditions, necessitating models that are robust to illumination changes. Although prior studies have addressed this issue by constructing separate low-light datasets, such approaches face scalability challenges due to the high cost of data collection. To overcome this issue, the proposed approach leverages DARK-ISP(Low-light Image Synthesis Pipeline) from the prior work DAI-Net to construct a synthetic low-light image dataset. This approach eliminates the need to collect real dark data. Retraining a model from scratch to handle low-light conditions is not cost-effective. A more practical approach is to leverage high-performance models pretrained in bright industrial environments. While fine-tuning such models is a possible solution, it often suffers from performance degradation due to domain shift. This study adopts a teacher-student framework to perform back-bone feature-level domain alignment between a teacher model trained on well-lit images and a student model trained on low-light images. The alignment is achieved using MMD(Maximum Mean Discrepancy) loss, which effectively mitigates the domain shift problem and reduces the representational gap between the two models. The detection model is based on RT-DETRv2, a lightweight ViT-based architecture that achieves both real-time performance and high object detection accuracy. Building upon the aligned features, a knowledge distillation method based on KD-DETR was applied. This method, specifically tailored for DETR architectures, further enhanced detection performance under low-light conditions. Experiments conducted on ROBOne recyclable waste dataset show that the proposed method achieves a 6.4% higher mAP compared to DAI-Net and a 2.2% improvement over simple fine-tuning on the target domain. Minseung Kim, Jehwan Choi, Jongchae Lee, Kyubin Hwang, Kang-Hyun Jo |
HSI | 1 |
| 2025 | 3D Reconstruction Using Colorized LiDAR Point CloudsabstractThis paper proposes a 3D Gaussian Splatting(3DGS) pipeline that directly utilizes LiDAR point data, replacing the sparse point cloud created with COLMAP. Only camera poses are extracted through COLMAP, LiDAR cloud point data is mapped to the 3D space, and downsampled. Each LiDAR point is added by sampling and averaging RGB values by projecting it as a multi-view image. Proposed method initializes the Gaussian Splatting model using the color LiDAR point cloud as the initial Gaussian position and color. As a result of the KITTI dataset experiment, the proposed method showed better performance and higher detail compared to the traditional COLMAP-based initialization, and also performance improvement was achieved in PSNR/SSIM/LPIPS measurements. Jongchae Lee, Minseung Kim, Kang-Hyun Jo |
HSI | 2 |
| 2025 | Integrated DNN-Based Parameter Estimation for Multichannel Speech Enhancement
Sein Cheong, Minseung Kim, Jong Won Shin |
IEEE Signal Process. Lett. | 2 |
| 2023 | DNN-based Parameter Estimation for MVDR Beamforming and Post-filtering
Minseung Kim, Sein Cheong, Jong Won Shin |
INTERSPEECH | 1 |
| 2022 | iDeepMMSE: An improved deep learning approach to MMSE speech and noise power spectrum estimation for speech enhancement
Minseung Kim, Hyungchan Song, Sein Cheong, Jong Won Shin |
INTERSPEECH | 1 |
| 2022 | Dual Microphone Speech Enhancement Based on Statistical Modeling of Interchannel Phase DifferenceabstractThe interchannel phase difference (IPD) may be one of the most widely-used spatial cues in multichannel speech processing, and has been used in beamformers and post filters for speech enhancement. The coherence, which is also used as a feature for speech enhancement, can provide information on the reliability of the IPD for the estimation of the speech presence probability (SPP). In this paper, we propose dual microphone speech enhancement adoptinga posterioriSPP estimation based on statistical modeling of the IPD. The marginal distribution of the IPD is derived from the distribution of the relative transfer function which is parameterized with the IPD and coherence, with a single assumption that the observed discrete Fourier transform (DFT) coefficients in each frequency are distributed according to a complex bivariate Gaussian distribution. Given the direction of arrival of the desired signal, thea posterioriSPP is obtained using the IPD distributions with and without the information on the location of the interfering source, and is applied to speech enhancement. Experimental results for various types and locations of noise, signal-to-noise ratios, reverberation times, and locations of the target source showed that the proposed method outperformed previously proposed approaches utilizing IPD information. Soojoong Hwang, Minseung Kim, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Improved Speech Enhancement Considering Speech PSD UncertaintyabstractSpeech enhancement based on statistical models has been studied for several decades. Recently, the speech enhancement adopting a speech power spectral density (PSD) uncertainty model has been proposed. This approach distinguishes the true speech PSD from its estimate and considers both as random variables. It incorporates a prior distribution of speech spectra and speech PSD estimators to derive the PSD uncertainty-aware counterpart to conventional clean speech estimators, which results in performance improvement. However, the speech PSD uncertainty model has not yet been adopted for parameter estimations such asspeech presence probability, noise PSD, and speech power spectra estimations in the speech enhancement framework. In this paper, we incorporate the speech PSD uncertainty model to all the components of the statistical model-based speech enhancement framework by deriving PSD uncertainty-aware counterparts to conventional parameter estimators. Specifically, we derive thespeech presence probability (SPP) where the likelihood function for each hypothesis is based on the speech PSD uncertainty. With thisSPP, a novel SPP-based noise PSD estimator is derived. Also, we derive the minimum mean-square error (MMSE) estimator for the power spectrum of the clean speech in the current frame under speech PSD uncertainty which is exploited to refine the speech PSD estimator. Finally, the refined speech PSD estimator is incorporated into the spectral gain function based on the speech PSD uncertainty model. The proposed approach showed improved noise PSD estimation performance in terms of the averaged logarithmic error distance, and improved speech enhancement performance in terms of the noise reduction, segmental signal-to-noise ratio, perceptual evaluation of speech quality (PESQ) scores and short-time objective intelligibility in our experiments. It also exhibited comparable performance with a real-time deep learning-based speech enhancement system in terms of the PESQ scores and composite measures for the VoiceBank-DEMAND dataset. Minseung Kim, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | DeepPep: Deep proteome inference from peptide profilesabstractProtein inference, the identification of the protein set that is the origin of a given peptide profile, is a fundamental challenge in proteomics. We present DeepPep, a deep-convolutional neural network framework that predicts the protein set from a proteomics mixture, given the sequence universe of possible proteins and a target peptide profile. In its core, DeepPep quantifies the change in probabilistic score of peptide-spectrum matches in the presence or absence of a specific protein, hence selecting as candidate proteins with the largest impact to the peptide profile. Application of the method across datasets argues for its competitive predictive ability (AUC of 0.80±0.18, AUPR of 0.84±0.28) in inferring proteins without need of peptide detectability on which the most competitive methods rely. We find that the convolutional neural network architecture outperforms the traditional artificial neural network architectures without convolution layers in protein inference. We expect that similar deep learning architectures that allow learning nonlinear patterns can be further extended to problems in metagenome profiling and cell type inference. The source code of DeepPep and the benchmark datasets used in this study are available at https://deeppep.github.io/DeepPep/. Minseung Kim, Ameen Eetemadi, Ilias Tagkopoulos |
PLoS Comput. Biol. | 1 |
| 2015 | Microbial Forensics: Predicting Phenotypic Characteristics and Environmental Conditions from Large-Scale Gene Expression ProfilesabstractA tantalizing question in cellular physiology is whether the cellular state and environmental conditions can be inferred by the expression signature of an organism. To investigate this relationship, we created an extensive normalized gene expression compendium for the bacterium Escherichia coli that was further enriched with meta-information through an iterative learning procedure. We then constructed an ensemble method to predict environmental and cellular state, including strain, growth phase, medium, oxygen level, antibiotic and carbon source presence. Results show that gene expression is an excellent predictor of environmental structure, with multi-class ensemble models achieving balanced accuracy between 70.0% (±3.5%) to 98.3% (±2.3%) for the various characteristics. Interestingly, this performance can be significantly boosted when environmental and strain characteristics are simultaneously considered, as a composite classifier that captures the inter-dependencies of three characteristics (medium, phase and strain) achieved 10.6% (±1.0%) higher performance than any individual models. Contrary to expectations, only 59% of the top informative genes were also identified as differentially expressed under the respective conditions. Functional analysis of the respective genetic signatures implicates a wide spectrum of Gene Ontology terms and KEGG pathways with condition-specific information content, including iron transport, transferases, and enterobactin synthesis. Further experimental phenotypic-to-genotypic mapping that we conducted for knock-out mutants argues for the information content of top-ranked genes. This work demonstrates the degree at which genome-scale transcriptional information can be predictive of latent, heterogeneous and seemingly disparate phenotypic and environmental characteristics, with far-reaching applications. Minseung Kim, Violeta Zorraquino, Ilias Tagkopoulos |
PLoS Comput. Biol. | 1 |