VLDB 2026 Research / reviewers in the wild / expert
Changmin Kim
dblp:02/1823
· DBLP profile ↗
7ranked-venue papers
0as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An On-device Robust Sound Recognition System for Real-time Context Awareness of RobotsabstractThis paper suggests an on-device robust sound recognition system for robots in real-time. The proposed system is designed to enable the robot to detect a variety of sound events in a variety of locations, including noisy and reverberant sound environments. To use suggested system on target robots, two VGGish models are trained on sever-side and the pre-trained models infer using an audio topic from a on-device real-time buffer handling system. The buffer handling system and the training system of deep learning model are designed to get almost silmilar input audio stream with each normalization system. To get robust performance in various environments, we use log-mel feature for general environments and per-chennal energy normalization for noisy and reverberant environments. Each feature is switched and used in real time on the robot depending on the sound environment mode. Several experimental results demonstrate the robust performance of the proposed real-time robust sound recognition system on a target robot. Ju-Man Song, Changmin Kim, Jungkwan Son |
RO-MAN | 2 |
| 2022 | Hybrid CTC-Attention Network-Based End-to-End Speech Recognition System for Korean LanguageabstractIn this study, an automatic end-to-end speech recognition system based on hybrid CTC-attention network for Korean language is proposed. Deep neural network/hidden Markov model (DNN/HMM)-based speech recognition system has driven dramatic improvement in this area. However, it is difficult for non-experts to develop speech recognition for new applications. End-to-end approaches have simplified speech recognition system into a single-network architecture. These approaches can develop speech recognition system that does not require expert knowledge. In this paper, we propose hybrid CTC-attention network as end-to-end speech recognition model for Korean language. This model effectively utilizes a CTC objective function during attention model training. This approach improves the performance in terms of speech recognition accuracy as well as training speed. In most languages, end-to-end speech recognition uses characters as output labels. However, for Korean, character-based end-to-end speech recognition is not an efficient approach because Korean language has 11,172 possible numbers of characters. The number is relatively large compared to other languages. For example, English has 26 characters, and Japanese has 50 characters. To address this problem, we utilize Korean 49 graphemes as output labels. Experimental result shows 10.02% character error rate (CER) when 740 hours of Korean training data are used. Hosung Park, Changmin Kim, Hyunsoo Son, Soonshin Seo |
J. Web Eng. | 2 |
| 2022 | Convolutional Neural Networks Using Log Mel-Spectrogram Separation for Audio Event Classification with Unknown DevicesabstractAudio event classification refers to the detection and classification of non-verbal signals, such as dog and horn sounds included in audio data, by a computer. Recently, deep neural network technology has been applied to audio event classification, exhibiting higher performance when compared to existing models. Among them, a convolutional neural network (CNN)-based training method that receives audio in the form of a spectrogram, which is a two-dimensional image, has been widely used. However, audio event classification has poor performance on test data when it is recorded by a device (unknown device) different from that used to record training data (known device). This is because the frequency range emphasized is different for each device used during recording, and the shapes of the resulting spectrograms generated by known devices and those generated by unknown devices differ. In this study, to improve the performance of the event classification system, a CNN based on the log mel-spectrogram separation technique was applied to the event classification system, and the performance of unknown devices was evaluated. The system can classify 16 types of audio signals. It receives audio data at 0.4-s length, and measures the accuracy of test data generated from unknown devices with a model trained via training data generated from known devices. The experiment showed that the performance compared to the baseline exhibited a relative improvement of up to 37.33%, from 63.63% to 73.33% based on Google Pixel, and from 47.42% to 65.12% based on the LG V50. Soonshin Seo, Changmin Kim |
J. Web Eng. | 2 |
| 2019 | Shortcut Connections Based Deep Speaker Embeddings for End-to-End Speaker Verification System
Soonshin Seo, Daniel Jun Rim, Minkyu Lim, Donghyun Lee 0001, Hosung Park, Junseok Oh, Changmin Kim |
INTERSPEECH | 7 |
| 2014 | Classification of major construction materials in construction environments using ensemble classifiers
Hyojoo Son, Changmin Kim, Nahyae Hwang, Changwan Kim, Youngcheol Kang |
Adv. Eng. Informatics | 2 |
| 2006 | Goal Programming Approach to Compose the Web Service Quality of Service
Daerae Cho, Changmin Kim, Moonwon Choo, Suk-Ho Kang, Wookey Lee |
ICCSA (4) | 2 |
| 2002 | Self-maintainable Data Warehouse Views Using Differential Files
Wookey Lee, Yonghun Hwang, Suk-Ho Kang, Sanggeun Kim, Changmin Kim, Yunsun Lee |
DEXA | 5 |