EDBT 2026 Demo / reviewers in the wild / expert
Yong-Hwa Park
dblp:238/4022
· DBLP profile ↗
13ranked-venue papers
0as first author
13since 2021 · last 2025
0000-0003-3519-1321ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Transfer Learning Based Motor Fault Diagnosis Using Motor Current Signals Robust to Speed, Load, and Capacity VariationsabstractThis study presents a transfer learning-based approach for fault diagnosis of AC motors. Specifically, it addresses the diagnosis of motor faults under conditions where motor capacity changes, random speed variations (5%–15%), and load fluctuations exist, using only current signals. Various motor faults were experimentally analyzed using a laboratory testbed. Transfer learning was applied to extract characteristic features that are robust to capacity, speed, and load variations, ensuring that these features are unaffected by fluctuations in the loss function. The results demonstrate that transfer learning can effectively diagnose faults even in changing environments, suggesting its potential application in motor condition monitoring in real industrial settings. Won-Ho Jung, Chanseung Yang, Jaewan Kim, Yong-Hwa Park |
IECON | 5 |
| 2024 | Acoustic Signal Based Ball Bearing Fault Diagnosis Using Adaptive Wavelet DenoisingabstractThis paper presents a non-contact fault diagnostic method for ball bearing using adaptive wavelet denoising, statistical-spectral acoustic features, and one-dimensional (1D) convolutional neural networks (CNN). The health conditions of the ball bearing are monitored by microphone under noisy condition. To eliminate noise, adaptive wavelet denoising method based on kurtosis-entropy (KE) index is proposed. Multiple acoustic features are extracted base on expert knowledge. The 1D ResNet is used to classify the health conditions of the bearings. Case study is presented to examine the proposed method’s capability to monitor the condition of ball bearings. The fault diagnosis results were compared with and without the adaptive wavelet denoising. The results show its effectiveness of the proposed fault diagnostic method using acoustic signals. Won-Ho Jung, Yong-Hwa Park |
IECON | 2 |
| 2024 | Performance Metric for Multiple Anomaly Score Distributions with Discrete Severity LevelsabstractThe rise of smart factories has heightened the demand for automated maintenance, and normal-data-based anomaly detection has proved particularly effective in environments where anomaly data are scarce. This method, which does not require anomaly data during training, has prompted researchers to focus not only on detecting anomalies but also on classifying severity levels by using anomaly scores. However, the existing performance metrics, such as the area under the receiver operating characteristic curve (AUROC), do not effectively reflect the performance of models in classifying severity levels based on anomaly scores. To address this limitation, we propose the weighted sum of the area under the receiver operating characteristic curve (WS-AUROC), which combines AUROC with a penalty for severity level differences. We conducted various experiments using different penalty assignment methods: uniform penalty regardless of severity level differences, penalty based on severity level index differences, and penalty based on actual physical quantities that cause anomalies. The latter method was the most sensitive. Additionally, we propose an anomaly detector that achieves clear separation of distributions and outperforms the ablation models on the WS-AUROC and AUROC metrics. Wonjun Yi, Won-Ho Jung, Yong-Hwa Park |
IECON | 3 |
| 2024 | Diversifying and Expanding Frequency-Adaptive Convolution Kernels for Sound Event DetectionabstractFrequency dynamic convolution (FDY conv) has shown the state-of-the-art performance in sound event detection (SED) using frequency-adaptive kernels obtained by frequency-varying combination of basis kernels.However, FDY conv lacks an explicit mean to diversify frequency-adaptive kernels, potentially limiting the performance.In addition, size of basis kernels is limited while time-frequency patterns span larger spectro-temporal range.Therefore, we propose dilated frequency dynamic convolution (DFD conv) which diversifies and expands frequency-adaptive kernels by introducing different dilation sizes to basis kernels.Experiments showed advantages of varying dilation sizes along frequency dimension, and analysis on attention weight variance proved dilated basis kernels are effectively diversified.By adapting class-wise median filter with intersection-based F1 score, proposed DFD-CRNN outperforms FDY-CRNN by 3.12% in terms of polyphonic sound detection score (PSDS). Hyeonuk Nam, Seong-Hu Kim, Deokki Min, Junhyeok Lee 0001, Yong-Hwa Park |
INTERSPEECH | 5 |
| 2024 | Thermal-Infrared Remote-Target Detection System for Maritime Rescue Using 3-D Game-Based Data Augmentation With GANabstractThis article proposes a deep learning-based thermal-infrared (TIR) remote target detection system for maritime rescue with a self-collected real TIR dataset and corresponding data augmentation method based on generative adversarial network (GAN). We have collected and established a real field TIR dataset consisting of multiple scenes imitating actual human rescue scenarios using a TIR camera (FLIR M364C). In addition, synthetic TIR data from a game (ARMA3) to augment the real TIR data are further collected to address dataset scarcity and improve the model performance. However, a significant domain gap exists between the real and synthetic TIR datasets. Hence, a proper domain adaptation (DA) algorithm is essential to overcome the gap. We suggest a target-background separation (TBS) scheme during the DA to mitigate this gap while preserving the shapes and locations of the small-size targets even after the domain transfer. Furthermore, a fixed-pattern kernel module inserted at the network front is proposed to improve the signal-to-noise ratio (SNR) as TIR remote targets inherently suffer from unclear boundaries and heavy clutters. The experimental results reveal that the segmentation network trained on both real and domain-translated synthetic TIR data shows improved performance compared to that trained on only real TIR data. Moreover, the segmentation network with the fixed-weight (FW) kernel module shows better performance than state-of-the-art methods in terms of every evaluation metric. Sungjin Cheong, Won-Ho Jung, Yoon-Seop Lim, Yong-Hwa Park |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Temporal Dynamic Convolutional Neural Network for Text-Independent Speaker Verification and Phonemic AnalysisabstractIn the field of text-independent speaker recognition, dynamic models that adapt along the time axis have been proposed to consider the phoneme-varying characteristics of speech. However, a detailed analysis of how dynamic models work depending on phonemes is insufficient. In this paper, we propose temporal dynamic CNN (TDY-CNN) that considers temporal variation of phonemes by applying kernels optimally adapting to each time bin. These kernels adapt to time bins by applying weighted sum of trained basis kernels. Then, an analysis of how adaptive kernels work on different phonemes in various layers is carried out. TDY-ResNet-38(×0.5) using six basis kernels improved an equal error rate (EER), the speaker verification performance, by 17.3% compared to the baseline model ResNet-38(×0.5). In addition, we showed that adaptive kernels depend on phoneme groups and are more phoneme-specific at early layers. The temporal dynamic model adapts itself to phonemes without explicitly given phoneme information during training, and results show the necessity to consider phoneme variation within utterances for more accurate and robust text-independent speaker verification. Seong-Hu Kim, Hyeonuk Nam, Yong-Hwa Park |
ICASSP | 3 |
| 2022 | Filteraugment: An Acoustic Environmental Data Augmentation MethodabstractAcoustic environments affect acoustic characteristics of sound to be recognized by physically interacting with sound wave propagation. Thus, training acoustic models for audio and speech tasks requires regularization on various acoustic environments in order to achieve robust performance in real life applications. We propose FilterAugment, a data augmentation method for regularization of acoustic models on various acoustic environments. FilterAugment mimics acoustic filters by applying different weights on frequency bands, therefore enables model to extract relevant information from wider frequency region. It is an improved version of frequency masking which masks information on random frequency bands. FilterAugment improved sound event detection (SED) model performance by 6.50% while frequency masking only improved 2.13% in terms of polyphonic sound detection score (PSDS). It achieved equal error rate (EER) of 1.22% when applied to a text-independent speaker verification model, outperforming model used frequency masking with EER of 1.26%. Prototype of FilterAugment was applied in our participation in DCASE 2021 challenge task 4, and played a major role in achieving the 3rd rank. Hyeonuk Nam, Seong-Hu Kim, Yong-Hwa Park |
ICASSP | 3 |
| 2022 | Fault Diagnosis of Inter-turn Short Circuit in Permanent Magnet Synchronous Motors with Current Signal Imaging and Semi-Supervised LearningabstractThis paper proposes machine-independent feature engineering for winding inter-turn short circuit fault that uses electrical current signals. Electrical current signal collected from permanent magnet synchronous motor (PMSM) is subjected to different environmental and operational conditions. To solve these problems, robust current signal imaging method and deep learning-based feature extraction method are developed. The overall procedure includes the following three key steps: (1) transformation of a one-dimensional time-series current signal to a two-dimensional image, (2) extracting features using convolutional neural networks, and (3) calculating a health indicator using Mahalanobis distance. Transformation of the time-series signal is based on recurrence plots (RP). The proposed RP method develops from feature engineering that provides the dominant fault feature representations in a robust way. The proposed RP is designed that maximizes the features of inter-turn short fault and minimizes the effect of noise from systems with various capacities. To demonstrate the validity of the proposed method, two case studies are conducted using an artificial fault seeded testbed with two different capacities of motor. By calculating the feature using only the electrical current signal of the motor without the parameters related to the capacity of the motor, the proposed feature can be applied to motors with different capacities while maintaining the same performance. Won-Ho Jung, Sung-Hyun Yun, Yoon-Seop Lim, Sungjin Cheong, Jaewoong Bae, Yong-Hwa Park |
IECON | 6 |
| 2022 | Fault Diagnosis of Ball Bearing Using Dynamic Convolutional Neural Networks Under Varying Speed ConditionabstractThe driving speed of bearing in rotating machines is usually variable rather than constant, so methods for accurate fault diagnosis under varying speed condition is required. In this paper, we propose the fault diagnosis model using dynamic convolutional neural network (DY-CNN) that considers variation of fault frequency characteristics by utilizing content-adaptive kernels for fault diagnosis of bearing under varying speed condition. As the input of model, 1-second intervals of vibration data with varying speed condition were used. These kernels adapt to short interval vibration data with varying speed condition by applying weighted sum of trained basis kernels. DY-CNN-based fault diagnosis model improved diagnosis accuracy by 7.07% compared to CNN-based fault diagnosis model. In addition, we showed that the adaptive kernels changed depending on fault types. DY-CNN-based fault diagnosis model adapted itself to fault types, and it performed accurate and robust fault diagnosis of ball bearing under varying speed condition. Seong-Hu Kim, Won-Ho Jung, Daegeun Lim, Yong-Hwa Park |
IECON | 4 |
| 2022 | Distortion Correction using Virtual PCG Pattern for Precise Stereo-based Large-scale 3D MeasurementabstractThree-dimensional (3D) measurement is an essential procedure in various manufacturing industries including shipbuilding. Since binocular systems are convenient and time-saving, they are proposed for ship block measurement. However, because of the very large scale of the ship blocks, working distance is about 10 m, resulting in a fatal limitation of calibration: the size of image portion corresponding to the checkerboard for calibration included in the whole image is extremely small. This prevents the distortion parameter of camera lens which most affects 3D reconstruction accuracy, from being accurately estimated in the calibration. To overcome this limitation, this paper proposes a method that pre-estimates the distortion correction map that covers the entire image area. A phase-shift circular grating (PCG) pattern displayed on a monitor is captured by the camera set to large scale. Since PCG patterns are generated by computer software, infinite number of patterns corresponding to desired orientations and positions can be generated, which are useful to measure the center of distortion and more accurate vanishing points. Based on the estimated vanishing points, the accurate distortion correction is performed using perspective projection invariants, and the distortion values in pixels are measured for each grid points to estimate the distortion correction map of the entire image area. An experimental 3D measurement was conducted with an estimated distortion correction map. As a result, Mean and standard deviation of 3D reconstruction error by proposed method were improved by 15.84% and 6.77% compared with the Zhang’s method, respectively. Jaeduck Lee, Zoohwan Hah, Yong-Hwa Park |
IECON | 4 |
| 2022 | Frequency Dynamic Convolution: Frequency-Adaptive Pattern Recognition for Sound Event Detectionabstract2D convolution is widely used in sound event detection (SED) to recognize two dimensional time-frequency patterns of sound events.However, 2D convolution enforces translation equivariance on sound events along both time and frequency axis while frequency is not shift-invariant dimension.In order to improve physical consistency of 2D convolution on SED, we propose frequency dynamic convolution which applies kernel that adapts to frequency components of input.Frequency dynamic convolution outperforms the baseline by 6.3% in DESED validation dataset in terms of polyphonic sound detection score (PSDS).It also significantly outperforms other pre-existing contentadaptive methods on SED.In addition, by comparing class-wise F1 scores of baseline and frequency dynamic convolution, we showed that frequency dynamic convolution is especially more effective for detection of non-stationary sound events with intricate time-frequency patterns.From this result, we verified that frequency dynamic convolution is superior in recognizing frequency-dependent patterns. Hyeonuk Nam, Seong-Hu Kim, Byeong-Yun Ko, Yong-Hwa Park |
INTERSPEECH | 4 |
| 2022 | Deep learning based cough detection camera using enhanced featuresabstractCoughing is a typical symptom of COVID-19. To detect and localize coughing sounds remotely, a convolutional neural network (CNN) based deep learning model was developed in this work and integrated with a sound camera for the visualization of the cough sounds. The cough detection model is a binary classifier of which the input is a two second acoustic feature and the output is one of two inferences (Cough or Others). Data augmentation was performed on the collected audio files to alleviate class imbalance and reflect various background noises in practical environments. For effective featuring of the cough sound, conventional features such as spectrograms, mel-scaled spectrograms, and mel-frequency cepstral coefficients (MFCC) were reinforced by utilizing their velocity (V) and acceleration (A) maps in this work. VGGNet, GoogLeNet, and ResNet were simplified to binary classifiers, and were named V-net, G-net, and R-net, respectively. To find the best combination of features and networks, training was performed for a total of 39 cases and the performance was confirmed using the test F1 score. Finally, a test F1 score of 91.9% (test accuracy of 97.2%) was achieved from G-net with the MFCC-V-A feature (named Spectroflow), an acoustic feature effective for use in cough detection. The trained cough detection model was integrated with a sound camera (i.e., one that visualizes sound sources using a beamforming microphone array). In a pilot test, the cough detection camera detected coughing sounds with an F1 score of 90.0% (accuracy of 96.0%), and the cough location in the camera image was tracked in real time. Gyeong-Tae Lee, Hyeonuk Nam, Seong-Hu Kim, Sang-Min Choi, Youngkey Kim, Yong-Hwa Park |
Expert Syst. Appl. | 6 |
| 2021 | Adaptive Convolutional Neural Network for Text-Independent Speaker Recognition
Seong-Hu Kim, Yong-Hwa Park |
Interspeech | 2 |