Sangwook Park 0002

dblp:65/364-2 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
5since 2021 · last 2024
0000-0002-6817-4846ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2024 Biomimetic Mappings for Active Sonar Object Recognition in Clutter
abstract
SONAR technology plays a pivotal role in terrain exploration and specifically identification of objects of interest. However, it grapples with a recurring challenge of clutter and noise which limits the performance of target recognition models. The challenge of noisy observations renders the choice of robust signal representations critical. Inspired by mammalian representation in the midbrain of echolocating bats, the present study evaluates the robustness of a decomposition of echo measurements that matches the statistics of natural vocalizations. This representation is contrasted with equally rich generic mappings as well as digital sonar images based on time-frequency representations. The study shows the clear advantage of the naturally optimized representation for object recognition in presence of background noise and clutter, and further underscores the potential of bio-inspired approaches in advancing SONAR technology.
Sangwook Park 0002, Angeles Salles, Kathryne Allen, Cynthia F. Moss, Mounya Elhilali
ICASSP1
2023 Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
abstract
Sound event detection is an important facet of audio tagging that aims to identify sounds of interest and define both the sound category and time boundaries for each sound event in a continuous recording. With advances in deep neural networks, there has been tremendous improvement in the performance of sound event detection systems, although at the expense of costly data collection and labeling efforts. In fact, current state-of-the-art methods employ supervised training methods that leverage large amounts of data samples and corresponding labels in order to facilitate identification of sound category and time stamps of events. As an alternative, the current study proposes a semi-supervised method for generating pseudo-labels from unsupervised data using a student-teacher scheme that balances self-training and cross-training. Additionally, this paper explores post-processing which extracts sound intervals from network prediction, for further improvement in sound event detection performance. The proposed approach is evaluated on sound event detection task for the DCASE2020 challenge. The results of these methods on both "validation" and "public evaluation" sets of DESED database show significant improvement compared to the state-of-the art systems in semi-supervised learning.
Sangwook Park 0002, David K. Han, Mounya Elhilali
IEEE Trans. Multim.1
2022 Time-Balanced Focal Loss for Audio Event Detection
abstract
Sound Event Detection (SED) tackles the challenge of identifying sound events in an audio recording by delimiting both their temporal boundaries as well as sound category. With recent advances in deep learning, current systems are able to leverage availability of large datasets to train sophisticated and highly effective SED models. Nonetheless, sound sources and acoustic characteristics of different classes vary greatly in their prevalence as well as representation in labeled datasets. The challenge with data imbalance in the case of SED stems not only from the representation (number of samples) across classes but also the natural asymmetry in time duration across different events varying from short transient events such as the clacking of dishes to more sustained events such as vacuuming. This variability results in an inherent disproportional representation of effective training samples. To address this compounded imbalance issue, this work proposes a balanced focal learning function that introduces a novel time-sensitive classwise weight. The proposed loss is applied to SED in the context of DCASE2021 challenge, and reports a notable improvement over the baseline, particularly in the case of shorter sound events.
Sangwook Park 0002, Mounya Elhilali
ICASSP1
2022 Temporal coding with magnitude-phase regularization for sound event detection
Sangwook Park 0002, Sandeep Kothinti, Mounya Elhilali
INTERSPEECH1
2021 Self-Training for Sound Event Detection in Audio Mixtures
abstract
Sound event detection (SED) takes on the task of identifying presence of specific sound events in a complex audio recording. SED has tremendous implications in video analytics, smart speaker algorithms and audio tagging. Recent advances in deep learning have afforded remarkable advances in performance of SED systems; albeit at the cost of extensive labeling efforts to train supervised methods using fully described sound class labels and timestamps. In order to address limitations in availability of training data, this work proposes a self-training technique to leverage unlabeled datasets in supervised learning using pseudo label estimation. This approach proposes a dual-term objective function: a classification loss for the original labels and expectation loss for pseudo labels. The proposed self training technique is applied to sound event detection in the context of the DCASE 2020 challenge, and reports a notable improvement over the baseline system for this task. The self-training approach is particularly effective in extending the labeled database with concurrent sound events.
Sangwook Park 0002, Ashwin Bellur, David K. Han, Mounya Elhilali
ICASSP1
2020 Amphibian Sounds Generating Network Based on Adversarial Learning
abstract
This letter proposes a generative network based on adversarial learning for synthesizing short-time audio streams and investigates the effectiveness of data augmentation for amphibian call sounds classification. Based on Fourier analysis, the generator is designed by a multi-layer perceptron composed of frequency basis learning layers and an output layer, and a discriminator is constructed by a convolutional neural network. Additionally, regularization on weights is introduced to train the networks with practical data that includes some disturbances. Synthetic audio streams are evaluated by quantitative comparison using inception score, and classification results are compared for real versus synthetic data. In conclusion, the proposed generative network is shown to produce realistic sounds and therefore useful for data augmentation.
Sangwook Park 0002, Mounya Elhilali, David K. Han, Hanseok Ko
IEEE Signal Process. Lett.1
2019 A Study of a Cross-Language Perception Based on Cortical Analysis Using Biomimetic STRFs
Sangwook Park 0002, David K. Han, Mounya Elhilali
INTERSPEECH1
2018 Man-Made Radio Frequency Interference Suppression for Compact HF Surface Wave Radar
abstract
High-frequency surface wave radar (HFSWR) suffers from a man-made interference because its amplitude is high enough to mask the Bragg scattering signal. Although several methods have been proposed for resolving this problem, they are inapplicable to compact HFSWR due to their antenna structures. This letter proposes an effective method of suppressing man-made radio frequency interference for compact HFSWR. The proposed method is composed of man-made interference detection and suppression by using regression based on probabilistic signal model. The proposed method is demonstrated in comparison with conventional methods in terms of root-mean-square error in experiments using synthetic and real data. The results show that the proposed method outperforms other methods in both simulated and practical situations.
Younglo Lee, Sangwook Park 0002, Chul Jin Cho, Bonhwa Ku, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.2
2017 Subspace projection cepstral coefficients for noise robust acoustic event recognition
abstract
In this paper, a novel feature for noise robust sound event recognition is proposed. The proposed feature is obtained by a two-step procedure. First, a subspace bank is established via target event analysis in complex vector space. Then, by projecting observation vectors onto the subspace bank, noise effect can be reduced while generating discriminant characters originated from differing event subspaces. To demonstrate robustness of the proposed feature, experiments with several classifiers were conducted with varying SNR cases under four noisy environments. According to the experimental results, the proposed method has shown superior performance over prominent conventional methods.
Sangwook Park 0002, Younglo Lee, David K. Han, Hanseok Ko
ICASSP1
2017 Compact HF Surface Wave Radar Data Generating Simulator for Ship Detection and Tracking
abstract
Toward a maritime surveillance objective, many ship detection and tracking algorithms have been investigated but are faced with poor performance in practical ocean environments. Compact high-frequency (HF) radar has also faced critical issues due to its long coherent processing interval and varying response from its orthogonal antenna structure. Hence, a simulator based on compact HF radar is proposed in this letter to provide a guideline for effective assessment of ship detection and tracking algorithms while considering these practical issues. To validate the proposed simulator, the simulator generated data has been compared with real data obtained by the compact HF radar sites.
Sangwook Park 0002, Chul Jin Cho, Bonhwa Ku, Hanseok Ko
IEEE Geosci. Remote. Sens. Lett.1
2015 Acoustic event recognition using dominant spectral basis vectors
Woohyun Choi, Sangwook Park 0002, David K. Han, Hanseok Ko
INTERSPEECH2