Jung-Woo Choi

dblp:116/2314 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-7264-6017ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
abstract
We propose a multichannel-to-multichannel target sound extraction (M2M-TSE) framework for separating multichannel target signals from a multichannel mixture of sound sources. Target sound extraction (TSE) isolates a specific target signal using user-provided clues, typically focusing on single-channel extraction with class labels or temporal activation maps. However, to preserve and utilize spatial information in multichannel audio signals, it is essential to extract multichannel signals of a target sound source. Moreover, the clue for extraction can also include spatial or temporal cues like direction-of-arrival (DoA) or timestamps of source activation. To address these challenges, we present an M2M framework that extracts a multichannel sound signal based on spatio-temporal clues.We demonstrate that our transformer-based architecture can successively accomplish the M2M-TSE task for multichannel signals synthesized from audio signals of diverse classes in different room environments. Furthermore, we show that the multichannel extraction task introduces sufficient inductive bias in the DNN, allowing it to directly handle DoA clues without utilizing hand-crafted spatial features.
Dayun Choi, Jung-Woo Choi
ICASSP2
2025 DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
abstract
This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed framework, DeFT-Mamba, utilizes the dense frequency-time attentive network (DeFTAN) combined with Mamba to extract sound objects, capturing the local time-frequency relations through gated convolution block and the global time-frequency relations through position-wise Hybrid Mamba. DeFT-Mamba surpasses existing separation and classification networks by a large margin, particularly in complex scenarios involving in-class polyphony. Additionally, a classification-based source counting method is introduced to identify the presence of multiple sources, outperforming conventional threshold-based approaches. Separation refinement tuning is also proposed to improve performance further. The proposed framework is trained and tested on a multichannel universal sound separation dataset developed in this work, designed to mimic realistic environments with moving sources and varying onsets and offsets of polyphonic events.
Jung-Woo Choi
ICASSP2
2025 DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis
abstract
We propose DeepASA, a multi-purpose model for auditory scene analysis that performs multi-input multi-output (MIMO) source separation, dereverberation, sound event detection (SED), audio classification, and direction-of-arrival estimation (DoAE) within a unified framework. DeepASA is designed for complex auditory scenes where multiple, often similar, sound sources overlap in time and move dynamically in space. To achieve robust and consistent inference across tasks, we introduce an object-oriented processing (OOP) strategy. This approach encapsulates diverse auditory features into object-centric representations and refines them through a chain-of-inference (CoI) mechanism. The pipeline comprises a dynamic temporal kernel-based feature extractor, a transformer-based aggregator, and an object separator that yields per-object features. These features feed into multiple task-specific decoders. Our object-centric representations naturally resolve the parameter association ambiguity inherent in traditional track-wise processing. However, early-stage object separation can lead to failure in downstream ASA tasks. To address this, we implement temporal coherence matching (TCM) within the chain-of-inference, enabling multi-task fusion and iterative refinement of object features using estimated auditory parameters. We evaluate DeepASA on representative spatial audio benchmark datasets, including ASA2, MC-FUSS, and STARSS23. Experimental results show that our model achieves state-of-the-art performance across all evaluated tasks, demonstrating its effectiveness in both source separation and auditory parameter estimation under diverse spatial auditory scenes. The demo video, samples and code are available at https://huggingface.co/spaces/donghoney22/DeepASA.
Younghoo Kwon, Jung-Woo Choi
NeurIPS3
2024 Noisy-Arcmix: Additive Noisy Angular Margin Loss Combined With Mixup For Anomalous Sound Detection
abstract
Unsupervised anomalous sound detection (ASD) aims to identify anomalous sounds by learning the features of normal operational sounds and sensing their deviations. Recent approaches have focused on the self-supervised task utilizing the classification of normal data, and advanced models have shown that securing representation space for anomalous data is important through representation learning yielding compact intra-class and well-separated intra-class distributions. However, we show that conventional approaches often fail to ensure sufficient intra-class compactness and exhibit angular disparity between samples and their corresponding centers. In this paper, we propose a training technique aimed at ensuring intra-class compactness and increasing the angle gap between normal and anomalous samples. Furthermore, we present an architecture that extracts features for important temporal regions, enabling the model to learn which time frames should be emphasized or suppressed. Experimental results demonstrate that the proposed method achieves the best performance giving 0.90%, 0.83%, and 2.16% improvement in terms of AUC, pAUC, and mAUC, respectively, compared to the state-of-the-art method on DCASE 2020 Challenge Task2 dataset. The source codes are available at https://github.com/soonhyeon/Noisy-ArcMix
Soonhyeon Choi, Jung-Woo Choi
ICASSP2
2024 CST-Former: Transformer with Channel-Spectro-Temporal Attention for Sound Event Localization and Detection
abstract
Sound event localization and detection (SELD) is a task for the classification of sound events and the localization of direction of arrival (DoA) utilizing multichannel acoustic signals. Prior studies employ spectral and channel information as the embedding for temporal attention. However, this usage limits the deep neural network from extracting meaningful features from the spectral or spatial domains. Therefore, our investigation in this paper presents a novel framework termed the Channel-Spectro-Temporal Transformer (CST-former) that bolsters SELD performance through the independent application of attention mechanisms to distinct domains. The CST-former architecture employs distinct attention mechanisms to independently process channel, spectral, and temporal information. In addition, we propose an unfolded local embedding (ULE) technique for channel attention (CA) to generate informative embedding vectors including local spectral and temporal information. Empirical validation through experimentation on the 2022 and 2023 DCASE Challenge task3 datasets affirms the efficacy of employing attention mechanisms separated across each domain and the benefit of ULE, in enhancing SELD performance.
Yusun Shul, Jung-Woo Choi
ICASSP2
2024 DeFTAN-AA: Array Geometry Agnostic Multichannel Speech Enhancement
Jung-Woo Choi
INTERSPEECH2
2024 DeFTAN-II: Efficient Multichannel Speech Enhancement With Subgroup Processing
abstract
In this work, we present DeFTAN-II, an efficient multichannel speech enhancement model based on transformer architecture and subgroup processing. Despite the success of transformers in speech enhancement, they face challenges in capturing local relations, reducing the high computational complexity, and lowering memory usage. To address these limitations, we introduce subgroup processing in our model, combining subgroups of locally emphasized features with other subgroups containing original features. The subgroup processing is implemented in several blocks of the proposed network. In the proposed split dense blocks extracting spatial features, a pair of subgroups is sequentially concatenated and processed by convolution layers to effectively reduce the computational complexity and memory usage. For the F- and T-transformers extracting temporal and spectral relations, we introduce cross-attention between subgroups to identify relationships between locally emphasized and non-emphasized features. The dual-path feedforward network then aggregates attended features in terms of the gating of local features processed by dilated convolutions. Through extensive comparisons with state-of-the-art multichannel speech enhancement models, we demonstrate that DeFTAN-II with subgroup processing outperforms existing methods at significantly lower computational complexity. Moreover, we evaluate the model's generalization capability on real-world data without fine-tuning, which further demonstrates its effectiveness in practical scenarios.
Jung-Woo Choi
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
abstract
Accurate estimation of indoor space geometries is vital for constructing precise digital twins, whose broad industrial applications include navigation in unfamiliar environments and efficient evacuation planning, particularly in low-light conditions. This study introduces EchoScan, a deep neural network model that utilizes acoustic echoes to perform room geometry inference. Conventional sound-based techniques rely on estimating geometry-related room parameters such as wall position and room size, thereby limiting the diversity of inferable room geometries. Contrarily, EchoScan overcomes this limitation by directly inferring room floorplan maps and height maps, thereby enabling it to handle rooms with complex shapes, including curved walls. The segmentation task for predicting floorplan and height maps enables the model to leverage both low- and high-order reflections. The use of high-order reflections further allows EchoScan to infer complex room shapes when some walls of the room are unobservable from the position of an audio device. Herein, EchoScan was trained and evaluated using RIRs synthesized from complex environments, including the Manhattan and Atlanta layouts, employing a practical audio device configuration compatible with commercial, off-the-shelf devices.
Inmo Yeon, Iljoo Jeong, Seungchul Lee, Jung-Woo Choi
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 DeFT-AN RT: Real-time Multichannel Speech Enhancement using Dense Frequency-Time Attentive Network and Non-overlapping Synthesis Window
Dayun Choi, Jung-Woo Choi
INTERSPEECH3
2023 DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement
abstract
In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and reverberation embedded in the short-time Fourier transform (STFT) of an input signal. The proposed mask estimation network incorporates three different types of blocks for aggregating information in the spatial, spectral, and temporal dimensions. It utilizes a spectral transformer with a modified feed-forward network and a temporal conformer with sequential dilated convolutions. The use of dense blocks and transformers dedicated to the three different characteristics of audio signals enables more comprehensive enhancement in noisy and reverberant environments. The remarkable performance of DeFT-AN over state-of-the-art multichannel models is demonstrated based on two popular noisy and reverberant datasets in terms of various metrics for speech quality and intelligibility.
Jung-Woo Choi
IEEE Signal Process. Lett.2
2022 Multiarray Eigenbeam-ESPRIT for 3D Sound Source Localization With Multiple Spherical Microphone Arrays
abstract
A 3D sound source localization (SSL) technique for multiple spherical arrays is proposed. Spherical arrays have been popularly used for direction of arrival (DoA) estimation, which serves as important prior knowledge for source separation, dereverberation, and binaural rendering for VR applications. In this work, a parametric 3D SSL technique is proposed that can detect multiple sources using multiple spherical array recordings. From the enriched information by using multiple arrays, we show that the parametric estimation of 3D positions is possible without scanning candidate positions in 3D space. To mitigate the problem of the ambiguous association between DoAs estimated from different arrays, a total covariance matrix including both auto- and cross-covariance matrices between arrays is utilized such that DoAs from arrays are automatically paired to uniquely localize 3D source positions. The joint parameter estimation is based on the generalized joint Schur decomposition of multiple matrices combined with a geometric projection robustly guiding the convergence of the proposed iteration algorithm. Simulations conducted in anechoic and reverberant room conditions reveal that the proposed technique can accurately determine multiple source positions in 3D space without ambiguities.
Jung-Woo Choi, Franz Zotter, Byeongho Jo, Jae-Hyoun Yoo
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Iterative Echo Labeling Algorithm With Convex Hull Expansion for Room Geometry Estimation
abstract
Estimating room geometry is an important problem in audio signal processing, with various applications such as dereverberation and spatial audio. Several methods have been proposed to localize the positions of the walls given the times of arrival (TOAs) from acoustic room impulse responses (RIRs). To localize a reflector, reflections from the same wall should be identified first. When multiple microphones are widely installed inside a room, however, it is hard to identify the reflections generated by the same walls, which leads to the echo labeling problem. In this paper, we propose an iterative echo labeling algorithm to solve the echo labeling problem. We construct a convex hull of ellipses with first-arriving reflections at the initial iteration and expand the convex hull by utilizing later reflections at subsequent iterations. While expanding the convex hull, we sequentially localize the walls by inspecting tangents of the convex hull. Finally, the iteration is terminated when the expanded convex hull of ellipses becomes a superset of a room. As a result, the proposed algorithm can solve the echo labeling problem without prior knowledge of the number of walls and high computational complexity.
Sooyeon Park, Jung-Woo Choi
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Robust Sound Source Localization considering Similarity of Back-Propagation Signals
abstract
We present a novel, robust sound source localization algorithm considering back-propagation signals. Sound propagation paths are estimated by generating direct and reflection acoustic rays based on ray tracing in a backward manner. We then compute the back-propagation signals by designing and using the impulse response of the backward sound propagation based on the acoustic ray paths. For identifying the 3D source position, we use a well-established Monte Carlo localization method. Candidates for a source position are determined by identifying convergence regions of acoustic ray paths. Those candidates are validated by measuring similarities between back-propagation signals, under the assumption that the back-propagation signals of different acoustic ray paths should be similar near the ground-truth sound source position. Thanks to considering similarities of back-propagation signals, our approach can localize a source position with an averaged error of 0.55 m in a room of 7 m by 7 m area with 3 m height in tested environments. We also place additional 67 dB and 77 dB white noise at the background, to test the robustness of our approach. Overall, we observe a 7 % to 100 % improvement in accuracy over the state-of-the-art method.
Inkyu An, Byeongho Jo, Youngsun Kwon, Jung-Woo Choi, Sung-Eui Yoon
ICRA4
2020 Real-Time Demonstration of Personal Audio and 3D Audio Rendering Using Line Array Systems
Jung-Woo Choi
MMM (2)1
2020 Extended Vector-Based EB-ESPRIT Method
abstract
The estimation of direction of arrivals (DoAs) from spherical microphone array data is one of the key issues in extracting source information from all-around audio recordings. One such technique is the eigenbeam estimation of signal parameters via the rotational invariance technique (EB-ESPRIT), which separates the signal subspace related to the stationary sound field and then directly estimates DoAs of multiple sound sources. EB-ESPRIT has been evolved in many different ways by involving different types of recurrence relations of spherical harmonics, all of which are able to identify DoAs of a limited number of sources that are noticeably smaller than the number of finite-order spherical harmonic coefficients recorded. In this work, we report that it is possible to go beyond the known limits of detectable sources. The proposed formula is also based on conventional recurrence relations and probably permits to reach the ultimate limit by additional constraints of the signal parameters that can better exploit the highest-order coefficients. Monte-Carlo simulations conducted with various source positions and signal-to-noise ratios (SNRs) reveal that the proposed technique can detect more sources with insignificant loss in estimation performance and robustness.
Byeongho Jo, Franz Zotter, Jung-Woo Choi
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Subband Optimization and Filtering Technique for Practical Personal Audio Systems
abstract
The implementation of personal audio systems in enclosed spaces, such as a car cabin, suffers from severe echoes from surrounding boundaries. In order to focus sound energy on a single seat position, echoes should be controlled by long multichannel filters with up to a few thousand taps for each channel, which leads to increased memory size and computational complexity. In an attempt to design a practical personal audio system, a subband based optimization and filtering technique are proposed. The design of optimal filters for downsampled low frequency responses enables finer control of low frequency echoes without significantly increasing the number of filter taps, while the broadband response of high frequency components can be controlled with a fewer number of filter taps. Experiments conducted in a real car cabin demonstrate that more than 20 dB SPL difference can be achieved across different seat positions with only half the number of filter taps.
Hyungjun So, Jung-Woo Choi
ICASSP2
2019 Diffraction-Aware Sound Localization for a Non-Line-of-Sight Source
abstract
We present a novel sound localization algorithm for a non-line-of-sight (NLOS) sound source in indoor environments. Our approach exploits the diffraction properties of sound waves as they bend around a barrier or an obstacle in the scene. We combine a ray tracing-based sound propagation algorithm with a Uniform Theory of Diffraction (UTD) model, which simulate bending effects by placing a virtual sound source on a wedge in the environment. We precompute the wedges of a reconstructed mesh of an indoor scene and use them to generate diffraction acoustic rays to localize the 3D position of the source. Our method identifies the convergence region of those generated acoustic rays as the estimated source position based on a particle filter. We have evaluated our algorithm in multiple scenarios consisting of static and dynamic NLOS sound sources. In our tested cases, our approach can localize a source position with an average accuracy error of 0.7m, measured by the L2 distance between estimated and actual source locations in a 7m×7m×3m room. Furthermore, we observe 37% to 130% improvement in accuracy over a state-of-the-art localization method that does not model diffraction effects, especially when a sound source is not visible to the robot.
Inkyu An, Doheon Lee, Jung-Woo Choi, Dinesh Manocha, Sung-Eui Yoon
ICRA3
2017 Smart Loudspeaker Arrays for Self-Coordination and User Tracking
Jungju Jee, Jung-Woo Choi
MMM (2)2
2017 Spherical Harmonic Smoothing for Localizing Coherent Sound Sources
abstract
Correlated or coherent sources can cause the localization performance of subspace-based beamformers to deteriorate. To solve this problem, various smoothing techniques have been proposed for the localization of multiple coherent sound sources. A common principle of smoothing techniques is to increase the rank of a covariance matrix by constructing multiple subarrays in the space, time, or frequency domain. The construction of such subarrays, however, requires the satisfaction of strong assumptions regarding the microphone positions or the temporal/spectral structures of the signals. In this paper, we propose a spherical harmonic smoothing technique that can perform smoothing in terms of spherical harmonic coefficients only. Unlike other smoothing techniques, the proposed technique constructs subarrays of spherical harmonic coefficients and uses them to increase the number of linearly independent observations in a signal subspace. Subarray construction in the spherical harmonics domain enables the accurate localization of multiple coherent sources even at a single frequency. The proposed technique can be applied to an arbitrary microphone array as long as the spherical harmonic coefficients can be measured up to a finite order.
Byeongho Jo, Jung-Woo Choi
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Distance perception of a virtual sound source synthesized near the listener position
Dong-Soo Kang, Jung-Woo Choi, William L. Martens
Multim. Tools Appl.2
2014 Mobile maestro: enabling immersive multi-speaker audio applications on commodity mobile devices
abstract
The goal of this work is to provide an abstraction of ideal sound environments to a new emerging class of Mobile Multi-speaker Audio (MMA) applications. Typically, it is challenging for MMA applications to implement advanced sound features (e.g., surround sound) accurately in mobile environments, especially due to unknown, irregular loudspeaker configurations. Towards an illusion that MMA applications run over specific loudspeaker configurations (i.e., speaker type, layout), this work proposes AMAC, a new Adaptive Mobile Audio Coordination system that senses the acoustic characteristics of mobile environments and controls individual loud-speakers adaptively and accurately. The prototype of AMAC implemented on commodity smartphones shows that it provides the coordination accuracy in sound arrival time in several tens of microseconds and reduces the variance in sound level substantially.
Hyosu Kim, Jung-Woo Choi, Hwidong Bae, Junehwa Song, Insik Shin
UbiComp3
2013 Sound Field Reproduction of a Virtual Source Inside a Loudspeaker Array With Minimal External Radiation
abstract
We derived an integral equation for reproducing the sound field of a virtual source inside an array of loudspeakers with reduced radiation to the outside. Reproduction of a sound field over a finite interior region inevitably generates sound waves that propagate outside the region. This undesirable radiation is reflected from walls and can induce artifacts in the interior region. In principle, the Kirchhoff-Helmholtz (KH) integral can be used to reproduce the interior sound field from an exterior virtual source without any external radiation. However, if there is a virtual source inside the array, the integral formula does not explicitly demonstrate how one can reproduce the sound field or minimize the external radiation. In this work, we derive an explicit formula for reproducing a sound field with minimal external radiation when a virtual source is located inside a loudspeaker array. The theory shows that external radiation can be effectively reduced without solving any inverse problem. The proposed formula follows the form of the KH integral and thus requires monopole and dipole sources. Although dipole sources are difficult to build in practice, the theory predicts that sound field reproduction with minimal external radiation is possible and that the room dependency of the sound field reproduction system can be decreased.
Jung-Woo Choi, Yang-Hann Kim
IEEE Trans. Speech Audio Process.1
2012 Integral Approach for Reproduction of Virtual Sound Source Surrounded by Loudspeaker Array
abstract
A method for reproducing a desired sound field by using an array of loudspeakers is proposed. A virtual sound source that generates the desired sound field can be located at either the outside or inside a loudspeaker array; however, a complete reproduction of the virtual source located inside the array is physically not possible because the sound field from the internal virtual source should satisfy the inhomogeneous wave equation. When time reversal is used for removing such inhomogeneity, convergent waves toward the location of the virtual source always exist. The removal of these artifacts is an objective of our study. Most of the theoretical forms for the suppression of artifacts have been derived from the approximated Kirchhoff-Helmholtz integral incorporating a curved or line array. In this paper, we aim to develop a method based on a complete three-dimensional integral formula, which does not require an inversion of sound fields. Using the formula, we can directly predict the behavior of artifacts and design the excitation of the loudspeakers that effectively suppresses the artifacts. This paper also highlights a single-layer formula that only incorporates monopole arrays for the reproduction of the virtual source inside them.
Jung-Woo Choi, Yang-Hann Kim
IEEE Trans. Speech Audio Process.1
2012 Robustness and Regularization of Personal Audio Systems
abstract
As well as being able to reproduce sound in one region of space, it would be useful to reduce the level of reproduced sound in other spatial regions, with a “personal audio” system. For mobile devices this is motivated by issues of privacy for the user and the need to reduce annoyance for other people nearby. Such personal audio systems can be realized with arrays of loudspeakers that become superdirectional at low frequencies, when the array dimensions are small compared with the acoustic wavelength. The design of the array then becomes a compromise between performance and array effort, defined as the sum of mean squared driving signals. Various methods of formulating this tradeoff as a regularization problem have been suggested and the connection between these formulations is discussed. Large array efforts are due to strongly self-cancelling multipole arrays. A concern is then the robustness of such an array to variations in the acoustic environment and driver sensitivity and position. The design of an array that is robust to these uncertainties then leads to a generalization of regularization.
Stephen J. Elliott, Jordan Cheer, Jung-Woo Choi, Youngtae Kim
IEEE Trans. Speech Audio Process.3