Seungkwon Beack

dblp:36/3718 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
14since 2021 · last 2024
0000-0002-6254-2062ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
YearPublicationVenuePosition
2024 Personalized Neural Speech Codec
abstract
In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is involved on each side of the communication, personalizing a codec for the specific user has rarely been explored in the literature. First, we assume speakers can be grouped into smaller subsets based on their perceptual similarity. Then, we also postulate that a group-specific codec can focus on the group’s speech characteristics to improve its perceptual quality and computational efficiency. To this end, we first develop a Siamese network that learns the speaker embeddings from the LibriSpeech dataset, which are then grouped into underlying speaker clusters. Finally, we retrain the LPCNet-based speech codec baselines on each of the speaker clusters. Subjective listening tests show that the proposed personalization scheme introduces model compression while maintaining speech quality. In other words, with the same model complexity, personalized codecs produce better speech quality.
Inseon Jang, Haici Yang, Wootaek Lim, Seungkwon Beack
ICASSP4
2024 Pre-Echo Reduction in Transform Audio Coding via Temporal Envelope Control with Machine Learning Based Estimation
abstract
This paper proposes a new method for pre-echo reduction in transform-based audio coding by controlling the temporal envelope of the waveform. The proposed method comprises two operating modes: temporal envelope flattening and temporal envelope correction of a target signal. The proposed method estimates signal levels with a low temporal resolution from side information using machine learning and converts them into a signal to be applied to the target signal to flatten and correct the temporal envelope. It also adjusts the signals to maintain signal continuity between the non-transient and transient frames. The proposed method differs from conventional methods in that it directly modifies the waveform before encoding and after decoding, which makes it useful as a new coding tool for legacy codecs. A subjective performance evaluation confirms that the proposed method uses fewer bits to provide sound quality equivalent to that of the short-window transform.1
Byeongho Jo, Seungkwon Beack, Hochong Park
ICASSP3
2024 Quantization Noise Masking in Perceptual Neural Audio Coder
abstract
This study investigates the implication of utilizing the psychoacoustic model (PAM) within the neural audio coder (NAC), specifically focusing on the masking of quantization noise. We introduce a novel training strategy to incorporate the PAM into the NAC more accurately. This method involves a discriminator that directly or indirectly measures the PAM loss. For the indirect measurement, a multi-scale STFT discriminator (MS-STFTD) is incorporated to introduce an auxiliary loss term in addition to the existing PAM loss. Conversely, for the direct measurement, we have designed a multi-scale PAM discriminator (MS-PAMD) that quantifies PAM-specific parameters. Experimental results show that adding the discriminator masks the quantization noise better than the previous NAC, and it obtains audio quality comparable to the commercial AAC in both objective and subjective scores.
Seungmin Shin, Joon Byun, Jongmo Sung, Seungkwon Beack, Young-Cheol Park
ICASSP4
2024 Representations of the Complex-Valued Frequency-Domain LPC for Audio Coding
abstract
For decades, linear predictive (LP) analysis of the real-valued time-domain signals has been developed for speech coding, analysis, and synthesis. Recently, the complex-valued frequency-domain LP coding was developed to enhance the estimation performance of the temporal envelope. To apply it for audio coding, a suitable representation for the complex-valued frequency-domain LP coefficients (CLPC) is required before quantization, but there was no efficient way to represent it. To address the problem, we propose efficient CLPC representations that retain some useful properties of conventional LPC representations. Through quantitative and qualitative evaluations, we demonstrate that our proposed representations increase quantization efficiency and improve audio coding performance.
Byeongho Jo, Seungkwon Beack
IEEE Signal Process. Lett.2
2024 Efficient Complex Immittance Spectral Frequency With the Perceptual-Metric-Based Codebook Search
abstract
Complex-valued frequency-domain linear predictive coding (CLPC) has been developed for audio coding. Recently, representations for efficiently quantizing CLPC coefficients have been proposed, including the complex immittance spectral frequency (CISF). The CISF has limitations in that it requires signalling the sequential information to eliminate ambiguity and the highest-order coefficient (HOC) for reconstructing the CLPC coefficients. This study developed a modified CISF-based method that eliminates the need for additional information by utilizing intermediate complex polynomial properties. Furthermore, a perceptual-metric-based codebook search was proposed to improve quantization efficiency. The experimental results show robust quantization performance, while listening tests demonstrate superior audio quality compared to MPEG-D USAC long TCX at 12 kbps.
Byeongho Jo, Seungkwon Beack
IEEE Signal Process. Lett.2
2023 A Perceptual Neural Audio Coder with a Mean-Scale Hyperprior
abstract
This paper proposes an end-to-end neural audio coder based on a mean-scale hyperprior model together with a perceptual optimization using a psychoacoustic model (PAM)-based loss function. The proposed coder estimates the mean and scale hyperpriors using a sub-network after assuming that the probability distribution of latent samples is Gaussian. The main network is an autoencoder based on Resnet-type gated linear units (ResGLUs), each comprising a generalized divisive normalization (GDN) layer. We train both networks to optimize perceptual attributes estimated using a multi-timescale scheme to obtain high perceptual quality. Experimental results show that the proposed model accurately predicts the mean and scale hyperpriors. Also, it obtains consistently higher audio quality than the commercial MP3 audio coder at all bitrates.
Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack
ICASSP5
2023 Audio Coding With Unified Noise Shaping And Phase Contrast Control
abstract
Over the past decade, audio coding technology has seen standardization and the development of many frameworks incorporated with linear predictive coding (LPC). As LPC reduces information in the frequency domain, LP-based frequency-domain noise-shaping (FDNS) was previously proposed. To code transient signals effectively, FDNS with temporal noise shaping (TNS) has emerged. However, these mainly operated in the modified discrete cosine transform domain, which essentially accompanies time domain aliasing. In this paper, a unified noise-shaping (UNS) framework including FDNS and complex LPC-based TNS (CTNS) in the DFT domain is proposed to overcome the aliasing issues. Additionally, a modified polar quantizer with phase contrast control is proposed, which saves phase bits depending on the frequency envelope information. The core coding feasibility at low bit rates is verified through various objective metrics and subjective listening evaluations.
Byeongho Jo, Seungkwon Beack, Taejin Lee 0003
ICASSP2
2023 Perceptual Improvement of Deep Neural Network (DNN) Speech Coder Using Parametric and Non-parametric Density Models
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park
INTERSPEECH4
2022 Deep Neural Network (DNN) Audio Coder Using A Perceptually Improved Training Method
abstract
A new end-to-end audio coder based on a deep neural network (DNN) is proposed. To compensate for the perceptual distortion that occurred by quantization, the proposed coder is optimized to minimize distortions in both signal and perceptual domains. The distortion in the perceptual domain is measured using the psychoacoustic model (PAM), and a loss function is obtained through the two-stage compensation approach. Also, the scalar uniform quantization was approximated using a uniform stochastic noise, together with a compression-decompression scheme, which provides simpler but more stable learning without an additional penalty than the softmax quantizer. Test results showed that the proposed coder achieves more accurate noise-masking than the previous PAM-based method and better perceptual quality then the MP3 audio coder.
Seungmin Shin, Joon Byun, Young-Cheol Park, Jongmo Sung, Seungkwon Beack
ICASSP5
2022 Optimization of Deep Neural Network (DNN) Speech Coder Using a Multi Time Scale Perceptual Loss Function
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park
INTERSPEECH4
2022 Highly Efficient Audio Coding With Blind Spectral Recovery Based on Machine Learning
abstract
This letter proposes a new method for audio coding that utilizes blind spectral recovery to improve the coding efficiency without compromising performance. The proposed method transmits only a fraction of the spectral coefficients, thereby reducing the coding bit rate. Then, it recovers the remaining coefficients in the decoder using the transmitted coefficients as input. The proposed method is differentiated from conventional spectral recovery in that the coefficients to be recovered are interleaved with the transmitted coefficients to obtain the most data correlation. Further, it enhances the transmitted coefficients, which are degraded by quantization errors, to deliver better information to the recovery process. The spectral recovery is conducted recursively on a band basis such that information recovered in one band is used for the recovery in subsequent bands. An improved level correction for the recovered coefficients and a new sign coding are also developed. A subjective performance evaluation confirms that the proposed method at 40 kbps provides statistically equivalent sound quality to a state-of-the-art coding method at 48 kbps for speech and music categories.
Seungkwon Beack, Wootaek Lim, Hochong Park
IEEE Signal Process. Lett.2
2022 Scalable and Efficient Neural Speech Coding: A Hybrid Design
abstract
We present a scalable and efficient neural waveform coding system for speech compression. We formulate the speech coding problem as an autoencoding task, where a convolutional neural network (CNN) performs encoding and decoding as a neural waveform codec (NWC) during its feedforward routine. The proposed NWC also defines quantization and entropy coding as a trainable module, so the coding artifacts and bitrate control are handled during the optimization process. We achieve efficiency by introducing compact model components to NWC, such as gated residual networks and depthwise separable convolution. Furthermore, the proposed models are with a scalable architecture, cross-module residual learning (CMRL), to cover a wide range of bitrates. To this end, we employ the residual coding concept to concatenate multiple NWC autoencoding modules, where each NWC module performs residual coding to restore any reconstruction loss that its preceding modules have created. CMRL can scale down to cover lower bitrates as well, for which it employs linear predictive coding (LPC) module as its first autoencoder. The hybrid design integrates LPC and NWC by redefining LPC’s quantization as a differentiable process, making the system training an end-to-end manner. The decoder of proposed system is with either one NWC (0.12 million parameters) in low to medium bitrate ranges (12 to 20 kbps) or two NWCs in the high bitrate (32 kbps). Although the decoding complexity is not yet as low as that of conventional speech codecs, it is significantly reduced from that of other neural speech coders, such as a WaveNet-based vocoder. For wide-band speech coding quality, our system yields comparable or superior performance to AMR-WB and Opus on TIMIT test utterances at low and medium bitrates. The proposed system can scale up to higher bitrates to achieve near transparent performance.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, Minje Kim 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Source-Aware Neural Speech Coding for Noisy Speech Compression
abstract
This paper introduces a novel neural network-based speech coding system that can process noisy speech effectively. The proposed source-aware neural audio coding (SANAC) system harmonizes a deep autoencoder-based source separation model and a neural coding system, so that it can explicitly perform source separation and coding in the latent space. An added benefit of this system is that the codec can allocate a different amount of bits to the underlying sources, so that the more important source sounds better in the decoded signal. We target a new use case where the user on the receiver side cares about the quality of the non-speech components in the speech communication, while the speech source still carries the most important information. Both objective and subjective evaluation tests show that SANAC can recover the original noisy speech better than the baseline neural audio coding system, which is with no source-aware coding mechanism, and two conventional codecs.
Haici Yang, Kai Zhen, Seungkwon Beack
ICASSP3
2021 Development of a Psychoacoustic Loss Function for the Deep Neural Network (DNN)-Based Speech Coder
Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack
Interspeech5
2020 Enhanced Method of Audio Coding Using CNN-Based Spectral Recovery with Adaptive Structure
abstract
A process of spectral recovery can enhance the performance of transform-based audio coding by transmitting only a portion of spectral data and recovering the missing spectral data in the decoder. This study proposes an enhanced method of audio coding based on spectral recovery with an adaptive structure that yields improved sound quality compared with the previous method. The spectral data to be recovered are arranged in an adaptive pattern depending on the difficulty of recovery. In addition, according to the spectral characteristics, prior information associated with these spectral data is selectively transmitted that helps a neural network improve the performance of magnitude recovery. Prior information also provides the signs of recovered magnitudes. A subjective performance evaluation shows that, for mono coding without window switching at 40 kbps, the proposed coding method provides better sound quality than the conventional method on average.
Seong-Hyeon Shin, Seungkwon Beack, Wootaek Lim, Hochong Park
ICASSP2
2020 Efficient and Scalable Neural Residual Waveform Coding with Collaborative Quantization
abstract
Scalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC coefficients and the corresponding residuals. CQ does not simply shoehorn LPC to a neural network, but bridges the computational capacity of advanced neural network models and traditional, yet efficient and domain-specific digital signal processing methods in an integrated manner. We demonstrate that CQ achieves much higher quality than its predecessor at 9 kbps with even lower model complexity. We also show that CQ can scale up to 24 kbps where it outperforms AMR-WB and Opus. As a neural waveform codec, CQ models are with less than 1 million parameters, significantly less than many other generative models.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack
ICASSP4
2020 Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
abstract
Conventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs, however, the objective nature of the loss function usually leads to suboptimal sound quality as well as high run-time complexity due to the large model size. In this work, we present a psychoacoustic calibration scheme to re-define the loss functions of neural audio coding systems so that it can decode signals more perceptually similar to the reference, yet with a much lower model complexity. The proposed loss function incorporates the global masking threshold, allowing the reconstruction error that corresponds to inaudible artifacts. Experimental results show that the proposed model outperforms the baseline neural codec twice as large and consuming 23.4% more bits per second. With the proposed method, a lightweight neural codec, with only 0.9 million parameters, performs near-transparent audio coding comparable with the commercial MPEG-1 Audio Layer III codec at 112 kbps.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack, Minje Kim 0001
IEEE Signal Process. Lett.4
2019 Audio Coding Based on Spectral Recovery by Convolutional Neural Network
abstract
This study proposes a new method of audio coding based on spectral recovery, which can enhance the performance of transform audio coding. An encoder represents spectral information of an input in a time-frequency domain and transmits only a portion of it so that the remaining spectral information can be recovered based on the transmitted information. A decoder recovers the magnitudes of missing spectral information using a convolutional neural network. The signs of missing spectral information are either transmitted or randomly assigned, according to their importance. By combining transmission and recovery of spectral information, the proposed method can enhance the coding performance, compared with conventional transform coding. The subjective performance evaluation shows that, for mono coding at 39.4 kbps, the proposed method provides higher sound quality than the USAC, by an average MUSHRA score of 8.5.
Seong-Hyeon Shin, Seungkwon Beack, Taejin Lee 0003, Hochong Park
ICASSP2
2019 Cascaded Cross-Module Residual Learning Towards Lightweight End-to-End Speech Coding
abstract
Speech codecs learn compact representations of speech signals to facilitate data transmission.Many recent deep neural network (DNN) based end-to-end speech codecs achieve low bitrates and high perceptual quality at the cost of model complexity.We propose a cross-module residual learning (CMRL) pipeline as a module carrier with each module reconstructing the residual from its preceding modules.CMRL differs from other DNN-based speech codecs, in that rather than modeling speech compression problem in a single large neural network, it optimizes a series of less-complicated modules in a two-phase training scheme.The proposed method shows better objective performance than AMR-WB and the state-of-the-art DNNbased speech codec with a similar network architecture.As an end-to-end model, it takes raw PCM signals as an input, but is also compatible with linear predictive coding (LPC), showing better subjective quality at high bitrates than AMR-WB and OPUS.The gain is achieved by using only 0.9 million trainable parameters, a significantly less complex architecture than the other DNN-based codecs in the literature.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack
INTERSPEECH4
2011 Spatial Audio Object Coding With Two-Step Coding Structure for Interactive Audio Service
abstract
An interactive audio service is a new conceptual audio service that provides the users with opportunities for a variety of experiences on the alternative and advanced audio services. In the interactive audio service, users can freely control various audio objects to make their own audio sounds. A spatial audio object coding (SAOC) is a useful technology that can support most parts of the interactive audio service with a relatively low bit-rate, but is very poor to perfect gain control of a certain audio object, i.e., the target audio object. In this paper, the SAOC with a two-step coding structure is proposed to efficiently handle the target audio object as well as the normal audio objects. A transform coded excitation (TCX) based residual coding scheme is presented in the context of the sound quality enhancement. From experimental results, it can be noted that the various audio objects can be successfully handled with respect to the bit-rate and the sound quality by using the proposed two-step coding structure SAOC.
Kwang-Ki Kim, Jeongil Seo, Seungkwon Beack, Kyeongok Kang, Minsoo Hahn
IEEE Trans. Multim.3
2007 Spatial Cue Based Sound Scene Control for MPEG Surround
abstract
Spatial audio coding scheme is advanced technology in the sense of saving bitrate without any scarifications of sound quality. In this paper, the method to utilizing the spatial cues is present. This utilization is for supporting a new useful functionality, i.e., flexible sound scene control. The function of sound scene control may result in providing user preferable audio service by interaction and be utilized in upcoming applications such as multi-view video service and virtual reality applications. It will be shown that even with a little addition of computational operations; this useful functionality will be achievable maintaining the corresponding sound quality.
Seungkwon Beack, Jeongil Seo, Taejin Lee 0003, Dae-young Jang
ICME1
2005 Sound source location cue coding system for compact representation of multi-channel audio
abstract
Binaural cue coding (BCC) has been introduced for compact representation of multi-channel audio. It exploits binaural cue parameters for capturing the spatial image of multi-channel audio. Recently, it has been standardized within MPEG as the name of "MPEG Surround." In this paper, we propose a sound source location cue coding (SSLCC) system for compressing multi-channel audio to be suitable at the narrow bandwidth transmission environment. To improve the compression ability of the conventional BCC, the SSLCC system utilizes the virtual source location information (VSLI) as a spatial cue parameter instead of the inter-channel level difference (ICLD) of the BCC system. Also the SSLCC system adopts enhanced pre/post processing algorithms to improve perceptual sound quality. Objective and subjective assessment results show that the proposed SSLCC system reveals better performance than the conventional BCC system.
Inseon Jang, Jeongil Seo, Seungkwon Beack, Kyeongok Kang, Han-Gil Moon
ACM Multimedia3