Inseon Jang

dblp:70/4879 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0003-2237-2668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2025 Neural Spectral Band Generation for Audio Coding
Woongjib Choi, Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
INTERSPEECH4
2025 Towards an Ultra-Low-Delay Neural Audio Coding with Computational Efficiency
Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
INTERSPEECH3
2024 Personalized Neural Speech Codec
abstract
In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is involved on each side of the communication, personalizing a codec for the specific user has rarely been explored in the literature. First, we assume speakers can be grouped into smaller subsets based on their perceptual similarity. Then, we also postulate that a group-specific codec can focus on the group’s speech characteristics to improve its perceptual quality and computational efficiency. To this end, we first develop a Siamese network that learns the speaker embeddings from the LibriSpeech dataset, which are then grouped into underlying speaker clusters. Finally, we retrain the LPCNet-based speech codec baselines on each of the speaker clusters. Subjective listening tests show that the proposed personalization scheme introduces model compression while maintaining speech quality. In other words, with the same model complexity, personalized codecs produce better speech quality.
Inseon Jang, Haici Yang, Wootaek Lim, Seungkwon Beack
ICASSP1
2024 Generative De-Quantization for Neural Speech Codec Via Latent Diffusion
abstract
End-to-end speech coding models achieve high coding gains by learning compact yet expressive features and a powerful decoder in a single network. A challenging problem as such results in unwelcome complexity increase and inferior speech quality. In this paper, we propose to separate the representation learning and information reconstruction tasks. We leverage an end-to-end codec for learning low-dimensional discrete tokens. Instead of using its decoder, we employ a latent diffusion model to de-quantize coded features into a high-dimensional continuous space, relieving the decoder’s burden of de-quantizing and upsampling. To mitigate the issue of over-smooth generation, we introduce midway-infilling with less noise reduction and stronger conditioning. We investigate the hyperparameters for midway-infilling and latent diffusion space with different dimensions in ablation studies. Subjective listening tests show that our model outperforms the state-of-the-art at two low bitrates, 1.5 and 3 kbps. We open-source the project for reproducibility1.
Haici Yang, Inseon Jang
ICASSP2
2023 Individual Sub-Band Estimation Approach to Bandwidth Extension and Enhancement of Coded Speech
abstract
The streaming Sound EnhAncement Network (SEANet) has demonstrated impressive performance for speech bandwidth extension (BWE) with low latency and computational complexity. Although the streaming SEANet was designed for voice communication systems, it was not tested with decoded signals that included coding artifacts. Our preliminary experiment showed that the output of the streaming SEANet for the decoded speech had room for improvement even if it was trained with decoded speeches, possibly because it should perform the BWE and the coded speech enhancement (CSE) at once. In this work, we propose to utilize two streaming SEANets in parallel, which are dedicated to the narrowband CSE and the generation of the upper band speech signal, respectively. Experimental results showed that the proposed model outperformed a bigger streaming SEANet trained to carry out both tasks in terms of the PESQ scores and the MUSHRA test.
Youngwon Choi, Eunkyun Lee, Inseon Jang, Jong Won Shin
ICASSP3
2023 Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator
abstract
In this paper, we improve the efficiency of the progressive multi-stage neural audio codec (PR-Codec) by utilizing perceptually motivated training criteria. Although our baseline PR-Codec successfully reconstructs full-band signals by progressively decoding the pre-defined subband signals, transparent quality can only be guaranteed in high bit-rates. To reduce bit-rates while maintaining perceptually transparent quality, we adopt a psychoacoustic model (PAM)-based loss and propose a perceptual weighting discriminator (PWD), which enables us to synthesize and discriminate audio signals in the perceptually motivated domain. We also introduce a scalar quantization with an entropy model to further enhance the quantization efficiency. Our experimental results show that our proposed model significantly improves perceptual reconstruction quality at the expense of the waveform disparity in the time-domain, compared to our previous model.
Byeong Hyeon Kim, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
ICASSP4
2023 End-to-End Neural Audio Coding in the MDCT Domain
abstract
Modern deep neural network (DNN)-based audio coding approaches utilize complicated non-linear functions (e.g., convolutional neural networks and non-linear activations), which leads to high complexity and memory usage. However, their decoded audio quality is still not much higher than that of signal processing-based legacy codecs. In this paper, we propose an effective frequency-domain neural audio coding paradigm that adopts the modified discrete cosine transform (MDCT) for analysis and synthesis and DNNs for the quantization of variables. It includes an efficient method to encode MDCT bins as well as a mechanism to adapt the quantization level of each bin. Our neural audio codec is trained in an end-to-end manner with the help of psychoacoustics-based perceptual loss, removing the burden of module-by-module fine-tuning. Experimental results show that our proposed model’s performance is comparable with the MP3 codec at around 64 and 48 kbps bit-rates for mono signals.
Hyungseob Lim, Byeong Hyeon Kim, Inseon Jang, Hong-Goo Kang
ICASSP4
2023 Native Multi-Band Audio Coding Within Hyper-Autoencoded Reconstruction Propagation Networks
abstract
Spectral sub-bands do not portray the same perceptual relevance. In audio coding, it is therefore desirable to have independent control over each of the constituent bands so that bitrate assignment and signal reconstruction can be achieved efficiently. In this work, we present a novel neural audio coding network that natively supports a multi-band coding paradigm. Our model extends the idea of compressed skip connections in the U-Net-based codec, allowing for independent control over both core and high band-specific reconstructions and bit allocation. Our system reconstructs the full-band signal mainly from the condensed core-band code, therefore exploiting and showcasing its bandwidth extension capabilities to its fullest. Meanwhile, the low-bitrate high-band code helps the high-band reconstruction similarly to MPEG audio codecs' spectral bandwidth replication. MUSHRA tests show that the proposed model not only improves the quality of the core band by explicitly assigning more bits to it but retains a good quality in the high-band as well.
Darius Petermann, Inseon Jang, Minje Kim 0001
ICASSP2
2022 Progressive Multi-Stage Neural Audio Coding with Guided References
abstract
In this paper, we propose an effective multi-stage neural audio coding algorithm that encodes full-band audio signals (up to 20 kHz) using an end-to-end training criterion. By predefining several dyadic subband signals as training targets, we progressively encode input audio signals in each stage such that deeper stages of the network encode the residual error terms from the previous encoding stage. Our proposed audio codec successfully decodes full-band audio signals by using an effective multi-stage vector quantization scheme to represent key encoding features extracted in the latent space. Subjective listening tests show that the decoded outputs of the proposed audio codec achieve almost transparent quality at an average bitrate of 132 kbps.
Chanwoo Lee, Hyungseob Lim, Inseon Jang, Hong-Goo Kang
ICASSP4
2022 Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator
abstract
In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each component independently. The rationale behind this idea is that conventional discriminators cannot reliably capture subtle distortions in audio signals, which have complicated time-frequency characteristics. By considering the time-frequency resolution of audio signals, our proposed method encourages the generator to better reconstruct harmonic and percussive features, both of which are critical for the quality of the generated signals. Listening tests show that our framework significantly enhances the stability of pitches and generates clearer piano samples compared to a baseline.
Hyungseob Lim, Chanwoo Lee, Inseon Jang, Hong-Goo Kang
ICASSP4
2022 Alias-and-Separate: Wideband Speech Coding Using Sub-Nyquist Sampling and Speech Separation
abstract
Decimation of a discrete-time signal below the Nyquist rate without applying an appropriate lowpass filter results in a distortion called aliasing. If wideband speech sampled at 16 kHz is decimated by 2 to result in a signal sampled at 8 kHz with aliasing, the decimated signal would be the summation of two speech-like signals, which are the narrowband speech covering 0-4 kHz and the spectrally flipped aliasing component coming from 8-4 kHz. Recently, the performance of speech separation has been remarkably improved with deep learning-based approaches, implying that the narrowband and aliasing components may be able to be separated. In this letter, we propose a novel method for low-rate wideband speech coding utilizing a standard narrowband codec. Instead of coding wideband speech using a wideband codec with a limited bitrate, we propose to decimate the input wideband speech incurring aliasing, and then encode it with a narrowband codec by allocating all the allowed bitrate to 0-4 kHz. After decoding the encoded bitstream, we apply a speech separation technique to obtain the narrowband and aliasing signals, which are then used to reconstruct the wideband speech by expansion, low/highpass filtering, and summation. Experimental results showed that the proposed method could achieve subjective quality comparable to the speeches coded by wideband codecs at higher bitrates in a subjective MUSHRA test.
Soojoong Hwang, Eunkyun Lee, Inseon Jang, Jong Won Shin
IEEE Signal Process. Lett.3
2021 Coded Speech Enhancement Using Neural Network-Based Vector-Quantized Residual Features
Youngju Cheon, Soojoong Hwang, Sangwook Han, Inseon Jang, Jong Won Shin
Interspeech4
2020 Emotional Speech Synthesis with Rich and Granularized Control
abstract
This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing the TTS input. We introduce an inter-to-intra emotional distance ratio algorithm to the embedding vectors that can minimize the distance to the target emotion category while maximizing its distance to the other emotion categories. To further enhance the expressiveness of a target speech, we also introduce an effective interpolation technique that enables the intensity of a target emotion to be gradually changed to that of neutral speech. Subjective evaluation results in terms of emotional expressiveness and controllability show the superiority of the proposed algorithm to the conventional methods.
Seyun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
ICASSP4
2019 An Effective Style Token Weight Control Technique for End-to-End Emotional Speech Synthesis
abstract
In this letter, we propose a high-quality emotional speech synthesis system, using emotional vector space, i.e., the weighted sum of global style tokens (GSTs). Our previous research verified the feasibility of GST-based emotional speech synthesis in an end-to-end text-to-speech synthesis framework. However, selecting appropriate reference audio (RA) signals to extract emotion embedding vectors to the specific types of target emotions remains problematic. To ameliorate the selection problem, we propose an effective way of generating emotion embedding vectors by utilizing the trained GSTs. By assuming that the trained GSTs represent an emotional vector space, we first investigate the distribution of all the training samples depending on the type of each emotion. We then regard the centroid of the distribution as an emotion-specific weighting value, which effectively controls the expressiveness of synthesized speech, even without using the RA for guidance, as it did before. Finally, we confirm that the proposed controlled weight-based method is superior to the conventional emotion label-based methods in terms of perceptual quality and emotion classification accuracy.
Ohsung Kwon, Inseon Jang, Chunghyun Ahn, Hong-Goo Kang
IEEE Signal Process. Lett.2
2017 Enhanced Feature Extraction for Speech Detection in Media Audio
Inseon Jang, Chunghyun Ahn, Jeongil Seo, Younseon Jang
INTERSPEECH1
2014 Semi-automatic DVS Authoring Method
Inseon Jang, Chunghyun Ahn, Younseon Jang
ICCHP (1)1
2014 Dynamic Subtitle Authoring Method Based on Audio Analysis for the Hearing Impaired
Wootaek Lim, Inseon Jang, Chunghyun Ahn
ICCHP (1)2
2005 Sound source location cue coding system for compact representation of multi-channel audio
abstract
Binaural cue coding (BCC) has been introduced for compact representation of multi-channel audio. It exploits binaural cue parameters for capturing the spatial image of multi-channel audio. Recently, it has been standardized within MPEG as the name of "MPEG Surround." In this paper, we propose a sound source location cue coding (SSLCC) system for compressing multi-channel audio to be suitable at the narrow bandwidth transmission environment. To improve the compression ability of the conventional BCC, the SSLCC system utilizes the virtual source location information (VSLI) as a spatial cue parameter instead of the inter-channel level difference (ICLD) of the BCC system. Also the SSLCC system adopts enhanced pre/post processing algorithms to improve perceptual sound quality. Objective and subjective assessment results show that the proposed SSLCC system reveals better performance than the conventional BCC system.
Inseon Jang, Jeongil Seo, Seungkwon Beack, Kyeongok Kang, Han-Gil Moon
ACM Multimedia1