VLDB 2026 Research / reviewers in the wild / expert
Jongmo Sung
dblp:90/8793
· DBLP profile ↗
11ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-3396-1137ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Quantization Noise Masking in Perceptual Neural Audio CoderabstractThis study investigates the implication of utilizing the psychoacoustic model (PAM) within the neural audio coder (NAC), specifically focusing on the masking of quantization noise. We introduce a novel training strategy to incorporate the PAM into the NAC more accurately. This method involves a discriminator that directly or indirectly measures the PAM loss. For the indirect measurement, a multi-scale STFT discriminator (MS-STFTD) is incorporated to introduce an auxiliary loss term in addition to the existing PAM loss. Conversely, for the direct measurement, we have designed a multi-scale PAM discriminator (MS-PAMD) that quantifies PAM-specific parameters. Experimental results show that adding the discriminator masks the quantization noise better than the previous NAC, and it obtains audio quality comparable to the commercial AAC in both objective and subjective scores. Seungmin Shin, Joon Byun, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
ICASSP | 3 |
| 2023 | A Perceptual Neural Audio Coder with a Mean-Scale HyperpriorabstractThis paper proposes an end-to-end neural audio coder based on a mean-scale hyperprior model together with a perceptual optimization using a psychoacoustic model (PAM)-based loss function. The proposed coder estimates the mean and scale hyperpriors using a sub-network after assuming that the probability distribution of latent samples is Gaussian. The main network is an autoencoder based on Resnet-type gated linear units (ResGLUs), each comprising a generalized divisive normalization (GDN) layer. We train both networks to optimize perceptual attributes estimated using a multi-timescale scheme to obtain high perceptual quality. Experimental results show that the proposed model accurately predicts the mean and scale hyperpriors. Also, it obtains consistently higher audio quality than the commercial MP3 audio coder at all bitrates. Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
ICASSP | 4 |
| 2023 | Perceptual Improvement of Deep Neural Network (DNN) Speech Coder Using Parametric and Non-parametric Density Models
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
INTERSPEECH | 3 |
| 2022 | Deep Neural Network (DNN) Audio Coder Using A Perceptually Improved Training MethodabstractA new end-to-end audio coder based on a deep neural network (DNN) is proposed. To compensate for the perceptual distortion that occurred by quantization, the proposed coder is optimized to minimize distortions in both signal and perceptual domains. The distortion in the perceptual domain is measured using the psychoacoustic model (PAM), and a loss function is obtained through the two-stage compensation approach. Also, the scalar uniform quantization was approximated using a uniform stochastic noise, together with a compression-decompression scheme, which provides simpler but more stable learning without an additional penalty than the softmax quantizer. Test results showed that the proposed coder achieves more accurate noise-masking than the previous PAM-based method and better perceptual quality then the MP3 audio coder. Seungmin Shin, Joon Byun, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
ICASSP | 4 |
| 2022 | Optimization of Deep Neural Network (DNN) Speech Coder Using a Multi Time Scale Perceptual Loss Function
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
INTERSPEECH | 3 |
| 2022 | Scalable and Efficient Neural Speech Coding: A Hybrid DesignabstractWe present a scalable and efficient neural waveform coding system for speech compression. We formulate the speech coding problem as an autoencoding task, where a convolutional neural network (CNN) performs encoding and decoding as a neural waveform codec (NWC) during its feedforward routine. The proposed NWC also defines quantization and entropy coding as a trainable module, so the coding artifacts and bitrate control are handled during the optimization process. We achieve efficiency by introducing compact model components to NWC, such as gated residual networks and depthwise separable convolution. Furthermore, the proposed models are with a scalable architecture, cross-module residual learning (CMRL), to cover a wide range of bitrates. To this end, we employ the residual coding concept to concatenate multiple NWC autoencoding modules, where each NWC module performs residual coding to restore any reconstruction loss that its preceding modules have created. CMRL can scale down to cover lower bitrates as well, for which it employs linear predictive coding (LPC) module as its first autoencoder. The hybrid design integrates LPC and NWC by redefining LPC’s quantization as a differentiable process, making the system training an end-to-end manner. The decoder of proposed system is with either one NWC (0.12 million parameters) in low to medium bitrate ranges (12 to 20 kbps) or two NWCs in the high bitrate (32 kbps). Although the decoding complexity is not yet as low as that of conventional speech codecs, it is significantly reduced from that of other neural speech coders, such as a WaveNet-based vocoder. For wide-band speech coding quality, our system yields comparable or superior performance to AMR-WB and Opus on TIMIT test utterances at low and medium bitrates. The proposed system can scale up to higher bitrates to achieve near transparent performance. Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, Minje Kim 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Development of a Psychoacoustic Loss Function for the Deep Neural Network (DNN)-Based Speech Coder
Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
Interspeech | 4 |
| 2020 | Efficient and Scalable Neural Residual Waveform Coding with Collaborative QuantizationabstractScalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC coefficients and the corresponding residuals. CQ does not simply shoehorn LPC to a neural network, but bridges the computational capacity of advanced neural network models and traditional, yet efficient and domain-specific digital signal processing methods in an integrated manner. We demonstrate that CQ achieves much higher quality than its predecessor at 9 kbps with even lower model complexity. We also show that CQ can scale up to 24 kbps where it outperforms AMR-WB and Opus. As a neural waveform codec, CQ models are with less than 1 million parameters, significantly less than many other generative models. Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack |
ICASSP | 3 |
| 2020 | Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio CodingabstractConventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs, however, the objective nature of the loss function usually leads to suboptimal sound quality as well as high run-time complexity due to the large model size. In this work, we present a psychoacoustic calibration scheme to re-define the loss functions of neural audio coding systems so that it can decode signals more perceptually similar to the reference, yet with a much lower model complexity. The proposed loss function incorporates the global masking threshold, allowing the reconstruction error that corresponds to inaudible artifacts. Experimental results show that the proposed model outperforms the baseline neural codec twice as large and consuming 23.4% more bits per second. With the proposed method, a lightweight neural codec, with only 0.9 million parameters, performs near-transparent audio coding comparable with the commercial MPEG-1 Audio Layer III codec at 112 kbps. Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack, Minje Kim 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Cascaded Cross-Module Residual Learning Towards Lightweight End-to-End Speech CodingabstractSpeech codecs learn compact representations of speech signals to facilitate data transmission.Many recent deep neural network (DNN) based end-to-end speech codecs achieve low bitrates and high perceptual quality at the cost of model complexity.We propose a cross-module residual learning (CMRL) pipeline as a module carrier with each module reconstructing the residual from its preceding modules.CMRL differs from other DNN-based speech codecs, in that rather than modeling speech compression problem in a single large neural network, it optimizes a series of less-complicated modules in a two-phase training scheme.The proposed method shows better objective performance than AMR-WB and the state-of-the-art DNNbased speech codec with a similar network architecture.As an end-to-end model, it takes raw PCM signals as an input, but is also compatible with linear predictive coding (LPC), showing better subjective quality at high bitrates than AMR-WB and OPUS.The gain is achieved by using only 0.9 million trainable parameters, a significantly less complex architecture than the other DNN-based codecs in the literature. Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack |
INTERSPEECH | 2 |
| 2011 | G.711.1 Annex D and G.722 Annex B - New ITU-T superwideband codecsabstractThis paper presents high quality monaural superwideband extensions to G.711.1 and G.722, recently standardized as Recommendations ITU-T G.711.1 Annex D and G.722 Annex B. The superwideband (50-14000 Hz) functionality is achieved using embedded scalable structure that adds extension layers on top of the wideband core codecs. The bit rates are extended to 96/112/128 and 64/80/96 kbit/s for G.711.1 and G.722, respectively. The main technologies include lower and higher band (0-4 kHz and 4-8 kHz) enhancements, 8-14 kHz bandwidth extension and transform coding based on algebraic vector quantization. The codecs' performance is illustrated with listening test results extracted from formal ITU-T Characterization tests. Lei Miao 0004, Zexin Liu, Vaclav Eksler, Stéphane Ragot, Claude Lamblin, Balázs Kövesi, Jongmo Sung, Masahiro Fukui, Shigeaki Sasaki, Yusuke Hiwasaki |
ICASSP | 8 |