Mi Suk Lee

dblp:65/4844 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
1since 2021 · last 2022
0000-0001-8951-5032ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › speech coding
neural speech codec
0.612022
Scalable and Efficient Neural Speech Coding: A Hybrid Design · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Audio and music processing › speech coding
scalable speech coding
0.612022
Scalable and Efficient Neural Speech Coding: A Hybrid Design · IEEE ACM Trans. Audio Speech Lang. Process. 2022
Audio and music processing
speech coding
0.612022
Scalable and Efficient Neural Speech Coding: A Hybrid Design · IEEE ACM Trans. Audio Speech Lang. Process. 2022

Methods — techniques the papers use, named apart from their topics

residual learning · 0.6linear predictive coding · 0.6gated residual networks · 0.6depthwise separable convolution · 0.6convolutional neural network · 0.6autoencoder · 0.6
YearPublicationVenuePosition
2022 Scalable and Efficient Neural Speech Coding: A Hybrid Design
abstract
We present a scalable and efficient neural waveform coding system for speech compression. We formulate the speech coding problem as an autoencoding task, where a convolutional neural network (CNN) performs encoding and decoding as a neural waveform codec (NWC) during its feedforward routine. The proposed NWC also defines quantization and entropy coding as a trainable module, so the coding artifacts and bitrate control are handled during the optimization process. We achieve efficiency by introducing compact model components to NWC, such as gated residual networks and depthwise separable convolution. Furthermore, the proposed models are with a scalable architecture, cross-module residual learning (CMRL), to cover a wide range of bitrates. To this end, we employ the residual coding concept to concatenate multiple NWC autoencoding modules, where each NWC module performs residual coding to restore any reconstruction loss that its preceding modules have created. CMRL can scale down to cover lower bitrates as well, for which it employs linear predictive coding (LPC) module as its first autoencoder. The hybrid design integrates LPC and NWC by redefining LPC’s quantization as a differentiable process, making the system training an end-to-end manner. The decoder of proposed system is with either one NWC (0.12 million parameters) in low to medium bitrate ranges (12 to 20 kbps) or two NWCs in the high bitrate (32 kbps). Although the decoding complexity is not yet as low as that of conventional speech codecs, it is significantly reduced from that of other neural speech coders, such as a WaveNet-based vocoder. For wide-band speech coding quality, our system yields comparable or superior performance to AMR-WB and Opus on TIMIT test utterances at low and medium bitrates. The proposed system can scale up to higher bitrates to achieve near transparent performance.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack, Minje Kim 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 A Dual-Staged Context Aggregation Method towards Efficient End-to-End Speech Enhancement
abstract
In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in the time domain without time-frequency transformation or mask estimation. However, aggregating contextual information from a high-resolution time domain signal with an affordable model complexity still remains challenging. In this paper, we propose a densely connected convolutional and recurrent network (DCCRN), a hybrid architecture, to enable dual-staged temporal context aggregation. With the dense connectivity and cross-component identical shortcut, DCCRN consistently outperforms competing convolutional baselines with an average STOI improvement of 0.23 and PESQ of 1.38 at three SNR levels. The proposed method is computationally efficient with only 1.38 million parameters. The generalizability performance on the unseen noise types is still decent considering its low complexity, although it is relatively weaker comparing to Wave-U-Net with 7.25 times more parameters.
Kai Zhen, Mi Suk Lee
ICASSP2
2020 Efficient and Scalable Neural Residual Waveform Coding with Collaborative Quantization
abstract
Scalability and efficiency are desired in neural speech codecs, which supports a wide range of bitrates for applications on various devices. We propose a collaborative quantization (CQ) scheme to jointly learn the codebook of LPC coefficients and the corresponding residuals. CQ does not simply shoehorn LPC to a neural network, but bridges the computational capacity of advanced neural network models and traditional, yet efficient and domain-specific digital signal processing methods in an integrated manner. We demonstrate that CQ achieves much higher quality than its predecessor at 9 kbps with even lower model complexity. We also show that CQ can scale up to 24 kbps where it outperforms AMR-WB and Opus. As a neural waveform codec, CQ models are with less than 1 million parameters, significantly less than many other generative models.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack
ICASSP2
2020 Psychoacoustic Calibration of Loss Functions for Efficient End-to-End Neural Audio Coding
abstract
Conventional audio coding technologies commonly leverage human perception of sound, or psychoacoustics, to reduce the bitrate while preserving the perceptual quality of the decoded audio signals. For neural audio codecs, however, the objective nature of the loss function usually leads to suboptimal sound quality as well as high run-time complexity due to the large model size. In this work, we present a psychoacoustic calibration scheme to re-define the loss functions of neural audio coding systems so that it can decode signals more perceptually similar to the reference, yet with a much lower model complexity. The proposed loss function incorporates the global masking threshold, allowing the reconstruction error that corresponds to inaudible artifacts. Experimental results show that the proposed model outperforms the baseline neural codec twice as large and consuming 23.4% more bits per second. With the proposed method, a lightweight neural codec, with only 0.9 million parameters, performs near-transparent audio coding comparable with the commercial MPEG-1 Audio Layer III codec at 112 kbps.
Kai Zhen, Mi Suk Lee, Jongmo Sung, Seungkwon Beack, Minje Kim 0001
IEEE Signal Process. Lett.2
2019 Cascaded Cross-Module Residual Learning Towards Lightweight End-to-End Speech Coding
abstract
Speech codecs learn compact representations of speech signals to facilitate data transmission.Many recent deep neural network (DNN) based end-to-end speech codecs achieve low bitrates and high perceptual quality at the cost of model complexity.We propose a cross-module residual learning (CMRL) pipeline as a module carrier with each module reconstructing the residual from its preceding modules.CMRL differs from other DNN-based speech codecs, in that rather than modeling speech compression problem in a single large neural network, it optimizes a series of less-complicated modules in a two-phase training scheme.The proposed method shows better objective performance than AMR-WB and the state-of-the-art DNNbased speech codec with a similar network architecture.As an end-to-end model, it takes raw PCM signals as an input, but is also compatible with linear predictive coding (LPC), showing better subjective quality at high bitrates than AMR-WB and OPUS.The gain is achieved by using only 0.9 million trainable parameters, a significantly less complex architecture than the other DNN-based codecs in the literature.
Kai Zhen, Jongmo Sung, Mi Suk Lee, Seungkwon Beack
INTERSPEECH3
2010 Superwideband extension of g.718 and g.729.1 speech codecs
abstract
This communication presents the recently standardized superwideband (SWB) extensions of ITU-T G.718 and G.729.1. These extensions were standardized as G.718 annex B and G.729.1 annex E. The SWB functionality is implemented using embedded scalable layers on top of the wideband (WB) core codecs, and it extends the bit rate of the codecs to 48 and 64 kbit/s for the G.718 and G.729.1, respectively. The main technology is a two-mode SWB coding method of the high frequencies. In addition, the G.729.1 SWB extension enhances the lower frequency range. The codec performance is illustrated with some listening test results extracted from the ITU-T Characterization phase.
Lasse Laaksonen, Mikko Tammi, Vladimir Malenovsky, Tommy Vaillancourt, Mi Suk Lee, Tomofumi Yamanashi, Masahiro Oshikiri, Claude Lamblin, Balázs Kövesi, Lei Miao 0004, Deming Zhang, Jon Gibbs, Holly Francois
INTERSPEECH5
2007 An 8-12 Kbit/S Embedded CELP Coder Interoperable with ITU-T G.729 CIDER: First Stage of the New G.729.1 Standard
abstract
ITU-T G.729.1 is a scalable coder recently standardized in ITU-T for wideband telephony and voice over IP (VoIP) applications. Composed of three stages, this codec provides a scalable bitstream between 8 and 32 kbit/s both in narrowband and wideband. This paper describes the first stage which is a narrowband embedded CELP coder at 8 and 12 kbit/s. The 8 kbit/s layer ensures interoperability with ITU-T G.729 standard with a reduced complexity, and with a quality better than G.729 Annex A. At 12 kbit/s, G.729.1 reaches the quality level of the 11.8 kbit/s G.729 Annex E in spite of the embedded structure. The modifications brought to the original G.729 scheme to achieve this performance are explained and formal test results provided.
Dominique Massaloux, Romain Trilling, Claude Lamblin, Stéphane Ragot, Hiroyuki Ehara, Mi Suk Lee, Do-Young Kim, Bruno Bessette
ICASSP (4)6
2007 ITU-T G.729.1: AN 8-32 Kbit/S Scalable Coder Interoperable with G.729 for Wideband Telephony and Voice Over IP
abstract
This paper describes the scalable coder - G.729.1 - which has been recently standardized by ITU-T for wideband telephony and voice over IP (VoIP) applications. G.729.1 can operate at 12 different bit rates from 32 down to 8 kbit/s with wideband quality starting at 14 kbit/s. This coder is a bitstream interoperable extension of ITU-T G.729 based on three embedded stages: narrowband cascaded CELP coding at 8 and 12 kbit/s, time-domain bandwidth extension (TDBWE) at 14 kbit/s, and split-band MDCT coding with spherical vector quantization (VQ) and pre-echo reduction from 16 to 32 kbit/s. Side information - consisting of signal class, phase, and energy - is transmitted at 12, 14 and 16 kbit/s to improve the resilience and recovery of the decoder in case of frame erasures. The quality, delay, and complexity of G.729.1 are summarized based on ITU-T results.
Stéphane Ragot, Balázs Kövesi, Romain Trilling, David Virette, Nicolas Duc, Dominique Massaloux, Stéphane Proust, Bernd Geiser, Martin Gartner, Stefan Schandl, Hervé Taddei, Eyal Shlomot, Hiroyuki Ehara, Koji Yoshida, Tommy Vaillancourt, Redwan Salami, Mi Suk Lee, Do-Young Kim
ICASSP (4)18
2001 A new distortion measure for spectral quantization based on the LSF intermodel interlacing property
Mi Suk Lee, Hong Kook Kim, Hwang Soo Lee
Speech Commun.1
1999 A 4 kbps adaptive fixed code-excited linear prediction speech coder
abstract
We propose an adaptive fixed code-excited linear prediction (AF-CELP) speech coder operating at 4 kbps. By exploiting the fact that a fixed codebook contribution to the speech signal is also periodic as the corresponding adaptive codebook contribution, the adaptive fixed codebook model efficiently represents excitation signals. In order to overcome the quality degradation caused by the coarse quantization of excitation, a paired pulse algebraic codebook structure is also applied to the excitation model. Additionally, a pitch prefiltering, a noise spreading, and a harmonic enhancement technique are adopted in the decoding process. The spectrogram reading and informal listening tests proved that the AF-CELP reproduces high quality speech.
Hong Kook Kim, Mi Suk Lee, Hwang Soo Lee
ICASSP2