VLDB 2026 Research / reviewers in the wild / expert
Joon Byun
dblp:315/6297
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0002-5114-5459ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Quantization Noise Masking in Perceptual Neural Audio CoderabstractThis study investigates the implication of utilizing the psychoacoustic model (PAM) within the neural audio coder (NAC), specifically focusing on the masking of quantization noise. We introduce a novel training strategy to incorporate the PAM into the NAC more accurately. This method involves a discriminator that directly or indirectly measures the PAM loss. For the indirect measurement, a multi-scale STFT discriminator (MS-STFTD) is incorporated to introduce an auxiliary loss term in addition to the existing PAM loss. Conversely, for the direct measurement, we have designed a multi-scale PAM discriminator (MS-PAMD) that quantifies PAM-specific parameters. Experimental results show that adding the discriminator masks the quantization noise better than the previous NAC, and it obtains audio quality comparable to the commercial AAC in both objective and subjective scores. Seungmin Shin, Joon Byun, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
ICASSP | 2 |
| 2023 | A Perceptual Neural Audio Coder with a Mean-Scale HyperpriorabstractThis paper proposes an end-to-end neural audio coder based on a mean-scale hyperprior model together with a perceptual optimization using a psychoacoustic model (PAM)-based loss function. The proposed coder estimates the mean and scale hyperpriors using a sub-network after assuming that the probability distribution of latent samples is Gaussian. The main network is an autoencoder based on Resnet-type gated linear units (ResGLUs), each comprising a generalized divisive normalization (GDN) layer. We train both networks to optimize perceptual attributes estimated using a multi-timescale scheme to obtain high perceptual quality. Experimental results show that the proposed model accurately predicts the mean and scale hyperpriors. Also, it obtains consistently higher audio quality than the commercial MP3 audio coder at all bitrates. Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
ICASSP | 1 |
| 2023 | Perceptual Improvement of Deep Neural Network (DNN) Speech Coder Using Parametric and Non-parametric Density Models
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
INTERSPEECH | 1 |
| 2022 | Deep Neural Network (DNN) Audio Coder Using A Perceptually Improved Training MethodabstractA new end-to-end audio coder based on a deep neural network (DNN) is proposed. To compensate for the perceptual distortion that occurred by quantization, the proposed coder is optimized to minimize distortions in both signal and perceptual domains. The distortion in the perceptual domain is measured using the psychoacoustic model (PAM), and a loss function is obtained through the two-stage compensation approach. Also, the scalar uniform quantization was approximated using a uniform stochastic noise, together with a compression-decompression scheme, which provides simpler but more stable learning without an additional penalty than the softmax quantizer. Test results showed that the proposed coder achieves more accurate noise-masking than the previous PAM-based method and better perceptual quality then the MP3 audio coder. Seungmin Shin, Joon Byun, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
ICASSP | 2 |
| 2022 | Optimization of Deep Neural Network (DNN) Speech Coder Using a Multi Time Scale Perceptual Loss Function
Joon Byun, Seungmin Shin, Jongmo Sung, Seungkwon Beack, Young-Cheol Park |
INTERSPEECH | 1 |
| 2021 | Development of a Psychoacoustic Loss Function for the Deep Neural Network (DNN)-Based Speech Coder
Joon Byun, Seungmin Shin, Young-Cheol Park, Jongmo Sung, Seungkwon Beack |
Interspeech | 1 |