VLDB 2026 Research / reviewers in the wild / expert
Joon-Young Yang
dblp:263/4921
· DBLP profile ↗
12ranked-venue papers
4as first author
10since 2021 · last 2024
0000-0003-0096-4371ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Learning Flexible Body Collision Dynamics with Hierarchical Contact Mesh TransformerabstractRecently, many mesh-based graph neural network (GNN) models have been proposed for modeling complex high-dimensional physical systems. Remarkable achievements have been made in significantly reducing the solving time compared to traditional numerical solvers. These methods are typically designed to i) reduce the computational cost in solving physical dynamics and/or ii) propose techniques to enhance the solution accuracy in fluid and rigid body dynamics. However, it remains under-explored whether they are effective in addressing the challenges of flexible body dynamics, where instantaneous collisions occur within a very short timeframe. In this paper, we present Hierarchical Contact Mesh Transformer (HCMT), which uses hierarchical mesh structures and can learn long-range dependencies (occurred by collisions) among spatially distant positions of a body --- two close positions in a higher-level mesh correspond to two distant positions in a lower-level mesh. HCMT enables long-range interactions, and the hierarchical mesh structure quickly propagates collision effects to faraway positions. To this end, it consists of a contact mesh Transformer and a hierarchical mesh Transformer (CMT and HMT, respectively). Lastly, we propose a flexible body dynamics dataset, consisting of trajectories that reflect experimental settings frequently used in the display industry for product designs. We also compare the performance of several baselines using well-known benchmark datasets. Our results show that HCMT provides significant performance improvements over existing methods. Our code is available at https://github.com/yuyudeep/hcmt. Youn-Yeol Yu, Jeongwhan Choi 0002, Kookjin Lee, Nayong Kim, Kiseok Chang, ChangSeung Woo, Ilho Kim, SeokWoo Lee, Joon-Young Yang, Sooyoung Yoon, Noseong Park |
ICLR | 10 |
| 2024 | Relational Proxy Loss for Audio-Text based Keyword Spotting
Youngmoon Jung, Joon-Young Yang, Jaeyoung Roh, Chang Woo Han, Hoonyoung Cho |
INTERSPEECH | 3 |
| 2024 | Efficient Lightweight Speaker Verification With Broadcasting CNN-Transformer and Knowledge Distillation Training of Self-Attention MapsabstractDeveloping a lightweight speaker embedding extractor (SEE) is crucial for the practical implementation of automatic speaker verification (ASV) systems. To this end, we recently introducedbroadcasting convolutional neural networks (CNNs)-meet-vision-Transformers(BC-CMT), a lightweight SEE that utilizes broadcasted residual learning (BRL) within the hybrid CNN-Transformer architecture to maintain a small number of model parameters. We proposed three BC-CMT-based SEE with three different sizes: BC-CMT-Tiny, -Small, and -Base. In this study, we extend our previously proposed BC-CMT by introducing an improved model architecture and a training strategy based on knowledge distillation (KD) using self-attention (SA) maps. First, to reduce the computational costs and latency of the BC-CMT, the two-dimensional (2D) SA operations in the BC-CMT, which calculate the SA maps in the frequency–time dimensions, are simplified to 1D SA operations that consider only temporal importance. Moreover, to enhance the SA capability of the BC-CMT, the group convolution layers in the SA block are adjusted to have smaller number of groups and are combined with the BRL operations. Second, to improve the training effectiveness of the modified BC-CMT-Tiny, the SA maps of a pretrained large BC-CMT-Base are used for the KD to guide those of a smaller BC-CMT-Tiny. Because the attention map sizes of the modified BC-CMT models do not depend on the number of frequency bins or convolution channels, the proposed strategy enables KD between feature maps with different sizes. The experimental results demonstrate that the proposed BC-CMT-Tiny model having 271.44K model parameters achieved 36.8% and 9.3% reduction in floating point operations on 1s signals and equal error rate (EER) on VoxCeleb 1 testset, respectively, compared to the conventional BC-CMT-Tiny. The CPU and GPU running time of the proposed BC-CMT-Tiny ranges of 1 to 10 s signals were 29.07 to 146.32 ms and 36.01 to 206.43 ms, respectively. The proposed KD further reduced the EER by 15.5% with improved attention capability. Jeong-Hwan Choi, Joon-Young Yang, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Improving Transformer-Based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention HeadsabstractTransformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training effectiveness of SA-EEND models, we propose the use of auxiliary losses for the SA heads of the transformer layers. Specifically, we assume that the attention weight matrices of an SA layer are redundant if their patterns are similar to those of the identity matrix. We then explicitly constrain such matrices to exhibit specific speaker activity patterns relevant to voice activity detection or overlapped speech detection tasks. Consequently, we expect the proposed auxiliary losses to guide the transformer layers to exhibit more diverse patterns in the attention weights, thereby reducing the assumed redundancies in the SA heads. The effectiveness of the proposed method is demonstrated using the simulated and CALLHOME datasets for two-speaker diarization tasks, reducing the diarization error rate of the conventional SA-EEND model by 32.58% and 17.11%, respectively. Ye-Rin Jeoung, Joon-Young Yang, Jeong-Hwan Choi, Joon-Hyuk Chang |
ICASSP | 2 |
| 2023 | Deeply Supervised Curriculum Learning for Deep Neural Network-based Sound Source Localization
Min-Sang Baek, Joon-Young Yang, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | Improved CNN-Transformer using Broadcasted Residual Learning for Text-Independent Speaker Verification
Jeong-Hwan Choi, Joon-Young Yang, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | HYU Submission for the SASV Challenge 2022: Reforming Speaker Embeddings with Spoofing-Aware Conditioning
Jeong-Hwan Choi, Joon-Young Yang, Ye-Rin Jeoung, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2022 | VACE-WPE: Virtual Acoustic Channel Expansion Based on Neural Networks for Weighted Prediction Error-Based Speech DereverberationabstractSpeech dereverberation is an important issue for many real-world speech processing applications. Among the techniques developed, the weighted prediction error (WPE) algorithm has been widely adopted and advanced over the last decade, which blindly cancels out the late reverberation component from the reverberant mixture of microphone signals. In this study, we extend the neural-network-based virtual acoustic channel expansion (VACE) framework for the WPE-based speech dereverberation, a variant of the WPE that we recently proposed to enable the use of dual-channel WPE algorithm in a single-microphone speech dereverberation scenario. Based on the previous study, some ablation studies are conducted regarding the constituents of the VACE-WPE in an offline processing scenario. These studies reveal the characteristics of the system, thereby simplifying the architecture and leading to the introduction of new strategies for training the neural network for the VACE. Experimental results demonstrate that VACE-WPE (our PyTorch implementation and pre-trained models are available fromhttps://github.com/dreadbird06/vace_wpe) considerably outperforms its single-channel counterpart in simulated noisy reverberant environments in terms of objective speech quality and is superior to the single-channel WPE as well as several fully neural speech dereverberation methods when employed as the front-end for the far-field automatic speech recognizer. Joon-Young Yang, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Task-Specific Optimization of Virtual Channel Linear Prediction-Based Speech Dereverberation Front-End for Far-Field Speaker VerificationabstractDeveloping a single-microphone speech denoising or dereverberation front-end for robust automatic speaker verification (ASV) in noisy far-field speaking scenarios is challenging. To address this problem, we present a novel front-end design that involves a recently proposed extension of the weighted prediction error (WPE) speech dereverberation algorithm, the virtual acoustic channel expansion (VACE)-WPE. It is demonstrated experimentally in this study that unlike the conventional WPE algorithm, the VACE-WPE can be explicitly trained to cancel out both late reverberation and background noise. To build the front-end, the VACE-WPE is first (pre)trained to preserve the noise components in the input signals and produce “noisy” dereverberated output signals, thus making the front-end to be inductively biased to preserve as much noise components as possible and perform dereverberation only. Subsequently, given a pretrained speaker embedding model, the VACE-WPE is additionally fine-tuned within a task-specific optimization (TSO) framework, causing the speaker embedding extracted from the processed signal to be similar to that extracted from the “noise-free” target signal. Consequently, the front-end is optimized not to perform unnecessarily excessive denoising, thus achieving “generally safe” dereverberation and denoising for far-field ASV. Moreover, to prevent the front-end from adversely affecting the unconstrained “in-the-wild” ASV performance under more general, non-far-field conditions, we propose a distortion regularization method within the TSO framework. The effectiveness of the proposed approach is verified on both far-field and in-the-wild ASV benchmarks, demonstrating its superiority over fully neural front-ends and other TSO methods in various cases. Joon-Young Yang, Joon-Hyuk Chang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Short-Utterance Embedding Enhancement Method Based on Time Series Forecasting Technique for Text-Independent Speaker VerificationabstractShort-utterance embedding, which is a speaker embedding extracted from a short utterance, shows poor speaker verification performance due to insufficient speaker information. To address the problem, we propose a method to map the set of short-utterance embeddings to a set of long-utterance embeddings based on a neural network. Specifically, a speech utterance is cropped into multiple segments whose durations are gradually increasing, and the speaker embeddings are extracted from the sequence of cropped segments using a pre-trained speaker embedding extractor. Subsequently, the sequence of embeddings is divided into a group of short-utterances embeddings and that of long-utterance embeddings. In our method, a sequence-to-sequence model based forecasting technique is employed, where an encoder transforms the group of short-utterance embeddings to a fixed-dimensional vector, and then a decoder converts the vector into a group of long-utterance embeddings. Experimental results on the VoxCeleb and Speakers in the Wild datasets show that our method improves the text-independent speaker verification performance under short utterance condition. Jeong-Hwan Choi, Joon-Young Yang, Joon-Hyuk Chang |
ASRU | 2 |
| 2020 | Virtual Acoustic Channel Expansion Based on Neural Networks for Weighted Prediction Error-Based Speech Dereverberation
Joon-Young Yang, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2019 | Joint Optimization of Neural Acoustic Beamforming and Dereverberation with x-Vectors for Robust Speaker Verification
Joon-Young Yang, Joon-Hyuk Chang |
INTERSPEECH | 1 |