Jae-Jin Jeon

dblp:01/5883 · also Jaejin Jeon · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 RoboLoc: A Benchmark Dataset for Point Place Recognition and Localization in Indoor-Outdoor Integrated Environments
abstract
ABSTRACT Robust place recognition is essential for reliable localization in robotics, particularly in complex environments with frequent indoor–outdoor transitions. However, existing LiDAR‐based datasets often focus on outdoor scenarios and lack seamless domain shifts. In this paper, we propose RoboLoc, a benchmark dataset designed for GPS‐free place recognition in indoor–outdoor environments with floor transitions. RoboLoc features real‐world robot trajectories, diverse elevation profiles, and transitions between structured indoor and unstructured outdoor domains. We benchmark a variety of state‐of‐the‐art models, point‐based, voxel‐based, and BEV‐based architectures, highlighting their generalizability domain shifts. RoboLoc provides a realistic testbed for developing multi‐domain localization systems in robotics and autonomous navigation.
Jae-Jin Jeon, Seonghoon Ryoo, Sang-duck Lee, Soomok Lee, Seungwoo Jeong
IET Image Process.1
2023 Knowledge Distillation From Offline to Streaming Transducer: Towards Accurate and Fast Streaming Model by Matching Alignments
abstract
Sequence transducer is a popular end-to-end automatic speech recognition model for streaming scenarios: While, there is a trade-off between accuracy and latency. Latency regularization methods such as FastEmit can reduce latency, but the more they try to reduce latency, the worse accuracy tends to be. Conversely, knowledge distillation (KD) is only used to improve accuracy, and latency is not considered. In this paper, we propose an effective method that combines FastEmit with the KD to reduce latency and improve the accuracy of offline model in scenarios where the latency gap between offline and streaming models gets small. This method reduce the latency gap by applying with FastEmit to both the offline and streaming models. Experimental results on the LibriSpeech dataset show that the model with the best trade-off between accuracy and latency achieves a relative error reduction rate of 7.5% and reduces the latency by $130 \mathrm{~ms}$ compared with the streaming conformer transducer.
Ji-Hwan Mo, Jae-Jin Jeon, Mun-Hak Lee, Joon-Hyuk Chang
ASRU2
2022 Automatic Pronunciation Assessment using Self-Supervised Speech Representation Learning
abstract
Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community.In particular, speech representations learned by SSL models have been shown to be effective for encoding various speech-related characteristics.In this context, we propose a novel automatic pronunciation assessment method based on SSL models.First, the proposed method fine-tunes the pre-trained SSL models with connectionist temporal classification to adapt the English pronunciation of English-as-a-second-language (ESL) learners in a data environment.Then, the layer-wise contextual representations are extracted from all across the transformer layers of the SSL models.Finally, the automatic pronunciation score is estimated using bidirectional long short-term memory with the layer-wise contextual representations and the corresponding text.We show that the proposed SSL model-based methods outperform the baselines, in terms of the Pearson correlation coefficient, on datasets of Korean ESL learner children and Speechocean762.Furthermore, we analyze how different representations of transformer layers in the SSL model affect the performance of the pronunciation assessment task.
Eesung Kim, Jae-Jin Jeon, Hyeji Seo
INTERSPEECH2
2021 Multitask Learning and Joint Optimization for Transformer-RNN-Transducer Speech Recognition
abstract
Recently, several types of end-to-end speech recognition methods named transformer-transducer were introduced. According to those kinds of methods, transcription networks are generally modeled by transformer-based neural networks, while prediction networks could be modeled by either transformers or recurrent neural networks (RNN). In this paper, we propose novel multitask learning, joint optimization, and joint decoding methods for transformer-RNN-transducer systems. Our proposed methods have the main advantage in that the model can maintain information on the large text corpus eliminating the necessity of an external language model (LM). We prove their effectiveness by performing experiments utilizing the well-known ESPNET toolkit for the widely used Librispeech datasets. We also show that the proposed methods can reduce word error rate (WER) by 16.6 % and 13.3 % for test-clean and test-other datasets, respectively, without changing the overall model structure nor exploiting an external LM.
Jae-Jin Jeon, Eesung Kim
ICASSP1
2021 U-Convolution Based Residual Echo Suppression with Multiple Encoders
abstract
In this paper, we propose an efficient end-to-end neural network that can estimate near-end speech using a U-convolution block by exploiting various signals to achieve residual echo suppression (RES). Specifically, the proposed model employs multiple encoders and an integration block to utilize complete signal information in an acoustic echo cancellation system and also applies the U-convolution blocks to separate near-end speech efficiently. The proposed network affords an improvement in the perceptual evaluation of speech quality (PESQ) and the short-time objective intelligibility (STOI), as compared to baselines, in scenarios involving smart audio devices. The experimental results show that the proposed method outperforms the baselines for various types of mismatched background noise and environmental reverberation, while requiring low computational resources.
Eesung Kim, Jae-Jin Jeon, Hyeji Seo
ICASSP2
2007 Adaptive step-size widely linear linearly constrained constant modulus algorithm for DS-CDMA receivers in nonstationary interference environments
Jun-Seok Lim, Ki-Young Han, Jae-Jin Jeon
Signal Process.3
2006 Recursive Complex Extreme Learning Machine with Widely Linear Processing for Nonlinear Channel Equalizer
Jun-Seok Lim, Jae-Jin Jeon
ISNN (2)2
2005 Adaptive widely linear minimum output energy algorithm for DS-CDMA systems
abstract
A novel blind widely linear (WL) minimum output energy (MOE) algorithm is proposed for the CDMA receiver. Whereas the conventional WL algorithms are only applicable to real-valued modulation, the proposed receiver is applicable to complex-valued modulation. The blind WL-MOE filter is analyzed, and the convergence properties of an adaptive implementation are derived. Simulation results confirm that the proposed algorithm has higher steady-state SINR and faster convergence speed than the conventional MOE algorithm but cannot outperform the conventional WL-MOE due to severer constraints.
Jae-Jin Jeon, Jeffrey G. Andrews, Koeng-Mo Sung
ICC1