Jaeuk Lee

dblp:170/1373 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0003-6038-9839ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Static to Interactive: Authoring Interactive Visualizations via Natural Language
Can Liu 0003, Jaeuk Lee, Tianhe Chen, Zhibang Jiang, Xiaolin Wen, Yong Wang 0021
PacificVis2
2026 A 6-bit 3.6-GS/s 4-Channel Time-Interleaved ADC With Front-Rank MSB Decision and 2-Then-1.5-bit/Cycle Architecture
abstract
This article presents a 6-bit, 3.6-GS/s, 4-channel time-interleaved (TI) analog-to-digital converter (ADC) with a front-rank most significant bit (MSB) decision (FRMD) technique and a 2-then−1.5-bit/cycle successive approximation register (SAR) architecture. The proposed FRMD technique determines the MSB concurrently with sampling, enhancing the conversion speed of the sub-ADCs. The proposed 2-then−1.5-bit/cycle architecture, which aligns well with the FRMD technique, incorporates comparator background offset calibration without consuming extra cycles, zeroing the offset. Since the comparator offset is the dominant source of inter-channel offset mismatch in TI-ADCs, this approach naturally eliminates the need for extra inter-channel offset calibration. A prototype ADC was fabricated in a 28-nm CMOS process, achieving 33.9 dB signal-to-noise-and-distortion ratio (SNDR) and 49.8 dB spurious-free dynamic range (SFDR) at the Nyquist frequency. The ADC consumes 6.15 mW at 3.6 GS/s, resulting in a Walden figure of merit (FoMW) of 41.8 fJ/conversion-step.
Changjoo Kim, Sooho Park, Minkyun Shim, Yohan Choi, Donghwi Seo, Jaeuk Lee, Chulwoo Kim
IEEE Trans. Circuits Syst. I Regul. Pap.8
2025 Vector Field Decomposition-Based Flow Matching for Zero-Shot Cross-Lingual Text-to-Speech
abstract
Zero-shot text-to-speech (TTS) has recently achieved remarkable performance by leveraging a speech prompt instead of a speaker embedding, as it provides richer information. However, zero-shot cross-lingual tasks synthesize speech in multiple languages according to a given language ID, regardless of the language of the speech prompt. Consequently, the inherent language-specific characteristics of the speech prompt may conflict with the language ID, potentially affecting the accuracy of language representation in speech. Thus, we propose vector field decomposition-based flow matching that decomposes the vector field into speaker and language components. These components are trained to be activated in different frequency bins, as speaker and language identity are distributed across distinct frequency ranges in speech. This approach is particularly effective for cross-lingual TTS, as it minimizes conflicts between speech prompts and language IDs. As a result, the summation of the two components directly forms the vector field that represents the probability path from a Gaussian distribution to the target data distribution (e.g., mel spectrogram). Experimental results demonstrate that the proposed method outperforms the conventional method in terms of both subjective and objective evaluations.
Jaeuk Lee, Nam-Seok Song, Joon-Hyuk Chang
IEEE Signal Process. Lett.1
2025 Tokenized Generative Speech Enhancement With Language Model and Flow Matching
abstract
We propose a novel generative speech enhancement (SE) framework that integrates a language model (LM) and a flow-matching model. To utilize an LM with discrete tokens, we introduce dMel, which discretizes Mel spectrograms into a predefined set of quantized values on a linear-scale without requiring additional neural networks. dMel preserves both semantic and acoustic characteristics, providing a compact and effective token-based alternative to Mel spectrograms. We design the first encoder-decoder LM for SE, which learns to map noisy dMel to enhanced ones. Subsequently, flow-matching de-quantizes enhanced dMel into continuous representation and refines it by learning the optimal transport-based probability path, improving perceptual quality. This unified approach enables structured reconstruction while effectively suppressing noise. Experimental results demonstrate the effectiveness of our method in enhancing speech quality, establishing a new paradigm for generative SE without reliance on neural codec-based representations.
Da-Hee Yang, Jaeuk Lee, Joon-Hyuk Chang
IEEE Signal Process. Lett.2
2024 Neural ATSM: Fully Neural Network-based Adaptive Time-Scale Modification Using Sentence-Specific Dynamic Control
Jaeuk Lee, Sohee Jang, Joon-Hyuk Chang
INTERSPEECH1
2024 Differentiable Duration Refinement Using Internal Division for Non-Autoregressive Text-to-Speech
abstract
Most non-autoregressive text-to-speech (TTS) models acquire target phoneme duration (target duration) from internal or external aligners. They transform the speech-phoneme alignment produced by the aligner into the target duration. Since this transformation is not differentiable, the gradient of the loss function that maximizes the TTS model's likelihood of speech (e.g., mel spectrogram or waveform) cannot be propagated to the target duration. In other words, the target duration is produced regardless of the TTS model's likelihood of speech. Hence, we introduce a differentiable duration refinement that produces a learnable target duration for maximizing the likelihood of speech. The proposed method uses an internal division to locate the phoneme boundary, which is determined to improve the performance of the TTS model. Additionally, we propose a duration distribution loss to enhance the performance of the duration predictor. Our baseline model is JETS, a representative end-to-end TTS model, and we apply the proposed methods to the baseline model. Experimental results show that the proposed method outperforms the baseline model in terms of subjective naturalness and character error rate.
Jaeuk Lee, Yoonsoo Shin, Joon-Hyuk Chang
IEEE Signal Process. Lett.1
2022 One-Shot Speaker Adaptation Based on Initialization by Generative Adversarial Networks for TTS
Jaeuk Lee, Joon-Hyuk Chang
INTERSPEECH1
2022 Advanced Speaker Embedding with Predictive Variance of Gaussian Distribution for Speaker Adaptation in TTS
Jaeuk Lee, Joon-Hyuk Chang
INTERSPEECH1