Changheon Lee

dblp:339/2654 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%
Artificial intelligence
1 paper
Speech recognition and synthesis · 67% Representation and self-supervised learning · 33%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms › quantum computer architecture
quantum compilation
1.012026
D'ArQ: A QOC Framework with Causality-Aware Grouping and Basis Selection · HPCA 2026
Emerging computing paradigms
quantum computer architecture
1.012026
D'ArQ: A QOC Framework with Causality-Aware Grouping and Basis Selection · HPCA 2026
Emerging computing paradigms › quantum control
quantum optimal control
1.012026
D'ArQ: A QOC Framework with Causality-Aware Grouping and Basis Selection · HPCA 2026
Machine learning › Representation and self-supervised learning › representation learning
discrete representation learning
0.812024
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis · ICML 2024
Natural language and speech › Speech recognition and synthesis › speaker recognition
speaker embedding
0.812024
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis · ICML 2024
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.812024
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis · ICML 2024

Methods — techniques the papers use, named apart from their topics

random unitary initialization · 1.0heuristic cost model · 1.0GOAT algorithm · 1.0DAG-based grouping · 1.0speaker conditioning · 0.8feature discretization · 0.8
YearPublicationVenuePosition
2026 D'ArQ: A QOC Framework with Causality-Aware Grouping and Basis Selection
abstract
Quantum Optimal Control (QOC) frameworks are powerful tools for compiling quantum circuits into low-latency hardware control pulses, but recent studies suffer from two critical limitations: lengthy compilation times and potential logical inconsistencies from flawed gate grouping strategies. In this work, we introduce d'ArQ, a novel QOC framework that solves these challenges. (i) We identify and resolve the causality problem, a flaw in greedy partitioning that can produce invalid schedules, by introducing a DAG-based grouping algorithm with assigning mergeability to each group so that it guarantees logical correctness. (ii) To mitigate compilation times, we use a pre-computed library of pulses derived from random unitary matrices to provide a high-quality random initialization for pulse optimization. (iii) Diverging from prior work based on GRAPE, d'ArQ is built on the GOAT algorithm. We demonstrate that the choice of analytic basis is a critical hyperparameter and introduce a heuristic cost model to dynamically select the optimal basis for each synthesis task, improving pulse performance. When evaluated against the state-of-the-art baseline PAQOC on a realistic, inhomogeneous hardware model, d'ArQ demonstrates superior performance. Notably, d'ArQ reduces circuit latency up to 22.8% and compilation time up to 56.8%, establishing a more robust and physically realistic path for circuit compilation.
Changheon Lee, Hyungseok Kim 0003, Seungwoo Choi 0001, Youngmin Kim 0005, Won Woo Ro
HPCA1
2024 ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
abstract
In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker’s dataset. Although various works with similar purposes have been actively studied, their performance has not yet reached that of trained multi-speaker models due to their fundamental limitations. To overcome previous limitations, we propose effective methods for feature learning and representing target speakers’ speech characteristics by discretizing the features and conditioning them to a speech synthesis model. Our method obtained a significantly higher similarity mean opinion score (SMOS) in subjective similarity evaluation than seen speakers of a high-performance multi-speaker model, even with unseen speakers. The proposed method also outperforms a zero-shot method by significant margins. Furthermore, our method shows remarkable performance in generating new artificial speakers. In addition, we demonstrate that the encoded latent features are sufficiently informative to reconstruct an original speaker’s speech completely. It implies that our method can be used as a general methodology to encode and reconstruct speakers’ characteristics in various tasks.
Jungil Kong, Jeongmin Kim 0002, Beomjeong Kim, Dohee Kong, Changheon Lee
ICML7
2022 Conformer-Based on-Device Streaming Speech Recognition with KD Compression and Two-Pass Architecture
abstract
This paper introduces a two-pass on-device automatic speech recognition (ASR) system, which is developed for commercialized devices. The first pass of the system is based on a causal Conformer-transducer model to generate partial results from the input audio stream. After processing an entire input utterance in the first pass, the candidates for the final result are rescored with a full-context attention model in the second pass. To minimize the computational overhead from rescoring, we compress the full-context model by applying knowledge distillation (KD). The total model size is reduced by 35% after KD with a 0.02% absolute loss in word error rate (WER). We also introduce decoding techniques to boost the accuracy on the test cases mismatched with the distribution of the training set. The techniques include on-device personal adaptation, spell correction and handling incorrectly segmented speech, which solve the critical issues for production-grade systems. The whole system including the two-pass end-to-end (E2E) model and a language model (LM) occupies 72MB in storage after 8-bit quantization. We demonstrate the entire system on mobile devices and report results on test sets collected from the production environment. The developed system achieves 5.65% WER which surpasses the baseline system with 39% relative WER improvement.
Jinhwan Park, Sichen Jin, Junmo Park, Dhairya Sandhyana, Changheon Lee, Myoungji Han, Jungin Lee, Seokyeong Jung, Chanwoo Kim 0001
SLT6