Kyu J. Han

dblp:375/3541 · DBLP profile ↗
← Back
15ranked-venue papers
0as first author
15since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 13 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021
YearPublicationVenuePosition
2026 LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding
abstract
Zhivar Sourati, Zheng Wang, Marianne Menglin Liu, Yazhe Hu, Mengqing Guo, Sujeeth Bharadwaj, Kyu J. Han, Tao Sheng, Sujith Ravi, Morteza Dehghani, Dan Roth. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhivar Sourati, Marianne Menglin Liu, Yazhe Hu, Mengqing Guo, Sujeeth Bharadwaj, Kyu J. Han, Sujith Ravi, Morteza Dehghani, Dan Roth 0001
ACL (1)7
2025 Hyper-adapter for Parameter-Efficient Multilingual ASR Adaptation
abstract
This work proposes a new parameter-efficient adaptation approach for multilingual ASR based on the hyper-network. Existing multilingual ASR adaptation methods apply either one residual adapter for all the languages, or language dependent adapters for each individual language. The residual adapter cannot compete the full finetuning in terms of WER, because it is agnostic to language information. Whereas the language dependent adapters introduce high parameter overhead without a parameter sharing strategy. In contrast, we leverage a hyper-network to generate the weights for the adapters across different languages. To achieve the best parameter sharing strategy that scales with a large number of languages, we propose multi-level conditioning vector fusion, orthogonal regularization to improve the hyper-network output diversity, and language loss weighting during the model training. The proposed approach demonstrates comparable or better WER and better parameter efficiency compared to previous multilingual ASR adaptation approaches on commonly used multilingual ASR benchmarks.
Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han
ICASSP3
2025 Speech Retrieval-Augmented Generation without Automatic Speech Recognition
abstract
One common approach for question answering over speech data is to first transcribe speech using automatic speech recognition (ASR) and then employ text-based retrieval-augmented generation (RAG) on the transcriptions. While this cascaded pipeline has proven effective in many practical settings, ASR errors can propagate to the retrieval and generation steps. To overcome this limitation, we introduce SpeechRAG, a novel framework designed for open-question answering over spoken data. Our proposed approach fine-tunes a pre-trained speech encoder into a speech adapter fed into a frozen large language model (LLM)–based retrieval model. By aligning the embedding spaces of text and speech, our speech retriever directly retrieves audio passages from text-based queries, leveraging the retrieval capacity of the frozen text retriever. Our retrieval experiments on spoken question answering datasets show that direct speech retrieval does not degrade over the text-based baseline, and outperforms the cascaded systems using ASR. For generation, we use a speech language model (SLM) as a generator, conditioned on audio passages rather than transcripts. Without fine-tuning of the SLM, this approach outperforms cascaded text-based models when there is high WER in the transcripts.
Do June Min, Karel Mundnich, Andy Lapastora, Erfan Soltanmohammadi, Srikanth Ronanki, Kyu J. Han
ICASSP6
2025 Zero-resource Speech Translation and Recognition with LLMs
abstract
Despite recent advancements in speech processing, zero-resource speech translation (ST) and automatic speech recognition (ASR) remain challenging problems. In this work, we propose to leverage a multilingual Large Language Model (LLM) to perform ST and ASR in languages for which the model has never seen paired audio-text data. We achieve this by using a pre-trained multilingual speech encoder, a multilingual LLM, and a lightweight adaptation module that maps the audio representations to the token embedding space of the LLM. We perform several experiments both in ST and ASR to understand how to best train the model and what data has the most impact on performance in previously unseen languages. In ST, our best model is capable to achieve BLEU scores over 23 in CoVoST2 for two previously unseen languages, while in ASR, we achieve WERs of up to 28.2%. We finally show that the performance of our system is bounded by the ability of the LLM to output text in the desired language.
Karel Mundnich, Xing Niu 0001, Prashant Mathur, Srikanth Ronanki, Brady Houston, Veera Raghavendra Elluru, Nilaksh Das, Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff
ICASSP12
2025 Knowledge Distillation From Ensemble for Spoken Language Identification
abstract
Spoken language identification (LID) has seen substantial performance gains with the rise of large-scale models. However, these models are often computationally expensive and impractical for many real-world applications. In this work, we propose a novel knowledge distillation from ensemble framework to address this challenge. By distilling an ensemble of large LID models into a single, more efficient student, we achieve comparable or even superior performance while reducing computational cost by 67%. Our approach yields a student model with less than 10% the size of a 200M+ parameter teacher ensemble, yet outperforming a 140M parameter teacher by 13% relative. Additionally, combining our distillation technique with decoupled knowledge distillation leads to substantial gains (50% relative), especially for confusable and low-resource languages in the FLEURS dataset.
Raghuveer Peri, Seyed Omid Sadjadi, Daniel Garcia-Romero, Srikanth Vishnubhotla, Kyu J. Han
ICASSP5
2025 Contextual ASR with Retrieval Augmented Large Language Model
abstract
Automatic speech recognition (ASR) systems can benefit from incorporating contextual information to improve recognition accuracy, especially for uncommon words or phrases. Current approaches like custom vocabularies or prompting with previous transcript segments provide limited contextual control. Compared to existing context biasing methods, RAG promises more flexible and scalable contextual control by leveraging LLMs’ broad knowledge. To this end, we propose leveraging large language models (LLMs) and retrieval-augmented generation (RAG) to enhance the contextual capabilities of ASR systems. Specifically, we propose systems based on text and audio LLMs to perform contextual error correction with context retrieved by querying a text-based retriever using the ASR module’s firstpass ASR hypotheses and a frequency-based custom vocabulary (CV) list. Our experiments reveal that the fine-tuned system has effectively learned to extract the relevant context to perform error correction while maintaining robustness against noise.
Cihan Xiao, Zejiang Hou, Daniel Garcia-Romero, Kyu J. Han
ICASSP4
2025 The Interspeech 2025 Speech Accessibility Project Challenge
Xiuwen Zheng 0003, Bornali Phukan, Jonghwan Na, Edward Cutrell, Kyu J. Han, Mark Hasegawa-Johnson, Pan-Pan Jiang, Aadhrik Kuila, Colin Lea, Bob MacDonald, Gautam Varma Mantena, Venkatesh Ravichandran, Leda Sari, Katrin Tomanek, Chang Dong Yoo, Chris Zwilling
INTERSPEECH5
2025 Defending Speech-enabled LLMs Against Adversarial Jailbreak Threats
Antonios Alexos, Raghuveer Peri, Sai Muralidhar Jayanthi, Metehan Cekic, Srikanth Vishnubhotla, Kyu J. Han, Srikanth Ronanki
INTERSPEECH6
2025 Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
abstract
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Jianfeng He, Yi Nian, Amy Wing-mei Wong, Kyu J. Han, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Mahnaz Koupaee, Jake W. Vincent, Saab Mansour, Igor Shalyminov, Han He, Hwanjun Song, Raphael Shu, Yi Nian, Amy Wing-mei Wong, Kyu J. Han
NAACL (Long Papers)11
2024 Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic, Marcello Federico, Kyu J. Han
ECCV (79)6
2024 Tackling Missing Modalities in Audio-Visual Representation Learning Using Masked Autoencoders
Georgios Chochlakis, Chandrashekhar Lavania, Prashant Mathur, Kyu J. Han
INTERSPEECH4
2024 Revisiting Convolution-free Transformer for Speech Recognition
Zejiang Hou, Goeric Huybrechts, Anshu Bhatia, Daniel Garcia-Romero, Kyu J. Han, Katrin Kirchhoff
INTERSPEECH5
2024 Improving Multilingual ASR Robustness to Errors in Language Input
Brady Houston, Omid Sadjadi, Zejiang Hou, Srikanth Vishnubhotla, Kyu J. Han
INTERSPEECH5
2024 SWAN: SubWord Alignment Network for HMM-free word timing estimation in end-to-end automatic speech recognition
Woo Hyun Kang, Srikanth Vishnubhotla, Rudolf Braun, Yogesh Virkar, Raghuveer Peri, Kyu J. Han
INTERSPEECH6
2023 Utility-Preserving Privacy-Enabled Speech Embeddings for Emotion Detection
Chandrashekhar Lavania, Sanjiv Das, Kyu J. Han
INTERSPEECH4