EDBT 2026 Demo / reviewers in the wild / expert
Ho-Lam Chung
dblp:276/6018
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized DataabstractWe propose a self-refining framework that enhances ASR performance with only unlabeled datasets. The process starts with an existing ASR model generating pseudo-labels on unannotated speech, which are then used to train a highfidelity text-to-speech (TTS) system. Then, synthesized speech text pairs are bootstrapped into the original ASR system, completing the closed-loop self-improvement cycle. We demonstrated the effectiveness of the framework on Taiwanese-Mandarin speech. Leveraging 6,000 hours of unlabeled speech, a moderate amount of text data, and synthetic content from the AI models, we adapt Whisper-large-v2 into a specialized model, Twister. Twister reduces error rates by up to $20 \%$ on Mandarin and $50 \%$ on Mandarin-English code-switching benchmarks compared to Whisper. Results highlight the framework as a compelling alternative to pseudo-labeling self-distillation approaches and provides a practical pathway for improving ASR performance in lowresource or domain-specific settings. Cheng-Kang Chou, Chan-Jan Hsu, Ho-Lam Chung, Liang-Hsuan Tseng, Hsi-Chun Cheng, Yu-Kuan Fu, Kuan-Po Huang, Hung-yi Lee |
ASRU | 3 |
| 2025 | Researchy Questions: A Dataset of Multi-Perspective, Decompositional Questions for Deep ResearchabstractExisting question answering (QA) datasets are no longer challenging to most powerful Large Language Models (LLMs). Traditional QA benchmarks like TriviaQA, NaturalQuestions, ELI5 and HotpotQA mainly study ''known unknowns'' with clear indications of both what information is missing, and how to find it to answer the question. A yet unmet need of the NLP community is a bank of non-factoid, multi-perspective questions involving a great deal of unclear information needs, i.e. ''unknown unknowns''. We claim we can find such questions in search engine logs, which is surprising because most question-intent queries are indeed factoid. Furthermore, recent products like Google's DeepResearch (announced a year after this resource was released publicly) specifically address such queries, retrieving hundreds of documents to synthesize report-style responses. We present Researchy Questions, the world's first, only and largest public dataset of ''Deep Research'' questions filtered from real search engine logs to be non-factoid, ''decompositional'' and multi-perspective. We show that users spend substantial ''effort'' on these questions in terms of signals like clicks and session length. We also show that ''slow thinking'' answering techniques, like decomposition into sub-questions shows benefit over answering directly. We release (at https://huggingface.co/datasets/corbyrosset/researchy_questions) about 100k Researchy Questions with a permissive CDLA-2.0 license, along with click histograms on over 350k Clueweb22 URLs that were clicked for each question. Corbin Rosset, Ho-Lam Chung, Guanghui Qin, Ethan C. Chau, Ahmed Awadallah 0001, Jennifer Neville, Nikhil Rao 0001 |
SIGIR | 2 |
| 2024 | GSQA: An End-to-End Model for Generative Spoken Question Answering
Min-Han Shih, Ho-Lam Chung, Yu-Chi Pai, Ming-Hao Hsu, Guan-Ting Lin, Shang-Wen Li 0001, Hung-yi Lee |
INTERSPEECH | 2 |
| 2024 | Codec-Superb @ SLT 2024: A Lightweight Benchmark For Neural Audio Codec ModelsabstractNeural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 20241, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge’s rules, datasets, participant systems, results, and findings.1https://codecsuperb.github.io/ Xuanjun Chen, Yi-Cheng Lin, Kai-Wei Chang 0001, Jiawei Du 0003, Ke-Han Lu, Alexander H. Liu, Ho-Lam Chung, Yuan-Kuei Wu, Dongchao Yang, Songxiang Liu, Yi-Chiao Wu, Xu Tan 0003, James R. Glass, Shinji Watanabe 0001, Hung-yi Lee |
SLT | 8 |
| 2024 | Handover QG: Question Generation by Decoder Fusion and Reinforcement LearningabstractIn recent years, Question Generation (QG) has gained significant attention as a research topic, particularly in the context of its potential to support automatic reading comprehension assessment preparation. However, current QG models are mostly trained on factoid-type datasets, which tend to produce questions that are too simple for assessing advanced abilities. One promising alternative is to train QG models on exam-type datasets, which contain questions that require content reasoning. Unfortunately, there is a shortage of such training data compared to factoid-type questions. To address this issue and improve the quality of QG for generating advanced questions, we propose theHandover QGframework. This framework involves the joint training of exam-type QG and factoid-type QG, and controls the question generation process by interleavingly using the exam-type QG decoder and the factoid-type QG decoder. Furthermore, we employ reinforcement learning to enhance QG performance. Our experimental evaluation shows that our model significantly outperforms the compared baselines, with a BLEU-4 score increase from 5.31 to 6.48. Human evaluation also confirms that the questions generated by our model are answerable and appropriately difficult. Overall, theHandover QGframework offers a promising solution for improving QG performance in generating advanced questions for reading comprehension assessment. Ho-Lam Chung, Ying-Hong Chan, Yao-Chung Fan |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2023 | Towards General-Purpose Text-Instruction-Guided Voice ConversionabstractThis paper introduces a novel voice conversion (VC) model, guided by text instructions such as “articulate slowly with a deep tone“ or “speak in a cheerful boyish voice”. Unlike traditional methods that rely on reference utterances to determine the attributes of the converted speech, using text instruction adds versatility and specificity to voice conversion. The proposed VC model is a neural codec language model which processes a sequence of discrete codes, resulting in the code sequence of converted speech. It utilizes text instructions as style prompts to modify the prosody and emotional information of the given speech. In contrast to previous approaches, which often rely on employing separate encoders like prosody and content encoders to handle different aspects of the source speech, our model handles various information of speech in an end-to-end manner. Experiments have demonstrated the impressive capabilities of our model in comprehending instructions and delivering reasonable results1. Chun-Yi Kuan, Chen-An Li, Tsu-Yuan Hsu, Tse-Yang Lin, Ho-Lam Chung, Kai-Wei Chang 0001, Shuo-Yiin Chang, Hung-yi Lee |
ASRU | 5 |
| 2023 | T5lephone: Bridging Speech and Text Self-Supervised Models for Spoken Language Understanding Via Phoneme Level T5abstractIn Spoken language understanding (SLU), a natural solution is concatenating pre-trained speech models (e.g. HuBERT) and pretrained language models (PLM, e.g. T5). Most previous works use pre-trained language models with subword-based tokenization. However, the granularity of input units affects the alignment of speech model outputs and language model inputs, and PLM with character-based tokenization is underexplored. In this work, we conduct extensive studies on how PLMs with different tokenization strategies affect spoken language understanding task including spoken question answering (SQA) and speech translation (ST).We further extend the idea to create T5lephone1, a variant of T5 that is pretrained using phonemicized text. We initialize T5lephone with existing PLMs to pretrain it using relatively lightweight computational resources. We reached state-of-the-art on NMSQA, and the T5lephone model exceeds T5 with other types of units on end-to-end SQA and ST. Our code is publicly available.2 Chan-Jan Hsu, Ho-Lam Chung, Hung-yi Lee, Yu Tsao 0001 |
ICASSP | 2 |
| 2023 | Bridging Speech and Textual Pre-Trained Models With Unsupervised ASRabstractSpoken language understanding (SLU) is a task aiming to extract high-level semantics from spoken utterances. Previous works have investigated the use of speech self-supervised models and textual pre-trained models, which have shown reasonable improvements to various SLU tasks. However, because of the mismatched modalities between speech signals and text tokens, previous methods usually need complex designs of the frameworks. This work proposes a simple yet efficient unsupervised paradigm that connects speech and textual pre-trained models, resulting in an unsupervised speech-to-semantic pre-trained model for various tasks in SLU. To be specific, we propose to use unsupervised automatic speech recognition (ASR) as a connector that bridges different modalities used in speech and textual pre-trained models. Our experiments show that unsupervised ASR itself can improve the representations from speech self-supervised models. More importantly, it is shown as an efficient connector between speech and textual pre-trained models, improving the performances of five different SLU tasks. Notably, on spoken question answering, we reach the state-of-the-art result over the challenging NMSQA benchmark. Jiatong Shi, Chan-Jan Hsu, Ho-Lam Chung, Dongji Gao, L. Paola García-Perera, Shinji Watanabe 0001, Ann Lee 0001, Hung-yi Lee |
ICASSP | 3 |
| 2023 | ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
Jiatong Shi, Dan Berrebbi, En-Pei Hu, Wei-Ping Huang, Ho-Lam Chung, Xuankai Chang, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Shinji Watanabe 0001 |
INTERSPEECH | 6 |
| 2022 | DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question AnsweringabstractSpoken Question Answering (SQA) is to find the answer from a spoken document given a question, which is crucial for personal assistants when replying to the queries from the users.Existing SQA methods all rely on Automatic Speech Recognition (ASR) transcripts.Not only does ASR need to be trained with massive annotated data that are time and cost-prohibitive to collect for low-resourced languages, but more importantly, very often the answers to the questions include name entities or out-of-vocabulary words that cannot be recognized correctly.Also, ASR aims to minimize recognition errors equally over all words, including many function words irrelevant to the SQA task.Therefore, SQA without ASR transcripts (textless) is always highly desired, although known to be very difficult.This work proposes Discrete Spoken Unit Adaptive Learning (DUAL), leveraging unlabeled data for pre-training and finetuned by the SQA downstream task.The time intervals of spoken answers can be directly predicted from spoken documents.We also release a new SQA benchmark corpus, NMSQA, for data with more realistic scenarios.We empirically showed that DUAL yields results comparable to those obtained by cascading ASR and text QA model and robust to real-world data.Our code and model will be open-sourced 1 . Guan-Ting Lin, Yung-Sung Chuang, Ho-Lam Chung, Shu-Wen Yang, Hsuan-Jui Chen, Shuyan Dong, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Lin-Shan Lee |
INTERSPEECH | 3 |
| 2022 | Misleading Inference Generation via Proximal Policy Optimization
Hsien-Yung Peng, Ho-Lam Chung, Ying-Hong Chan, Yao-Chung Fan |
PAKDD (1) | 2 |