Wei-Ping Huang

dblp:140/8512 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 SUTA-LM: Bridging Test-Time Adaptation and Language Model Rescoring for Robust ASR
abstract
Despite progress in end-to-end ASR, real-world domain mismatches still cause performance drops, which Test-Time Adaptation (TTA) aims to mitigate by adjusting models during inference. Recent work explores combining TTA with external language models, using techniques like beam search rescoring or generative error correction. In this work, we identify a previously overlooked challenge: TTA can interfere with language model rescoring, revealing the nontrivial nature of effectively combining the two methods. Based on this insight, we propose SUTA-LM, a simple yet effective extension of SUTA, an entropy-minimization-based TTA approach, with language model rescoring. SUTALM first applies a controlled adaptation process guided by an auto-step selection mechanism leveraging both acoustic and linguistic information, followed by language model rescoring to refine the outputs. Experiments on 18 diverse ASR datasets show that SUTA-LM achieves robust results across a wide range of domains.11The source code is available at https://github.com/hhhaaahhhaa/ASR-TTA
Wei-Ping Huang, Guan-Ting Lin, Hung-yi Lee
ASRU1
2024 Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
Chun-Yi Kuan, Wei-Ping Huang, Hung-yi Lee
INTERSPEECH2
2024 Speech-Copilot: Leveraging Large Language Models for Speech Processing Via Task Decomposition, Modularization, and Program Generation
abstract
In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot builds speech processing-specific toolsets by analyzing pre-collected task instructions and breaking tasks into manageable sub-tasks. It features a flexible agent that performs tasks through program generation based on large language models. Our approach achieves state-of the-art performance on the Dynamic-SUPERB benchmark, demonstrating its effectiveness across diverse speech-processing tasks. Key contributions include: 1) developing an innovative framework for speech processing-specific toolset construction, 2) establishing a high-performing agent based on large language models, and 3) offering a new perspective on addressing challenging instruction-oriented speech-processing tasks. Without additional training required by end-to-end approaches, our method provides a flexible and extendable solution for a wide range of speech-processing applications.
Chun-Yi Kuan, Chih-Kai Yang, Wei-Ping Huang, Ke-Han Lu, Hung-yi Lee
SLT3
2023 Maximizing Data Efficiency for Cross-Lingual TTS Adaptation by Self-Supervised Representation Mixing and Embedding Initialization
abstract
This paper presents an effective transfer learning framework for language adaptation in text-to-speech systems, with a focus on achieving language adaptation using minimal labeled and unlabeled data. While many works focus on reducing the usage of labeled data, very few consider minimizing the usage of unlabeled data. By utilizing self-supervised features in the pretraining stage, replacing the noisy portion of pseudo labels with these features during fine-tuning, and incorporating an embedding initialization trick, our method leverages more information from unlabeled data compared to conventional approaches. Experimental results show that our framework is able to synthesize intelligible speech in unseen languages with only 4 utterances of labeled data and 15 minutes of unlabeled data. Our methodology continues to surpass conventional techniques, even when a greater volume of data is accessible. These findings highlight the potential of our data-efficient language adaptation framework.
Wei-Ping Huang, Sung-Feng Huang, Hung-yi Lee
ASRU1
2023 Findings of the 2023 ML-Superb Challenge: Pre-Training And Evaluation Over More Languages And Beyond
abstract
The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identification. The challenge comprises a research track focused on applying ML-SUPERB to specific multilingual subjects, a Challenge Track for model submissions, and a New Language Track where language resource researchers can contribute and evaluate their low-resource language data in the context of the latest progress in multilingual speech recognition. The challenge garnered 12 model submissions and 54 language corpora, resulting in a comprehensive benchmark encompassing 154 languages. The findings indicate that merely scaling models is not the definitive solution for multilingual speech tasks, and a variety of speech/voice types present significant challenges in multilingual speech processing.
Jiatong Shi, Dan Berrebbi, Hsiu-Hsuan Wang, Wei-Ping Huang, En-Pei Hu, Ho-Lam Chuang, Xuankai Chang, Yuxun Tang, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Shinji Watanabe 0001
ASRU5
2023 Why We Should Report the Details in Subjective Evaluation of TTS More Rigorously
Cheng-Han Chiang, Wei-Ping Huang, Hung-yi Lee
INTERSPEECH2
2023 ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
Jiatong Shi, Dan Berrebbi, En-Pei Hu, Wei-Ping Huang, Ho-Lam Chung, Xuankai Chang, Shang-Wen Li 0001, Abdel-rahman Mohamed, Hung-yi Lee, Shinji Watanabe 0001
INTERSPEECH5
2022 Few Shot Cross-Lingual TTS Using Transferable Phoneme Embedding
abstract
This paper studies a transferable phoneme embedding framework that aims to deal with the cross-lingual text-to-speech (TTS) problem under the few-shot setting.Transfer learning is a common approach when it comes to few-shot learning since training from scratch on few-shot training data is bound to overfit.Still, we find that the naive transfer learning approach fails to adapt to unseen languages under extremely fewshot settings, where less than 8 minutes of data is provided.We deal with the problem by proposing a framework that consists of a phoneme-based TTS model and a codebook module to project phonemes from different languages into a learned latent space.Furthermore, by utilizing phoneme-level averaged selfsupervised learned features, we effectively improve the quality of synthesized speeches.Experiments show that using 4 utterances, which is about 30 seconds of data, is enough to synthesize intelligible speech when adapting to an unseen language using our framework.
Wei-Ping Huang, Po-Chun Chen, Sung-Feng Huang, Hung-yi Lee
INTERSPEECH1
2022 On the Utility of Self-Supervised Models for Prosody-Related Tasks
abstract
Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in speech signals. However, relatively little is known about the suitability of such models for prosody-related tasks or the extent to which they encode prosodic information. We present a new evaluation framework, “SUPERB-prosody,” consisting of three prosody-related downstream tasks and two pseudo tasks. We find that 13 of the 15 SSL models outperformed the baseline on all the prosody-related tasks. We also show good performance on two pseudo tasks: prosody reconstruction and future prosody prediction. We further analyze the layerwise contributions of the SSL models. Overall we conclude that SSL speech models are highly effective for prosody-related tasks. We release our code11https://github.com/JSALT-2022-SSL/superb-prosody for the community to support further investigation of SSL models' utility for prosody.
Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen-An Li, Hung-yi Lee, Nigel G. Ward
SLT3
2013 Application of Data Mining on HPLC Fingerprints of Szechwan Lovage Rhizome Analysis
abstract
Based on the integration of Java language and open-source R software environment, the article was developed a Traditional Chinese Medicine fingerprints analysis and visualization system and taken Szechwan Lovage Rhizome HPLC fingerprints as the study object to conduct data processing, information analyzing, and data mining research. In the article, 24 batches of Szechwan Lovage Rhizome from three different growth regions together with 3 standard samples were selected to make experiment detection. Data from HPLC fingerprints were processed by principal component analysis (PCA), and were completed the regional difference analysis for the main active components of the medicine from the different growth regions, and then with 3D visualization of the result, the growth regions were significantly distinguished from system. In addition, embedded with the GIS technology, the system was initially accomplished the correlation analysis between the Szechwan Lovage Rhizome fingerprints data and the geographical space data. Therefore the fingerprints data quantitative analysis method and system developed here can be regarded as an efficient way for quality detection and analysis automatically and intelligently of the traditional Chinese medicine.
Chun-Yi Huang, Sheng-Feng Shi, Wei-Ping Huang
MSN4