EDBT 2026 Demo / reviewers in the wild / expert
Chien-Yu Huang
dblp:81/6994 · also Chien-yu Huang
· DBLP profile ↗
17ranked-venue papers
8as first author
9since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style CaptioningabstractInstruction-based speech processing is becoming popular. Studies show that training with multiple tasks boosts performance, but collecting diverse, large-scale tasks and datasets is expensive. Thus, it is highly desirable to design a fundamental task that benefits other downstream tasks. This paper introduces a multi-talker speaking style captioning task to enhance the understanding of speaker and prosodic information. We used large language models to generate descriptions for multi-talker speech. Then, we trained our model with pre-training on this captioning task followed by instruction tuning. Evaluation on Dynamic-SUPERB shows our model outperforming the baseline pre-trained only on single-talker tasks, particularly in speaker and emotion recognition. The code and dataset are available at https://github.com/cyhuang-tw/speechcaps. Chien-Yu Huang, Min-Han Shih, Ke-Han Lu, Chi-Yuan Hsiao, Hung-yi Lee |
ICASSP | 1 |
| 2025 | Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 TasksabstractMultimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wide range of natural language instructions is critical for bridging communication gaps and facilitating more intuitive interactions. However, the absence of a comprehensive evaluation benchmark poses a significant challenge. We present Dynamic-SUPERB Phase-2, an open and evolving benchmark for the comprehensive evaluation of instruction-based universal speech models. Building upon the first generation, this second version incorporates 125 new tasks contributed collaboratively by the global research community, expanding the benchmark to a total of 180 tasks, making it the largest benchmark for speech and audio evaluation. While the first generation of Dynamic-SUPERB was limited to classification tasks, Dynamic-SUPERB Phase-2 broadens its evaluation capabilities by introducing a wide array of novel and diverse tasks, including regression and sequence generation, across speech, music, and environmental audio. Evaluation results show that no model performed well universally. SALMONN-13B excelled in English ASR and Qwen2-Audio-7B-Instruct showed high accuracy in emotion recognition, but current models still require further innovations to handle a broader range of tasks. We open-source all task data and the evaluation pipeline at https://github.com/dynamic-superb/dynamic-superb. Chien-Yu Huang, Wei-Chih Chen, Shu-Wen Yang, Andy T. Liu, Chen-An Li, Yu-Xiang Lin, Wei-Cheng Tseng, Anuj Diwan, Yi-Jen Shih, Jiatong Shi, Chih-Kai Yang, Xuanjun Chen, Chi-Yuan Hsiao, Puyuan Peng, Shih-Heng Wang, Chun-Yi Kuan, Ke-Han Lu, Kai-Wei Chang 0001, Fabian Ritter Gutierrez |
ICLR | 1 |
| 2024 | Dynamic-Superb: Towards a Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark For SpeechabstractText language models have shown remarkable zero-shot capability in generalizing to unseen tasks when provided with well-formulated instructions. However, existing studies in speech processing primarily focus on limited or specific tasks. Moreover, the lack of standardized benchmarks hinders a fair comparison across different approaches. Thus, we present Dynamic-SUPERB, a benchmark designed for building universal speech models capable of leveraging instruction tuning to perform multiple tasks in a zero-shot fashion. To achieve comprehensive coverage of diverse speech tasks and harness instruction tuning, we invite the community to collaborate and contribute, facilitating the dynamic growth of the benchmark. To initiate, Dynamic-SUPERB features 55 evaluation instances by combining 33 tasks and 22 datasets. This spans a broad spectrum of dimensions, providing a comprehensive platform for evaluation. Additionally, we propose several approaches to establish benchmark baselines. These include the utilization of speech models, text language models, and the multimodal encoder. Evaluation results indicate that while these baselines perform reasonably on seen tasks, they struggle with unseen ones. We release all materials to the public and welcome researchers to collaborate on the project, advancing technologies in the field together1. Chien-Yu Huang, Ke-Han Lu, Shih-Heng Wang, Chi-Yuan Hsiao, Chun-Yi Kuan, Siddhant Arora, Kai-Wei Chang 0001, Jiatong Shi, Yifan Peng 0003, Roshan S. Sharma, Shinji Watanabe 0001, Bhiksha Raj, Shady Shehata, Hung-yi Lee |
ICASSP | 1 |
| 2024 | Fusion Of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech RecognitionabstractSelf-supervised learning (SSL) models have shown exceptional capabilities across various speech-processing tasks. Continuous SSL representations are effective but suffer from high computational and storage demands. On the other hand, discrete SSL representations, although with degraded performance, reduce transmission and storage costs, and improve input sequence efficiency through de-duplication and subword-modeling. To boost the performance of discrete representations for ASR, we introduce a novel fusion mechanism that integrates two discrete representations. The fusion mechanism preserves all the benefits of discrete representation while enhancing the model’s performance by integrating complementary information. Additionally, we explore “self-augmented” discrete representations, which apply transformations to a single continuous SSL representation, eliminating the fusion mechanism’s dependency on multiple SSL models and further decreasing its inference costs. Experimental results on benchmarks, including LibriSpeech and ML-SUPERB, indicate up to 19% and 24% relative character error rate improvement compared with the non-fusion baseline, validating the effectiveness of our proposed methods. Shih-Heng Wang, Jiatong Shi, Chien-Yu Huang, Shinji Watanabe 0001, Hung-yi Lee |
SLT | 3 |
| 2023 | Prompting and Adapter Tuning For Self-Supervised Encoder-Decoder Speech ModelabstractPrompting and adapter tuning have emerged as efficient alternatives to fine-tuning (FT) methods. However, existing studies on speech prompting focused on classification tasks and failed on more complex sequence generation tasks. Besides, adapter tuning is primarily applied with a focus on encoder-only self-supervised models. Our experiments show that prompting on Wav2Seq, a self-supervised encoder-decoder model, surpasses previous works in sequence generation tasks. It achieves a remarkable 53% relative improvement in word error rate for ASR and a 27% in F1 score for slot filling. Additionally, prompting competes with the FT method in the low-resource scenario. Moreover, we show the transferability of prompting and adapter tuning on Wav2Seq in cross-lingual ASR. When limited trainable parameters are involved, prompting and adapter tuning consistently outperform conventional FT across 7 languages. Notably, in the low-resource scenario, prompting consistently outperforms adapter tuning. Kai-Wei Chang 0001, Ming-Hsin Chen, Yun-Ping Lin, Jing Neng Hsu, Paul Kuo-Ming Huang, Chien-Yu Huang, Shang-Wen Li 0001, Hung-yi Lee |
ASRU | 6 |
| 2022 | Toward Degradation-Robust Voice ConversionabstractAny-to-any voice conversion technologies convert the vocal timbre of an utterance to any speaker even unseen during training. Although there have been several state-of-the-art any-to-any voice conversion models, they were all based on clean utterances to convert successfully. However, in real-world scenarios, it is difficult to collect clean utterances of a speaker, and they are usually degraded by noises or reverberations. It thus becomes highly desired to understand how these degradations affect voice conversion and build a degradation-robust model. We report in this paper the first comprehensive study on the degradation robustness of any-to-any voice conversion. We show that the performance of state-of-the-art models nowadays was severely hampered given degraded utterances. To this end, we then propose speech enhancement concatenation and denoising training to improve the robustness. In addition to common degradations, we also consider adversarial noises, which alter the model output significantly yet are human-imperceptible. It was shown that both concatenations with off-the-shelf speech enhancement models and denoising training on voice conversion models could improve the robustness, while each of them had pros and cons. Chien-Yu Huang, Kai-Wei Chang 0001, Hung-yi Lee |
ICASSP | 1 |
| 2021 | Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-SpeechabstractThe few-shot multi-speaker multi-style voice cloning task is to synthesize utterances with voice and speaking style similar to a reference speaker given only a few reference samples. In this work, we investigate different speaker representations and proposed to integrate pretrained and learnable speaker representations. Among different types of embeddings, the embedding pretrained by voice conversion achieves the best performance. The FastSpeech 2 model combined with both pretrained and learnable speaker representations shows great generalization ability on few-shot speakers and achieved 2nd place in the one-shot track of the ICASSP 2021 M2VoC challenge. Chung-Ming Chien, Jheng-Hao Lin, Chien-Yu Huang, Po-Chun Hsu, Hung-yi Lee |
ICASSP | 3 |
| 2021 | Utilizing Self-Supervised Representations for MOS PredictionabstractSpeech quality assessment has been a critical issue in speech processing for decades. Existing automatic evaluations usually require clean references or parallel ground truth data, which is infeasible when the amount of data soars. Subjective tests, on the other hand, do not need any additional clean or parallel data and correlates better to human perception. However, such a test is expensive and time-consuming because crowd work is necessary. It thus becomes highly desired to develop an automatic evaluation approach that correlates well with human perception while not requiring ground truth data. In this paper, we use self-supervised pre-trained models for MOS prediction. We show their representations can distinguish between clean and noisy audios. Then, we fine-tune these pre-trained models followed by simple linear layers in an end-to-end manner. The experiment results showed that our framework outperforms the two previous state-of-the-art models by a significant improvement on Voice Conversion Challenge 2018 and achieves comparable or superior performance on Voice Conversion Challenge 2016. We also conducted an ablation study to further investigate how each module benefits the task. The experiment results are implemented and reproducible with publicly available toolkits. Wei-Cheng Tseng, Chien-Yu Huang, Wei-Tsung Kao, Yist Y. Lin, Hung-yi Lee |
Interspeech | 2 |
| 2021 | Defending Your Voice: Adversarial Attack on Voice ConversionabstractSubstantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance. Nonetheless, the improved conversion technologies also led to concerns about privacy and authentication. It thus becomes highly desired to be able to prevent one's voice from being improperly utilized with such voice conversion technologies. This is why we report in this paper the first known attempt to perform adversarial attack on voice conversion. We introduce human imperceptible noise into the utterances of a speaker whose voice is to be defended. Given these adversarial examples, voice conversion models cannot convert other utterances so as to sound like being produced by the defended speaker. Preliminary experiments were conducted on two currently state-of-the-art zero-shot voice conversion models. Objective and subjective evaluation results in both white-box and black-box scenarios are reported. It was shown that the speaker characteristics of the converted utterances were made obviously different from those of the defended speaker, while the adversarial examples of the defended speaker are not distinguishable from the authentic utterances. Chien-Yu Huang, Yist Y. Lin, Hung-yi Lee, Lin-Shan Lee |
SLT | 1 |
| 2019 | Holes in the Outline: Subject-dependent Abstract Quality and its Implications for Scientific Literature SearchabstractScientific literature search engines typically index abstracts instead of the full-text of publications. The expectation is that the abstract provides a comprehensive summary of the article, enumerating key points for the reader to assess whether their information needs could be satisfied by reading the full-text. Furthermore, from a practical standpoint, obtaining the full-text is more complicated due to licensing issues, in the case of commercial publishers, and resource limitations of public repositories and pre-print servers. Chien-Yu Huang, Arlene Casey, Dorota Glowacka, Alan Medlar |
CHIIR | 1 |
| 2018 | General floorplanning methodology for 3D ICs with an arbitrary bonding styleabstractThis paper proposes a general floorplanning methodology which can be applied to 3D ICs with an arbitrary bonding style. Some researches have shown that a 3D IC with the hybrid bonding style, which includes face-to-back and face-to-face, may obtain better results than that simply using the face-to-back bonding style. We respectively present an approach to assign modules to tiers for each kind of bonding style. Further, a new utilization function, called cosine-shaped function, is proposed to estimate utilizations of bins required by the analytical-based approach. Our experimental results show the cosine-shaped function can obtain a little better result than the bell-shaped function on IBM benchmarks for 2D floorplanning. We also show that the proposed 3D floorplanning methodology consumes less TSVs and induces shorter wirelength compared to previous work in the hybrid bonding style. Jai-Ming Lin, Chien-Yu Huang |
DATE | 2 |
| 2018 | Co-synthesis of floorplanning and powerplanning in 3D ICs for multiple supply voltage designsabstractThis paper addresses a 3D floorplanning methodology, which considers floorplanning and powerplanning at the same time for Multiple Supply Voltage (MSV) circuits. Physical design becomes more complex for MSV designs since modules with the same power domain have to be placed at close locations in 3D space to facilitate powerplanning and reduce IR-drop, which would deteriorate wirelength. By properly partitioning modules of the same power domain into several voltage islands and increasing overlap area of the voltage islands in contiguous dies, we can reduce routing resource usage without increasing wirelength significantly. Further, unlike previous works, our approach not only can handle a netlist with soft modules and hard modules but also can meet the fixed-outline constraint. The experimental results show that our methodology gets better results than other approach in designs with single voltage domain and is also promising for MSV designs. Jai-Ming Lin, Chien-Yu Huang, Jhih-Ying Yang |
DATE | 2 |
| 2009 | Evaluating the process of a genetic algorithm to improve the back-propagation network: A Monte Carlo study
Chien-Yu Huang, Long-Hui Chen, Yueh-Li Chen, Fengming M. Chang |
Expert Syst. Appl. | 1 |
| 2008 | Optimizing back-propagation networks via a calibrated heuristic algorithm with an orthogonal array
Tai-Yue Wang, Chien-Yu Huang |
Expert Syst. Appl. | 2 |
| 2007 | Applying optimized BPN to a chaotic time series problem
Tai-Yue Wang, Chien-Yu Huang |
Expert Syst. Appl. | 2 |
| 2006 | Two-dimensional pilot-aided channel estimation for wireless OFDM systems over severe frequency-selective fading environmentsabstractOrthogonal frequency division multiplexing (OFDM) technique is effective and powerful in high data rate digital transmission due to its spectral efficiency, robustness in multipath propagation environments and ability to cope with intersymbol interference. Channel estimation is a crucial problem in coherent OFDM systems, and the various estimation techniques with particular pilot arrangements have been investigated recently. In this paper, a two-dimensional pilot-aided channel estimation technique is developed for wireless OFDM communications under severe frequency-selective fading channels. By applying least-square (LS) estimation for the leading pilot symbol in time domain, the channel statistics can be obtained and preserved. And then, LS estimation is also applied for the following regular pilot tones in the OFDM data symbol. Based on the variation of the two LS estimates in the time domain, a novel two-dimensional pilotaided channel estimation algorithm is presented. From the simulation results, the proposed channel estimation scheme has superior performance in the severe frequency-selective fading channels, especially in high SNR. Even if the mobile speed rises, the proposed algorithm still achieves excellent performance. Chien-Yu Huang, Wen-Jeng Lin, Jia-Chin Lin 0001, Jung-Shan Lin |
IWCMC | 1 |
| 2006 | LMI-based Integral fuzzy control of DC-DC convertersabstractIn this paper, we propose a T-S fuzzy controller which combines the merits of: i) the capability for dealing with nonlinear systems; ii) the powerful LMI approach to obtain control gains; iii) the high performance of integral controllers; iv) the workable rigorous proof for exponential convergence of error signals; and v) the flexibility on tuning decay rate. The output regulation problems of a basic buck converter and a zero-voltage-transition (ZVT) buck converter are used as application examples to illustrate the control performance of the proposed methodology. First, we consider a general nonlinear system which can represent the large-signal models of the converters. After introducing an added integral state of output regulation error and taking coordinate translation on an equilibrium point, the resulting augmented system is represented into a Takagi-Sugeno (T-S) fuzzy model. Then, the concept of parallel distributed compensation is applied to design the control law whereby the control gains are obtained by solving linear matrix inequalities (LMIs). An interesting result is that the obtained control law is formed only by the linear state feedback signals weighted by grade functions. In addition, the robustness analysis is carried out when uncertainty and disturbance are taken into consideration. The performance of numerical simulations and practical experiments results is satisfactory. Kuang-Yow Lian, Jeih-Jang Liou, Chien-Yu Huang |
IEEE Trans. Fuzzy Syst. | 3 |