VLDB 2026 Research / reviewers in the wild / expert
Yist Y. Lin
dblp:265/6549
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2024
0009-0000-9054-1596ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Large-Scale Evaluation of Speech Foundation ModelsabstractThe foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific data collection and modeling. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. To bridge this gap, we establish the Speech processing Universal PERformance Benchmark (SUPERB). SUPERB represents an ecosystem designed to evaluate foundation models across a wide range of speech processing tasks, facilitating the sharing of results on an online leaderboard and fostering collaboration through a community-driven benchmark database that aids in new development cycles. We present a unified learning framework for solving the speech processing tasks in SUPERB with the frozen foundation model followed by task-specialized lightweight prediction heads. Combining our results with community submissions, we verify that the framework is simple yet effective, as the best-performing foundation model shows competitive generalizability across most SUPERB tasks. Finally, we conduct a series of analyses to offer an in-depth understanding of SUPERB and speech foundation models, including information flows across tasks inside the models and the statistical significance and robustness of the benchmark. Shu-Wen Yang, Heng-Jui Chang, Zili Huang, Andy T. Liu, Cheng-I Lai, Jiatong Shi, Xuankai Chang, Hsiang-Sheng Tsai, Wen-Chin Huang, Tzu-hsun Feng, Po-Han Chi, Yist Y. Lin, Yung-Sung Chuang, Tzu-Hsien Huang, Wei-Cheng Tseng, Kushal Lakhotia, Shang-Wen Li 0001, Abdel-rahman Mohamed, Shinji Watanabe 0001, Hung-yi Lee |
IEEE ACM Trans. Audio Speech Lang. Process. | 13 |
| 2023 | Improving Frame-level Classifier for Word Timings with Non-peaky CTC in End-to-End Automatic Speech Recognition
Xianzhao Chen, Yist Y. Lin, Zejun Ma 0001 |
INTERSPEECH | 2 |
| 2023 | Random Utterance Concatenation Based Data Augmentation for Improving Short-video Speech RecognitionabstractOne of limitations in end-to-end automatic speech recognition (ASR) framework is its performance would be compromised if train-test utterance lengths are mismatched.In this paper, we propose an on-the-fly random utterance concatenation (RUC) based data augmentation method to alleviate train-test utterance length mismatch issue for short-video ASR task.Specifically, we are motivated by observations that our human-transcribed training utterances tend to be much shorter for short-video spontaneous speech (∼3 seconds on average), while our test utterance generated from voice activity detection front-end is much longer (∼10 seconds on average).Such a mismatch can lead to suboptimal performance.Empirically, it's observed the proposed RUC method significantly improves long utterance recognition without performance drop on short one.Overall, it achieves 5.72% word error rate reduction on average for 15 languages and improved robustness to various utterance length. Yist Y. Lin, Haihua Xu 0001, Van Tung Pham, Yerbolat Khassanov, Tze Yuang Chong, Lu Lu 0015, Zejun Ma 0001 |
INTERSPEECH | 1 |
| 2021 | Fragmentvc: Any-To-Any Voice Conversion by End-To-End Extracting and Fusing Fine-Grained Voice Fragments with AttentionabstractAny-to-any voice conversion aims to convert the voice from and to any speakers even unseen during training, which is much more challenging compared to one-to-one or many-to-many tasks, but much more attractive in real-world scenarios. In this paper we proposed FragmentVC, in which the latent phonetic structure of the utterance from the source speaker is obtained from Wav2Vec 2.0, while the spectral features of the utterance(s) from the target speaker are obtained from log mel-spectrograms. By aligning the hidden structures of the two different feature spaces with a two-stage training process, FragmentVC is able to extract fine-grained voice fragments from the target speaker utterance(s) and fuse them into the desired utterance, all based on the attention mechanism of Transformer as verified with analysis on attention maps, and is accomplished end-to-end. This approach is trained with reconstruction loss only without any disentanglement considerations between content and speaker information and doesn't require parallel data. Objective evaluation based on speaker verification and subjective evaluation with MOS both showed that this approach outperformed SOTA approaches, such as AdaIN-VC and AutoVC. Yist Y. Lin, Chung-Ming Chien, Jheng-Hao Lin, Hung-yi Lee, Lin-Shan Lee |
ICASSP | 1 |
| 2021 | S2VC: A Framework for Any-to-Any Voice Conversion with Self-Supervised Pretrained RepresentationsabstractAny-to-any voice conversion (VC) aims to convert the timbre of utterances from and to any speakers seen or unseen during training.Various any-to-any VC approaches have been proposed like AUTOVC, AdaINVC, and FragmentVC.AUTOVC, and AdaINVC utilize source and target encoders to disentangle the content and speaker information of the features.Frag-mentVC utilizes two encoders to encode source and target information and adopts cross attention to align the source and target features with similar phonetic content.Moreover, pretrained features are adopted.AUTOVC used d-vector to extract speaker information, and self-supervised learning (SSL) features like wav2vec 2.0 is used in FragmentVC to extract the phonetic content information.Different from previous works, we proposed S2VC that utilizes Self-Supervised features as both source and target features for the VC model.Supervised phoneme posteriorgram (PPG), which is believed to be speaker-independent and widely used in VC to extract content information, is chosen as a strong baseline for SSL features.The objective evaluation and subjective evaluation both show models taking SSL feature CPC as both source and target features outperforms that taking PPG as source feature, suggesting that SSL features have great potential in improving VC. Jheng-Hao Lin, Yist Y. Lin, Chung-Ming Chien, Hung-yi Lee |
Interspeech | 2 |
| 2021 | Utilizing Self-Supervised Representations for MOS PredictionabstractSpeech quality assessment has been a critical issue in speech processing for decades. Existing automatic evaluations usually require clean references or parallel ground truth data, which is infeasible when the amount of data soars. Subjective tests, on the other hand, do not need any additional clean or parallel data and correlates better to human perception. However, such a test is expensive and time-consuming because crowd work is necessary. It thus becomes highly desired to develop an automatic evaluation approach that correlates well with human perception while not requiring ground truth data. In this paper, we use self-supervised pre-trained models for MOS prediction. We show their representations can distinguish between clean and noisy audios. Then, we fine-tune these pre-trained models followed by simple linear layers in an end-to-end manner. The experiment results showed that our framework outperforms the two previous state-of-the-art models by a significant improvement on Voice Conversion Challenge 2018 and achieves comparable or superior performance on Voice Conversion Challenge 2016. We also conducted an ablation study to further investigate how each module benefits the task. The experiment results are implemented and reproducible with publicly available toolkits. Wei-Cheng Tseng, Chien-Yu Huang, Wei-Tsung Kao, Yist Y. Lin, Hung-yi Lee |
Interspeech | 4 |
| 2021 | SUPERB: Speech Processing Universal PERformance BenchmarkabstractSelf-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various tasks with minimal adaptation.However, the speech processing community lacks a similar setup to systematically explore the paradigm.To bridge this gap, we introduce Speech processing Universal PERformance Benchmark (SUPERB).SUPERB is a leaderboard to benchmark the performance of a shared model across a wide range of speech processing tasks with minimal architecture changes and labeled data.Among multiple usages of the shared model, we especially focus on extracting the representation learned from SSL for its preferable re-usability.We present a simple framework to solve SUPERB tasks by learning task-specialized lightweight prediction heads on top of the frozen shared model.Our results demonstrate that the framework is promising as SSL representations show competitive generalizability and accessibility across SUPERB tasks.We release SUPERB as a challenge with a leaderboard 1 and a benchmark toolkit 2 to fuel the research in representation learning and general speech processing. Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li 0001, Shinji Watanabe 0001, Abdel-rahman Mohamed, Hung-yi Lee |
Interspeech | 6 |
| 2021 | Defending Your Voice: Adversarial Attack on Voice ConversionabstractSubstantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance. Nonetheless, the improved conversion technologies also led to concerns about privacy and authentication. It thus becomes highly desired to be able to prevent one's voice from being improperly utilized with such voice conversion technologies. This is why we report in this paper the first known attempt to perform adversarial attack on voice conversion. We introduce human imperceptible noise into the utterances of a speaker whose voice is to be defended. Given these adversarial examples, voice conversion models cannot convert other utterances so as to sound like being produced by the defended speaker. Preliminary experiments were conducted on two currently state-of-the-art zero-shot voice conversion models. Objective and subjective evaluation results in both white-box and black-box scenarios are reported. It was shown that the speaker characteristics of the converted utterances were made obviously different from those of the defended speaker, while the adversarial examples of the defended speaker are not distinguishable from the authentic utterances. Chien-Yu Huang, Yist Y. Lin, Hung-yi Lee, Lin-Shan Lee |
SLT | 2 |