Jun-You Wang

dblp:296/4519 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-9119-9259ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Similarity-based Accent Recognition with Continuous and Discrete Self-supervised Speech Representations
abstract
The primary challenge in accent recognition lies in data scarcity due to the high diversity of accents, which make the collection of large-scale training data for each accent almost impossible in practice. To overcome this challenge, we propose a simple solution that leverages both continuous and discrete feature representations from pretrained speech self-supervised learning (SSL) models. Our model is simplified to a linear projection layer and a set of trainable accent class embeddings. Cosine similarity between the accent embeddings and the latent features of an audio sample is used to predict its accent class. This approach enables the model to access features that contain rich accent-related information while reducing the risk of model overfitting. Our method provides a practical and efficient way to tackle accent recognition, especially in low-resource scenarios. Experimental results on English accent recognition show that our best model achieves an accuracy of 84.0% on the AESRC 2020 dataset and an Unweighted Average Recall (UAR) of 50.0% on the VCTK corpus, setting new state-of-the-art results on both datasets.
Jun-You Wang, Sheng Li 0010, Li-An Lu, Sydney Chia-Chun Kao, Jyh-Shing Roger Jang
ICASSP1
2025 Many-to-Many Singing Performance Style Transfer on Pitch and Energy Contours
abstract
Singing voice conversion (SVC) aims to convert the singer identity of a singing voice to that of another singer. However, most existing SVC systems only perform the conversion of timbre information, while leaving other information unchanged. This approach does not consider other aspects of singer identity, particularly a singer's performance style, which is reflected in the pitch (F0) and the energy (volume dynamics) contours of singing. To address this issue, this paper proposes a many-to-many singing performance style transfer system that converts the pitch and energy contours of one singer's style to another singer's. To achieve this target, we utilize two AutoVC-like autoencoders with an information bottleneck to automatically disentangle performance style from other musical contents, one for the pitch contour while another for the energy contour. Experiment results suggested that the proposed model can perform singing performance style transfer in a many-to-many conversion scenario, resulting in improved singer identity similarity to the target singer.
Yu-Teng Hsu, Jun-You Wang, Jyh-Shing Roger Jang
IEEE Signal Process. Lett.2
2024 MIR-MLPop: A Multilingual Pop Music Dataset with Time-Aligned Lyrics and Audio
abstract
We introduce MIR-MLPop, a publicly available multilingual pop music dataset designed for automatic lyrics transcription and lyrics alignment in polyphonic music. The dataset comprises 90 pop music tracks in Mandarin, Cantonese, and Taiwanese Hokkien, with manually annotated time-aligned lyrics with both characters and pronunciation labels. To the best of our knowledge, this is the first ever singing dataset for Cantonese and Taiwanese Hokkien. In the experiments, using the pretrained Whisper model as the backbone, we develop lyrics transcription and lyrics alignment models for all three languages. Overall, the results are promising for both tasks, but show clear differences among the languages. Our models perform significantly better on languages that have been seen by Whisper during pretraining than on the language unseen by Whisper. This finding highlights the potential challenge in lyrics transcription and alignment for low-resource languages that have not been covered by pretrained speech models.
Jun-You Wang, Chung-Che Wang, Chon-In Leong, Jyh-Shing Roger Jang
ICASSP1
2023 Zero-Shot Singing Voice Synthesis from Musical Score
abstract
Zero-shot singing voice synthesis (SVS), the task to synthesize the singing voice of an arbitrary target singer, has gained increasing attentions in the past few years. Several recently proposed systems have demonstrated promising results on this task. However, these systems require detailed musical features at the frame level as the musical content. To deal with this issue, we propose a model that performs zero-shot SVS with only musical score as the musical content condition. To help model training, we build an acoustic encoder that extracts linguistic features from audio, and train it with the lyrics transcription objective. The output of the acoustic encoder serves as an alternative to the musical score, allowing the SVS model to learn from weakly labeled data. Results suggest that the proposed method outperforms baseline semi-supervised method in both subjective and objective tests.
Jun-You Wang, Hung-yi Lee, Jyh-Shing Roger Jang, Li Su 0004
ASRU1
2023 Adapting Pretrained Speech Model for Mandarin Lyrics Transcription and Alignment
abstract
The tasks of automatic lyrics transcription and lyrics alignment have witnessed significant performance improvements in the past few years. However, most of the previous works only focus on English in which large-scale datasets are available. In this paper, we address lyrics transcription and alignment of polyphonic Mandarin pop music in a low-resource setting. To deal with the data scarcity issue, we adapt pretrained Whisper model and fine-tune it on a monophonic Mandarin singing dataset. With the use of data augmentation and source separation model, results show that the proposed method achieves a character error rate of less than 18% on a Mandarin polyphonic dataset for lyrics transcription, and a mean absolute error of 0.071 seconds for lyrics alignment. Our results demonstrate the potential of adapting a pretrained speech model for lyrics transcription and alignment in low-resource scenarios.
Jun-You Wang, Chon-In Leong, Li Su 0004, Jyh-Shing Roger Jang
ASRU1
2023 Training a Singing Transcription Model Using Connectionist Temporal Classification Loss and Cross-Entropy Loss
abstract
In this paper, we propose a method that uses a combination of the Connectionist Temporal Classification (CTC) loss and the cross-entropy loss to train a note-level singing transcription model. By considering the task as predicting a note sequence of the input audio, we can compute the CTC loss between the prediction and the groundtruth note sequence, and further use it with the traditional cross-entropy loss to optimize the transcription model. By comparing the proposed method with a baseline that only utilizes the cross-entropy loss, the results show improved model performance on all the evaluation metrics. Furthermore, using the CTC loss allows the transcription model to learn from weakly labeled data, which is easier to annotate than traditional strongly labeled data. Moreover, we point out the issue of the intrinsic global time shift on the onset labels between datasets. By automatically estimating and calibrating the global time shift of the training dataset, the performance of the singing transcription model is then not affected by the global time shift in the cross-dataset evaluation scenario.
Jun-You Wang, Jyh-Shing Roger Jang
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 On the Preparation and Validation of a Large-Scale Dataset of Singing Transcription
abstract
This paper proposes a large-scale dataset for singing transcription, along with some methods for fine-tuning and validating its contents. The dataset is named MIR-ST500, which consists of more than 160,000 notes from 500 pop songs. To create this large-scale dataset, we set some labeling criteria and ask non-experts to label notes. We also perform some adjustments on the annotation to correct minor errors. Finally, to validate the dataset, we train a singing transcription model on MIR-ST500 dataset and evaluate it on various datasets. The result shows that we can certainly construct a better singing transcription model for various purposes using MIR-ST500, which is properly labeled and validated.
Jun-You Wang, Jyh-Shing Roger Jang
ICASSP1