EDBT 2026 Demo / reviewers in the wild / expert
Jyh-Shing Roger Jang
dblp:81/829
· DBLP profile ↗
92ranked-venue papers
12as first author
22since 2021 · last 2026
0000-0002-7319-9095ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 43 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorComputer networks · 2Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Social Sentiment Improve Bitcoin Dollar-Cost Averaging? Evidence from VADER and CryptoBERT
Hsiu Yuan Liang, Jyh-Shing Roger Jang, Ming-Hua Hsieh, Chien-Ping Chung |
ICBC | 2 |
| 2025 | Towards Generalized Source Tracing for Codec-Based Deepfake SpeechabstractRecent attempts at source tracing for codecbased deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-theart performance on the CoSG test set of CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing. I-Ming Lin, Xuanjun Chen, Lin Zhang 0054, Hung-yi Lee, Jyh-Shing Roger Jang |
ASRU | 6 |
| 2025 | Enhancing LLM Question Answering with RAG through Dense Vector Search and Re-RankingabstractRetrieval-Augmented Generation (RAG) has emerged as a powerful framework for enhancing Large Language Models (LLMs) by incorporating external knowledge through information retrieval (IR) techniques. However, in question-answering tasks, RAG often retrieves documents that are only semantically similar to the query, which may not provide the most relevant information for generating accurate responses. To address this limitation, we propose an improved retrieval pipeline that combines dense vector search with a re-ranking mechanism to more effectively identify and extract highly relevant knowledge from the retrieved content. We evaluated our approach on two Chinese datasets, TTQA and TMMLU+, using 17 different LLMs. Experimental results show that our method improves performance by up to 21.24% over baseline approaches, particularly on two finance-related subsets, after incorporating domain-specific financial regulations to enhance the knowledge base used in the TMMLU+ dataset. Te-Lun Yang, Jyi-Shane Liu, Yuen-Hsien Tseng, Jyh-Shing Roger Jang, Ming-Ching Chang, Wei-Chao Chen |
AVSS | 4 |
| 2025 | Similarity-based Accent Recognition with Continuous and Discrete Self-supervised Speech RepresentationsabstractThe primary challenge in accent recognition lies in data scarcity due to the high diversity of accents, which make the collection of large-scale training data for each accent almost impossible in practice. To overcome this challenge, we propose a simple solution that leverages both continuous and discrete feature representations from pretrained speech self-supervised learning (SSL) models. Our model is simplified to a linear projection layer and a set of trainable accent class embeddings. Cosine similarity between the accent embeddings and the latent features of an audio sample is used to predict its accent class. This approach enables the model to access features that contain rich accent-related information while reducing the risk of model overfitting. Our method provides a practical and efficient way to tackle accent recognition, especially in low-resource scenarios. Experimental results on English accent recognition show that our best model achieves an accuracy of 84.0% on the AESRC 2020 dataset and an Unweighted Average Recall (UAR) of 50.0% on the VCTK corpus, setting new state-of-the-art results on both datasets. Jun-You Wang, Sheng Li 0010, Li-An Lu, Sydney Chia-Chun Kao, Jyh-Shing Roger Jang |
ICASSP | 5 |
| 2025 | Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
Xuanjun Chen, I-Ming Lin, Lin Zhang 0054, Jiawei Du 0003, Hung-yi Lee, Jyh-Shing Roger Jang |
INTERSPEECH | 7 |
| 2025 | Many-to-Many Singing Performance Style Transfer on Pitch and Energy ContoursabstractSinging voice conversion (SVC) aims to convert the singer identity of a singing voice to that of another singer. However, most existing SVC systems only perform the conversion of timbre information, while leaving other information unchanged. This approach does not consider other aspects of singer identity, particularly a singer's performance style, which is reflected in the pitch (F0) and the energy (volume dynamics) contours of singing. To address this issue, this paper proposes a many-to-many singing performance style transfer system that converts the pitch and energy contours of one singer's style to another singer's. To achieve this target, we utilize two AutoVC-like autoencoders with an information bottleneck to automatically disentangle performance style from other musical contents, one for the pitch contour while another for the energy contour. Experiment results suggested that the proposed model can perform singing performance style transfer in a many-to-many conversion scenario, resulting in improved singer identity similarity to the target singer. Yu-Teng Hsu, Jun-You Wang, Jyh-Shing Roger Jang |
IEEE Signal Process. Lett. | 3 |
| 2024 | Enhancing 3D Human Pose Estimation with Bone Length Adjustment
Chih-Hsiang Hsu, Jyh-Shing Roger Jang |
ACCV (1) | 2 |
| 2024 | Multimodal Transformer Distillation for Audio-Visual SynchronizationabstractAudio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact information. However, it requires high computing resources, making it impractical for real-world applications. This paper proposed an MTD-VocaLiST model, which is trained by our proposed multimodal Transformer distillation (MTD) loss. MTD loss enables MTDVocaLiST model to deeply mimic the cross-attention distribution and value-relation in the Transformer of VocaLiST. Additionally, we harness uncertainty weighting to fully exploit the interaction information across all layers. Our proposed method is effective in two aspects: From the distillation method perspective, MTD loss outperforms other strong distillation baselines. From the distilled model’s performance perspective: 1) MTDVocaLiST outperforms similar-size SOTA models, SyncNet, and Perfect Match models by 15.65% and 3.35%; 2) MTDVocaLiST reduces the model size of VocaLiST by 83.52%, yet still maintaining similar performance. Xuanjun Chen, Chung-Che Wang, Hung-yi Lee, Jyh-Shing Roger Jang |
ICASSP | 5 |
| 2024 | MIR-MLPop: A Multilingual Pop Music Dataset with Time-Aligned Lyrics and AudioabstractWe introduce MIR-MLPop, a publicly available multilingual pop music dataset designed for automatic lyrics transcription and lyrics alignment in polyphonic music. The dataset comprises 90 pop music tracks in Mandarin, Cantonese, and Taiwanese Hokkien, with manually annotated time-aligned lyrics with both characters and pronunciation labels. To the best of our knowledge, this is the first ever singing dataset for Cantonese and Taiwanese Hokkien. In the experiments, using the pretrained Whisper model as the backbone, we develop lyrics transcription and lyrics alignment models for all three languages. Overall, the results are promising for both tasks, but show clear differences among the languages. Our models perform significantly better on languages that have been seen by Whisper during pretraining than on the language unseen by Whisper. This finding highlights the potential challenge in lyrics transcription and alignment for low-resource languages that have not been covered by pretrained speech models. Jun-You Wang, Chung-Che Wang, Chon-In Leong, Jyh-Shing Roger Jang |
ICASSP | 4 |
| 2024 | Neural Codec-based Adversarial Sample Detection for Speaker Verification
Xuanjun Chen, Jiawei Du 0003, Jyh-Shing Roger Jang, Hung-yi Lee |
INTERSPEECH | 4 |
| 2024 | Leveraging Phonemic Transcription and Whisper toward Clinically Significant Indices for Automatic Child Speech Assessment
Yeh-Sheng Lin, Shu-Chuan Tseng, Jyh-Shing Roger Jang |
INTERSPEECH | 3 |
| 2024 | DFADD: The Diffusion and Flow-Matching Based Audio Deepfake DatasetabstractMainstream zero-shot TTS production systems like Voicebox and Seed-TTS achieve human parity speech by leveraging Flow-matching and Diffusion models, respectively. Unfortunately, human-level audio synthesis leads to identity misuse and information security issues. Currently, many anti-spoofing models have been developed against deepfake audio. However, the efficacy of current state-of-the-art anti-spoofing models in countering audio synthesized by diffusion and flow-matching based TTS systems remains unknown. In this paper, we proposed the Diffusion and Flow-matching based Audio Deepfake (DFADD) dataset. The DFADD dataset collected the deepfake audio based on advanced diffusion and flowmatching TTS models. Additionally, we reveal that current anti-spoofing models lack sufficient robustness against highly human-like audio generated by diffusion and flow-matching TTS systems. The proposed DFADD dataset addresses this gap and provides a valuable resource for developing more resilient anti-spoofing models. Jiawei Du 0003, I-Ming Lin, I-Hsiang Chiu, Xuanjun Chen, Wenze Ren, Yu Tsao 0001, Hung-yi Lee, Jyh-Shing Roger Jang |
SLT | 9 |
| 2024 | Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition SystemsabstractSpeech emotion recognition (SER) is an essential technology for human-computer interaction systems. However, the previous study reveals that 80.77% of SER papers yield results that cannot be reproduced on the well-known IEMOCAP dataset. The main reason for reproducibility challenges is that the database did not provide standard data splits (e.g., train, development, and test sets). Prior papers could define its partition, but they did not provide details of the partition or source code for processing the partition. Therefore, this work aims to make SER open and reproducible to everyone. We develop the EMO-SUPERB, shorted for EMOtion Speech Universal PERformance Benchmark, including a user-friendly codebase to leverage 16 state-of-the-art (SOTA) speech self-supervised learning models for exhaustive evaluation plus one SOTA SER model across 6 open-source SER datasets in English and Chinese. We make all resources open-source to facilitate future developments in SER. Researchers can easily upload their systems or datasets to EMO-SUPERB, and we name the project “Open-Emotion”. Huang-Cheng Chou, Kai-Wei Chang 0001, Lucas Goncalves, Jiawei Du 0003, Jyh-Shing Roger Jang, Chi-Chun Lee, Hung-yi Lee |
SLT | 6 |
| 2024 | WC-SBERT: Zero-Shot Topic Classification Using SBERT and Light Self-Training on Wikipedia CategoriesabstractIn natural language processing (NLP), zero-shot topic classification requires machines to understand the contextual meanings of texts in a downstream task without using the corresponding labeled texts for training, which is highly desirable for various applications. In this article, we propose a novel approach to construct a zero-shot task-specific model called WC-SBERT with satisfactory performance. The proposed approach is highly efficient since it uses light self-training requiring target labels (target class names of downstream tasks) only, which is distinct from other research that uses both the target labels and the unlabeled texts for training. In particular, during the pre-training stage, WC-SBERT uses contrastive learning with multiple negative ranking losses to construct the pre-trained model based on the similarity between Wiki categories. For the self-training stage, online contrastive loss is utilized to reduce the distance between a target label and Wiki categories of similar Wiki pages to the label. Experimental results indicate that compared to existing self-training models, WC-SBERT achieves rapid inference on approximately 6.45 million Wiki text entries by utilizing pre-stored Wikipedia text embeddings, significantly reducing inference time per sample by a factor of 2,746 to 16,746. During the fine-tuning step, the time required for each sample is reduced by a factor of 23–67. Overall, the total training time shows a maximum reduction of 27.5 times across different datasets. Most importantly, our model has achieved state-of-the-art (SOTA) accuracy on two of the three commonly used datasets for evaluating zero-shot classification, namely the AG News (0.84) and Yahoo! Answers (0.64) datasets. The code for WC-SBERT is publicly available on GitHub, 1 and the dataset can also be accessed on Hugging Face. 2 Te-Yu Chi, Jyh-Shing Roger Jang |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | Zero-Shot Singing Voice Synthesis from Musical ScoreabstractZero-shot singing voice synthesis (SVS), the task to synthesize the singing voice of an arbitrary target singer, has gained increasing attentions in the past few years. Several recently proposed systems have demonstrated promising results on this task. However, these systems require detailed musical features at the frame level as the musical content. To deal with this issue, we propose a model that performs zero-shot SVS with only musical score as the musical content condition. To help model training, we build an acoustic encoder that extracts linguistic features from audio, and train it with the lyrics transcription objective. The output of the acoustic encoder serves as an alternative to the musical score, allowing the SVS model to learn from weakly labeled data. Results suggest that the proposed method outperforms baseline semi-supervised method in both subjective and objective tests. Jun-You Wang, Hung-yi Lee, Jyh-Shing Roger Jang, Li Su 0004 |
ASRU | 3 |
| 2023 | Adapting Pretrained Speech Model for Mandarin Lyrics Transcription and AlignmentabstractThe tasks of automatic lyrics transcription and lyrics alignment have witnessed significant performance improvements in the past few years. However, most of the previous works only focus on English in which large-scale datasets are available. In this paper, we address lyrics transcription and alignment of polyphonic Mandarin pop music in a low-resource setting. To deal with the data scarcity issue, we adapt pretrained Whisper model and fine-tune it on a monophonic Mandarin singing dataset. With the use of data augmentation and source separation model, results show that the proposed method achieves a character error rate of less than 18% on a Mandarin polyphonic dataset for lyrics transcription, and a mean absolute error of 0.071 seconds for lyrics alignment. Our results demonstrate the potential of adapting a pretrained speech model for lyrics transcription and alignment in low-resource scenarios. Jun-You Wang, Chon-In Leong, Li Su 0004, Jyh-Shing Roger Jang |
ASRU | 5 |
| 2023 | Noise-Robust Bandwidth Expansion for 8K Speech Recordings
Yin-Tse Lin, Bo-Hao Su, Chi-Han Lin, Shih-Chan Kuo, Jyh-Shing Roger Jang, Chi-Chun Lee |
INTERSPEECH | 5 |
| 2023 | Training a Singing Transcription Model Using Connectionist Temporal Classification Loss and Cross-Entropy LossabstractIn this paper, we propose a method that uses a combination of the Connectionist Temporal Classification (CTC) loss and the cross-entropy loss to train a note-level singing transcription model. By considering the task as predicting a note sequence of the input audio, we can compute the CTC loss between the prediction and the groundtruth note sequence, and further use it with the traditional cross-entropy loss to optimize the transcription model. By comparing the proposed method with a baseline that only utilizes the cross-entropy loss, the results show improved model performance on all the evaluation metrics. Furthermore, using the CTC loss allows the transcription model to learn from weakly labeled data, which is easier to annotate than traditional strongly labeled data. Moreover, we point out the issue of the intrinsic global time shift on the onset labels between datasets. By automatically estimating and calibrating the global time shift of the training dataset, the performance of the singing transcription model is then not affected by the global time shift in the cross-dataset evaluation scenario. Jun-You Wang, Jyh-Shing Roger Jang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer ModelabstractIn this paper, we propose a new dataset named EGDB, that contains transcriptions of the electric guitar performance of 240 tablatures rendered with different tones. Moreover, we benchmark the performance of two well-known transcription models proposed originally for the piano on this dataset, along with a multi-loss Transformer model that we newly propose. Our evaluation on this dataset and a separate set of real-world recordings demonstrate the influence of timbre on the accuracy of guitar sheet transcription, the potential of using multiple losses for Transformers, as well as the room for further improvement for this task. Yu-Hua Chen, Wen-Yi Hsiao, Tsu-Kuang Hsieh, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
ICASSP | 4 |
| 2022 | Push-Pull: Characterizing the Adversarial Robustness for Audio-Visual Active Speaker DetectionabstractAudio-visual active speaker detection (AVASD) is well-developed, and now is an indispensable front-end for several multi-modal applications. However, to the best of our knowledge, the adversarial robustness of AVASD models hasn't been investigated, not to mention the effective defense against such attacks. In this paper, we are the first to reveal the vulnerability of AVASD models under audio-only, visual-only, and audio-visual adversarial attacks through extensive experiments. What's more, we also propose a novel audio-visual interaction loss (AVIL) for making attackers difficult to find feasible adversarial examples under an allocated attack budget. The loss aims at pushing the inter-class embeddings to be dispersed, namely non-speech and speech clusters, sufficiently disentangled, and pulling the intra-class embeddings as close as possible to keep them compact. Experimental results show the AVIL outperforms the adversarial training by 33.14 mAP (%) under multi-modal attacks. Xuanjun Chen, Helen M. Meng, Hung-yi Lee, Jyh-Shing Roger Jang |
SLT | 5 |
| 2021 | Mandarin Electrolaryngeal Speech Voice Conversion with Sequence-to-Sequence ModelingabstractThe electrolaryngeal speech (EL speech) is typically spoken with an electrolarynx device that generates excitation signals to substitute human vocal fold vibrations. Because the excitation signals cannot perfectly characterize sound sources generated by vocal folds, the naturalness and intelligibility of the EL speech are inevitably worse than that of the natural speech (NL speech). To improve speech naturalness, statistical models, such as Gaussian mixture models and deep-learning-based models, have been employed for EL speech voice conversion (ELVC). The ELVC task aims to convert EL speech into NL speech through an ELVC model. To implement a frame-wise ELVC system, accurate feature alignment is crucial for model training. However, the abnormal acoustic characteristics of the EL speech cause misalignments and accordingly limit the ELVC performance. To address this issue, we propose a novel ELVC system based on sequence-to-sequence (seq2seq) modeling with text-to-speech (TTS) pretraining. The seq2seq model involves an attention mechanism to concurrently perform representation learning and alignment. Meanwhile, TTS pretraining provides efficient training with limited data. Experimental results show that the proposed ELVC system yields notable improvements in terms of standardized evaluation metrics and subjective listening tests over a well-known frame-wise ELVC system. Ming-Chi Yen, Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Shu-Wei Tsai, Yu Tsao 0001, Tomoki Toda, Jyh-Shing Roger Jang, Hsin-Min Wang |
ASRU | 8 |
| 2021 | On the Preparation and Validation of a Large-Scale Dataset of Singing TranscriptionabstractThis paper proposes a large-scale dataset for singing transcription, along with some methods for fine-tuning and validating its contents. The dataset is named MIR-ST500, which consists of more than 160,000 notes from 500 pop songs. To create this large-scale dataset, we set some labeling criteria and ask non-experts to label notes. We also perform some adjustments on the annotation to correct minor errors. Finally, to validate the dataset, we train a singing transcription model on MIR-ST500 dataset and evaluate it on various datasets. The result shows that we can certainly construct a better singing transcription model for various purposes using MIR-ST500, which is properly labeled and validated. Jun-You Wang, Jyh-Shing Roger Jang |
ICASSP | 2 |
| 2020 | Fast Tensor Factorization for Large-Scale Context-Aware Recommendation from Implicit FeedbackabstractThis paper presents a fast Tensor Factorization (TF) algorithm for context-aware recommendation from implicit feedback. For such a recommendation problem, the observed data indicate the (positive) association between users and items in some given contexts. For better accuracy, it has been shown essential to include unobserved data that indicate the negative user-item-context associations. As such unobserved data greatly outnumber the observed ones, for efficiency existing algorithms usually use only a small part of the unobserved data for model training. We show in this paper that it is possible, and beneficial, to use all the unobserved data in training a TF based context-aware recommender system. This is achieved by two technical innovations. First, we scrutinize the matrix computation of the closed-form solution and accelerate the computation by memorizing the repetitive computation. Second, we further boost the generalization and scalability by dropping out some pairwise interactions when updating user, item or context factors to prevent overfitting and to reduce the training time. The resulting whole-data based learning algorithm, referred to as DropTF in the paper, is efficient and scale well. Our evaluation on two small benchmark datasets and a million-scale large dataset demonstrates improved accuracy over some existing algorithms for context-aware recommendation. Szu-Yu Chou, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
IEEE Trans. Big Data | 2 |
| 2020 | AutoRhythm: A Music Game With Automatic Hit-Timing Generation and Percussion IdentificationabstractThis article describes a music rhythm game called AutoRhythm, which can automatically generate the hit timing as game contents from a given piece of music, and identify user-defined percussion of real objects in real time for gameplay. More specifically, AutoRhythm can generate the hit timing of a piece of music based on onset detection, so the user can use any music from their own collection for the rhythm game. Moreover, to make the game more realistic, AutoRhythm also allows the user to interact with the game via any object that can produce percussion sounds, such as a pen or a chopstick hitting against a table. AutoRhythm can identify the percussions in real time to replace tapping on the screen. This real-time user percussion identification is achieved based on the frame-based power spectrum of the filtered recording after background music reduction, which is performed based on the concept of active noise cancellation, with the estimated noisy playback music being subtracted from the original recording. Based on a test data set of 100 recordings, our experiment indicates that our system can achieve an F-measure of 78.22%, which outperforms other well-known classifiers and is quite satisfactory for the purpose of gameplay. Tzu-Chun Yeh, Jyh-Shing Roger Jang |
IEEE Trans. Games | 2 |
| 2020 | Backpropagation With $N$ -D Vector-Valued Neurons Using Arbitrary Bilinear ProductsabstractVector-valued neural learning has emerged as a promising direction in deep learning recently. Traditionally, training data for neural networks (NNs) are formulated as a vector of scalars; however, its performance may not be optimal since associations among adjacent scalars are not modeled. In this article, we propose a new vector neural architecture called the Arbitrary BIlinear Product NN (ABIPNN), which processes information as vectors in each neuron, and the feedforward projections are defined using arbitrary bilinear products. Such bilinear products can include circular convolution, 7-D vector product, skew circular convolution, reversed-time circular convolution, or other new products that are not seen in the previous work. As a proof-of-concept, we apply our proposed network to multispectral image denoising and singing voice separation. Experimental results show that ABIPNN obtains substantial improvements when compared to conventional NNs, suggesting that associations are learned during training. Zhe-Cheng Fan, Tak-Shing Chan, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Improving ResNet-based Feature Extractor for Face Recognition via Re-ranking and Approximate Nearest NeighborabstractThis paper proposes a framework for face recognition based on feature extractor from ResNet, together with other steps for performance improvement, including face detection, face alignment, face verification/identification, and re-ranking via Approximate Nearest Neighbor Search (ANNS). First, we evaluate two face detection algorithms, MTCNN, and FaceBoxes on three common face detection benchmarks, and then summarize the best usage scenario for each approach. Second, with certain preprocessing and postprocessing, our system selects the ResNet-based feature extractor, which achieves 99.33% verification accuracy on the LFW benchmark. Third, we use the penalty curve to determine the best configuration and obtain improved results of face verification. Based on the proposed preprocessing and post-processing, our method not only boosts accuracy from 84.3% to 86.5% in large inter-class variation datasets (CASIA - WebFace) but improves Rank-l accuracy from 86.6% to 87.7% in large intra-class variation datasets (FG-NET). Sheng-Hsing Hsiao, Jyh-Shing Roger Jang |
AVSS | 2 |
| 2019 | K-Same-Siamese-GAN: K-Same Algorithm with Generative Adversarial Network for Facial Image De-identification with Hyperparameter Tuning and Mixed Precision TrainingabstractFor a data holder, such as a hospital or a government entity, who has a privately held collection of personal data, in which the revealing and/or processing of the personal identifiable data is restricted and prohibited by law. Then, “how can we ensure the data holder does conceal the identity of each individual in the imagery of personal data while still preserving certain useful aspects of the data after de-identification?” becomes a challenge issue. In this work, we propose an approach towards high-resolution facial image de-identification, called k-Same-Siamese-GAN, which leverages the k-Same-Anonymity mechanism, the Generative Adversarial Network, and the hyperparameter tuning methods. Moreover, to speed up model training and reduce memory consumption, the mixed precision training technique is also applied to make kSS-GAN provide guarantees regarding privacy protection on close-form identities and be trained much more efficiently as well. Finally, to validate its applicability, the proposed work has been applied to actual datasets - RafD and CelebA for performance testing. Besides protecting privacy of high-resolution facial images, the proposed system is also Justified for its ability in automating parameter tuning and breaking through the limitation of the number of adjustable parameters. Yi-Lun Pan, Min-Jhih Haung, Kuo-Teng Ding, Ja-Ling Wu, Jyh-Shing Roger Jang |
AVSS | 5 |
| 2019 | Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound RecognitionabstractIn this paper, we introduce a novel attentional similarity module for the problem of few-shot sound recognition. Given a few examples of an unseen sound event, a classifier must be quickly adapted to recognize the new sound event without much fine-tuning. The proposed attentional similarity module can be plugged into any metric-based learning method for few-shot learning, allowing the resulting model to especially match related short sound events. Extensive experiments on two datasets show that the proposed module consistently improves the performance of five different metric-based learning methods for few-shot sound recognition. The relative improvement ranges from +4.1% to +7.7% for 5-shot 5-way accuracy for the ESC-50 dataset, and from +2.1% to +6.5% for noiseESC-50. Qualitative results demonstrate that our method contributes in particular to the recognition of transient sound events. Szu-Yu Chou, Kai-Hsiang Cheng, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
ICASSP | 3 |
| 2019 | Deep Cyclic Group NetworksabstractWe propose a new network architecture called deep cyclic group network (DCGN) that uses the cyclic group algebra for convolutional vector-neuron learning. The input to DCGN is a three-way tensor, where the mode-3 dimension corresponds to the dimensionality of the input data, e.g., three for RGB images. To handle vector-valued inputs, we replace scalar multiplication with circular convolution for the feedforward and backpropagation processes. As a result, every feature map and kernel map is a three-way tensor with the same mode-3 dimension as the input data. This way, DCGN may capture more of the relations among different data dimensions, especially for regression tasks where the target output has the same dimensionality as the input data. Moreover, DCGN can deal with input data of arbitrary dimensions, a property that existing architectures such as deep complex networks and deep quaternion networks (DQN) lack. Experiments show that DCGN indeed performs better than convolutional neural networks and DQN for two regression tasks, namely color image inpainting and multispectral image denoising. Zhe-Cheng Fan, Tak-Shing Chan, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
IJCNN | 4 |
| 2019 | An effective method for audio-to-score alignment using onsets and modified constant Q spectra
Chun-Ta Chen, Jyh-Shing Roger Jang |
Multim. Tools Appl. | 2 |
| 2018 | SVSGAN: Singing Voice Separation Via Generative Adversarial NetworkabstractSeparating two sources from an audio mixture is an important task with many applications. It is a challenging problem since only one signal channel is available for analysis. In this paper, we propose a novel framework for singing voice separation using the generative adversarial network (GAN) with a time-frequency masking function. The mixture spectra is considered to be a distribution and is mapped to the clean spectra which is also considered a distribution. The approximation of distributions between mixture spectra and clean spectra is performed during the adversarial training process. In contrast with current deep learning approaches for source separation, the parameters of the proposed framework are first initialized in a supervised setting and then optimized by the training procedure of GAN in an unsupervised setting. Experimental results on three datasets (MIR-1K, iKala and DSD100) show that performance can be improved by the proposed framework consisting of conventional networks. Zhe-Cheng Fan, Yen-Lin Lai, Jyh-Shing Roger Jang |
ICASSP | 3 |
| 2018 | Learning to Recognize Transient Sound Events using Attentional SupervisionabstractMaking sense of the surrounding context and ongoing events through not only the visual inputs but also acoustic cues is critical for various AI applications. This paper presents an attempt to learn a neural network model that recognizes more than 500 different sound events from the audio part of user generated videos (UGV). Aside from the large number of categories and the diverse recording conditions found in UGV, the task is challenging because a sound event may occur only for a short period of time in a video clip. Our model specifically tackles this issue by combining a main subnet that aggregates information from the entire clip to make clip-level predictions, and a supplementary subnet that examines each short segment of the clip for segment-level predictions. As the labeled data available for model training are typically on the clip level, the latter subnet learns to pay attention to segments selectively to facilitate attentional segment-level supervision. We call our model the M&mnet, for it leverages both “M”acro (clip-level) supervision and “m”icro (segment-level) supervision derived from the macro one. Our experiments show that M&mnet works remarkably well for recognizing sound events, establishing a new state-of-theart for DCASE17 and AudioSet data sets. Qualitative analysis suggests that our model exhibits strong gains for short events. In addition, we show that the micro subnet is computationally light and we can use multiple micro subnets to better exploit information in different temporal scales. Szu-Yu Chou, Jyh-Shing Roger Jang, Yi-Hsuan Yang |
IJCAI | 2 |
| 2018 | A hierarchical linguistic information-based model of English prosody: L2 data analysis and implications for computer-assisted language learning
Chao-yu Su, Chiu-yu Tseng, Jyh-Shing Roger Jang, Tanya Visceglia |
Comput. Speech Lang. | 3 |
| 2017 | Conditional preference nets for user and item cold start problems in music recommendationabstractA great amount of data is usually needed for a recommender system to learn the associations between users and items. However, in practical applications, new users and new items emerge everyday, and the system has to react to them promptly. The ability to recommend proper items to new users affects the users' first impression and accordingly the retention rate, whereas recommending new items to proper users contributes to the freshness of the recommendation. In this paper, we propose a deep learning model called the conditional preference nets (CPN) to deal with both new users and items under the same model framework. CPN employs an introductory user survey to learn about new users, and content features automatically extracted from items for the item side. Through a new idea called the preference vectors and an existing content embedding technique, the same model can capitalize the observed associations between known (i.e. old) users and items, thereby benefiting from the cumulative knowledge of user behavior gained over time. We validate the superiority of CPN over prior arts using the Million Song Dataset. We also demonstrate how CPN allows a user to pick either genres, artists, or the combination of them in the introductory survey. Szu-Yu Chou, Li-Chia Yang, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
ICME | 4 |
| 2016 | An efficient method for polyphonic audio-to-score alignment using onset detection and constant Q transformabstractThis paper proposes an innovative method that aligns a polyphonic audio recording of music to its corresponding symbolic score. In the first step, we perform onset detection and then apply constant Q transform around each onset. A similarity matrix is computed by using a scoring function which evaluates the similarity between notes in the music score and onsets in the audio recording. At last, we use dynamic programming to extract the best alignment path in the similarity matrix. We compared two onset detectors and two note matching methods. Our method is more efficient and has higher precision than the traditional chroma-based DTW method. Our algorithm achieved the best precision, which are 10% higher than the compared traditional algorithm when the tolerance window is 50 ms. Chun-Ta Chen, Jyh-Shing Roger Jang, Wen-Shan Liu, Chi-Yao Weng |
ICASSP | 2 |
| 2016 | Addressing Cold Start for Next-song RecommendationabstractThe cold start problem arises in various recommendation applications. In this paper, we propose a tensor factorization-based algorithm that exploits content features extracted from music audio to deal with the cold start problem for the emerging application next-song recommendation. Specifically, the new algorithm learns sequential behavior to predict the next song that a user would be interested in based on the last song the user just listened to. A unique characteristic of the algorithm is that it learns and updates the mapping between the audio feature space and the item latent space each time during the iterations of the factorization process. This way, the content features can be better exploited in forming the latent features for both users and items, leading to more effective solutions for cold-start recommendation. Evaluation on a large-scale music recommendation dataset shows that the recommendation result of the proposed algorithm exhibits not only higher accuracy but also better novelty and diversity, suggesting its applicability in helping a user explore new items in next-item recommendation. Our implementation is available at https://github.com/fearofchou/ALMM. Szu-Yu Chou, Yi-Hsuan Yang, Jyh-Shing Roger Jang, Yu-Ching Lin |
RecSys | 3 |
| 2015 | Vocal activity informed singing voice separation with the iKala datasetabstractA new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available iKala dataset that features longer durations and higher quality than the existing MIR-1K dataset for singing voice separation. Part of it will be used in the MIREX Singing Voice Separation task. Experimental results on both the MIR-1K dataset and the new iKala dataset confirmed that the more informed the algorithm is, the better the separation results are. Tak-Shing Chan, Tzu-Chun Yeh, Zhe-Cheng Fan, Hung-Wei Chen, Li Su 0004, Yi-Hsuan Yang, Jyh-Shing Roger Jang |
ICASSP | 7 |
| 2015 | AutoRhythm: A music game with automatic hit-time generation and percussion identificationabstractThis paper describes a music rhythm game called AutoRhythm, which can automatically generate the hit time for a rhythm game from a given piece of music, and identify user-defined percussions in real time when a user is playing the game. More specifically, AutoRhythm can automatically generate the hit time of the given music, either locally or via server-based computation, such that users can use the user-supplied music for the game directly. Moreover, to make the rhythm game more realistic, AutoRhythm allows users to interact with the game via any objects that can produce percussion sound, such as a pen or a chopstick hitting on the table. AutoRhythm can identify the percussions in real time while the music is playing. The identification is based on the power spectrum of each frame of the recording which combines percussions and playback music. Based on a test dataset of 12 recordings (with 2455 percussions of 4 types), our experiment indicates an F-measure of 96.79%, which is satisfactory for the purpose of the game. The flexibility of being able to use any user-supplied music for the game and to identify user-defined percussions from any objects available at hand makes the game innovative and unique of its kind. Pei-Pei Chen, Tzu-Chun Yeh, Jyh-Shing Roger Jang, Wenshan Liou |
ICME | 3 |
| 2015 | Automatic Music Mood Classification Based on Timbre and Modulation FeaturesabstractIn recent years, many short-term timbre and long-term modulation features have been developed for content-based music classification. However, two operations in modulation analysis are likely to smooth out useful modulation information, which may degrade classification performance. To deal with this problem, this paper proposes the use of a two-dimensional representation of acoustic frequency and modulation frequency to extract joint acoustic frequency and modulation frequency features. Long-term joint frequency features, such as acoustic-modulation spectral contrast/valley (AMSC/AMSV), acoustic-modulation spectral flatness measure (AMSFM), and acoustic-modulation spectral crest measure (AMSCM), are then computed from the spectra of each joint frequency subband. By combining the proposed features, together with the modulation spectral analysis of MFCC and statistical descriptors of short-term timbre features, this new feature set outperforms previous approaches with statistical significance. Jia-Min Ren, Ming-Ju Wu, Jyh-Shing Roger Jang |
IEEE Trans. Affect. Comput. | 3 |
| 2015 | Automatic Pronunciation Scoring with Score Combination by Learning to Rank and Class-Normalized DP-Based QuantizationabstractThis paper proposes an automatic pronunciation scoring framework using learning to rank and class-normalized, dynamic-programming-based quantization. The goal is to train a model that is able to grade the pronunciation of a second language learner, such that the predicted score is as close as possible to the one given by a human teacher. Under this framework, each utterance is given a score of 1 to 5 by human raters, which is treated as a ground truth rank for the training algorithm. The corpus was rated by qualified English teachers in Taiwan (nonnative speakers). Nine phone-level scores are computed and converted into word-level scores through four conversion methods. We select the 16 best performing scores as the input features to train the learning-to-rank function. The output of the function is then quantized to a discrete rank on a 1-5 scale. The quantization is done with class normalization to alleviate the problem of data imbalance over different classes. Experimental results show that the proposed framework achieves a higher correlation to the human scores than other methods, along with higher accuracy in detecting instances of mispronunciation. We also release a new version of our nonnative corpus with human rankings. Liang-Yu Chen 0006, Jyh-Shing Roger Jang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Improving Query-by-Singing/Humming by Combining Melody and Lyric InformationabstractThis paper proposes a novel method for improving query-by-singing/humming systems by using both melody and lyric information. First, singing/humming discrimination is performed to distinguish between singing and humming queries, which is achieved by considering the similarity between acoustic models. For the humming queries, a pitch-only melody recognition method that was ranked first among the MIREX (Music Information Retrieval Evaluation eXchange) query-by-singing/humming task submissions is applied. For the singing queries, a lyric similarity is computed using speech recognition techniques; the computed similarity is subsequently combined with the melody distance to exploit additional information in the lyrics. Several methods for combining melody distance and lyric similarity are investigated. Under the optimal experimental settings, the proposed query-by-singing/humming system achieves 51.19% error rate reduction for the top-10 retrieved results, indicating the feasibility of the proposed method. Chung-Che Wang, Jyh-Shing Roger Jang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Audio Musical Dice Game: A User-Preference-Aware Medley Generating SystemabstractThis article proposes a framework for creating user-preference-aware music medleys from users' music collections. We treat the medley generation process as an audio version of a musical dice game. Once the user's collection has been analyzed, the system is able to generate various pleasing medleys. This flexibility allows users to create medleys according to the specified conditions, such as the medley structure or the must-use clips. Even users without musical knowledge can compose medley songs from their favorite tracks. The effectiveness of the system has been evaluated through both objective and subjective experiments on individual components in the system. Yin-Tzu Lin, I-Ting Liu, Jyh-Shing Roger Jang, Ja-Ling Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2015 | Combining Acoustic and Multilevel Visual Features for Music Genre ClassificationabstractMost music genre classification approaches extract acoustic features from frames to capture timbre information, leading to the common framework of bag-of-frames analysis. However, time-frequency analysis is also vital for modeling music genres. This article proposes multilevel visual features for extracting spectrogram textures and their temporal variations. A confidence-based late fusion is proposed for combining the acoustic and visual features. The experimental results indicated that the proposed method achieved an accuracy improvement of approximately 14% and 2% in the world's largest benchmark dataset (MASD) and Unique dataset, respectively. In particular, the proposed approach won the Music Information Retrieval Evaluation eXchange (MIREX) music genre classification contests from 2011 to 2013, demonstrating the feasibility and necessity of combining acoustic and visual features for classifying music genres. Ming-Ju Wu, Jyh-Shing Roger Jang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Improved score-performance alignment algorithms on polyphonic musicabstractAutomated symbolic music alignment is a challenging task due to the variation of performance by different performers. It becomes more complicated when dealing with polyphonic music because note events could occur at the same time. The goal of this study is to find an efficient algorithm for aligning two polyphonic symbolic representations (MIDI files, for instance) of the same music. To this end, we design two methods for such score-performance alignment that matches the performance with its corresponding score. The first method applies a string matching algorithm based on dynamic programming. The second method is based on the principle of "divide and conquer" that performs efficient alignment recursively. To evaluate the algorithms, we have collected a set of 21 MIDI pairs of classic piano performance with human corrected note-level mapping as ground truth. We have released the dataset as a public resource. Both the proposed algorithms achieved a precision and recall higher than 96% in our experiment, outperforming the most recently proposed method [7] in the literature. Besides, the execution time of proposed methods is much faster the method of [7]. Chun-Ta Chen, Jyh-Shing Roger Jang, Wenshan Liou |
ICASSP | 2 |
| 2014 | Bridging Music via Sound EffectsabstractThe prevalence of digital technologies allows people to easily create and share their own media contents, but sometimes we do not have handy tools to manipulate the media we want to create. For example, while creating personal films, a user may separately find the music segment that matches each part of the video, and then concatenate the segments to create the soundtrack that matches the visual contents along the timeline. However, one may lose the temporal coherence between consecutive music segments if the user just choose music segments that best match parts of the video contents. In this study, we focus on how to smoothly connect not-so-coherent music clips and make the transition natural and pleasant to hear. In particular, we improve the temporal smoothness by bar alignment and dual tempo adjustment. To further fit in with the transition between clips, we incorporate "sound effect insertion" which is a commonly used technique in popular song composition/remixing. In order to provide pleasant listening experience and systematically analyze the effectiveness of the proposed music bridging method, we have conducted specifically designed experiments to collect subjective opinions and reduce the cognitive loads of the participants. The experimental results indicate that with proper arrangement for creating smooth transition via tempo adjustment and sound effect insertion, the listening experience can be largely enhanced. Yin-Tzu Lin, Chuan-Lung Lee, Jyh-Shing Roger Jang, Ja-Ling Wu |
ISM | 3 |
| 2014 | Phone Boundary Annotation in Conversational Speech
Yi-Fen Liu, Shu-Chuan Tseng, Jyh-Shing Roger Jang |
LREC | 3 |
| 2012 | Accelerating query by singing/humming on GPU: Optimization for web deploymentabstractThis paper presents the use of GPU for implementing a parallelized comparison method of linear scaling in a query by singing/humming system, which can compare a user's acoustic input to the database containing about 13,000 songs. We focus on the comparison from anywhere in a song, and the optimum setting is found through 3 different schemes of parallelization. With a speedup factor of 66, the proposed scheme with the optimum setting has been successfully implemented in a public QBSH system that is available from the internet. Chung-Che Wang, Chieh-Hsing Chen, Chin-Yang Kuo, Li-Ting Chiu, Jyh-Shing Roger Jang |
ICASSP | 5 |
| 2012 | A hybrid approach to singing pitch extraction based on trend estimation and hidden Markov modelsabstractIn this paper, we propose a hybrid method for singing pitch extraction from polyphonic audio music. We have observed several kinds of pitch errors made by a previously proposed algorithm based on trend estimation. We also noticed that other pitch tracking methods tend to have other types of pitch error. Then it becomes intuitive to combine the results of several pitch trackers to achieve a better accuracy. In this paper, we adopt 3 methods as a committee to determine the pitch, including the trend-estimation-based method for forward and backward signals, and training-based HMM method. Experimental results demonstrate that the proposed approach outperforms the best algorithm for the task of audio melody extraction in MIREX 2010. Tzu-Chun Yeh, Ming-Ju Wu, Jyh-Shing Roger Jang, Wei-Lun Chang, I-Bin Liao |
ICASSP | 3 |
| 2012 | Improvement in Automatic Pronunciation Scoring using Additional Basic Scores and Learning to Rank
Liang-Yu Chen 0006, Jyh-Shing Roger Jang |
INTERSPEECH | 2 |
| 2012 | A Tandem Algorithm for Singing Pitch Extraction and Voice Separation From Music AccompanimentabstractSinging pitch estimation and singing voice separation are challenging due to the presence of music accompaniments that are often nonstationary and harmonic. Inspired by computational auditory scene analysis (CASA), this paper investigates a tandem algorithm that estimates the singing pitch and separates the singing voice jointly and iteratively. Rough pitches are first estimated and then used to separate the target singer by considering harmonicity and temporal continuity. The separated singing voice and estimated pitches are used to improve each other iteratively. To enhance the performance of the tandem algorithm for dealing with musical recordings, we propose a trend estimation algorithm to detect the pitch ranges of a singing voice in each time frame. The detected trend substantially reduces the difficulty of singing pitch detection by removing a large number of wrong pitch candidates either produced by musical instruments or the overtones of the singing voice. Systematic evaluation shows that the tandem algorithm outperforms previous systems for pitch extraction and singing voice separation. Chao-Ling Hsu, DeLiang Wang, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Discovering Time-Constrained Sequential Patterns for Music Genre ClassificationabstractA music piece can be considered as a sequence of sound events which represent both short-term and long-term temporal information. However, in the task of automatic music genre classification, most of text-categorization-based approaches could only capture temporal local dependencies (e.g., unigram and bigram-based occurrence statistics) to represent music contents. In this paper, we propose the use of time-constrained sequential patterns (TSPs) as effective features for music genre classification. First of all, an automatic language identification technique is performed to tokenize each music piece into a sequence of hidden Markov model indices. Then TSP mining is applied to discover genre-specific TSPs, followed by the computation of occurrence frequencies of TSPs in each music piece. Finally, support vector machine classifiers are employed based on these occurrence frequencies to perform the classification task. Experiments conducted on two widely used datasets for music genre classification, GTZAN and ISMIR2004Genre, show that the proposed method can discover more discriminative temporal structures and achieve a better recognition accuracy than the unigram and bigram-based statistical approach. Jia-Min Ren, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A trend estimation algorithm for singing pitch detection in musical recordingsabstractDetecting pitch values for singing voice in the presence of music accompaniment is challenging but useful for many applications. We propose a trend estimation algorithm to detect the pitch ranges of a singing voice in each time frame. The detected trend substantially reduces the difficulty of singing pitch detection by reducing a large number of wrong pitch candidates either produced by musical instruments or the overtones of the singing voice. The proposed algorithm can be applied to improve the performance of singing pitch detection. Quantitative evaluations show that proposed trend estimation improves an existing algorithm significantly. The results from the MIREX 2010 competition show that our system achieves the best overall raw-pitch accuracy for vocal songs. Chao-Ling Hsu, DeLiang Wang, Jyh-Shing Roger Jang |
ICASSP | 3 |
| 2011 | Time-constrained sequential pattern discovery for music genre classificationabstractMusic consists of both local and long-term temporal information. However, for a genre classification task, most of the text categorization based approaches only capture local temporal dependences (e.g. statistics of unigrams and bigrams). In our previous work, we use sequential patterns to capture long-term temporal information from the tokenized sequences of music pieces. In this paper, we propose the use of time-constrained sequential patterns (TSPs) to enhance the mined long-term temporal structures so that these TSPs can fit more closely to the human perception. Experimental results show that the proposed method can discover more temporal structures than statistical language modeling approaches and achieves better recognition accuracy. Jia-Min Ren, Jyh-Shing Roger Jang |
ICASSP | 2 |
| 2011 | A Kernel Framework for Content-Based Artist Recommendation System in MusicabstractThis paper proposes a content-based artist recommendation framework which learns relationships between users' preference and music contents through ordinal regression. In particular, an artist is characterized by the parameters of its corresponding acoustical model which is adapted from a universal background model. These artist-specific acoustic features together with their preference rankings are then used as input vectors for the proposed order preserving projection (OPP) algorithm which tries to find a suitable subspace such that the desired ranking order of the data after projection can be kept as much as possible. The proposed linear OPP can be kernelized to learn the nonlinear relationship between music contents and users' artist rank orders. Under the proposed framework of kernelized OPP (KOPP), we can derive the nonlinear relationship and, more importantly, efficiently fuse acoustic and symbolic features obtained from the artist recommended meta-data. Experimental results demonstrate that OPP attains comparable results with those obtained with a conventional ordinal regression method, Prank. Moreover, by exploring the nonlinear relationship among training examples and combining acoustic and symbolic features, KOPP outperforms previous approaches to artist recommendation. Zhi-Sheng Chen, Jyh-Shing Roger Jang |
IEEE Trans. Multim. | 2 |
| 2010 | Using Tangible Learning Companions in English EducationabstractThe tangible learning companions were made to interact with learners by speaking up to 15 basic daily greetings such as “What's your name?”, “Where do you live?”, “How is the weather today?” and so forth. In this study, we did the investigation and evaluation of learners' attitude and learning effect by using tangible learning companions in learning English conversation. There were 61 learners participated in this study. The results of this study revealed the learning effect of using tangible learning companions in English conversation was generally positive for the elementary school. The speaking ability of young students could be improved. Besides, the learners' learning engagement and motivation was enhanced after they practiced English conversation with tangible learning companions. Shelley S. C. Young, Jyh-Shing Roger Jang |
ICALT | 3 |
| 2010 | On the use of sequential patterns mining as temporal features for music genre classificationabstractMusic can be viewed as a sequence of sound events. However, most of current approaches to genre classification either ignore temporal information or only capture local structures within the music under analysis. In this paper, we propose the use of a song tokenization method (which transforms the music into a sequence of units) in conjunction with a data mining technique for investigating the long-term structures (also known as sequential patterns) for music genre classification. Experimental results show that the introduction of sequential patterns can effectively outperform previous approach that considers local temporal features only for music genre classification. Jia-Min Ren, Zhi-Sheng Chen, Jyh-Shing Roger Jang |
ICASSP | 3 |
| 2010 | Automatic pronunciation scoring using learning to rank and DP-based score segmentation
Liang-Yu Chen 0006, Jyh-Shing Roger Jang |
INTERSPEECH | 2 |
| 2010 | Coping imbalanced prosodic unit boundary detection with linguistically-motivated prosodic featuresabstractContinuous speech input for ASR processing is usually presegmented into speech stretches by pauses. In this paper, we propose that smaller, prosodically defined units can be identified by tackling the problem on imbalanced prosodic unit boundary detection using five machine learning techniques. A parsimonious set of linguistically motivated prosodic features has been proven to be useful to characterize prosodic boundary information. Furthermore, BMPM is prone to have true positive rate on the minority class, i.e. the defined prosodic units. As a whole, the decision tree classifier, C4.5, reaches a more stable performance than the other algorithms. Index Terms: prosodic unit, machine learning, biased minimax probability machine Yi-Fen Liu, Shu-Chuan Tseng, Jyh-Shing Roger Jang, Alvin Cheng-Hsien Chen |
INTERSPEECH | 3 |
| 2010 | On the Improvement of Singing Voice Separation for Monaural Recordings Using the MIR-1K DatasetabstractMonaural singing voice separation is an extremely challenging problem. While efforts in pitch-based inference methods have led to considerable progress in voiced singing voice separation, little attention has been paid to the incapability of such methods to separate unvoiced singing voice due to its in harmonic structure and weaker energy. In this paper, we proposed a systematic approach to identify and separate the unvoiced singing voice from the music accompaniment. We have also enhanced the performance of separating voiced singing via a spectral subtraction method. The proposed system follows the framework of computational auditory scene analysis (CASA) which consists of the segmentation stage and the grouping stage. In the segmentation stage, the input song signals are decomposed into small sensory elements in different time-frequency resolutions. The unvoiced sensory elements are then identified by Gaussian mixture models. The experimental results demonstrated that the quality of the separated singing voice is improved for both the unvoiced and voiced parts. Moreover, to deal with the problem of lack of a publicly available dataset for singing voice separation, we have constructed a corpus called MIR-1K (multimedia information retrieval lab, 1000 song clips) where all singing voices and music accompaniments were recorded separately. Each song clip comes with human-labeled pitch values, unvoiced sounds and vocal/non-vocal segments, and lyrics, as well as the speech recording of the lyrics. Chao-Ling Hsu, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Evaluation of Tangible Learning Companion/Robot for English Language LearningabstractThe purpose of this multiyear study is to combine the voice recognition technology with a tangible learning companion/robot which our research team developed to enhance English learning conversation for beginners. In this study, we collected the userspsila needs analysis of using tangible learning companion first then did a pilot test to analyze its usability. The results show that the learners improved their pronunciation to achieve right voice recognition through practicing with the learning companion and the tangible learning companion can promote the less proficient learnerspsilaconfidence. Moreover, we also noticed that learning with the tangible learning companion could increase the pleasure of learning speaking English and tactile interactions between learners. Shelley S. C. Young, Jyh-Shing Roger Jang |
ICALT | 3 |
| 2009 | On the Use of Anti-Word Models for Audio Music Annotation and RetrievalabstractQuery-by-semantic-description (QBSD) is a natural way for searching/annotating music in a large database. To improve QBSD, we propose the use of anti-words for each annotation word based on the concept of supervised multiclass labeling (SML). More specifically, words that are highly associated with the opposite semantic meaning of a word constitute its anti-word set. By modeling both a word and its anti-word set, our annotation system can achieve 31.1% of equal mean per-word precision and recall, while the original SML model achieves 27.8%. Moreover, by constructing the models of the anti-word explicitly, the performance is also significantly improved for the retrieval system, especially when the query keyword is the antonym of an existing annotation word. Zhi-Sheng Chen, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Minimum phone error discriminative training for Mandarin Chinese speaker adaptationabstractSpeaker adaptation is an efficient way to model a new speaker from an existing speaker-independent model with limited speaker-dependent data. In this paper, we investigate the use of discriminative training schemes based on the minimum phone error (MPE) criterion to improve a well-known speaker adaptation technique, a combination of transform-based adaptation and Bayesian adaptation. Furthermore, a new approach utilizing the statistics of the model-based regression tree for controlling the interpolation between maximum likelihood (ML) and MPE objective functions is also presented. Several comparative experiments were conducted on a continuous speech recognition task for Mandarin Chinese. Experimental results show that the proposed approach can further improve the performance of the original hybrid adaptation. Liang-Yu Chen 0006, Chun-Jen Lee, Jyh-Shing Roger Jang |
INTERSPEECH | 3 |
| 2008 | TRUES: Tone Recognition Using Extended SegmentsabstractTone recognition has been a basic but important task for speech recognition and assessment of tonal languages, such as Mandarin Chinese. Most previously proposed approaches adopt a two-step approach where syllables within an utterance are identified via forced alignment first, and tone recognition using a variety of classifiers---such as neural networks, Gaussian mixture models (GMM), hidden Markov models (HMM), support vector machines (SVM)---is then performed on each segmented syllable to predict its tone. However, forced alignment does not always generate accurate syllable boundaries, leading to unstable voiced-unvoiced detection and deteriorating performance in tone recognition. Aiming to alleviate this problem, we propose a robust approach called Tone Recognition Using Extended Segments (TRUES) for HMM-based continuous tone recognition. The proposed approach extracts an unbroken pitch contour from a given utterance based on dynamic programming over time-domain acoustic features of average magnitude difference function (AMDF). The pitch contour of each syllable is then extended for tri-tone HMM modeling, such that the influence from inaccurate syllable boundaries is lessened. Our experimental results demonstrate that the proposed TRUES achieves 49.13% relative error rate reduction over that of the recently proposed supratone modeling, which is deemed the state of the art of tone recognition that outperforms several previously proposed approaches. The encouraging improvement demonstrates the effectiveness and robustness of the proposed TRUES, as well as the corresponding pitch determination algorithm which produces unbroken pitch contours. Jiang-Chun Chen, Jyh-Shing Roger Jang |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | A General Framework of Progressive Filtering and Its Application to Query by Singing/HummingabstractThis paper presents the mathematical formulation and design methodology of progressive filtering (PF) for multimedia information retrieval, and discusses its application to the so-called query by singing/humming (QBSH), or more formally, melody recognition. The concept of PF and the corresponding dynamic programming-based design method are applicable to large multimedia retrieval systems for striking a balance between efficiency (in terms of response time) and effectiveness (in terms of recognition rate). The application of the proposed PF to a five-stage QBSH system is reported, and the experimental results demonstrate the feasibility of the proposed approach. Jyh-Shing Roger Jang, Hong-Ru Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | An effective initial/final duration prediction method for corpus-based singing voice synthesis of Mandarin ChineseabstractIn this paper, we propose an effective method for predicting initial/final duration for corpus-based singing voice synthesis of Mandarin Chinese. The goal of the method is to improve the naturalness and clarity of the synthesized singing voices. To achieve this goal, we construct an individual initial/final (I/F) duration prediction model for each category of consonants. Support vector machine is used for duration prediction in each model. In order to achieve better accuracy, we use both linguistic/phonetic attributes and music-score information as the input features for the I/F duration prediction model. Experimental results demonstrate that the proposed method is effective in predicting the I/F duration for singing voice synthesis. Cheng-Yuan Lin, Pei-Chi Jao, Jyh-Shing Roger Jang |
INTERSPEECH | 3 |
| 2007 | Automatic Phonetic Segmentation by Score Predictive Model for the Corpora of Mandarin Singing VoicesabstractThis paper proposes the concept of a score predictive model (SPM) that can refine the phoneme boundaries obtained by a hidden Markov model (HMM) and dynamic time warping (DTW) for a Mandarin singing voice corpus. An SPM is constructed by using support vector regression. It predicts the score of a phoneme boundary according to the boundary's 58-dimensional feature vector. The correctly identified boundaries of a singing corpus can then be used for corpus-based singing voice synthesis. Several experiments with different settings, including the use of different initial estimates, different acoustic features, and various regression approaches, were designed to verify the feasibility of the proposed approach. Experimental results demonstrate that the proposed SPM is able to effectively refine the results of the HMM and DTW. Cheng-Yuan Lin, Jyh-Shing Roger Jang |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Formant-based English vowel assessment for Chinese in TaiwanabstractThis paper proposes a formant-based approach for computer-assisted English vowel assessment. Various studies in formant-based speech synthesis have suggested the importance of formant coefficients; this motivates us to investigate pronunciation assessment using formant information instead of MFCC (Mel-frequency cepstral coefficients) alone. In particular, we explore the multistream HMM with the addition of formant information to improve the phoneme segmentation. We then propose the use of PCN (pronunciation confusion network) together with a formant-based confidence measure to improve error detection rates. Furthermore, the pros and cons of using cross-word phone model for both native speakers and L2 learners are discussed. Experimental results demonstrate the feasibility of the proposed approach for automatic vowel pronunciation assessment. Index Terms: computer assisted pronunciation training, formant, assessment, pronunciation confusion network, speech recognition Jiang-Chun Chen, Wei-Tang Hsu, Jyh-Shing Roger Jang, Ren-Yuan Lyu, Yuang-Chin Chiang |
INTERSPEECH | 3 |
| 2006 | Automatic phonetic segmentation by using a SPM-based approach for a Mandarin singing voice corpusabstractThis paper proposes a score predictive model (SPM) based approach to integrate two segmentation results obtained by HMM and DTW for a Mandarin singing voice corpus. The SPM can predict the score of a boundary according to its corresponding 14 dimensional feature vector. In order to verify the performance of the proposed method, several experiments were performed. The experimental results demonstrate the feasibility of the proposed approach. Index Terms: automatic phonetic segmentation, boundary refinement, score predictive model Cheng-Yuan Lin, Jyh-Shing Roger Jang |
INTERSPEECH | 2 |
| 2006 | Admission control schemes for proportional differentiated services enabled internet servers using machine learning techniques
Chenn-Jung Huang, Chih-Lun Cheng, Yi-Ta Chuang, Jyh-Shing Roger Jang |
Expert Syst. Appl. | 4 |
| 2006 | Extraction of transliteration pairs from parallel corpora using a statistical transliteration model
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
Inf. Sci. | 3 |
| 2006 | Alignment of bilingual named entities in parallel corpora using statistical models and multiple knowledge sourcesabstractNamed entity (NE) extraction is one of the fundamental tasks in natural language processing (NLP). Although many studies have focused on identifying NEs within monolingual documents, aligning NEs in bilingual documents has not been investigated extensively due to the complexity of the task. In this article we introduce a new approach to aligning bilingual NEs in parallel corpora by incorporating statistical models with multiple knowledge sources. In our approach, we model the process of translating an English NE phrase into a Chinese equivalent using lexical translation/transliteration probabilities for word translation and alignment probabilities for word reordering. The method involves automatically learning phrase alignment and acquiring word translations from a bilingual phrase dictionary and parallel corpora, and automatically discovering transliteration transformations from a training set of name-transliteration pairs. The method also involves language-specific knowledge functions, including handling abbreviations, recognizing Chinese personal names, and expanding acronyms. At runtime, the proposed models are applied to each source NE in a pair of bilingual sentences to generate and evaluate the target NE candidates; the source and target NEs are then aligned based on the computed probabilities. Experimental results demonstrate that the proposed approach, which integrates statistical models with extra knowledge sources, is highly feasible and offers significant improvement in performance compared to our previous work, as well as the traditional approach of IBM Model 4. Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2005 | A hybrid approach to automatic segmentation and labeling for Mandarin Chinese speech corpusabstractIn this paper, we propose a hybrid approach to refine the phonetic boundaries in a Mandarin speech corpus. This approach employs different sets of acoustic features for different categories of phonetic transitions, except for the most difficult case of “periodic voiced + periodic voiced”, which is therefore handled by a heuristic scheme. Several experiments are designed to demonstrate the feasibility of the proposed approach. Cheng-Yuan Lin, Kuan-Ting Chen, Jyh-Shing Roger Jang |
INTERSPEECH | 3 |
| 2005 | A corpus-based singing voice synthesis system for mandarin ChineseabstractIn this paper, the design and implementation of a corpus-based singing voice synthesis (SVS) system for Mandarin Chinese was introduced. The design rules of three corpora for singing voice synthesis were proposed. After that, two distance functions were defined and the Viterbi search algorithm was applied to identify the optimal combinations of synthesis units from the three corpora. For better performance, several sound effects with synthesized outputs were combined. Finally, we conduct a listening experiment to demonstrate the feasibility of this system. Cheng-Yuan Lin, Tzu-Ying Lin, Jyh-Shing Roger Jang |
ACM Multimedia | 3 |
| 2004 | Automatic pronunciation assessment for Mandarin ChineseabstractThis work describes the algorithms used in a prototypical software system for automatic pronunciation assessment of Mandarin Chinese. The system uses Viterbi decoding to isolate each syllable and find the log probability of a given utterance based on HMM (hidden Markov models). The isolated syllables are then sent to a GMM (Gaussian mixture model) for tone recognition. Based on the log probability and the result from tone recognition, a parametric scoring function, using a neural network, is constructed to approximate the scoring results from human experts. The experimental results demonstrate the system can consistently gives scores that are close to those from human's subjective evaluation. Jiang-Chun Chen, Jyh-Shing Roger Jang, Ming-Chun Wu |
ICME | 2 |
| 2004 | i-Ring: a system for humming transcription and chord generationabstractThis paper describes the construction of a system called i-Ring that can generate a polyphonic ringtone based on a user's humming input. Algorithms used in the system for music transcription and chord generation include various techniques in speech and music signal processing, such as pitch tracking and dynamic programming, which are explained in the paper. Experimental results demonstrate the feasibility of i-Ring as a handy tool for computer assisted music composition for polyphonic ringtones. Hong-Ru Lee, Jyh-Shing Roger Jang |
ICME | 2 |
| 2004 | A two-phase pitch marking method for TD-PSOLA synthesisabstractAbstract. This paper describes a robust two-phase pitch marking method based on peak-valley decision and dynamic programming. In the first phase, we select either peaks or valleys for pitch mark candidates according to its similarity to an estimated pitch curve. In the second phase, we define state and transition probabilities, and then employ dynamic programming to find the most likely pitch marks. We have also designed different tests to demonstrate the feasibility of the proposed approach. 1 Cheng-Yuan Lin, Jyh-Shing Roger Jang |
INTERSPEECH | 2 |
| 2004 | Research and developments of a multi-modal MIR engine for commercial applications in East AsiaabstractAbstract This article describes the research and development of an efficient Music Information Retrieval (MIR) engine that is embedded in a karaoke software package targeted for Asian people's need of music retrieval. The MIR engine has a multi‐modal interface that allows queries by singing, humming, tapping, speaking, and writing. In particular, we discuss the design philosophy, technical barriers, and performance evaluation of such an engine, as well as its current and potential commercial applications. Feedbacks and feature requests from users, which greatly influence our future work, are also addressed. Jyh-Shing Roger Jang, Hong-Ru Lee, Jiang-Chuen Chen, Cheng-Yuan Lin |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | New refinement schemes for voice conversionabstractNew refinement schemes for voice conversion are proposed in this paper. We take mel-frequency cepstral coefficients (MFCC) as the basic feature and adopt cepstral mean subtraction to compensate the channel effects. We propose S/U/V (silence/unvoiced/voiced) decision rule such that two sets of codebooks are used to capture the difference between unvoiced and voiced segments of the source speaker. Moreover, we apply three schemes to refine the synthesized voice, including pitch refinement with PSOLA, energy equalization, and frame concatenation based on synchronized pitch marks. The satisfactory performance of the voice conversion system can be demonstrated through ABX listening test and MOS grade. Cheng-Yuan Lin, Jyh-Shing Roger Jang |
ICME | 2 |
| 2003 | Microcontroller implementation of melody recognition: a prototypeabstractThis demo presents a 16-bit microcontroller implementation of a content-based music retrieval system that can take a user's acoustic input (5-second clip of singing or humming) and then retrieve the intended song from 20 candidate songs. Performance evaluation based on 192 clips shows that the system has a satisfactory top-1 recognition rate of 92%. This system demonstrates the feasibility of microcontroller based melody recognition for music retrieval, which can be used in consumer electronics such as melody-activated interactive toys, query engines for MP3 players or karaoke machines, and so on. Jyh-Shing Roger Jang, Yung-Sen Jang |
ACM Multimedia | 1 |
| 2003 | An automatic singing voice rectifier designabstractThis paper proposes a new approach to automatic singing voice rectification. There are two components in the rectifier; one is the recognizer based on dynamic time warping and the other is the synthesizer based PSOLA (Pitch Synchronous Overlap and Add) for pitch shifting. The purpose of the recognizer is to identify the locations of off-key parts of the user's acoustic input. Then with the target music score, the synthesizer tries to correct the off-key parts by appropriate pitch shifting to match the give music score. We also attempt some singing and listening experiments for evaluating the feasibility of the rectifier and the results exhibit the satisfactory performance. Cheng-Yuan Lin, Jyh-Shing Roger Jang, Mao-Yuan Hsu |
ACM Multimedia | 2 |
| 2003 | A Statistical Approach to Chinese-to-English Back-Transliteration
Chun-Jen Lee, Jason S. Chang, Jyh-Shing Roger Jang |
PACLIC | 3 |
| 2001 | Content-based Music Retrieval Using Linear Scaling and Branch-and-bound Tree Searchabstractpaper presents the use of linear scaling and tree search in a content-based music retrieval system that can take a user's acoustic input (8-second clip of singing or humming) via a microphone and then retrieve the intended song from over 3000 candidate songs in the database. The system, known as Super MBox, demonstrates the feasibility of real-time content-based music retrieval with a high recognition rate. Super MBox first takes the user's acoustic input from a microphone and converts it into a pitch vector. Then a fast comparison engine using linear scaling and tree search is employed to compute the similarity scores. We have tested Super MBox and found the top-20 recognition rate is about 73% with about 1000 clips of test inputs from people with mediocre singing skills. Jyh-Shing Roger Jang, Hong-Ru Lee, Ming-Yang Kao |
ICME | 1 |
| 2001 | Hierarchical filtering method for content-based music retrieval via acoustic inputabstractThis paper presents an implementation of a content-based music retrieval system that can take a user's acoustic input (8-second clip of singing or humming) via a microphone and then retrieve the intended song from a database containing over 3000 candidate songs. The system, known as Super MBox, demonstrates the feasibility of real-time music retrieval with a high success rate. Super MBox first takes the user's acoustic input from a microphone and converts it into a pitch vector. Then a hierarchical filtering method (HFM) is used to first filter out 80% unlikely candidates and then compare the query input with the remaining 20% candidates in a detailed manner. The output of Super MBox is a ranked song list according to the computed similarity scores. A brief mathematical analysis of the two-step HFM is given in the paper to explain how to derive the optimum parameters of the comparison engine. The proposed HFM and its analysis framework can be directly applied to other multimedia information retrieval systems. We have tested Super MBox extensively and found the top-20 success rate is over 85%, based on a dataset of about singing/humming 2000 clips from people with mediocre singing skills. Our studies demonstrate the feasibility of using Super MBox as a prototype for music search engines over the Internet and/or query engines in digital music libraries. Jyh-Shing Roger Jang, Hong-Ru Lee |
ACM Multimedia | 1 |
| 2001 | Super MBox: an efficient/effective content-based music retrieval systemabstractThis demo presents an implementation of a content-based music retrieval system that can take a user's acoustic input (8-second clip of singing or humming) via a microphone and then retrieve the intended song from a database containing 13,000 candidate songs. The system, known as Super MBox, demonstrates the feasibility of real-time music retrieval with a high recognition rate, which can be used for music search engines over the Internet and/or query engines in digital music libraries or karaoke machines. Jyh-Shing Roger Jang, Hong-Ru Lee, Jiang-Chun Chen |
ACM Multimedia | 1 |
| 2000 | Evolving color recipesabstractThis paper highlights an evolutionary computing intelligence for a computerized color recipe prediction that requires function approximation and combinatorial solution of colorants to produce color recipes for a given target color sample. We attack this real challenging problem in the color (paint) industry by using an evolutionary computing system that consists of a problem-specific knowledge and three principal constituents of soft-computing: neural networks, a fuzzy system, and a genetic algorithm. Departing from the recipe results obtained by neural networks (NN) approaches, the evolutionary system attempts to improve them in conjunction with fuzzy classification, a knowledge base and neural fitness functions. All components function synergistically in obtaining precise color recipe outputs through simulation of color paint manufacturing process. Such computational intelligence can be useful, especially when an exact mathematical model of the real-world process under consideration is not available explicitly. Eiji Mizutani, Hideyuki Takagi, David M. Auslander, Jyh-Shing Roger Jang |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 1998 | Author's Reply
Jyh-Shing Roger Jang |
IEEE Trans. Neural Networks | 1 |
| 1995 | Neuro-fuzzy modeling and controlabstractFundamental and advanced developments in neuro-fuzzy synergisms for modeling and control are reviewed. The essential part of neuro-fuzzy synergisms comes from a common framework called adaptive networks, which unifies both neural networks and fuzzy models. The fuzzy models under the framework of adaptive networks is called adaptive-network-based fuzzy inference system (ANFIS), which possess certain advantages over neural networks. We introduce the design methods for ANFIS in both modeling and control applications. Current problems and future directions for neuro-fuzzy approaches are also addressed.> Jyh-Shing Roger Jang, Chuen-Tsai Sun |
Proc. IEEE | 1 |
| 1993 | Functional equivalence between radial basis function networks and fuzzy inference systemsabstractIt is shown that, under some minor restrictions, the functional behavior of radial basis function networks (RBFNs) and that of fuzzy inference systems are actually equivalent. This functional equivalence makes it possible to apply what has been discovered (learning rule, representational power, etc.) for one of the models to the other, and vice versa. It is of interest to observe that two models stemming from different origins turn out to be functionally equivalent. Jyh-Shing Roger Jang, Chuen-Tsai Sun |
IEEE Trans. Neural Networks | 1 |
| 1993 | ANFIS: adaptive-network-based fuzzy inference systemabstractThe architecture and learning procedure underlying ANFIS (adaptive-network-based fuzzy inference system) is presented, which is a fuzzy inference system implemented in the framework of adaptive networks. By using a hybrid learning procedure, the proposed ANFIS can construct an input-output mapping based on both human knowledge (in the form of fuzzy if-then rules) and stipulated input-output data pairs. In the simulation, the ANFIS architecture is employed to model nonlinear functions, identify nonlinear components on-line in a control system, and predict a chaotic time series, all yielding remarkable results. Comparisons with artificial neural networks and earlier work on fuzzy modeling are listed and discussed. Other extensions of the proposed ANFIS and promising applications to automatic control and signal processing are also suggested.> Jyh-Shing Roger Jang |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1992 | Self-learning fuzzy controllers based on temporal backpropagationabstractA generalized control strategy that enhances fuzzy controllers with self-learning capability for achieving prescribed control objectives in a near-optimal manner is presented. This methodology, termed temporal backpropagation, is model-sensitive in the sense that it can deal with plants that can be represented in a piecewise-differentiable format, such as difference equations, neural networks, GMDH structures, and fuzzy models. Regardless of the numbers of inputs and outputs of the plants under consideration, the proposed approach can either refine the fuzzy if-then rules of human experts or automatically derive the fuzzy if-then rules if human experts are not available. The inverted pendulum system is employed as a testbed to demonstrate the effectiveness of the proposed control scheme and the robustness of the acquired fuzzy controller. Jyh-Shing Roger Jang |
IEEE Trans. Neural Networks | 1 |
| 1991 | Fuzzy Modeling Using Generalized Neural Networks and Kalman Filter Algorithm
Jyh-Shing Roger Jang |
AAAI | 1 |
| 1990 | A hierarchical approach to designing approximate reasoning-based controllers for dynamic physical systems
Hamid R. Berenji, Yung-Yaw Chen, Chuen-Chien Lee, Jyh-Shing Roger Jang, S. Murugesan |
UAI | 4 |