EDBT 2026 Demo / reviewers in the wild / expert
Hsin-Min Wang
dblp:28/5019
· DBLP profile ↗
222ranked-venue papers
7as first author
52since 2021 · last 2026
0000-0003-3599-5071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 175 · 5 first-author · 43 since 2021Artificial intelligence and machine learning · 128 · 4 first-author · 34 since 2021Databases, data management, data science and information retrieval · 5Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing
An-Ci Peng, Kuan-Tang Huang, Tien-Hong Lo, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 5 |
| 2026 | TG-ASR: Translation-Guided Learning with Parallel Gated Cross Attention for Low-Resource Automatic Speech Recognition
Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
LREC | 5 |
| 2026 | DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality AlignmentabstractEvaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS). Most automatic MOS estimators are trained with point-wise regression or distributional classification. These objectives do not directly optimize rank-based metrics and provide weak geometric constraints for cross-modal coherence. To address these gaps, we propose DeRA-MOS, a decoupled optimization framework for TTM evaluation. For MI, we introduce a batch-aware listwise ranking loss that models relative order within each mini-batch and better aligns with evaluation based on Spearman's rank correlation coefficient (SRCC). For TA, we introduce a score-anchored modality alignment loss that maps human scores to target audio-text similarity and regularizes the latent space before fusion. By effectively mitigating the point-wise training mismatch and modality drift, experiments on MusicEval demonstrate that our decoupled framework yields substantial improvements in both MI and TA ranking metrics, establishing a robust paradigm for large-scale TTM evaluation. Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
IEEE Signal Process. Lett. | 3 |
| 2026 | Toward Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model RepresentationsabstractPerceptual voice quality assessment plays a vital role in diagnosing and monitoring voice disorders.Traditional methods, such as the Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) and the Grade, Roughness, Breathiness, Asthenia, and Strain (GRBAS) scales, rely on expert raters and are prone to inter-rater variability, emphasizing the need for objective solutions. This study introduces the Voice Quality Assessment Network (VOQANet), a deep learning framework that employs an attention mechanism and Speech Foundation Model (SFM) embeddings to extract high-level features. To further enhance performance, we propose VOQANet+, which integrates self-supervised SFM embeddings with low-level acoustic descriptors-namely jitter, shimmer, and harmonics-to-noise ratio (HNR). Unlike previous approaches that focus solely on vowel-based phonation (PVQD-A), our models are evaluated on both vowel-level and sentence-level speech (PVQD-S) to assess generalizability. Experimental results demonstrate that sentence-based inputs yield higher accuracy, particularly at the patient level. Overall, VOQANet consistently outperforms baseline models in terms of root mean squared error (RMSE) and Pearson correlation coefficient across CAPE-V and GRBAS dimensions, with VOQANet+ achieving even greater performance gains. Additionally, VOQANet+ maintains consistent performance under noisy conditions, suggesting enhanced robustness for real-world and telehealth applications. This work highlights the value of combining SFM embeddings with low-level features for accurate and robust pathological voice assessment. Whenty Ariyanti, Kuan-Yu Chen 0002, Sabato Marco Siniscalchi, Hsin-Min Wang, Yu Tsao 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Revealing the Role of Audio Channels in ASR Performance DegradationabstractPre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences. Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ASRU | 5 |
| 2025 | HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality AssessmentabstractModern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When tested on audio with a higher sampling rate, these models can produce biased scores. We present HighRateMOS, the first non-intrusive mean opinion score (MOS) model that explicitly considers sampling rate. HighRateMOS ensembles three model variants that exploit the following information: (i) a learnable embedding of speech sampling rate, (ii) Wav2vec 2.0 selfsupervised embeddings, (iii) multi-scale CNN spectral features, and (iv) MFCC features. In AudioMOS 2025 Track 3, HighRateMOS ranked first in five of eight metrics. Our experiments confirm that modeling sampling rate leads to more robust and sampling-rate-agnostic speech quality predictions. Wenze Ren, Yi-Cheng Lin, Wen-Chin Huang, Ryandhimas E. Zezario, Szu-Wei Fu, Sung-Feng Huang, Erica Cooper, Hung-Yu Wei 0001, Hsin-Min Wang, Hung-yi Lee, Yu Tsao 0001 |
ASRU | 10 |
| 2025 | QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation SystemsabstractEvaluating audio generation systems, including text-to-music (TTM), text-to-speech (TTS), and text-to-audio (TTA), remains challenging due to the subjective and multi-dimensional nature of human perception. Existing methods treat mean opinion score (MOS) prediction as a regression problem, but standard regression losses overlook the relativity of perceptual judgments. To address this limitation, we introduce QAMRO, a novel Quality-aware Adaptive Margin Ranking Optimization framework that seamlessly integrates regression objectives from different perspectives, aiming to highlight perceptual differences and prioritize accurate ratings. Our framework leverages pre-trained audio-text models such as CLAP and Audiobox-Aesthetics, and is trained exclusively on the official AudioMOS Challenge 2025 dataset. It demonstrates superior alignment with human evaluations across all dimensions, significantly outperforming robust baseline models. Chien-Chun Wang, Kuan-Tang Huang, Cheng-Yeh Yang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ASRU | 5 |
| 2025 | Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised EmbeddingsabstractWe present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores—Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness—for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain robust audio quality assessment without synthetic training data. Dyah A. M. G. Wisnu, Ryandhimas E. Zezario, Stefano Rini, Hsin-Min Wang, Yu Tsao 0001 |
ASRU | 4 |
| 2025 | MSECG: Incorporating Mamba for Robust and Efficient ECG Super-ResolutionabstractElectrocardiogram (ECG) signals play a crucial role in diagnosing cardiovascular diseases. To reduce power consumption in wearable or portable devices used for long-term ECG monitoring, super-resolution (SR) techniques have been developed, enabling these devices to collect and transmit signals at a lower sampling rate. In this study, we propose MSECG, a compact neural network model designed for ECG SR. MSECG combines the strength of the recurrent Mamba model with convolutional layers to capture both local and global dependencies in ECG waveforms, allowing for the effective reconstruction of high-resolution signals. We also assess the model’s performance in real-world noisy conditions by utilizing ECG data from the PTB-XL database and noise data from the MIT-BIH Noise Stress Test Database. Experimental results show that MSECG outperforms two contemporary ECG SR models under both clean and noisy conditions while using fewer parameters, offering a more powerful and robust solution for long-term ECG monitoring applications. I Chiu, Kuan-Chen Wang, Kai-Chun Liu, Hsin-Min Wang, Ping-Cheng Yeh, Yu Tsao 0001 |
ICASSP | 5 |
| 2025 | Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech EnhancementabstractIn multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving very promis- ing performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information. Wenze Ren, Yi-Cheng Lin, Xuanjun Chen, Rong Chao, Kuo-Hsuan Hung, You-Jin Li, Wen-Yuan Ting, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 9 |
| 2025 | Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech RecognitionabstractWhile pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics. Chien-Chun Wang, Cheng-Kang Chou, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
ICASSP | 6 |
| 2025 | A Study on Zero-shot Non-intrusive Speech Assessment using Large Language ModelsabstractThis work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the text’s naturalness via targeted prompt engineering. We evaluate the assessment metrics predicted by GPT-4o and GPT-Whisper, examining their correlation with human-based quality and intelligibility assessments and the character error rate (CER) of automatic speech recognition. Experimental results show that GPT-4o alone is less effective for audio analysis, while GPT-Whisper achieves higher prediction accuracy, has moderate correlation with speech quality and intelligibility, and has higher correlation with CER. Compared to SpeechLMScore and DNSMOS, GPT-Whisper excels in intelligibility metrics, but performs slightly worse than SpeechLMScore in quality estimation. Furthermore, GPT-Whisper outperforms supervised non-intrusive models MOS-SSL and MTI-Net in Spearman’s rank correlation for Whisper’s CER. These findings validate GPT-Whisper’s potential for zero-shot speech assessment without requiring additional training data. Ryandhimas E. Zezario, Sabato Marco Siniscalchi, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 3 |
| 2025 | A Study on Speech Assessment with Visual Cues
Shafique Ahmed, Ryandhimas E. Zezario, Nasir Saleem, Amir Hussain 0001, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 5 |
| 2025 | A Comparative Study on Proactive and Passive Detection of Deepfake Speech
Chia-Hua Wu, Wanying Ge, Xin Wang 0037, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 6 |
| 2025 | Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing AidsabstractGiven the critical role of non-intrusive speech intelligibility assessment in hearing aids (HA), this paper enhances its performance by introducing Feature Importance across Domains (FiDo). We estimate feature importance on spectral and time-domain acoustic features as well as latent representations of Whisper. Importance weights are calculated per frame, and based on these weights, features are projected into new spaces, allowing the model to focus on important areas early. Next, feature concatenation is performed to combine the features before the assessment module processes them. Experimental results show that when FiDo is incorporated into the improved multi-branched speech intelligibility model MBI-Net+, RMSE can be reduced by 7.62% (from 26.10 to 24.11). MBI-Net+ with FiDo also achieves a relative RMSE reduction of 3.98% compared to the best system in the 2023 Clarity Prediction Challenge. These results validate FiDo's effectiveness in enhancing neural speech assessment in HA. Ryandhimas E. Zezario, Sabato Marco Siniscalchi, Fei Chen 0011, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2025 | AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face VideosabstractMultimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely detection is crucial. The damage to either modality (i.e., visual or audio) can only be discovered through multimodal models that can exploit both pieces of information simultaneously. However, previous methods mainly adopt unimodal video forensics and use supervised pretraining for forgery detection. This study proposes a new method based on a multimodal self-supervised-learning (SSL) feature extractor to exploit inconsistency between audio and visual modalities for multimodal video forgery detection. We use the transformer-based SSL pretrained Audio-Visual HuBERT (AV-HuBERT) model as a visual and acoustic feature extractor and a multiscale temporal convolutional neural network to capture the temporal correlation between the audio and visual modalities. Since AV-HuBERT only extracts visual features from the lip region, we also adopt another transformer-based video model to exploit facial features and capture spatial and temporal artifacts caused during the deepfake generation process. Experimental results show that our model outperforms all existing models and achieves new state-of-the-art performance on the FakeAVCeleb and DeepfakeTIMIT datasets. Sahibzada Adil Shahzad, Ammarah Hashmi, Yan-Tsung Peng, Yu Tsao 0001, Hsin-Min Wang |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2024 | Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment ModelabstractThis study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multitask learning. The 3QUEST metrics, namely Speech-MOS (S-MOS), Noise-MOS (N-MOS), and General-MOS (G-MOS), are the assessment targets. The pretrained MOSA-Net model is utilized to estimate three pseudo labels: perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI), and speech distortion index (SDI). Multi-task learning is then employed to train MTQ-Net by combining a supervised loss (derived from the difference between the estimated score and the ground-truth label) and a semi-supervised loss (derived from the difference between the estimated score and the pseudo label), where the Huber loss is employed as the loss function. Experimental results first demonstrate the advantages of MPL compared to training a model from scratch and using a direct knowledge transfer mechanism. Second, the benefit of the Huber loss for improving the predictive ability of MTQ-Net is verified. Finally, the MTQ-Net with the MPL approach exhibits higher overall predictive power compared to other SSL-based speech assessment models. Ryandhimas E. Zezario, Bo-Ren Bai, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 4 |
| 2024 | A Study On Incorporating Whisper For Robust Speech AssessmentabstractThis research introduces an enhanced version of the multi-objective speech assessment model–MOSA-Net+, by leveraging the acoustic features from Whisper, a large-scaled weakly supervised model. We first investigate the effectiveness of Whisper in deploying a more robust speech assessment model. After that, we explore combining representations from Whisper and SSL models. The experimental results reveal that Whisper’s embedding features can contribute to more accurate prediction performance. Moreover, combining the embedding features from Whisper and SSL models only leads to marginal improvement. As compared to intrusive methods, MOSA-Net, and other SSL-based speech assessment models, MOSA-Net+ yields notable improvements in estimating subjective quality and intelligibility scores across all evaluation metrics in Taiwan Mandarin Hearing In Noise test - Quality & Intelligibility (TMHINT-QI) dataset. To further validate its robustness, MOSA-Net+ was tested in the noisy-and-enhanced track of the VoiceMOS Challenge 2023, where it obtained the top-ranked performance among nine systems. Ryandhimas E. Zezario, Yuwen Chen 0006, Szu-Wei Fu, Yu Tsao 0001, Hsin-Min Wang, Chiou-Shann Fuh |
ICME | 5 |
| 2024 | Learnable Layer Selection and Model Fusion for Speech Self-Supervised Learning Models
Sheng-Chieh Chiu, Chia-Hua Wu, Jih-Kang Hsieh, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 5 |
| 2024 | SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
Chun Yin, Tai-Shih Chi, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2024 | Non-Intrusive Speech Intelligibility Prediction for Hearing Aids using Whisper and Metadata
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2024 | The Voicemos Challenge 2024: Beyond Speech Quality PredictionabstractWe present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of “zoomed-in” high-quality samples from speech synthesis systems. The second track was to predict ratings of samples from singing voice synthesis and voice conversion with a large variety of systems, listeners, and languages. The third track was semi-supervised quality prediction for noisy, clean, and enhanced speech, where a very small amount of labeled training data was provided. Among the eight teams from both academia and industry, we found that many were able to outperform the baseline systems. Successful techniques included retrieval-based methods and the use of non-self-supervised representations like spectrograms and pitch histograms. These results showed that the challenge has advanced the field of subjective speech rating prediction. Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E. Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, Yu Tsao 0001 |
SLT | 6 |
| 2024 | Effective Noise-Aware Data Simulation For Domain-Adaptive Speech Enhancement Leveraging Dynamic Stochastic PerturbationabstractCross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation. Chien-Chun Wang, Hung-Shin Lee, Berlin Chen, Hsin-Min Wang |
SLT | 5 |
| 2023 | The Voicemos Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple DomainsabstractWe present the second edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthesized and processed speech. This year, we emphasize real-world and challenging zero-shot out-of-domain MOS prediction with three tracks for three different voice evaluation scenarios. Ten teams from industry and academia in seven different countries participated. Surprisingly, we found that the two sub-tracks of French text-to-speech synthesis had large differences in their predictability, and that singing voice-converted samples were not as difficult to predict as we had expected. Use of diverse datasets and listener information during training appeared to be successful approaches. Erica Cooper, Wen-Chin Huang, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda, Junichi Yamagishi |
ASRU | 4 |
| 2023 | LC4SV: A Denoising Framework Learning to Compensate for Unseen Speaker Verification ModelsabstractThe performance of speaker verification (SV) models may drop dramatically in noisy environments. A speech enhancement (SE) module can be used as a front-end strategy. However, existing SE methods may fail to bring performance improvements to downstream SV systems due to artifacts in the predicted signals of SE models. To compensate for artifacts, we propose a generic denoising framework named LC4SV, which can serve as a pre-processor for various unknown downstream SV models. In LC4SV, we employ a learning-based interpolation agent to automatically generate the appropriate coefficients between the enhanced signal and its noisy input to improve SV performance in noisy environments. Our experimental results demonstrate that LC4SV consistently improves the performance of various unseen SV systems. To the best of our knowledge, this work is the first attempt to develop a learning-based interpolation scheme aiming at improving SV performance in noisy environments. Chi-Chang Lee, Chu-Song Chen, Hsin-Min Wang, Tsung-Te Liu, Yu Tsao 0001 |
ASRU | 4 |
| 2023 | D4AM: A General Denoising Framework for Downstream Acoustic Models
Chi-Chang Lee, Yu Tsao 0001, Hsin-Min Wang, Chu-Song Chen |
ICLR | 3 |
| 2023 | Mandarin Electrolaryngeal Speech Voice Conversion using Cross-domain Features
Hsin-Hao Chen 0006, Yung-Lun Chien, Ming-Chi Yen, Shu-Wei Tsai, Tai-Shih Chi, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 6 |
| 2023 | A Training and Inference Strategy Using Noisy and Enhanced Speech as Target for Speech Enhancement without Clean Speech
Yao-Fei Cheng, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 5 |
| 2023 | Audio-Visual Mandarin Electrolaryngeal Speech Voice Conversion
Yung-Lun Chien, Hsin-Hao Chen 0006, Ming-Chi Yen, Shu-Wei Tsai, Hsin-Min Wang, Yu Tsao 0001, Tai-Shih Chi |
INTERSPEECH | 5 |
| 2023 | Multi-Target Extractor and Detector for Unknown-Number Speaker DiarizationabstractStrong representations of target speakers can help extract important information about speakers and detect corresponding temporal regions in multi-speaker conversations. In this study, we propose a neural architecture that simultaneously extracts speaker representations consistent with the speaker diarization objective and detects the presence of each speaker on a frame-by-frame basis regardless of the number of speakers in a conversation. A speaker representation (called z-vector) extractor and a time-speaker contextualizer, implemented by a residual network and processing data in both temporal and speaker dimensions, are integrated into a unified framework. Tests on the CALLHOME corpus show that our model outperforms most of the methods proposed so far. Evaluations in a more challenging case with simultaneous speakers ranging from 2 to 7 show that our model achieves 6.4% to 30.9% relative diarization error rate reductions over several typical baselines. Chin-Yi Cheng, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
IEEE Signal Process. Lett. | 4 |
| 2023 | Generalization Ability Improvement of Speaker Representation and Anti-Interference for Speaker VerificationabstractThe ability to generalize to mismatches between training and testing conditions and resist interference from other speakers is crucial for the performance of speaker verification. In this paper, we propose two novel approaches to improve the generalization ability to deal with the mismatched recorded scenarios and languages in test conditions and to reduce the influence of interference from other speakers on the similarity measurement of two speaker embeddings. First, parent embedding learning (PEL) is used for model training, which exploits the generalization ability of the shared structure to improve the representation of speaker embeddings. Second, partial adaptive score normalization (PAS-Norm) is used to reduce the influence of interference from other speakers on embedding-based similarity measures. In the experiments, the speaker embedding models are trained using the VoxCeleb2 dataset, and the performance is evaluated on four other datasets under different conditions, including VoxCeleb1, Librispeech, SITW, and CN-Celeb datasets. In the experiments on VoxCeleb1, evaluation results considering a large number of verification speakers and identity restrictions show that the proposed PEL-based system reduces the EER by 6.0% and 4.9% in these two cases, respectively, compared to the state-of-the-art (SOTA) system. Furthermore, in the experiments evaluating speaker verification in mismatch conditions on SITW and CN-Celeb, the proposed PEL-based system also outperforms the SOTA system. In the language mismatched conditions, the EER is reduced by 8.3%. For the evaluation of the influence of interference from other speakers, the EER is significantly reduced by 24.4% when PAS-Norm is used instead of the baseline AS-Norm score normalization method. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Decomposition and Reorganization of Phonetic Information for Speaker Embedding LearningabstractSpeech content is closely related to the stability of speaker embeddings in speaker verification tasks. In this paper, we propose a novel architecture based on self-constraint learning (SCL) and reconstruction task (RT) to remove the influence of phonetic information on speaker embedding generation. First, SCL is used to reduce the divergence of frame-level features, which can avoid ambiguity between the resulting embeddings of the two utterances being compared. Second, RT is used to further remove phonetic information in frame-level layers, focusing on speaker-discriminative feature transformation. In our experiments, the speaker embedding models were trained on the VoxCeleb2 dataset and evaluated on the VoxCeleb1, Librispeech, SITW and VoxMovies datasets. Experimental results on VoxCeleb1 show that the proposed DROP-TDNN system reduced the EER by 7.5%, compared to the state-of-the-art ECAPA-TDNN system. Furthermore, the proposed DROP-TDNN system also outperformed the ECAPA-TDNN system in the experiments on SITW, Librispeech and VoxMovies under cross-dataset conditions. In the experiments on SITW, the proposed system reduced the EER by 3.4% compared to the ECAPA-TDNN system. In the experiments on Librispeech, the proposed system demonstrated the advantage of removing phonetic information under the clean speech condition, with a significant reduction of 25.5% in EER compared to the ECAPA-TDNN system. In the experiments on VoxMovies, the proposed system reduced the EER by up to 7.9% compared to the ECAPA-TDNN system under different pronunciation and background conditions. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Deep Learning-Based Non-Intrusive Multi-Objective Speech Assessment Model With Cross-Domain FeaturesabstractThis study proposes a cross-domain multi-objective speech assessment model, called MOSA-Net, which can simultaneously estimate the speech quality, intelligibility, and distortion assessment scores of an input speech signal. MOSA-Net comprises a convolutional neural network and bidirectional long short-term memory architecture for representation extraction, and a multiplicative attention layer and a fully connected layer for each assessment metric prediction. Additionally, cross-domain features (spectral and time-domain features) and latent representations from self-supervised learned (SSL) models are used as inputs to combine rich acoustic information to obtain more accurate assessments. Experimental results show that in both seen and unseen noise environments, MOSA-Net can improve the linear correlation coefficient (LCC) scores in perceptual evaluation of speech quality (PESQ) prediction, compared to Quality-Net, an existing single-task model for PESQ prediction, and improve LCC scores in short-time objective intelligibility (STOI) prediction, compared to STOI-Net, an existing single-task model for STOI prediction. Moreover, MOSA-Net can be used as a pre-trained model to be effectively adapted to an assessment model for predicting subjective quality and intelligibility scores with a limited amount of training data. Experimental results show that MOSA-Net can improve LCC scores in mean opinion score (MOS) predictions, compared to MOS-SSL, a strong single-task model for MOS prediction. We further adopt the latent representations of MOSA-Net to guide the speech enhancement (SE) process and derive a quality-intelligibility (QI)-aware SE (QIA-SE) approach. Experimental results show that QIA-SE outperforms the baseline SE system with improved PESQ scores in both seen and unseen noise environments over a baseline SE model. Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | EMGSE: Acoustic/EMG Fusion for Multimodal Speech EnhancementabstractMultimodal learning has been proven to be an effective method to improve speech enhancement (SE) performance, especially in challenging situations such as low signal-to-noise ratios, speech noise, or unseen noise types. In previous studies, several types of auxiliary data have been used to construct multimodal SE systems, such as lip images, electropalatography, or electromagnetic midsagittal articulography. In this paper, we propose a novel EMGSE framework for multimodal SE, which integrates audio and facial electromyography (EMG) signals. Facial EMG is a biological signal containing articulatory movement information, which can be measured in a non-invasive way. Experimental results show that the proposed EMGSE system can achieve better performance than the audio-only SE system. The benefits of fusing EMG signals with acoustic signals for SE are notable under challenging circumstances. Furthermore, this study reveals that cheek EMG is sufficient for SE. Kuan-Chen Wang, Kai-Chun Liu, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 3 |
| 2022 | Partially Fake Audio Detection by Self-Attention-Based Fake Span DiscoveryabstractThe past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoof challenge mainly focuses on synthesized audios by advanced speech synthesis and voice conversion models, and replay attacks. Recently, the first Audio Deep Synthesis Detection challenge (ADD 2022) extends the attack scenarios into more aspects. Also, ADD 2022 is the first challenge to propose the partially fake audio detection task. Such brand new attacks are dangerous and how to tackle such attacks remains an open question. Thus, we propose a novel framework by introducing the question-answering (fake span discovery) strategy with the self-attention mechanism to detect partially fake audios. The proposed fake span detection module tasks the anti-spoofing model to predict the start and end positions of the fake clip within the partially fake audio, address the model’s attention into discovering the fake spans rather than other shortcuts with less generalization, and finally equips the model with the discrimination capacity between real and partially fake audios. Our submission ranked second in the partially fake audio detection track of ADD 2022. Heng-Cheng Kuo, Naijun Zheng, Kuo-Hsuan Hung, Hung-yi Lee, Yu Tsao 0001, Hsin-Min Wang, Helen M. Meng |
ICASSP | 7 |
| 2022 | The VoiceMOS Challenge 2022abstractWe present the first edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthetic speech. This challenge drew 22 participating teams from academia and industry who tried a variety of approaches to tackle the problem of predicting human ratings of synthesized speech. The listening test data for the main track of the challenge consisted of samples from 187 different text-to-speech and voice conversion systems spanning over a decade of research, and the out-of-domain track consisted of data from more recent systems rated in a separate listening test. Results of the challenge show the effectiveness of fine-tuning self-supervised speech models for the MOS prediction task, as well as the difficulty of predicting MOS ratings for unseen speakers and listeners, and for unseen systems in the out-of-domain setting. Wen-Chin Huang, Erica Cooper, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2022 | Chain-based Discriminative Autoencoders for Speech RecognitionabstractIn our previous work, we proposed a discriminative autoencoder (DcAE) for speech recognition.DcAE combines two training schemes into one.First, since DcAE aims to learn encoderdecoder mappings, the squared error between the reconstructed speech and the input speech is minimized.Second, in the code layer, frame-based phonetic embeddings are obtained by minimizing the categorical cross-entropy between ground truth labels and predicted triphone-state scores.DcAE is developed based on the Kaldi toolkit by treating various TDNN models as encoders.In this paper, we further propose three new versions of DcAE.First, a new objective function that considers both categorical cross-entropy and mutual information between ground truth and predicted triphone-state sequences is used.The resulting DcAE is called a chain-based DcAE (c-DcAE).For application to robust speech recognition, we further extend c-DcAE to hierarchical and parallel structures, resulting in hc-DcAE and pc-DcAE.In these two models, both the error between the reconstructed noisy speech and the input noisy speech and the error between the enhanced speech and the reference clean speech are taken into the objective function.Experimental results on the WSJ and Aurora-4 corpora show that our DcAE models outperform baseline systems. Hung-Shin Lee, Pin-Tuan Huang, Yao-Fei Cheng, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2022 | NASTAR: Noise Adaptive Speech Enhancement with Target-Conditional ResamplingabstractFor deep learning-based speech enhancement (SE) systems, the training-test acoustic mismatch can cause notable performance degradation.To address the mismatch issue, numerous noise adaptation strategies have been derived.In this paper, we propose a novel method, called noise adaptive speech enhancement with target-conditional resampling (NASTAR), which reduces mismatches with only one sample (one-shot) of noisy speech in the target environment.NASTAR uses a feedback mechanism to simulate adaptive training data via a noise extractor and a retrieval model.The noise extractor estimates the target noise from the noisy speech, called pseudo-noise.The noise retrieval model retrieves relevant noise samples from a pool of noise signals according to the noisy speech, called relevant-cohort.The pseudo-noise and the relevant-cohort set are jointly sampled and mixed with the source speech corpus to prepare simulated training data for noise adaptation.Experimental results show that NASTAR can effectively use one noisy speech sample to adapt an SE model to a target condition.Moreover, both the noise extractor and the noise retrieval model contribute to model adaptation.To our best knowledge, NASTAR is the first work to perform one-shot noise adaptation through noise extraction and retrieval. Chi-Chang Lee, Cheng-Hung Hu, Yuchen Lin 0003, Chu-Song Chen, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 5 |
| 2022 | Disentangling the Impacts of Language and Channel Variability on Speech Separation NetworksabstractBecause the performance of speech separation is excellent for speech in which two speakers completely overlap, research attention has been shifted to dealing with more realistic scenarios.However, domain mismatch between training/test situations due to factors, such as speaker, content, channel, and environment, remains a severe problem for speech separation.Speaker and environment mismatches have been studied in the existing literature.Nevertheless, there are few studies on speech content and channel mismatches.Moreover, the impacts of language and channel in these studies are mostly tangled.In this study, we create several datasets for various experiments.The results show that the impacts of different languages are small enough to be ignored compared to the impacts of different channels.In our experiments, training on data recorded by Android phones leads to the best generalizability.Moreover, we provide a new solution for channel mismatch by evaluating projection, where the channel similarity can be measured and used to effectively select additional training data to improve the performance of in-the-wild test data. Fan-Lin Wang, Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2022 | MBI-Net: A Non-Intrusive Multi-Branched Speech Intelligibility Prediction Model for Hearing Aids
Ryandhimas E. Zezario, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2022 | MTI-Net: A Multi-Target Speech Intelligibility Prediction ModelabstractRecently, deep learning (DL)-based non-intrusive speech assessment models have attracted great attention.Many studies report that these DL-based models yield satisfactory assessment performance and good flexibility, but their performance in unseen environments remains a challenge.Furthermore, compared to quality scores, fewer studies elaborate deep learning models to estimate intelligibility scores.This study proposes a multi-task speech intelligibility prediction model, called MTI-Net, for simultaneously predicting human and machine intelligibility measures.Specifically, given a speech utterance, MTI-Net is designed to predict human subjective listening test results and word error rate (WER) scores.We also investigate several methods that can improve the prediction performance of MTI-Net.First, we compare different features (including low-level features and embeddings from self-supervised learning (SSL) models) and prediction targets of MTI-Net.Second, we explore the effect of transfer learning and multi-tasking learning on training MTI-Net.Finally, we examine the potential advantages of fine-tuning SSL embeddings.Experimental results demonstrate the effectiveness of using cross-domain features, multi-task learning, and fine-tuning SSL embeddings.Furthermore, it is confirmed that the intelligibility and WER scores predicted by MTI-Net are highly correlated with the ground-truth scores. Ryandhimas E. Zezario, Szu-Wei Fu, Fei Chen 0011, Chiou-Shann Fuh, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 5 |
| 2022 | SVSNet: An End-to-End Speaker Voice Similarity Assessment ModelabstractNeural evaluation metrics derived for numerous speech generation tasks have recently attracted great attention. In this paper, we propose SVSNet, the first end-to-end neural network model to assess the speaker voice similarity between converted speech and natural speech for voice conversion tasks. Unlike most neural evaluation metrics that use hand-crafted features, SVSNet directly takes the raw waveform as input to more completely utilize speech information for prediction. SVSNet consists of encoder, co-attention, distance calculation, and prediction modules and is trained in an end-to-end manner. The experimental results on the Voice Conversion Challenge 2018 and 2020 (VCC2018 and VCC2020) datasets show that SVSNet outperforms well-known baseline systems in the assessment of speaker similarity at the utterance and system levels. Cheng-Hung Hu, Yu-Huai Peng, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang |
IEEE Signal Process. Lett. | 5 |
| 2022 | Improved Lite Audio-Visual Speech EnhancementabstractNumerous studies have investigated the effectiveness of audio-visual multimodal learning for speech enhancement (AVSE) tasks, seeking a solution that uses visual data as auxiliary and complementary input to reduce the noise of noisy speech signals. Recently, we proposed a lite audio-visual speech enhancement (LAVSE) algorithm for a car-driving scenario. Compared to conventional AVSE systems, LAVSE requires less online computation and to some extent solves the user privacy problem on facial data. In this study, we extend LAVSE to improve its ability to address three practical issues often encountered in implementing AVSE systems, namely, the additional cost of processing visual data, audio-visual asynchronization, and low-quality visual data. The proposed system is termed improved LAVSE (iLAVSE), which uses a convolutional recurrent neural network architecture as the core AVSE model. We evaluate iLAVSE on the Taiwan Mandarin speech with video dataset. Experimental results confirm that compared to conventional AVSE systems, iLAVSE can effectively overcome the aforementioned three practical issues and can improve enhancement performance. The results also confirm that iLAVSE is suitable for real-world scenarios, where high-quality audio-visual sensors may not always be available. Shang-Yi Chuang, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | HASA-Net: A Non-Intrusive Hearing-Aid Speech Assessment NetworkabstractWithout the need of a clean reference, non-intrusive speech assessment methods have caught great attention for objective evaluations. Recently, deep neural network (DNN) models have been applied to build non-intrusive speech assessment approaches and confirmed to provide promising performance. However, most DNN-based approaches are designed for normal-hearing listeners without considering hearing-loss factors. In this study, we propose a DNN-based hearing aid speech assessment network (HASA-Net), formed by a bidirectional long short-term memory (BLSTM) model, to predict speech quality and intelligibility scores simultaneously according to input speech signals and specified hearing-loss patterns. To the best of our knowledge, HASA-Net is the first work to incorporate quality and intelligibility assessments utilizing a unified DNN-based non-intrusive model for hearing aids. Experimental results show that the predicted speech quality and intelligibility scores of HASA-Net are highly correlated to two well-known intrusive hearing-aid evaluation metrics, hearing aid speech quality index (HASQI) and hearing aid speech perception index (HASPI), respectively. Hsin-Tien Chiang, Yi-Chiao Wu, Tomoki Toda, Hsin-Min Wang, Yih-Chun Hu, Yu Tsao 0001 |
ASRU | 5 |
| 2021 | Mandarin Electrolaryngeal Speech Voice Conversion with Sequence-to-Sequence ModelingabstractThe electrolaryngeal speech (EL speech) is typically spoken with an electrolarynx device that generates excitation signals to substitute human vocal fold vibrations. Because the excitation signals cannot perfectly characterize sound sources generated by vocal folds, the naturalness and intelligibility of the EL speech are inevitably worse than that of the natural speech (NL speech). To improve speech naturalness, statistical models, such as Gaussian mixture models and deep-learning-based models, have been employed for EL speech voice conversion (ELVC). The ELVC task aims to convert EL speech into NL speech through an ELVC model. To implement a frame-wise ELVC system, accurate feature alignment is crucial for model training. However, the abnormal acoustic characteristics of the EL speech cause misalignments and accordingly limit the ELVC performance. To address this issue, we propose a novel ELVC system based on sequence-to-sequence (seq2seq) modeling with text-to-speech (TTS) pretraining. The seq2seq model involves an attention mechanism to concurrently perform representation learning and alignment. Meanwhile, TTS pretraining provides efficient training with limited data. Experimental results show that the proposed ELVC system yields notable improvements in terms of standardized evaluation metrics and subjective listening tests over a well-known frame-wise ELVC system. Ming-Chi Yen, Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Shu-Wei Tsai, Yu Tsao 0001, Tomoki Toda, Jyh-Shing Roger Jang, Hsin-Min Wang |
ASRU | 9 |
| 2021 | Speech Recognition by Simply Fine-Tuning BertabstractWe propose a simple method for automatic speech recognition (ASR) by fine-tuning BERT, which is a language model (LM) trained on large-scale unlabeled text data and can generate rich contextual representations. Our assumption is that given a history context sequence, a powerful LM can narrow the range of possible choices and the speech signal can be used as a simple clue. Hence, comparing to conventional ASR systems that train a powerful acoustic model (AM) from scratch, we believe that speech recognition is possible by simply fine-tuning a BERT model. As an initial study, we demonstrate the effectiveness of the proposed idea on the AISHELL dataset and show that stacking a very simple AM on top of BERT can yield reasonable performance. Wen-Chin Huang, Chia-Hua Wu, Shang-Bao Luo, Kuan-Yu Chen 0002, Hsin-Min Wang, Tomoki Toda |
ICASSP | 5 |
| 2021 | Melody Harmonization Using Orderless Nade, Chord Balancing, and Blocked Gibbs SamplingabstractCoherence and interestingness are two criteria for evaluating the performance of melody harmonization, which aims to generate a chord progression from a symbolic melody. In this study, we apply the concept of orderless NADE, which takes the melody and its partially masked chord sequence as the input of the BiLSTM-based networks to learn the masked ground truth, to the training process. In addition, the class weights are used to compensate for some reasonable chord labels that are rarely seen in the training set. Consistent with the stochasticity in training, blocked Gibbs sampling with proper numbers of masking/generating loops is used in the inference phase to progressively trade the coherence of the generated chord sequence off against its interestingness. The experiments were conducted on a dataset of 18,005 melody/chord pairs. Our proposed model outperforms the state-of-the-art system MTHarmonizer in five of six different objective metrics based on chord/melody harmonicity and chord progression. The subjective test results with more than 100 participants also show the superiority of our model. Chung-En Sun, Yi-Wei Chen, Hung-Shin Lee, Yen-Hsing Chen, Hsin-Min Wang |
ICASSP | 5 |
| 2021 | AlloST: Low-Resource Speech Translation Without Source TranscriptionabstractThe end-to-end architecture has made promising progress in speech translation (ST). However, the ST task is still challenging under low-resource conditions. Most ST models have shown unsatisfactory results, especially in the absence of word information from the source speech utterance. In this study, we survey methods to improve ST performance without using source transcription, and propose a learning framework that utilizes a language-independent universal phone recognizer. The framework is based on an attention-based sequence-to-sequence model, where the encoder generates the phonetic embeddings and phone-aware acoustic representations, and the decoder controls the fusion of the two embedding streams to produce the target token sequence. In addition to investigating different fusion strategies, we explore the specific usage of byte pair encoding (BPE), which compresses a phone sequence into a syllable-like segmented sequence. Due to the conversion of symbols, a segmented sequence represents not only pronunciation but also language-dependent information lacking in phones. Experiments conducted on the Fisher Spanish-English and Taigi-Mandarin drama corpora show that our method outperforms the conformer-based baseline, and the performance is close to that of the existing best method using source transcription. Yao-Fei Cheng, Hung-Shin Lee, Hsin-Min Wang |
Interspeech | 3 |
| 2021 | A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice ConversionabstractWe propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a dysarthria patient are nearly impossible to collect, previous work failed to recover the individuality of the patient. In light of this, we suggest a novel, two-stage approach for DVC, which is highly flexible in that no normal speech of the patient is required. First, a powerful parallel sequence-to-sequence model converts the input dysarthric speech into a normal speech of a reference speaker as an intermediate product, and a nonparallel, frame-wise VC model realized with a variational autoencoder then converts the speaker identity of the reference speech back to that of the patient while assumed to be capable of preserving the enhanced quality. We investigate several design options. Experimental evaluation results demonstrate the potential of our approach to improving the quality of the dysarthric speech while maintaining the speaker identity. Wen-Chin Huang, Kazuhiro Kobayashi, Yu-Huai Peng, Ching-Feng Liu, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda |
Interspeech | 6 |
| 2021 | Dual-Path Filter Network: Speaker-Aware Modeling for Speech SeparationabstractSpeech separation has been extensively studied to deal with the cocktail party problem in recent years.All related approaches can be divided into two categories: time-frequency domain methods and time domain methods.In addition, some methods try to generate speaker vectors to support source separation.In this study, we propose a new model called dualpath filter network (DPFN).Our model focuses on the postprocessing of speech separation to improve speech separation performance.DPFN is composed of two parts: the speaker module and the separation module.First, the speaker module infers the identities of the speakers.Then, the separation module uses the speakers' information to extract the voices of individual speakers from the mixture.DPFN constructed based on DPRNN-TasNet is not only superior to DPRNN-TasNet, but also avoids the problem of permutation-invariant training (PIT). Fan-Lin Wang, Yu-Huai Peng, Hung-Shin Lee, Hsin-Min Wang |
Interspeech | 4 |
| 2021 | Relational Data Selection for Data Augmentation of Speaker-Dependent Multi-Band MelGAN VocoderabstractNowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available.Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impractical to collect a large amount of data of a specific target speaker for most realworld applications.To tackle the problem of limited target data, a data augmentation method based on speaker representation and similarity measurement of speaker verification is proposed in this paper.The proposed method selects utterances that have similar speaker identity to the target speaker from an external corpus, and then combines the selected utterances with the limited target data for SD vocoder adaptation.The evaluation results show that, compared with the vocoder adapted using only limited target data, the vocoder adapted using augmented data improves both the quality and similarity of synthesized speech. Yi-Chiao Wu, Cheng-Hung Hu, Hung-Shin Lee, Yu-Huai Peng, Wen-Chin Huang, Yu Tsao 0001, Hsin-Min Wang, Tomoki Toda |
Interspeech | 7 |
| 2021 | Learning to Visualize Music Through Shot Sequence for Automatic Concert Video MashupabstractAn experienced director usually switches among different types of shots to make visual storytelling more touching. When filming a musical performance, appropriate switching shots can produce some special effects, such as enhancing the expression of emotion or heating up the atmosphere. However, while the visual storytelling technique is often used in making professional recordings of a live concert, amateur recordings of audiences often lack such storytelling concepts and skills when filming the same event. Thus a versatile system that can perform video mashup to create a refined high-quality video from such amateur clips is desirable. To this end, we aim at translating the music into an attractive shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. To achieve the task, we first introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. We then distill the knowledge in MF-RNNs with film-language into a lightweight RNN, which is more efficient and easier to deploy. The results from objective and subjective experiments demonstrate that both MF-RNNs with film-language and lightweight RNN can generate attractive shot sequences for music, thereby enhancing the viewing and listening experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hsiao-Rong Tyan, Hsin-Min Wang, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2020 | Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker VerificationabstractThis paper aims to improve speaker embedding representation based on x-vector for extracting more detailed information for speaker verification. We propose a statistics pooling time delay neural network (TDNN), in which the TDNN structure integrates statistics pooling for each layer, to consider the variation of temporal context in frame-level transformation. The proposed feature vector, named as statsvector, are compared with the baseline x-vector features on the VoxCeleb dataset and the Speakers in the Wild (SITW) dataset for speaker verification. The experimental results showed that the proposed stats-vector with score fusion achieved the best performance on VoxCeleb1 dataset. Furthermore, considering the interference from other speakers in the recordings, we found that the proposed statsvector efficiently reduced the interference and improved the speaker verification performance on the SITW dataset. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang, Chien-Lin Huang |
ICASSP | 3 |
| 2020 | Combining Deep Embeddings of Acoustic and Articulatory Features for Speaker IdentificationabstractIn this study, deep embedding of acoustic and articulatory features are combined for speaker identification. First, a convolutional neural network (CNN)-based universal background model (UBM) is constructed to generate acoustic feature (AC) embedding. In addition, as the articulatory features (AFs) represent some important phonological properties during speech production, a multilayer perceptron (MLP)-based AF embedding extraction model is also constructed for AF embedding extraction. The extracted AC and AF embeddings are concatenated as a combined feature vector for speaker identification using a fully-connected neural network. This proposed system was evaluated by three corpora consisting of King-ASR, LibriSpeech and SITW, and the experiments were conducted according to the properties of the datasets. We adopted all three corpora to evaluate the effect of AF embedding, and the results showed that combining AF embedding into the input feature vector improved the performance of speaker identification. The LibriSpeech corpus was used to evaluate the effect of the number of enrolled speakers. The proposed system achieved an EER of 7.80% outperforming the method based on x-vector with PLDA (8.25%). And we further evaluated the effect of signal mismatch using the SITW corpus. The proposed system achieved an EER of 25.19%, which outperformed the other baseline methods. Qian-Bei Hong, Chung-Hsien Wu 0001, Hsin-Min Wang, Chien-Lin Huang |
ICASSP | 3 |
| 2020 | Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech EnhancementabstractNonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for unseen noisy conditions are not guaranteed, which is a common weak point of supervised learning approaches. In this study, we proposed an unsupervised learning approach for speech enhancement, i.e., denoising autoencoder with linear regression decoder (DAELD) model for speech enhancement. The DAELD is trained with noisy speech as both input and target output in a self-supervised learning manner. In addition, with properly setting a shrinkage threshold for internal hidden representations, noise could be removed during the reconstruction from the hidden representations via the linear regression decoder. Speech enhancement experiments were carried out to test the proposed model. Results confirmed that the proposed DAELD could achieve comparable and sometimes even better enhancement performance as compared to the conventional supervised speech enhancement approaches, in both seen and unseen noise environments. Moreover, we observe that higher performances tend to achieve by DAELD when the training data cover more diverse noise types and signal-tonoise-ratio (SNR) levels. Ryandhimas E. Zezario, Tassadaq Hussain, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
ICASSP | 4 |
| 2020 | Lite Audio-Visual Speech Enhancementabstractstatus: Published Shang-Yi Chuang, Yu Tsao 0001, Chen-Chou Lo, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2020 | SERIL: Noise Adaptive Speech Enhancement Using Regularization-Based Incremental LearningabstractNumerous noise adaptation techniques have been proposed to fine-tune deep-learning models in speech enhancement (SE) for mismatched noise environments. Nevertheless, adaptation to a new environment may lead to catastrophic forgetting of the previously learned environments. The catastrophic forgetting issue degrades the performance of SE in real-world embedded devices, which often revisit previous noise environments. The nature of embedded devices does not allow solving the issue with additional storage of all pre-trained models or earlier training data. In this paper, we propose a regularization-based incremental learning SE (SERIL) strategy, complementing existing noise adaptation strategies without using additional storage. With a regularization constraint, the parameters are updated to the new noise environment while retaining the knowledge of the previous noise environments. The experimental results show that, when faced with a new noise domain, the SERIL model outperforms the unadapted SE model. Meanwhile, compared with the current adaptive technique based on fine-tuning, the SERIL model can reduce the forgetting of previous noise environments by 52%. The results verify that the SERIL model can effectively adjust itself to new noise environments while overcoming the catastrophic forgetting issue. The results make SERIL a favorable choice for real-world SE applications, where the noise environment changes frequently. Chi-Chang Lee, Yuchen Lin 0003, Hsuan-Tien Lin, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2020 | ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling |
Comput. Speech Lang. | 16 |
| 2020 | WaveCRN: An Efficient Convolutional Recurrent Neural Network for End-to-End Speech EnhancementabstractDue to the simple design pipeline, end-to-end (E2E) neural models for speech enhancement (SE) have attracted great interest. In order to improve the performance of the E2E model, the local and sequential properties of speech should be efficiently taken into account when modelling. However, in most current E2E models for SE, these properties are either not fully considered or are too complex to be realized. In this letter, we propose an efficient E2E SE model, termed WaveCRN. Compared with models based on convolutional neural networks (CNN) or long short-term memory (LSTM), WaveCRN uses a CNN module to capture the speech locality features and a stacked simple recurrent units (SRU) module to model the sequential property of the locality features. Different from conventional recurrent neural networks and LSTM, SRU can be efficiently parallelized in calculation, with even fewer model parameters. In order to more effectively suppress noise components in the noisy speech, we derive a novel restricted feature masking approach, which performs enhancement on the feature maps in the hidden layers; this is different from the approaches that apply the estimated ratio mask to the noisy spectral features, which is commonly used in speech separation methods. Experimental results on speech denoising and compressed speech restoration tasks confirm that with the SRU and the restricted feature map, WaveCRN performs comparably to other state-of-the-art approaches with notably reduced model complexity and inference time. Tsun-An Hsieh, Hsin-Min Wang, Xugang Lu, Yu Tsao 0001 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Subspace-Based Representation and Learning for Phonotactic Spoken Language RecognitionabstractPhonotactic constraints can be employed to distinguish languages by representing a speech utterance as a multinomial distribution or phone events. In the present study, we propose a new learning mechanism based on subspace-based representation, which can extract concealed phonotactic structures from utterances, for language verification and dialect/accent identification. The framework mainly involves two successive parts. The first part involves subspace construction. Specifically, it decodes each utterance into a sequence of vectors filled with phone-posteriors and transforms the vector sequence into a linear orthogonal subspace based on low-rank matrix factorization or dynamic linear modeling. The second part involves subspace learning based on kernel machines, such as support vector machines and the newly developed subspace-based neural networks (SNNs). The input layer of SNNs is specifically designed for the sample represented by subspaces. The topology ensures that the same output can be derived from identical subspaces by modifying the conventional feed-forward pass to fit the mathematical definition of subspace similarity. Evaluated on the “General LR” test of NIST LRE 2007, the proposed method achieved up to 52%, 46%, 56%, and 27% relative reductions in equal error rates over the sequence-based PPR-LM, PPR-VSM, and PPR-IVEC methods and the lattice-based PPR-LM method, respectively. Furthermore, on the dialect/accent identification task of NIST LRE 2009, the SNN-based system performed better than the aforementioned four baseline methods. Hung-Shin Lee, Yu Tsao 0001, Shyh-Kang Jeng, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Multichannel Speech Enhancement by Raw Waveform-Mapping Using Fully Convolutional NetworksabstractIn recent years, waveform-mapping-based speech enhancement (SE) methods have garnered significant attention. These methods generally use a deep learning model to directly process and reconstruct speech waveforms. Because both the input and output are in waveform format, the waveform-mapping-based SE methods can overcome the distortion caused by imperfect phase estimation, which may be encountered in spectral-mapping-based SE systems. So far, most waveform-mapping-based SE methods have focused on single-channel tasks. In this article, we propose a novel fully convolutional network (FCN) with Sinc and dilated convolutional layers (termed SDFCN) for multichannel SE that operates in the time domain. We also propose an extended version of SDFCN, called the residual SDFCN (termed rSDFCN). The proposed methods are evaluated on three multichannel SE tasks, namely the dual-channel inner-ear microphones SE task, the distributed microphones SE task, and the CHiME-3 dataset. The experimental results confirm the outstanding denoising capability of the proposed SE systems on the three tasks and the benefits of using the residual architecture on the overall SE performance. Chang-Le Liu, Sze-Wei Fu, You-Jin Li, Jen-Wei Huang, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2020 | Speech Enhancement Based on Denoising Autoencoder With Multi-Branched EncodersabstractDeep learning-based models have greatly advanced the performance of speech enhancement (SE) systems. However, two problems remain unsolved, which are closely related to model generalizability to noisy conditions: (1) mismatched noisy condition during testing, i.e., the performance is generally sub-optimal when models are tested with unseen noise types that are not involved in the training data; (2) local focus on specific noisy conditions, i.e., models trained using multiple types of noises cannot optimally remove a specific noise type even though the noise type has been involved in the training data. These problems are common in real applications. In this article, we propose a novel denoising autoencoder with a multi-branched encoder (termed DAEME) model to deal with these two problems. In the DAEME model, two stages are involved: training and testing. In the training stage, we build multiple component models to form a multi-branched encoder based on a decision tree (DSDT). The DSDT is built based on prior knowledge of speech and noisy conditions (the speaker, environment, and signal factors are considered in this paper), where each component of the multi-branched encoder performs a particular mapping from noisy to clean speech along the branch in the DSDT. Finally, a decoder is trained on top of the multi-branched encoder. In the testing stage, noisy speech is first processed by each component model. The multiple outputs from these models are then integrated into the decoder to determine the final enhanced speech. Experimental results show that DAEME is superior to several baseline models in terms of objective evaluation metrics, automatic speech recognition results, and quality in subjective human listening tests. Ryandhimas E. Zezario, Syu-Siang Wang, Jonathan Sherman, Yi-Yen Hsieh, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Spoken Multiple-Choice Question Answering Using Multimodal Convolutional Neural NetworksabstractIn a spoken multiple-choice question answering (MCQA) task, where passages, questions, and choices are given in the form of speech, usually only the auto-transcribed text is considered in system development. The acoustic-level information may contain useful cues for answer prediction. However, to the best of our knowledge, only a few studies focus on using the acoustic-level information or fusing the acoustic-level information with the text-level information for a spoken MCQA task. Therefore, this paper presents a hierarchical multistage multimodal (HMM) framework based on convolutional neural networks (CNNs) to integrate text- and acoustic-level statistics into neural modeling for spoken MCQA. Specifically, the acoustic-level statistics are expected to offset text inaccuracies caused by automatic speech recognition (ASR) systems or representation inadequacy lurking in word embedding generators, thereby making the spoken MCQA system robust. In the proposed HMM framework, two modalities are first manipulated to separately derive the acoustic- and text-level representations for the passage, question, and choices. Next, these clever features are jointly involved in inferring the relationships among the passage, question, and choices. Then, a final representation is derived for each choice, which encodes the relationship of the choice to the passage and question. Finally, the most likely answer is determined based on the individual final representations of all choices. Evaluated on the data of “Formosa Grand Challenge - Talk to AI”, a Mandarin Chinese spoken MCQA contest held in 2018, the proposed HMM framework achieves remarkable improvements in accuracy over the text-only baseline. Shang-Bao Luo, Hung-Shin Lee, Kuan-Yu Chen 0002, Hsin-Min Wang |
ASRU | 4 |
| 2019 | Reinforcement Learning Based Speech Enhancement for Robust Speech RecognitionabstractConventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is to minimize the recognition error, the recognition results should be used to design the objective function for optimizing the SE model. However, the structure of an ASR system, which consists of multiple units, such as acoustic and language models, is usually complex and not differentiable. In this study, we propose to adopt the reinforcement learning (RL) algorithm to optimize the SE model based on the recognition results. We evaluated the proposed RL-based SE system on the Mandarin Chinese broadcast news corpus (MATBN). Experimental results demonstrate that the proposed SE system can effectively improve the ASR results with a notable 12:40% and 19:23% error rate reductions for signal to noise ratio (SNR) at 0 dB and 5 dB conditions, respectively. Yih-Liang Shen, Chao-Yuan Huang, Syu-Siang Wang, Yu Tsao 0001, Hsin-Min Wang, Tai-Shih Chi |
ICASSP | 5 |
| 2019 | Exploring the Encoder Layers of Discriminative Autoencoders for LVCSR
Pin-Tuan Huang, Hung-Shin Lee, Syu-Siang Wang, Kuan-Yu Chen 0002, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 6 |
| 2019 | Investigation of F0 Conditioning and Fully Convolutional Networks in Variational Autoencoder Based Voice ConversionabstractIn this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship between vocoder features extracted using the high quality vocoders adopted in conventional VC systems, and hypothesize that the spectral features are in fact F0 dependent. Such hypothesis implies that during the conversion phase, the latent codes and the converted features in VAE based VC are in fact source F0 dependent. To this end, we propose to utilize the F0 as an additional input of the decoder. The model can learn to disentangle the latent code from the F0 and thus generates converted F0 dependent converted features. Second, to better capture temporal dependencies of the spectral features and the F0 pattern, we replace the frame wise conversion structure in the original VAE based VC framework with a fully convolutional network structure. Our experiments demonstrate that the degree of disentanglement as well as the naturalness of the converted speech are indeed improved. Wen-Chin Huang, Yi-Chiao Wu, Chen-Chou Lo, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 9 |
| 2019 | Noise Adaptive Speech Enhancement Using Domain Adversarial TrainingabstractIn this study, we propose a novel noise adaptive speech enhancement (SE) system, which employs a domain adversarial training (DAT) approach to tackle the issue of a noise type mismatch between the training and testing conditions. Such a mismatch is a critical problem in deep-learning-based SE systems. A large mismatch may cause a serious performance degradation to the SE performance. Because we generally use a well-trained SE system to handle various unseen noise types, a noise type mismatch commonly occurs in real-world scenarios. The proposed noise adaptive SE system contains an encoder-decoder-based enhancement model and a domain discriminator model. During adaptation, the DAT approach encourages the encoder to produce noise-invariant features based on the information from the discriminator model and consequentially increases the robustness of the enhancement model to unseen noise types. Herein, we regard stationary noises as the source domain (with the ground truth of clean speech) and non-stationary noises as the target domain (without the ground truth). We evaluated the proposed system on TIMIT sentences. The experiment results show that the proposed noise adaptive SE system successfully provides significant improvements in PESQ (19.0%), SSNR (39.3%), and STOI (27.0%) over the SE system without an adaptation. Chien-Feng Liao, Yu Tsao 0001, Hung-yi Lee, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2019 | MOSNet: Deep Learning-Based Objective Assessment for Voice ConversionabstractExisting objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In this paper, we propose deep learning-based assessment models to predict human ratings of converted speech. We adopt the convolutional and recurrent neural network models to build a mean opinion score (MOS) predictor, termed as MOSNet. The proposed models are tested on large-scale listening test results of the Voice Conversion Challenge (VCC) 2018. Experimental results show that the predicted scores of the proposed MOSNet are highly correlated with human MOS ratings at the system level while being fairly correlated with human MOS ratings at the utterance level. Meanwhile, we have modified MOSNet to predict the similarity scores, and the preliminary results show that the predicted scores are also fairly correlated with human ratings. These results confirm that the proposed models could be used as a computational evaluator to measure the MOS of VC systems to reduce the need for expensive human rating. Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang 0037, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 7 |
| 2019 | Specialized Speech Enhancement Model Selection Based on Learned Non-Intrusive Quality Assessment Metric
Ryandhimas E. Zezario, Szu-Wei Fu, Xugang Lu, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 4 |
| 2018 | Essence Vector-Based Query Modeling for Spoken Document RetrievalabstractSpoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods and investigating the effectiveness of these methods on the SDR task. This paper first presents a novel paragraph embedding method, named the essence vector (EV) model, which aims at inferring a representation for a given paragraph by encapsulating the most representative information from the paragraph and excluding the general background information at the same time. On top of the EV model, we develop three query language modeling mechanisms to improve the retrieval performance. A series of empirical SDR experiments conducted on two benchmark collections demonstrate the good efficacy of the proposed framework, compared to several existing strong baseline systems. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 4 |
| 2018 | Seethevoice: Learning from Music to Visual Storytelling of ShotsabstractTypes of shots in the language of film are considered the key elements used by a director for visual storytelling. In filming a musical performance, manipulating shots could stimulate desired effects such as manifesting the emotion or deepening the atmosphere. However, while the visual storytelling technique is often employed in creating professional recordings of a live concert, audience recordings of the same event often lack such sophisticated manipulations. Thus it would be useful to have a versatile system that can perform video mashup to create a refined video from such amateur clips. To this end, we propose to translate the music into a near-professional shot (type) sequence by learning the relation between music and visual storytelling of shots. The resulting shot sequence can then be used to better portray the visual storytelling of a song and guide the concert video mashup process. Our method introduces a novel probabilistic-based fusion approach, named as multi-resolution fused recurrent neural networks (MF-RNNs) with film-language, which integrates multi-resolution fused RNNs and a film-language model for boosting the translation performance. The results from objective and subjective experiments demonstrate that MF-RNNs with film-language can generate an appealing shot sequence with better viewing experience. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICME | 5 |
| 2018 | Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model Based on BLSTMabstractNowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech with its clean counterpart.The need of a "golden" reference considerably restricts the practicality of such assessment tools in real-world scenarios since the clean reference usually cannot be accessed.On the other hand, human beings can readily evaluate the speech quality without any reference (e.g., mean opinion score (MOS) tests), implying the existence of an objective and non-intrusive (no clean reference needed) quality assessment mechanism.In this study, we propose a novel endto-end, non-intrusive speech quality evaluation model, termed Quality-Net, based on bidirectional long short-term memory.The evaluation of utterance-level quality in Quality-Net is based on the frame-level assessment.Frame constraints and sensible initializations of forget gate biases are applied to learn meaningful frame-level quality assessment from the utterancelevel quality label.Experimental results show that Quality-Net can yield high correlation to PESQ (0.9 for the noisy speech and 0.84 for the speech processed by speech enhancement).We believe that Quality-Net has potential to be used in a wide variety of applications of speech signal processing. Szu-Wei Fu, Yu Tsao 0001, Hsin-Te Hwang, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2018 | Exemplar-Based Spectral Detail Compensation for Voice Conversion
Yu-Huai Peng, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 5 |
| 2018 | An Information Distillation Framework for Extractive SummarizationabstractIn the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some realistic tasks such as document summarization. Nevertheless, classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions in this paper are threefold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph of interest. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. Third, a new summarization framework, which can take both relevance and redundancy information into account simultaneously, is also introduced. We evaluate the proposed embedding methods (i.e., EV and D-EV) and the summarization framework on two benchmark summarization corpora. The experimental results demonstrate the effectiveness and applicability of the proposed framework in relation to several well-practiced and state-of-the-art summarization methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Coherent Deep-Net Fusion To Classify Shots In Concert VideosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director. The technique is often used in creating professional recordings of a live concert, but meanwhile may not be appropriately applied in audience recordings of the same event. Such variations could cause the task of classifying shots in concert videos, professional or amateur, very challenging. To achieve more reliable shot classification, we propose a novel probabilistic-based approach, named as coherent classification net (CC-Net), by addressing three crucial issues. First, we focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pretrained on a large-scale data set for object recognition. Second, we introduce a frame-wise classification scheme, the error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently, but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for a classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification in each image frame. Third, we feed the frame-wise classification results to a linear-chain conditional random field module to refine the shot predictions by taking into account the global and temporal regularities. We provide extensive experimental results on a data set of live concert videos to demonstrate the advantage of the proposed CC-Net over existing popular fusion approaches for shot classification. Jen-Chun Lin, Wen-Li Wei, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
IEEE Trans. Multim. | 5 |
| 2017 | Neural relevance-aware query modeling for spoken document retrievalabstractSpoken document retrieval (SDR) is becoming a much-needed application due to that unprecedented volumes of audio-visual media have been made available in our daily life. As far as we are aware, most of the wide variety of SDR methods mainly focus on exploring robust indexing and effective retrieval methods to quantify the relevance degree between a pair of query and document. However, similar to information retrieval (IR), a fundamental challenge facing SDR is that a query is usually too short to convey a user's information need, such that a retrieval system cannot always achieve prospective efficacy when with the existing retrieval methods. In order to further boost retrieval performance, several studies turn their attention to reformulating the original query by leveraging an online pseudo-relevance feedback (PRF) process, which often comes at the price of taking significant time. Motivated by these observations, this paper presents a novel extension of the general line of SDR research and its contribution is at least two-fold. First, building on neural network-based techniques, we put forward a neural relevance-aware query modeling (NRM) framework, which is designed to not only infer a discriminative query language model automatically for a given query, but also get around the time-consuming PRF process. Second, the utility of the methods instantiated from our proposed framework and several widely-used retrieval methods are extensively analyzed and compared on a standard SDR task, which suggests the superiority of our methods. Tien-Hong Lo, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
ASRU | 4 |
| 2017 | A locality-preserving essence vector modeling framework for spoken document retrievalabstractBecause unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machine learning applications owing largely to its excellent performance. In the context of natural language processing, the pioneering work can date back to the word embedding methods. However, learning of paragraph (or sentence and document) representations is more reasonable and suitable for some tasks, such as information retrieval and document summarization. Nevertheless, as far as we are aware, there is relatively less work focusing on launching paragraph embedding methods into SDR. Motivated by these observations, this paper proposes a novel paragraph embedding method, named the locality-preserving essence vector (LPEV) model. LPEV is designed with consideration to two aspects. First, the model aims at not only distilling the most representative information from a paragraph but also getting rid of the general background information. Second, inspired by the local invariance perspective, which is a celebrated principle used in manifold learning techniques, LPEV also manages to preserve semantic locality in the learned low-dimensional embedding space for producing more informative and discriminative vector representations of paragraphs. On top of the proposed framework, a series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good efficacy of our SDR methods as compared to existing strong baselines. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 4 |
| 2017 | Discriminative autoencoders for speaker verificationabstractThis paper presents a learning and scoring framework based on neural networks for speaker verification. The framework employs an autoencoder as its primary structure while three factors are jointly considered in the objective function for speaker discrimination. The first one, relating to the sample reconstruction error, makes the structure essentially a generative model, which benefits to learn most salient and useful properties of the data. Functioning in the middlemost hidden layer, the other two attempt to ensure that utterances spoken by the same speaker are mapped into similar identity codes in the speaker discriminative subspace, where the dispersion of all identity codes are maximized to some extent so as to avoid the effect of over-concentration. Finally, the decision score of each utterance pair is simply computed by cosine similarity of their identity codes. Dealing with utterances represented by i-vectors, the results of experiments conducted on the male portion of the core task in the NIST 2010 Speaker Recognition Evaluation (SRE) significantly demonstrate the merits of our approach over the conventional PLDA method. Hung-Shin Lee, Yu-Ding Lu, Chin-Cheng Hsu, Yu Tsao 0001, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 5 |
| 2017 | Leveraging manifold learning for extractive broadcast news summarizationabstractExtractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech summarization. Manifold learning has experienced a surge of research interest in various domains concerned with dimensionality reduction and data representation recently, but has so far been largely under-explored in extractive text or speech summarization. Our contributions in this paper are at least twofold. First, we explore the use of several manifold learning algorithms to capture the latent semantic information of sentences for enhanced extractive speech summarization, including isometric feature mapping (ISOMAP), locally linear embedding (LLE) and Laplacian eigenmap. Second, the merits of our proposed summarization methods and several widely-used methods are extensively analyzed and compared. The empirical results demonstrate the effectiveness of our unsupervised summarization methods, in relation to several state-of-the-art methods. In particular, a synergy of the manifold learning based methods and state-of-the-art methods, such as the integer linear programming (ILP) method, contributes to further gains in summarization performance. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu |
ICASSP | 4 |
| 2017 | Speech emotion recognition with skew-robust neural networksabstractWe propose a neural-network training algorithm that is robust to data imbalance in classification. In our proposed algorithm, weights are introduced to training examples, effectively modifying the trajectory traversed in the parameter space during the learning process. Furthermore, the proposed algorithm would reduce to the normal stochastic gradient decent learning if the data is balanced. On the FAU-Aibo database, which is known to be used in Interspeech Emotion Challenge, the proposed method achieves an unweighted average (UA) recall rate of 45.3% on the 5-class speech emotion recognition task. Within the static modeling framework, where each example is represented as a fixed-length vector, this performance is one of the best performance ever achieved on the 5-class task. Po-Yuan Shih, Chia-Ping Chen, Hsin-Min Wang |
ICASSP | 3 |
| 2017 | Deep-net fusion to classify shots in concert videosabstractVarying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. We first focus on learning more effective features by fusing the layer-wise outputs extracted from a deep convolutional neural network (CNN), pre-trained on a large-scale dataset for object recognition. We then introduce a probabilistic fusion model, termed as error weighted deep cross-correlation model (EW-Deep-CCM), to boost the classification accuracy. Specifically, the deep neural network-based cross-correlation model (Deep-CCM) is constructed to not only model the extracted feature hierarchies of CNN independently but also relate the statistical dependencies of paired features from different layers. Then, a Bayesian error weighting scheme for classifier combination is adopted to explore the contributions from individual Deep-CCM classifiers to enhance the accuracy of shot classification. We provide extensive experimental results on a dataset of live concert videos to demonstrate the advantage of the proposed EW-Deep-CCM over existing popular fusion approaches. The video demos can be found at https://sites.google.com/site/ewdeepccm2/demo. Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Yi-Hsuan Yang, Hsin-Min Wang, Hsiao-Rong Tyan, Hong-Yuan Mark Liao |
ICASSP | 5 |
| 2017 | A locally linear embbeding based postfiltering approach for speech enhancementabstractThis paper presents a novel postfiltering approach based on the locally linear embedding (LLE) algorithm for speech enchantment (SE). The aim of the proposed LLE-based postfiltering approach is to further remove the residual noise components from the SE-processed speech signals through a spectral conversion process, thereby increasing the signal-to-noise ratio (SNR) and speech quality. The proposed postfiltering approach consists of two phases. In the offline phase, paired SE-processed and clean speech exemplars are prepared for dictionary construction. In the online phase, the LLE algorithm is adopted to convert the SE-processed speech signals to the clean ones. The present study integrates the LLE-based postfiltering approach with a deep denoising autoencoder (DDAE) SE method, which has been confirmed to provide outstanding capability for noise reduction. Experimental results show that the proposed postfiltering approach can notably enhance the DDAE-based SE processed speech signals in different noise types and SNR levels. Yi-Chiao Wu, Hsin-Te Hwang, Syu-Siang Wang, Chin-Cheng Hsu, Ying-Hui Lai, Yu Tsao 0001, Hsin-Min Wang |
ICASSP | 7 |
| 2017 | Exploring the Use of Significant Words Language Modeling for Spoken Document Retrieval
Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
INTERSPEECH | 3 |
| 2017 | Voice Conversion from Unaligned Corpora Using Variational Autoencoding Wasserstein Generative Adversarial NetworksabstractBuilding a voice conversion (VC) system from non-parallel speech corpora is challenging but highly valuable in real application scenarios.In most situations, the source and the target speakers do not repeat the same texts or they may even speak different languages.In this case, one possible, although indirect, solution is to build a generative model for speech.Generative models focus on explaining the observations with latent variables instead of learning a pairwise transformation function, thereby bypassing the requirement of speech frame alignment.In this paper, we propose a non-parallel VC framework with a variational autoencoding Wasserstein generative adversarial network (VAW-GAN) that explicitly considers a VC objective when building the speech model.Experimental results corroborate the capability of our framework for building a VC system from unaligned data, and demonstrate improved conversion quality. Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 5 |
| 2017 | Wavelet Speech Enhancement Based on Robust Principal Component Analysis
Chia-Lung Wu, Hsiang-Ping Hsu, Syu-Siang Wang, Jeih-Weih Hung, Ying-Hui Lai, Hsin-Min Wang, Yu Tsao 0001 |
INTERSPEECH | 6 |
| 2017 | A Post-Filtering Approach Based on Locally Linear Embedding Difference Compensation for Speech Enhancement
Yi-Chiao Wu, Hsin-Te Hwang, Syu-Siang Wang, Chin-Cheng Hsu, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 6 |
| 2017 | Discriminative Autoencoders for Acoustic Modeling
Ming-Han Yang, Hung-Shin Lee, Yu-Ding Lu, Kuan-Yu Chen 0002, Yu Tsao 0001, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 7 |
| 2017 | Automatic Music Video Generation Based on Simultaneous Soundtrack Recommendation and Video EditingabstractAn automated process that can suggest a soundtrack to a user-generated video (UGV) and make the UGV a music-compliant professional-like video is challenging but desirable. To this end, this paper presents an automatic music video (MV) generation system that conducts soundtrack recommendation and video editing simultaneously. Given a long UGV, it is first divided into a sequence of fixed-length short (e.g., 2 seconds) segments, and then a multi-task deep neural network (MDNN) is applied to predict the pseudo acoustic (music) features (or called the pseudo song) from the visual (video) features of each video segment. In this way, the distance between any pair of video and music segments of same length can be computed in the music feature space. Second, the sequence of pseudo acoustic (music) features of the UGV and the sequence of the acoustic (music) features of each music track in the music collection are temporarily aligned by the dynamic time warping (DTW) algorithm with a pseudo-song-based deep similarity matching (PDSM) metric. Third, for each music track, the video editing module selects and concatenates the segments of the UGV based on the target and concatenation costs given by a pseudo-song-based deep concatenation cost (PDCC) metric according to the DTW-aligned result to generate a music-compliant professional-like video. Finally, all the generated MVs are ranked, and the best MV is recommended to the user. The MDNN for pseudo song prediction and the PDSM and PDCC metrics are trained by an annotated official music video (OMV) corpus. The results of objective and subjective experiments demonstrate that the proposed system performs well and can generate appealing MVs with better viewing and listening experiences. Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang, Hong-Yuan Mark Liao |
ACM Multimedia | 4 |
| 2017 | A Position-Aware Language Modeling Framework for Extractive Broadcast News Speech SummarizationabstractExtractive summarization, a process that automatically picks exemplary sentences from a text (or spoken) document with the goal of concisely conveying key information therein, has seen a surge of attention from scholars and practitioners recently. Using a language modeling (LM) approach for sentence selection has been proven effective for performing unsupervised extractive summarization. However, one of the major difficulties facing the LM approach is to model sentences and estimate their parameters more accurately for each text (or spoken) document. We extend this line of research and make the following contributions in this work. First, we propose a position-aware language modeling framework using various granularities of position-specific information to better estimate the sentence models involved in the summarization process. Second, we explore disparate ways to integrate the positional cues into relevance models through a pseudo-relevance feedback procedure. Third, we extensively evaluate various models originated from our proposed framework and several well-established unsupervised methods. Empirical evaluation conducted on a broadcast news summarization task further demonstrates performance merits of the proposed summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2016 | Learning to Distill: The Essence Vector Modeling FrameworkabstractIn the context of natural language processing, representation learning has emerged as a newly active research subject because of its excellent performance in many applications. Learning representations of words is a pioneering study in this school of research. However, paragraph (or sentence and document) embedding learning is more suitable/reasonable for some tasks, such as sentiment classification and document summarization. Nevertheless, as far as we are aware, there is only a dearth of research focusing on launching unsupervised paragraph embedding methods. Classic paragraph embedding methods infer the representation of a given paragraph by considering all of the words occurring in the paragraph. Consequently, those stop or function words that occur frequently may mislead the embedding learning process to produce a misty paragraph representation. Motivated by these observations, our major contributions are twofold. First, we propose a novel unsupervised paragraph embedding method, named the essence vector (EV) model, which aims at not only distilling the most representative information from a paragraph but also excluding the general background information to produce a more informative low-dimensional vector representation for the paragraph. We evaluate the proposed EV model on benchmark sentiment classification and multi-document summarization tasks. The experimental results demonstrate the effectiveness and applicability of the proposed embedding method. Second, in view of the increasing importance of spoken content processing, an extension of the EV model, named the denoising essence vector (D-EV) model, is proposed. The D-EV model not only inherits the advantages of the EV model but also can infer a more robust representation for a given spoken paragraph against imperfect speech recognition. The utility of the D-EV model is evaluated on a spoken document summarization task, confirming the effectiveness of the proposed embedding method in relation to several well-practiced and state-of-the-art summarization methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
COLING | 4 |
| 2016 | Improved spoken document summarization with coverage modeling techniquesabstractExtractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods designed to model coverage can be characterized by either reducing redundancy or increasing diversity in the summary. Maximal margin relevance (MMR) is a widely-cited method since it takes both relevance and redundancy into account when generating a summary for a given document. In addition to MMR, there is only a dearth of research concentrating on reducing redundancy or increasing diversity for the spoken document summarization task, as far as we are aware. Motivated by these observations, two major contributions are presented in this paper. First, in contrast to MMR, which considers coverage by reducing redundancy, we propose two novel coverage-based methods, which directly increase diversity. With the proposed methods, a set of representative sentences, which not only are relevant to the given document but also cover most of the important sub-themes of the document, can be selected automatically. Second, we make a step forward to plug in several document/sentence representation methods into the proposed framework to further enhance the summarization performance. A series of empirical evaluations demonstrate the effectiveness of our proposed methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ICASSP | 4 |
| 2016 | DEMV-matchmaker: Emotional temporal course representation and deep similarity matching for automatic music video generationabstractThis paper presents a deep similarity matching-based emotion-oriented music video (MV) generation system, called DEMV-matchmaker, which utilizes an emotion-oriented deep similarity matching (EDSM) metric as a bridge to connect music and video. Specifically, we adopt an emotional temporal course model (ETCM) to respectively learn the relationship between music and its emotional temporal phase sequence and the relationship between video and its emotional temporal phase sequence from an emotion-annotated MV corpus. An emotional temporal structure preserved histogram (ETPH) representation is proposed to keep the recognized emotional temporal phase sequence information for EDSM metric construction. A deep neural network (DNN) is then applied to learn an EDSM metric based on the ETPHs for the given positive (official) and negative (artificial) MV examples. For MV generation, the EDSM metric is applied to measure the similarity between ETPHs of video and music. The results of objective and subjective experiments demonstrate that DEMV-matchmaker performs well and can generate appealing music videos that can enhance the viewing and listening experience. Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang |
ICASSP | 3 |
| 2016 | Minimization of Regression and Ranking Losses with Shallow Neural Networks on Automatic Sincerity Evaluation
Hung-Shin Lee, Yu Tsao 0001, Chi-Chun Lee, Hsin-Min Wang, Wei-Chen Chen, Shan-Wen Hsiao, Shyh-Kang Jeng |
INTERSPEECH | 4 |
| 2016 | Exploring Word Mover's Distance and Semantic-Aware Embedding Techniques for Extractive Broadcast News Summarization
Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 5 |
| 2016 | Locally Linear Embedding for Exemplar-Based Spectral Conversion
Yi-Chiao Wu, Hsin-Te Hwang, Chin-Cheng Hsu, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 5 |
| 2016 | Novel Word Embedding and Translation-based Language Modeling for Extractive Speech SummarizationabstractWord embedding methods revolve around learning continuous distributed vector representations of words with neural networks, which can capture semantic and/or syntactic cues, and in turn be used to induce similarity measures among words, sentences and documents in context. Celebrated methods can be categorized as prediction-based and count-based methods according to the training objectives and model architectures. Their pros and cons have been extensively analyzed and evaluated in recent studies, but there is relatively less work continuing the line of research to develop an enhanced learning method that brings together the advantages of the two model families. In addition, the interpretation of the learned word representations still remains somewhat opaque. Motivated by the observations and considering the pressing need, this paper presents a novel method for learning the word representations, which not only inherits the advantages of classic word embedding methods but also offers a clearer and more rigorous interpretation of the learned word representations. Built upon the proposed word embedding method, we further formulate a translation-based language modeling framework for the extractive speech summarization task. A series of empirical evaluations demonstrate the effectiveness of the proposed word representation learning and language modeling techniques in extractive speech summarization. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen |
ACM Multimedia | 4 |
| 2016 | Automatic Music Video Generation Based on Emotion-Oriented Pseudo Song Prediction and MatchingabstractThe main difficulty in automatic music video (MV) generation lies in how to match two different media (i.e., video and music). This paper proposes a novel content-based MV generation system based on emotion-oriented pseudo song prediction and matching. We use a multi-task deep neural network (MDNN) to jointly learn the relationship among music, video, and emotion from an emotion-annotated MV corpus. Given a queried video, the MDNN is applied to predict the acoustic (music) features from the visual (video) features, i.e., the pseudo song corresponding to the video. Then, the pseudo acoustic (music) features are matched with the acoustic (music) features of each music track in the music collection according to a pseudo-song-based deep similarity matching (PDSM) metric given by another deep neural network (DNN) trained on the acoustic and pseudo acoustic features of the positive (official), less-positive (artificial), and negative (artificial) MV examples. The results of objective and subjective experiments demonstrate that the proposed pseudo-song-based framework performs well and can generate appealing MVs with better viewing and listening experiences. Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang |
ACM Multimedia | 3 |
| 2016 | Exploring the use of unsupervised query modeling techniques for speech recognition and summarization
Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Hsin-Hsi Chen |
Speech Commun. | 4 |
| 2016 | Alignment of Lyrics With Accompanied Singing Audio Based on Acoustic-Phonetic Vowel Likelihood ModelingabstractThis study addresses the task of aligning lyrics with accompanied singing recordings. With a vowel-only representation of lyric syllables, our approach evaluates likelihood scores of vowel types with glottal pulse shapes and formant frequencies extracted from a small set of singing examples. The proposed vowel likelihood model is used in conjunction with a prior model of frame-wise syllable sequence in determining an optimal evolution of syllabic position. In lyrics alignment experiments, we optimized numerical parameters on two independent development sets and then tested the optimized system on two other datasets. New objective performance measures are introduced in the evaluation to provide further insight into the quality of alignment. Use of glottal pulse shapes and formant frequencies is shown by a controlled experiment to account for a 0.07 difference in average normalized alignment error. Another controlled experiment demonstrates that, with a difference of 0.03, F0-invariant glottal pulse shape gives a lower average normalized alignment error than does F0-invariant spectrum envelope, the latter being assumed by MFCC-based timbre models. Yu-Ren Chien, Hsin-Min Wang, Shyh-Kang Jeng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Incorporating paragraph embeddings and density peaks clustering for spoken document summarizationabstractRepresentation learning has emerged as a newly active research subject in many machine learning applications because of its excellent performance. As an instantiation, word embedding has been widely used in the natural language processing area. However, as far as we are aware, there are relatively few studies investigating paragraph embedding methods in extractive text or speech summarization. Extractive summarization aims at selecting a set of indicative sentences from a source document to express the most important theme of the document. There is a general consensus that relevance and redundancy are both critical issues for users in a realistic summarization scenario. However, most of the existing methods focus on determining only the relevance degree between sentences and a given document, while the redundancy degree is calculated by a post-processing step. Based on these observations, three contributions are proposed in this paper. First, we comprehensively compare the word and paragraph embedding methods for spoken document summarization. Next, we propose a novel summarization framework which can take both relevance and redundancy information into account simultaneously. Consequently, a set of representative sentences can be automatically selected through a one-pass process. Third, we further plug in paragraph embedding methods into the proposed framework to enhance the summarization performance. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Kai-Wun Shih, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang |
ASRU | 5 |
| 2015 | I-vector based language modeling for query representationabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Following the research tendency, many efforts have been devoted towards developing indexing and modeling techniques for representing spoken documents, but only few have been made on improving query formulation for better representing users' information needs. The i-vector based language modeling (IVLM) framework, stemming from the state-of-the-art i-vector framework for language identification and speaker recognition, has been proposed and formulated to represent documents in SDR with good promise recently. However, a major challenge of using IVLM for query modeling is that a query usually consists of only a few words; thus, it is hard to learn a reliable representation accordingly. In this paper, we focus our attention on query reformulation and propose three novel methods on top of IVLM to more accurately represent users' information needs. In addition, we also explore the use of multi-levels of index features, including word- and subword-level units, to work in concert with the proposed methods. A series of empirical SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the good effectiveness of our proposed methods as compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 2 |
| 2015 | A histogram density modeling approach to music emotion recognitionabstractMusic emotion recognition is concerned with developing predictive models that comprehend the affective content of musical signals. Recently, a growing number of attempts has been made to model the music emotion as a probability distribution in the valence-arousal (VA) space to better account for the subjectivity. In this paper, we present a novel histogram density modeling approach that models the emotion distribution by a 2-D histogram over the quantized VA space and learns a set of latent histograms to predict the emotion probability density of a song from audio. The proposed model is free from parametric distribution assumptions over the VA space, easy to implement, and extremely fast to train. We also extend our model to deal with the temporal dynamics of time-varying emotion labels. Comprehensive performance study on two larger-scale datasets demonstrates that our approach achieves comparable performance to the state-of-the-art ones, but with much better training and testing efficiency. Ju-Chiang Wang, Hsin-Min Wang, Gert R. G. Lanckriet |
ICASSP | 2 |
| 2015 | Leveraging word embeddings for spoken document summarizationabstractOwing to the rapidly growing multimedia content available on the Internet, extractive spoken document summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document to concisely express the most important theme of the document, has been an active area of research and experimentation. On the other hand, word embedding has emerged as a newly favorite research subject because of its excellent performance in many natural language processing (NLP)-related tasks. However, as far as we are aware, there are relatively few studies investigating its use in extractive text or speech summarization. A common thread of leveraging word embeddings in the summarization process is to represent the document (or sentence) by averaging the word embeddings of the words occurring in the document (or sentence). Then, intuitively, the cosine similarity measure can be employed to determine the relevance degree between a pair of representations. Beyond the continued efforts made to improve the representation of words, this paper focuses on building novel and efficient ranking models based on the general word embedding methods for extractive speech summarization. Experimental results demonstrate the effectiveness of our proposed methods, compared to existing state-of-the-art methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
INTERSPEECH | 3 |
| 2015 | Positional language modeling for extractive broadcast news speech summarizationabstractExtractive summarization, with the intention of automatically selecting a set of representative sentences from a text (or spoken) document so as to concisely express the most important theme of the document, has been an active area of experimentation and development.A recent trend of research is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing extractive summarization in an unsupervised fashion.However, one of the major challenges facing the LM approach is how to formulate the sentence models and estimate their parameters more accurately for each text (or spoken) document to be summarized.This paper extends this line of research and its contributions are three-fold.First, we propose a positional language modeling framework using different granularities of position-specific information to better estimate the sentence models involved in summarization.Second, we also explore to integrate the positional cues into relevance modeling through a pseudo-relevance feedback procedure.Third, the utilities of the various methods originated from our proposed framework and several well-established unsupervised methods are analyzed and compared extensively.Empirical evaluations conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 4 |
| 2015 | EMV-matchmaker: Emotional Temporal Course Modeling and Matching for Automatic Music Video GenerationabstractThis paper presents a novel content-based emotion-oriented music video (MV) generation system, called EMV-matchmaker, which utilizes the emotional temporal phase sequence of the multimedia content as a bridge to connect music and video. Specifically, we adopt an emotional temporal course model (ETCM) to respectively learn the relationship between music and its emotional temporal phase sequence and the relationship between video and its emotional temporal phase sequence from an emotion-annotated MV corpus. Then, given a video clip (or a music clip), the visual (or acoustic) ETCM is applied to predict its emotional temporal phase sequence in a valence-arousal (VA) emotional space from the corresponding low-level visual (or acoustic) features. For MV generation, string matching is applied to measure the similarity between the emotional temporal phase sequences of video and music. The results of objective and subjective experiments demonstrate that EMV-matchmaker performs well and can generate appealing music videos that can enhance the viewing and listening experience. Jen-Chun Lin, Wen-Li Wei, Hsin-Min Wang |
ACM Multimedia | 3 |
| 2015 | Modeling the Affective Content of Music with a Gaussian Mixture ModelabstractModeling the association between music and emotion has been considered important for music information retrieval and affective human computer interaction. This paper presents a novel generative model called acoustic emotion Gaussians (AEG) for computational modeling of emotion. Instead of assigning a music excerpt with a deterministic (hard) emotion label, AEG treats the affective content of music as a (soft) probability distribution in the valence-arousal space and parameterizes it with a Gaussian mixture model (GMM). In this way, the subjective nature of emotion perception is explicitly modeled. Specifically, AEG employs two GMMs to characterize the audio and emotion data. The fitting algorithm of the GMM parameters makes the model learning process transparent and interpretable. Based on AEG, a probabilistic graphical structure for predicting the emotion distribution from music audio data is also developed. A comprehensive performance study over two emotion-labeled datasets demonstrates that AEG offers new insights into the relationship between music and emotion (e.g., to assess the “affective diversity” of a corpus) and represents an effective means of emotion modeling. Readers can easily implement AEG via the publicly available codes. As the AEG model is generic, it holds the promise of analyzing any signal that carries affective or other highly subjective information. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
IEEE Trans. Affect. Comput. | 3 |
| 2015 | A Probabilistic Framework for Chinese Spelling CheckabstractChinese spelling check (CSC) is still an unsolved problem today since there are many homonymous or homomorphous characters. Recently, more and more CSC systems have been proposed. To the best of our knowledge, language modeling is one of the major components among these systems because of its simplicity and moderately good predictive power. After deeply analyzing the school of research, we are aware that most of the systems only employ the conventional n -gram language models. The contributions of this article are threefold. First, we propose a novel probabilistic framework for CSC, which naturally combines several important components, such as the substitution model and the language model, to inherit their individual merits as well as to overcome their limitations. Second, we incorporate the topic language models into the CSC system in an unsupervised fashion. The topic language models can capture the long-span semantic information from a word (character) string while the conventional n -gram language models can only preserve the local regularity information. Third, we further integrate Web resources with the proposed framework to enhance the overall performance. Our rigorously empirical experiments demonstrate the consistent and utility performance of the proposed framework in the CSC task. Kuan-Yu Chen 0002, Hsin-Min Wang, Hsin-Hsi Chen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Extractive Broadcast News Summarization Leveraging Recurrent Neural Network Language Modeling TechniquesabstractExtractive text or speech summarization manages to select a set of salient sentences from an original document and concatenate them to form a summary, enabling users to better browse through and understand the content of the document. A recent stream of research on extractive summarization is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each sentence in the document to be summarized. In view of this, our work in this paper explores a novel use of recurrent neural network language modeling (RNNLM) framework for extractive broadcast news summarization. On top of such a framework, the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within broadcast news documents, getting around the need for the strict bag-of-words assumption. Furthermore, different model complexities and combinations are extensively analyzed and compared. Experimental results demonstrate the performance merits of our summarization methods when compared to several well-studied state-of-the-art unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Ea-Ee Jan, Wen-Lian Hsu, Hsin-Hsi Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | An Acoustic-Phonetic Model of F0 Likelihood for Vocal Melody ExtractionabstractThis paper presents a novel approach to extraction of vocal melodies from accompanied singing recordings. Central to our approach is a model of vocal fundamental frequency (F0) likelihood that integrates acoustic-phonetic knowledge and real-world data. This model consists of a timbral fitness score and a loudness measure of each F0 candidate. Timbral fitness is measured for the partial amplitudes of an F0 candidate, with respect to a small set of vocal timbre examples. This F0-specific measurement of timbral fitness depends on an acoustic-phonetic F0 modification of each timbre example. In the loudness part of the likelihood model, sinusoids are detected, tracked, and pruned to give loudness values that minimize interference from the accompaniment. A final F0 estimate is determined by a prior model of F0 sequence in addition to the likelihood model. Melody extraction is completed by detecting voiced time positions according to the singing voice loudness variations given by the estimated F0 sequence. The numerical parameters involved in our approach were optimized on three development sets from different sources before the system was evaluated on ten test sets separate from these development sets. Controlled experiments show that use of the timbral fitness score accounts for a 13% difference in overall accuracy. Yu-Ren Chien, Hsin-Min Wang, Shyh-Kang Jeng |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Combining Relevance Language Modeling and Clarity Measure for Extractive Speech SummarizationabstractExtractive speech summarization, which purports to select an indicative set of sentences from a spoken document so as to succinctly represent the most important aspects of the document, has garnered much research over the years. In this paper, we cast extractive speech summarization as an ad-hoc information retrieval (IR) problem and investigate various language modeling (LM) methods for important sentence selection. The main contributions of this paper are four-fold. First, we explore a novel sentence modeling paradigm built on top of the notion of relevance, where the relationship between a candidate summary sentence and a spoken document to be summarized is discovered through different granularities of context for relevance modeling. Second, not only lexical but also topical cues inherent in the spoken document are exploited for sentence modeling. Third, we propose a novel clarity measure for use in important sentence selection, which can help quantify the thematic specificity of each individual sentence that is deemed to be a crucial indicator orthogonal to the relevance measure provided by the LM-based methods. Fourth, in an attempt to lessen summarization performance degradation caused by imperfect speech recognition, we investigate making use of different levels of index features for LM-based sentence modeling, including words, subword-level units, and their combination. Experiments on broadcast news summarization seem to demonstrate the performance merits of our methods when compared to several existing well-developed and/or state-of-the-art methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Leveraging Effective Query Modeling Techniques for Speech Recognition and SummarizationabstractStatistical language modeling (LM) that purports to quantify the acceptability of a given piece of text has long been an interesting yet challenging research area.In particular, language modeling for information retrieval (IR) has enjoyed remarkable empirical success; one emerging stream of the LM approach for IR is to employ the pseudo-relevance feedback process to enhance the representation of an input query so as to improve retrieval effectiveness.This paper presents a continuation of such a general line of research and the main contribution is threefold.First, we propose a principled framework which can unify the relationships among several widely-used query modeling formulations.Second, on top of the successfully developed framework, we propose an extended query modeling formulation by incorporating critical query-specific information cues to guide the model estimation.Third, we further adopt and formalize such a framework to the speech recognition and summarization tasks.A series of empirical experiments reveal the feasibility of such an LM framework and the performance merits of the deduced models on these two tasks. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Ea-Ee Jan, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
EMNLP | 5 |
| 2014 | I-vector based language modeling for spoken document retrievalabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. The i-vector based framework has been proposed and introduced to language identification (LID) and speaker recognition (SR) tasks recently. The major contribution of the i-vector framework is to reduce a series of acoustic feature vectors of a speech utterance to a low-dimensional vector representation, and then numbers of well-developed postprocessing techniques (such as probabilistic linear discriminative analysis, PLDA) can be readily and effectively used. However, to our best knowledge, there is no research up to date on applying the i-vector framework for SDR or information retrieval (IR). In this paper, we make a step forward to formulate an i-vector based language modeling (IVLM) framework for SDR. Furthermore, we evaluate the proposed IVLM framework with both inductive and transductive learning strategies. We also exploit multi-levels of index features, including word- and subword-level units, in concert with the proposed framework. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection demonstrate the performance merits of our proposed framework when compared to several existing approaches. Kuan-Yu Chen 0002, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 3 |
| 2014 | Speaker verification using kernel-based binary classifiers with binary operation derived featuresabstractIn this paper, we study the use of two kinds of kernel-based discriminative models, namely support vector machine (SVM) and deep neural network (DNN), for speaker verification. We treat the verification task as a binary classification problem, in which a pair of two utterances, each represented by an i-vector, is assumed to belong to either the “within-speaker” group or the “between-speaker” group. To solve the problem, we employ various binary operations to retain the basic relationship between any pair of i-vectors to form a single vector for training the discriminative models. This study also investigates the correlation of achievable performances with the number of training pairs and the various combinations of basic binary operations, using the SVM and DNN binary classifiers. The experiments are conducted on the male portion of the core task in the NIST 2005 Speaker Recognition Evaluation (SRE), and the results are competitive or even better, in terms of normalized decision cost function (minDCF) and equal error rate (EER), while compared to other non-probabilistic based models, such as the conventional speaker SVMs and the LDA-based cosine distance scoring. Hung-Shin Lee, Yu Tso, Yun-Fan Chang, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 4 |
| 2014 | Effective pseudo-relevance feedback for language modeling in extractive speech summarizationabstractExtractive speech summarization, aiming to automatically select an indicative set of sentences from a spoken document so as to concisely represent the most important aspects of the document, has become an active area for research and experimentation. An emerging stream of work is to employ the language modeling (LM) framework along with the Kullback-Leibler divergence measure for extractive speech summarization, which can perform important sentence selection in an unsupervised manner and has shown preliminary success. This paper presents a continuation of such a general line of research and its main contribution is two-fold. First, by virtue of pseudo-relevance feedback, we explore several effective sentence modeling formulations to enhance the sentence models involved in the LM-based summarization framework. Second, the utilities of our summarization methods and several widely-used methods are analyzed and compared extensively, which demonstrates the effectiveness of our methods. Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
ICASSP | 5 |
| 2014 | Improving music auto-tagging by intra-song instance baggingabstractBagging is one the most classic ensemble learning techniques in the machine learning literature. The idea is to generate multiple subsets of the training data via bootstrapping (random sampling with replacement), and then aggregate the output of the models trained from each subset via voting or averaging. As music is a temporal signal, we propose and study two bagging methods in this paper: the inter-song instance bagging that bootstraps song-level features, and the intra-song instance bagging that draws bootstrapping samples directly from short-time features for each training song. In particular, we focus on the latter method, as it better exploits the temporal information of music signals. The bagging methods result in surprisingly effective models for music auto-tagging: incorporating the idea to a simple linear support vector machine (SVM) based system yields accuracies that are comparable or even superior to state-of-the-art, possibly more sophisticated methods for three different datasets. As the bagging method is a meta algorithm, it holds the promise of improving other MIR systems. Chin-Chia Michael Yeh, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICASSP | 4 |
| 2014 | A recurrent neural network language modeling framework for extractive speech summarizationabstractExtractive speech summarization, with the purpose of automatically selecting a set of representative sentences from a spoken document so as to concisely express the most important theme of the document, has been an active area of research and development. A recent school of thought is to employ the language modeling (LM) approach for important sentence selection, which has proven to be effective for performing speech summarization in an unsupervised fashion. However, one of the major challenges facing the LM approach is how to formulate the sentence models and accurately estimate their parameters for each spoken document to be summarized. This paper presents a continuation of this general line of research and its contribution is two-fold. First, we propose a novel and effective recurrent neural network language modeling (RNNLM) framework for speech summarization, on top of which the deduced sentence models are able to render not only word usage cues but also long-span structural information of word co-occurrence relationships within spoken documents, getting around the need for the strict bag-of-words assumption. Second, the utilities of the method originated from our proposed framework and several widely-used unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a broadcast news summarization task seem to demonstrate the performance merits of our summarization method when compared to several state-of-the-art existing unsupervised methods. Kuan-Yu Chen 0002, Shih-Hung Liu, Berlin Chen, Hsin-Min Wang, Wen-Lian Hsu, Hsin-Hsi Chen |
ICME | 4 |
| 2014 | Towards time-varying music auto-tagging based on CAL500 expansionabstractMusic auto-tagging refers to automatically assigning semantic labels (tags) such as genre, mood and instrument to music so as to facilitate text-based music retrieval. Although significant progress has been made in recent years, relatively little research has focused on semantic labels that are time-varying within a track. Existing approaches and datasets usually assume that different fragments of a track share the same tag labels, disregarding the tags that are time-varying (e.g., mood) or local in time (e.g., instrument solo). In this paper, we present a new dataset dedicated to time-varying music auto-tagging. The dataset, called CAL500exp, is an enriched version of the well-known CAL500 dataset used for conventional track-level tagging. Given the tag set of CAL500, eleven subjects with strong music background were recruited to annotate the time-varying tag labels. A new user interface for annotation is developed to reduce the subject's annotation effort yet increase the quality of labels. Moreover, we present an empirical evaluation that demonstrates the performance improvement CAL500exp brings about for time-varying music auto-tagging. By providing more accurate and consistent descriptions of music content in a finer granularity, CAL500exp may open new opportunities to understand and to model the temporal context of musical semantics. Shuo-Yang Wang, Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang |
ICME | 4 |
| 2014 | Ensemble of machine learning algorithms for cognitive and physical speaker load detectionabstractWe present our methods and results on participating in the Interspeech 2014 Computational Paralinguistics ChallengE (ComParE) of which the goal is to detect certain type of load of a speaker using acoustic features. There are in total seven classification models contributing to our final prediction, namely, neural network with rectified linear unit and dropout (ReLUNet), conditional restricted Boltzmann machine (CRBM), logistic regression (LR), support vector machine (SVM), Gaussian discriminant analysis (GDA), k-nearest neighbors (KNN), and random forest (RF). When linearly blending the predictions of these models, we are able to get significant improvements over the challenge baseline. Index Terms: Physical Load Detection, Cognitive Load Detection, Neural Network, Classification Models How Jing, Ting-Yao Hu, Hung-Shin Lee, Wei-Chen Chen, Chi-Chun Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 7 |
| 2014 | Clustering-based i-vector formulation for speaker recognition
Hung-Shin Lee, Yu Tsao 0001, Hsin-Min Wang, Shyh-Kang Jeng |
INTERSPEECH | 3 |
| 2014 | Enhanced language modeling for extractive speech summarization with sentence relatedness informationabstractExtractive summarization is intended to automatically select a set of representative sentences from a text or spoken document that can concisely express the most important topics of the document. Language modeling (LM) has been proven to be a promising framework for performing extractive summarization in an unsupervised manner. However, there remain two fundamental challenges facing existing LM-based methods. One is how to construct sentence models involved in the LM framework more accurately without resorting to external information sources. The other is how to additionally take into account the sentence-level structural relationships embedded in a document for important sentence selection. To address these two challenges, in this paper we explore a novel approach that generates overlapped clusters to extract sentence relatedness information from the document to be summarized, which can be used not only to enhance the estimation of various sentence models but also to allow for the sentence-level structural relationships for better summarization performance. Further, the utilities of our proposed methods and several state-of-the-art unsupervised methods are analyzed and compared extensively. A series of experiments conducted on a Mandarin broadcast news summarization task demonstrate the effectiveness and viability of our method. Index Terms: speech summarization, language modeling, clustering, relevance, sentence relatedness Shih-Hung Liu, Kuan-Yu Chen 0002, Yu-Lun Hsieh, Berlin Chen, Hsin-Min Wang, Hsu-Chun Yen, Wen-Lian Hsu |
INTERSPEECH | 5 |
| 2014 | Generalized k-Labelsets Ensemble for Multi-Label and Cost-Sensitive ClassificationabstractLabel powerset (LP) method is one category of multi-label learning algorithm. This paper presents a basis expansions model for multi-label classification, where a basis function is an LP classifier trained on a random k-labelset. The expansion coefficients are learned to minimize the global error between the prediction and the ground truth. We derive an analytic solution to learn the coefficients efficiently. We further extend this model to handle the cost-sensitive multi-label classification problem, and apply it in social tagging to handle the issue of the noisy training set by treating the tag counts as the misclassification costs. We have conducted experiments on several benchmark datasets and compared our method with other state-of-the-art multi-label learning methods. Experimental results on both multi-label classification and cost-sensitive social tagging demonstrate that our method has better performance than other methods. Hung-Yi Lo, Shou-De Lin, Hsin-Min Wang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Effective pseudo-relevance feedback for spoken document retrievalabstractWith the exponential proliferation of multimedia associated with spoken documents, research on spoken document retrieval (SDR) has emerged and attracted much attention in the past two decades. Apart from much effort devoted to developing robust indexing and modeling techniques for representing spoken documents, a recent line of thought targets at the improvement of query modeling for better reflecting the user's information need. Pseudo-relevance feedback is by far the most commonly-used paradigm for query reformulation, which assumes that a small amount of top-ranked feedback documents obtained from the initial round of retrieval are relevant and can be utilized for this purpose. Nevertheless, simply taking all of the top-ranked feedback documents obtained from the initial retrieval for query modeling (reformulation) does not always work well, especially when the top-ranked documents contain much redundant or non-relevant information. In the view of this, we explore in this paper an interesting problem of how to effectively glean useful cues from the top-ranked documents so as to achieve more accurate query modeling. To do this, different kinds of information cues are considered and integrated into the process of feedback document selection so as to improve query effectiveness. Experiments conducted on the TDT (Topic Detection and Tracking) task show the advantages of our retrieval methods for SDR. Yi-Wen Chen, Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen |
ICASSP | 3 |
| 2013 | Weighted matrix factorization for spoken document retrievalabstractSince more and more multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research subject in the past two decades. Recently, topic models have been successfully used in SDR as well as general information retrieval (IR). These models fall into two categories: probabilistic topic models (PTM) and non-probabilistic topic models (NPTM). One major difference between PTM and NPTM is that the former only takes the words occurring in a document into account, whereas the latter, such as latent semantic analysis (LSA), explicitly models all the words in the vocabulary (including both occurring and non-occurring words). We believe that the non-occurring words can provide additional information that is also useful for SDR. However, to our best knowledge, there is a dearth of work investigating the effectiveness of the non-occurring words for SDR and IR. In order to make effective use of those non-occurring words of documents for semantic analysis, we propose a weighted matrix factorization (WMF) framework, in which the impact of the non-occurring words on the semantic analysis can be modulated properly. The results of SDR experiments conducted on the TDT-2 (Topic Detection and Tracking) collection highlight the performance merits of our proposed framework when compared to several existing topic models. Kuan-Yu Chen 0002, Hsin-Min Wang, Berlin Chen, Hsin-Hsi Chen |
ICASSP | 2 |
| 2013 | Subspace-based phonotactic language recognition using multivariate dynamic linear modelsabstractPhonotactics, dealing with permissible phone patterns and their frequencies of occurrence in a specific language, is acknowledged to be related to spoken language recognition (SLR). With the assistance of phone recognizers, each speech utterance can be decoded into an ordered sequence of phone vectors filled with likelihood scores contributed by all possible phone models. In this paper, we propose a novel approach to dig the concealed phonotactic structure out of the phone-likelihood vectors through a kind of multivariate time series analysis: dynamic linear models (DLM). In these models, treating the generation of phone patterns in each utterance as a dynamic system, the relationship between adjacent vectors is linearly and time-invariantly modeled, and unobserved states are introduced to capture a temporal coherence intrinsic in the system. Each utterance expressed by the DLM is further transformed into a fixed-dimensional linear subspace so that well-developed distance measures between two subspaces can be applied to linear discriminant analysis (LDA) in a dissimilarity-based fashion. The results of SLR experiments on the OGI-TS corpus demonstrate that the proposed framework outperforms the well-known vector space modeling (VSM)-based methods and achieves comparable performance to our previous subspace-based method. Hung-Shin Lee, Yu-Chin Shih, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 3 |
| 2013 | Semantic Naïve Bayes Classifier for Document Classification
How Jing, Yu Tsao 0001, Kuan-Yu Chen 0002, Hsin-Min Wang |
IJCNLP | 4 |
| 2013 | Alleviating the over-smoothing problem in GMM-based voice conversion with discriminative trainingabstractIn this paper, we propose a discriminative training (DT) method to alleviate the muffled sound effect caused by over smoothing in the Gaussian mixture model (GMM)-based voice conversion (VC). For the conventional GMM-based VC, we often observed a large degree of ambiguities among acoustic classes (generative classes), determined by the source feature vectors for generating the converted feature vectors, causing the “muffled sound” effect on the converted voice. The proposed DT method is applied to refine the parameters in the maximum likelihood (ML)-trained joint density GMM (JDGMM) in the training stage to reduce the ambiguities among acoustic classes (generative classes) to alleviate the muffled sound effect. Experimental results demonstrate that the DT method significantly enhances the discriminative power between acoustic classes (generative classes) in the objective evaluation and effectively alleviates the muffled sound effect in the subjective evaluation. Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Yih-Ru Wang, Sin-Horng Chen |
INTERSPEECH | 3 |
| 2013 | Non-reference audio quality assessment for online live music recordingsabstractImmensely popular video sharing websites such as YouTube have become the most important sources of music information for Internet users and the most prominent platform for sharing live music. The audio quality of this huge amount of live music recordings, however, varies significantly due to factors such as environmental noise, location, and recording device. However, most video search engines do not take audio quality into consideration when retrieving and ranking results. Given the fact that most users prefer live music videos with better audio quality, we propose the first automatic, non-reference audio quality assessment framework for live music video search online. We first construct two annotated datasets of live music recordings. The first dataset contains 500 human-annotated pieces, and the second contains 2,400 synthetic pieces systematically generated by adding noise effects to clean recordings. Then, we formulate the assessment task as a ranking problem and try to solve it using a learning-based scheme. To validate the effectiveness of our framework, we perform both objective and subjective evaluations. Results show that our framework significantly improves the ranking performance of live music recording retrieval and can prove useful for various real-world music applications. Ju-Chiang Wang, Jingli Cai, Zhiyan Duan, Hsin-Min Wang, Ye Wang 0007 |
ACM Multimedia | 5 |
| 2013 | Query-Document Relevance Topic Models
Meng-Sung Wu, Chia-Ping Chen, Hsin-Min Wang |
PAKDD (2) | 3 |
| 2012 | Generalized k-labelset ensemble for multi-label classificationabstractLabel powerset (LP) method is one category of multi-label learning algorithms. It reduces the multi-label classification problem to a multi-class classification problem by treating each distinct combination of labels in the training set as a different class. This paper proposes a basis expansion model for multi-label classification, where a basis function is a LP classifier trained on a random k-labelset. The expansion coefficients are learned to minimize the global error between the prediction and the multi-label ground truth. We derive an analytic solution to learn the coefficients efficiently. We have conducted experiments using several benchmark datasets and compared our method with other state-of-the-art multi-label learning methods. The results show that our method has better or competitive performance against other methods. Hung-Yi Lo, Shou-De Lin, Hsin-Min Wang |
ICASSP | 3 |
| 2012 | Playing with tagging: A real-time tagging music playerabstractVisualizing audio signals during playback has long been a fundamental function of music players. However, most visual effects are generated by audio signal processing directly and render meaningless or incomprehensible displays to users. In this paper, we present an intelligent music player called the Playing with Tagging (PWT) music player. By integrating a real-time music tagger, the PWT player can display dynamic tag distributions via a set of tag bars that move in sync with the music. To synchronize the tag distributions, the music tagger must be able to online recognize the music tags. We utilize a Gaussian mixture model (GMM) as an auditory feature encoding reference and a mixture of tag-based aspect models (TBAMs) to predict the tag distribution for a short sliding chunk of the music played. To evaluate the real-time tagging function, we simulate tag prediction on short music chunks. The results of experiments on the MajorMiner dataset demonstrate the potential and effectiveness of the proposed music tagging method. Ju-Chiang Wang, Hsin-Min Wang, Shyh-Kang Jeng |
ICASSP | 2 |
| 2012 | Term relevance dependency model for text classification
Meng-Sung Wu, Hsin-Min Wang |
ICPR | 2 |
| 2012 | Word Relevance Modeling for Speech RecognitionabstractLanguage models for speech recognition tend to be brittle across domains, since their performance is vulnerable to changes in the genre or topic of the text on which they are trained. A number of adaptation methods, discovering either lexical co-occurrence or topic cues, have been developed to mitigate this problem with varying degrees of success. Among them, a more recent thread of work is the relevance modeling approach, which has shown promise to capture the lexical co-occurrence relationship between the entire search history and an upcoming word. However, a potential downside to such an approach is the need of resorting to a retrieval procedure to obtain relevance information; this is usually complex and time-consuming for practical applications. In this paper, we propose a word relevance modeling framework, which introduces a novel use of relevance information for dynamic language model adaptation in speech recognition. It not only inherits the merits of several existing techniques but also provides a flexible yet systematic way to render the lexical, topical, and proximity relationships between the search history and the upcoming word. Experiments on large vocabulary continuous speech recognition demonstrate the performance merits of the methods instantiated from this framework when compared to several existing methods. Kuan-Yu Chen 0002, Hao-Chin Chang, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2012 | A Study of Mutual Information for GMM-Based Spectral ConversionabstractThe Gaussian mixture model (GMM)-based method has dominated the field of voice conversion (VC) for last decade. However, the converted spectra are excessively smoothed and thus produce muffled converted sound. In this study, we improve the speech quality by enhancing the dependency between the source (natural sound) and converted feature vectors (converted sound). It is believed that enhancing this dependency can make the converted sound closer to the natural sound. To this end, we propose an integrated maximum a posteriori and mutual information (MAPMI) criterion for parameter generation on spectral conversion. Experimental results demonstrate that the quality of converted speech by the proposed MAPMI method outperforms that by the conventional method in terms of formal listening test. Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Yih-Ru Wang, Sin-Horng Chen |
INTERSPEECH | 3 |
| 2012 | Subspace-Based Feature Representation and Learning for Language RecognitionabstractThis paper presents a novel subspace-based approach for phonotactic language recognition. The whole framework is divided into two parts: the speech feature representation and the subspacebased learning algorithm. First, the phonetic information as well as the contextual relationship, possessed by spoken utterances, are more abundantly retrieved by likelihood computation and feature concatenation through the decoding processed by an automatic speech recognizer. It is assumed that the extracted phone frames reside in a lower dimensional eigen-subspace, in which the structure of data can be approximately captured. Each utterance is further represented by a fixed-dimensional linear subspace. Second, to measure the similarity between two utterances, suitable non-Euclidean metrics are explored and applied to non-linear discriminant analysis in a kernel fashion, followed by a back-end classifier, such as the k-nearest neighbor (K-NN) classifier. The results of experiments on the OGI-TS database demonstrate that the proposed framework outperforms the well-known vector space modeling based method with relative reductions of 38.90% and 27.13% on the 1-to-50-second and 3-second data sets respectively in equal error rate (EER). Yu-Chin Shih, Hung-Shin Lee, Hsin-Min Wang, Shyh-Kang Jeng |
INTERSPEECH | 3 |
| 2012 | The acousticvisual emotion guassians model for automatic generation of music videoabstractThis paper presents a novel content-based system that utilizes the perceived emotion of multimedia content as a bridge to connect music and video. Specifically, we propose a novel machine learning framework, called Acousticvisual Emotion Guassians (AVEG), to jointly learn the tripartite relationship among music, video, and emotion from an emotion-annotated corpus of music videos. For a music piece (or a video sequence), the AVEG model is applied to predict its emotion distribution in a stochastic emotion space from the corresponding low-level acoustic (resp. visual) features. Finally, music and video are matched by measuring the similarity between the two corresponding emotion distributions, based on a distance measure such as KL divergence. Ju-Chiang Wang, Yi-Hsuan Yang, I-Hong Jhuo, Yen-Yu Lin, Hsin-Min Wang |
ACM Multimedia | 5 |
| 2012 | The acoustic emotion gaussians model for emotion-based music annotation and retrievalabstractOne of the most exciting but challenging endeavors in music research is to develop a computational model that comprehends the affective content of music signals and organizes a music collection according to emotion. In this paper, we propose a novel acoustic emotion Gaussians (AEG) model that defines a proper generative process of emotion perception in music. As a generative model, AEG permits easy and straightforward interpretations of the model learning processes. To bridge the acoustic feature space and music emotion space, a set of latent feature classes, which are learned from data, is introduced to perform the end-to-end semantic mappings between the two spaces. Based on the space of latent feature classes, the AEG model is applicable to both automatic music emotion annotation and emotion-based music retrieval. To gain insights into the AEG model, we also provide illustrations of the model learning process. A comprehensive performance study is conducted to demonstrate the superior accuracy of AEG over its predecessors, using two emotion annotated music corpora MER60 and MTurk. Our results show that the AEG model outperforms the state-of-the-art methods in automatic music emotion annotation. Moreover, for the first time a quantitative evaluation of emotion-based music retrieval is reported. Ju-Chiang Wang, Yi-Hsuan Yang, Hsin-Min Wang, Shyh-Kang Jeng |
ACM Multimedia | 3 |
| 2012 | A Term Association Translation Model for Naive Bayes Text Classification
Meng-Sung Wu, Hsin-Min Wang |
PAKDD (1) | 2 |
| 2011 | Cost-sensitive stacking for audio tag annotation and retrievalabstractAudio tags correspond to keywords that people use to de scribe different aspects of a music clip, such as the genre, mood, and instrumentation. Since social tags are usually as signed by people with different levels of musical knowledge, they inevitably contain noisy information. By treating the tag counts as costs, we can model the audio tagging problem as a cost-sensitive classification problem. In addition, tag correlation is another useful information for automatic audio tagging since some tags often co-occur. By considering the co-occurrences of tags, we can model the audio tagging problem as a multi-label classification problem. To exploit the tag count and correlation information jointly, we formulate the audio tagging task as a novel cost-sensitive multi-label (CSML) learning problem. The results of audio tag annotation and retrieval experiments demonstrate that the new approach outperforms our MIREX 2009 winning method. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang, Shou-De Lin |
ICASSP | 3 |
| 2011 | Automatic annotation of Web videosabstractMost Web videos are captured in uncontrolled environments (e.g. videos captured by freely-moving cameras with low resolution); this makes automatic video annotation very difficult. To address this problem, we present a robust moving foreground object detection method followed by the integration of features collected from heterogeneous domains. We advance SIFT feature matching and present a probabilistic framework to construct consensus foreground object templates (CFOT). The CFOT can detect moving foreground objects of interest across video frames, and this allows us to extract visual features from foreground regions of interest. Together with the use of audio features, we are able to improve resulting annotation accuracy. We conduct experiments and achieve promising results on a Web video dataset collected from YouTube. Shih-Wei Sun, Yu-Chiang Frank Wang, Yao-Ling Hung, Chia-Ling Chang, Kuan-Chieh Chen, Shih-Sian Cheng, Hsin-Min Wang, Hong-Yuan Mark Liao |
ICME | 7 |
| 2011 | Query by multi-tags with multi-level preferences for content-based music retrievalabstractThis paper presents a novel content-based music retrieval system that accepts a query containing multiple tags with multiple levels of preference (denoted as an MTML query) to retrieve music from an untagged music database. We select a limited number of popular music tags to form the tag space and design an interface for users to input queries by operating the scroll bars. To effect MTML content-based music retrieval, we introduce a tag-based music aspect model that jointly models the auditory features and tag-based text features of a song. Two indexing methods and their corresponding matching methods, namely pseudo song-based matching and tag co-occurrence pattern-based matching, are incorporated into the pre-learned tag-based music aspect model. Finally, we evaluate the proposed system on the Major Miner dataset. The results demonstrate the potential of using MTML queries to retrieve music from an untagged music database. Ju-Chiang Wang, Meng-Sung Wu, Hsin-Min Wang, Shyh-Kang Jeng |
ICME | 3 |
| 2011 | Colorizing tags in tag cloud: a novel query-by-tag music search systemabstractThis paper presents a novel content-based query-by-tag music search system for an untagged music database. We design a new tag query interface that allows users to input multiple tags with multiple levels of preference (denoted as an MTML query) by colorizing desired tags in a web-based tag cloud interface. When a user clicks and holds the left mouse button (or presses and holds his/her finger on a touch screen) on a desired tag, the color of the tag will change cyclically according to a color map (from dark blue to bright red), which represents the level of preference (from 0 to 1). In this way, the user can easily organize and check the query of multiple tags with multiple levels of preference through the colored tags. To effect the MTML content-based music retrieval, we introduce a probabilistic fusion model (denoted as GMFM), which consists of two mixture models, namely a Gaussian mixture model and a multinomial mixture model. GMFM can jointly model the auditory features and tag labels of a song. Two indexing methods and their corresponding matching methods, namely pseudo song-based matching and tag affinity-based matching, are incorporated into the pre-learned GMFM. We evaluate the proposed system on the MajorMiner and CAL-500 datasets. The experimental results demonstrate the effectiveness of GMFM and the potential of using MTML queries to search music from an untagged music database. Ju-Chiang Wang, Yu-Chin Shih, Meng-Sung Wu, Hsin-Min Wang, Shyh-Kang Jeng |
ACM Multimedia | 4 |
| 2011 | Audio Tag Annotation and Retrieval Using Tag Count Information
Hung-Yi Lo, Shou-De Lin, Hsin-Min Wang |
MMM (1) | 3 |
| 2011 | Cost-Sensitive Multi-Label Learning for Audio Tag Annotation and RetrievalabstractAudio tags correspond to keywords that people use to describe different aspects of a music clip. With the explosive growth of digital music available on the Web, automatic audio tagging, which can be used to annotate unknown music or retrieve desirable music, is becoming increasingly important. This can be achieved by training a binary classifier for each tag based on the labeled music data. Our method that won the MIREX 2009 audio tagging competition is one of this kind of methods. However, since social tags are usually assigned by people with different levels of musical knowledge, they inevitably contain noisy information. By treating the tag counts as costs, we can model the audio tagging problem as a cost-sensitive classification problem. In addition, tag correlation information is useful for automatic audio tagging since some tags often co-occur. By considering the co-occurrences of tags, we can model the audio tagging problem as a multi-label classification problem. To exploit the tag count and correlation information jointly, we formulate the audio tagging task as a novel cost-sensitive multi-label (CSML) learning problem and propose two solutions to solve it. The experimental results demonstrate that the new approach outperforms our MIREX 2009 winning method. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang, Shou-De Lin |
IEEE Trans. Multim. | 3 |
| 2010 | Background music identification through content filtering and min-hash matchingabstractA novel framework for background music identification is proposed in this paper. Given a piece of audio signals that mixes background music with speech/noise, we identify the music part with source music data. Conventional methods that take the whole audio signals for identification are inappropriate in terms of efficiency and accuracy. In our framework, the audio content is filtered through speech center cancellation and noise removal to extract clear music segments. To identify these music segments, we use a compact feature representation and efficient similarity measurement based on the min-hash theory. The results of experiments on the RWC music database show a promising direction. Chih-Yi Chiu, Dimitrios Bountouridis, Ju-Chiang Wang, Hsin-Min Wang |
ICASSP | 4 |
| 2010 | Detecting pitching frames in baseball game video using Markov random walkabstractPitching is the starting point of an event in baseball games. Hence, locating pitching shots is a critical step in content analysis of a baseball game video. However, pitching frames vary with innings and games. Existing methods that require a great deal of effort to construct empirical rules or label training data do not capture the characteristics of various pitching frames very well. In this paper, we present an unsupervised method for pitching frame detection by using Markov random walk. A video stream is first divided into content-homogeneous shots, and these shots are merged into states through hierarchical agglomerative clustering. Then, the state with the highest visit probability according to the Markov random walk theory is deemed the set of pitching frames. Finally, a model trained on the pitching frames in the pitching state is further used to detect the remaining potential pitching frames in other states. Our experiments demonstrate that the proposed method yields satisfactory results in a variety of MLB games. Chih-Yi Chiu, Po-Chih Lin, Wei-Ming Chang, Hsin-Min Wang, Shi-Nine Yang |
ICIP | 4 |
| 2010 | Homogeneous segmentation and classifier ensemble for audio tag annotation and retrievalabstractAudio tags describe different types of musical information such as genre, mood, and instrument. This paper aims to automatically annotate audio clips with tags and retrieve relevant clips from a music database by tags. Given an audio clip, we divide it into several homogeneous segments by using an audio novelty curve, and then extract audio features from each segment with respect to various musical information, such as dynamics, rhythm, timbre, pitch, and tonality. The features in frame-based feature vector sequence format are further represented by their mean and standard deviation such that they can be combined with other segment-based features to form a fixed-dimensional feature vector for a segment. We train an ensemble classifier, which consists of SVM and AdaBoost classifiers, for each tag. For the audio annotation task, the individual classifier outputs are transformed into calibrated probability scores such that probability ensemble can be employed. For the audio retrieval task, we propose using ranking ensemble. We participated in the MIREX 2009 audio tag classification task and our system was ranked first in terms of F-measure and the area under the ROC curve given a tag. Hung-Yi Lo, Ju-Chiang Wang, Hsin-Min Wang |
ICME | 3 |
| 2010 | A Discriminative and Heteroscedastic Linear Feature Transformation for Multiclass ClassificationabstractThis paper presents a novel discriminative feature transformation, named full-rank generalized likelihood ratio discriminant analysis (fGLRDA), on the grounds of the likelihood ratio test (LRT). fGLRDA attempts to seek a feature space, which is linearly isomorphic to the original n-dimensional feature space and is characterized by a full-rank (n×n) transformation matrix, under the assumption that all the class-discrimination information resides in a d-dimensional subspace (d <; n), through making the most confusing situation, described by the null hypothesis, as unlikely as possible to happen without the homoscedastic assumption on class distributions. Our experimental results demonstrate that fGLRDA can yield moderate performance improvements over other existing methods, such as linear discriminant analysis (LDA) for the speaker identification task. Hung-Shin Lee, Hsin-Min Wang, Berlin Chen |
ICPR | 2 |
| 2010 | Phonetic subspace mixture model for speaker diarizationabstractThis paper presents an improved distance measure for speaker clustering in speaker diarization systems. The proposed phonetic subspace mixture (PSM) model introduces phonetic information to the ΔBIC distance measure. Therefore, the new PSM model-based ΔBIC distance measure can remove the effect of phonetic content on the diarization results. The typical ΔBIC distance measure can be seen as a special case of the new ΔBIC distance measure. Our experiment results show that the new distance measurement consistently improves the speaker diarization performance on three datasets. Index Terms: BIC, phonetic information, speaker diarization 1. I-Fan Chen, Shih-Sian Cheng, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2010 | Bayesian speaker recognition using Gaussian mixture model and laplace approximationabstractThis paper presents a Bayesian approach for Gaussian mix-ture model (GMM)-based speaker identification. Some ap-proaches evaluate the speaker score of a test speech utterance using a single data likelihood over the GMM learned by point estimation methods according to the maximum likelihood or maximum a posteriori criteria. In contrast, the Bayesian ap-proach evaluates the score by using the expectation of the data likelihood over the posterior distribution of the model parame-ters, which is depicted by Bayesian integration. However, as the integration can not be derived analytically, we apply Laplace approximation to the derivations. Theoretically, we show that the proposed Bayesian approach is equivalent to the GMM-UBM approach when infinite training data is available for each speaker. The results of speaker identification experiments on the TIMIT corpus show that the proposed Bayesian approach consistently outperforms the GMM-UBM approach under very limited training data conditions, although the improvement is not very significant. Index Terms: speaker identification, speaker recognition, Bayesian inference, GMM-UBM Shih-Sian Cheng, I-Fan Chen, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2010 | Exploiting semantic associative information in topic modelingabstractTopic modeling has been widely applied in a variety of text modeling tasks as well as in speech recognition systems for effectively capturing the semantic and statistic information in documents or speech utterances. Most topic models rely on the bag-of-words assumption that results in learned latent topics composed of lists of individual words. Unfortunately, these words may convey topical information but lack accurate semantic knowledge of the text. In this paper, we present the semantic associative topic model, where the concept of the semantic association terms is extended to topic modeling, which provides guidance on modeling the semantic associations that occur among single words by expressing a document as an association of multiple words. Further, the pointwise KL-divergence metric is used to measure the significance of the association. We also integrate original PLSA and SATM models, which have mixed feature representations. Experimental results on WSJ and AP datasets show that the proposed approaches achieved higher performance compared to other methods. Meng-Sung Wu, Hung-Shin Lee, Hsin-Min Wang |
SLT | 3 |
| 2010 | BIC-Based Speaker Segmentation Using Divide-and-Conquer Strategies With Application to Speaker DiarizationabstractIn this paper, we propose three divide-and-conquer approaches for Bayesian information criterion (BlC)-based speaker segmentation. The approaches detect speaker changes by recursively partitioning a large analysis window into two sub-windows and recursively verifying the merging of two adjacent audio segments using DeltaBIC, a widely-adopted distance measure of two audio segments. We compare our approaches to three popular distance-based approaches, namely, Chen and Gopalakrishnan's window-growing-based approach, Siegler et al.'s fixed-size sliding window approach, and Delacourt and Wellekens's DISTBIC approach, by performing computational cost analysis and conducting speaker change detection experiments on two broadcast news data sets. The results show that the proposed approaches are more efficient and achieve higher segmentation accuracy than the compared distance-based approaches. In addition, we apply the segmentation approaches discussed in this paper to the speaker diarization task. The experiment results show that a more effective segmentation approach leads to better diarization accuracy. Shih-Sian Cheng, Hsin-Min Wang, Hsin-Chia Fu |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Time-Series Linear Search for Video Copies Based on Compact Signature Manipulation and Containment Relation ModelingabstractThis paper presents a novel time-series linear search (TLS) method for detecting video copies. The method utilizes a sliding window to locate window sequences that are near-duplicates of a given query sequence. We address two issues of the conventional TLS method in order to strengthen its video copy detection capability. First, to accelerate the TLS process, we use a sequence-level signature as a compact representation of a video sequence based on the min-hash theory, and develop an efficient heap manipulation technique for fast generation of each window sequence's signature. Second, to improve the robustness of the TLS method, we use two techniques, namely, window length estimation and threshold transform, to resolve the containment relation problem caused by various types of video transformation and editing, such as frame cropping and speed change. The results of experiments on the MUSCLE-VCD-2007 dataset demonstrate that the proposed method is efficient and robust against different types of video transformation and editing. Chih-Yi Chiu, Hsin-Min Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Fast min-hashing indexing and robust spatio-temporal matching for detecting video copiesabstractThe increase in the number of video copies, both legal and illegal, has become a major problem in the multimedia and Internet era. In this article, we propose a novel method for detecting various video copies in a video sequence. To achieve fast and robust detection, the method fully integrates several components, namely the min-hashing signature to compactly represent a video sequence, a spatio-temporal matching scheme to accurately evaluate video similarity compiled from the spatial and temporal aspects, and some speedup techniques to expedite both min-hashing indexing and spatio-temporal matching. The results of experiments demonstrate that, compared to several baseline methods with different feature descriptors and matching schemes, the proposed method which combines both global and local feature descriptors yields the best performance when encountering a variety of video transformations. The method is very fast, requiring approximately 0.06 seconds to search for copies of a thirty-second video clip in a six-hour video sequence. Chih-Yi Chiu, Hsin-Min Wang, Chu-Song Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2009 | Learning to rank from Bayesian decision inferenceabstractRanking is a key problem in many information retrieval (IR) applications, such as document retrieval and collaborative filtering. In this paper, we address the issue of learning to rank in document retrieval. Learning-based methods, such as RankNet, RankSVM, and RankBoost, try to create ranking functions automatically by using some training data. Recently, several learning to rank methods have been proposed to directly optimize the performance of IR applications in terms of various evaluation measures. They undoubtedly provide statistically significant improvements over conventional methods; however, from the viewpoint of decision-making, most of them do not minimize the Bayes risk of the IR system. In an attempt to fill this research gap, we propose a novel framework that directly optimizes the Bayes risk related to the ranking accuracy in terms of the IR evaluation measures. The results of experiments on the LETOR collections demonstrate that the framework outperforms several existing methods in most cases. Jen-Wei Kuo, Pu-Jen Cheng, Hsin-Min Wang |
CIKM | 3 |
| 2009 | Articulatory feature asynchrony analysis and compensation in detection-based ASRabstractThis paper investigates the effects of two types of imperfection, namely detection errors and articulatory feature asynchrony, of the front-end articulatory feature detector on the performance of a detection-based ASR system. Based on a set of variable-controlled experiments, we find that articulatory feature asynchrony is the major issue that should be addressed in detection-based ASR. To this end, we propose several methods to reduce the asynchrony or the effects of asynchrony. The results are quite promising; for example, currently, we can achieve 67.67 % phone accuracy in the TIMIT free phone recognition task with only 11 binary-valued articulatory features. Index Terms: articulatory feature asynchrony, detection-based ASR, speech recognition I-Fan Chen, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2009 | Speaker diarization using divide-and-conquerabstractSpeaker diarization systems usually consist of two core components: speaker segmentation and speaker clustering. The current state-of-the-art speaker diarization systems usually apply hierarchical agglomerative clustering (HAC) for speaker clustering after segmentation. However, HAC’s quadratic computational complexity with respect to the number of data samples inevitably limits its application in large-scale data sets. In this paper, we propose a divide-and-conquer (DAC) framework for speaker diarization. It recursively partitions the input speech stream into two sub-streams, performs diarization on them separately, and then combines the diarization results obtained from them using HAC. The results of experiments conducted on RT-02 and RT-03 broadcast news data show that the proposed framework is faster than the conventional segmentation and clustering-based approach while achieving comparable diarization accuracy. Moreover, the proposed framework obtains a higher speedup over the conventional approach on a larger test data set. Index Terms: speaker diarization, speaker segmentation, speaker clustering, divide-and-conquer Shih-Sian Cheng, Chun-Han Tseng, Chia-Ping Chen, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2009 | Improving GMM-UBM speaker verification using discriminative feedback adaptation
Yi-Hsiang Chao, Wei-Ho Tsai, Hsin-Min Wang |
Comput. Speech Lang. | 3 |
| 2009 | Evolutionary minimization of the Rand index for speaker clustering
Wei-Ho Tsai, Hsin-Min Wang |
Comput. Speech Lang. | 2 |
| 2009 | Improving the characterization of the alternative hypothesis via minimum verification error training with applications to speaker verification
Yi-Hsiang Chao, Wei-Ho Tsai, Hsin-Min Wang, Ruei-Chuan Chang |
Pattern Recognit. | 3 |
| 2009 | A Comparative Study of Probabilistic Ranking Models for Chinese Spoken Document SummarizationabstractExtractive document summarization automatically selects a number of indicative sentences, passages, or paragraphs from an original document according to a target summarization ratio, and sequences them to form a concise summary. In this article, we present a comparative study of various probabilistic ranking models for spoken document summarization, including supervised classification-based summarizers and unsupervised probabilistic generative summarizers. We also investigate the use of unsupervised summarizers to improve the performance of supervised summarizers when manual labels are not available for training the latter. A novel training data selection approach that leverages the relevance information of spoken sentences to select reliable document-summary pairs derived by the probabilistic generative summarizers is explored for training the classification-based summarizers. Encouraging initial results on Mandarin Chinese broadcast news data are demonstrated. Shih-Hsiang Lin, Berlin Chen, Hsin-Min Wang |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2009 | A Probabilistic Generative Framework for Extractive Broadcast News Speech SummarizationabstractIn this paper, we consider extractive summarization of broadcast news speech and propose a unified probabilistic generative framework that combines the sentence generative probability and the sentence prior probability for sentence ranking. Each sentence of a spoken document to be summarized is treated as a probabilistic generative model for predicting the document. Two matching strategies, namely literal term matching and concept matching, are thoroughly investigated. We explore the use of the language model (LM) and the relevance model (RM) for literal term matching, while the sentence topical mixture model (STMM) and the word topical mixture model (WTMM) are used for concept matching. In addition, the lexical and prosodic features, as well as the relevance information of spoken sentences, are properly incorporated for the estimation of the sentence prior probability. An elegant feature of our proposed framework is that both the sentence generative probability and the sentence prior probability can be estimated in an unsupervised manner, without the need for handcrafted document-summary pairs. The experiments were performed on Chinese broadcast news collected in Taiwan, and very encouraging results were obtained. Berlin Chen, Hsin-Min Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Model-Based Clustering by Probabilistic Self-Organizing MapsabstractIn this paper, we consider the learning process of a probabilistic self-organizing map (PbSOM) as a model-based data clustering procedure that preserves the topological relationships between data clusters in a neural network. Based on this concept, we develop a coupling-likelihood mixture model for the PbSOM that extends the reference vectors in Kohonen's self-organizing map (SOM) to multivariate Gaussian distributions. We also derive three expectation-maximization (EM)-type algorithms, called the SOCEM, SOEM, and SODAEM algorithms, for learning the model (PbSOM) based on the maximum-likelihood criterion. SOCEM is derived by using the classification EM (CEM) algorithm to maximize the classification likelihood; SOEM is derived by using the EM algorithm to maximize the mixture likelihood; and SODAEM is a deterministic annealing (DA) variant of SOCEM and SOEM. Moreover, by shrinking the neighborhood size, SOCEM and SOEM can be interpreted, respectively, as DA variants of the CEM and EM algorithms for Gaussian model-based clustering. The experimental results show that the proposed PbSOM learning algorithms achieve comparable data clustering performance to that of the deterministic annealing EM (DAEM) approach, while maintaining the topology-preserving property. Shih-Sian Cheng, Hsin-Chia Fu, Hsin-Min Wang |
IEEE Trans. Neural Networks | 3 |
| 2008 | BIC-based audio segmentation by divide-and-conquerabstractAudio segmentation has received increasing attention in recent years for its potential applications in automatic indexing and transcription of audio data. Among existing audio segmentation approaches, the BIC-based approach proposed by Chen and Gopalakrishnan is most well-known for its high accuracy. However, this window-growing-based segmentation approach suffers from the high computation cost. In this paper, we propose using the efficient divide-and-conquer strategy in audio segmentation. Our approaches detect acoustic changes by recursively partitioning an analysis window into two sub-windows using ΔBIC. The results of experiments conducted on the broadcast news data demonstrate that our approaches not only have a lower computation cost but also achieve a higher segmentation accuracy than window-growing-based segmentation. Shih-Sian Cheng, Hsin-Min Wang, Hsin-Chia Fu |
ICASSP | 2 |
| 2008 | A comparative study of probabilistic ranking models for spoken document summarizationabstractThe purpose of extractive document summarization is to automatically select a number of indicative sentences, passages, or paragraphs from the original document according to a target summarization ratio and then sequence them to form a concise summary. In the paper, we present a comparative study of various supervised and unsupervised probabilistic ranking models for spoken document summarization on the Chinese broadcast news. Moreover, we also investigate the possibility of using unsupervised summarizers to boost the performance of supervised summarizers when manual labels are not available for the training of supervised summarizers. Encouraging results were initially demonstrated. Shih-Hsiang Lin, Hsin-Min Wang |
ICASSP | 3 |
| 2008 | Using Kernel Discriminant Analysis to Improve the Characterization of the Alternative Hypothesis for Speaker VerificationabstractSpeaker verification can be viewed as a task of modeling and testing two hypotheses: the null hypothesis and the alternative hypothesis. Since the alternative hypothesis involves unknown impostors, it is usually hard to characterizeapriori. In this paper, we propose improving the characterization of the alternative hypothesis by designing two decision functions based, respectively, on a weighted arithmetic combination and a weighted geometric combination of discriminative information derived from a set of pretrained background models. The parameters associated with the combinations are then optimized using two kernel discriminant analysis techniques, namely, the kernel Fisher discriminant (KFD) and support vector machine (SVM). The proposed approaches have two advantages over existing methods. The first is that they embed a trainable mechanism in the decision functions. The second is that they convert variable-length utterances into fixed-dimension characteristic vectors, which are easily processed by kernel discriminant analysis. The results of speaker-verification experiments conducted on two speech corpora show that the proposed methods outperform conventional likelihood ratio-based approaches. Yi-Hsiang Chao, Wei-Ho Tsai, Hsin-Min Wang, Ruei-Chuan Chang |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | A Query-by-Singing System for Retrieving Karaoke MusicabstractThis paper investigates the problem of retrieving karaoke music using query-by-singing techniques. Unlike regular CD music, where the stereo sound involves two audio channels that usually sound the same, karaoke music encompasses two distinct channels in each track: one is a mixture of the lead vocals and background accompaniment, and the other consists of accompaniment only. Although the two audio channels are distinct, the accompaniments in the two channels often resemble each other. We exploit this characteristic to: i) infer the background accompaniment for the lead vocals from the accompaniment-only channel, so that the main melody underlying the lead vocals can be extracted more effectively, and ii) detect phrase onsets based on the Bayesian information criterion (BIC) to predict the onset points of a song where a user's sung query may begin, so that the similarity between the melodies of the query and the song can be examined more efficiently. To further refine extraction of the main melody, we propose correcting potential errors in the estimated sung notes by exploiting a composition characteristic of popular songs whereby the sung notes within a verse or chorus section usually vary no more than two octaves. In addition, to facilitate an efficient and accurate search of a large music database, we employ multiple-pass dynamic time warping (DTW) combined with multiple-level data abstraction (MLDA) to compare the similarities of melodies. The results of experiments conducted on a karaoke database comprised of 1071 popular songs demonstrate the feasibility of query-by-singing retrieval for karaoke music. Hung-Ming Yu, Wei-Ho Tsai, Hsin-Min Wang |
IEEE Trans. Multim. | 3 |
| 2007 | Spoken document summarization using relevant informationabstractExtractive summarization usually automatically selects indicative sentences from a document according to a certain target summarization ratio, and then sequences them to form a summary. In this paper, we investigate the use of information from relevant documents retrieved from a contemporary text collection for each sentence of a spoken document to be summarized in a probabilistic generative framework for extractive spoken document summarization. In the proposed methods, the probability of a document being generated by a sentence is modeled by a hidden Markov model (HMM), while the retrieved relevant text documents are used to estimate the HMM’s parameters and the sentence’s prior probability. The results of experiments on Chinese broadcast news compiled in Taiwan show that the new methods outperform the previous HMM approach. Shih-Hsiang Lin, Hsin-Min Wang, Berlin Chen |
ASRU | 3 |
| 2007 | Improved Methods for Characterizing the Alternative Hypothesis using Minimum Verification Error Training for LLR-Based Speaker VerificationabstractSpeaker verification based on the log-likelihood ratio (LLR) is essentially a task of modeling and testing two hypotheses: the null hypothesis and the alternative hypothesis. Since the alternative hypothesis involves unknown imposters, it is usually hard to characterize a priori. In this paper, we propose a framework to better characterize the alternative hypothesis with the goal of optimally separating client speakers from imposters. The proposed framework is built on either a weighted arithmetic combination or a weighted geometric combination of useful information extracted from a set of pre-trained anti-speaker models. The parameters associated with the combinations are then optimized using minimum verification error training such that both the false acceptance probability and the false rejection probability are minimized. Our experiment results show that the proposed framework outperforms conventional LLR-based approaches. Yi-Hsiang Chao, Wei-Ho Tsai, Hsin-Min Wang, Ruei-Chuan Chang |
ICASSP (4) | 3 |
| 2007 | Phonetic Boundary Refinement using Support Vector MachineabstractIn this paper, we propose using support vector machine (SVM) to refine the hypothesized phone transition boundaries given by the HMM-based Viterbi forced alignment. We conducted experiments on the TIMIT speech corpus. The phone transitions were automatically partitioned into 46 clusters according to their acoustic characteristics and the cross-validation using the training data; hence, 46 phone-transition-dependent SVM classifiers were used for phone boundary refinement. The proposed HMM-SVM approach performs as well as the recent discriminative HMM-based segmentation. The best accuracies achieved are 81.23% within a tolerance of 10 ms and 92.47% within a tolerance of 20 ms. The mean boundary distance is 7.73 ms. Hung-Yi Lo, Hsin-Min Wang |
ICASSP (4) | 2 |
| 2007 | Speaker Clustering Based on Minimum Rand IndexabstractThis paper presents an effective method for clustering unknown speech utterances based on their associated speakers. The proposed method jointly optimizes the generated clusters and the number of clusters by estimating and minimizing the Rand index of the clustering. The Rand index, which reflects clustering errors that utterances from the same speaker are placed in different clusters, or utterances from different speakers are placed in the same cluster, reaches its minimal value only when the number of clusters is equal to the true speaker population size. We approximate the Rand index by a function of the similarity measures between utterances and employ the genetic algorithm to determine the cluster where each utterance should be located, such that the overall clustering errors are minimized. The experimental results show that the proposed speaker-clustering method outperforms the conventional method based on hierarchical agglomerative clustering in conjunction with the Bayesian information criterion to determine the number of clusters. Wei-Ho Tsai, Hsin-Min Wang |
ICASSP (4) | 2 |
| 2007 | Cascading Multimodal Verification using Face, Voice and Iris InformationabstractIn this paper we propose a novel fusion strategy which fuses information from multiple physical traits via a cascading verification process. In the proposed system users are verified by each individual modules sequentially in turns of face, voice and iris, and would be accepted once he/she is verified by one of the modules without performing the rest of the verifications. Through adjusting thresholds for each module, the proposed approach exhibits different behavior with respect to security and user convenience. We provide a criterion to select thresholds for different requirements and we also design an user interface which helps users find the personalized thresholds intuitively. The proposed approach is verified with experiments on our in-house face-voice-iris database. The experimental results indicate that besides the flexibility between security and convenience, the proposed system also achieves better accuracy than its most accurate module. Ping-Han Lee, Lu-Jong Chu, Yi-Ping Hung, Sheng-Wen Shih, Chu-Song Chen, Hsin-Min Wang |
ICME | 6 |
| 2007 | Evolutionary minimum verification error learning of the alternative hypothesis model for LLR-based speaker verificationabstractIt is usually difficult to characterize the alternative hypothesis precisely in a log-likelihood ratio (LLR)-based speaker verification system. In a previous work, we proposed using a weighted arithmetic combination (WAC) or a weighted geometric combination (WGC) of the likelihoods of the background models instead of heuristic combinations, such as the arithmetic mean and the geometric mean, to better characterize the alternative hypothesis. In this paper, we further propose learning the parameters associated with WAC or WGC via an evolutionary minimum verification error (MVE) training method, such that both the false acceptance probability and the false rejection probability can be minimized. Our experiment results show that the proposed methods outperform conventional LLR-based approaches. Index Terms: genetic algorithm, log-likelihood ratio, minimum verification error training, speaker verification Yi-Hsiang Chao, Wei-Ho Tsai, Shih-Sian Cheng, Hsin-Min Wang, Ruei-Chuan Chang |
INTERSPEECH | 4 |
| 2007 | A unified probabilistic generative framework for extractive spoken document summarizationabstractIn this paper, we consider extractive summarization of Chinese broadcast news speech. A unified probabilistic generative framework that combined the sentence generative probability and the sentence prior probability for sentence ranking was proposed. Each sentence of a spoken document to be summarized was treated as a probabilistic generative model for predicting the document. Two different matching strategies, i.e., literal term matching and concept matching, were extensively investigated. We explored the use of the hidden Markov model (HMM) and relevance model (RM) for literal term matching, while the word topical mixture model (WTMM) for concept matching. On the other hand, the confidence scores, structural features, and a set of prosodic features were properly incorporated together using the whole sentence maximum entropy model (WSME) for the estimation of the sentence prior probability. The experiments were performed on the Chinese broadcast news collected in Taiwan. Very promising and encouraging results were initially obtained. Hsuan-Sheng Chiu, Hsin-Min Wang, Berlin Chen |
INTERSPEECH | 3 |
| 2007 | Improved HMM/SVM methods for automatic phoneme segmentationabstractThis paper presents improved HMM/SVM methods for a two-stage phoneme segmentation framework, which tries to imitate the human phoneme segmentation process. The first stage per-forms hidden Markov model (HMM) forced alignment accord-ing to the minimum boundary error (MBE) criterion. The objec-tive is to align a phoneme sequence of a speech utterance with its acoustic signal counterpart based on MBE-trained HMMs and explicit phoneme duration models. The second stage uses the support vector machine (SVM) method to refine the hy-pothesized phoneme boundaries derived by HMM-based forced alignment. The efficacy of the proposed framework has been validated on two speech databases: the TIMIT English database Jen-Wei Kuo, Hung-Yi Lo, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2007 | Automatic Speaker Clustering Using a Voice Characteristic Reference Space and Maximum Purity EstimationabstractThis paper investigates the problem of automatically grouping unknown speech utterances based on their associated speakers. In attempts to determine which utterances should be grouped together, it is necessary to measure the voice similarities between utterances. Since most existing methods measure the inter-utterance similarities based directly on the spectrum-based features, the resulting clusters may not be well-related to speakers, but to various acoustic classes instead. This study remedies this shortcoming by projecting utterances onto a reference space trained to cover the generic voice characteristics underlying the whole utterance collection. The resultant projection vectors naturally reflect the relationships of voice similarities among all the utterances, and hence are more robust against interference from nonspeaker factors. Then, a clustering method based on maximum purity estimation is proposed, with the aim of maximizing the similarities between utterances within all the clusters. This method employs a genetic algorithm to determine the cluster to which each utterance should be assigned, which overcomes the limitation of conventional hierarchical clustering that the final result can only reach the local optimum. In addition, the proposed clustering method adapts a Bayesian information criterion to determine how many clusters should be created Wei-Ho Tsai, Shih-Sian Cheng, Hsin-Min Wang |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | On Maximizing the Within-Cluster Homogeneity of Speaker Voice Characteristics For Speech Utterance ClusteringabstractThis paper investigates the problem of how to partition unknown speech utterances into clusters, such that the overall within-cluster homogeneity of speakers' voice characteristics can be maximized. The within-cluster homogeneity is characterized by the likelihood probability that a cluster model, trained using all the utterances within a cluster, matches each of the within-cluster utterances. Such probability is then maximized by using a genetic algorithm, which determines the best cluster where each utterance should be located. For greater computational efficiency, also proposed is an alternative solution that approximates the likelihood probability with a divergence-based model similarity. The method is further designed to estimate the optimal number of clusters automatically. Wei-Ho Tsai, Hsin-Min Wang |
ICASSP (1) | 2 |
| 2006 | Improving the characterization of the alternative hypothesis via kernel discriminant analysis for likelihood ratio-based speaker verificationabstractThe performance of a likelihood ratio-based speaker verification system is highly dependent on modeling of the target speaker’s voice (the null hypothesis) and characterization of non-target speakers ’ voices (the alternative hypothesis). To better characterize the ill-defined alternative hypothesis, this study proposes a new likelihood ratio measure based on a composite-structure Gaussian mixture model, the so-called GMM2. Motivated by the combined use of a variety of background models to represent the alternative hypothesis, GMM2 is designed with an inner set of mixture weights connected to the significance of each individual Gaussian density, and an outer set of mixture weights connected to the significance of each individual background model. Through the use of kernel discriminant analysis namely, Kernel Fisher Discriminant (KFD) or Support Vector Machine (SVM), GMM2 is trained in such a manner that the utterances of the null hypothesis can be optimally separated from those of the alternative hypothesis. Index Terms: speaker verification, likelihood ratio, kernel Fisher discriminant, support vector machine Yi-Hsiang Chao, Wei-Ho Tsai, Hsin-Min Wang, Ruei-Chuan Chang |
INTERSPEECH | 3 |
| 2006 | Minimum boundary error training for automatic phonetic segmentationabstractAnnotated speech corpora are indispensable to various areas of speech research. In this paper, we present a novel discriminative training approach for HMM-based automatic phonetic segmentation. The objective of the proposed minimum boundary error (MBE) discriminative training approach is to minimize the expected boundary errors over a set of phonetic alignments represented as a phonetic lattice. This approach is inspired by the recently proposed minimum phone error (MPE) training algorithm for automatic speech recognition. To evaluate the MBE training approach, we conducted automatic phonetic segmentation experiments on the TIMIT acoustic-phonetic continuous speech corpus. The MBE-trained HMMs can identify 79.75 % of human-labeled phone boundaries within a tolerance of 10 ms, compared to 71.23 % identified by the conventional ML-trained HMMs. Moreover, by using the MBE-trained HMMs, only 7.89 % of automatically labeled phone boundaries have errors larger than 20 ms. Index Terms: minimum boundary error, automatic phonetic segmentation, HMM, forced alignment Jen-Wei Kuo, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2006 | Automatic singer recognition of popular music recordings via estimation and modeling of solo vocal signalsabstractIn this paper, we investigate the problem of automatic singer identification, detection and tracking in popular music recordings with one or multiple singers. This problem reflects an important issue in multimedia applications that require the transcription and indexing of music data to meet the increasing demand for content-based information retrieval. The major challenges for this study arise from the fact that a singer's voice tends to be arbitrarily altered from time to time and is inextricably intertwined with the signal of the background accompaniment. To determine who is singing, or whether or when a particular singer is present in a music recording, methods are presented for separating vocal from nonvocal regions, for isolating singers' vocal characteristics from background music, and for distinguishing singers from one another. Experimental evaluations conducted on a pop music database consisting of solo and duet tracks confirm the validity of the proposed methods. Wei-Ho Tsai, Hsin-Min Wang |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Gmm-Based Bhattacharyya Kernel Fisher Discriminant Analysis For Speaker RecognitionabstractClearly, the linear discriminant classifier is not robust enough to cope with most real-world data classification problems. Kernel Fisher discriminant analysis (KFDA) tries to increase the expressiveness of the discriminant based on the high order statistics of the data set. In this paper, we propose the GMM-based KFDA with the Bhattacharyya kernel to obtain a transformation, called a speaker eigenspace, based on which the transformed MFCC features are more discriminative for speaker recognition. In our approach, the eigenspace is directly constructed from the complete GMM parameter set, rather than the supervectors considering mean vectors only as in the eigenvoice approach. Moreover, FDA, which is believed to be more appropriate for classification accuracy than principal component analysis (PCA), is applied for eigenspace construction. The speaker identification experiments show that the new features outperform the MFCC features, in particular when the amount of enrollment data for each speaker is very small. Yi-Hsiang Chao, Hsin-Min Wang, Ruei-Chuan Chang |
ICASSP (1) | 2 |
| 2005 | Clustering Speech Utterances by Speaker Using Eigenvoice-Motivated Vector Space ModelsabstractThe paper investigates the problem of automatically grouping unknown speech utterances based on their associated speakers. The proposed method utilizes the vector space model, which was originally developed in document-retrieval research, to characterize each utterance as a tf-idf-based vector of acoustic terms, thereby deriving a reliable measurement of similarity between utterances. To define the required acoustic terms that are most representative in terms of voice characteristics, the Eigenvoice approach is applied to the utterances to be clustered, which creates a set of eigenvector-based terms. To further improve speaker-clustering performance, the proposed method encompasses a mechanism of blind relevance feedback for refining the inter-utterance similarity measure. Wei-Ho Tsai, Shih-Sian Cheng, Yi-Hsiang Chao, Hsin-Min Wang |
ICASSP (1) | 4 |
| 2005 | An Efficient Approach to Multimodal Person Identity Verification by Fusing Face and Voice InformationabstractThis paper presents an effective method to combine speech recognition, speaker verification and face verification for biometric authentication. Our method provides a light-weight enrollment process and an easy-to-use verification interface. A multi-face/single-sentence strategy is used to combine voice and face verification modules, and support vector machine is employed for information fusion. Experimental results show that our method can achieve high verification accuracies Hsien-Ting Cheng, Yi-Hsiang Chao, Shih-Liang Yeh, Chu-Song Chen, Hsin-Min Wang, Yi-Ping Hung |
ICME | 5 |
| 2005 | Speaker clustering of unknown utterances based on maximum purity estimationabstractThis paper addresses the problem of automatically grouping unknown speech utterances that are from the same speaker. A clustering method based on maximum purity estimation is proposed, with the aim of maximizing the similarities of voice characteristics between utterances within all the clusters. This method employs a genetic algorithm to determine the cluster where each utterance should be located, which overcomes the limitation of conventional hierarchical clustering that the final result can only reach the local optimum. The proposed clustering method also incorporates a Bayesian information criterion to determine how many clusters should be created. 1. Wei-Ho Tsai, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2005 | Fluent speech prosody: Framework and modeling
Chiu-yu Tseng, Shao-huang Pin, Yehlin Lee, Hsin-Min Wang, Yong-cheng Chen |
Speech Commun. | 4 |
| 2004 | Automatic detection and tracking of target singer in multi-singer music recordingsabstractIn this paper, we investigate the problem of automatically detecting and tracking a specified person's singing portions within a music recording with multiple simultaneous or nonsimultaneous singers. This problem reflects an important issue in multimedia applications which require transcription and indexing of music data in response to the increasing demand nowadays for content-based information retrieval. The major challenges of this study arise from the fact that the singer's voices are inextricably intertwined with the signal of the background accompaniment. To determine whether or when an accompanied voice is present and from a sought singer, methods are presented for separating vocal from non-vocal regions, for extracting and modeling singers' vocal characteristics, and for distinguishing among vocal regions performed by the target singer and other simultaneous or non-simultaneous singers. Wei-Ho Tsai, Hsin-Min Wang |
ICASSP (4) | 2 |
| 2004 | A query-by-example framework to retrieve music documents by singerabstractWe present a framework for music document retrieval that allows users to retrieve a specified singer's music recordings from an unlabeled database by submitting a fragment of music as a query to the system. Such a framework can be of great use for those wishing to know more about a particular singer but having no idea about what the name of this singer is. In order for the searched documents to be relevant to the query, methods are proposed to compare the similarity between document and query based on the automatic extraction of a singer's voice characteristics from a music recording. Wei-Ho Tsai, Hsin-Min Wang |
ICME | 2 |
| 2004 | Statistical Chinese spoken document retrieval using latent topical informationabstractInformation retrieval which aims to provide people with easy access to all kinds of information is now becoming more and more emphasized. However, most approaches to information retrieval are primarily based on literal term matching and operate in a deterministic manner. Thus their performance is often limited due to the problems of vocabulary mismatch and not able to be steadily improved through use. In order to overcome these drawbacks as well as to enhance the retrieval performance, in this paper we explore the use of topical mixture model for statistical Chinese spoken document retrieval. Various kinds of model structures and learning approaches were extensively investigated. In addition, the retrieval capabilities were verified by comparison with the conventional vector space model and latent semantic indexing model, as well as our previously presented HMM/N-gram retrieval model. The experiments were performed on the TDT2 Chinese collection. Noticeable improvements in retrieval performance were obtained. Jen-Wei Kuo, Yao-Min Huang, Berlin Chen, Hsin-Min Wang |
INTERSPEECH | 4 |
| 2004 | Speaker clustering of speech utterances using a voice characteristic reference spaceabstractThis paper presents an effective technique for clustering speech utterances based on their associated speaker. In attempts to determine which utterances are from the same speakers, a prerequisite is to measure the similarity of voice characteristics between utterances. Since the vast majority of existing methods evaluate the inter-utterance similarity by taking only the information from the spectrum-based features of utterance pairs into account, the resulting clusters may not be well relevant to speaker, but instead likely to the environmental conditions or other acoustic classes. To compensate for this shortcoming, this study proposes to project utterances from their spectrum-based feature representation onto a reference space trained to cover the generic voice characteristics inherently in all of the utterances to be clustered. The resultant projection vectors naturally reflect the relationships between all the utterances and are more robust against the interference from non-speaker factors. We exemplarily present three distinct implementations for reference space creation. 1. Wei-Ho Tsai, Shih-Sian Cheng, Hsin-Min Wang |
INTERSPEECH | 3 |
| 2004 | METRIC-SEQDAC: a hybrid approach for audio segmentationabstractThis paper presents a hybrid approach for audio segmentation, in which the metric-based segmentation with long sliding windows is applied first to segment an audio stream into shorter sub-segments, and then the divide-and-conquer segmentation is applied to a fixed-length window that slides from the beginning to the end of each sub-segment to sequentially detect the remaining acoustic changes. The experimental results on five one-hour broadcast news shows show that our approach outperforms the existing metric-based and model-selection-based approaches. Hsin-Min Wang, Shih-Sian Cheng |
INTERSPEECH | 1 |
| 2004 | Mandarin-English Information (MEI): investigating translingual speech retrieval
Helen M. Meng, Berlin Chen, Sanjeev Khudanpur, Gina-Anne Levow, Wai Kit Lo, Douglas W. Oard, Patrick Schone, Karen Tang, Hsin-Min Wang, Jianqiang Wang 0002 |
Comput. Speech Lang. | 9 |
| 2004 | A discriminative HMM/N-gram-based retrieval approach for mandarin spoken documentsabstractIn recent years, statistical modeling approaches have steadily gained in popularity in the field of information retrieval. This article presents an HMM/N-gram-based retrieval approach for Mandarin spoken documents. The underlying characteristics and the various structures of this approach were extensively investigated and analyzed. The retrieval capabilities were verified by tests with word- and syllable-level indexing features and comparisons to the conventional vector-space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. Fusion of information via indexing word- and syllable-level features was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2003 | A sequential metric-based audio segmentation method via the Bayesian information criterionabstractIn this paper, we propose a sequential metric-based audio segmentation method that has the advantage of low computation cost of metric-based methods and the advantage of high accuracy of model-selection-based methods. There are two major differences between our method and the conventional metricbased methods:(1) Each changing point has multiple chances to be detected by different pairs of windows, rather than only once by its neighboring acoustic information.(2) By introducing the Bayesian Information Criterion(BIC) into the distance computation of two windows, we can deal with the thresholding issue more easily. We used five one-hour broadcast news shows for experiments, and the experimental results show that our method performs as well as the model-selection-based methods, but with a lower computation cost. Shih-Sian Cheng, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2003 | Multi-scale document expansion in English-Mandarin cross-language spoken document retrievalabstractThis paper presents the application of document expansion using a side collection to a cross-language spoken document retrieval (CL-SDR) task to improve retrieval performance. Document expansion is applied to a series of English-Mandarin CL-SDR experiments using selected retrieval models (probabilistic belief network, vector space model, and HMM-based retrieval model). English textual queries are used to retrieve relevant documents from an archive of Mandarin radio broadcast news. We have devised a multiscale approach for document expansion – a process that enriches the Mandarin spoken document collection in order to improve overall retrieval performance. A document is expanded by (i) first retrieving related documents on a character bigram scale, (ii) then extracting word units from such related documents as expansion terms to augment the original document and (iii) finally indexing all documents in the collection by means of character bigrams and those expanded terms by within-word character bigrams to prepare for future retrieval. Hence the document expansion approach is multi-scale as it involves both word and subword scales. Experimental results show that this approach achieves performance improvements up to 14 % across several retrieval models. 1. Wai Kit Lo, Yuk-Chi Li, Gina-Anne Levow, Hsin-Min Wang, Helen M. Meng |
INTERSPEECH | 4 |
| 2003 | Automatic singer identification of popular music recordings via estimation and modeling of solo vocal signalabstractThis study presents an effective technique for automatically identifying the singer of a music recording. Since the vast majority of popular music contains background accompaniment during most or all vocal passages, directly acquiring isolated solo voice data for extracting the singer’s vocal characteristics is usually infeasible. To eliminate the interference of background music for singer identification, we leverage statistical estimation of a piece’s musical background to build a reliable model for the solo voice. Validity of the proposed singer identification system is confirmed via the experimental evaluations conducted on a 23-singer pop music database. 1. Wei-Ho Tsai, Hsin-Min Wang, Dwight Rodgers |
INTERSPEECH | 2 |
| 2002 | A hierarchical tag-graph search scheme with layered grammar rules for spontaneous speech understanding
Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
Pattern Recognit. Lett. | 3 |
| 2002 | Discriminating capabilities of syllable-based features and approaches of utilizing them for voice retrieval of speech information in Mandarin ChineseabstractWith the rapidly growing use of the audio and multimedia information over the Internet, the technology for retrieving speech information using voice queries is becoming more and more important. In this paper, considering the monosyllabic structure of the Chinese language, a whole class of syllable-based indexing features, including overlapping segments of syllables and syllable pairs separated by a few syllables, is extensively investigated based on a Mandarin broadcast news database. The strong discriminating capabilities of such syllable-based features were verified by comparing with the word- or character-based features. Good approaches for better utilizing such capabilities, including fusion with the word- and character-level information and improved approaches to obtain better syllable-based features and query expressions, were extensively investigated. Very encouraging experimental results were obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Eigenspace-based maximum a posteriori linear regression for rapid speaker adaptationabstractWe present an eigenspace-based approach toward prior density selection for the MAPLR framework. The proposed eigenspace-based MAPLR approach was developed by introducing a priori knowledge analysis on the training speakers via probabilistic principal component analysis (PPCA), so as to construct an eigenspace for speaker-specific full regression matrices as well as to derive a set of bases called eigen-matrices. The priors of MAPLR transformations for each outside speaker are then chosen in the space spanned by the first K eigen-matrices. By incorporating the PPCA model into the MAPLR scheme, the number of free parameters in choosing the priors can be effectively reduced, while the underlying structure of the acoustic space as well as the precise modeling of the inter-dimensional correlation among the model parameters can be well preserved. Both supervised and unsupervised adaptation experiments showed that the proposed approach significantly outperformed the conventional maximum likelihood linear regression (MLLR) approach using either diagonal or full regression matrices. Kuan-Ting Chen, Hsin-Min Wang |
ICASSP | 2 |
| 2001 | Multi-scale-audio indexing for translingual spoken document retrievalabstractMEI (Mandarin-English Information) is an English-Chinese crosslingual spoken document retrieval (CL-SDR) system developed during the Johns Hopkins University Summer Workshop 2000. We integrate speech recognition, machine translation, and information retrieval technologies to perform CL-SDR. MEI advocates a multi-scale paradigm, where both Chinese words and subwords (characters and syllables) are used in retrieval. The use of subword units can complement the word unit in handling the problems of Chinese word tokenization ambiguity, Chinese homophone ambiguity, and out-of-vocabulary words in audio indexing. This paper focuses on multi-scale audio indexing in MEI. Experiments are based on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3), where we indexed Voice of America Mandarin news broadcasts by speech recognition on both the word and subword scales. We discuss the development of the MEI syllable recognizer, the representations of spoken documents using overlapping subword n-grams and lattice structures. Results show that augmenting words with subwords is beneficial to CL-SDR performance. Hsin-Min Wang, Helen M. Meng, Patrick Schone, Berlin Chen, Wai Kit Lo |
ICASSP | 1 |
| 2001 | Improved spoken document retrieval by exploring extra acoustic and linguistic cuesabstractIn this paper, we explored the use of various extra information to improve the performance of spoken document retrieval (SDR). From the speech recognition perspective, we incorporated the acoustic stress and word confusion information into the audio indexing. From the linguistic perspective, we applied the part-of-speech information in both the audio indexing and the query representation. From the information retrieval perspective, we integrated techniques such as the query expansion by word associations and the blind relevance feedback into the retrieval process. The SDR experiments were based on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). We used the Chinese newswire text stories as query exemplars and the Mandarin Chinese audio news stories as the spoken documents. With all the above acoustic and linguistic cues applied, the average precision was improved from 0.5122 to 0.6312 for the TDT-2 collection and from 0.6216 to 0.7172 for the TDT-3 collection. 1. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2001 | An HMM/n-gram-based linguistic processing approach for Mandarin spoken document retrievalabstractIn this paper an HMM/N-gram-based linguistic processing approach for Mandarin spoken document retrieval is presented. The underlying characteristics and different structures of this approach were extensively investigated. The retrieval capabilities were verified by tests with indexing features of word- and syllable(subword)-levels and comparison with the conventional vector space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. The information fusion of indexing features of word- and syllable-levels was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2001 | Comparative analysis for data-driven temporal filters obtained via principal component analysis (PCA) and linear discriminant analysis (LDA) in speech recognitionabstractThe Linear Discriminant Analysis (LDA) has been widely used to derive the data-driven temporal filtering of speech feature vectors. In this paper, we proposed that the Principal Component Analysis (PCA) can also be used in the optimization process just as LDA to obtain the temporal filters, and detailed comparative analysis between these two approaches are presented and discussed. It's found that the PCA-derived temporal filters significantly improve the recognition performance of the original MFCC features as LDA-derived filters do. Also, while PCA/LDA filters are combined with the conventional temporal filters, RASTA or CMS, the recognition performance will be further improved regardless the training and testing environments are matched or mismatched, compressed or noise corrupted. 1. Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2000 | Retrieval of broadcast news speech in Mandarin Chinese collected in Taiwan using syllable-level statistical characteristicsabstractSpoken document retrieval has been extensively studied over the years because of its high potential in various applications in the near future. Considering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of spoken documents in Mandarin Chinese using syllable-level statistical characteristics has been studied, and very encouraging experimental results on retrieval of broadcast news speech collected in Taiwan were obtained. This paper reports some interesting initial results and findings obtained in this research. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ICASSP | 2 |
| 2000 | Fast speaker adaptation using eigenspace-based maximum likelihood linear regressionabstractThis paper presents an eigenspace-based fast speaker adaptation approach which can improve the modeling accuracy of the conventional maximum likelihood linear regression (MLLR) techniques when only very limited adaptation data is available. The proposed eigenspace-based MLLR approach was developed by introducing a priori knowledge analysis on the training speakers via PCA, so as to construct an eigenspace for MLLR full regression matrices as well as to derive a set of bases called eigen-matrices. The full regression matrices for each outside speaker are then constrained to be located in the space spanned by the first K eigen-matrices. The proposed eigenspace-based regression matrices, serving as an initial estimate of the speaker-specific MLLR transformation, effectively reduces the number of free parameters, while precise modeling for the inter-dimensional correlation among the model parameters by full matrices was maintained. Experimental results showed that for supervised adaptation... Kuan-Ting Chen, Wen-Wei Liau, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 3 |
| 2000 | Retrieval of mandarin broadcast news using spoken queriesabstractConsidering the monosyllabic structure of the Chinese language, a whole class of indexing features for retrieval of Mandarin broadcast news using syllable-level statistical characteristics has been previously investigated. This paper presents the improvements achieved over the previous results. The major differences are: (1) Multi-scale character- and word-level indexing terms have been integrated with the syllable-level information. (2) Information cues from the contemporary newswire text corpus have been used to create more accurate syllable indexing terms. (3) Automatic document expansion, blind relevance feedback, and query expansion via the term association matrix have been applied in retrieval. With all these schemes, the average precision can be improved from 55.46 % to 71.29%. 1. Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2000 | Automatic metric-based speech segmentation for broadcast news via principal component analysisabstractIn this paper, we proposed an algorithm used to improve the performance of the metric-based segmentation techniques, by which the segmentation points are found at maxima of a distance measured between two contiguous windows shifted along the stream of speech features. In our proposed method, the PCA processes are first performed on the speech features to obtain more robust features, and then the above metric-based segmentation was applied on the PCA-derived features to decide the segmentation points. Experiment results show that our proposed method can efficiently improve the detection rates of the segmentation points up to 7% while the false alarm rates remain unchanged. Jeih-Weih Hung, Hsin-Min Wang, Lin-Shan Lee |
INTERSPEECH | 2 |
| 2000 | Syllable-Based Chinese Text/Spoken Document Retrieval Using Text/Speech QueriesabstractIn light of the rapid growth of Chinese information resources on the Internet, this study investigates a novel approach that deals with the problem of Chinese text and spoken document retrieval using both text and speech queries. By properly utilizing the monosyllabic structure of the Chinese language, the proposed approach estimates the statistical similarity between the text/speech queries and the text/spoken documents at the phonetic level using the syllable-based statistical information. The investigation successfully implemented a prototype system with an interface supporting some user-friendly functions and the initial test results demonstrate the feasibility of the proposed approach. Bo-Ren Bai, Berlin Chen, Hsin-Min Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2000 | A spoken-access approach for chinese text and speech information retrievalabstractThis paper presents an efficient spoken-access approach for both Chinese text and Mandarin speech information retrieval. The proposed approach is developed not only to deal with the retrieval of spoken documents, but also to improve the capability of human-computer interaction via voice input for information-retrieval systems. Based on utilization of the monosyllabic structure of the Chinese language, the proposed approach can tolerate speech recognition errors by performing speech query recognition and approximate information retrieval at the syllable-level. Furthermore, with the help of automatic term suggestion and relevance feedback techniques, the proposed approach is robust in enabling users using voice input to interact with IR systems at each stage of the retrieval process. Extensive experiments show that the proposed approach can improve the effectiveness of information retrieval via speech interaction. The encouraging results suggest that a Mandarin speech interface for information retrieval and digital library systems can, therefore, be developed. Lee-Feng Chien, Hsin-Min Wang, Bo-Ren Bai, Sun-Chien Lin |
J. Am. Soc. Inf. Sci. | 2 |
| 2000 | Mandarin spoken document retrieval based on syllable lattice matching
Hsin-Min Wang |
Pattern Recognit. Lett. | 1 |
| 2000 | Experiments in syllable-based retrieval of broadcast news speech in Mandarin Chinese
Hsin-Min Wang |
Speech Commun. | 1 |
| 1999 | Consistent dialogue across concurrent topics based on an expert system modelabstractThis paper describes the development and evaluation of objective methods for testing synthetic intonation. While subjective methods are available for assessing the quality of synthetic intonation, such tests consume time and resources, and are not useful for day-to-day model development. Therefore, objective measures of F0 modelling are necessary. Currently, objective evaluation of synthetic intonation involves the use of Root Mean Squared Error and Correlation. However, it is unclear how large an improvement in either score must be before it is reflected perceptually. It is also unclear how detailed an analysis these metrics provide. Therefore, two other metrics are to be tested, both of which are similar to a basic RMSE measurement. All of the evaluation results are compared to a perceptual study in order to determine how the objective measures relate to perceived differences in the contours. 1. INTRODUCTION One difficulty in building models for synthesizing intonation is determinin... Bor-Shen Lin, Hsin-Min Wang, Lin-Shan Lee |
EUROSPEECH | 2 |
| 1999 | Automatic selection of phonetically distributed sentence sets for speaker adaptation with application to large vocabulary Mandarin speech recognition
Jia-Lin Shen, Hsin-Min Wang, Ren-Yuan Lyu, Lin-Shan Lee |
Comput. Speech Lang. | 2 |
| 1998 | A*-admissible key-phrase spotting with sub-syllable level utterance verificationabstractIn this paper, we propose an A*-admissible key-phrase spotting framework, which needs little domain knowledge and is capable of extracting salient key-phrase fragments from an input utterance in real-time. There are two key features in our approach. Firstly, the acoustic models and the search framework are specially designed such that very high degree vocabulary flexibility can be achieved for any desired application tasks. Secondly, the search framework uses an efficient two-pass A* search to generate N-best key-phrase candidates and then several sub-syllable level verification functions are properly weighted and used to further improve the recognition accuracy. Experimental results show that the A*-admissible key-phrase spotting with sub-word level utterance method outperforms the baseline methods used in common approaches. 1. INTRODUCTION In recent years, various spoken dialog systems have been widely investigated for the fast growing demand for real-world applications. It is diff... Berlin Chen, Hsin-Min Wang, Lee-Feng Chien, Lin-Shan Lee |
ICSLP | 2 |
| 1998 | Hierarchical tag-graph search for spontaneous speech understanding in spoken dialog systemsabstractIt has been relatively difficult to develop natural language parsers for spoken dialog systems, not only because of the possible recognition errors, pauses, hesitations, out-ofvocabulary words, and the grammatically incorrect sentence structures, but because of the great efforts required to develop a general enough grammar with satisfactory coverage and flexibility to handle different applications. In this paper, a new hierarchical graph-based search scheme with layered structure is presented, which is shown to provide more robust and flexible spontaneous speech understanding for spoken dialog systems. 1.INTRODUCTION Traditionally, natural language understanding is integrated with the speech recognizer with a N-best interface in spoken dialog systems [1][2], that is, the recognizer sequentially generates its best N sentence hypotheses until any one is accepted by the natural language understanding part. However, for spontaneous speech with fragments, disfluencies, OOV words, and ill-... Bor-Shen Lin, Berlin Chen, Hsin-Min Wang, Lin-Shan Lee |
ICSLP | 3 |
| 1998 | Towards a Mandarin voice memo systemabstractUsing voice memos in stead of text memos is believed to be more natural, convenient, and attractive. This paper presents a working Mandarin voice memo system that provides functions of automatic notification and voice retrieval. The main techniques include the content-based spoken document retrieval approach and the date-time expression detection and understanding approach. Extensive preliminary experiments were performed and encouraging results were demonstrated. 1. INTRODUCTION Using voice memos in stead of text memos is believed to be more natural, convenient, and attractive because it is definitely much easier for people to speak memos than to write down memos or to type memos into computers using a keyboard. Furthermore, users are more likely to record detailed information if all they need to do is just speak. However, it's far more difficult to retrieve these voice memos than to retrieve the text ones. With advances in the speech recognition technology, voice retrieval of spoke... Hsin-Min Wang, Bor-Shen Lin, Berlin Chen, Bo-Ren Bai |
ICSLP | 1 |
| 1997 | Internet Chinese information retrieval using unconstrained Mandarin speech queries based on a client-server architecture and a PAT-tree-based language modelabstractIn order to pursue high performance of Chinese information access on the Internet, this paper presents an attractive approach with a successful integration of efficient speech recognition and information retrieval techniques. A working system based on the proposed approach for speech retrieval of real-time Chinese netnews services has been implemented and tested. Very exciting performance has been achieved. Lee-Feng Chien, Sung-Chien Lin, Jenn-Chau Hong, Ming-Chiuan Chen, Hsin-Min Wang, Jia-Lin Shen, Keh-Jiann Chen, Lin-Shan Lee |
ICASSP | 5 |
| 1997 | Complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary using limited training dataabstractThis correspondence presents the first known results of complete recognition of continuous Mandarin speech for the Chinese language with very large vocabulary but very limited training data. Various acoustic and linguistic processing techniques were developed, and a prototype system of a continuous speech Mandarin dictation machine has been successfully implemented. The best recognition accuracy achieved is 92.2% for finally decoded Chinese characters. Hsin-Min Wang, Tai-Hsuan Ho, Rung-Chiung Yang, Jia-Lin Shen, Bo-Ren Bai, Jenn-Chau Hong, Wei-Peng Chen, Tong-Lo Yu, Lin-Shan Lee |
IEEE Trans. Speech Audio Process. | 1 |
| 1996 | Frameworks for recognition of Mandarin syllables with tones using sub-syllabic units
Chih-Heng Lin, Pei-Yih Ting, Hsin-Min Wang |
Speech Commun. | 4 |
| 1995 | Complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary but limited training dataabstractThis paper presents the first known results for complete recognition of continuous Mandarin speech for Chinese language with very large vocabulary but very limited training data. Although some isolated-syllable-based or isolated-word-based large-vocabulary Mandarin speech recognition systems have been successfully developed, a continuous-speech-based system of this kind has never been reported before. For successful development of this system, several important techniques have been used, including acoustic modeling of a set of sub-syllabic models for base syllable recognition and another set of context-dependent models for tone recognition, a multiple candidate searching technique based on a concatenated syllable matching algorithm to synchronize base syllable and tone recognition, and a word-class-based Chinese language model for linguistic decoding. The best recognition accuracy achieved is 88.69% for finally decoded Chinese characters, with 88.69%, 91.57%, and 81.37% accuracy for base syllables, tones, and tonal syllables respectively. Hsin-Min Wang, Jia-Lin Shen, Yen-Ju Yang, Chiu-yu Tseng, Lin-Shan Lee |
ICASSP | 1 |
| 1995 | Fast and accurate continuous speech recognition for Chinese language with very large vocabulary
Tai-Hsuan Ho, Hsin-Min Wang, Lee-Feng Chien, Keh-Jiann Chen, Lin-Shan Lee |
EUROSPEECH | 2 |
| 1994 | An initial study on a segmental probability model approach to large-vocabulary continuous Mandarin speech recognitionabstractThis paper presents an initial study to perform large-vocabulary continuous Mandarin speech recognition based on a segmental probability model (SPM) approach. SPM was first proposed for recognition of isolated Mandarin syllables, in which every syllable must be equally segmented before recognition. A concatenated syllable matching algorithm is therefore introduced in place of the conventional Viterbi search algorithm to perform the recognition process based on SPM. In addition, a training procedure is also proposed to reestimate the SPM parameters for continuous speech. Preliminary simulation results indicate that significant improvements in both recognition rates and speed can be achieved as compared to the conventional HMM-based Viterbi search approaches.> Jia-Lin Shen, Hsin-Min Wang, Bo-Ren Bai, Lin-Shan Lee |
ICASSP (2) | 2 |
| 1994 | Incremental speaker adaptation using phonetically balanced training sentences for Mandarin syllable recognition based on segmental probability models
Jia-Lin Shen, Hsin-Min Wang, Ren-Yuan Lyu, Lin-Shan Lee |
ICSLP | 2 |
| 1993 | Golden Mandarin (II)-an improved single-chip real-time Mandarin dictation machine for Chinese language with very large vocabulary
Lin-Shan Lee, Chiu-yu Tseng, Keh-Jiann Chen, I-Jung Hung, Ming-Yu Lee, Lee-Feng Chien, Yumin Lee, Ren-Yuan Lyu, Hsin-Min Wang, Yung-Chuan Wu, Tung-Sheng Lin, Hung-Yan Gu, Chi-ping Nee, Chun-Yi Liao, Yeng-Ju Yang, Yuan-Cheng Chang, Rung-Chiung Yang |
ICASSP (2) | 9 |