EDBT 2026 Demo / reviewers in the wild / expert
Ting-Yao Hu
dblp:76/10031
· DBLP profile ↗
19ranked-venue papers
8as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Corpus Synthesis for Zero-Shot ASR Domain Adaptation Using Large Language ModelsabstractWhile Automatic Speech Recognition (ASR) systems are widely used in many real-world applications, they often do not generalize well to new domains and need to be fine-tuned on data from these domains. However, target-domain data usually are not readily available in many scenarios. In this paper, we propose a new strategy for adapting ASR models to new target domains without any text or speech from those domains. To accomplish this, we propose a novel data synthesis pipeline that uses a Large Language Model (LLM) to generate a target domain text corpus, and a state-of-the-art controllable speech synthesis model to generate the corresponding speech. We propose a simple yet effective in-context instruction fine-tuning strategy to increase the effectiveness of LLM in generating text corpora for new domains. Experiments on the SLURP dataset show that the proposed method achieves an average relative word error rate improvement of 28% on unseen target domains without any performance drop in source domains. Hsuan Su, Ting-Yao Hu, Hema Swetha Koppula, Raviteja Vemulapalli, Jen-Hao Rick Chang, Karren D. Yang, Gautam Varma Mantena, Oncel Tuzel |
ICASSP | 2 |
| 2023 | I See What You Hear: A Vision-Inspired Method to Localize WordsabstractThis paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contemporary literature for visual data. Noting that an audio can be interpreted as a 1-dimensional image, object localization techniques can be fundamentally useful for word localization. Building upon this idea, we propose a lightweight solution for word detection and localization. We use bounding box regression for word localization, which enables our model to detect the occurrence, offset, and duration of keywords in a given audio stream. We experiment with LibriSpeech and train a model to localize 1000 words. Compared to existing work [1], our method reduces model size by 94%, and improves the F1 score by 6.5%. Mohammad Samragh Razlighi, Arnav Kundu, Ting-Yao Hu, Aman Chadha, Ashish Srivastava, Minsik Cho, Oncel Tuzel, Devang Naik |
ICASSP | 3 |
| 2023 | Text is all You Need: Personalizing ASR Models Using Controllable Speech SynthesisabstractAdapting generic speech recognition models to specific individuals is a challenging problem due to the scarcity of personalized data. Recent works have proposed boosting the amount of training data using personalized text-to-speech synthesis. Here, we ask two fundamental questions about this strategy: when is synthetic data effective for personalization, and why is it effective in those cases? To address the first question, we adapt a state-of-the-art automatic speech recognition (ASR) model to target speakers from four benchmark datasets representative of different speaker types. We show that ASR personalization with synthetic data is effective in all cases, but particularly when (i) the target speaker is underrepresented in the global data, and (ii) the capacity of the global model is limited. To address the second question of why personalized synthetic data is effective, we use controllable speech synthesis (CSS) to generate speech with varied styles and content. Surprisingly, we find that the text content of the synthetic data, rather than style, is important for speaker adaptation. These results lead us to propose a data selection strategy for ASR personalization based on speech content. Karren D. Yang, Ting-Yao Hu, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel |
ICASSP | 2 |
| 2022 | SYNT++: Utilizing Imperfect Synthetic Data to Improve Speech RecognitionabstractWith recent advances in speech synthesis, synthetic data is becoming a viable alternative to real data for training speech recognition models. However, machine learning with synthetic data is not trivial due to the gap between the synthetic and the real data distributions. Synthetic datasets may contain artifacts that do not exist in real data such as structured noise, content errors, or unrealistic speaking styles. Moreover, the synthesis process may introduce a bias due to uneven sampling of the data manifold. We propose two novel techniques during training to mitigate the problems due to the distribution gap: (i) a rejection sampling algorithm and (ii) using separate batch normalization statistics for the real and the synthetic samples. We show that these methods significantly improve the training of speech recognition models using synthetic data. We evaluate the proposed approach on keyword detection and Automatic Speech Recognition (ASR) tasks, and observe up to 18% and 13% relative error reduction, respectively, compared to naively using the synthetic data. Ting-Yao Hu, Mohammadreza Armandpour, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Oncel Tuzel |
ICASSP | 1 |
| 2021 | Project RISE: Recognizing Industrial Smoke EmissionsabstractIndustrial smoke emissions pose a significant concern to human health. Prior works have shown that using Computer Vision (CV) techniques to identify smoke as visual evidence can influence the attitude of regulators and empower citizens to pursue environmental justice. However, existing datasets are not of sufficient quality nor quantity to train the robust CV models needed to support air quality advocacy. We introduce RISE, the first large-scale video dataset for Recognizing Industrial Smoke Emissions. We adopted a citizen science approach to collaborate with local community members to annotate whether a video clip has smoke emissions. Our dataset contains 12,567 clips from 19 distinct views from cameras that monitored three industrial facilities. These daytime clips span 30 days over two years, including all four seasons. We ran experiments using deep neural networks to establish a strong performance baseline and reveal smoke recognition challenges. Our survey study discussed community feedback, and our data analysis displayed opportunities for integrating citizen scientists and crowd workers into the application of Artificial Intelligence for Social Impact. Yen-Chia Hsu, Ting-Hao 'Kenneth' Huang, Ting-Yao Hu, Paul Dille, Sean Prendi, Ryan Hoffman, Anastasia Tsuhlares, Jessica Pachuta, Randy Sargent, Illah R. Nourbakhsh |
AAAI | 3 |
| 2021 | Statistical Distance Metric Learning for Image Set RetrievalabstractMeasuring similarity between two image sets is instrumental in many computer vision tasks, such as video face recognition, multi-shot person re-identification and gait recognition. In most of the recent works, it is done by aggregating the embedding features of images as a fixed size vector, and calculating a metric in vector space (i.e. Euclidean distance). The embedding feature function can be learned by deep metric learning (DML) technique. However, methods relying on feature aggregation fail to capture the diversity and uncertainty within image sets. In this paper, we obviate the need of feature aggregation and propose a novel Statistical Distance Metric Learning (SDML) framework, which represents each image set as a probability distribution in embedding feature space and compares two image sets by statistical distance between their distributions. Among all types of statistical distance, we choose Jeffrey’s divergence (JD), which can be obtained from two embedding feature sets by kNN based density estimator. We also design a statistical centroid loss function to enhance the discriminative power of training process. Our SDML framework naturally preserves the diversity within an image set, and the relation between two sets. We evaluate our proposed approach on gait recognition and multi-shot person re-id. The experiment results show that SDML outperforms conventional DML, and also receives competitive/superior performance comparing to the previous state-of-the-arts on the aforementioned tasks. Ting-Yao Hu, Alex Hauptmann 0001 |
ICASSP | 1 |
| 2021 | SapAugment: Learning A Sample Adaptive Policy for Data AugmentationabstractData augmentation methods usually apply the same augmentation (or a mix of them) to all the training samples. For example, to perturb data with noise, the noise is sampled from a Normal distribution with a fixed standard deviation, for all samples. We hypothesize that a hard sample with high training loss already provides strong training signal to update the model parameters and should be perturbed with mild or no augmentation. Perturbing a hard sample with a strong augmentation may also make it too hard to learn from. Furthermore, a sample with low training loss should be perturbed by a stronger augmentation to provide more robustness to a variety of conditions. To formalize these intuitions, we propose a novel method to learn a Sample-Adaptive Policy for Augmentation – SapAugment. Our policy adapts the augmentation parameters based on the training loss of the data samples. In the example of Gaussian noise, a hard sample will be perturbed with a low variance noise and an easy sample with a high variance noise. Furthermore, the proposed method combines multiple augmentation methods into a methodical policy learning framework and obviates hand-crafting augmentation parameters by trial-and-error. We apply our method on an automatic speech recognition (ASR) task, and combine existing and novel augmentations using the proposed framework. We show substantial improvement, up to 21% relative reduction in word error rate on LibriSpeech dataset, over the state-of-the-art speech augmentation method. Ting-Yao Hu, Ashish Shrivastava 0001, Jen-Hao Rick Chang, Hema Swetha Koppula, Kyuyeon Hwang, Ozlem Kalinli, Oncel Tuzel |
ICASSP | 1 |
| 2021 | Pose Guided Person Image Generation With Hidden P-Norm RegressionabstractIn this paper, we propose a novel approach to solve the pose guided person image generation task. We assume that the relation between pose and appearance information can be described by a simple matrix operation in hidden space. Based on this assumption, our method estimates a pose-invariant feature matrix for each identity, and uses it to predict the target appearance conditioned on the target pose. The estimation process is formulated as a p-norm regression problem in hidden space. By utilizing the differentiation of the solution of this regression problem, the parameters of the whole framework can be trained in an end-to-end manner. While most previous works are only applicable to the supervised training and single-shot generation scenario, our method can be easily adapted to unsupervised training and multi-shot generation. Extensive experiments on the challenging Market-1501 dataset show that our method yields competitive performance in all the aforementioned variant scenarios. Ting-Yao Hu, Alex Hauptmann 0001 |
ICIP | 1 |
| 2020 | Unsupervised Style and Content Separation by Minimizing Mutual Information for Speech SynthesisabstractWe present a method to generate speech from input text and a style vector that is extracted from a reference speech signal in an unsupervised manner, i.e., no style annotation, such as speaker information, is required. Existing unsupervised methods, during training, generate speech by computing style from the corresponding ground truth sample and use a decoder to combine the style vector with the input text. Training the model in such a way leaks content information into the style vector. The decoder can use the leaked content and ignore some of the input text to minimize the reconstruction loss. At inference time, when the reference speech does not match the content input, the output may not contain all of the content of the input text. We refer to this problem as "content leakage", which we address by explicitly estimating and minimizing the mutual information between the style and the content through an adversarial training formulation. We call our method MIST - Mutual Information based Style Content Separation. The main goal of the method is to preserve the input content in the synthesized speech signal, which we measure by the word error rate (WER) and show substantial improvements over state-of-the-art unsupervised speech synthesis methods. Ting-Yao Hu, Ashish Shrivastava 0001, Oncel Tuzel, Chandra Dhir |
ICASSP | 1 |
| 2019 | Multi-shot Person Re-identification through Set Distance with Visual Distributional RepresentationabstractPerson re-identification aims to identify a specific person at distinct times and locations. It is challenging because of occlusion, illumination, and viewpoint change in camera views. Recently, multi-shot person re-id task receives more attention since it is closer to real-world application. A key point of a good algorithm for multi-shot person re-id is the temporal aggregation of the person appearance features. While most of the current approaches apply pooling strategies and obtain a fixed-size vector representation, these may lose the matching evidence between examples. In this work, we propose the idea of visual distributional representation, which interprets an image set as samples drawn from an unknown distribution in appearance feature space. Based on the supervision signals from a downstream task of interest, the method reshapes the appearance feature space and further learns the unknown distribution of each image set. In the context of multi-shot person re-id, we apply this novel concept along with Wasserstein distance and jointly learn a distributional set distance function between two image sets. In this way, the proper alignment between two image sets can be discovered naturally in a non-parametric manner. Our experiment results on three public datasets show the advantages of our proposed method compared to other state-of-the-art approaches. Ting-Yao Hu, Alex Hauptmann 0001 |
ICMR | 1 |
| 2017 | Integrating Verbal and Nonvebval Input into a Dynamic Response Spoken Dialogue SystemabstractIn this work, we present a dynamic response spoken dialogue system (DRSDS). It is capable of understanding the verbal and nonverbal language of users and making instant, situation-aware response. Incorporating with two external systems, MultiSense and email summarization, we built an email reading agent on mobile device to show the functionality of DRSDS. Ting-Yao Hu, Chirag Raman, Salvador Medina Maza, Liangke Gui, Tadas Baltrusaitis, Robert E. Frederking, Louis-Philippe Morency, Alan W. Black, Maxine Eskénazi |
AAAI | 1 |
| 2015 | Ensemble environment modeling using affine transform group
Yu Tsao 0001, Payton Lin, Ting-Yao Hu, Xugang Lu |
Speech Commun. | 3 |
| 2014 | Ensemble of machine learning algorithms for cognitive and physical speaker load detectionabstractWe present our methods and results on participating in the Interspeech 2014 Computational Paralinguistics ChallengE (ComParE) of which the goal is to detect certain type of load of a speaker using acoustic features. There are in total seven classification models contributing to our final prediction, namely, neural network with rectified linear unit and dropout (ReLUNet), conditional restricted Boltzmann machine (CRBM), logistic regression (LR), support vector machine (SVM), Gaussian discriminant analysis (GDA), k-nearest neighbors (KNN), and random forest (RF). When linearly blending the predictions of these models, we are able to get significant improvements over the challenge baseline. Index Terms: Physical Load Detection, Cognitive Load Detection, Neural Network, Classification Models How Jing, Ting-Yao Hu, Hung-Shin Lee, Wei-Chen Chen, Chi-Chun Lee, Yu Tsao 0001, Hsin-Min Wang |
INTERSPEECH | 2 |
| 2014 | Incorporating local information of the acoustic environments to MAP-based feature compensation and acoustic model adaptationabstractThe maximum a posteriori (MAP) criterion is popularly used for feature compensation (FC) and acoustic model adaptation (MA) to reduce the mismatch between training and testing data sets. MAP-based FC and MA require prior densities of mapping function parameters, and designing suitable prior densities plays an important role in obtaining satisfactory performance. In this paper, we propose to use an environment structuring framework to provide suitable prior densities for facilitating MAP-based FC and MA for robust speech recognition. The framework is constructed in a two-stage hierarchical tree structure using environment clustering and partitioning processes. The constructed framework is highly capable of characterizing local information about complex speaker and speaking acoustic conditions. The local information is utilized to specify hyper-parameters in prior densities, which are then used in MAP-based FC and MA to handle the mismatch issue. We evaluated the proposed framework on Aurora-2, a connected digit recognition task, and Aurora-4, a large vocabulary continuous speech recognition (LVCSR) task. On both tasks, experimental results showed that with the prepared environment structuring framework, we could obtain suitable prior densities for enhancing the performance of MAP-based FC and MA. Yu Tsao 0001, Xugang Lu, Paul R. Dixon, Ting-Yao Hu, Shigeki Matsuda, Chiori Hori |
Comput. Speech Lang. | 4 |
| 2013 | Ensemble of machine learning and acoustic segment model techniques for speech emotion and autism spectrum disorders recognition
Hung-yi Lee, Ting-Yao Hu, How Jing, Yun-Fan Chang, Yu Tsao 0001, Yu-Cheng Kao, Tsang-Long Pao |
INTERSPEECH | 2 |
| 2012 | Discriminative Fuzzy Clustering Maximum a Posterior Linear Regression for Speaker AdaptationabstractWe propose a discriminative fuzzy clustering maximum a posterior linear regression (DFCMAPLR) model adaptation approach to compensate the acoustic mismatch due to speaker variability. The DFCMAPLR approach adopts the MAP criterion and a discriminative objective function to estimate shared affine transform and fuzzy weight sets, respectively. Then, through a linear combination of the calculated fuzzy weights and shared affine transforms, more specific affine transforms are formed for model adaptation. By incorporating the MAP criterion and the discriminative information, DFCMAPLR can calculate shared affine transforms reliably and enhance the discriminative power of the adapted acoustic model. Based on the experimental results on the ASTTEL200 Mandarin corpus, we verified that DFCMAPLR outperforms not only the conventional maximum likelihood linear regression (MLLR) but also the fuzzy clustering MLLR(FCMLLR), which estimates the shared affine transform and fuzzy weight sets both based on the maximum likelihood criterion. Moreover, when compared to the baseline result, DFCMAPLR provides a clear improvement of 9.86% (24.04% to 21.67%) relative average phone error rate (PER) reduction. Ting-Yao Hu, Yu Tsao 0001, Lin-Shan Lee |
INTERSPEECH | 1 |
| 2011 | Learning context-aware sparse representation for single image super-resolutionabstractThis paper presents a novel learning-based method for single image super-resolution (SR). Given an input low-resolution image and its image pyramid, we propose to perform context-constrained image segmentation and construct an image segment dataset with different context categories. By learning context-specific image sparse representation, our method aims to model the relationship between the interpolated image patches and their ground truth pixel values from different context categories via support vector regression (SVR). To synthesize the final SR output, we upsample the input image by bicubic interpolation, followed by the refinement of each image patch using the SVR model learned from the associated context category. Unlike prior learning-based SR methods, our approach does not require the reoccurrence of similar image patches (within or across image scales), and we do not need to collect training low and high-resolution image data in advance either. Empirical results show that our proposed method is quantitatively and qualitatively more effective than existing interpolation or learning-based SR approaches. Ming-Chun Yang, Chang-Heng Wang, Ting-Yao Hu, Yu-Chiang Frank Wang |
ICIP | 3 |
| 2011 | Learning of context-aware single image super-resolutionabstractWe propose a novel learning-based method for single image super-resolution (SR). Given a low-resolution input image and its image pyramid, we advance a context-constrained image segmentation to construct a super-pixel database with different context categories for learning purposes. By utilizing context-specific image sparse representation, our method aims at modeling the relationship between the interpolated image patches and their ground truth pixels from different context categories via support vector regression (SVR). To produce the final SR output, we upsample the low-resolution input, followed by the refinement of each image patch using the SVR models observed from the associated context categories. Unlike prior learning-based SR methods, our approach advances a self-learning technique and does not assume the reoccurrence of image patches (within or across image scales). We do not need to collect training low/high-resolution image data in advance either. Empirical results verify the effectiveness of our SR approach, which quantitatively and qualitatively outperforms existing interpolation or learning-based SR methods in most cases. Ming-Chun Yang, Ting-Yao Hu, Chang-Heng Wang, Yu-Chiang Frank Wang |
VCIP | 2 |
| 2010 | Automatic Transcription for Music with Two Timbres from Monaural Sound SourceabstractA new approach to automatic music transcription for music with two timbres from a monaural sound source is proposed in this paper. The system is mainly composed of two parts, a fundamental frequency detector and a timbre discriminator. In the fundamental frequency detector, the short time Fourier transform (STFT) and a peak detection algorithm are used. By combining the relative magnitude of the fundamental frequency and its harmonics, a characteristic timbre vector is computed. The timbre discrimination is done by the classification of the timbre vectors with the support vector machine (SVM). Particularly, by designing the training and classification procedure in SVM, the mixed-timbre signal can also be classified in this system. Using a small database of polyphonic music of two timbres as the testing input, a 73% hit rate is achieved by this system. Yuh Shyang Wang, Ting-Yao Hu, Shyh-Kang Jeng |
ISM | 2 |