Ricardo Gutierrez-Osuna

dblp:06/5593 · DBLP profile ↗
← Back
94ranked-venue papers
4as first author
22since 2021 · last 2025
0000-0003-2817-2085ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 47 · 1 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2Computer networks · 1
YearPublicationVenuePosition
2025 DarkStream: real-time speech anonymization with low latency
abstract
We propose DarkStream, a streaming speech synthesis model for real-time speaker anonymization. To improve content encoding under strict latency constraints, DarkStream combines a causal waveform encoder, a short lookahead buffer, and transformer-based contextual layers. To further reduce inference time, the model generates waveforms directly via a neural vocoder, thus removing intermediate mel-spectrogram conversions. Finally, DarkStream anonymizes speaker identity by injecting a GAN-generated pseudo-speaker embedding into linguistic features from the content encoder. Evaluations show our model achieves strong anonymization, yielding close to 50% speaker verification EER (near-chance performance) on the lazy-informed attack scenario, while maintaining acceptable linguistic intelligibility (WER within 9%). By balancing low-latency, robust privacy, and minimal intelligibility degradation, DarkStream provides a practical solution for privacy-preserving real-time speech communication.
Waris Quamer, Ricardo Gutierrez-Osuna
ASRU2
2025 Health-Driven Personalized Metabolic Models of Postprandial Glucose Responses to Mixed Meals
Anurag Das, Ghady Nasrallah, Sicong Huang 0002, Bobak Mortazavi, Ricardo Gutierrez-Osuna
BSN5
2025 Does Gamified Breath-Biofeedback Promote Adherence, Relaxation, and Skill Transfer in the Wild?
abstract
This paper investigates whether gamification of deep breathing (DB) exercises promotes relaxation, skill transfer, and adherence to treatment in ambulatory settings. We designed a game-biofeedback (GBF) intervention where users perform DB exercises while playing a video game, and the game adapts according to the user's breathing rate using negative reinforcement instrumental conditioning. As a control, we developed an interactive paced-breathing treatment (PACE) where users follow a visual signal with their breathing and touch. In a user study, 30 participants were randomly assigned to GBF or PACE, and were allowed to practice at their leisure over the course of three days. Results show that the GBF group practiced their treatments significantly more often, achieved better skill transfer at post-test, and obtained a higher increase in self-reported positivity and relaxation during treatment. Our findings suggest that the use of negative reinforcement coupled with a fun casual game can be used as an alternative tool to promote relaxation and improve adherence to stress management interventions.
Dennis Rodrigo Da Cunha Silva, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.2
2024 Macronutrient Constraints and Priors Improve Carbohydrate Predictions from Continuous Glucose Monitors
abstract
We propose an approach to estimate the macronu-trients in a meal automatically by analyzing the meal's glucose response using off-the-shelf wearable sensors (continuous glucose monitors). We rely on the fact that the shape of the glucose response to a meal depends on all the macronutrients in the meal, not just its carbohydrates (carbs). However, protein, fat, and fiber tend to affect the glucose response in similar ways, so recovering their individual amounts is numerically ill-conditioned. To address this problem, our approach compresses macronutrients into a latent variable that captures their correlated effects on glucose. Then, we train a machine learning model to predict the latent variable from the glucose response of a meal. Finally, we recover the amount of the original macronutrients by incorporating prior knowledge of how they co-occur in conventional meals. Using experimental data from 45 participants, we show that predicting carbs indirectly (through the latent variable) reduces the prediction error when compared to predicting carbs directly, i.e., without considering the protein and fats in the meal.
Anurag Das, Edmund Do, Namino Glanz, Wendy Bevier, Rony Santiago, David Kerr, Bobak Mortazavi, Ricardo Gutierrez-Osuna
BSN8
2024 End-To-End Streaming Model For Low-Latency Speech Anonymization
abstract
Speaker anonymization aims to conceal cues to speaker identity while preserving linguistic content. Current machine learning based approaches require substantial computational resources, hindering real-time streaming applications. To address these concerns, we propose a streaming model that achieves speaker anonymization with low latency. The system is trained in an end-to-end autoencoder fashion using a lightweight content encoder that extracts HuBERT-like information, a pretrained speaker encoder that extract speaker identity, and a variance encoder that injects pitch and energy information. These three disentangled representations are fed to a decoder that re-synthesizes the speech signal. We present evaluation results from two implementations of our system, a full model that achieves a latency of 230 ms, and a lite version (0.1 x in size) that further reduces latency to 66 ms while maintaining state-of-the-art performance in naturalness, intelligibility, and privacy preservation.
Waris Quamer, Ricardo Gutierrez-Osuna
SLT2
2024 Improving Mispronunciation Detection Using Speech Reconstruction
abstract
Training related machine learning tasks simultaneously can lead to improved performance on both tasks. Text- to-speech (TTS) and mispronunciation detection and diagnosis (MDD) both operate using phonetic information and we wanted to examine whether a boost in MDD performance can be by two tasks. We propose a network that reconstructs speech from the phones produced by the MDD system and computes a speech reconstruction loss. We hypothesize that the phones produced by the MDD system will be closer to the ground truth if the reconstructed speech sounds closer to the original speech. To test this, we first extract wav2vec features from a pre-trained model and feed it to the MDD system along with the text input. The MDD system then predicts the target annotated phones and then synthesizes speech from the predicted phones. The system is therefore trained by computing both a speech reconstruction loss as well as an MDD loss. Comparing the proposed systems against an identical system but without speech reconstruction and another state-of-the-art baseline, we found that the proposed system achieves higher mispronunciation detection and diagnosis (MDD) scores. On a set of sentences unseen during training, the and speaker verification simultaneously can lead to improve proposed system achieves higher MDD scores, which suggests that reconstructing the speech signal from the predicted phones helps the system generalize to new test sentences. We also tested whether the system can generate accented speech when the input phones have mispronunciations. Results from our perceptual experiments show that speech generated from phones containing mispronunciations sounds more accented and less intelligible than phones without any mispronunciations, which suggests that the system can identify differences in phones and generate the desired speech signal.
Anurag Das, Ricardo Gutierrez-Osuna
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Modeling the effect of non-exercise activity on peak post-prandial glucose in diabetes
abstract
The timing, intensity, and duration of postprandial exercise are important factors that reduce glucose excursions. When exercise is of moderate intensity, performed between 25 and 55 minutes after a meal, it results in greater attenuation of glucose. However, the potential glucose reduction for shorter-duration, non-exercise activity thermogenesis (NEAT) (such as activities of daily living) may also be beneficial, particularly in cases where exercise is neither feasible nor prudent. Therefore, we designed a system to capture blood glucose and activity intensity through internet of medical things devices and modeled the impact of the timing and duration of NEAT on peak glucose. This work designed a linear mixed effects model to evaluate the impact of NEAT on peak, postprandial glucose in a study of data captured on varied participants with or without diabetes. We found at least 25 minutes of NEAT starting 30 minutes after the meal most effectively reduced peak post-prandial glucose.Clinical Relevance— This work establishes the impact of NEAT on reducing post-prandial peak glucose in free-living environments as another method of controlling glucose surges
Edmund Do, Anurag Das, Namino Glanz, Wendy Bevier, Rony Santiago, David Kerr, Ricardo Gutierrez-Osuna, Bobak Mortazavi
BSN7
2023 Detection of glycemic excursions using morphological and time-domain ECG features
abstract
Managing diabetes often involves monitoring blood glucose in real time to detect excursions (e.g., hypoglycemia and hyperglycemia). Continuous glucose monitors (CGMs) are generally used for this purpose, but CGMs are both expensive and invasive (they require inserting a flexible needle under the skin). To address this issue, we examine whether non-invasive devices, such as electrocardiograms (ECG), can be used to predict glucose excursions. In particular, we consider two types of cardiac information: (1) heartbeat morphology, which generally requires ECG recordings, and (2) heartbeat timing, which can be obtained from inexpensive wrist-worn devices, such as fitness trackers. We use convolutional networks to analyze beat morphology, and recurrent networks and feature engineering to analyze the inter-beat interval (IBI) time series. Then, we validate individual models and their combinations on an experimental dataset containing ECG and CGM recordings for then young adults with type 1 diabetes. We find that beat morphology outperforms beat timing in hypoglycemia prediction, but the reverse happens for hyperglycemia prediction. In both prediction problems, combining morphology and time-domain information outperforms using each source of information independently.
Kathan Vyas, Carolina Villegas, Elizabeth Kubota-Mishra, Darpit Dave, Madhav Erraguntla, Gerard L. Coté, Daniel J. Desalvo, Siripoom McKay, Ricardo Gutierrez-Osuna
BSN9
2023 Decoupling Segmental and Prosodic Cues of Non-native Speech through Vector Quantization
Waris Quamer, Anurag Das, Ricardo Gutierrez-Osuna
INTERSPEECH3
2023 Towards Participant-Independent Stress Detection Using Instrumented Peripherals
abstract
Methods to measure work stress generally rely on subjective measures from questionnaires or require dedicated sensors that are cumbersome to wear and interfere with the task. To address this problem, we propose a method to detect stress unobtrusively using commodity devices (keyboards, mice) instrumented with pressure sensors. We propose a minimalist design that can be easily replicated by other researchers using off-the-shelf and low-cost hardware. We validate the design in a laboratory experiment that simulates office tasks and mild stressors while avoiding methodological limitations of previous studies. We compare stress-detection performance when using conventional features reported in the literature (keystroke dynamics, mouse trajectories) augmented with information from pressure sensors. Our results indicate that pressure provides additional information for stress discrimination; adding pressure information to keystroke dynamics and mouse trajectories improves classification performance by 6% and 3%, respectively. These results show how devices that are already part of the modern workplace may be used and enhanced to automatically and unobtrusively detect stress.
Dennis Rodrigo Da Cunha Silva, Zelun Wang, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.3
2022 Minimizing Residuals for Native-Nonnative Voice Conversion in a Sparse, Anchor-Based Representation of Speech
abstract
We present a dictionary-learning algorithm for reducing the sparse coding residual of an exemplar-based method for native-to-nonnative voice conversion (VC). The proposed algorithm iteratively updates the source and target speaker dictionaries to reduce both the residual and voice conversion error, thereby increasing synthesis quality. We evaluate the method on speech from the ARCTIC and L2-ARCTIC corpora and compare it to a baseline exemplar-based VC algorithm. The proposed algorithm significantly improves synthesis quality to more than double that of the baseline system while using two orders of magnitude fewer atoms. Additionally, the proposed algorithm significantly reduces both the VC error and the residual magnitude. We discuss the implications of the algorithm for broad exemplar-based VC systems.
Christopher Liberatore, Ricardo Gutierrez-Osuna
ICASSP2
2022 Joint Hypoglycemia Prediction and Glucose Forecasting via Deep Multi-Task Learning
abstract
We present a multitask learning approach to the problem of hypoglycemia (HG) prediction in diabetes. The approach is based on a state-of-the-art time series forecasting model, N-BEATS, and extends it by adding a classification task so that the model performs both glucose forecasting (i.e., predicting future glucose values) and HG prediction (i.e., probability of future HG events sometime within the prediction horizon). We also propose an alternative loss function that penalizes forecasting errors in the HG range. We evaluate the approach on a dataset containing over 1.6M recordings from 112 patients with type 1 diabetes who wore a continuous glucose monitor (CGM) for 90 days. Our results show that the classification branch significantly outperforms the forecasting branch on the problem of HG prediction, and that the new loss function is more effective at reducing forecasting errors in the HG range than multi-task learning.
Mu Yang, Darpit Dave, Madhav Erraguntla, Gerard L. Coté, Ricardo Gutierrez-Osuna
ICASSP5
2022 Zero-Shot Foreign Accent Conversion without a Native Reference
Waris Quamer, Anurag Das, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna
INTERSPEECH5
2022 Accentron: Foreign accent conversion to arbitrary non-native speakers using zero-shot learning
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna
Comput. Speech Lang.3
2022 Predicting the Macronutrient Composition of Mixed Meals From Dietary Biomarkers in Blood
abstract
Diet monitoring is an essential intervention component for a number of diseases, from type 2 diabetes to cardiovascular diseases. However, current methods for diet monitoring are burdensome and often inaccurate. In prior work, we showed that continuous glucose monitors (CGMs) may be used to predict meal macronutrients (e.g., carbohydrates, protein, fat) by analyzing the shape of the post-prandial glucose response. In this study, we examine a number of additional dietary biomarkers in blood by their ability to improve macronutrient prediction, compared to using CGMs alone. For this purpose, we conducted a nutritional study where (n = 10) participants consumed nine different mixed meals with varied but known macronutrient amounts, and we analyzed the concentration of 33 dietary biomarkers (including amino acids, insulin, triglycerides, and glucose) at various times post-prandially. Then, we built machine learning models to predict macronutrient amounts from (1) individual biomarkers and (2) their combinations. We find that the additional blood biomarkers provide complementary information, and more importantly, achieve lower normalized root mean squared error (NRMSE) for the three macronutrients (carbohydrates: 22.9%; protein: 23.4%; fat: 32.3%) than CGMs alone (carbohydrates: 28.9%, t(18) =1.64, p =0.060; protein: 46.4%, t(18) =5.38, p 0.001; fat: 40.0%, t(18) =2.09, p =0.025). Our main conclusion is that augmenting CGMs to measure these additional dietary biomarkers improves macronutrient prediction performance, and may ultimately lead to the development of automated methods to monitor nutritional intake. This work is significant to biomedical research as it provides a potential solution to the long-standing problem of diet monitoring, facilitating new interventions for a number of diseases.
Anurag Das, Bobak Mortazavi, Seyedhooman Sajjadi, Theodora Chaspari, Laura Ruebush, Nicolaas E. P. Deutz, Gerard L. Coté, Ricardo Gutierrez-Osuna
IEEE J. Biomed. Health Informatics8
2021 A Sparse Coding Approach to Automatic Diet Monitoring with Continuous Glucose Monitors
abstract
Measuring dietary intake is a major challenge in the management of chronic diseases. Current methods rely on self-report measures, which are cumbersome to obtain and often unreliable. This article presents an approach to estimate dietary intake automatically by analyzing the post-prandial glucose response (PPGR) of a meal, as measured with continuous glucose monitors. In particular, we propose a sparse-coding technique that can be used to estimate the amounts of macronutrients (carbohydrates, protein, fat) in a meal from the meal’s PPGR. We use Lasso regularization to represent the PPGR of a new meal as a sparse combination of PPGRs in a dictionary, then combine the sparse weights with the macronutrient amounts in the dictionary’s meals to estimate the macronutrients in the new meal. We evaluate the approach on a dataset containing nine standardized meals and their corresponding PPGRs, consumed by fifteen participants. The proposed technique consistently outperforms two baseline systems based on ridge regression and nearest-neighbors, in terms of correlation and normalized root mean square error of the predictions.
Anurag Das, Seyedhooman Sajjadi, Bobak Mortazavi, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Ricardo Gutierrez-Osuna
ICASSP8
2021 Towards The Development of Subject-Independent Inverse Metabolic Models
abstract
Diet monitoring is an important component of interventions in type 2 diabetes, but is time intensive and often inaccurate. To address this issue, we describe an approach to monitor diet automatically, by analyzing fluctuations in glucose after a meal is consumed. In particular, we evaluate three standardization techniques (baseline correction, feature normalization, and model personalization) that can be used to compensate for the large individual differences that exist in food metabolism. Then, we build machine learning models to predict the amounts of macronutrients in a meal from the associated glucose responses. We evaluate the approach on a dataset containing glucose responses for 15 participants who consumed 9 meals. Three techniques improve the accuracy of the models: subtracting the baseline glucose, performing z-score normalization, and scaling the amount of macronutrients by each individuals’ body mass index.
Seyedhooman Sajjadi, Anurag Das, Ricardo Gutierrez-Osuna, Theodora Chaspari, Projna Paromita, Laura Ruebush, Nicolaas E. P. Deutz, Bobak Mortazavi
ICASSP3
2021 Assessing Posterior-Based Mispronunciation Detection on Field-Collected Recordings from Child Speech Therapy Sessions
Adam Hair, Guanlong Zhao, Beena Ahmed, Kirrie J. Ballard, Ricardo Gutierrez-Osuna
Interspeech5
2021 An Exemplar Selection Algorithm for Native-Nonnative Voice Conversion
Christopher Liberatore, Ricardo Gutierrez-Osuna
Interspeech2
2021 Effects of Voice Type and Task on L2 Learners' Awareness of Pronunciation Errors
abstract
Research suggests learners may improve their second language (L2) pronunciation by imitating voices with similar acoustic profiles. However, previously reported improvements have been in suprasegmentals (prosodic features such as intonation). It remains unclear if voice similarity applies to L2 segmentals (consonants and vowels). To address this issue, this study investigates how voice similarity facilitates awareness of pronunciation errors, a necessary step in pronunciation improvement. In two experiments, advanced L2 learners identified their pronunciation errors by comparing their production to the production of a resynthesized model voice using learners’ voices as the base (Golden Speaker voice), or to an unfamiliar resynthesized voice with the same gender as the learner (Silver Speaker voice). In Experiment 1, L2 learners identified all syllables with vowel and consonant errors when comparing their production to the model voice. Their choices were compared to identifications by expert judges. In Experiment 2, learners were told how many errors the expert judges had identified before identifying the same number of errors. Results did not support facilitative effects of Golden Speaker voices in either experiment, but Experiment 2 resulted in higher identification percentages. Discussion of the challenges in self-identification of errors in relation to voice similarity are offered.
Alif Silpachai, Ivana Rehman, Taylor Anne Barriuso, John Levis, Evgeny Chukharev-Hudilainen, Guanlong Zhao, Ricardo Gutierrez-Osuna
Interspeech7
2021 Partial Reinforcement in Game Biofeedback for Relaxation Training
abstract
This paper investigates the effect of reinforcement schedules on biofeedback games for stress self-regulation. In particular, it examines whether partial reinforcement can improve resistance to extinction of relaxation behaviors, i.e., once biofeedback is removed. Namely, we compare two types of reinforcement schedules (partial and continuous) in a mobile biofeedback game that encourages players to slow their breathing during gameplay. The game uses a negative-reinforcement instrumental conditioning paradigm, removing an aversive stimulus (random actions in the game) if players slows down their breathing. We conducted an experimental trial with 24 participants to compare the two reinforcement schedules against a control condition. Our results indicate that partial reinforcement improves resistance to extinction, as measured by breathing rate and skin conductance post-treatment. In addition, based on linear regression and correlation analysis we found that participants in the partial reinforcement learned to slow their breathing at the same pace as those under continuous reinforcement. The article discusses the implications of these results and directions for future work.
Avinash Parnandi 0001, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.2
2021 Converting Foreign Accent Speech Without a Reference
abstract
Foreign accent conversion (FAC) is the problem of generating a synthetic voice that has the voice identity of a second-language (L2) learner and the pronunciation patterns of a native (L1) speaker. This synthetic voice has been referred to as a “golden-speaker” in the pronunciation-training literature. FAC is generally achieved by building a voice-conversion model that maps utterances from a source (L1) speaker onto the target (L2) speaker. As such, FAC requires that a reference utterance from the L1 speaker be available at synthesis time. This greatly restricts the application scope of the FAC system. In this work, we propose a “reference-free” FAC system that eliminates the need for reference L1 utterances at synthesis time, and transforms L2 utterances directly. The system is trained in two steps. First, a conventional FAC procedure is used to create a golden-speaker using utterances from a reference L1 speaker (which are then discarded) and the L2 speaker. Second, a pronunciation-correction model is trained to convert L2 utterances to match the golden-speaker utterances obtained in the first step. At synthesis time, the pronunciation-correction model directly transforms a novel L2 utterance into its golden-speaker counterpart. Our results show that the system reduces foreign accents in novel L2 utterances, achieving a 20.5% relative reduction in word-error-rate of an American English automatic speech recognizer and a 19% reduction in perceptual ratings of foreign accentedness obtained through listening tests. Over 73% of the listeners also rated golden-speaker utterances as having the same voice identity as the original L2 utterances.
Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Emotional Footprints of Email Interruptions
abstract
Working in an environment with constant interruptions is known to affect stress, but how do interruptions affect emotional expression? Emotional expression can have significant impact on interactions among coworkers. We analyzed the video of 26 participants who performed an essay task in a laboratory while receiving either continual email interruptions or receiving a single batch of email. Facial videos of the participants were run through a convolutional neural network to determine the emotional mix via decoding of facial expressions. Using a novel co-occurrence matrix analysis, we showed that with batched email, a neutral emotional state is dominant with sadness being a distant second, and with continual interruptions, this pattern is reversed, and sadness is mixed with fear. We discuss the implications of these results for how interruptions can impact employees' well-being and organizational climate.
Christopher Blank, Shaila Zaman, Amanveer Wesley, Panagiotis Tsiamyrtzis, Dennis Rodrigo Da Cunha Silva, Ricardo Gutierrez-Osuna, Gloria Mark, Ioannis Pavlidis
CHI6
2020 Understanding the Effect of Voice Quality and Accent on Talker Similarity
abstract
This paper presents a methodology to study the role of nonnative accents on talker recognition by humans. The methodology combines a state-of-the-art accent-conversion system to resynthesize the voice of a speaker with a different accent of her/his own, and a protocol for perceptual listening tests to measure the relative contribution of accent and voice quality on speaker similarity. Using a corpus of non-native and native speakers, we generated accent conversions in two different directions: non-native speakers with native accents, and native speakers with non-native accents. Then, we asked listeners to rate the similarity between 50 pairs of real or synthesized speakers. Using a linear mixed effects model, we find that (for our corpus) the effect of voice quality is five times as large as that of non-native accent, and that the effect goes away when speakers share the same (native) accent. We discuss the potential significance of this work in earwitness identification and sociophonetics.
Anurag Das, Guanlong Zhao, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna
INTERSPEECH5
2020 Improving the Speaker Identity of Non-Parallel Many-to-Many Voice Conversion with Adversarial Speaker Recognition
Shaojin Ding, Guanlong Zhao, Ricardo Gutierrez-Osuna
INTERSPEECH3
2020 Gaming Away Stress: Using Biofeedback Games to Learn Paced Breathing
abstract
Biofeedback games are an attractive alternative to standard techniques for learning short-term relaxation skills. In this paper, we present the design, implementation, and evaluation of three respiratory biofeedback games. To validate these games, we compared breathing rate across 100 male participants (23 years ± 3.2 years) playing biofeedback and audio pacing versions of these games as well as a paced breathing app. The games were placed between repeat runs of a cognitively stressful Stroop-based task and the impact of the games on breathing and cognitive performance in the task also assessed. Our results showed that 1) differences in gameplay did not impact player performance; 2) biofeedback not only led to better breath control during play but also during the subsequent cognitively stressful task; and 3) biofeedback led to better attentional-cognitive performance in the subsequent task. Our multi-game experiments show that using respiratory biofeedback in video games is an effective strategy to learn paced breathing—on par with the standalone technique of paced breathing—and to self-regulate stress levels in later stressful scenarios. Furthermore, owing to its entertainment value, our relaxation solution has the potential to be more engaging and accessible than standalone paced breathing, for use over longer durations.
M. Abdullah Zafar, Beena Ahmed, Rami G. Al Rihawi, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.4
2020 Learning Structured Sparse Representations for Voice Conversion
abstract
Sparse-coding techniques for voice conversion assume that an utterance can be decomposed into a sparse code that only carries linguistic contents, and a dictionary of atoms that captures the speakers' characteristics. However, conventional dictionary-construction and sparse-coding algorithms rarely meet this assumption. The result is that the sparse code is no longer speaker-independent, which leads to lower voice-conversion performance. In this paper, we propose a Cluster-Structured Sparse Representation (CSSR) that improves the speaker independence of the representations. CSSR consists of two complementary components: a Cluster-Structured Dictionary Learning module that groups atoms in the dictionary into clusters, and a Cluster-Selective Objective Function that encourages each speech frame to be represented by atoms from a small number of clusters. We conducted four experiments on the CMU ARCTIC corpus to evaluate the proposed method. In a first ablation study, results show that each of the two CSSR components enhances speaker independence, and that combining both components leads to further improvements. In a second experiment, we find that CSSR uses increasingly larger dictionaries more efficiently than phoneme-based representations by allowing finer-grained decompositions of speech sounds. In a third experiment, results from objective and subjective measurements show that CSSR outperforms prior voice-conversion methods, improving the acoustic quality of the synthesized speech while retaining the target speaker's voice identity. Finally, we show that the CSSR captures latent (i.e., phonetic) information in the speech signal.
Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Evaluating Automatic Speech Recognition for Child Speech Therapy Applications
abstract
Automatic speech recognition (ASR) technology can be a useful tool in mobile apps for child speech therapy, empowering children to complete their practice with limited caregiver supervision. However, little is known about the feasibility of performing ASR on mobile devices, particularly when training data is limited. In this study, we investigated the performance of two low-resource ASR systems on disordered speech from children. We compared the open-source PocketSphinx (PS) recognizer using adapted acoustic models and a custom template-matching (TM) recognizer. TM and the adapted models significantly out-perform the default PS model. On average, maximum likelihood linear regression and maximum a posteriori adaptation increased PS accuracy from 59.4% to 63.8% and 80.0%, respectively, suggesting that the models successfully captured speaker-specific word production variations. TM reached a mean accuracy of 75.8%
Adam Hair, Kirrie J. Ballard, Beena Ahmed, Ricardo Gutierrez-Osuna
ASSETS4
2019 Email Makes You Sweat: Examining Email Interruptions and Stress Using Thermal Imaging
abstract
Workplace environments are characterized by frequent interruptions that can lead to stress. However, measures of stress due to interruptions are typically obtained through self-reports, which can be affected by memory and emotional biases. In this paper, we use a thermal imaging system to obtain objective measures of stress and investigate personality differences in contexts of high and low interruptions. Since a major source of workplace interruptions is email, we studied 63 participants while multitasking in a controlled office environment with two different email contexts: managing email in batch mode or with frequent interruptions. We discovered that people who score high in Neuroticism are significantly more stressed in batching environments than those low in Neuroticism. People who are more stressed finish emails faster. Last, using Linguistic Inquiry Word Count on the email text, we find that higher stressed people in multitasking environments use more anger in their emails. These findings help to disambiguate prior conflicting results on email batching and stress.
Fatema Akbar 0001, A. Elvan Bayraktaroglu, Pradeep Buddharaju, Dennis Rodrigo Da Cunha Silva, Ge Gao 0001, Ted Grover, Ricardo Gutierrez-Osuna, Nathan Cooper Jones, Gloria Mark, Ioannis Pavlidis, Kevin M. Storer, Zelun Wang, Amanveer Wesley, Shaila Zaman
CHI7
2019 Group Latent Embedding for Vector Quantized Variational Autoencoder in Non-Parallel Voice Conversion
Shaojin Ding, Ricardo Gutierrez-Osuna
INTERSPEECH2
2019 Foreign Accent Conversion by Synthesizing Speech from Phonetic Posteriorgrams
Guanlong Zhao, Shaojin Ding, Ricardo Gutierrez-Osuna
INTERSPEECH3
2019 Golden speaker builder - An interactive tool for pronunciation training
Shaojin Ding, Christopher Liberatore, Sinem Sonsaat, Ivana Lucic, Alif Silpachai, Guanlong Zhao, Evgeny Chukharev-Hudilainen, John Levis, Ricardo Gutierrez-Osuna
Speech Commun.9
2019 Visual Biofeedback and Game Adaptation in Relaxation Skill Transfer
abstract
This paper compares the effectiveness of two biofeedback mechanisms to promote acquisition and transfer of deep-breathing skills using a casual videogame. The first biofeedback mechanism, game adaptation, delivers respiratory information by altering an internal parameter of the game; the second, visual biofeedback, displays respiratory information explicitly without altering the game. We conduct a user study that examines visual biofeedback and game adaptation as independent variables with electrodermal activity, heart rate variability, and respiration as dependent variables. In particular, we evaluate these forms of biofeedback by their ability to facilitate acquisition of relaxation skills and promote skill transfer to subsequent stressful tasks. Our results indicate that game adaptation promotes skill acquisition and transfer more effectively than visual biofeedback, but that a combination of the two outperforms either in isolation. Combining visual and game biofeedback also results in faster learning of deep-breathing skills than either channel alone. Our study suggests that the two forms of biofeedback play different roles, with game adaptation being more effective in encouraging deep breathing, and visual channels helping players maintain the target breathing rate.
Avinash Parnandi 0001, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.2
2019 Using Phonetic Posteriorgram Based Frame Pairing for Segmental Accent Conversion
abstract
Accent conversion (AC) aims to transform non-native utterances to sound as if the speaker had a native accent. This can be achieved by mapping source speech spectra from a native speaker into the acoustic space of the target non-native speaker. In prior work, we proposed an AC approach that matches frames between the two speakers based on their acoustic similarity after compensating for differences in vocal tract length. In this paper, we propose a new approach that matches frames between the two speakers based on their phonetic (rather than acoustic) similarity. Namely, we map frames from the two speakers into a phonetic posteriorgram using speaker-independent acoustic models trained on native speech. We thoroughly evaluate the approach on a speech corpus containing multiple native and non-native speakers. The proposed algorithm outperforms the prior approach, improving ratings of acoustic quality (22% increase in mean opinion score) and native accent (69% preference) while retaining the voice quality of the non-native speaker. Furthermore, we show that the approach can be used in the reverse conversion direction, i.e., generating speech with a native speaker's voice quality and a non-native accent. Finally, we show that this approach can be applied to non-parallel training data, achieving the same accent conversion performance.
Guanlong Zhao, Ricardo Gutierrez-Osuna
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Apraxia world: a speech therapy game for children with speech sound disorders
abstract
This paper presents Apraxia World, a remote therapy tool for speech sound disorders that integrates speech exercises into an engaging platformer-style game. In Apraxia World, the player controls the avatar with virtual buttons/joystick, whereas speech input is associated with assets needed to advance from one level to the next. We tested performance and child preference of two strategies for delivering speech exercises: during each level, and after it. Most children indicated that doing exercises after completing each level was less disruptive and preferable to doing exercises scattered through the level. We also found that children liked having perceived control over the game (character appearance, exercise behavior). Our results indicate that (i) a familiar style of game successfully engages children, (ii) speech exercises function well when decoupled from game control, and (iii) children are willing to complete required speech exercises while playing a game they enjoy.
Adam Hair, Penelope Monroe, Beena Ahmed, Kirrie J. Ballard, Ricardo Gutierrez-Osuna
IDC5
2018 Voice Conversion Through Residual Warping in a Sparse, Anchor-Based Representation of Speech
abstract
In previous work we presented a Sparse, Anchor-Based Representation of speech (SABR) that uses phonemic “anchors” to represent an utterance with a set of sparse non-negative weights. SABR is speaker-independent: combining weights from a source speaker with anchors from a target speaker can be used for voice conversion. Here, we present an extension of the original SABR that significantly improves voice conversion synthesis. Namely, we take the residual signal from the SABR decomposition of the source speaker's utterance, and warp it to the target speaker's space using a weighted warping function learned from pairs of source-target anchors. Using subjective and objective evaluations, we examine the performance of adding the warped residual (SABR+Res) to the original synthesis (SABR). Specifically, listeners rated SABR+Res with an average mean opinion score (MOS) of 3.6, a significant improvement compared to 2.2 MOS for SABR alone (p <; 0.01) and 2.5 MOS for a baseline GMM method (p <; 0.01). In an XAB speaker identity test, listeners correctly identified the identity of SABR+Res (81 %) and SABR (84%) as frequently as a GMM method (82%) (p = 0.70, P = 0.35). These results indicate that adding the warped residual can dramatically improve synthesis while retaining the desirable independent qualities of SABR models.
Christopher Liberatore, Guanlong Zhao, Ricardo Gutierrez-Osuna
ICASSP3
2018 Accent Conversion Using Phonetic Posteriorgrams
abstract
Accent conversion (AC) aims to transform non-native speech to sound as if the speaker had a native accent. This can be achieved by mapping source spectra from a native speaker into the acoustic space of the non-native speaker. In prior work, we proposed an AC approach that matches frames between the two speakers based on their acoustic similarity after compensating for differences in vocal tract length. In this paper, we propose an approach that matches frames between the two speakers based on their phonetic (rather than acoustic) similarity. Namely, we map frames from the two speakers into a phonetic posteriorgram using speaker-independent acoustic models trained on native speech. We evaluate the proposed algorithm on a corpus containing multiple native and non-native speakers. Compared to the previous AC algorithm, the proposed algorithm improves the ratings of acoustic quality (20% increase in mean opinion score) and native accent (69% preference) while retaining the voice identity of the non-native speaker.
Guanlong Zhao, Sinem Sonsaat, John Levis, Evgeny Chukharev-Hudilainen, Ricardo Gutierrez-Osuna
ICASSP5
2018 Learning Structured Dictionaries for Exemplar-based Voice Conversion
Shaojin Ding, Christopher Liberatore, Ricardo Gutierrez-Osuna
INTERSPEECH3
2018 Improving Sparse Representations in Exemplar-Based Voice Conversion with a Phoneme-Selective Objective Function
Shaojin Ding, Guanlong Zhao, Christopher Liberatore, Ricardo Gutierrez-Osuna
INTERSPEECH4
2018 L2-ARCTIC: A Non-native English Speech Corpus
abstract
In this paper, we introduce L2-ARCTIC, a speech corpus of non-native English that is intended for research in voice conversion, accent conversion, and mispronunciation detection. This initial release includes recordings from ten non-native speakers of English whose first languages (L1s) are Hindi, Korean, Mandarin, Spanish, and Arabic, each L1 containing recordings from one male and one female speaker. Each speaker recorded approximately one hour of read speech from the Carnegie Mellon University ARCTIC prompts, from which we generated orthographic and forced-aligned phonetic transcriptions. In addition, we manually annotated 150 utterances per speaker to identify three types of mispronunciation errors: substitutions, deletions, and additions, making it a valuable resource not only for research in voice conversion and accent conversion but also in computer-assisted pronunciation training. The corpus is publicly accessible at https://psi.engr.tamu.edu/l2-arctic-corpus/.
Guanlong Zhao, Sinem Sonsaat, Alif Silpachai, Ivana Lucic, Evgeny Chukharev-Hudilainen, John Levis, Ricardo Gutierrez-Osuna
INTERSPEECH7
2018 BioPad: Leveraging off-the-Shelf Video Games for Stress Self-Regulation
abstract
This paper presents an approach to use commercial videogames for biofeedback training. It consists of intercepting signals from the game controller and adapting them in real-time based on physiological measurements from the player. We present three sample implementations and a case study for teaching stress self-regulation via an immersive car racing game. We use a crossover gaming device to manipulate controller signals, and a respiratory sensor to monitor the players' breathing rate. We then alter the speed of the car to encourage slow deep breathing, in this way, allowing players to reduce their arousal while playing the game. We evaluate the approach against an alternative form of biofeedback that uses a graphic overlay to convey physiological information, and a control condition (playing the game without biofeedback). Experimental results show that our approach can promote deep breathing during gameplay, and also during a subsequent task, once biofeedback is removed. Our results also indicate that delivering biofeedback through subtle changes in gameplay can be as effective as delivering them directly through a visual display. These results open the possibility to develop low-cost and engaging biofeedback interventions using a variety of commercial videogames to promote adherence.
Zelun Wang, Avinash Parnandi 0001, Ricardo Gutierrez-Osuna
IEEE J. Biomed. Health Informatics3
2017 Speed-Accuracy Tradeoffs for Detecting Sign Language Content in Video Sharing Sites
abstract
Sign language is the primary medium of communication for many people who are deaf or hard of hearing. Members of this community access online sign language (SL) content posted on video sharing sites to stay informed. Unfortunately, locating SL videos can be difficult since the text-based search on video sharing sites is based on metadata rather than on the video content. Low cost or real-time video classification techniques would be invaluable for improving access to this content. Our prior work developed a technique to identify SL content based on video features alone but is computationally expensive. Here we describe and evaluate three optimization strategies that have the potential to reduce the computation time without overly impacting precision and recall. Two optimizations reduce the cost of face-detection, whereas the third focuses on analyzing shorter segments of the video. Our results identify a combination of these techniques that yields a 96% reduction in computation time while losing only 1% in F1 score. To further reduce computation, we additionally explore a keyframe-based approach that achieves comparable recall but lower precision than the above techniques, making it appropriate as an early filter in a staged classifier.
Frank M. Shipman III, Satyakiran Duggina, Caio D. D. Monteiro, Ricardo Gutierrez-Osuna
ASSETS4
2017 Exemplar selection methods in voice conversion
abstract
Exemplar-based methods for voice conversion often use a large number of randomly-selected exemplars to ensure good coverage. As a result, the factorization step can be costly. This paper presents two algorithms that can be used to construct compact sets of exemplars. The first algorithm uses a forward selection procedure to build the exemplar set sequentially, selecting exemplar pairs that minimize the joint reconstruction error on source and target frames. The second algorithm uses a backward elimination procedure to remove exemplars that contribute the least to the factorization. We evaluate both selection strategies on voice conversion tasks using the ARCTIC corpus. Our results using objective measures and subjective listening tests show that both strategies can significantly reduce the size of the exemplar set (five-fold, in our experiments) while achieving the same performance on voice conversion.
Guanlong Zhao, Ricardo Gutierrez-Osuna
ICASSP2
2017 Physiological Modalities for Relaxation Skill Transfer in Biofeedback Games
abstract
We present an adaptive biofeedback game for teaching self-regulation of stress. Our approach consists of monitoring the user's physiology during gameplay and adapting the game using a positive feedback loop that rewards relaxing behaviors and penalizes states of high arousal. We evaluate the approach using a casual game under three biofeedback modalities: electrodermal activity, heart rate variability, and breathing rate. The three biosignals can be measured noninvasively with wearable sensors, and represent different degrees of voluntary control and selectivity toward arousal. We conducted an experiment trial with 25 participants to compare the three modalities against a standard treatment (deep breathing) and a control condition (the game without biofeedback). Our results indicate that breathing-based game biofeedback is more effective in inducing relaxation during treatment than the other four groups. Participants in this group also showed greater retention of the relaxation skills (without biofeedback) during a subsequent stressor.
Avinash Parnandi 0001, Ricardo Gutierrez-Osuna
IEEE J. Biomed. Health Informatics2
2016 Classification of bisyllabic lexical stress patterns in disordered speech using deep learning
abstract
Technology-based therapy tools can be of great benefit to children with developmental speech disabilities as they typically require sustained practice with a speech therapist for several years. Towards this aim, over the past 4 years we have developed speech processing tools to automatically detect common errors in disordered speech. This paper presents an automated technique to identify incorrect lexical stress. Specifically, we describe a deep neural network (DNN) that can be used to classify the four different bisyllabic stress patterns: strong-weak (SW), weak-strong (WS), strong-strong (SS) and weak-weak (WW). We derive input features for the DNN from the duration, pitch, intensity and spectral energy on each of the two consecutive syllables. Using these features, we achieve 93% correct classification between SW/WS stress patterns and 88% correct classification of the four bisyllabic patterns on speech from typically developing children, while we obtain 73.4% classification between SW/WS in disordered speech. These figures represent a two-fold reduction in error rates compared to our prior work, which used a DNN with differential features from consecutive syllables.
Mostafa Shahin, Ricardo Gutierrez-Osuna, Beena Ahmed
ICASSP2
2016 Comparing Articulatory and Acoustic Strategies for Reducing Non-Native Accents
Sandesh Aryal, Ricardo Gutierrez-Osuna
INTERSPEECH2
2016 Generating Gestural Scores from Acoustics Through a Sparse Anchor-Based Representation of Speech
Christopher Liberatore, Ricardo Gutierrez-Osuna
INTERSPEECH2
2016 Detecting and Identifying Sign Languages through Visual Features
abstract
The popularity of video sharing sites has encouraged the creation and distribution of sign language (SL) content. Unfortunately, locating SL videos on a desired topic is not a straightforward task. Retrieval depends on the existence and correctness of metadata to indicate that the video contains SL. This problem gets worse when considering a particular type of sign language (e.g. American Sign Language - ASL, British Sign Language - BSL, French Sign Language – LSF, etc.), where metadata needs to be even more specific. To address this problem, we have expanded a previous SL classifier to distinguish videos in different SLs. The new classifier achieves an F1 score of 98% when discriminating between BSL and LSF videos with static backgrounds, and a 70% F1 score when distinguishing between ASL and BSL videos found on popular video sharing sites. Such accuracy with visual features alone is possible when comparing languages with one-handed and two-handed manual alphabets.
Caio D. D. Monteiro, Christy Maria Mathew, Ricardo Gutierrez-Osuna, Frank M. Shipman III
ISM3
2016 Data driven articulatory synthesis with deep neural networks
Sandesh Aryal, Ricardo Gutierrez-Osuna
Comput. Speech Lang.2
2016 ReBreathe: A Calibration Protocol that Improves Stress/Relax Classification by Relabeling Deep Breathing Relaxation Exercises
abstract
Training stress-prediction models is challenging due to the difficulty in reliably eliciting stress and relaxation responses in participants. For example, a task intended to elicit a relaxation response (e.g., deep breathing) can have the opposite effect depending on the participant's appraisal of and familiarity with the exercise. Including such instances in a training set undermines the accuracy of the resulting prediction model. This paper presents a technique, ReBreathe, to identify such instances based on respiratory patterns and determine their accurate stress/relax labels. We compared this relabeling approach against two labeling techniques: 1) nominal labels obtained from the experimental protocol and 2) labels obtained from subjective assessments. We then trained generalized estimating equation regression models to predict the resulting stress/relax labels from measures of heart rate variability and electrodermal activity. Training the model using protocol labels achieved a classification rate of 0.53 on participants not included in the training set. Relabeling the exercises based on each participant's subjective ratings increased classification rates but only marginally (0.61). In contrast, relabeling the exercises based on respiratory patterns increased classification rates to 0.88, or a four-fold reduction in error rates. These results illustrate the unreliability of protocol and subjective labels during stress/relax exercises and the potential benefits of ReBreathe.
Beena Ahmed, Hira Khan, Jongyong Choi, Ricardo Gutierrez-Osuna
IEEE Trans. Affect. Comput.4
2015 Automatic Assessment of OCR Quality in Historical Documents
abstract
Mass digitization of historical documents is a challenging problem for optical character recognition (OCR) tools. Issues include noisy backgrounds and faded text due to aging, border/marginal noise, bleed-through, skewing, warping, as well as irregular fonts and page layouts. As a result, OCR tools often produce a large number of spurious bounding boxes (BBs) in addition to those that correspond to words in the document. This paper presents an iterative classification algorithm to automatically label BBs (i.e., as text or noise) based on their spatial distribution and geometry. The approach uses a rule-base classifier to generate initial text/noise labels for each BB, followed by an iterative classifier that refines the initial labels by incorporating local information to each BB, its spatial location, shape and size. When evaluated on a dataset containing over 72,000 manually-labeled BBs from 159 historical documents, the algorithm can classify BBs with 0.95 precision and 0.96 recall. Further evaluation on a collection of 6,775 documents with ground-truth transcriptions shows that the algorithm can also be used to predict document quality (0.7 correlation) and improve OCR transcriptions in 85% of the cases.
Ricardo Gutierrez-Osuna, Matthew Christy, Boris Capitanu, Loretta Auvil, Liz Grumbach, Richard Furuta, Laura Mandell
AAAI2
2015 Joint optimization of anatomical and gestural parameters in a physical vocal tract model
abstract
We describe a method for adapting a physical vocal tract model's anatomical and gestural parameters using acoustic information to match a target speaker. Physical vocal tract models are hard to adjust to match a speaker, as doing so requires information which is difficult to capture, such as X-Ray or MRI information. We propose an analysis-by-synthesis approach to adjust the parameters of the VocalTractLab (VTL) physical vocal tract model, optimizing on an acoustic distance objective function. We compare our method with one which does not adjust anatomy parameters, just gestural parameters, and find that the proposed method results in a net improvement. We also test our method's ability to recreate a synthetic speaker for which the ground truth parameters are known, and find that the method can reproduce the speaker if parameters pertaining to teeth and lips are fixed.
Christopher Liberatore, Ricardo Gutierrez-Osuna
ICASSP2
2015 Articulatory-based conversion of foreign accents with deep neural networks
Sandesh Aryal, Ricardo Gutierrez-Osuna
INTERSPEECH2
2015 SABR: sparse, anchor-based representation of the speech signal
abstract
We present SABR (Sparse, Anchor-Based Representation), an analysis technique to decompose the speech signal into speaker-dependent and speaker-independent components. Given a collection of utterances for a particular speaker, SABR uses the centroid for each phoneme as an acoustic “anchor, ” then applies Lasso regularization to represent each speech frame as a sparse non-negative combination of the anchors. We illustrate the performance of the method on a speaker-independent phoneme recognition task and a voice conversion task. Using a linear classifier, SABR weights achieve significantly higher phoneme recognition rates than Mel frequency Cepstral coefficients. SABR weights can also be used directly to perform accent conversion without the need to train a speaker-to-speaker regression model.
Christopher Liberatore, Sandesh Aryal, Zelun Wang, Seth Polsley, Ricardo Gutierrez-Osuna
INTERSPEECH5
2015 Tabby Talks: An automated tool for the assessment of childhood apraxia of speech
Mostafa Shahin, Beena Ahmed, Avinash Parnandi 0001, Virendra Karappa, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna
Speech Commun.7
2014 Accent conversion through cross-speaker articulatory synthesis
abstract
Accent conversion (AC) seeks to transform second-language (L2) utterances to appear as if produced with a native (L1) accent. In the acoustic domain, AC is difficult due to the complex interaction between linguistic content and voice quality. Alternatively, AC can be performed in the articulatory domain by building a mapping from L2 articulators to L2 acoustics, and then driving the model with L1 articulators. However, collecting articulatory data for each L2 learner is impractical. Here we propose an approach that avoids this expensive step. Our method builds a cross-speaker forward mapping (CSFM) to generate L2 acoustic observations directly from L1 articulatory trajectories. We evaluated the CSFM against a baseline articulatory synthesizer trained with L2 articulators. Subjective listening tests show that both methods perform comparably in terms of accent reduction and ability to preserve the voice quality of the L2 speaker, with only a small impact in acoustic quality.
Sandesh Aryal, Ricardo Gutierrez-Osuna
ICASSP2
2014 Can voice conversion be used to reduce non-native accents?
abstract
Voice-conversion (VC) techniques aim to transform utterances from a source speaker to sound as if a target speaker had produced them. For this reason, VC is generally ill-suited for accent-conversion (AC) purposes, where the goal is to capture the regional accent of the source while preserving the voice quality of the target. In this paper, we propose a modification of the conventional training process for VC that allows it to perform as an AC transform. Namely, we pair source and target vectors based not on their ordering within a parallel corpus, as is commonly done in VC, but based on their linguistic similarity. We validate the approach on a corpus containing native-accented and Spanish-accented utterances, and compare it against conventional VC through a series of listening tests. We also analyze whether phonological differences between the two languages (Spanish and American English) help predict the performance of the two methods.
Sandesh Aryal, Ricardo Gutierrez-Osuna
ICASSP2
2014 Normalization of articulatory data through Procrustes transformations and analysis-by-synthesis
abstract
We describe and compare three methods that can be used to normalize articulatory data across speakers. The methods seek to explain systematic anatomical differences between a source and target speaker without modifying the articulatory velocities of the source speaker. The first method is the classical Procrustes transform, which allows for a global translation, rotation, and scaling of articulator positions. We present an extension to the Procrustes transform that allows independent translations of each articulator. The additional parameters provide a 35% increase in articulatory similarity between pairs of speakers when compared to classical Procrustes. The proposed extension is finally coupled with a data-driven articulatory synthesizer in an analysis-by-synthesis loop to select model parameters that best explain the predicted acoustic (rather than articulatory) differences. This normalization method is able to increase acoustic similarity between source and the target speaker by 34%. However, it also reduces articulatory similarity by 22%, which suggest that improvements in acoustic similarity do not necessarily require an increase in articulatory similarity.
Daniel Felps, Sandesh Aryal, Ricardo Gutierrez-Osuna
ICASSP3
2014 Detection of sign-language content in video through polar motion profiles
abstract
Locating sign language (SL) videos on video sharing sites (e.g., YouTube) is challenging because search engines generally do not use the visual content of videos for indexing. Instead, indexing is done solely based on textual content (e.g., title, description, metadata). As a result, untagged SL videos do not appear in the search results. In this paper, we present and evaluate a classification approach to detect SL videos based on their visual content. The approach uses an ensemble of Haar-based face detectors to define regions of interest (ROI), and a background model to segment movements in the ROI. The two-dimensional (2D) distribution of foreground pixels in the ROI is then reduced to two 1D polar motion profiles by means of a polar-coordinate transformation, and then classified by means of an SVM. When evaluated on a dataset of user-contributed YouTube videos, the approach achieves 81% precision and 94% recall.
Virendra Karappa, Caio D. D. Monteiro, Frank M. Shipman III, Ricardo Gutierrez-Osuna
ICASSP4
2014 A comparison of GMM-HMM and DNN-HMM based pronunciation verification techniques for use in the assessment of childhood apraxia of speech
abstract
This paper introduces a pronunciation verification method to be used in an automatic assessment therapy tool of child disordered speech. The proposed method creates a phonebased search lattice that is flexible enough to cover all probable mispronunciations. This allows us to verify the correctness of the pronunciation and detect the incorrect phonemes produced by the child. We compare between two different acoustic models, the conventional GMM-HMM and the hybrid DNN-HMM. Results show that the hybrid DNNHMM outperforms the conventional GMM-HMM for all experiments on both normal and disordered speech. The total correctness accuracy of the system at the phoneme level is above 85% when used with disordered speech.
Mostafa Shahin, Beena Ahmed, Jacqueline McKechnie, Kirrie J. Ballard, Ricardo Gutierrez-Osuna
INTERSPEECH5
2014 Context-sensitive intra-class clustering
Yingwei Yu, Ricardo Gutierrez-Osuna, Yoonsuck Choe
Pattern Recognit. Lett.2
2013 Contactless Measurement of Heart Rate Variability from Pupillary Fluctuations
abstract
The ability to measure a person's physiological parameters in a contact less fashion (i.e., without attaching electrodes to the skin) has tremendous potential in a number of applications, from affective interfaces to healthcare delivery. In this paper, we present a proof-of-concept method for measuring one such vital parameter, heart rate variability (HRV), in a contact less fashion from spontaneous fluctuations in pupillary diameter. Our approach uses a remote eye tracker for imaging and an integro-differential algorithm for segmenting the pupil-iris boundary. We then estimate HRV from the relative distribution of energy in the low frequency (0.04 to 0.15 Hz) and high frequency (0.15 to 0.4 Hz) bands of the power spectrum of the time series of pupillary fluctuations. We validated the method under a range of breathing conditions and under different illumination levels. Our results show a high degree of agreement between our pupillary estimate of HRV and ground truth measurements from an ECG-grade heart rate monitor. These results support the feasibility of estimating HRV in a non-contact, non-invasive fashion.
Avinash Parnandi 0001, Ricardo Gutierrez-Osuna
ACII2
2013 A Control-Theoretic Approach to Adaptive Physiological Games
abstract
We present an adaptive biofeedback game that aims to maintain the player's arousal level by monitoring physiological signals. We use concepts from control theory to model the interaction between human physiology and game difficulty during game play. We validate the approach on a car-racing game with real-time adaptive game mechanics. Specifically, we use car speed, road visibility, and steering jitter as three mechanisms to manipulate game difficulty. We propose quantitative measures to characterize the effectiveness of these game adaptations in manipulating the player's arousal. For this purpose, we use electro dermal activity (EDA) as a physiological correlate of arousal. Experimental trials with 20 subjects in both open-loop (no feedback) and closed-loop (negative feedback) conditions show statistically significant differences among the three game mechanics in terms of their effectiveness. Specifically, manipulating car speed provides higher arousal levels than changing road visibility or vehicle steering. Finally, we discuss the theoretical and practical implications of our approach.
Avinash Parnandi 0001, Youngpyo Son, Ricardo Gutierrez-Osuna
ACII3
2013 Architecture of an automated therapy tool for childhood apraxia of speech
abstract
We present a multi-tier system for the remote administration of speech therapy to children with apraxia of speech. The system uses a client-server architecture model and facilitates task-oriented remote therapeutic training in both in-home and clinical settings. Namely, the system allows a speech therapist to remotely assign speech production exercises to each child through a web interface, and the child to practice these exercises on a mobile device. The mobile app records the child's utterances and streams them to a back-end server for automated scoring by a speech-analysis engine. The therapist can then review the individual recordings and the automated scores through a web interface, provide feedback to the child, and adapt the training program as needed. We validated the system through a pilot study with children diagnosed with apraxia of speech, and their parents and speech therapists. Here we describe the overall client-server architecture, middleware tools used to build the system, the speech-analysis tools for automatic scoring of recorded utterances, and results from the pilot study. Our results support the feasibility of the system as a complement to traditional face-to-face therapy through the use of mobile tools and automated speech analysis algorithms.
Avinash Parnandi 0001, Virendra Karappa, Youngpyo Son, Mostafa Shahin, Jacqueline McKechnie, Kirrie J. Ballard, Beena Ahmed, Ricardo Gutierrez-Osuna
ASSETS8
2013 Articulatory inversion and synthesis: Towards articulatory-based modification of speech
abstract
Certain speech modifications, such as changes in foreign/regional accents or articulatory styles, are performed more effectively in the articulatory domain than in the acoustic domain. Though measuring articulators is cumbersome, articulatory parameters may be estimated from acoustics through inversion. In this paper, we study the impact on synthesis quality when articulators predicted from acoustics are used in articulatory synthesis. For this purpose, we trained a GMM articulatory synthesizer and drove it with articulators predicted with an RBF-based inversion model. Using inverted instead of measured articulators degraded synthesis quality, as measured through Mel cepstral distortion and subjective tests. However, retraining the synthesizer with predicted articulators not only reversed the effect of errors introduced during inversion but also improved synthesis quality relative to using measured articulators. These results suggest that inverted articulators do not compromise synthesis quality, and open up the possibility of performing speech modification in the articulatory domain through inversion.
Sandesh Aryal, Ricardo Gutierrez-Osuna
ICASSP2
2013 Active analysis of chemical mixtures with multi-modal sparse non-negative least squares
abstract
New sensor technologies such as Fabry-Pérot interferometers (FPI) offer low-cost and portable alternatives to traditional infrared absorption spectroscopy for chemical analysis. However, with FPIs the absorption spectrum has to be measured one wavelength at a time. In this work, we propose an active-sensing framework to select a subset of wavelengths that best separates the specific components of a chemical mixture. Compared to passive feature-selection approaches, in which the subset is selected offline, active sensing selects the next feature on-the-fly based on previous measurements so as to reduce uncertainty. We propose a novel multi-modal non-negative least squares method (MM-NNLS) to solve the underlying linear system, which has multiple near-optimal solutions. We tested the framework on mixture problems of up to 10 components from a library of 100 chemicals. MM-NNLS can solve complex mixtures using only a small number of measurements, and outperforms passive approaches in terms of sensing efficiency and stability.
Ricardo Gutierrez-Osuna
ICASSP2
2013 SILK: Scale-space integrated Lucas-Kanade image registration for super-resolution from video
abstract
Registration between low-resolution images is a crucial step in super-resolution. Conventional methods tend to separate scale estimation from translation and rotation estimation. This is because the scale parameter is inherently related to the image resolution. In this paper, we present an area-based image registration technique that can simultaneously estimate translation, rotation, and scale parameters and also take into account differences in resolution between two images. We first develop a scale-space model that relates each reference pixel to a single observation pixel with a scale parameter. This model is then easily generalized to include x-y shift and rotation parameters. By integrating the scale-space model into a non-linear least squares method, the method can iteratively estimate the transformation (x-y shift, rotation, and scale) in an accurate and efficient manner. We compare our proposed scale-space integrated Lucas-Kanade's method (SILK) against Lucas-Kanade's optical flow and scale-invariant feature transform (SIFT) matching and show that our method is suitable for super-resolution from very low resolution image sequences.
Joseph Lee, Ricardo Gutierrez-Osuna, S. Susan Young
ICASSP2
2013 Foreign accent conversion through voice morphing
abstract
We present a voice morphing strategy that can be used to generate a continuum of accent transformations between a foreign speaker and a native speaker. The approach performs a cepstral decomposition of speech into spectral slope and spectral detail. Accent conversions are then generated by combining the spectral slope of the foreign speaker with a morph of the spectral detail of the native speaker. Spectral morphing is achieved by representing the spectral detail through pulse density modulation and averaging pulses in a pair-wise fashion. The technique is evaluated on parallel recordings from two ARCTIC speakers using objective measures of acoustic quality, speaker identity and foreign accent that have been recently shown to correlate with perceptual results from listening tests. Index Terms: voice morphing, accent conversion.
Sandesh Aryal, Daniel Felps, Ricardo Gutierrez-Osuna
INTERSPEECH3
2012 Design and evaluation of classifier for identifying sign language videos in video sharing sites
abstract
Video sharing sites provide an opportunity for the collection and use of sign language presentations about a wide range of topics. Currently, locating sign language videos (SL videos) in such sharing sites relies on the existence and accuracy of tags, titles or other metadata indicating the content is in sign language. In this paper, we describe the design and evaluation of a classifier for distinguishing between sign language videos and other videos. A test collection of SL videos and videos likely to be incorrectly recognized as SL videos (likely false positives) was created for evaluating alternative classifiers. Five video features thought to be potentially valuable for this task were developed based on common video analysis techniques. A comparison of the relative value of the five video features shows that a measure of the symmetry of movement relative to the face is the best feature for distinguishing sign language videos. Overall, an SVM classifier provided with all five features achieves 82% precision and 90% recall when tested on the challenging test collection. The performance would be considerably higher when applied to the more varied collections of large video sharing sites.
Caio D. D. Monteiro, Ricardo Gutierrez-Osuna, Frank M. Shipman III
ASSETS2
2012 Foreign Accent Conversion Through Concatenative Synthesis in the Articulatory Domain
abstract
We propose a concatenative synthesis approach to the problem of foreign accent conversion. The approach consists of replacing the most accented portions of nonnative speech with alternative segments from a corpus of the speaker's own speech based on their similarity to those from a reference native speaker. We propose and compare two approaches for selecting units, one based on acoustic similarity [e.g., mel frequency cepstral coefficients (MFCCs)] and a second one based on articulatory similarity, as measured through electromagnetic articulography (EMA). Our hypothesis is that articulatory features provide a better metric for linguistic similarity across speakers than acoustic features. To test this hypothesis, we recorded an articulatory-acoustic corpus from a native and a nonnative speaker, and evaluated the two speech representations (acoustic versus articulatory) through a series of perceptual experiments. Formal listening tests indicate that the approach can achieve a 20% reduction in perceived accent, but also reveal a strong coupling between accent and speaker identity. To address this issue, we disguised original and resynthesized utterances by altering their average pitch and normalizing vocal tract length. An additional listening experiment supports the hypothesis that articulatory features are less speaker dependent than acoustic features.
Daniel Felps, Christian Geng, Ricardo Gutierrez-Osuna
IEEE Trans. Speech Audio Process.3
2011 Reverse caricatures effects on three-dimensional facial reconstructions
Jobany Rodriguez, Ricardo Gutierrez-Osuna
Image Vis. Comput.2
2010 Relying on critical articulators to estimate vocal tract spectra in an articulatory-acoustic database
abstract
We present a new phone-dependent feature weighting scheme that can be used to map articulatory configurations (e.g. EMA) onto vocal tract spectra (e.g. MFCC) through table lookup. The approach consists of assigning feature weights according to a feature's ability to predict the acoustic distance between frames. Since an articulator's predictive accuracy is phone-dependent (e.g., lip location is a better predictor for bilabial sounds than for palatal sounds), a unique weight vector is found for each phone. Inspection of the weights reveals a correspondence with the expected critical articulators for many phones. The proposed method reduces overall cepstral error by 6\% when compared to a uniform weighting scheme. Vowels show the greatest benefit, though improvements occur for 80\% of the tested phones.
Daniel Felps, Christian Geng, Korin Richmond, Ricardo Gutierrez-Osuna
INTERSPEECH5
2010 Developing Objective Measures of Foreign-Accent Conversion
abstract
Various methods have recently appeared to transform foreign-accented speech into its native-accented counterpart. Evaluation of these accent conversion methods requires extensive listening tests across a number of perceptual dimensions. This article presents three objective measures that may be used to assess the acoustic quality, degree of foreign accent, and speaker identity of accent-converted utterances. Accent conversion generates novel utterances: those of a foreign speaker with a native accent. Therefore, the acoustic quality in accent conversion cannot be evaluated with conventional measures of spectral distortion, which assume that a clean recording of the speech signal is available for comparison. Here we evaluate a single-ended measure of speech quality, ITU-T recommendation P.563 for narrow-band telephony. We also propose a measure of foreign accent that exploits a weakness of automatic speech recognizers: their sensitivity to foreign accents. Namely, we use phoneme-level match scores given by the HTK recognizer trained on a large number of English American speakers to obtain a measure of native accent. Finally, we propose a measure of speaker identity that projects acoustic vectors (e.g., Mel cepstral, F0) onto the linear discriminant that maximizes separability for a given pair of source and target speakers. The three measures are evaluated on a corpus of accent-converted utterances that had been previously rated through perceptual tests. Our results show that the three measures have a high degree of correlation with their corresponding subjective ratings, suggesting that they may be used to accelerate the development of foreign-accent conversion tools. Applications of these measures in the context of computer assisted pronunciation training and voice conversion are also discussed.
Daniel Felps, Ricardo Gutierrez-Osuna
IEEE Trans. Speech Audio Process.2
2009 High-Resolution Speech Signal Reconstruction in Wireless Sensor Networks
abstract
Data streaming is an emerging class of applications for sensor networks that has very high bandwidth and processing power requirements. In this paper, a new approach for speech data streaming is proposed, which is based on a distributed scheme. This scheme focuses on balancing the energy consumption among nodes in a sensor network by allowing low- resolution streams from multiple nodes to be fused at a central processing node in order to produce an enhanced resolution speech signal. Simulations and experimental results with real microphone signals are presented.
Andria Pazarloglou, Radu Stoleru, Ricardo Gutierrez-Osuna
CCNC3
2009 Demo abstract: Signal reconstruction with subnyquist sampling using wireless sensor networks
Andria Pazarloglou, Stephen M. George, Radu Stoleru, Ricardo Gutierrez-Osuna
IPSN4
2009 Foreign accent conversion in computer assisted pronunciation training
Daniel Felps, Heather Bortfeld, Ricardo Gutierrez-Osuna
Speech Commun.3
2008 Reducing the other-race effect through caricatures
abstract
We recognize faces from our own race better than those from another race. Although the relative contribution of different mechanisms (e.g. contact vs. attention) remains elusive, it is generally agreed that the other-race effect results from the fact that discriminatory facial features are race-dependent. Previous research has also shown that facial recognition improves when viewers are first familiarized with faces whose most distinctive features have been caricaturized. In this study, we sought to determine the extent to which familiarization with caricaturized faces could also be used to reduce other-race effects. Using an old/new face recognition paradigm, Caucasian subjects were first familiarized with a set of faces from multiple races, and then asked to recognize those faces among a set of confounders. Participants who were familiarized with and then asked to recognize veridical versions of the faces showed a significant other-race effect on Indian faces. In contrast, participants who were familiarized with caricaturized versions of the same faces, and then asked to recognize their veridical versions, showed no other-race effects on Indian faces. This result suggests that caricaturization may be used to help individuals focus their attention to features that are useful for recognition of other-race faces.
Jobany Rodriguez, Heather Bortfeld, Ricardo Gutierrez-Osuna
FG3
2008 Kernel oriented discriminant analysis for speaker-independent phoneme spaces
abstract
Speaker independent feature extraction is a critical problem in speech recognition. Oriented principal component analysis (OPCA) is a potential solution that can find a subspace robust against noise of the data set. The objective of this paper is to find a speaker-independent subspace by generalizing OPCA in two steps: First, we find a nonlinear subspace with the help of a kernel trick, which we refer to as kernel OPCA. Second, we generalize OPCA to problems with more than two phonemes, which leads to oriented discriminant analysis (ODA). In addition, we equip ODA with the kernel trick again, which we refer to as kernel ODA. The models are tested on the CMU ARCTIC speech database. Our results indicate that our proposed kernel methods can outperform linear OPCA and linear ODA at finding a speaker-independent phoneme space.
Heeyoul Choi, Ricardo Gutierrez-Osuna, Seungjin Choi 0001, Yoonsuck Choe
ICPR2
2007 Elimination of junk document surrogate candidates through pattern recognition
abstract
A surrogate is an object that stands for a document and enables navigation to that document. Hypermedia is often represented with textual surrogates, even though studies have shown that image and text surrogates facilitate the formation of mental models and overall understanding. Surrogates may be formed by breaking a document down into a set of smaller elements, each of which is a surrogate candidate. While processing these surrogate candidates from an HTML document, relevant information may appear together with less useful junk material, such as navigation bars and advertisements.
Eunyee Koh, Daniel Caruso, Andruid Kerne, Ricardo Gutierrez-Osuna
ACM Symposium on Document Engineering4
2006 Contrast enhancement and background suppression of chemosensor array patterns with the KIII model
abstract
Inspired by the ability of the olfactory bulb to enhance the contrast between odor representations, we propose a new hebbian learning rule that is able to increase the separability of odor patterns from gas sensor arrays. The proposed learning rule employs a hebbian term to build associations within odors and an anti-hebbian term to reduce correlated activity across odors. In addition to increasing the separability of patterns, the new learning rule can also achieve odor background suppression when combined with a habituation term. These two functions are demonstrated on Freeman's KIII, a neurodynamics model of the olfactory system. The system is first characterized on synthetic data, and also validated on experimental data from an array of chemical sensors exposed to organic solvents. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 937–953, 2006.
Agustin Gutierrez-Galvez, Ricardo Gutierrez-Osuna
Int. J. Intell. Syst.2
2006 A comparison of acoustic coding models for speech-driven facial animation
Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Ricardo Gutierrez-Osuna
Speech Commun.4
2006 Processing of chemical sensor arrays with a biologically inspired model of olfactory coding
abstract
This paper presents a computational model for chemical sensor arrays inspired by the first two stages in the olfactory pathway: distributed coding with olfactory receptor neurons and chemotopic convergence onto glomerular units. We propose a monotonic concentration-response model that maps conventional sensor-array inputs into a distributed activation pattern across a large population of neuroreceptors. Projection onto glomerular units in the olfactory bulb is then simulated with a self-organizing model of chemotopic convergence. The pattern recognition performance of the model is characterized using a database of odor patterns from an array of temperature modulated chemical sensors. The chemotopic code achieved by the proposed model is shown to improve the signal-to-noise ratio available at the sensor inputs while being consistent with results from neurobiology.
Baranidharan Raman, P. A. Sun, Agustin Gutierrez-Galvez, Ricardo Gutierrez-Osuna
IEEE Trans. Neural Networks4
2005 Mixture segmentation and background suppression in chemosensor arrays with a model of olfactory bulb-cortex interaction
abstract
We present a model of olfactory bulb-cortex interaction for the purpose of mixture processing with gas sensor arrays. The olfactory bulb is modeled with a neurodynamic model whose lateral inhibitory connections are learned through a modified Hebbian-anti-Hebbian rule. Bulbar outputs are then projected in a non-topographic fashion onto the olfactory cortex. Associational connections within cortex using Hebbian learning form a content addressable memory. Finally, inhibitory feedback from cortex is used to modulate bulbar activity. Depending on the form of feedback, Hebbian or anti-Hebbian, the model is able to perform background suppression or mixture segmentation. The model is validated on experimental data from a gas sensor array.
Baranidharan Raman, Ricardo Gutierrez-Osuna
IJCNN2
2005 Audio/visual mapping with cross-modal hidden Markov models
abstract
The audio/visual mapping problem of speech-driven facial animation has intrigued researchers for years. Recent research efforts have demonstrated that hidden Markov model (HMM) techniques, which have been applied successfully to the problem of speech recognition, could achieve a similar level of success in audio/visual mapping problems. A number of HMM-based methods have been proposed and shown to be effective by the respective designers, but it is yet unclear how these techniques compare to each other on a common test bed. In this paper, we quantitatively compare three recently proposed cross-modal HMM methods, namely the remapping HMM (R-HMM), the least-mean-squared HMM (LMS-HMM), and HMM inversion (HMMI). The objective of our comparison is not only to highlight the merits and demerits of different mapping designs, but also to study the optimality of the acoustic representation and HMM structure for the purpose of speech-driven facial animation. This paper presents a brief overview of these models, followed by an analysis of their mapping capabilities on a synthetic dataset. An empirical comparison on an experimental audio-visual dataset consisting of 75 TIMIT sentences is finally presented. Our results show that HMMI provides the best performance, both on synthetic and experimental audio-visual data.
Shengli Fu, Ricardo Gutierrez-Osuna, Anna Esposito, Praveen K. Kakumanu, Oscar N. Garcia
IEEE Trans. Multim.2
2005 Speech-driven facial animation with realistic dynamics
abstract
This work presents an integral system capable of generating animations with realistic dynamics, including the individualized nuances, of three-dimensional (3-D) human faces driven by speech acoustics. The system is capable of capturing short phenomena in the orofacial dynamics of a given speaker by tracking the 3-D location of various MPEG-4 facial points through stereovision. A perceptual transformation of the speech spectral envelope and prosodic cues are combined into an acoustic feature vector to predict 3-D orofacial dynamics by means of a nearest-neighbor algorithm. The Karhunen-Loe/spl acute/ve transformation is used to identify the principal components of orofacial motion, decoupling perceptually natural components from experimental noise. We also present a highly optimized MPEG-4 compliant player capable of generating audio-synchronized animations at 60 frames/s. The player is based on a pseudo-muscle model augmented with a nonpenetrable ellipsoidal structure to approximate the skull and the jaw. This structure adds a sense of volume that provides more realistic dynamics than existing simplified pseudo-muscle-based approaches, yet it is simple enough to work at the desired frame rate. Experimental results on an audiovisual database of compact TIMIT sentences are presented to illustrate the performance of the complete system.
Ricardo Gutierrez-Osuna, Praveen K. Kakumanu, Anna Esposito, Oscar N. Garcia, Adriana Bojórquez, José Luis Castillo, Isaac Rudomín
IEEE Trans. Multim.1
2004 Sensor-based machine olfaction with a neurodynamics model of the olfactory bulb
abstract
We propose a biologically inspired model of olfactory processing for chemosensor arrays. The model captures three functions in the early olfactory pathway: chemotopic convergence of receptor neurons onto the olfactory bulb, center on-off surround lateral interactions, and adaptation to sustained stimuli. The projection of ORNs onto glomerular units is simulated with a self-organizing model of chemotopic convergence, which leads to odor specific spatial patterning. This information serves as an input to a network of mitral cells with center on-off surround lateral inhibition, which enhances the initial contrast among odors and decouples odor identity from intensity. Finally, slow adaptation of mitral cells adds a temporal dimension to the spatial patterns that further enhances odor discrimination. The model is validated using experimental data from an array of temperature-modulated metal-oxide sensors.
Baranidharan Raman, Agustin Gutierrez-Galvez, Alexandre Perera-Lluna, Ricardo Gutierrez-Osuna
IROS4
2004 Chemosensory Processing in a Spiking Model of the Olfactory Bulb: Chemotopic Convergence and Center Surround Inhibition
abstract
This paper presents a neuromorphic model of two olfactory signal- processing primitives: chemotopic convergence of olfactory receptor neurons, and center on-off surround lateral inhibition in the olfactory bulb. A self-organizing model of receptor convergence onto glomeruli is used to generate a spatially organized map, an olfactory image. This map serves as input to a lattice of spiking neurons with lateral connections. The dynamics of this recurrent network transforms the initial olfactory image into a spatio-temporal pattern that evolves and stabilizes into odor- and intensity-coding attractors. The model is validated using experimental data from an array of temperature-modulated gas sensors. Our results are consistent with recent neurobiological findings on the antennal lobe of the honeybee and the locust.
Baranidharan Raman, Ricardo Gutierrez-Osuna
NIPS2
2003 Coherent oscillations as a neural code in a model of the olfactory system
abstract
This paper presents an investigation of two odor-coding mechanisms in Freeman's KIII neurodynamics model. Motivated by experimental evidence that supports the existence of a neural code based on synchronous oscillations, we propose an analogy between synchronization in neural populations and phase locking in KIII channels. The information carried by the phase is compared against the conventional amplitude code in terms of pattern-recovery capabilities. First, the scalar invariance of the KIII with respect to phase information is established. Symmetries and redundancies in the associative memory matrices are then exploited to perform an exhaustive evaluation of patterns on an 8-channel model. Simulation results show that phase information outperforms amplitude information in the recovery of odor patterns from incomplete or corrupted sensory stimulus.
Agustin Gutierrez-Galvez, Ricardo Gutierrez-Osuna
IJCNN2
2003 Pattern completion through phase coding in population neurodynamics
Agustin Gutierrez-Galvez, Ricardo Gutierrez-Osuna
Neural Networks2
2003 Attribute bagging: improving accuracy of classifier ensembles by using random feature subsets
Robert K. Bryll, Ricardo Gutierrez-Osuna, Francis K. H. Quek
Pattern Recognit.2
2003 Habituation in the KIII olfactory model with chemical sensor arrays
abstract
This paper presents a novel combination of chemical sensors and the KIII model for simulating mixture perception with a habituation process triggered by local activity. Stimuli are generated by partitioning feature space with labeled lines. Pattern completion is demonstrated through coherent oscillations across granule populations using experimental odor mixtures.
Ricardo Gutierrez-Osuna, Agustin Gutierrez-Galvez
IEEE Trans. Neural Networks1
2001 Chemosensory Adaptation in an Electronic Nose
abstract
This article presents a computational mechanism inspired by the process of chemosensory adaptation in the mammalian olfactory system. The algorithm operates on multiple subsets of the sensory space, generating a family of discriminant functions for different volatile compounds. A set of selectivity coefficients is associated to each discriminant function on the basis of its behavior in the presence of mixtures. These coefficients are employed to form a weighted average of the discriminant functions and establish a feedback signal that reduces the contribution of certain sensory inputs, inhibiting the overall selectivity of the system to previously detected analytes. The algorithm is validated on a database of organic solvents using an array of temperature-modulated metal-oxide chemoresistors.
Ricardo Gutierrez-Osuna, Nilesh U. Powar, P. Sun
BIBE1
1999 A method for evaluating data-preprocessing techniques for odour classification with an array of gas sensors
abstract
The performance of a pattern recognition system is dependent on, among other things, an appropriate data-preprocessing technique, In this paper, we describe a method to evaluate the performance of a variety of these techniques for the problem of odour classification using an array of gas sensors, also referred to as an electronic nose. Four experimental odour databases with different complexities are used to score the data-preprocessing techniques. The performance measure used is the cross-validation estimate of the classification rate of a K nearest neighbor voting rule operating on Fisher's linear discriminant projection subspace.
Ricardo Gutierrez-Osuna, H. Troy Nagle
IEEE Trans. Syst. Man Cybern. Part B1
1995 Global self-localization for autonomous mobile robots using self-organizing Kohonen neural networks
abstract
An approach to global self-localization for autonomous mobile robots has been developed using self-organizing Kohonen neural networks. This approach categorizes discrete regions of space using mapped sonar data corrupted by noise of varied sources and ranges. Our approach is similar to optical character recognition (OCR) in that the mapped sonar data can, over time, assume the form of a character unique to that room. Hence, it is believed that an autonomous vehicle can be capable of determining which room it is in based on mapped sensory data ascertained by wandering through and exploring that room. With some pre-processing and a robust explore routine, the solution becomes time-, translation- and rotation-invariant.
Jason A. Janét, Ricardo Gutierrez-Osuna, Troy A. Chase, Mark W. White, Ren C. Luo
IROS (3)2