Harish Arsikere

dblp:76/9876 · DBLP profile ↗
← Back
26ranked-venue papers
12as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 12 first-author · 6 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 AMuSE: Attentive Multilingual Speech Encoding for Zero-Prior ASR
abstract
Multilingual ASR offers training, deployment and overall performance benefits, but models trained via simple data pooling are known to suffer from cross-lingual interference. Oracle language information (exact-prior) and language-specific parameters are usually leveraged to overcome this, but such approaches cannot enable seamless, truly multilingual experiences. Existing methods try to overcome this limitation by relying on inferred language information or language agnostic mixture-of-experts, but they incur additional runtime complexity and/or training cost in addition to being less effective in streaming scenarios. Building on previous studies where models were trained to handle mixed-prior (knowledge that the underlying language belongs to a known group), we propose Attentive Multilingual Speech Encoding (AMuSE), a training framework designed to match exact-prior performance even in the absence of underlying language information at runtime (zero-prior), thereby making the model prior-agnostic. Leveraging AMuSE, we build a zero-prior enabled LLM-based ASR system that outperforms several exact-prior driven state-of-the-art benchmarks.
Ashutosh Varshney, Debmalya Chakrabarty, Akshat Jaiswal, Harish Arsikere, Swayambhu Nath Ray, Frederick Weber, Prantik Sen, Garima Lalwani, Sambuddha Bhattacharya, Sri Garimella
ICASSP4
2025 DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
Prabash Reddy Male, Swayambhu Nath Ray, Harish Arsikere, Akshat Jaiswal, Prakhar Swarup, Prantik Sen, Debmalya Chakrabarty, K. V. Vijay Girish, Nikhil Bhave, Frederick Weber, Sambuddha Bhattacharya, Sri Garimella
INTERSPEECH3
2021 REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling
abstract
Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DAT). We unveil the magic behind DAT and provide, for the first time, a theoretical guarantee that DAT learns accent-invariant representations. We also prove that performing the gradient reversal in DAT is equivalent to minimizing the Jensen-Shannon divergence between domain output distributions. Motivated by the proof of equivalence, we introduce reDAT, a novel technique based on DAT, which relabels data using either unsupervised clustering or soft labels. Experiments on 23K hours of multi-accent data show that DAT achieves competitive results over accent-specific baselines on both native and non-native English accents but up to 13% relative WER reduction on unseen accents; our reDAT yields further improvements over DAT by 3% and 8% relatively on non-native accents of American and British English.
Hu Hu, Xuesong Yang, Zeynab Raeesy, Jinxi Guo, Gökçe Keskin, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Roland Maas
ICASSP6
2021 Joint ASR and Language Identification Using RNN-T: An Efficient Approach to Dynamic Language Switching
abstract
Conventional dynamic language switching enables seamless multilingual interactions by running several monolingual ASR systems in parallel and triggering the appropriate downstream components using a standalone language identification (LID) service. Since this solution is neither scalable nor cost- and memory-efficient, especially for on-device applications, we propose end-to-end, streaming, joint ASR-LID architectures based on the recurrent neural network transducer framework. Two key formulations are explored: (1) joint training using a unified output space for ASR and LID vocabularies, and (2) joint training viewed as multi-task optimization. We also evaluate the benefit of using auxiliary language information obtained on-the-fly from an acoustic LID classifier. Experiments with the English-Hindi language pair show that: (a) multi-task architectures perform better overall, and (b) the best joint architecture surpasses monolingual ASR (6.4–9.2% word error rate reduction) and acoustic LID (53.9–56.1% error rate reduction) baselines while reducing the overall memory footprint by up to 46%.
Surabhi Punjabi, Harish Arsikere, Zeynab Raeesy, Chander Chandak, Nikhil Bhave, Ankish Bansal, Sergio Murillo, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, Sri Garimella, Roland Maas, Mathieu Hans, Athanasios Mouchtaris, Siegfried Kunzmann
ICASSP2
2021 Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End
abstract
Comprehending the overall intent of an utterance helps a listener recognize the individual words spoken. Inspired by this fact, we perform a novel study of the impact of explicitly incorporating intent representations as additional information to improve a recurrent neural network-transducer (RNN-T) based automatic speech recognition (ASR) system. An audio-to-intent (A2I) model encodes the intent of the utterance in the form of embeddings or posteriors, and these are used as auxiliary inputs for RNN-T training and inference. Experimenting with a 50k-hour far-field English speech corpus, this study shows that when running the system in non-streaming mode, where intent representation is extracted from the entire utterance and then used to bias streaming RNN-T search from the start, it provides a 5.56% relative word error rate reduction (WERR). On the other hand, a streaming system using per-frame intent posteriors as extra inputs for the RNN-T ASR system yields a 3.33% relative WERR. A further detailed analysis of the streaming system indicates that our proposed method brings especially good gain on media-playing related intents (e.g. 9.12% relative WERR on PlayMusicIntent).
Swayambhu Nath Ray, Minhua Wu, Anirudh Raju, Pegah Ghahremani, Raghavendra Bilgi, Milind Rao, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Jasha Droppo
Interspeech7
2021 Efficient Large Scale Semi-Supervised Learning for CTC Based Acoustic Models
abstract
Semi-supervised learning (SSL) is an active area of research which aims to utilize unlabeled data to improve the accuracy of speech recognition systems. While the previous studies have established the efficacy of various SSL methods on varying amounts of data, this paper presents largest ASR SSL experiment ever conducted till date where 75K hours of labeled and 1.2 million hours of unlabeled data is used for model training. In addition, the paper introduces couple of novel techniques to facilitate such a large scale experiment: 1) a simple scalable Teacher-Student based SSL method for connectionist temporal classification (CTC) objective and 2) effective data selection mechanisms for leveraging massive amounts of unlabeled data to boost the performance of student models. Further, we apply SSL in all stages of the acoustic model training, including final stage sequence discriminative training. Our experiments indicate encouraging word error rate (WER) gains up to 14% in such a large transcribed data regime due to the SSL training.
Prakhar Swarup, Debmalya Chakrabarty, Ashtosh Sapru, Hitesh Tulsiani, Harish Arsikere, Sri Garimella
SLT5
2020 Improved Training Strategies for End-to-End Speech Recognition in Digital Voice Assistants
Hitesh Tulsiani, Ashtosh Sapru, Harish Arsikere, Surabhi Punjabi, Sri Garimella
INTERSPEECH3
2019 Language Model Bootstrapping Using Neural Machine Translation for Conversational Speech Recognition
abstract
Building conversational speech recognition systems for new languages is constrained by the availability of utterances capturing user-device interactions. Data collection is expensive and limited by speed of manual transcription. In order to address this, we advocate the use of neural machine translation as a data augmentation technique for bootstrapping language models. Machine translation (MT) offers a systematic way of incorporating collections from mature, resource-rich conversational systems that may be available for a different language. However, ingesting raw translations from a general purpose MT system may not be effective owing to the presence of named entities, intra sentential code-switching and the domain mismatch between the conversational data being translated and the parallel text used for MT training. To circumvent this, we explore following domain adaptation techniques: (a) sentence embedding based data selection for MT training, (b) model finetuning, and (c) rescoring and filtering translated hypotheses. Using Hindi language as the experimental testbed, we supplement transcribed collections with translated US English utterances. We observe a relative word error rate reduction of 7.8-15.6%, depending on the bootstrapping phase. Fine grained analysis reveals that translation particularly aids the interaction scenarios underrepresented in the transcribed data.
Surabhi Punjabi, Harish Arsikere, Sri Garimella
ASRU2
2019 Multi-Dialect Acoustic Modeling Using Phone Mapping and Online i-Vectors
Harish Arsikere, Ashtosh Sapru, Sri Garimella
INTERSPEECH1
2017 Robust Online i-Vectors for Unsupervised Adaptation of DNN Acoustic Models: A Study in the Context of Digital Voice Assistants
Harish Arsikere, Sri Garimella
INTERSPEECH1
2016 Novel acoustic features for automatic dialog-act tagging
abstract
This paper presents 57 new acoustic features for automatic dialog-act tagging. The features are intended to be richer than and complementary to the traditional cumulative statistics of intonation. Some of our novel contributions include feature normalization with respect to neighboring utterances, incorporation of periodicity and formant features, modeling of cognitive phenomena such as hesitations, and utterance-level aggregation of short-term acoustic effects. The proposed features are applied to 3-way dialog-act tagging and question detection using two databases (British-English call-center conversations and Switchboard), and compared with a popular cumulative-statistics baseline using logistic-regression models. Our features are found to be significantly better than and complementary to the baseline, on average, achieving an absolute performance gain of ~5-6%. Combined feature ranking reveals that about 75% of the top 20 features belong to the proposed feature set, and that the two corpora differ in their feature preferences despite similar overall performance.
Harish Arsikere, Arunasish Sen, Prathosh A. P., Vivek Tyagi
ICASSP1
2016 Speaker Verification Using Short Utterances with DNN-Based Estimation of Subglottal Acoustic Features
Jinxi Guo, Gary Yeung, Deepak Muralidharan, Harish Arsikere, Amber Afshan, Abeer Alwan
INTERSPEECH4
2015 Stylex: a corpus of educational videos for research on speaking styles and their impact on engagement and learning
Harish Arsikere, Sonal Patil, Kundan Srivastava, Om Deshmukh
INTERSPEECH1
2015 Age-dependent height estimation and speaker normalization for children's speech using the first three subglottal resonances
abstract
This paper proposes an age-dependent scheme for automatic height estimation and speaker normalization of children’s speech, using the first three subglottal resonances (SGRs). Similar to previous work, our analysis indicates that children above the age of 11 years show different acoustic properties from those under 11. Therefore, an age-dependent model is investigated. The estimation algorithms for the first three SGRs are motivated by our previous research for adults. The algorithms for the first two SGRs have been applied to children’s speech before. This paper proposes a similar approach to estimate Sg3 for children. The algorithm is trained and evaluated on 46 children, aged between 6-17 years, using cross-validation. Average RMS errors in estimating Sg1, Sg2 and Sg3 using the age-dependent model are 51, 128 and 168 Hz, respectively. The height estimation algorithm employs a negative correlation between SGRs and height, and the mean absolute height estimation error was found to be less than 3.8cm for the younger children and 4.9cm for the older children. In addition, using TIDIGITS, a linear frequency warping scheme using age-dependent Sg3 gives statisticallysignificant word error rate reductions (up to 26%) relative to conventional VTLN.
Jinxi Guo, Rohit Paturi, Gary Yeung, Steven M. Lulich, Harish Arsikere, Abeer Alwan
INTERSPEECH5
2015 Acoustic stress detection for improved navigation of educational videos
Sonal Patil, Harish Arsikere, Om Deshmukh
INTERSPEECH2
2015 Content-driven Multi-modal Techniques for Non-linear Video Navigation
abstract
The growth of Massive Open Online Courses (MOOCs) has been remarkable in the last few years. A significant amount of MOOCs content is in the form of videos and participants often use non-linear navigation to browse through a video. This paper proposes the design of a system that provides non-linear navigation in educational videos using features derived from a combination of audio and visual content of a video. It provides multiple dimensions for quickly navigating to a given point of interest in a video i.e., customized dynamic time-aware word-cloud, video pages, and a 2-D timeline. In word-cloud, the relative placement of the words indicates their temporal ordering in the video whereas color codes are used to represent acoustic stress. The 2-D timeline is used to present multiple occurrences of a keyword/concept in the video in response to user click in the word-cloud. Additionally, visual content is analyzed to identify frames with "maximum written content", known as video pages. We conducted a user study with 20 users to evaluate the proposed system and compared it with transcription-based interfaces used by major MOOC providers. Our findings suggest that the proposed system leads to statistically significant navigation time savings especially on multimodal navigation tasks.
Kundan Srivastava, S. Mohana Prasad, Harish Arsikere, Sonal Patil, Om Deshmukh
IUI4
2014 Frequency warping using subglottal resonances: Complementarity with VTLN and robustness to additive noise
abstract
Based on our recently-proposed frequency-warping scheme using subglottal resonances (SGRs), this paper addresses two well-known limitations of conventional vocal-tract length normalization (VTLN): (1) its sub-optimal nature owing to the lack of frequency-dependent scaling, and (2) sensitivity to noise. Based on the idea of filter-bank interpolation, a novel approach is proposed to realize the combined effect of VTLN and SGR-based warping (which provides frequency-dependent scaling). Using the Wall Street Journal database and the conventional MFCC front end, SGR warping is shown to be complementary to VTLN in performance. Since SGR warping depends more on the given signal and less on models trained a priori, we argue that SGR warping is less sensitive to noise than VTLN. Through experiments on the AURORA-4 database with power-normalized cepstral coefficients as noise-robust front-end features, we show that SGR warping is better than VTLN, in clean as well as multi-conditional training.
Harish Arsikere, Abeer Alwan
ICASSP1
2014 Computationally-efficient endpointing features for natural spoken interaction with personal-assistant systems
abstract
Current speech-input systems typically use a nonspeech threshold for end-of-utterance detection. While usually sufficient for short utterances, the approach can cut speakers off during pauses in more complex utterances. We elicit personal-assistant speech (reminders, calendar entries, messaging, search) using a recognizer with a dramatically increased endpoint threshold, and find frequent nonfinal pauses. A standard endpointer with a 500 ms threshold (latency) results in a 36% cutoff rate for this corpus. Based on the new data, we develop low-cost acoustic features to discriminate nonfinal from final pauses. Features capture periodicity, speaking rate, spectral constancy, duration/intensity, and pitch of prepausal speech - using no speech recognition, speaker or session information. Classification experiments yield 20% EER at a 100 ms latency, thereby reducing both cutoffs and latency compared with the threshold-only baseline. Additional results on computational cost, feature importance, and speaker differences are discussed.
Harish Arsikere, Elizabeth Shriberg, Umut Ozertem
ICASSP1
2014 Speaker recognition via fusion of subglottal features and MFCCs
abstract
Motivated by the speaker-specificity and stationarity of subglot-tal acoustics, this paper investigates the utility of subglottal cep-stral coefficients (SGCCs) for speaker identification (SID) and verification (SV). SGCCs can be computed using accelerom-eter recordings of subglottal acoustics, but such an approach is infeasible in real-world scenarios. To estimate SGCCs from speech signals, we adopt the Bayesian minimum mean squared error (MMSE) estimator proposed in the speech-to-articulatory inversion literature. The joint distribution of SGCCs and speech MFCCs is modeled using the WashU-UCLA corpus (containing simultaneous recordings of speech and subglottal acoustics), and the resulting model is used to obtain an MMSE estimate of SGCCs from unseen (test) MFCCs. Cross-validation experi-ments on the WashU-UCLA corpus show that the estimation ef-ficacy, on average, is speaker dependent. A score-level fusion of MFCC and SGCC systems outperforms the MFCC-only base-line in both SID and SV tasks. On the TIMIT database (SID), the relative reduction in identification error is 16, 40 and 51% for G.712-filtered (300–3400 Hz), narrowband (0–4000 Hz) and wideband (0–8000 Hz) speech, respectively. On the NIST 2008 database (SV), the relative reduction in equal error rate is 4 and 11 % for 10 and 5 second utterances, respectively. Index Terms: speaker recognition, subglottal acoustics, cep-stral coefficients, score combination, MMSE estimation
Harish Arsikere, Hitesh Anand Gupta, Abeer Alwan
INTERSPEECH1
2014 The relationship between the second subglottal resonance and vowel class, standing height, trunk length, and F0 variation for Mandarin speakers
abstract
The relationship between vowel formants and the second subglottal resonance (Sg2) has previously been explored in English, German, Hungarian and Korean. Results from these studies indicate that vowel space is categorically divided by Sg2 and that Sg2 correlates well with standing height. One of the goals of this work is to verify if the above findings hold true in Mandarin as well. The correlation between Sg2 and sitting height (trunk length) is also studied. Further, since Mandarin is a tonal language (with more pitch variations compared to English), we study the relationship between Sg2 and fundamental frequency (F0). A new corpus of simultaneous recordings of speech and subglottal acoustics was collected from 20 native Mandarin speakers. Results on this corpus indicate that Sg2 divides vowel space in Mandarin as well, and that it is more correlated with sitting height than standing height. Paired t-tests are conducted on the Sg2 measurements from different vowel parts, which represent different F0 regions. Preliminary results show that there is no statistically-significant variation of Sg2 with F0 within a tone. Index Terms: second subglottal resonance, Mandarin, vowel space, sitting height, tonal language
Jinxi Guo, Angli Liu, Harish Arsikere, Abeer Alwan, Steven M. Lulich
INTERSPEECH3
2013 Non-linear frequency warping for VTLN using subglottal resonances and the third formant frequency
abstract
This paper proposes a non-linear frequency warping scheme for VTLN. It is based on mapping the subglottal resonances (SGRs) and the third formant frequency (F3) of a given utterance to those of a reference speaker. SGRs are used because they relate to formants in specific ways while remaining phonetically invariant, and F3 is used because it is somewhat correlated to vocal-tract length. Given an utterance, the warping parameters (SGRs and F3) are determined by obtaining initial estimates from the signal, and refining the estimates with respect to a speaker-independent model. For children (TIDIGITS), the proposed method yields statistically-significant word error rate (WER) reductions (up to 15%) relative to conventional VTLN (linear warping) when: (1) speakers show poor baseline performance, and/or (2) training data are limited. For adults (Wall Street Journal), the WER reduction relative to conventional VTLN is 4-5%. Comparison with other non-linear warping techniques is also reported.
Harish Arsikere, Steven M. Lulich, Abeer Alwan
ICASSP1
2013 Automatic estimation of the first three subglottal resonances from adults' speech signals with application to speaker height estimation
Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan
Speech Commun.1
2012 Automatic height estimation using the second subglottal resonance
abstract
This paper presents an algorithm for automatically estimating speaker height. It is based on: (1) a recently-proposed model of the subglottal system that explains the inverse relation observed between subglottal resonances and height, and (2) an improved version of our previous algorithm for automatically estimating the second subglottal resonance (Sg2). The improved Sg2 estimation algorithm was trained and evaluated on recently-collected data from 30 and 20 adult speakers, respectively. Sg2 estimation error was found to reduce by 29%, on average, as compared to the previous algorithm. The height estimation algorithm, employing the inverse relation between Sg2 and height, was trained on data from the above-mentioned 50 adults. It was evaluated on 563 adult speakers in the TIMIT corpus, and the mean absolute height estimation error was found to be less than 5.6cm.
Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan
ICASSP1
2012 Automatic estimation of the first two subglottal resonances in children's speech with application to speaker normalization in limited-data conditions
abstract
This paper proposes an automatic algorithm for estimating the first two subglottal resonances (SGRs)—Sg1 and Sg2— from continuous speech of children, and applies it to automatic speaker normalization in mismatched, limited-data conditions. The proposed algorithm is based on the observation that Sg1 and Sg2 form phonological vowel feature boundaries, and is motivated by our recent SGR estimation algorithm for adults. The algorithm is trained and evaluated, respectively, on 25 and 9 children, aged between 7 and 18 years. The average RMS errors incurred in estimating Sg1 and Sg2 are 55 and 144 Hz, respectively. By applying the proposed algorithm to a connected digits speech recognition task, it is shown that: 1) a linear frequency warping using Sg1 or Sg2 is comparable to or better than maximum likelihood-based vocal tract length normalization (MLVTLN), 2) the performance of SGR-based frequency warping is less content dependent than that of ML-VTLN, and 3) SGRbased frequency warping can be integrated into ML-VTLN to yield a statistically-significant improvement in performance.
Harish Arsikere, Gary K. F. Leung, Steven M. Lulich, Abeer Alwan
INTERSPEECH1
2011 Automatic estimation of the second subglottal resonance from natural speech
abstract
This paper deal s with the automatic estimation of the second subglottal resonance (Sg2) from natural speech spoken by adults, since our previous work focused only on estimating Sg2 from isolated diphthongs. A new database comprising speech and subglottal data of native American English (AE) speakers and bilingual Spanish/English speakers was used for the analysis. Data from 11 speakers (6 females and 5 males) were used to derive an empirical relation among the second and third formant frequencies (F2 and F3) and Sg2. Using the derived relation, Sg2 was automatically estimated from voiced sounds in English and Spanish sentences spoken by 20 different speakers (10 males and 10 females). On average, the error in estimating Sg2 was less than 100 Hz in at least 9 isolated AE vowels and less than 40 Hz in continuous speech consisting of English or Spanish sentences.
Harish Arsikere, Steven M. Lulich, Abeer Alwan
ICASSP1
2011 Analysis and Automatic Estimation of Children's Subglottal Resonances
abstract
Models and measurements of subglottal resonances are gener-ally made from adult data, but there are several applications in which it would be useful to know about subglottal resonances in children. We therefore conducted an analysis of both new and old recordings of children’s subglottal acoustics in order 1) to produce a fuller picture of the variability of children’s subglottal resonances, and 2) to confirm that existing models of subglottal acoustics can be reasonably applied to children. We also tested the effectiveness of recent algorithms for estimating children’s subglottal resonances from speech formants and the fundamen-tal frequency, which were originally formulated based on adult data. It was found that these algorithms are effective for chil-dren at least 150cm tall. Index Terms: subglottal resonances, child speech, speech pro-duction, speaker normalization
Steven M. Lulich, Harish Arsikere, John R. Morton, Gary K. F. Leung, Abeer Alwan, Mitchell Sommers
INTERSPEECH2