Milind Rao

dblp:145/3392 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 On Retrieval of Long Audios with Complex Text Queries
Ruochu Yang, Milind Rao, Harshavardhan Sundar, Anirudh Raju, Aparna Khare, Srinath Tankasala, Di He 0004, Venkatesh Ravichandran
INTERSPEECH2
2023 Federated Self-Learning with Weak Supervision for Speech Recognition
abstract
Automatic speech recognition (ASR) models with low-footprint are increasingly being deployed on edge devices for conversational agents, which enhances privacy. We study the problem of federated continual incremental learning for recurrent neural network-transducer (RNN-T) ASR models in the privacy-enhancing scheme of learning on-device, without access to ground truth human transcripts or machine transcriptions from a stronger ASR model. In particular, we study the performance of a self-learning based scheme, with a paired teacher model updated through an exponential moving average of ASR models. Further, we propose using possibly noisy weak-supervision signals such as feedback scores and natural language understanding semantics determined from user behavior across multiple turns in a session of interactions with the conversational agent. These signals are leveraged in a multitask policy-gradient training approach to improve the performance of self-learning for ASR. Finally, we show how catastrophic forgetting can be mitigated by combining on-device learning with a memory-replay approach using selected historical datasets. These innovations allow for 10% relative improvement in WER on new use cases with minimal degradation on other test sets in the absence of strong-supervision signals such as ground-truth transcriptions.
Milind Rao, Gopinath Chennupati, Gautam Tiwari, Anit Kumar Sahu, Anirudh Raju, Ariya Rastrow, Jasha Droppo
ICASSP1
2023 Learning When to Trust Which Teacher for Weakly Supervised ASR
abstract
Automatic speech recognition (ASR) training can utilize multiple experts as teacher models, each trained on a specific domain or accent.Teacher models may be opaque in nature since their architecture may be not be known or their training cadence is different from that of the student ASR model.Still, the student models are updated incrementally using the pseudo-labels generated independently by the expert teachers.In this paper, we exploit supervision from multiple domain experts in training student ASR models.This training strategy is especially useful in scenarios where few or no human transcriptions are available.To that end, we propose a Smart-Weighter mechanism that selects an appropriate expert based on the input audio, and then trains the student model in an unsupervised setting.We show the efficacy of our approach using LibriSpeech and LibriLight benchmarks and find an improvement of 4 to 25% over baselines that uniformly weight all the experts, use a single expert model, or combine experts using ROVER.
Aakriti Agrawal, Milind Rao, Anit Kumar Sahu, Gopinath Chennupati, Andreas Stolcke
INTERSPEECH2
2022 On joint training with interfaces for spoken language understanding
abstract
Spoken language understanding (SLU) systems extract both text transcripts and semantics associated with intents and slots from input speech utterances.SLU systems usually consist of (1) an automatic speech recognition (ASR) module, (2) an interface module that exposes relevant outputs from ASR, and (3) a natural language understanding (NLU) module.Interfaces in SLU systems carry information on text transcriptions or richer information like neural embeddings from ASR to NLU.In this paper, we study how interfaces affect joint-training for spoken language understanding.Most notably, we obtain the state-of-theart results on the publicly available 50-hr SLURP [1] dataset.We first leverage large-size pretrained ASR and NLU models that are connected by a text interface, and then jointly train both models via a sequence loss function.For scenarios where pretrained models are not utilized, the best results are obtained through a joint sequence loss training using richer neural interfaces.Finally, we show the overall diminishing impact of leveraging pretrained models with increased training data size.
Anirudh Raju, Milind Rao, Gautam Tiwari, Pranav Dheram, Bryan Anderson, Chul Lee, Bach Bui, Ariya Rastrow
INTERSPEECH2
2022 ILASR: Privacy-Preserving Incremental Learning for Automatic Speech Recognition at Production Scale
abstract
Incremental learning is one paradigm to enable model building and updating at scale with streaming data. For end-to-end automatic speech recognition (ASR) tasks, the absence of human annotated labels along with the need for privacy preserving policies for model building makes it a daunting challenge. Motivated by these challenges, in this paper we use a cloud based framework for production systems to demonstrate insights from privacy preserving incremental learning for automatic speech recognition (ILASR). By privacy preserving, we mean, usage of ephemeral data which are not human annotated. This system is a step forward for production level ASR models for incremental/continual learning that offers near real-time test-bed for experimentation in the cloud for end-to-end ASR, while adhering to privacy-preserving policies. We show that the proposed system can improve the production models significantly ($3%$) over a new time period of six months even in the absence of human annotated labels with varying levels of weak supervision and large batch sizes in incremental learning. This improvement is $20%$ over test sets with new words and phrases in the new time period. We demonstrate the effectiveness of model building in a privacy-preserving incremental fashion for ASR while further exploring the utility of having an effective teacher model and use of large batch sizes.
Gopinath Chennupati, Milind Rao, Gurpreet Chadha, Aaron Eakin, Anirudh Raju, Gautam Tiwari, Anit Kumar Sahu, Ariya Rastrow, Jasha Droppo, Andy Oberlin, Buddha Nandanoor, Prahalad Venkataramanan, Pankaj Sitpure
KDD2
2022 Decentralized Optimization Over Noisy, Rate-Constrained Networks: Achieving Consensus by Communicating Differences
abstract
In decentralized optimization, multiple nodes in a network collaborate to minimize the sum of their local loss functions. The information exchange between nodes required for this task, is often limited by network connectivity. We consider a setting in which communication between nodes is hindered by both (i) a finite rate-constraint on the signal transmitted by any node, and (ii) additive noise corrupting the signal received by any node. We propose a novel algorithm for this scenario: Decentralized Lazy Mirror Descent with Differential Exchanges (DLMD-DiffEx), which guarantees convergence of the local estimates to the optimal solution under the given communication constraints. A salient feature of DLMD-DiffEx is the introduction of additional proxy variables that are maintained by the nodes to account for the disagreement in their estimates due to channel noise and rate-constraints. Convergence to the optimal solution is attained by having nodes iteratively exchange these disagreement terms until consensus is achieved. In order to prevent noise accumulation during this exchange, DLMD-DiffEx relies on two sequences: one controlling the power of the transmitted signal, and the other determining the consensus rate. We provide insights on the design of these two sequences which highlights the interplay between consensus rate and noise amplification. We investigate the performance of DLMD-DiffEx both from a theoretical perspective as well as through numerical evaluations on synthetic data and MNIST. MATLAB and Python implementations can be found athttps://github.com/rajarshisaha95/DLMD-DiffEx.
Rajarshi Saha, Stefano Rini, Milind Rao, Andrea J. Goldsmith
IEEE J. Sel. Areas Commun.3
2021 DO as I Mean, Not as I Say: Sequence Loss Training for Spoken Language Understanding
abstract
Spoken language understanding (SLU) systems extract transcriptions, as well as semantics of intent or named entities from speech, and are essential components of voice activated systems. SLU models, which either directly extract semantics from audio or are composed of pipelined automatic speech recognition (ASR) and natural language understanding (NLU) models, are typically trained via differentiable cross-entropy losses, even when the relevant performance metrics of interest are word or semantic error rates. In this work, we propose non-differentiable sequence losses based on SLU metrics as a proxy for semantic error and use the REINFORCE trick to train ASR and SLU models with this loss. We show that custom sequence loss training is the state-of-the-art on open SLU datasets and leads to 6% relative improvement in both ASR and NLU performance metrics on large proprietary datasets. We also demonstrate how the semantic sequence loss training paradigm can be used to update ASR and SLU models without transcripts, using semantic feedback alone.
Milind Rao, Pranav Dheram, Gautam Tiwari, Anirudh Raju, Jasha Droppo, Ariya Rastrow, Andreas Stolcke
ICASSP1
2021 Decentralized Optimization Over Noisy, Rate-Constrained Networks: How We Agree By Talking About How We Disagree
abstract
In decentralized optimization, multiple nodes in a network collaborate to minimize the sum of their local loss functions. The information exchange between nodes required for this task is often limited by network connectivity. We consider a generalization of this setting, in which communication is further hindered by (i) a finite data-rate constraint on the signal transmitted by any node, and (ii) an additive noise corrupting the signal received by any node. We develop a novel algorithm for this scenario: Decentralized Lazy Mirror Descent with Differential Exchanges (DLMD-DiffEx), which guarantees convergence of the local estimates to the optimal solution. A salient feature of DLMD-DiffEx is the introduction of additional proxy variables that are maintained by the nodes to account for the disagreement in their estimates due to channel noise and data-rate constraints. We investigate the performance of DLMD-DiffEx both from a theoretical perspective as well as through numerical evaluations.
Rajarshi Saha, Stefano Rini, Milind Rao, Andrea J. Goldsmith
ICASSP3
2021 Listen with Intent: Improving Speech Recognition with Audio-to-Intent Front-End
abstract
Comprehending the overall intent of an utterance helps a listener recognize the individual words spoken. Inspired by this fact, we perform a novel study of the impact of explicitly incorporating intent representations as additional information to improve a recurrent neural network-transducer (RNN-T) based automatic speech recognition (ASR) system. An audio-to-intent (A2I) model encodes the intent of the utterance in the form of embeddings or posteriors, and these are used as auxiliary inputs for RNN-T training and inference. Experimenting with a 50k-hour far-field English speech corpus, this study shows that when running the system in non-streaming mode, where intent representation is extracted from the entire utterance and then used to bias streaming RNN-T search from the start, it provides a 5.56% relative word error rate reduction (WERR). On the other hand, a streaming system using per-frame intent posteriors as extra inputs for the RNN-T ASR system yields a 3.33% relative WERR. A further detailed analysis of the streaming system indicates that our proposed method brings especially good gain on media-playing related intents (e.g. 9.12% relative WERR on PlayMusicIntent).
Swayambhu Nath Ray, Minhua Wu, Anirudh Raju, Pegah Ghahremani, Raghavendra Bilgi, Milind Rao, Harish Arsikere, Ariya Rastrow, Andreas Stolcke, Jasha Droppo
Interspeech6
2020 Speech to Semantics: Improve ASR and NLU Jointly via All-Neural Interfaces
abstract
We consider the problem of spoken language understanding (SLU) of extracting natural language intents and associated slot arguments or named entities from speech that is primarily directed at voice assistants. Such a system subsumes both automatic speech recognition (ASR) as well as natural language understanding (NLU). An end-to-end joint SLU model can be built to a required specification opening up the opportunity to deploy on hardware constrained scenarios like devices enabling voice assistants to work offline, in a privacy preserving manner, whilst also reducing server costs. We first present models that extract utterance intent directly from speech without intermediate text output. We then present a compositional model, which generates the transcript using the Listen Attend Spell ASR system and then extracts interpretation using a neural NLU model. Finally, we contrast these methods to a jointly trained end-to-end joint SLU model, consisting of ASR and NLU subsystems which are connected by a neural network based interface instead of text, that produces transcripts as well as NLU interpretation. We show that the jointly trained model shows improvements to ASR incorporating semantic information from NLU and also improves NLU by exposing it to ASR confusion encoded in the hidden layer.
Milind Rao, Anirudh Raju, Pranav Dheram, Bach Bui, Ariya Rastrow
INTERSPEECH1
2019 Distributed Convex Optimization with Limited Communications
abstract
In this paper, a distributed convex optimization algorithm, termed distributed coordinate dual averaging (DCDA) algorithm, is proposed. The DCDA algorithm addresses the scenario of a large distributed optimization problem with limited communication among nodes in the network. Currently known distributed subgradient descent methods, such as the distributed dual averaging or the distributed alternating direction method of multipliers, assume that nodes can exchange messages of large cardinality. Such an assumption on the network communication capabilities is not valid in many scenarios of practical relevance. To address this setting, we propose the DCDA algorithm as a distributed convex optimization algorithm in which the communication between nodes in each round is restricted to a fixed number of dimensions. We bound the rate of convergence under different communication protocols and network architectures for this algorithm. We also consider the extensions to the cases of imperfect gradient knowledge and when transmitted messages are corrupted by additive noise or are quantized. Numerical simulations demonstrating the performance of DCDA in these different settings are also provided.
Milind Rao, Stefano Rini, Andrea J. Goldsmith
ICASSP1
2018 Deep Learning for Joint Source-Channel Coding of Text
abstract
We consider the problem of joint source and channel coding of structured data such as natural language over a noisy channel. The typical approach to this problem in both theory and practice involves performing source coding to first compress the text and then channel coding to add robustness for the transmission across the channel. This approach is optimal in terms of minimizing end-to-end distortion with arbitrarily large block lengths of both the source and channel codes when transmission is over discrete memoryless channels. However, the optimality of this approach is no longer ensured for documents of finite length and limitations on the length of the encoding. We will show in this scenario that we can achieve lower word error rates by developing a deep learning based encoder and decoder. While the approach of separate source and channel coding would minimize bit error rates, our approach preserves semantic information of sentences by first embedding sentences in a semantic space where sentences closer in meaning are located closer together, and then performing joint source and channel coding on these embeddings.
Nariman Farsad, Milind Rao, Andrea J. Goldsmith
ICASSP2
2018 QVZ: lossy compression of quality values
abstract
Bioinformatics (2015) 31(19), 3122–3129 The authors of the above article wish to inform readers that a post-production correction has been made to add missing funding information: NIH grant U01 CA198943.
Greg Malysa, Mikel Hernaez, Idoia Ochoa, Milind Rao, Karthik Ganesan 0001, Tsachy Weissman
Bioinform.4
2017 Estimation in autoregressive processes with partial observations
abstract
We consider the problem of estimating the covariance matrix and the transition matrix of vector autoregressive (VAR) processes from partial measurements. This model encompasses settings where there are limitations in the data acquisition of the underlying measurement systems so that data is lost or corrupted by noise. An estimator for the covariance matrix of the observations is first presented. More refined estimators, factoring in structural constraints on the covariance matrix such as sparsity, bandedness, sparsity of the inverse and low-rankness are then introduced that are particularly useful in the high-dimensional regime. These estimates are then used to perform system identification by estimating the state transition matrix with or without further structural assumptions. Non-asymptotic guarantees are presented for all estimators.
Milind Rao, Tara Javidi, Yonina C. Eldar, Andrea J. Goldsmith
ICASSP1
2017 Fundamental estimation limits in autoregressive processes with compressive measurements
abstract
We consider the problem of estimating the parameters of a vector autoregressive (VAR) process from low-dimensional random projections of the observations. This setting covers the cases where we take compressive measurements of the observations or have limits in the data acquisition process associated with the measurement system and are only able to subsample. We first present fundamental bounds on the convergence of any estimator for the covariance or state-transition matrices with and without considering structural constraints of sparsity and low-rankness. We then construct an estimator for these matrices or the parameters of the VAR process and show that it is order optimal.
Milind Rao, Tara Javidi, Yonina C. Eldar, Andrea J. Goldsmith
ISIT1
2015 MGF approach to the capacity analysis of Generalized Two-Ray fading models
abstract
We propose a class of Generalized Two-Ray (GTR) fading channels that consists of two line of sight (LOS) components with random phase and a diffuse component. Observing that the GTR fading model can be expressed in terms of the underlying Rician distribution, we derive a closed-form expression for the moment generating function (MGF) of the signal-to-noise ratio (SNR) of this model. We then employ this approach to compute the ergodic capacity with receiver side information. The impact of the underlying phase difference between the LOS components on the average SNR of the signal received is also illustrated.
Milind Rao, Francisco Javier López-Martínez, Mohamed-Slim Alouini, Andrea J. Goldsmith
ICC1
2015 QVZ: lossy compression of quality values
abstract
MOTIVATION: Recent advancements in sequencing technology have led to a drastic reduction in the cost of sequencing a genome. This has generated an unprecedented amount of genomic data that must be stored, processed and transmitted. To facilitate this effort, we propose a new lossy compressor for the quality values presented in genomic data files (e.g. FASTQ and SAM files), which comprise roughly half of the storage space (in the uncompressed domain). Lossy compression allows for compression of data beyond its lossless limit. RESULTS: The proposed algorithm QVZ exhibits better rate-distortion performance than the previously proposed algorithms, for several distortion metrics and for the lossless case. Moreover, it allows the user to define any quasi-convex distortion function to be minimized, a feature not supported by the previous algorithms. Finally, we show that QVZ-compressed data exhibit better performance in the genotyping than data compressed with previously proposed algorithms, in the sense that for a similar rate, a genotyping closer to that achieved with the original quality values is obtained. AVAILABILITY AND IMPLEMENTATION: QVZ is written in C and can be downloaded from https://github.com/mikelhernaez/qvz. CONTACT: [email protected] or [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Greg Malysa, Mikel Hernaez, Idoia Ochoa, Milind Rao, Karthik Ganesan 0001, Tsachy Weissman
Bioinform.4
2015 MGF Approach to the Analysis of Generalized Two-Ray Fading Models
abstract
We analyze a class of generalized two-ray (GTR) fading channels that consist of two line-of-sight (LOS) components with random phase plus a diffuse component. We derive a closed-form expression for the moment-generating function of the signal-to-noise ratio (SNR) for this model, which greatly simplifies its analysis. This expression arises from the observation that the GTR fading model can be expressed in terms of a conditional underlying Rician distribution. We illustrate the approach to derive simple expressions for statistics and performance metrics of interest, such as the amount of fading, the level crossing rate, the symbol error rate, and the ergodic capacity in GTR fading channels. We also show that the effect of considering a more general distribution for the phase difference between the LOS components has an impact on the average SNR.
Milind Rao, Francisco Javier López-Martínez, Mohamed-Slim Alouini, Andrea J. Goldsmith
IEEE Trans. Wirel. Commun.1