VLDB 2026 Research / reviewers in the wild / expert
Rakshith Sharma Srinivasa
dblp:235/5670
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter NetworkabstractMultichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottleneck in denoising speech within one stage. In contrast, conventional signal processing algorithms often feature a two-stage design, where the first stage focuses on spatially aligning the received signals with respect to the speech source, followed by the second stage to filter out noise. In this paper, we introduce Align-and-Filter network (AFnet) for deep learning-based SE that decouples the primal denoising problem into two sub-problems, which imitates the alignment-followed-by-filtering wisdom from signal processing. The key is to leverage the relative transfer functions (RTFs) that encode meaningful spatial information via a tactically designed alignment strategy. Experimental results show that by leveraging the proposed RTF-based spatial alignment supervision, AFnet learns interpretable directional features to better exploit spatial separability of sound sources for improved SE performance. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICASSP | 4 |
| 2025 | MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword SpottingabstractDeep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation of existing State-of-the-Art (SOTA) keyword spotting models on OOD data w.r.t. in-domain testing data, and propose a training mechanism to improve performance on OOD data. Specifically, we propose a novel combination of Mixup and Information Bottleneck, called MIB, to achieve SOTA performance on OOD data. Considering on-device applications, we show across multiple models ranging from sizes of 12.5K parameters to 350K parameters, that MIB achieves as much as 2.5% (absolute) improvement in performance over OOD data. Further, in the more realistic case where OOD keywords are uttered in the presence of OOD noise, MIB achieves as much as 10% (absolute) performance improvement over SOTA models. The proposed MIB is model-agnostic, i.e., it can be applied to enhance the training of any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 2 |
| 2025 | RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned PriorabstractDenoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass certain useful information about the target clean data distribution. However, existing adoption of the standard Gaussian as the prior distribution in turn discards such information when shaping the prior, resulting in sub-optimal performance. In this paper, we propose to improve conditional DDPMs for signal restoration by leveraging a more informative prior that is jointly learned with the diffusion model. The proposed framework, called RestoreGrad, seamlessly integrates DDPMs into the variational autoencoder (VAE) framework, taking advantage of the correlation between the degraded and clean signals to encode a better diffusion prior. On speech and image restoration tasks, we show that RestoreGrad demonstrates faster convergence (5-10 times fewer training steps) to achieve better quality of restored signals over existing DDPM baselines and improved robustness to using fewer sampling steps in inference time (2-2.5 times fewer), advocating the advantages of leveraging jointly learned prior for efficiency improvements in the diffusion process. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICML | 5 |
| 2024 | End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG SignalsabstractCuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this paper, we introduce a novel end-to-end model based on transformers. Further, we also introduce a novel contrastive loss-based loss function for robust training. To study the limits of performance for our proposed ideas, we first study personalized models trained on large subject-specific datasets, and achieve an average mean absolute error of 1.08/0.68 mmHg for systolic (SBP) and diastolic BP (DBP) across all subjects while achieving a best case of 0.29/0.19 mmHg. Further, in the case where subject-specific data is scarce, we leverage transfer learning using multi-subject data, and show that our model outperforms State-of-the-Art (SOTA) methods across varying amounts of subject-specific data. Suhas BN, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 2 |
| 2024 | Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language ModelabstractZero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that addresses an issue in the original CL which treats non-corresponding pairs indiscriminately. We confirm that SSCL outperforms CL through experiments. Then, we explore how including text or speech in-domain data during the SSCL training affects the out-of-domain intent classification.During the zero-shot classification, embeddings for a set of classes in the new domain are generated to calculate the similarities between each class embedding and an input utterance embedding, after which the most similar class is predicted for the utterance’s intent. Although manually-collected text sentences per class can be used to generate the class embedding, the data collection can be costly. Thus, we explore how to generate better class embeddings without human-collected text data in the target domain. The best proposed method employing an instruction-tuned Llama2, a public large language model, shows the performance comparable to the case where the human-collected text data was used, implying the importance of accurate class embedding generation. Rakshith Sharma Srinivasa, Ching Hua Lee, Yashas Malur Saidutta, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 2 |
| 2024 | Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech EnhancementabstractDeep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perform adequately in another environment, the target (new) domain of unseen acoustic or noise conditions. Even though we can improve the target domain performance by leveraging paired data in that domain, in reality, noisy data is more straightforward to collect. Therefore, it is worth studying unsupervised domain adaptation techniques for SE that utilize only noisy data from the target domain, together with exploiting the knowledge available from the source domain paired data, for improved SE in the new domain. In this paper, we present a novel adaptation framework for SE by leveraging self-supervised learning (SSL) based speech models. SSL models are pre-trained with large amount of raw speech data to extract representations rich in phonetic and acoustics information. We explore the potential of leveraging SSL representations for effective SE adaptation to new domains. To our knowledge, it is the first attempt to apply SSL models for domain adaptation in SE. Ching Hua Lee, Chouchang Yang, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Yilin Shen, Hongxia Jin |
ICASSP | 3 |
| 2024 | CIFD: Controlled Information Flow to Enhance Knowledge DistillationabstractKnowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediately sized models, Teacher Assistants (TAs) to help the transfer process. However, training TAs is expensive, as training these models is a knowledge transfer task in itself. Further, these TAs are larger than the student model and training them especially in large data settings can be computationally intensive. In this paper, we propose a novel framework called Controlled Information Flow for Knowledge Distillation (CIFD) consisting of two components. First, we propose a significantly smaller alternatives to TAs, the Rate-Distortion Module (RDM) which uses the teacher's penultimate layer embedding and a information rate-constrained bottleneck layer to replace the Teacher Assistant model. RDMs are smaller and easier to train than TAs, especially in large data regimes, since they operate on the teacher embeddings and do not need to relearn low level input feature extractors. Also, by varying the information rate across the bottleneck, RDMs can replace TAs of different sizes. Secondly, we propose the use of Information Bottleneck Module in the student model, which is crucial for regularization in the presence of a large number of RDMs. We show comprehensive state-of-the-art results of the proposed method over large datasets like Imagenet. Further, we show the significant improvement in distilling CLIP like models over a huge 12M image-text dataset. It outperforms CLIP specialized distillation methods across five zero-shot classification datasets and two zero-shot image-text retrieval datasets. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
NeurIPS | 2 |
| 2023 | To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive RefinementabstractKeyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 2 |
| 2023 | Robust Keyword Spotting for Noisy Environments by Leveraging Speech Enhancement and Speech Presence Probability
Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Yilin Shen, Hongxia Jin |
INTERSPEECH | 3 |
| 2023 | CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive LossabstractThis paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar to Contrastive Language-Image Pre-training (CLIP) and Locked-image Tuning (LiT) that have recently gained considerable attention. Classical contrastive training employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a new contrastive loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to transfer the structure of the embedding space from one modality to another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. By using publicly available datasets, we achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification. Rakshith Sharma Srinivasa, Chouchang Yang, Yashas Malur Saidutta, Ching Hua Lee, Yilin Shen, Hongxia Jin |
NeurIPS | 1 |
| 2021 | STAN: spatio-temporal attention network for pandemic prediction using real-world evidenceabstractOBJECTIVE: The COVID-19 pandemic has created many challenges that need immediate attention. Various epidemiological and deep learning models have been developed to predict the COVID-19 outbreak, but all have limitations that affect the accuracy and robustness of the predictions. Our method aims at addressing these limitations and making earlier and more accurate pandemic outbreak predictions by (1) using patients' EHR data from different counties and states that encode local disease status and medical resource utilization condition; (2) considering demographic similarity and geographical proximity between locations; and (3) integrating pandemic transmission dynamics into deep learning models. MATERIALS AND METHODS: We proposed a spatio-temporal attention network (STAN) for pandemic prediction. It uses an attention-based graph convolutional network to capture geographical and temporal trends and predict the number of cases for a fixed number of days into the future. We also designed a physical law-based loss term for enhancing long-term prediction. STAN was tested using both massive real-world patient data and open source COVID-19 statistics provided by Johns Hopkins university across all U.S. counties. RESULTS: STAN outperforms epidemiological modeling methods such as SIR and SEIR and deep learning models on both long-term and short-term predictions, achieving up to 87% lower mean squared error compared to the best baseline prediction model. CONCLUSIONS: By using information from real-world patient data and geographical data, STAN can better capture the disease status and medical resource utilization information and thus provides more accurate pandemic modeling. With pandemic transmission law based regularization, STAN also achieves good long-term prediction performance. Rakshith Sharma Srinivasa, Cheng Qian 0001, Lucas Glass, Jeffrey Spaeder, Justin K. Romberg, Jimeng Sun 0001, Cao Xiao |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Sample complexity bounds for localized sketchingabstractWe consider sketched approximate matrix multiplication and ridge regression in the novel setting of localized sketching, where at any given point, only part of the data matrix is available. This corresponds to a block diagonal structure on the sketching matrix. We show that, under mild conditions, block diagonal sketching matrices require only $O(\sr / \epsilon^2)$ and $O(\sd_{\lambda}/\epsilon)$ total sample complexity for matrix multiplication and ridge regression, respectively. This matches the state-of-the-art bounds that are obtained using global sketching matrices. The localized nature of sketching considered allows for different parts of the data matrix to be sketched independently and hence is more amenable to computation in distributed and streaming settings and results in a smaller memory and computational footprint. Rakshith Sharma Srinivasa, Mark A. Davenport, Justin K. Romberg |
AISTATS | 1 |
| 2020 | Trading Beams for Bandwidth: Imaging with Randomized BeamformingabstractWe study the problem of actively imaging a range-limited far-field scene using an antenna array. We describe how the range limit imposes structure in the measurements across multiple wavelengths. This structure allows us to introduce a novel trade-off: the number of spatial array measurements (i.e., beams that have to be formed) can be reduced to be significantly lower than the number array elements if the scene is illuminated with a broadband source. To take advantage of this trade-off, we use a small number of “generic” linear combinations of the array outputs, instead of the phase offsets used in conventional beamforming. We provide theoretical justification for the proposed trade-off without making any strong structural assumptions on the target scene (such as sparsity) except that it is range-limited. In proving our theoretical results, we take inspiration from the sketching literature. We also provide simulation results to establish the merit of the proposed signal acquisition strategy. Our proposed method results in a reduction in the number of required spatial measurements in an array imaging system and hence can directly impact their speed and cost of operation. Rakshith Sharma Srinivasa, Mark A. Davenport, Justin K. Romberg |
SIAM J. Imaging Sci. | 1 |
| 2019 | Decentralized sketching of low rank matricesabstractWe address a low-rank matrix recovery problem where each column of a rank-r matrix X of size (d1,d2) is compressed beyond the point of recovery to size L with L << d1. Leveraging the joint structure between the columns, we propose a method to recover the matrix to within an epsilon relative error in the Frobenius norm from a total of O(r(d1 + d2)\log^6(d1 + d2)/\epsilon^2) observations. This guarantee holds uniformly for all incoherent matrices of rank r. In our method, we propose to use a novel matrix norm called the mixed-norm along with the maximum l2 norm of the columns to design a novel convex relaxation for low-rank recovery that is tailored to our observation model. We also show that our proposed mixed-norm, the standard nuclear norm, and the max-norm are particular instances of convex regularization of low-rankness via tensor norms. Finally, we provide a scalable ADMM algorithm for the mixed-norm based method and demonstrate its empirical performance via large-scale simulations. Rakshith Sharma Srinivasa, Kiryung Lee, Marius Junge, Justin K. Romberg |
NeurIPS | 1 |