VLDB 2026 Research / reviewers in the wild / expert
Yashas Malur Saidutta
dblp:229/0916
· DBLP profile ↗
19ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0002-8487-0219ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMsabstractAvinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni, Srinivas Chappidi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Avinash Amballa, Yashas Malur Saidutta, Chi-Heng Lin, Vivek Kulkarni, Srinivas Chappidi |
ACL (1) | 2 |
| 2025 | SemTexIB: Semantic Text Communication with Information Bottleneck: Integrating Rate and Semantic Similarity into Training ObjectivesabstractRecent major developments in semantic communication systems stem from integration of deep learning (DL) techniques. Following the discovery of capacity achieving codes, the primary motivation for adopting the semantic approach, which retrieves meaning without requiring an exact reconstruction, is its potential to further conserve resources such as bandwidth and power. In this paper, we propose a novel semantic communication framework for textual data over additive white Gaussian noise (AWGN) channels via DL. Our framework leverages the information bottleneck (IB) principle to balance minimizing bit transmission under wireless channel rate constraints with maximizing semantic information retention. Unlike previous works, we integrate the bilingual evaluation understudy (BLEU) sentence similarity score into the training objective to enhance model performance. In particular, inspired by knowledge distillation, we utilize large language models (LLMs) during training to transfer their knowledge of text semantics into our model. Using IB principle, we train a neural semantic encoder at the transmitter and a neural semantic decoder at the receiver that incorporates into its objective function the rate constraint together with the BLEU score and the knowledge encoded in the soft probabilities produced by the LLM. Through extensive experiments, our proposed framework demonstrates a notable improvement of up to 45% in text semantic similarity compared to state-of-the-art benchmarks operating at the same channel capacity, significantly outperforming traditional communication systems. Moreover, it exhibits robustness to variations in signal-to-noise ratio (SNR) and achieves significant gains across both low and medium SNR regimes. Abdulrahman Alamoudi, Ahmet Faruk Saz, Yashas Malur Saidutta, Faramarz Fekri |
GLOBECOM | 3 |
| 2025 | Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter NetworkabstractMultichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottleneck in denoising speech within one stage. In contrast, conventional signal processing algorithms often feature a two-stage design, where the first stage focuses on spatially aligning the received signals with respect to the speech source, followed by the second stage to filter out noise. In this paper, we introduce Align-and-Filter network (AFnet) for deep learning-based SE that decouples the primal denoising problem into two sub-problems, which imitates the alignment-followed-by-filtering wisdom from signal processing. The key is to leverage the relative transfer functions (RTFs) that encode meaningful spatial information via a tactically designed alignment strategy. Experimental results show that by leveraging the proposed RTF-based spatial alignment supervision, AFnet learns interpretable directional features to better exploit spatial separability of sound sources for improved SE performance. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICASSP | 3 |
| 2025 | MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword SpottingabstractDeep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation of existing State-of-the-Art (SOTA) keyword spotting models on OOD data w.r.t. in-domain testing data, and propose a training mechanism to improve performance on OOD data. Specifically, we propose a novel combination of Mixup and Information Bottleneck, called MIB, to achieve SOTA performance on OOD data. Considering on-device applications, we show across multiple models ranging from sizes of 12.5K parameters to 350K parameters, that MIB achieves as much as 2.5% (absolute) improvement in performance over OOD data. Further, in the more realistic case where OOD keywords are uttered in the presence of OOD noise, MIB achieves as much as 10% (absolute) performance improvement over SOTA models. The proposed MIB is model-agnostic, i.e., it can be applied to enhance the training of any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2025 | RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned PriorabstractDenoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass certain useful information about the target clean data distribution. However, existing adoption of the standard Gaussian as the prior distribution in turn discards such information when shaping the prior, resulting in sub-optimal performance. In this paper, we propose to improve conditional DDPMs for signal restoration by leveraging a more informative prior that is jointly learned with the diffusion model. The proposed framework, called RestoreGrad, seamlessly integrates DDPMs into the variational autoencoder (VAE) framework, taking advantage of the correlation between the degraded and clean signals to encode a better diffusion prior. On speech and image restoration tasks, we show that RestoreGrad demonstrates faster convergence (5-10 times fewer training steps) to achieve better quality of restored signals over existing DDPM baselines and improved robustness to using fewer sampling steps in inference time (2-2.5 times fewer), advocating the advantages of leveraging jointly learned prior for efficiency improvements in the diffusion process. Ching Hua Lee, Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Yilin Shen, Hongxia Jin |
ICML | 4 |
| 2024 | Distributed Functional Compression for Independent Component Analysis in Wireless NetworksabstractIn this paper, we consider distributed Independent Component Analysis (ICA) in wireless networks, where data from several geographically distributed wireless nodes (nodes) must be transmitted to a central server (server) to extract original sources through ICA. However, transmitting the vast amount of data over wireless channels to the server poses significant challenges due to limited bandwidth and privacy concerns. Our research addresses how to encode node data to meet channel rate constraints while providing maximally relevant information for ICA. Particularly, we propose a distributed functional compression framework for learning ICA over orthogonal AWGN channels. The framework leverages the Information Bottleneck (IB) principle to encode and compress the data to meet the channel rate constraint while maximally preserving the functionally relevant information for ICA. We train both neural encoders at the nodes and a neural decoder at the server in an unsupervised manner using the IB principle. We consider ICA for both linear and nonlinear mixing setups. Compared to the state-of-the-art, over real dataset, our proposed framework demonstrates a remarkable improvement of up to approximately 43% in accurately estimating the source signals in ICA while meeting the channels’ rate constraints. Finally, we propose a three-stage training algorithm, where the raw sensory data never leaves the nodes either for training or inference, to reduce the communication overhead. We show that our proposed training algorithm notably reduces channel use compared to the traditional cloud-based method, where the observed data from the nodes are compressed and transmitted to the cloud for learning ICA. A. Alamoudi, Yashas Malur Saidutta, Faramarz Fekri |
GLOBECOM | 2 |
| 2024 | End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG SignalsabstractCuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this paper, we introduce a novel end-to-end model based on transformers. Further, we also introduce a novel contrastive loss-based loss function for robust training. To study the limits of performance for our proposed ideas, we first study personalized models trained on large subject-specific datasets, and achieve an average mean absolute error of 1.08/0.68 mmHg for systolic (SBP) and diastolic BP (DBP) across all subjects while achieving a best case of 0.29/0.19 mmHg. Further, in the case where subject-specific data is scarce, we leverage transfer learning using multi-subject data, and show that our model outperforms State-of-the-Art (SOTA) methods across varying amounts of subject-specific data. Suhas BN, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 3 |
| 2024 | Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language ModelabstractZero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that addresses an issue in the original CL which treats non-corresponding pairs indiscriminately. We confirm that SSCL outperforms CL through experiments. Then, we explore how including text or speech in-domain data during the SSCL training affects the out-of-domain intent classification.During the zero-shot classification, embeddings for a set of classes in the new domain are generated to calculate the similarities between each class embedding and an input utterance embedding, after which the most similar class is predicted for the utterance’s intent. Although manually-collected text sentences per class can be used to generate the class embedding, the data collection can be costly. Thus, we explore how to generate better class embeddings without human-collected text data in the target domain. The best proposed method employing an instruction-tuned Llama2, a public large language model, shows the performance comparable to the case where the human-collected text data was used, implying the importance of accurate class embedding generation. Rakshith Sharma Srinivasa, Ching Hua Lee, Yashas Malur Saidutta, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 4 |
| 2024 | Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech EnhancementabstractDeep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perform adequately in another environment, the target (new) domain of unseen acoustic or noise conditions. Even though we can improve the target domain performance by leveraging paired data in that domain, in reality, noisy data is more straightforward to collect. Therefore, it is worth studying unsupervised domain adaptation techniques for SE that utilize only noisy data from the target domain, together with exploiting the knowledge available from the source domain paired data, for improved SE in the new domain. In this paper, we present a novel adaptation framework for SE by leveraging self-supervised learning (SSL) based speech models. SSL models are pre-trained with large amount of raw speech data to extract representations rich in phonetic and acoustics information. We explore the potential of leveraging SSL representations for effective SE adaptation to new domains. To our knowledge, it is the first attempt to apply SSL models for domain adaptation in SE. Ching Hua Lee, Chouchang Yang, Rakshith Sharma Srinivasa, Yashas Malur Saidutta, Yilin Shen, Hongxia Jin |
ICASSP | 4 |
| 2024 | CIFD: Controlled Information Flow to Enhance Knowledge DistillationabstractKnowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediately sized models, Teacher Assistants (TAs) to help the transfer process. However, training TAs is expensive, as training these models is a knowledge transfer task in itself. Further, these TAs are larger than the student model and training them especially in large data settings can be computationally intensive. In this paper, we propose a novel framework called Controlled Information Flow for Knowledge Distillation (CIFD) consisting of two components. First, we propose a significantly smaller alternatives to TAs, the Rate-Distortion Module (RDM) which uses the teacher's penultimate layer embedding and a information rate-constrained bottleneck layer to replace the Teacher Assistant model. RDMs are smaller and easier to train than TAs, especially in large data regimes, since they operate on the teacher embeddings and do not need to relearn low level input feature extractors. Also, by varying the information rate across the bottleneck, RDMs can replace TAs of different sizes. Secondly, we propose the use of Information Bottleneck Module in the student model, which is crucial for regularization in the presence of a large number of RDMs. We show comprehensive state-of-the-art results of the proposed method over large datasets like Imagenet. Further, we show the significant improvement in distilling CLIP like models over a huge 12M image-text dataset. It outperforms CLIP specialized distillation methods across five zero-shot classification datasets and two zero-shot image-text retrieval datasets. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
NeurIPS | 1 |
| 2023 | To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive RefinementabstractKeyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simple yet elegant solution to this problem that follows from the law of total probability. We show that existing deep keyword spotting mechanisms can be improved by Successive Refinement, where the system first classifies whether the input audio is speech or not, followed by whether the input is keyword-like or not, and finally classifies which keyword was uttered. We show across multiple models with size ranging from 13K parameters to 2.41M parameters, the successive refinement technique reduces FA by up to a factor of 8 on in-domain held-out FA data, and up to a factor of 7 on out-of-domain (OOD) FA data. Further, our proposed approach is "plug-and-play" and can be applied to any deep keyword spotting model. Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Chouchang Yang, Yilin Shen, Hongxia Jin |
ICASSP | 1 |
| 2023 | Robust Keyword Spotting for Noisy Environments by Leveraging Speech Enhancement and Speech Presence Probability
Chouchang Yang, Yashas Malur Saidutta, Rakshith Sharma Srinivasa, Ching Hua Lee, Yilin Shen, Hongxia Jin |
INTERSPEECH | 2 |
| 2023 | CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive LossabstractThis paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar to Contrastive Language-Image Pre-training (CLIP) and Locked-image Tuning (LiT) that have recently gained considerable attention. Classical contrastive training employs sets of positive and negative examples to align similar and repel dissimilar training data samples. However, similarity amongst training examples has a more continuous nature, thus calling for a more `non-binary' treatment. To address this, we propose a new contrastive loss function called Continuously Weighted Contrastive Loss (CWCL) that employs a continuous measure of similarity. With CWCL, we seek to transfer the structure of the embedding space from one modality to another. Owing to the continuous nature of similarity in the proposed loss function, these models outperform existing methods for 0-shot transfer across multiple models, datasets and modalities. By using publicly available datasets, we achieve 5-8% (absolute) improvement over previous state-of-the-art methods in 0-shot image classification and 20-30% (absolute) improvement in 0-shot speech-to-intent classification and keyword classification. Rakshith Sharma Srinivasa, Chouchang Yang, Yashas Malur Saidutta, Ching Hua Lee, Yilin Shen, Hongxia Jin |
NeurIPS | 4 |
| 2022 | A Machine Learning Framework for Privacy-Aware Distributed Functional Compression over AWGN ChannelsabstractIn many diverse fields, distributed IoT devices perform collaborative inference by communicating with an edge router. Often sensory data contains sensitive attributes that should not be revealed to the router. To address this, we develop, to the best of our knowledge, the first privacy-aware machine learning framework for distributed functional compression over AWGN channels. The key feature of our approach to privacy is that we focus only on sensitive attributes of data rather than paying a high cost to protect everything. Employing a mutual information based privacy constraint, we first propose a novel approximate upper bound to protect sensitive attributes in the compressed representations of the sensory data. Next, in conjunction with the upper bound, we propose an adversarial lower bound to enhance the protection further. Thirdly, we propose novel decompositions to these bounds such distributed edge devices can ensure overall privacy by independently privatizing their components. This allows us to propose an enhanced privacy-aware algorithm that protects sensitive information during training and inference. Our experiments show that the privacy-utility trade-off from our proposed methods is significantly better than existing mechanisms. Yashas Malur Saidutta, Faramarz Fekri, Afshin Abdi |
ITW | 1 |
| 2021 | Analog Joint Source-Channel Coding for Distributed Functional Compression using Deep Neural NetworksabstractIn this paper, we study Joint Source-Channel Coding (JSCC) for distributed analog functional compression over both Gaussian Multiple Access Channel (MAC) and AWGN channels. Notably, we propose a deep neural network based solution for learning encoders and decoders. We propose three methods of increasing performance. The first one frames the problem as an autoencoder; the second one incorporates the power constraint in the objective by using a Lagrange multiplier; the third method derives the objective from the information bottleneck principle. We show that all proposed methods are variational approximations to upper bounds on the indirect rate-distortion problem's minimization objective. Further, we show that the third method is the variational approximation of a tighter upper bound compared to the other two. Finally, we show empirical performance results for image classification. We compare with existing work and showcase the performance improvement yielded by the proposed methods. Yashas Malur Saidutta, Afshin Abdi, Faramarz Fekri |
ISIT | 1 |
| 2021 | Social event planning using hybrid pairwise Markov random fieldsabstractEvent-based social networks (EBSNs) have become increasingly popular, which provide online social event management platforms for event organizers to publish and share social events (e.g., outdoor activities). In EBSNs, a major challenge for a social event organizer is how to plan a social event to attract the maximum number of attendance. To organize an event, three essential elements are required, namely, what (i.e., event content), where (i.e., event location), and when (i.e., event time). In this paper, we focus on the social event planning problem, which selects a location and time to hold a social event for the organizer with the given event content, to maximize the total number of participants. The solution of the social event planning problem could support decision-making for social event organizers. For simplicity, we denote a location and time pair as an item in this paper. To solve the social event planning problem, we present a hybrid pairwise Markov random field (H-PMRF) model which takes latent preferences of users, latent attributes of items, similarities between users and similarities between items into consideration. In particular, we construct an undirected graph where each node represents a user's decision on a specific item and each edge represents the relationship between the nodes, define the node potentials and edge potentials which model the dependency relationships between nodes, and give a joint probability distribution over the graph. Further, we adopt the Loopy Belief Propagation algorithm to compute the posterior probability distribution of each node in H-PMRF and select the location and time to hold the event which could attract the maximum number of participants. We collect real-world data set from DoubanEvent website and conduct extensive experiments on it. Experimental results show that the proposed model outperforms several baselines. Xiao Li 0033, Yashas Malur Saidutta, Faramarz Fekri |
Int. J. Intell. Syst. | 2 |
| 2021 | Joint Source-Channel Coding Over Additive Noise Analog Channels Using Mixture of Variational AutoencodersabstractIn this paper, we present a learning scheme for Joint Source-Channel Coding (JSCC) over analog independent additive noise channels. We formulate the learning problem by showing that the minimization loss function from rate-distortion theory, is upper bounded by the loss function of the Variational Autoencoder (VAE). We show that when the source dimension is greater than the channel dimension, the encoding of two source samples in the neighborhood of each other need not be near each other. Such discontinuous projection needs to be accounted for by using multiple encoders and selecting an encoder to encode samples on a particular side of the discontinuity. We explore two selection methodologies, one based on an intuitive rule and the other where it is posed as a learning task in a Mixture-of-Experts (MoE) setup. We analyze the gradients of these methods and reason why the latter is better at avoiding local optima. We show the efficacy of the proposed methodology by simulating the performance of the system for JSCC of Gaussian sources over AWGN channels and showing that the learned solutions are close to or better than the ones proposed earlier. The proposed methodology is also naturally capable of generalizing to other source distributions which we showcase by simulating for Laplace sources. The learned systems are also robust to changes in channel conditions. Further, a single system can be trained to generalize over a range of channel conditions provided the channel conditions are known at both the transmitter and the receiver. Finally, we evaluate our proposed methodology on three different image datasets and showcase consistent improvement over existing methods due to the VAE formulation. Yashas Malur Saidutta, Afshin Abdi, Faramarz Fekri |
IEEE J. Sel. Areas Commun. | 1 |
| 2019 | M to 1 Joint Source-Channel Coding of Gaussian Sources via Dichotomy of the Input Space Based on Deep LearningabstractIn this paper, we propose a deep neural network framework for Joint Source-Channel Coding of an m dimensional i.i.d. Gaussian source for transmission over a single additive white Gaussian noise channel with no delay. The framework employs two neural encoder-decoder pairs that learn to split the input signal space into two disjoint support sets. The encoder and the decoder are jointly trained to minimize the mean square error subject to a power constraint on the signal transmitted across the channel. The proposed method achieves results as good as the state of the art for m=3,4 and is easily extendable to higher dimensions. The trained model, we discovered, assigns almost equal probability to the disjoint support sets. The results show that the scheme performance is within 1.9dB of the Shannon optimal limit over a wide range of Channel Signal to Noise Ratios (CSNR) from 0dB to 30dB for various values of m. The method is also robust, i.e. employing a model trained at CSNR+/-5dB is only 0.6dB worse than a model trained specifically for that CSNR. Yashas Malur Saidutta, Afshin Abdi, Faramarz Fekri |
DCC | 1 |
| 2019 | Joint Source-Channel Coding for Gaussian Sources over AWGN Channels using Variational AutoencodersabstractIn this paper, we study joint source-channel coding of gaussian sources over multiple AWGN channels where the source dimension is greater than the number of channels. We model our system as a Variational Autoencoder and show that its loss function takes up a form that is an upper bound on the optimization function got from rate-distortion theory. The constructed system employs two encoders that learn to split the source input space into almost half with no constraints. The system is jointly trained in a data-driven manner, end-to-end. We achieve state of the art results for certain configurations, some of which are 0.7dB better than previous works. We also showcase that the trained encoder/decoder is robust, i.e., even if the channel conditions change by +/-5dB, the performance of the system does not vary by more than 0.7dB w.r.t. a system trained at that channel condition. The trained system, to an extent, has the ability to generalize when a single input dimension is dropped and for some scenarios it is less than 1dB away from the system trained for that reduced dimension. Yashas Malur Saidutta, Afshin Abdi, Faramarz Fekri |
ISIT | 1 |