EDBT 2026 Demo / reviewers in the wild / expert
Rajesh M. Hegde
dblp:88/5193 · also Rajesh Mahanand Hegde
· DBLP profile ↗
66ranked-venue papers
7as first author
18since 2021 · last 2025
0000-0002-6142-7724ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 4 since 2021Computer networks · 15 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sparse Bayesian Integrated CNN Framework for Enhanced Acoustic Source LocalizationabstractThis paper presents a novel framework for super-resolution direction of arrival (DOA) estimation of acoustic sources in the spherical harmonics (SH) domain. The proposed approach combines sparse Bayesian learning (SBL) with a convolutional neural network (CNN). The CNN is utilized to classify DOA based on the spherical harmonics decomposition (SHD) of recordings from a spherical microphone array (SMA), providing coarse DOA estimates. These estimates are then refined using SBL, which operates on a densely sampled grid around the CNN-predicted DOA classes to achieve precise localization. The CNN component exhibits robustness in noisy and reverberant environments, while SBL specializes in high-resolution localization of multiple sparse sources. By leveraging the strengths of both methods, the SH-CNN-SBL framework enhances DOA estimation accuracy in challenging conditions. Extensive simulations and real-world experiments are performed to validate the effectiveness, of the proposed method in achieving a resolution of 1°. Priyadarshini Dwivedi, Gyanajyoti Routray, Rajesh M. Hegde |
ICASSP | 3 |
| 2025 | Optimal Device Selection and Resource Allocation in Federated LearningabstractWith the advent of federated learning, the development of privacy-preserving learning models has assumed significance in several applications. However, challenges arise due to the participation of a massive number of edge devices and limited resources in a network. In this context, this paper addresses a joint device selection and resource allocation problem to improve the performance of federated learning in resource-constrained edge networks. The proposed method enhances network performance while optimally allocating limited network resources among selected devices. The optimal device selection and resource allocation problem is formulated as maximizing the number of data samples under latency, energy, and power constraints. A computationally efficient solution to this problem is proposed to ensure an optimal solution in terms of device selection and network resources. Comparison with existing methods demonstrates its ability to find optimal solutions while significantly reducing computation time by 82%. Deepali Kushwaha, Rajesh M. Hegde |
ICASSP | 2 |
| 2025 | Device Selection for Resource-Efficient Edge Caching in a Federated Learning FrameworkabstractEdge caching enhances user experience and network efficiency by locally storing popular content. Using federated learning to find popular content enables model training directly on edge devices, eliminating the need to share raw content request data. However, involving multiple devices in training can be resource-intensive. This paper proposes a device selection method to enhance edge caching performance by accurately predicting content popularity while minimizing resource consumption. Experiments on the MovieLens 1M dataset indicate that 95.23% of achievable cache efficiency can be obtained with 70% of devices. Comparison with the existing device selection methods demonstrates the improved state-of-the-art performance of the proposed approach. Deepali Kushwaha, Meenal Narkhede, Archana Limaye, Niranjan Pol, Rajesh M. Hegde |
ICASSP | 5 |
| 2025 | GEE Maximization in UAV-Aided Mobile IoT Networks Using Deep Reinforcement LearningabstractThe rapid advancement of Internet of Things (IoT) technology has improved the connectivity of several applications. The recent introduction of mobile IoT devices (IoTDs) has further broadened the scope of conventional IoT networks, resulting in the Internet of Mobile Things (IoMT). However, the IoTDs’ limited data-storage capacity and dynamic mobility is challenging for efficient data collection in resource-constrained IoMT networks. In this work, a deep reinforcement learning (DRL) method is developed to efficiently collect IoTDs’ data using an Unmanned Aerial Vehicle (UAV). The proposed UAV scheduling is designed in the Deep Deterministic Policy Gradient (DDPG) framework, over the Twin Delayed Deep Deterministic Policy Gradient (TD3) approach. The DRL agent (UAV) learns data collection policies by adapting to different IoMT network uncertainties, such as IoTD data-storage levels, mobility patterns, and data transfer constraints. Simulation results indicate the superiority of the proposed TD3-based UAV scheduling method over other DRL approaches in UAV-IoTD data collection and continuing network functionality. They motivate using the proposed method in designing reliable and autonomous IoMT networks. Rajesh M. Hegde |
ICASSP | 2 |
| 2025 | ENADL: Towards Performance Improvement of IoT Networks Using Deep Learning-Based Node Fault PredictionabstractThe Internet of Things (IoT) has grown explosively with wireless technology integration. Several IoT applications require high data throughput, low data transmission latency, and high data gathering reliability. Since, the IoT network (IoTN) is generally dynamic and utilizes a multi-hop data transmission scheme for such applications, the throughput, latency, and network lifetime tend to degrade as the hops increase. Moreover, IoT devices (IoD) are low-cost, less computationally capable, and battery-limited, further impacting performance. A faulty IoD worsens network lifetime and throughput. Predicting faulty nodes and re-routing data can significantly enhance performance. This work proposes a node fault prediction framework to enhance data routing in dynamic IoTN, maximizing throughput and lifetime. The network is represented as a graph in which the IoD are the nodes. Then a novel deep learning model is proposed utilizing various node and edge features to predict the faulty IoDs. Particularly, the proposed edge and node features-accumulation deep learning (ENADL) method exploits features, such as Euclidean distance between nodes, residual energy level of nodes, and type and number of messages passed between edges to predict the forthcoming faulty IoD. Thereafter, data routing is performed over the updated network topology. Furthermore, to improve the network lifetime, the node's degree and betweenness centrality measures-based energy allocation method is also proposed. Finally, numerical results on simulated and real-field testbeds demonstrate the ENADL method.s effectiveness in predicting faulty nodes and re-routing data packets. This results in maximized network throughput and lifetime as compared to several existing methods. Shraddha Tripathi, Faheem Nizar, Om Jee Pandey, Tushar Sandhan, Rajesh M. Hegde |
IEEE Trans. Reliab. | 5 |
| 2024 | swCNN: A Small World Convolutional Neural Network for Efficient Training
Shubham Dwivedi, Tushar Sandhan, Om Jee Pandey, Rajesh M. Hegde |
ICPR (8) | 4 |
| 2024 | Data Distribution-Aware Model Aggregation for non-IID Data in a Federated Learning FrameworkabstractAn increase in dependence on data-driven technologies has raised user privacy concerns. Utilizing federated learning allows for the parallel processing of extensive amounts of data on edge devices, resulting in minimized latency and data privacy preservation. However, the issue arises when multiple devices with diverse datasets are involved, and conventional aggregation methods prove ineffective in achieving optimal global performance. Thus, an aggregation method is required to account for both heterogeneity in datasets and diverse data distributions to improve global model performance. This work presents a novel local model aggregation method for federated learning that recognizes the deviation between local and global data distribution to define local model aggregation weight. The Kullback-Leibler (KL) divergence measures the deviation in data distributions across classes. The proposed data distribution-aware aggregation (DDAA) method is evaluated on the Google Speech Commands (GKWS) dataset with three types of data distribution among devices: Independent and Identical Distribution (IID), semi-non-IID, and non-IID. Compared to the conventional Federated Average (FedAvg) aggregation method, the experimental results indicate reasonable improvements in classification accuracy for all three distributions. The proposed work is compared with several aggregation methods present in the literature, and the results show an improvement in F1 scores ranging from 0.02 to 0.13. Apart from performance improvements, the significance of the DDAA method is also elucidated by providing insights into the computational complexity compared to existing methods. Deepali Kushwaha, Ananya Mehrotra, Rajesh M. Hegde |
NOMS | 3 |
| 2024 | Socially Aware Network Clustering for Throughput Maximization in Mobile Wireless Sensor NetworksabstractMobile sensors, such as smart wearables, autonomous cars, cognitive healthcare devices, and intelligent drones, draw great attention due to their ability to provide a wide-range of services. Hence, a consistent connection of each sensor node to the access point is a crucial requirement to obtain high data throughput in such mobile wireless sensor networks (MWSN). Data interference among such densely populated mobile sensor nodes (MSN) must also be minimized to enhance the data gathering reliability at individual MSN. In this context, a novel socially aware network clustering and interference management technique for MWSN is proposed in this work. The proposed method considers the current and prediction of future encounters among MSN to compute the social relationship index (SRI)-factor in developing a novel clustering algorithm. Moreover, a frequency-separation (FS) distance between clusters is utilized to form the non-overlapping clusters. The FS distance aids in reducing the network interference. For further interference management and reliable data transfer, beamforming is also utilized over the clustered MWSN. Subsequently, an optimization problem is formulated to maximize the data throughput over the clustered MWSN with respect to antenna downtilt angle while accounting for the high-density mobile behavior of MSN. Finally, experiments are conducted to evaluate the performance of the proposed method over a time-varying MWSN. The obtained results demonstrate the effectiveness of the proposed when compared to the benchmark methods. The results also validate the utilization of the proposed method over medium and large-scale network applications. Shraddha Tripathi, Om Jee Pandey, Rajesh M. Hegde |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | A Novel Resource Management Framework for Blockchain-Based Federated Learning in IoT NetworksabstractAt present, the centralized learning models, used for IoT applications generating large amount of data, face several challenges such as bandwidth scarcity, more energy consumption, increased uses of computing resources, poor connectivity, high computational complexity, reduced privacy, and large latency towards data transfer. In order to address the aforementioned challenges, Blockchain-Enabled Federated Learning Networks (BFLNs) emerged recently, which deal with trained model parameters only, rather than raw data. BFLNs provide enhanced security along with improved energy-efficiency and Quality-of-Service (QoS). However, BFLNs suffer with the challenges of exponential increased action space in deciding various parameter levels towards training and block generation. Motivated by aforementioned challenges of BFLNs, in this work, we are proposing an actor-critic Reinforcement Learning (RL) method to model the Machine Learning Model Owner (MLMO) in selecting the optimal set of parameter levels, addressing the challenges of exponential grow of action space in BFLNs. Further, due to the implicit entropy exploration, actor-critic RL method balances the exploration-exploitation trade-off and shows better performance than most off-policy methods, on large discrete action spaces. Therefore, in this work, considering the mobile scenario of the devices, MLMO decides the data and energy levels that the mobile devices use for the training and determine the block generation rate. This leads to minimized system latency and reduced overall cost, while achieving the target accuracy. Specifically, we have used Proximal Policy Optimization (PPO) as an on-policy actor-critic method with it's two variants, one based on Monte Carlo (MC) returns and another based on Generalized Advantage Estimate (GAE). We analyzed that PPO has better exploration and sample efficiency, lesser training time, and consistently higher cumulative rewards, when compared to off-policy Deep Q-Network (DQN). Aman Mishra, Yash Garg, Om Jee Pandey, Mahendra Kumar Shukla, Athanasios V. Vasilakos, Rajesh M. Hegde |
IEEE Trans. Sustain. Comput. | 6 |
| 2023 | Optimal Device Selection in Federated Learning for Resource-Constrained Edge NetworksabstractLow latency, resource efficiency, and data privacy are some of the crucial requirements in modern communication networks. Federated learning can efficiently address these issues by utilizing the data at the network edge and processing massive amounts of data in parallel at the edge devices, thus ensuring data privacy and low latency. For a large-scale federated learning task spanning many devices, challenges arise due to device heterogeneity, data variability, and limited network resources. Optimal selection of edge devices participating in federated learning is essential to attaining resilient, reliable, and resource-efficient edge networks. In this context, this article proposes an optimal device selection method to minimize redundant data training and improve network resource utilization without affecting the performance of federated learning over resource-constrained edge networks. The proposed optimal device selection method aims to minimize network resource demands while maximizing data diversity within the aggregated model. The performance of the proposed federated learning framework is evaluated using a publicly available image data set of handwritten digits, EMNIST (an extended version of the MNIST data set). Experimental results indicate that the proposed framework can obtain accuracy convergence performance on par with conventional federated learning methods while significantly reducing device usage (up to$50\%$) and resource utilization (up to$30\%$) while reaching 99% of achievable accuracy. The proposed method can therefore be effectively applied to resource-constrained edge networks. Deepali Kushwaha, Surender Redhu, Christopher G. Brinton, Rajesh M. Hegde |
IEEE Internet Things J. | 4 |
| 2023 | Learning based method for near field acoustic range estimation in spherical harmonics domain using intensity vectors
Priyadarshini Dwivedi, Gyanajyoti Routray, Rajesh M. Hegde |
Pattern Recognit. Lett. | 3 |
| 2023 | A Socially-Aware Radio Map Framework for Improving QoS of UAV-Assisted MEC NetworksabstractThe expeditious growth of the Internet of Things (IoT) has accelerated the evolution of multi-access edge computing (MEC). MEC alleviates the challenges of conventional cloud computing, such as high data latency, poor data gathering reliability, increased network cost, and lack of network robustness. The primary objective of MEC is to facilitate a hierarchy of edge servers to address these quality-of-service (QoS) challenges, especially the information propagation issue due to the mobility of IoT devices (IoD). Further, social-relationship among mobile IoD is a critical parameter used to reduce the data transmission delay and queue size at the MEC. Specifically, in this work, a novel socially-aware radio map generation method is proposed to compute the fine-grained and accurate locations of QoS-deprived areas. Firstly, a novel method to compute the social relationship index (SRI) factor is proposed on the basis of current and future encounters among moving IoDs. Then the obtained SRI factor is used to form clusters of mobile IoD. The clusters’ signal to interference plus noise ratio (SINR) is then used to generate the socially-aware radio map. Following that, unmanned aerial vehicles (UAV) use this radio map, which contains rich and serviceable channel information, for 3D beamforming towards the mobile clusters. Using the obtained radio map, Kalman filter-based offline path planning of UAVs is proposed to minimize the UAVs flying distance from the initial to final locations. Furthermore, an optimization problem is formulated to assess the performance of the proposed method. Finally, the performance of the proposed method is compared with the existing methods, taking into account various network parameters such as optimum number of UAVs needed to cover the deployed area, data transmission delay, and received SINR. Shraddha Tripathi, Om Jee Pandey, Linga Reddy Cenkeramaddi, Rajesh M. Hegde |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2022 | Learning to Predict Speech in Silent Videos Via Audiovisual AnalogyabstractLipreading is a difficult task, even for humans. And synthesizing the original speech waveform from lipreading makes it even a more challenging problem. Towards this end, we present a deep learning framework that can be trained end-to-end to learn the mapping between the auditory and visual signals. In particular, in this paper, our interest is to design a model that can efficiently predict the speech signal in a given silent talking-face video. The proposed framework generates a speech signal by mapping the video frames in a sequence of feature vectors. However, unlike some recent methods that adopt a sequence-to-sequence approach for translation from the frame stream to the audio stream, we posit it as an analogy learning problem between the two modalities. In which each frame is mapped to the corresponding speech segment via a deep audio-visual analogy framework. We predict plausible audio stream by training adversarially against a discriminator network. Our experiments, both qualitative and quantitative, on the publicly available GRID dataset show that the proposed method outperforms prior work on existing evaluation benchmarks. Our user studies confirm that our generated samples are more natural and closely match the ground truth speech signal. Ravindra Yadav, Ashish Sardana, Vinay P. Namboodiri, Rajesh M. Hegde |
ICASSP | 4 |
| 2022 | Learning Speaker-specific Lip-to-Speech GenerationabstractUnderstanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every speaker has a different accent and speaking style, which can be inferred from their visual and speech features. This work aims to understand the correlation/mapping between speech and the sequence of lip movement of individual speakers in an unconstrained and large vocabulary. We model the frame sequence as a distribution of features from the transformer in an autoencoder setting and learn the embeddings jointly that exploits temporal properties of both audio and video. We learn temporal synchronization using deep metric learning, which guides the decoder to generate speech in sync with input lip movements. The predictive posterior thus gives us the generated speech in speaker speaking style. We have trained our model on the Grid and Lip2Wav Chemistry lecture dataset to evaluate single speaker natural speech generation tasks from lip movement in an unconstrained natural setting. Extensive evaluation using various qualitative and quantitative metrics with human evaluation also shows that our method outperforms on Lip2Wav Chemistry dataset (large vocabulary in an unconstrained setting) by a good margin across almost all evaluation metrics and marginally outperforms the state-of-the-art on GRID dataset. Munender Varshney, Ravindra Yadav, Vinay P. Namboodiri, Rajesh M. Hegde |
ICPR | 4 |
| 2022 | Group Delay based Methods for Detection and Recognition of Whispered SpeechabstractThe present study demonstrates the effectiveness of the group delay function for detection and recognition of whispered speech. The group delay function in its spectral form is able to differentiate phonated from the whispered speech. In particular, the mean height-bandwidth product of the formants in the lower and mid frequency regions of the short term spectrum of speech is used herein for whisper speech detection. The mean height-bandwidth vector is obtained herein across all frames and is further smoothed using a moving average filter. The smoothed temporal version on this vector is able to detect the phonated to whisper change points. Towards this end, the current study investigates the cepstral, linear prediction (LP), minimum variance distortionless response (MVDR) and numerator of the group delay based smoothing techniques for whisper change point detection. Experiments on whispered speech detection are performed on the CHAINS database. Experimental results are compared to various methods for whispered speech detection available in literature. Kishore Vedvyasan, Karan Nathwani, Rajesh M. Hegde |
ICPR | 3 |
| 2022 | Improving Quality-of-Service in Cluster-Based UAV-Assisted Edge NetworksabstractWith millions of devices connected together, the Internet of Things (IoT) has become an emerging technology for future wireless networks. The ever-increasing number of smart devices and data hungry applications demand a high Quality-of-Service (QoS) for IoT. In conventional networks, data being sent to cloud for computational purpose leads to poor QoS. In order to address QoS challenges, mobile edge networks have emerged as a promising solution. In edge networks, bringing the networks resources closer to the end devices results in improved QoS. The maneuverability and the ease of versatile deployment coupled with cost efficiency makes unmanned aerial vehicles (UAVs) a promising candidate for future edge networks. The UAVs can act as edge servers to provide computational capabilities and improved services to the edge devices. Due to the flying ability, UAVs can establish better line-of-sight link with the ground devices. In this paper, we consider that the edge devices in the area of interest have to be facilitated with a certain desired QoS, which is based on the notion of outage probability of the wireless link between the UAV and the edge devices. In this context, we first propose a novel method that computes the optimum height at which UAV should hover, resulting in maximum coverage radius with sufficiently small outage probability. Then the geographical area is divided in optimal number of clusters using a novel algorithm based on K-means clustering. The method computes the optimum number of UAVs required for covering the area of interest. Each of the UAVs utilizes 3D beamforming in order to cover its own coverage area. For this purpose, we are taking coordinate transformation of the original area and forming a wide beam to cover the desired area. The obtained results demonstrate the effectiveness of the proposed method when compared to existing methods, which validate the utilization of the proposed method over large scale network applications. Tushar Bose, Aala Suresh, Om Jee Pandey, Linga Reddy Cenkeramaddi, Rajesh M. Hegde |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2021 | A Cluster based Sensor-Selection Scheme for Energy-Efficient Agriculture Sensor NetworksabstractImproving the energy-efficiency of remotely deployed sensor nodes in agriculture wireless networks is very challenging due to a lack of access to energy grid. Network clustering and limiting the amount of sensor data are among the various methods to improve the lifetime of these sensor nodes. In this work, an optimal sensor-selection scheme is proposed to improve the Quality of Service in clustered agriculture networks. the proposed method selects a limited number of sensor nodes to be active in each cluster. It considers the estimator performance while selecting limited sensor nodes for environmental monitoring. the optimal sensor-selection process considers the information of remaining energy of sensor nodes in agriculture wireless networks. the proposed method selects a limited number of nodes in each cluster while maximizing the estimator performance at the receiver. Our cluster based sensor-selection scheme optimizes the energy-efficiency in a distributed manner. Network clustering also ensures a uniform sensor-selection over a network area. Extensive experiments are conducted to analyse the performance of the proposed optimal sensor-selection method over different network scenarios. Our experimental results indicate significant improvements in energy-efficiency of agriculture wireless sensor networks and motivate the proposed method in practice. Surender Redhu, Amrendra P. Singh, Rajesh M. Hegde, Baltasar Beferull-Lozano |
CCNC | 3 |
| 2021 | Speech Prediction in Silent Videos Using Variational AutoencodersabstractUnderstanding the relationship between the auditory and visual signals is crucial for many different applications ranging from computer-generated imagery (CGI) and video editing automation to assisting people with hearing or visual impairments. However, this is challenging since the distribution of both audio and visual modality is inherently multi-modal. Therefore, most of the existing methods ignore the multimodal aspect and assume that there only exists a deterministic one-to-one mapping between the two modalities. It can lead to low-quality predictions as the model collapses to optimizing the average behavior rather than learning the full data distributions. In this paper, we present a stochastic model for generating speech in a silent video. The proposed model combines recurrent neural networks and variational deep generative models to learn the auditory signal’s conditional distribution given the visual signal. We demonstrate the performance of our model on the GRID dataset based on standard benchmarks. Ravindra Yadav, Ashish Sardana, Vinay P. Namboodiri, Rajesh M. Hegde |
ICASSP | 4 |
| 2020 | Stochastic Talking Face Generation Using Latent Distribution MatchingabstractThe ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a variety of talking face generations based on single audio input. Indeed, just having the ability to generate a single talking face would make a system almost robotic in nature. In contrast, our unsupervised stochastic audio-to-video generation model allows for diverse generations from a single audio input. Particularly, we present an unsupervised stochastic audio-to-video generation model that can capture multiple modes of the video distribution. We ensure that all the diverse generations are plausible. We do so through a principled multi-modal variational autoencoder framework. We demonstrate its efficacy on the challenging LRW and GRID datasets and demonstrate performance better than the baseline, while having the ability to generate multiple diverse lip synchronized videos. Ravindra Yadav, Ashish Sardana, Vinay P. Namboodiri, Rajesh M. Hegde |
INTERSPEECH | 4 |
| 2020 | Bridged Variational Autoencoders for Joint Modeling of Images and AttributesabstractGenerative models have recently shown the ability to realistically generate data and model the distribution accurately. However, joint modeling of an image with the attribute that it is labeled with requires learning a cross modal correspondence between image and attribute data. Though the information present in a set of images and its attributes possesses completely different statistical properties altogether, there exists an inherent correspondence that is challenging to capture. Various models have aimed at capturing this correspondence either through joint modeling of a variational autoencoder or through separate encoder networks that are then concatenated. We present an alternative by proposing a bridged variational autoencoder that allows for learning cross-modal correspondence by incorporating cross-modal hallucination losses in the latent space. In comparison to the existing methods, we have found that by using a bridge connection in latent space we not only obtain better generation results, but also obtain highly parameter-efficient model which provide 40% reduction in training parameters for bimodal dataset and nearly 70% reduction for trimodal dataset. We validate the proposed method through comparison with state of the art methods and benchmarking on standard datasets. Ravindra Yadav, Ashish Sardana, Vinay P. Namboodiri, Rajesh M. Hegde |
WACV | 4 |
| 2020 | Optimal relay node selection in time-varying IoT networks using apriori contact pattern information
Surender Redhu, Rajesh M. Hegde |
Ad Hoc Networks | 2 |
| 2020 | A Deep Learning Framework for Robust DOA Estimation Using Spherical Harmonic DecompositionabstractSpherical harmonic decomposition facilitates decomposing the sound pressure at different microphones into independent functions of frequency, azimuth and elevation of the source and microphone locations. This decomposition facilitates the extraction of two sets of features containing different information about elevation and azimuth of the source for the direction of arrival (DOA) estimation. These features can be given as input to a learning approach for the estimation of azimuth and elevation separately. This approach aims at breaking down the problem of DOA estimation into azimuth and elevation estimation separately. An advantage of this is the reduction in computational complexity when compared with the joint DOA estimation. This facilitates a straightforward extension of this approach to denser DOA search grids. The contribution of this paper is threefold. First, we propose spherical harmonic magnitude and phase features and discuss the information present in these features regarding the azimuth and elevation of the source. Second, we propose the convolutional neural network architectures for DOA estimation. Third, we analyse the training, run-time computational complexities and propose to extend the DOA estimation approach to dense DOA search grid rather than restricting to a sparse DOA search grid. The performance of conventional DOA estimation approaches degrades in case of a noisy and reverberant environment. Several advancements to the existing DOA estimation approaches have been recently proposed. However, to the best of the authors' knowledge, learning approaches to DOA estimation with dense DOA search grids with few frames in the context of spherical arrays have not been proposed. Performance evaluation is carried out using simulated as well as real datasets. The proposed approach is also evaluated on LOCATA dataset in the context of a moving source. The results are motivating enough to consider the application of the proposed method in practical scenarios. Vishnuvardhan Varanasi, Rajesh M. Hegde |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Fault-Resilient Distributed Detection and Estimation Over a SW-WSN Using LCMV BeamformingabstractRecent technology advancement has resulted in optimistic view toward the practicability of wireless sensor networks (WSNs) in the context of Internet of Things (IoT) and Cyber Physical Systems (CPS). However, to realize their full benefits in a broad range of commercial applications, there are still many technical hitches that need to be overcome. In this paper, we address three vital technical issues in a WSN: (1) distributed event detection, (2) distributed parameter estimation, and (3) network's robustness. We make use of a recent development in social networks called small world characteristics and propose novel fault-resilient distributed detection and estimation methods over a small world WSN (SW-WSN). In particular, a small world WSN has been developed by mounting antenna arrays on sensor nodes for the purpose of beamforming. A low-complexity optimization problem for beamforming is formulated by introducing a new parameter Flow between node pairs. Additionally, a new beamforming algorithm is also proposed which optimizes this flow, leading to optimal beam parameters. The proposed method yields a lower average path length and a higher average clustering coefficient of the network. Experiments are conducted using simulations and real node deployments over a WSN testbed. Analysis and experimental results obtained demonstrate that the proposed SW-WSN model achieves faster convergence rates for both distributed detection and distributed estimation while being resilient to node failures when compared to results obtained using state-of-the-art methods. Om Jee Pandey, Ved Gautam, Ha H. Nguyen 0001, Mahendra Kumar Shukla, Rajesh M. Hegde |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2020 | Cooperative Network Model for Joint Mobile Sink Scheduling and Dynamic Buffer Management Using Q-LearningabstractDevelopment of energy-efficient wireless sensor networks is crucial in the deployment of IoT and IIoT for modern day applications like smart home, smart vehicles, and smart industries. Several methods like network clustering, mobile sink deployment and dynamic sensing rate have been used in improving the energy-efficiency of wireless sensor networks in IoT framework. However, these methods have been developed independently which can lead to certain network issues like reduced lifetime, network breakdown among others. In this work, an energy-efficient method that optimizes mobile sink scheduling while concurrently providing dynamic buffer management is proposed. A cooperative network model that incorporates node clustering and mobile sink deployment in variable node sensing rate scenario is first developed. However, in such cooperative network models, mobile sink scheduling and buffer overflow management which causes information loss become challenging. This is primarily due to limited buffer size, variable sensing rate of the nodes, and the unavailability of mobile sink at all times near a cluster. Therefore, a reinforcement Q-learning framework is developed for scheduling the mobile sink while minimizing the information loss caused by buffer overflow in each cluster of a clustered WSN. More specifically, the network behaviour is learnt in the context of buffer overflow using Q-learning approach. The proposed method computes the adaptive halt-times for the mobile sink based on information loss and buffer overflow in each cluster. Performance of the proposed joint mobile sink scheduling and dynamic buffer management method is evaluated on a medium scale WSN. A clustered wireless sensor network with a total of 600 sensor nodes is considered for performance evaluation. The proposed method is shown to learn the variable node sensing rate in a reasonable amount of time using convergence analysis. Numeric evaluations indicate that the proposed method minimizes the information loss in a medium scale wireless sensor network while improving the network lifetime simultaneously. The proposed cooperative network model also outperforms in terms of energy-efficiency when compared to conventional WSN. The results are motivating enough for the use of cooperative network model in practical WSNs for IoT applications. Surender Redhu, Rajesh M. Hegde |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Model Free Calibration of Wheeled Robots Using Gaussian ProcessabstractRobotic calibration allows for the fusion of data from multiple sensors such as odometers, cameras, etc., by providing appropriate relationships between the corresponding reference frames. For wheeled robots equipped with camera/lidar along with wheel encoders, calibration entails learning the motion model of the sensor or the robot in terms of the data from the encoders and generally carried out before performing tasks such as simultaneous localization and mapping (SLAM). This work puts forward a novel Gaussian Process-based non-parametric approach for calibrating wheeled robots with arbitrary or unknown drive configurations. The procedure is more general as it learns the entire sensor/robot motion model in terms of odometry measurements. Different from existing non-parametric approaches, our method relies on measurements from the onboard sensors and hence does not require the ground truth information from external motion capture systems. Alternatively, we propose a computationally efficient approach that relies on the linear approximation of the sensor motion model. Finally, we perform experiments to calibrate robots with un-modelled effects to demonstrate the accuracy, usefulness, and flexibility of the proposed approach. Mohan Krishna Nutalapati, Lavish Arora, Anway Bose, Ketan Rajawat, Rajesh M. Hegde |
IROS | 5 |
| 2019 | Network lifetime improvement using landmark-assisted mobile sink scheduling for cyber-physical system applications
Surender Redhu, Rajesh M. Hegde |
Ad Hoc Networks | 2 |
| 2019 | Near-Field Acoustic Source Localization Using Spherical Harmonic FeaturesabstractNear-field acoustic source localization and beamforming has hitherto not been investigated extensively in the spherical harmonic domain under reverberant conditions. In this paper, a novel method for the near-field direction of arrival (DOA) and range estimation using signal invariant and direction independent spherical harmonic features is proposed. A spatial pressure interpolation method that effectively captures the acoustic energy on the surface of the sphere is first developed in the spherical harmonic domain. Near-field DOA estimates are then computed using this pressure distribution. Spherical harmonic features that are signal invariant and direction independent are then extracted using the near field DOA estimates. Signal invariant features are obtained by normalizing spherical harmonic coefficients with a component that is proportional to the source signal strength. Direction independent features are obtained using two methods. Rotation of spherical harmonic functions over a sphere is performed using Wigner-D functions in one method, whereas in the other, the effect of DOA is compensated by spherical harmonic normalization. Using the signal invariant and direction independent features, a learning-based framework which utilizes a convolutional neural network and voicing activity detection is also developed to compute the range of the near-field source. Experiments are conducted both on simulated and real speech data for evaluating the performance of the proposed spherical harmonic features in the context of near-field localization as well as beamforming. Root mean square error of both near-field DOA and source range estimates are obtained. Objective evaluation of near-field beamformed acoustic outputs is also performed. Results obtained are motivating enough for the method to be used in practical near-field beamforming applications. Vishnuvardhan Varanasi, Ayushya Agarwal, Rajesh M. Hegde |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Joint Adaptive Impulse Response Estimation and Inverse Filtering for Enhancing In-Car AudioabstractPerformance of conventional audio equalization methods for improving in-car audio listening experience is limited by the uncertainties in computing the highly varying in-car channel response. Hence these methods generally compute the channel response which is then utilized in designing the inverse filter. In this paper, a novel adaptive equalization method is developed where the channel impulse response and inverse filter are jointly estimated. The method iteratively estimates the uncertainties in the channel response using a Kalman filter and updates the inverse filter gains at every step. The joint estimation method is thus adaptive and robust to the highly varying in-car acoustic conditions. Additional contributions of this work include the development of a car database that captures impulse responses and noise samples under various in-car conditions. Both subjective and objective evaluations are performed to show the performance improvements obtained using the proposed method. Ajay Dagar, Sai Nitish Satyavolu, Rajesh M. Hegde |
ICASSP | 3 |
| 2018 | Joint Mobile Sink Scheduling and Data Aggregation in Asynchronous Wireless Sensor Networks Using Q-LearningabstractEnergy-efficient data aggregation is a challenging problem in asynchronous wireless sensor networks. Asynchronous behaviour of sensor nodes is generally due to adaptive duty cycling and it leads to information loss, buffer overflow and poor quality of services. To overcome these issues, a joint mobile sink scheduling and data aggregation scheme is proposed in this work. A reinforcement learning framework is developed herein for budgeting the energy of mobile sink while minimizing the information loss in each cluster of a clustered WSN. More specifically, a Q-learning approach is used to learn the network behaviour over time and compute adaptive halt-times for the mobile sink based on active number of nodes in each cluster. Experiments on joint mobile sink scheduling and data aggregation are conducted on a medium scale WSN. Experimental results indicate that proposed method minimizes the information loss in an asynchronous wireless sensor network. It is also observed that mobile sink performs the data gathering operation with limited energy consumption while maximizing network lifetime. Surender Redhu, Pratyush Garg, Rajesh M. Hegde |
ICASSP | 3 |
| 2018 | Stochastic Online Dictionary Learning for Speech Source Localization and Separation in Spherical Harmonic DomainabstractFrequency and location dependent components in the speech signal can be decoupled by signal processing in the spherical harmonic domain. In this paper, a sparsity based method for joint source localization and separation method using online dictionary learning is proposed. Conventional sparsity based methods utilize an overcomplete dictionary to find a sparse linear combination of dictionary atoms. Online dictionary learning discussed herein, addresses the joint localization and separation problem by learning the dictionary atoms based on stochastic approximation. The location dependent terms present in the dictionary atoms at various frequencies are then clustered to find a robust estimate of number of sources and their locations. Using these estimates, the sources are separated from the mixture. Experiments on speech source localization and separation are conducted at various SNR. Performance evaluation scores like RMSE, log spectral distance and perceptual mean opinion scores indicate reasonable improvement over conventional methods for speech source separation. Vishnuvardhan Varanasi, Rajesh M. Hegde |
ICASSP | 2 |
| 2018 | Poster: Joint Data Latency and Packet Loss Optimization for Relay-Node Selection in Time-Varying IoT NetworksabstractThe mobility of smart and connected devices in Internet of Things brings the challenge of reliable data forwarding. In this work, a relay-node selection method is proposed for mobility-tolerant data forwarding. The proposed relay-node selection method uses the connectivity information in time-varying networks. The connectivity information of IoT devices is modelled using homogeneous Poisson point processes. The connectivity duration information of all the devices with their neighbours is updated continuously. Based on connectivity information, an optimal relay-node is selected as a data forwarding node. The online relay-node selection from a source node to the base station establishes a data forwarding path. The proposed method selects the relay-nodes based on joint optimization of two network parameters, data latency and packet loss risk. The simulation results show that the proposed method significantly improves the data latency and the packet loss risk for data forwarding over time-varying IoT networks. Surender Redhu, Mukund Maheshwari, Kshitij Yeotikar, Rajesh M. Hegde |
MobiCom | 4 |
| 2018 | Energy-efficient wake-up radio protocol using optimal sensor-selection for IoTabstractEnergy conservation in sensor nodes of an IoT framework is a challenging problem. Wake-up radio, a new technology being developed by IEEE, describes a broad solution to this challenge. In this paper, we address the issue of optimal sensor nodes selection to maximize the energy-efficiency of a WSN for IoT applications. A wake-up protocol is then applied on the optimally selected-sensor nodes. The node selection is formulated as an optimization problem with spatio-temporal correlation constraints thus minimizing redundant data transfer. Subsequently, a MLE method is used to solve the problem. Additional novelty of this work is the development of a wake-up protocol over hexagonal grids when a mobile sink is used for data aggregation leading to improved energy-efficiency. The impact of the proposed method on decision errors in IoT applications is also discussed. Extensive experiments are performed to analyse the energy-efficiency, estimation error, modelling accuracy, network lifetime of WSN in the context of IoT applications. Results indicate an improvement in energy-efficiency when compared to conventional methods. Surender Redhu, Rahul Mahavar, Rajesh M. Hegde |
WCNC | 3 |
| 2018 | Client-wise cohort set selection by combining speaker- and phoneme-specific I-vectors for speaker verification
Waquar Ahmad, Harish Karnick, Rajesh M. Hegde |
Multim. Tools Appl. | 3 |
| 2017 | Second order cone programming based localization method for Internet of ThingsabstractA novel method for device localization under mixed line-of-sight/non-line-of-sight (LOS/NLOS) conditions based on second order cone programming (SOCP) is presented in this paper. The devices can communicate cooperatively among themselves in a large internet of things (IoT) network. SOCP methods have, hitherto, not been utilized in the node localization under mixed LOS/NLOS conditions. Unlike semidefinite programming (SDP) formulation, SOCP is computationally efficient for resource constrained IoT network. The proposed method can work seamlessly in mixed LOS/NLOS conditions. The robustness of the method is due to the fair utilization of all measurements obtained under LOS and NLOS conditions. The computational complexity of this method is quadratic in the number of nearest neighbours of the unknown node. Cramér-Rao bound and localization error are analyzed to illustrate the effectiveness of the proposed method. The experimental results of the proposed method indicate a reasonable improvement when compared to recent state of the art methods. Sudhir Kumar 0002, Rishabh Dixit, Rajesh M. Hegde |
CoDIT | 3 |
| 2017 | Robust online direction of arrival estimation using low dimensional spherical harmonic featuresabstractSignal processing in spherical harmonic domain has the ability to decouple frequency dependent and location dependent components of the signal received. A method for low dimensional spherical harmonic feature extraction is proposed in this work for DOA estimation in noisy and reverberant environments. The features are extracted using frequency smoothing and a transformation which makes them frequency and signal invariant. Additionally an online manifold regularization framework is explored which utilizes the proposed spherical harmonic features to compute real time DOA estimates. This framework minimizes an instantaneous risk function and finds an inverse mapping function that maps spherical harmonic features to the DOA estimate. Performance of the proposed DOA estimation method is then compared with DOA estimates obtained from features such as generalized cross correlation and relative transfer function in a semi-supervised manifold regularization framework. Experimental results on DOA estimation in terms of root mean square error and probability of resolution indicate a reasonable improvement in the localization performance along with significant reduction in feature dimension. Vishnuvardhan Varanasi, Rajesh M. Hegde |
ICASSP | 2 |
| 2017 | Music Tempo Estimation Using Sub-Band SynchronyabstractTempo estimation aims at estimating the pace of a musical piece measured in beats per minute. This paper presents a new tempo estimation method that utilizes coherent energy changes across multiple frequency sub-bands to identify the onsets. A new measure, called the sub-band synchrony, is proposed to detect and quantify the coherent amplitude changes across multiple sub-bands. Given a musical piece, our method first detects the onsets using the sub-band synchrony measure. The periodicity of the resulting onset curve, measured using the autocorrelation function, is used to estimate the tempo value. The performance of the sub-band synchrony based tempo estimation method is evaluated on two music databases. Experimental results indicate a reasonable improvement in performance when compared to conventional methods of tempo estimation. Shreyan Chowdhury, Tanaya Guha, Rajesh M. Hegde |
INTERSPEECH | 3 |
| 2017 | Node localization over small world WSNs using constrained average path length reduction
Om Jee Pandey, Rajesh M. Hegde |
Ad Hoc Networks | 2 |
| 2016 | Radial filters for near field source separation in spherical harmonic domainabstractRadial filter design for processing near field speech sources over a spherical microphone array is challenging. Polynomial based radial filter design procedures have been proposed in earlier work. In this paper we address the issue of radial filter design using a family of orthogonal polynomials called the Gegenbauer polynomials. The radial filters designed using this approach indicate an improved radial response and greater efficiency in attenuating distant sources. Improved white noise gain and directivity index are also noted from experimental evaluations. The radial filters hence designed are used to separate directionally co-incident near field speech sources. Subjective evaluation is conducted on separated sources using measures like LSD, PESQ, and SDR. The subjective evaluation scores are motivating enough to be considered for practical speech and audio applications. Isha Agrawal, Rajesh M. Hegde |
ICASSP | 2 |
| 2016 | The spherical harmonics root-musicabstractSpherical harmonics root-MUSIC (MUltiple SIgnal Classification) technique for source localization using spherical microphone array is presented in this paper. Earlier work on root-MUSIC is limited to linear and planar arrays. Root-MUSIC for planar array utilizes the concept of manifold separation and beamspace transformation. In this paper, the Vandermonde structure of array manifold for a particular order is proved. Hence, the validity of root-MUSIC in the spherical harmonics domain is confirmed. The proposed method is evaluated by using simulated experiments on source localization. Root mean square error analysis and statistical analysis are presented. The experimental measures at various signal to noise ratios (SNRs) show the robustness of the proposed method. The method is also verified by using experiment on real signal acquired over spherical microphone array. Lalan Kumar, Guoan Bi, Rajesh M. Hegde |
ICASSP | 3 |
| 2016 | Gaussian Process Regression for Fingerprinting based Localization
Sudhir Kumar 0002, Rajesh M. Hegde, Agathoniki Trigoni |
Ad Hoc Networks | 2 |
| 2016 | Multi-sensor data fusion methods for indoor localization under collinear ambiguity
Sudhir Kumar 0002, Rajesh M. Hegde |
Pervasive Mob. Comput. | 2 |
| 2015 | Representation and modeling of spherical harmonics manifold for source localizationabstractSource localization has been studied in the spatial domain using differential geometry in earlier work. However, parameters of the sensor array manifold have hitherto not been investigated for source localization in spherical harmonics domain. The objective of this work is to represent and model the manifold surface using differential geometry. The system model for source localization over a spherical harmonic manifold is first formulated. Subsequently, the manifold parameters are modeled in the spherical harmonics domain. Source localization methods using MUSIC and MVDR over the spherical harmonics manifold are developed. Experiments on source localization using a spherical microphone array indicate high resolution in noise. Arun Parthasarathy, Saurabh Kataria 0001, Lalan Kumar, Rajesh M. Hegde |
ICASSP | 4 |
| 2015 | Hybrid maximum depth-kNN method for real time node tracking using multi-sensor dataabstractIn this paper, a hybrid maximum depth - k Nearest Neighbour (hybrid MD-kNN) method for real time sensor node tracking and localization is proposed. The method combines two individual location hypothesis functions obtained from generalized maximum depth and generalized kNN methods. The individual location hypothesis functions are themselves obtained from multiple sensors measuring visible light, humidity, temperature, acoustics, and link quality. The hybridMD-kNN method therefore combines the lower computational power of maximum depth and outlier rejection ability of kNN method to realize a robust real time tracking method. Additionally, this method does not require the assumption of an underlying distribution under non-line-of-sight (NLOS) conditions. Additional novelty of this method is the utilization of multivariate data obtained from multiple sensors which has hitherto not been used. The affine invariance property of the hybrid MD-kNN method is proved and its robustness is illustrated in the context of node localization. Experimental results on the Intel Berkeley research data set indicates reasonable improvements over conventional methods available in literature. Sudhir Kumar 0002, Abhay Kumar 0001, Rajesh M. Hegde |
ICC | 4 |
| 2015 | Joint source localization and separation in spherical harmonic domain using a sparsity based method
Sachin N. Kalkur, Sandeep Reddy C, Rajesh M. Hegde |
INTERSPEECH | 3 |
| 2015 | Sensor node tracking using semi-supervised Hidden Markov Models
Sudhir Kumar 0002, Shriman Narayan Tiwari, Rajesh M. Hegde |
Ad Hoc Networks | 3 |
| 2015 | Multi-sensor data fusion methods for indoor activity recognition using temporal evidence theory
Aseem Kushwah, Sudhir Kumar 0002, Rajesh M. Hegde |
Pervasive Mob. Comput. | 3 |
| 2015 | Joint source separation and dereverberation using constrained spectral divergence optimization
Karan Nathwani, Rajesh M. Hegde |
Signal Process. | 2 |
| 2015 | Robust acoustic echo cancellation using Kalman filter in double talk scenario
Sanchit Goel, Karan Nathwani, Rajesh M. Hegde |
Speech Commun. | 4 |
| 2015 | Stochastic Cramér-Rao Bound Analysis for DOA Estimation in Spherical Harmonics DomainabstractCramér-Rao bound (CRB) has been formulated in earlier work for linear, planar and 3-D array configurations. The formulations developed in prior work, make use of the standard spatial data model. In this paper, the existence of CRB for the spherical harmonics data model is first verified. Subsequently, an expression for stochastic CRB is derived for direction of arrival (DOA) estimation in spherical harmonics domain. The stochastic CRBs for azimuth and elevation are plotted at various signal to noise ratios (SNRs) and snapshots. It is noted that a lower bound on the CRB is attained at high SNR. A similar observation is made when larger number of snapshots are used. Lalan Kumar, Rajesh M. Hegde |
IEEE Signal Process. Lett. | 2 |
| 2014 | Fast modelling of pinna spectral notches from HRTFs using linear prediction residual cepstrumabstractDeveloping individualized head related transfer functions (HRTF) is an essential requirement for accurate virtualization of sound. However it is time consuming and complicated for both the subject and the developer. Obtaining the spectral notches which are the most prominent features of HRTF is very important to reconstruct the head related impulse response (HRIR) accurately. In this paper, a method suitable for fast computation of the frequencies of spectral notches is proposed. The linear prediction residual cepstrum is used to compute the spectral notches with a high degree of accuracy in this work. Subsequent use of Batteaus Reflection model to overlay the spectral notches on the pinna images indicate that the proposed method is able to provide finer contours. Experiments on reconstruction of the HRIR indicates that the method performs better than other methods. Chaitanya Ahuja, Rajesh M. Hegde |
ICASSP | 2 |
| 2014 | A sparse reconstruction method for speech source localization using partial dictionaries over a spherical microphone arrayabstractSparse reconstruction methods have been used extensively for source localization over uniform linear arrays and circular arrays. In this paper a sparse reconstruction method for speech source localization using partial dictionaries over a spherical microphone array is proposed. The source localization method proposed in this work addresses two important research issues. It formulates the source localization problem in the spherical harmonics domain as a sparse reconstruction problem. Subsequently, a low complexity method to estimate the direction of arrival (DOA) of multiple sources is also proposed by using partial elevation angle dictionaries. The use of such dictionaries reduces the complexity of the search involved in the two dimensional DOA estimation. Source localization experiments are conducted at different SNRs and compared with conventional DOA estimation methods like MUSIC and MVDR. The experimental results obtained from the proposed method indicate a reasonable reduction in the localization error. Kushagra Singhal, Rajesh M. Hegde |
INTERSPEECH | 2 |
| 2013 | Energy efficient optimal node-source localization using mobile beacon in ad-hoc sensor networksabstractIn this paper, a single mobile beacon based method to localize nodes using principle of maximum power reception is proposed. Optimal positioning of the mobile beacon for minimum energy consumption is also discussed. In contrast to existing methods, the node localization is done with prior location of only three nodes. There is no need of synchronization, as there is only one mobile anchor and each node communicates only with the anchor node. Also, this method is not constrained by a fixed sensor geometry. The localization is done in a distributed fashion, at each sensor node. Experiments on node-source localization are conducted by deploying sensors in an ad-hoc manner in both outdoor and indoor environments. Localization results obtained herein indicate a reasonable performance improvement when compared to conventional methods. Sudhir Kumar 0002, Vatsal Sharan, Rajesh M. Hegde |
GLOBECOM | 3 |
| 2013 | Significance of variable height-bandwidth group delay filters in the spectral reconstruction of speech
Devanshu Arya, Anant Raj, Rajesh M. Hegde |
INTERSPEECH | 3 |
| 2013 | Joint noise cancellation and dereverberation using multi-channel linearly constrained minimum variance filter
Karan Nathwani, Rajesh M. Hegde |
INTERSPEECH | 2 |
| 2013 | Group Delay Based Methods for Speaker Segregation and its Application in Multimedia Information RetrievalabstractA novel method of single channel speaker segregation using the group delay cross correlation function is proposed in this paper. The group delay function, which is the negative derivative of the phase spectrum, yields robust spectral estimates. Hence the group delay spectral estimates are first computed over frequency sub-bands after passing the speech signal through a bank of filters. The filter bank spacing is based on a multi-pitch algorithm that computes the pitch estimates of the competing speakers. An affinity matrix is then computed from the group delay spectral estimates of each frequency sub-band. This affinity matrix represents the correlations of the different sub-bands in the mixed broadband speech signal. The grouping of correlated harmonics present in the mixed speech signal is then carried out by using a new iterative graph cut method. The signals are reconstructed from the respective harmonic groups which represent individual speakers in the mixed speech signal. Spectrographic masks are then applied on the reconstructed signals to refine their perceptual quality. The quality of separated speech is evaluated using several objective and subjective criteria. Experiments on multi-speaker automatic speech recognition are conducted using mixed speech data from the GRID corpus. A cell phone based multimedia information retrieval system (MIRS) for multi-source meeting environments are also developed. Karan Nathwani, Pranav Pandit, Rajesh M. Hegde |
IEEE Trans. Multim. | 3 |
| 2010 | Significance of the MUSIC-group delay spectrum in speech acquisition from distant microphonesabstractConventionally the spectral magnitude of MUSIC is used for efficient beam forming and clean speech acquisition from distant microphones. The MUSIC method is unable to resolve closely spaced DOAs with a computationally plausible number of sensors. In this paper we propose the use of the group delay function computed from theMUSIC phase spectrum for efficient DOA estimation. The group delay function which has been hitherto used for temporal frequency processing of speech signals is computed on the phase spectrum of MUSIC and is found to resolve spatially contiguous speech sources. The additive property of the group delay function in the spatial domain is also discussed using root-MUSIC polynomial analysis. Experimental results on DOA estimation using a two channel microphone array show that the average error distribution of the MUSIC group delay spectrum is minimum when compared to MUSIC magnitude spectrum. Filter-Sum beam formers are trained using estimated DOAs on speech acquired from distant microphones. The results of speech recognition experiments conducted on meeting room data are used to illustrate the significance of the MUSIC group delay spectrum in speech acquisition from distant microphones. Mrityunjaya Shukla, Rajesh M. Hegde |
ICASSP | 2 |
| 2009 | On the Design and Prototype Implementation of a Multimodal Situation Aware SystemabstractIn this paper we describe the design concepts and prototype implementation of a situation aware ubiquitous computing system using multiple modalities such as National Marine Electronics Association (NMEA) data from Global Positioning System (GPS) receivers, text, speech, environmental audio, and handwriting inputs. While most mobile and communication devices know where and who they are, by accessing context information primarily in the form of location, time stamps, and user identity, the concept of sharing of this information in a reliable and intelligent fashion is crucial in many scenarios. A framework which takes the concept of context aware computing to the level of situation aware computing by intelligent information exchange between context aware devices is designed and implemented in this work. Four sensual modes of contextual information like text, speech, environmental audio, and handwriting are augmented to conventional contextual information sources like location from GPS, user identity based on IP addresses (IPA), and time stamps. Each device derives its context not necessarily using the same criteria or parameters but by employing selective fusion and fission of multiple modalities. The processing of each individual modality takes place at the client device followed by the summarization of context as a text file. Exchange of dynamic context information between devices is enabled in real time to create multimodal situation aware devices. A central repository of all user context profiles is also created to enable self-learning devices in the future. Based on the results of simulated situations and real field deployments it is shown that the use of multiple modalities like speech, environmental audio, and handwriting inputs along with conventional modalities can create devices with enhanced situational awareness. Rajesh M. Hegde, Joseph Kurniawan, Bhaskar D. Rao |
IEEE Trans. Multim. | 1 |
| 2008 | Significance of group delay based acoustic features in the linguistic search space for robust speech recognition
Rajesh M. Hegde, Hema A. Murthy |
INTERSPEECH | 2 |
| 2007 | Spectral Estimation of Voiced Speech using a Family of MVDR EstimatesabstractWe present a robust approach to modeling voiced speech using a family of minimum variance distortionless response (MVDR) spectral estimates. The method exploits the fact that for a fixed model order, for a sinusoidal signal in noise, the MVDR estimate at the sinusoidal frequency is approximately related to the sinusoidal and noise power in a simple linear manner with the coefficients being dependent on the model order. Modeling voiced speech as a sum of harmonic signals, we then use the aforementioned relationship along with a least squares approach to combine a family of MVDR estimates (MVDR estimates of different orders) and develop a robust approach for modeling voiced speech. Experimental results of spectral estimation of sinusoids, synthetic vowels, and actual speech signals at SNR of 0 dB and 5 dB using this approach indicate an increased resolution in the estimated MVDR spectra. The MFCC computed from the MVDR spectra using this approach are also used for speaker identification experiments on the TEMIT database at various SNR. The results indicate a reasonable improvement in recognition performance when compared to the MFCC and the fixed order MVDR-MFCC. Rajesh M. Hegde, Yuzhe Jin, Bhaskar D. Rao |
ICASSP (4) | 1 |
| 2007 | Significance of the Modified Group Delay Feature in Speech RecognitionabstractSpectral representation of speech is complete when both the Fourier transform magnitude and phase spectra are specified. In conventional speech recognition systems, features are generally derived from the short-time magnitude spectrum. Although the importance of Fourier transform phase in speech perception has been realized, few attempts have been made to extract features from it. This is primarily because the resonances of the speech signal which manifest as transitions in the phase spectrum are completely masked by the wrapping of the phase spectrum. Hence, an alternative to processing the Fourier transform phase, for extracting speech features, is to process the group delay function which can be directly computed from the speech signal. The group delay function has been used in earlier efforts, to extract pitch and formant information from the speech signal. In all these efforts, no attempt was made to extract features from the speech signal and use them for speech recognition applications. This is primarily because the group delay function fails to capture the short-time spectral structure of speech owing to zeros that are close to the unit circle in the z-plane and also due to pitch periodicity effects. In this paper, the group delay function is modified to overcome these effects. Cepstral features are extracted from the modified group delay function and are called the modified group delay feature (MODGDF). The MODGDF is used for three speech recognition tasks namely, speaker, language, and continuous-speech recognition. Based on the results of feature and performance evaluation, the significance of the MODGDF as a new feature for speech recognition is discussed Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Speech Processing Using Joint Features Derived from the Modified Group Delay FunctionabstractThe paper discusses the significance of joint cepstral features derived from the modified group delay function and MFCC in speech processing. We start with a definition of cepstral features derived from the modified group delay function called the modified group delay feature (MODGDF) which is derived from the Fourier transform phase. Robustness issues like similarities of the MODGDF to RASTA and cepstral mean subtraction are discussed. The efficiency with which formants can be reconstructed for noisy cellular speech using joint features derived from early fusion is illustrated. The joint features are used for four speech processing tasks phoneme, syllable, speaker, and language recognition. Based on the results of analysis and performance evaluation, the significance of joint features derived from the MODGDF and MFCC are discussed. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
ICASSP (1) | 1 |
| 2004 | Application of the modified group delay function to speaker identification and discriminationabstractIn this paper, we explore new methods by which speakers can be identified and discriminated, using features derived from the Fourier transform phase. The modified group delay feature (MODGDF) which is a parameterized form of the modified group delay function is used as a front end feature in this study. A Gaussian mixture model (GMM) based speaker identification system is built with the MODGDF as the front end feature. The system is tested on both clean (TIMIT) and noisy telephone (NTIMIT) speech. The results obtained are compared with traditional Mel frequency cepstral coefficients (MFCC) which is derived from the Fourier transform magnitude. When both MFCC and MODGDF were combined, the performance improved by about 4% indicating that both phase and magnitude contain complementary information. In an earlier paper (Murthy et al. (2003)), it was shown that the MODGDF does possess phoneme specific characteristics. In this paper we show that the MODGDF has speaker specific properties. We also make an attempt to understand speaker discriminating characteristics of the MODGDF using the nonlinear mapping technique based on Sammon mapping (Sammon (1969)) and find that the MODGDF empirically demonstrates a certain level of linear separability among speakers. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
ICASSP (1) | 1 |
| 2004 | Cluster and Intrinsic Dimensionality Analysis of the Modified Group Delay Feature for Speaker Classification
Rajesh M. Hegde, Hema A. Murthy |
ICONIP | 1 |
| 2004 | Continuous speech recognition using joint features derived from the modified group delay function and MFCCabstractFeature extraction and selection for continuous speech recognition is a complex task. State of the art speech recognition systems use features that are derived by ignoring the Fourier transform phase. In our earlier studies we have shown the efficacy of The Modified Group Delay Feature (MODGDF) derived from the Fourier transform phase for phoneme, syllable and speaker recognition. In this paper we use the MODGDF and the popular MFCC derived from Fourier transform magnitude to compute joint features for continuous speech recognition of two Indian languages Tamil and Telugu. A novel method of segmentation of the continuous speech signal into syllable like units followed by isolated style recognition using HMMs is used. We further use an innovative technique which transforms the problem of detecting the correct string of syllabic units with maximum likelihood to finding an optimal state sequence locally. The recognition system does not use any language models. The MODGDF gave promising recognition performance for the two languages and compared well with the MFCC. Joint features derived using MODGDF and MFCC gave a 10.6% improvement for both Tamil and Telugu languages. The improvement reinforces the hypothesis that MODGDF captures complementary information to that of the MFCC and can be used along with the MFCC to capture the complete information in the speech signal at functional level and help in avoiding heavy auditory and language models. Rajesh M. Hegde, Hema A. Murthy, Venkata Ramana Rao Gadde |
INTERSPEECH | 1 |
| 2004 | The modified group delay feature: a new spectral representation of speechabstractAutomatic recognition of speech by machines begins with extraction of meaningful features from the speech signal. Conventional features like the MFCC are derived from the Fourier transform magnitude spectrum, while totally ignoring the phase spectrum. The importance of the Modified group delay feature (MODGDF) derived from the Fourier transform phase spectrum for speaker and phoneme recognition has been presented in our previous efforts. In this paper we try to analyse the feature theoretically and provide justifications in terms of de-correlation, robustness to convolutional and white noise, cluster structures, separability in lower dimensional space, task independence and class separability. The results of speaker identification and continuous speech recognition using the MODGDF as the front end are also presented. Joint features derived from the MODGDF and MFCC gave significant improvements in recognition performance for both speaker and continuous speech recognition tasks. Using the analytical results in the first half of the paper and the results of performance evaluation in the second half, the MODGDF is proposed as an alternative spectral representation of speech. Hema A. Murthy, Rajesh M. Hegde, Venkata Ramana Rao Gadde |
INTERSPEECH | 2 |
| 2003 | Segmentation of speech into syllable-like unitsabstractIn the development of a syllable-centric ASR system, segmentation of the acoustic signal into syllabic units is an important stage. This paper presents a minimum phase group delay based approach to segment spontaneous speech into syllablelike units. Here, three different minimum phase signals are derived from the short term energy functions of three sub-bands of speech signals, as if it were a magnitude spectrum. The experiments are carried out on Switchboard and OGI-MLTS corpus and the error in segmentation is found to be utmost 40msec for 85% of the syllable segments. T. Nagarajan 0001, Hema A. Murthy, Rajesh M. Hegde |
INTERSPEECH | 3 |