Shashidhar G. Koolagudi

dblp:74/1930 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-6928-0237ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Evas: a unified deep learning framework for multi-dimensional evaluation of text summaries
Keerthan Kumar T. G., J. Shreyas, Videh Raj Nema, Shruthan Radhakrishna, Kinshuk Kashyap, Shashidhar G. Koolagudi
Knowl. Inf. Syst.6
2026 Use of timbre features from speech for speaker recognition in emotional environment
Shalini Tomar, Shashidhar G. Koolagudi
Speech Commun.2
2025 Blended-emotional speech for Speaker Recognition by using the fusion of Mel-CQT spectrograms feature extraction
Shalini Tomar, Shashidhar G. Koolagudi
Expert Syst. Appl.2
2025 Video forgery localization using inter-frame denoising and intra-frame segmentation
Debanik Banerjee, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi
Multim. Tools Appl.3
2024 End-to-end latent fingerprint enhancement using multi-scale Generative Adversarial Network
Pramukha R. N., Akhila P., Shashidhar G. Koolagudi
Pattern Recognit. Lett.3
2023 NORD: NOde Ranking-based efficient virtual network embedding over single Domain substrate networks
Keerthan Kumar T. G., Sourav Kanti Addya, Anurag Satpathy, Shashidhar G. Koolagudi
Comput. Networks4
2023 Acoustic scene classification using projection Kervolutional neural network
Manjunath Mulimani, Ritika Nandi, Shashidhar G. Koolagudi
Multim. Tools Appl.3
2023 Automatic diagnosis of COVID-19 related respiratory diseases from speech
Kushan Shekhar, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi
Multim. Tools Appl.3
2022 An Improved Transformer Transducer Architecture for Hindi-English Code Switched Speech Recognition
Ansen Antony, Sumanth Reddy Kota, Akhilesh Lade, Spoorthy Venkatesh, Shashidhar G. Koolagudi
INTERSPEECH5
2021 Dialect Identification using Chroma-Spectral Shape Features with Ensemble Technique
Nagaratna B. Chittaragi, Shashidhar G. Koolagudi
Comput. Speech Lang.2
2020 Polyphonic Sound Event Detection Using Transposed Convolutional Recurrent Neural Network
abstract
In this paper we propose a Transposed Convolutional Recurrent Neural Network (TCRNN) architecture for polyphonic sound event recognition. Transposed convolution layer, which caries out a regular convolution operation but reverts the spatial transformation and it is combined with a bidirectional Recurrent Neural Network (RNN) to get TCRNN. Instead of the traditional mel spectrogram features, the proposed methodology incorporates mel-IFgram (Instantaneous Frequency spectrogram) features. The performance of the proposed approach is evaluated on sound events of publicly available TUT-SED 2016 and Joint sound scene and polyphonic sound event recognition datasets. Results show that the proposed approach outperforms state-of-the-art methods.
Chandra Churh Chatterjee, Manjunath Mulimani, Shashidhar G. Koolagudi
ICASSP3
2020 A Deep Neural Network-Driven Feature Learning Method for Polyphonic Acoustic Event Detection from Real-Life Recordings
abstract
In this paper, a Deep Neural Network (DNN)-driven feature learning method for polyphonic Acoustic Event Detection (AED) is proposed. The proposed DNN is a combination of different layers used to characterize multiple overlapped acoustic events in the mixture. During training, DNN is able to learn the optimal set of discriminative spectral characteristics of the overlapped (polyphonic) acoustic events. The performance of the proposed method is evaluated on the TUT Sound Event 2016 (TUT-SED 2016) real-life dataset and joint Acoustic Scene Classification (ASC) and polyphonic AED dataset. Results show that proposed approach outperforms the state-of-the-art methods.
Manjunath Mulimani, Akash B. Kademani, Shashidhar G. Koolagudi
ICASSP3
2020 Semantic-Preserving Image Compression
abstract
Video traffic comprises a large majority of the total traffic on the internet today. Uncompressed visual data requires a very large data rate; lossy compression techniques are employed in order to keep the data-rate manageable. Increasingly, a significant amount of visual data being generated is consumed by analytics (such as classification, detection, etc.) residing in the cloud. Image and video compression can produce visual artifacts, especially at lower data-rates, which can result in a significant drop in performance on such analytic tasks. Moreover, standard image and video compression techniques aim to optimize perceptual quality for human consumption by allocating more bits to perceptually significant features of the scene. However, these features may not necessarily be the most suitable ones for semantic tasks. We present here an approach to compress visual data in order to maximize performance on a given analytic task. We train a deep auto-encoder using a multi-task loss to learn the relevant embeddings. An approximate differentiable model of the quantizer is used during training which helps boost the accuracy during inference. We apply our approach on an image classification problem and show that for a given level of compression, it achieves higher classification accuracy than that obtained by performing classification on images compressed using JPEG. Our approach also outperforms the relevant state-of-the-art approach by a significant margin.
Neel Patwa, Nilesh A. Ahuja, V. Srinivasa Somayazulu, Omesh Tickoo, Srenivas Varadarajan, Shashidhar G. Koolagudi
ICIP6
2020 Classification of aspirated and unaspirated sounds in speech using excitation and signal level information
Pravin B. Ramteke, Sujata Supanekar, Shashidhar G. Koolagudi
Comput. Speech Lang.3
2019 Locality-Constrained Linear Coding Based Fused Visual Features for Robust Acoustic Event Classification
Manjunath Mulimani, Shashidhar G. Koolagudi
INTERSPEECH2
2019 NITK Kids' Speech Corpus
Pravin B. Ramteke, Sujata Supanekar, Pradyoth Hegde, Hanna Nelson, Venkataraja Aithal, Shashidhar G. Koolagudi
INTERSPEECH6
2019 A Novel Approach to Video Steganography using a 3D Chaotic Map
abstract
In this paper, we introduce a novel approach for data-hiding in videos using 3-dimensional Chaotic Maps. A video is represented as a 3-dimensional image, with the third axis constituting the frames of the video. Existing chaotic map based data-hiding techniques on videos is confined to applications of 2-dimensional chaotic maps on a per-frame basis. In this paper, a 3-dimensional extension of the logistic chaos map is applied to identify pixels to encode information in the video's 3-dimensional space and 3-3-2 Least Significant Bit (LSB) substitution is used to encode 1 byte of information per pixel. We have implemented and presented a proof of concept that has been analyzed on a test video using various quality metrics. The chaotic map based data-hiding approach proposed in this paper is shown to be secured and the results observed are inline with the standard results for a video steganographic algorithm using LSB substitution.
Gurupungav Narayanan, Rishika Narayanan, Nihal Haneef, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi
TENCON5
2019 An approach for Mridanga stroke transcription in Carnatic music using HGCC
abstract
Mridanga is a percussion instrument used in Carnatic music, it is a two sided drum. Stroke is a process of striking the drum membrane leading to a unique sound. Stroke transcription is a process to identify and label different beat sounds produced by the percussion instrument in a track. It is an essential feature for music information retrieval (MIR) and auto content creation. In this paper a novel approach to Mridanga stroke transcription is proposed. Mridanga stroke transcription is similar to speech recognition in which, the approach is to use Mel-Frequency Cepstral Coefficients (MFCC) features, or variation of MFCC. To increase the classification in stroke transcription in Mridanga a new feature extraction method called Harmonic Grouping Cepstral Coefficient(HGCC) is introduced. The newly introduced method follows steps similar to MFCC during extraction the deviation lies in filters used for extraction. The proposed approach displays an accuracy of 80% for signal to noise ratio range of 10dB-40 dB a marginal gain to existing baseline MFCC.
Vishnu G. Swaroop, Shashidhar G. Koolagudi
TENCON2
2019 Segmentation and characterization of acoustic event spectrograms using singular value decomposition
Manjunath Mulimani, Shashidhar G. Koolagudi
Expert Syst. Appl.2
2019 Phoneme boundary detection from speech: A rule based approach
Pravin B. Ramteke, Shashidhar G. Koolagudi
Speech Commun.2
2018 Robust Acoustic Event Classification Using Bag-of-Visual-Words
Manjunath Mulimani, Shashidhar G. Koolagudi
INTERSPEECH2
2018 Robust Dialect Identification System using Spectro-Temporal Gabor Features
abstract
Automatic identification of dialects of a language is gaining popularity in the field of automatic speech recognition (ASR) systems. The present work proposes an automatic dialect identification (ADI) system using 2D Gabor and spectral features. A comprehensive study of the five dialects of a Dravidian Kannada language has been taken up. Gabor filters representing spectro-temporal modulations attempt in emulation of the human auditory system concerning signal processing strategies. Hence, they are able to well perceive human voices in tern recognize dialectal variations effectively. Also, spectral features Mel frequency cepstral coefficients (MFCC) are derived. A single classifier based support vector machine (SVM) and ensemble based extreme random forest (ERF) classification methods are employed for recognition. The effectiveness of the Gabor features for ADI system is demonstrated with proposed Kannada dialect dataset along with a standard intonation variation in English (IViE) dataset for British English dialects. The Gabor features have shown better performance over MFCC features with both datasets. Better recognition performance of 88.75% and 99.16% is achieved with Kannada and IViE dialect datasets respectively. Proposed Gabor features have demonstrated better performances even under noisy conditions.
Nagaratna B. Chittaragi, Siva Krishna P. Mothukuri, Pradyoth Hegde, Shashidhar G. Koolagudi
TENCON4
2018 Sobriety Testing Based on Thermal Infrared Images Using Convolutional Neural Networks
abstract
This paper proposes a method to test the sobriety of an individual using infrared images of the persons eyes, face, hand, and facial profile. The database we used consisted of images of forty different individuals. The process is broken down into two main stages. In the first stage, the data set was divided according to body part and each one was run through its own Convolutional Neural Network (CNN). We then tested the resulting network against a validation data set. The results obtained gave us an indication of which body parts were better suited for identifying signs of drunken state and sobriety. In the second stage, we took the weights of CNN giving best validation accuracy from the first stage. We then grouped the body parts according to the person they belong to. The body parts were fed together into a CNN using the weights obtained in the first stage. The result for each body part was passed to a simple back-propagation neural network (BPNN) to get final results. We tried to identify the most optimal configuration of neural networks for each stage of the process. The results we obtained showed that facial profile images tend to give very good indications of sobriety. The results also showed that combining the results of multiple body parts using a simple BPNN gives a higher accuracy than that of individual ones.
Aditya K. Kamath, A. Tarun Karthik, Leslie Monis, Manjunath Mulimani, Shashidhar G. Koolagudi
TENCON5
2018 Acoustic Event Classification Using Spectrogram Features
abstract
This paper investigates a new feature extraction method to extract different features from the spectrogram of an audio signal for Acoustic Event Classification (AEC). A new set of features is formulated and extracted from local spectrogram regions named blocks. The average recognition performance of proposed spectrogram based features and Mel-frequency cepstral coefficients (MFCCs) with their deltas and accelerations on Support Vector Machines (SVM) is compared. In this work, different categories of acoustic events are considered from the Freiburg-106 dataset. Proposed features show significantly improved performance over conventional Mel-frequency cepstral coefficients (MFCCs) for Acoustic Event Classification.
Manjunath Mulimani, Shashidhar G. Koolagudi
TENCON2
2018 Reconstruction of Edges from Fan-Beam Projections
abstract
The goal of computerised tomography is to reconstruct cross sectional image of the object under consideration from it's projections whereas edge detection is an image analysis problem of utmost importance in medical imaging to outline the boundaries of tumours, bones etc. In this paper, a technique to reconstruct the edges directly from fan-beam projections, using the Marr-Hildreth operator, is presented. To obtain the edge map of object under consideration, the divergent beam transform of Marr-Hildreth operator is convolved with ramp filter to yield an edge reconstruction filter which is finally convolved with the acquired fan-beam projections and back-projected, resulting in a convolution back-projection, to reconstruct the edges. The paper also discusses about the utilisation of state-of-the-art Noo's algorithm to reconstruct the edges directly from equi-angular fan beam projections. Finally, the proposed technique is simulated to make relevant conclusions and inferences.
Adapa Venkata Narasimhadhan, Shashidhar G. Koolagudi, G. V. S. S. K. R. Naganjaneyulu, Sure Avinash, Vinay Peddireddy, N. Bal Kishan, Jeny Rajan
TENCON3
2018 Classification of vocal and non-vocal segments in audio clips using genetic algorithm based feature selection (GAFS)
Vishnu Srinivasa Murthy Yarlagadda, Shashidhar G. Koolagudi
Expert Syst. Appl.2
2017 Image Processing Approach to Diagnose Eye Diseases
M. Prashasthi, K. S. Shravya, Ankit Deepak, Manjunath Mulimani, Shashidhar G. Koolagudi
ACIIDS (2)5
2015 Scalable and fair forwarding of elephant and mice traffic in software defined networks
Saumya Hegde, Shashidhar G. Koolagudi, Swapan Bhattacharya
Comput. Networks2
2011 Recognition of emotions from video using neural network models
K. Sreenivasa Rao, V. K. Saroj, Sudhamay Maity, Shashidhar G. Koolagudi
Expert Syst. Appl.4