VLDB 2026 Research / reviewers in the wild / expert
Shashidhar G. Koolagudi
dblp:74/1930
· DBLP profile ↗
29ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-6928-0237ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evas: a unified deep learning framework for multi-dimensional evaluation of text summaries
Keerthan Kumar T. G., J. Shreyas, Videh Raj Nema, Shruthan Radhakrishna, Kinshuk Kashyap, Shashidhar G. Koolagudi |
Knowl. Inf. Syst. | 6 |
| 2026 | Use of timbre features from speech for speaker recognition in emotional environment
Shalini Tomar, Shashidhar G. Koolagudi |
Speech Commun. | 2 |
| 2025 | Blended-emotional speech for Speaker Recognition by using the fusion of Mel-CQT spectrograms feature extraction
Shalini Tomar, Shashidhar G. Koolagudi |
Expert Syst. Appl. | 2 |
| 2025 | Video forgery localization using inter-frame denoising and intra-frame segmentation
Debanik Banerjee, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi |
Multim. Tools Appl. | 3 |
| 2024 | End-to-end latent fingerprint enhancement using multi-scale Generative Adversarial Network
Pramukha R. N., Akhila P., Shashidhar G. Koolagudi |
Pattern Recognit. Lett. | 3 |
| 2023 | NORD: NOde Ranking-based efficient virtual network embedding over single Domain substrate networks
Keerthan Kumar T. G., Sourav Kanti Addya, Anurag Satpathy, Shashidhar G. Koolagudi |
Comput. Networks | 4 |
| 2023 | Acoustic scene classification using projection Kervolutional neural network
Manjunath Mulimani, Ritika Nandi, Shashidhar G. Koolagudi |
Multim. Tools Appl. | 3 |
| 2023 | Automatic diagnosis of COVID-19 related respiratory diseases from speech
Kushan Shekhar, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi |
Multim. Tools Appl. | 3 |
| 2022 | An Improved Transformer Transducer Architecture for Hindi-English Code Switched Speech Recognition
Ansen Antony, Sumanth Reddy Kota, Akhilesh Lade, Spoorthy Venkatesh, Shashidhar G. Koolagudi |
INTERSPEECH | 5 |
| 2021 | Dialect Identification using Chroma-Spectral Shape Features with Ensemble Technique
Nagaratna B. Chittaragi, Shashidhar G. Koolagudi |
Comput. Speech Lang. | 2 |
| 2020 | Polyphonic Sound Event Detection Using Transposed Convolutional Recurrent Neural NetworkabstractIn this paper we propose a Transposed Convolutional Recurrent Neural Network (TCRNN) architecture for polyphonic sound event recognition. Transposed convolution layer, which caries out a regular convolution operation but reverts the spatial transformation and it is combined with a bidirectional Recurrent Neural Network (RNN) to get TCRNN. Instead of the traditional mel spectrogram features, the proposed methodology incorporates mel-IFgram (Instantaneous Frequency spectrogram) features. The performance of the proposed approach is evaluated on sound events of publicly available TUT-SED 2016 and Joint sound scene and polyphonic sound event recognition datasets. Results show that the proposed approach outperforms state-of-the-art methods. Chandra Churh Chatterjee, Manjunath Mulimani, Shashidhar G. Koolagudi |
ICASSP | 3 |
| 2020 | A Deep Neural Network-Driven Feature Learning Method for Polyphonic Acoustic Event Detection from Real-Life RecordingsabstractIn this paper, a Deep Neural Network (DNN)-driven feature learning method for polyphonic Acoustic Event Detection (AED) is proposed. The proposed DNN is a combination of different layers used to characterize multiple overlapped acoustic events in the mixture. During training, DNN is able to learn the optimal set of discriminative spectral characteristics of the overlapped (polyphonic) acoustic events. The performance of the proposed method is evaluated on the TUT Sound Event 2016 (TUT-SED 2016) real-life dataset and joint Acoustic Scene Classification (ASC) and polyphonic AED dataset. Results show that proposed approach outperforms the state-of-the-art methods. Manjunath Mulimani, Akash B. Kademani, Shashidhar G. Koolagudi |
ICASSP | 3 |
| 2020 | Semantic-Preserving Image CompressionabstractVideo traffic comprises a large majority of the total traffic on the internet today. Uncompressed visual data requires a very large data rate; lossy compression techniques are employed in order to keep the data-rate manageable. Increasingly, a significant amount of visual data being generated is consumed by analytics (such as classification, detection, etc.) residing in the cloud. Image and video compression can produce visual artifacts, especially at lower data-rates, which can result in a significant drop in performance on such analytic tasks. Moreover, standard image and video compression techniques aim to optimize perceptual quality for human consumption by allocating more bits to perceptually significant features of the scene. However, these features may not necessarily be the most suitable ones for semantic tasks. We present here an approach to compress visual data in order to maximize performance on a given analytic task. We train a deep auto-encoder using a multi-task loss to learn the relevant embeddings. An approximate differentiable model of the quantizer is used during training which helps boost the accuracy during inference. We apply our approach on an image classification problem and show that for a given level of compression, it achieves higher classification accuracy than that obtained by performing classification on images compressed using JPEG. Our approach also outperforms the relevant state-of-the-art approach by a significant margin. Neel Patwa, Nilesh A. Ahuja, V. Srinivasa Somayazulu, Omesh Tickoo, Srenivas Varadarajan, Shashidhar G. Koolagudi |
ICIP | 6 |
| 2020 | Classification of aspirated and unaspirated sounds in speech using excitation and signal level information
Pravin B. Ramteke, Sujata Supanekar, Shashidhar G. Koolagudi |
Comput. Speech Lang. | 3 |
| 2019 | Locality-Constrained Linear Coding Based Fused Visual Features for Robust Acoustic Event Classification
Manjunath Mulimani, Shashidhar G. Koolagudi |
INTERSPEECH | 2 |
| 2019 | NITK Kids' Speech Corpus
Pravin B. Ramteke, Sujata Supanekar, Pradyoth Hegde, Hanna Nelson, Venkataraja Aithal, Shashidhar G. Koolagudi |
INTERSPEECH | 6 |
| 2019 | A Novel Approach to Video Steganography using a 3D Chaotic MapabstractIn this paper, we introduce a novel approach for data-hiding in videos using 3-dimensional Chaotic Maps. A video is represented as a 3-dimensional image, with the third axis constituting the frames of the video. Existing chaotic map based data-hiding techniques on videos is confined to applications of 2-dimensional chaotic maps on a per-frame basis. In this paper, a 3-dimensional extension of the logistic chaos map is applied to identify pixels to encode information in the video's 3-dimensional space and 3-3-2 Least Significant Bit (LSB) substitution is used to encode 1 byte of information per pixel. We have implemented and presented a proof of concept that has been analyzed on a test video using various quality metrics. The chaotic map based data-hiding approach proposed in this paper is shown to be secured and the results observed are inline with the standard results for a video steganographic algorithm using LSB substitution. Gurupungav Narayanan, Rishika Narayanan, Nihal Haneef, Nagaratna B. Chittaragi, Shashidhar G. Koolagudi |
TENCON | 5 |
| 2019 | An approach for Mridanga stroke transcription in Carnatic music using HGCCabstractMridanga is a percussion instrument used in Carnatic music, it is a two sided drum. Stroke is a process of striking the drum membrane leading to a unique sound. Stroke transcription is a process to identify and label different beat sounds produced by the percussion instrument in a track. It is an essential feature for music information retrieval (MIR) and auto content creation. In this paper a novel approach to Mridanga stroke transcription is proposed. Mridanga stroke transcription is similar to speech recognition in which, the approach is to use Mel-Frequency Cepstral Coefficients (MFCC) features, or variation of MFCC. To increase the classification in stroke transcription in Mridanga a new feature extraction method called Harmonic Grouping Cepstral Coefficient(HGCC) is introduced. The newly introduced method follows steps similar to MFCC during extraction the deviation lies in filters used for extraction. The proposed approach displays an accuracy of 80% for signal to noise ratio range of 10dB-40 dB a marginal gain to existing baseline MFCC. Vishnu G. Swaroop, Shashidhar G. Koolagudi |
TENCON | 2 |
| 2019 | Segmentation and characterization of acoustic event spectrograms using singular value decomposition
Manjunath Mulimani, Shashidhar G. Koolagudi |
Expert Syst. Appl. | 2 |
| 2019 | Phoneme boundary detection from speech: A rule based approach
Pravin B. Ramteke, Shashidhar G. Koolagudi |
Speech Commun. | 2 |
| 2018 | Robust Acoustic Event Classification Using Bag-of-Visual-Words
Manjunath Mulimani, Shashidhar G. Koolagudi |
INTERSPEECH | 2 |
| 2018 | Robust Dialect Identification System using Spectro-Temporal Gabor FeaturesabstractAutomatic identification of dialects of a language is gaining popularity in the field of automatic speech recognition (ASR) systems. The present work proposes an automatic dialect identification (ADI) system using 2D Gabor and spectral features. A comprehensive study of the five dialects of a Dravidian Kannada language has been taken up. Gabor filters representing spectro-temporal modulations attempt in emulation of the human auditory system concerning signal processing strategies. Hence, they are able to well perceive human voices in tern recognize dialectal variations effectively. Also, spectral features Mel frequency cepstral coefficients (MFCC) are derived. A single classifier based support vector machine (SVM) and ensemble based extreme random forest (ERF) classification methods are employed for recognition. The effectiveness of the Gabor features for ADI system is demonstrated with proposed Kannada dialect dataset along with a standard intonation variation in English (IViE) dataset for British English dialects. The Gabor features have shown better performance over MFCC features with both datasets. Better recognition performance of 88.75% and 99.16% is achieved with Kannada and IViE dialect datasets respectively. Proposed Gabor features have demonstrated better performances even under noisy conditions. Nagaratna B. Chittaragi, Siva Krishna P. Mothukuri, Pradyoth Hegde, Shashidhar G. Koolagudi |
TENCON | 4 |
| 2018 | Sobriety Testing Based on Thermal Infrared Images Using Convolutional Neural NetworksabstractThis paper proposes a method to test the sobriety of an individual using infrared images of the persons eyes, face, hand, and facial profile. The database we used consisted of images of forty different individuals. The process is broken down into two main stages. In the first stage, the data set was divided according to body part and each one was run through its own Convolutional Neural Network (CNN). We then tested the resulting network against a validation data set. The results obtained gave us an indication of which body parts were better suited for identifying signs of drunken state and sobriety. In the second stage, we took the weights of CNN giving best validation accuracy from the first stage. We then grouped the body parts according to the person they belong to. The body parts were fed together into a CNN using the weights obtained in the first stage. The result for each body part was passed to a simple back-propagation neural network (BPNN) to get final results. We tried to identify the most optimal configuration of neural networks for each stage of the process. The results we obtained showed that facial profile images tend to give very good indications of sobriety. The results also showed that combining the results of multiple body parts using a simple BPNN gives a higher accuracy than that of individual ones. Aditya K. Kamath, A. Tarun Karthik, Leslie Monis, Manjunath Mulimani, Shashidhar G. Koolagudi |
TENCON | 5 |
| 2018 | Acoustic Event Classification Using Spectrogram FeaturesabstractThis paper investigates a new feature extraction method to extract different features from the spectrogram of an audio signal for Acoustic Event Classification (AEC). A new set of features is formulated and extracted from local spectrogram regions named blocks. The average recognition performance of proposed spectrogram based features and Mel-frequency cepstral coefficients (MFCCs) with their deltas and accelerations on Support Vector Machines (SVM) is compared. In this work, different categories of acoustic events are considered from the Freiburg-106 dataset. Proposed features show significantly improved performance over conventional Mel-frequency cepstral coefficients (MFCCs) for Acoustic Event Classification. Manjunath Mulimani, Shashidhar G. Koolagudi |
TENCON | 2 |
| 2018 | Reconstruction of Edges from Fan-Beam ProjectionsabstractThe goal of computerised tomography is to reconstruct cross sectional image of the object under consideration from it's projections whereas edge detection is an image analysis problem of utmost importance in medical imaging to outline the boundaries of tumours, bones etc. In this paper, a technique to reconstruct the edges directly from fan-beam projections, using the Marr-Hildreth operator, is presented. To obtain the edge map of object under consideration, the divergent beam transform of Marr-Hildreth operator is convolved with ramp filter to yield an edge reconstruction filter which is finally convolved with the acquired fan-beam projections and back-projected, resulting in a convolution back-projection, to reconstruct the edges. The paper also discusses about the utilisation of state-of-the-art Noo's algorithm to reconstruct the edges directly from equi-angular fan beam projections. Finally, the proposed technique is simulated to make relevant conclusions and inferences. Adapa Venkata Narasimhadhan, Shashidhar G. Koolagudi, G. V. S. S. K. R. Naganjaneyulu, Sure Avinash, Vinay Peddireddy, N. Bal Kishan, Jeny Rajan |
TENCON | 3 |
| 2018 | Classification of vocal and non-vocal segments in audio clips using genetic algorithm based feature selection (GAFS)
Vishnu Srinivasa Murthy Yarlagadda, Shashidhar G. Koolagudi |
Expert Syst. Appl. | 2 |
| 2017 | Image Processing Approach to Diagnose Eye Diseases
M. Prashasthi, K. S. Shravya, Ankit Deepak, Manjunath Mulimani, Shashidhar G. Koolagudi |
ACIIDS (2) | 5 |
| 2015 | Scalable and fair forwarding of elephant and mice traffic in software defined networks
Saumya Hegde, Shashidhar G. Koolagudi, Swapan Bhattacharya |
Comput. Networks | 2 |
| 2011 | Recognition of emotions from video using neural network models
K. Sreenivasa Rao, V. K. Saroj, Sudhamay Maity, Shashidhar G. Koolagudi |
Expert Syst. Appl. | 4 |