VLDB 2026 Research / reviewers in the wild / expert
Rajeev Rajan
dblp:136/5282
· DBLP profile ↗
22ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0001-5488-9026ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 12 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Influence of Oropharyngeal Esophageal Cavity Geometry and Beak Angle on Vocal Tract Resonance of Birds using Computational ModelingabstractBirds produce complex vocalizations through coordinated motor systems, with the syrinx producing sound and the trachea, beak, and other structures modifying it. The bird’s vocal tract includes resonating cavities such as the oropharyngeal esophageal cavity (OEC), which shapes the sound produced by the syrinx. We developed computational models of a bird’s upper vocal tract based on micro-CT scans, including three models. The first is the Primary Model, while the others are Refined Models with two different OEC geometries: a cylinder (RMC) and a truncated cone (RMT). Using COMSOL Multiphysics, we analyzed both the impact of the OEC on resonance and the influence of different OEC shapes on sound production. Validation with bird’s vocalizations showed that the truncated conical OEC most accurately matched the observed resonant frequencies. Noumida Abdul Kareem, Rajeev Rajan |
ICASSP | 2 |
| 2025 | Analysis of Avian Biphonic Vocalization Using Computational Modelling
Noumida A, Rajeev Rajan |
INTERSPEECH | 2 |
| 2025 | A Siamese Network-Based Framework for Voice Mimicry Proficiency Assessment Using X-Vector Embeddings
Bhasi K. C., Rajeev Rajan |
INTERSPEECH | 2 |
| 2025 | Focal Modulation Network: A Novel Solution for Polyphonic Music Instrument Recognition without Attention and Aggregation Strategy
Lekshmi Chandrika Reghunath, Rajeev Rajan |
INTERSPEECH | 2 |
| 2025 | Binary-class concrete surface crack detection using a transfer learning model
R. Ritzy, Umadevi V. A., K. Girija, Rajeev Rajan |
Knowl. Based Syst. | 4 |
| 2025 | Oktoechos classification in liturgical music using self attention based-stacked bi-directional networks
Rajeev Rajan, A. Noumida, T. V. Hridya Raj |
Multim. Tools Appl. | 1 |
| 2025 | Identifying overlapping bird species from raw field audio recordings by assembling Grouped Channel Feature Attention with multi-scale residual CBAM
Noumida Abdul Kareem, Rajeev Rajan |
Neural Comput. Appl. | 2 |
| 2025 | Beyond transformers: hierarchical contextualization and gated aggregation for multiple predominant instrument recognition in polyphonic music
Lekshmi Chandrika Reghunath, Rajeev Rajan |
J. Supercomput. | 2 |
| 2024 | Attention-augmented X-vectors for the Evaluation of Mimicked Speech Using Sparse Autoencoder-LSTM framework
Bhasi K. C., Rajeev Rajan, Noumida Abdul Kareem |
INTERSPEECH | 2 |
| 2024 | Multi-label Bird Species Classification from Field Recordings using Mel_Graph-GCN Framework
Noumida Abdul Kareem, Rajeev Rajan |
INTERSPEECH | 2 |
| 2024 | Automatic music mood classification using multi-modal attention framework
Sujeesha A. S., Mala J. B., Rajeev Rajan |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A review on tonic estimation algorithms in indian art music
Aiswarya M. A., Rajeev Rajan |
Multim. Tools Appl. | 2 |
| 2023 | Paraconsistent Feature Analysis for the Competency Evaluation of Voice ImpersonationabstractMimicry is the art of imitating public figures, animals, and instruments. It’s a popular profession in entertainment, where skilled individuals synchronize voice and gestures, but it is also used deceptively in voice-based crimes. We conducted a paraconsistent feature analysis (PFE) on acoustic features to gain a better understanding of their role in detecting highly competent voice impersonators. To validate our observations from the feature analysis, we performed an experiment on a mimicry dataset using four classifiers. Our proposed system evaluates the competence of artists in voice mimicking and ranks them based on scores obtained from a classifier. A ’hit’ is recorded when the system identifies the artist with the highest mean opinion score as rank-1. Performance is evaluated using the top X-hit criteria on the mimicry dataset. Through our experiments, we demonstrate the potential of PFE in selecting optimal features for assessing the best voice-mimicking candidate. Rajeev Rajan, Noumida Abdul Kareem, Sreelakshmi S |
ASRU | 1 |
| 2023 | Statistical Analysis of Speech Disorder Specific Features to Characterise Dysarthria Severity LevelabstractPoor coordination of the speech production subsystems due to any neurological injury or a neuro-degenerative disease leads to dysarthria, a neuro-motor speech disorder. Dysarthric speech impairments can be mapped to the deficits caused in phonation, articulation, prosody, and glottal functioning. With the aim of reducing the subjectivity in clinical evaluations, many automated systems are proposed in the literature to assess the dysarthria severity level using these features. This work aims to analyse the suitability of these features in determining the severity level. A detailed investigation is done to rank these features for their efficacy in modelling the pathological aspects of dysarthric speech, using the technique of paraconsistent feature engineering. The study used two dysarthric speech databases, UA-Speech and TORGO. It puts light into the fact that both the prosody and articulation features are best useful for dysarthria severity estimation, which was supported by the classification accuracies obtained on using different machine learning classifiers. Amlu Anna Joshy, P. N. Parameswaran, Siddharth R. Nair, Rajeev Rajan |
ICASSP | 4 |
| 2023 | Dysarthria severity classification using multi-head attention and multi-task learning
Amlu Anna Joshy, Rajeev Rajan |
Speech Commun. | 2 |
| 2022 | Oktoechos Classification in Liturgical Music Using SBU-LSTM/GRU
Rajeev Rajan, Ananya Ayasi |
INTERSPEECH | 1 |
| 2021 | Distance Metric Learnt Kernel-Based Music Classification Using Timbral DescriptorsabstractAutomatic music genre classification based on distance metric learning (DML) is proposed in this paper. Three types of timbral descriptors, namely, mel-frequency cepstral coefficient (MFCC) features, modified group delay features (MODGDF) and low-level timbral feature sets are combined at the feature level. We experimented with k nearest neighbor (kNN) and support vector machine (SVM)-based classifiers for standard and DML kernels (DMLK) using GTZAN and Folk music dataset. Standard kernel-based kNN and SVM-based classifiers report classification accuracy (in%) of 79.03 and 90.16, respectively, on GTZAN dataset and 86.60 and 92.26, respectively, for Folk music dataset, with the best performing RBF kernel. A further improvement was observed when DML kernels were used in place of standard kernels in the kernel kNN and SVM-based classifiers with an accuracy of 84.46%, 92.74% (GTZAN), 90.00 and 96.23 (Folk music dataset) for DMLK-kNN and DMLK-SVM, respectively. The results demonstrate the potential of DML kernels in music genre classification task. Rajeev Rajan, B. S. Shajee Mohan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2020 | Competency Evaluation in Voice Mimicking Using Acoustic Cues
Abhijith Girish, Adharsh Sabu, Akshay Prasannan Latha, Rajeev Rajan |
INTERSPEECH | 4 |
| 2020 | Poetic Meter Classification Using i-Vector-MTF Fusion
Rajeev Rajan, Aiswarya Vinod Kumar, Ben P. Babu |
INTERSPEECH | 1 |
| 2019 | Design and Development of a Multi-Lingual Speech Corpora (TaMaR-EmoDB) for Emotion Analysis
Rajeev Rajan, Haritha U. G., Sujitha A. C., Rejisha T. M. |
INTERSPEECH | 1 |
| 2017 | Two-pitch tracking in co-channel speech using modified group delay functions
Rajeev Rajan, Hema A. Murthy |
Speech Commun. | 1 |
| 2013 | Group delay based melody monopitch extraction from musicabstractIn this paper, we propose a modified group delay based method for melodic pitch extraction from heterophonic music. The power spectrum of the music signal is first flattened in order that the system characteristics are annihilated, while the characteristics of the source are emphasized. The modified group delay function of this signal produces peaks at multiples of the pitch period. The first 3 peaks are used to determine the actual pitch period. The performance of the proposed system was evaluated on two datasets ADC-2004, and LabROSA. The performance is comparable to that of other magnitude spectrum based approaches. The algorithms are also applied to heterophonic music, namely Carnatic Music. As ground truth is not available for Carnatic Music, the pitch contours were used to synthesize the music, which was evaluated for correctness by a professional musician. Rajeev Rajan, Hema A. Murthy |
ICASSP | 1 |