VLDB 2026 Research / reviewers in the wild / expert
Deepu Vijayasenan
dblp:90/7534
· DBLP profile ↗
25ranked-venue papers
13as first author
6since 2021 · last 2025
0009-0006-4444-7869ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Analysis of Speech Features in Identifying Client's Change Talk in Motivational InterviewingabstractMotivational Interviewing (MI) is a frequently used and effective psychotherapy approach for treating behavioral problems. MI is a collaborative interaction for understanding the client's own reasoning for a change in behavior. In this study, we analyzed an MI corpus in the nutrition and fitness domains, in which counselor and client utterances were categorized using Motivational Interviewing Skill Code (MISC). During the interaction when the client expresses the need or willingness to change is the Change Talk (CT). We aimed to analyze the speech features and proposed a BiLSTM multimodal neural network model for detecting the change talk or not a change talk. Our approach using speech and language information in detecting the change talk is par with other multimodal approaches(language and facial information) with an F1-score of 0.573 for CT. The proposed set of speech features shows the statistical significance in identifying the change talk or not a change talk. Shareef Babu Kalluri, Deepu Vijayasenan |
TENCON | 2 |
| 2024 | The Second DISPLACE Challenge: DIarization of SPeaker and LAnguage in Conversational Environments
Shareef Babu Kalluri, Prachi Singh, Pratik Roy Chowdhuri, Apoorva Kulkarni, Shikha Baghel, Pradyoth Hegde, Swapnil Sontakke, K. T. Deepak, S. R. Mahadeva Prasanna, Deepu Vijayasenan, Sriram Ganapathy |
INTERSPEECH | 10 |
| 2024 | Summary of the DISPLACE challenge 2023-DIarization of SPeaker and LAnguage in Conversational Environments
Shikha Baghel, Shreyas Ramoji, Somil Jain, Pratik Roy Chowdhuri, Prachi Singh, Deepu Vijayasenan, Sriram Ganapathy |
Speech Commun. | 6 |
| 2023 | The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments
Shikha Baghel, Shreyas Ramoji, Sidharth, Ranjana H, Prachi Singh, Somil Jain, Pratik Roy Chowdhuri, Kaustubh Kulkarni, Swapnil Padhi, Deepu Vijayasenan, Sriram Ganapathy |
INTERSPEECH | 10 |
| 2021 | NISP: A Multi-lingual Multi-accent Dataset for Speaker ProfilingabstractMany commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include physical traits of the speaker like height, age, and gender of the speaker along with the native language of the speaker. Many of the datasets available have only partial information for speaker profiling. In this paper, we attempt to overcome this limitation by developing a new dataset which has speech data from five different Indian languages along with English. The meta-data information for speaker profiling applications like linguistic information, regional information, and physical characteristics of a speaker are also collected. We call this dataset as NITK-IISc Multilingual Multi-accent Speaker Profiling (NISP) dataset. The description of the dataset, potential applications, and baseline results for speaker profiling on this dataset are provided in this paper. Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy, Ragesh Rajan M, Prashant Krishnan V |
ICASSP | 2 |
| 2021 | COVID-19 Detection from Spectral Features on the DiCOVA DatasetabstractIn this paper we investigate the cues of COVID-19 on sustained phonation of Vowel-/i/, deep breathing and number counting data of the DiCOVA dataset.We use an ensemble of classifiers trained on different features, namely, super-vectors, formants, harmonics and MFCC features.We fit a two-class Weighted SVM classifier to separate the COVID-19 audio from Non-COVID-19 audio.Weighted penalties help mitigate the challenge of class imbalance in the dataset.The results are reported on the stationary (breathing, Vowel-/i/) and nonstationary(counting data) data using individual and combination of features on each type of utterance.We find that the Formant information plays a crucial role in classification.The proposed system resulted in an AUC score of 0.734 for cross validation, and 0.717 for evaluation dataset. Kotra Venkata Sai Ritwik, Shareef Babu Kalluri, Deepu Vijayasenan |
Interspeech | 3 |
| 2020 | Automatic speaker profiling from short duration speech data
Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy |
Speech Commun. | 2 |
| 2019 | A Deep Neural Network Based End to End Model for Joint Height and Age Estimation from Short Duration SpeechabstractAutomatic height and age prediction of a speaker has a wide variety of applications in speaker profiling, forensics etc. Often in such applications only a few seconds of speech data is available to reliably estimate the speaker parameters. Traditionally, age and height were predicted separately using different estimation algorithms. In this work, we propose a unified DNN architecture to predict both height and age of a speaker for short durations of speech. A novel initialization scheme for the deep neural architecture is introduced, that avoids the requirement for a large training dataset. We evaluate the system on TIMIT dataset where the mean duration of speech segments is around 2.5s. The DNN system is able to improve the age RMSE by at least 0.6 years as compared to a conventional support vector regression system trained on Gaussian Mixture Model mean supervectors. The system achieves an RMSE error of 6.85 and 6.29 cm for male and female height prediction. In case of age estimation, the RMSE errors are 7.60 and 8.63 years for male and female respectively. Analysis of shorter speech segments reveals that even with 1 second speech input the performance degradation is at most 3% compared to the full duration speech files. Shareef Babu Kalluri, Deepu Vijayasenan, Sriram Ganapathy |
ICASSP | 2 |
| 2019 | An Integrated Deep Learning Approach towards Automatic Evaluation of Ki-67 Labeling IndexabstractKi-67 labeling index is a widely used biomarker for the diagnosis and monitoring of cancer. Many automated techniques have been proposed for evaluating Ki-67 index. In this paper, we introduce an integrated deep learning based approach. We use MobileUnet model for segmentation and classification and connected component based algorithm for the estimation of Ki-67 index in bladder cancer cases. The average F1 score is 0.92 and dice score is 0.96. The mean absolute error in the evaluated Ki-67 index is 2.1. We also explore possible pre-processing steps to generalize the segmentation model to at least one another type of cancer. Histogram matching and re-sizing improve the performance in breast cancer data by 12% in F1 score and 8% in dice score. Deepu Vijayasenan, David S. Sumam, Saraswathy Sreeram, Pooja K. Suresh |
TENCON | 2 |
| 2018 | Prediction of Aesthetic Elements in Karnatic Music: A Machine Learning Approach
Ragesh Rajan M, Ashwin Vijayakumar, Deepu Vijayasenan |
INTERSPEECH | 3 |
| 2012 | Speaker diarization of meetings based on large TDOA feature vectorsabstractThis paper investigates the use of large TDOA feature vectors together with acoustic information in speaker diarization of meetings. TDOAs are obtained by considering all possible microphones pairs and this approach is compared with conventional TDOA features extracted w.r.t. a reference channel. The study is carried using two systems, the first based on Gaussian Mixture Modeling and the second based on the Information Bottleneck approach. Results on NIST RT06/RT07/RT09 evaluation datasets show a large speaker error reduction of 30% relative going from 14.3% to 10.8% for the first and from 12.3% to 8.2% for the second whenever the feature weighting is properly handled. Furthermore results reveal that the IB system is more robust to different number of microphones even when all pairs large TDOA vectors are used thus outperforming the HMM/GMM by 25% relative (8.2% error compared to 10.8%). Deepu Vijayasenan, Fabio Valente |
ICASSP | 1 |
| 2012 | DiarTk : An Open Source Toolkit for Research in Multistream Speaker Diarization and its Application to Meetings RecordingsabstractThe speaker diarization task consists of inferring “who spoke when ” in an audio stream without any prior knowledge and has been object of several NIST international evaluation campaigns is last years. A common trend for improving performances has been the use of several different feature streams as diverse as speaker location features, visual features or noise robust acous-tic features. This paper describes an open source toolkit re-leased under GPL license aiming at facilitating research in mul-tistream speaker diarization and reproducing state-of-the-art re-sults. In contrary to other related diarization toolkits, it is ex-plicitly designed to handle an arbitrary number of features with very different statistics while limiting the computational com-plexity. The release includes a set of scripts to replicate bench-mark results on previous NIST evaluations and is intended to provide an easy to use software to study and include novel fea-tures into diarization systems. Deepu Vijayasenan, Fabio Valente |
INTERSPEECH | 1 |
| 2012 | Multistream speaker diarization of meetings recordings beyond MFCC and TDOA features
Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
Speech Commun. | 1 |
| 2011 | Speaker diarization of meetings based on speaker role n-gram modelsabstractSpeaker diarization of meeting recordings is generally based on acoustic information ignoring that meetings are instances of conversations. Several recent works have shown that the sequence of speakers in a conversation and their roles are related and statistically predictable. This paper proposes the use of speaker roles n-gram model to capture the conversation patterns probability and investigates its use as prior information into a state-of-the-art diarization system. Experiments are run on the AMI corpus annotated in terms of roles. The proposed technique reduces the diarization speaker error by 19% when the roles are known and by 17% when they are estimated. Furthermore the paper investigates how the n-gram models generalize to different settings like those from the Rich Transcription campaigns. Experiments on 17 meetings reveal that the speaker error can be reduced by 12% also in this case thus the n-gram can generalize across corpora. Fabio Valente, Deepu Vijayasenan, Petr Motlícek |
ICASSP | 2 |
| 2011 | Multistream speaker diarization through Information Bottleneck system outputs combinationabstractSpeaker diarization of meetings recorded with Multiple Distant Microphones makes extensive use of multiple feature streams like MFCC and Time Delay of Arrivals (TDOA). Typically the combination happens using separate models for each feature stream. This work investigates if the combination of multiple feature streams can happen through the combination of multiple diarization systems performed using those features. The paper extends the previously proposed Information Bottleneck method to handle the combination of several probabilistic diarization outputs. In contrast to the conventional model-based feature combination, this technique is referred as system-based combination. Furthermore the paper introduces an hybrid model-system combination. Experiments are run on data from the Rich Transcription campaigns and show that the system based combination largely outperforms the model based combination by 37% relative. The hybrid approaches improve by 10-20%. The analysis of errors shows that the improvements come from the recordings where the individual MFCC and TDOA systems provide very different performances. Deepu Vijayasenan, Fabio Valente, Petr Motlícek |
ICASSP | 1 |
| 2011 | An Information Theoretic Combination of MFCC and TDOA Features for Speaker DiarizationabstractThis correspondence describes a novel system for speaker diarization of meetings recordings based on the combination of acoustic features (MFCC) and time delay of arrivals (TDOAS). The first part of the paper analyzes differences between MFCC and TDOA features which possess completely different statistical properties. When Gaussian mixture models are used, experiments reveal that the diarization system is sensitive to the different recording scenarios (i.e., meeting rooms with varying number of microphones). In the second part, a new multistream diarization system is proposed extending previous work on information theoretic diarization. Both speaker clustering and speaker realignment steps are discussed; in contrary to current systems, the proposed method avoids to perform the feature combination averaging log-likelihood scores. Experiments on meetings data reveal that the proposed approach outperforms the GMM-based system when the recording is done with varying number of microphones. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Variational Bayesian speaker diarization of meeting recordingsabstractThis paper investigates the use of the Variational Bayesian (VB) framework for speaker diarization of meetings data extending previous related works on Broadcast News audio. VB learning aims at maximizing a bound, known as Free Energy, on the model marginal likelihood and allows joint model learning and model selection according to the same objective function. While the BIC is valid only in the asymptotic limit, the Free Energy is always a valid bound. The paper proposes the use of Free Energy as objective function in speaker diarization. It can be used to select dynamically without any supervision or tuning, elements that typically affect the diarization performance i.e. the inferred number of speakers, the size of the GMM and the initialization. The proposed approach is compared with a conventional state-of-the-art system on the RT06 evaluation data for meeting recordings diarization and shows an improvement of 8.4% relative in terms of speaker error. Fabio Valente, Petr Motlícek, Deepu Vijayasenan |
ICASSP | 3 |
| 2010 | Multistream speaker diarization beyond two acoustic feature streamsabstractSpeaker diarization for meetings data are recently converging towards multistream systems. The most common complementary features used in combination with MFCC are Time Delay of Arrival (TDOA). Also other features have been proposed although, there are no reported improvements on top of MFCC+TDOA systems. In this work we investigate the combination of other feature sets along with MFCC+TDOA. We discuss issues and problems related to the weighting of four different streams proposing a solution based on a smoothed version of the speaker error. Experiments are presented on NIST RT06 meeting diarization evaluation. Results reveal that the combination of four acoustic feature streams results in a 30% relative improvement with respect to the MFCC+TDOA feature combination. To the authors' best knowledge, this is the first successful attempt to improve the MFCC+TDOA baseline including other feature streams. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
ICASSP | 1 |
| 2010 | Advances in fast multistream diarization based on the information bottleneck frameworkabstractMultistream diarization is an effective way to improve the diarization performance, MFCC and Time Delay Of Arrivals (TDOA) being the most commonly used features. This paper extends our previous work on information bottleneck diarization aiming to include large number of features besides MFCC and TDOA while keeping computational costs low. At first HMM/GMM and IB systems are compared in case of two and four feature streams and analysis of errors is performed. Results on a dataset of 17 meetings show that, in spite of comparable oracle performances, the IB system is more robust to feature weight variations. Then a sequential optimization is introduced that further improves the speaker error by 5 − 8% relative. In the last part, computational issues are discussed. The proposed approach is significantly faster and its complexity marginally grows with the number of feature streams running in 0.75 realtime even with four streams achieving a speaker error equal to 6%. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
INTERSPEECH | 1 |
| 2009 | Mutual information based channel selection for speaker diarization of meetings dataabstractIn the meeting case scenario, audio is often recorded using Multiple Distance Microphones (MDM) in a non-intrusive manner. Typically a beamforming is performed in order to obtain a single enhanced signal out of the multiple channels. This paper investigates the use of mutual information for selecting the channel subset that produces the lowest error in a diarization system. Conventional systems perform channel selection on the basis of signal properties such as SNR, cross correlation. In this paper, we propose the use of a mutual information measure that is directly related to the objective function of the diarization system. The proposed algorithms are evaluated on the NIST RT 06 eval dataset. Channel selection improves the speaker error by 1.1% absolute (6.5% relative) w.r.t. the use of all channels. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
ICASSP | 1 |
| 2009 | KL realignment for speaker diarization with multiple feature streamsabstractThis paper aims at investigating the use of Kullback-Leibler (KL) divergence based realignment with application to speaker diarization. The use of KL divergence based realignment operates directly on the speaker posterior distribution estimates and is compared with traditional realignment performed using HMM/GMM system. We hypothesize that using posterior estimates to re-align speaker boundaries is more robust than gaussian mixture models in case of multiple feature streams with different statistical properties. Experiments are run on the NIST RT06 data. These experiments reveal that in case of conventional MFCC features the two approaches yields the same performance while the KL based system outperforms the HMM/GMM re-alignment in case of combination of multiple feature streams (MFCC and TDOA). Index Terms: speaker diarization, information bottleneck, feature combination Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
INTERSPEECH | 1 |
| 2009 | An Information Theoretic Approach to Speaker Diarization of Meeting DataabstractA speaker diarization system based on an information theoretic framework is described. The problem is formulated according to the information bottleneck (IB) principle. Unlike other approaches where the distance between speaker segments is arbitrarily introduced, the IB method seeks the partition that maximizes the mutual information between observations and variables relevant for the problem while minimizing the distortion between observations. This solves the problem of choosing the distance between speech segments, which becomes the Jensen-Shannon divergence as it arises from the IB objective function optimization. We discuss issues related to speaker diarization using this information theoretic framework such as the criteria for inferring the number of speakers, the tradeoff between quality and compression achieved by the diarization system, and the algorithms for optimizing the objective function. Furthermore, we benchmark the proposed system against a state-of-the-art system on the NIST RT06 (rich transcription) data set for speaker diarization of meetings. The IB-based system achieves a diarization error rate of 23.2% compared to 23.6% for the baseline system. This approach being mainly based on nonparametric clustering, it runs significantly faster than the baseline HMM/GMM based system, resulting in faster-than-real-time diarization. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Combination of agglomerative and sequential clustering for speaker diarizationabstractThis paper aims at investigating the use of sequential clustering for speaker diarization. Conventional diarization systems are based on parametric models and agglomerative clustering. In our previous work we proposed a non-parametric method based on the agglomerative information bottleneck for very fast diarization. Here we consider the combination of sequential and agglomerative clustering for avoiding local maxima of the objective function and for purification. Experiments are run on the RT06 eval data. Sequential Clustering with oracle model selection can reduce the speaker error by 10% w.r.t. agglomerative clustering. When the model selection is based on Normalized Mutual Information criterion, a relative improvement of 5% is obtained using a combination of agglomerative and sequential clustering. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
ICASSP | 1 |
| 2008 | Integration of TDOA features in information bottleneck framework for fast speaker diarizationabstractIn this paper we address the combination of multiple feature streams in a fast speaker diarization system for meeting recordings. Whenever Multiple Distant Microphones (MDM) are used, it is possible to estimate the Time Delay of Arrival (TDOA) for different channels. In \\cite{xavi_comb}, it is shown that TDOA can be used as additional features together with conventional spectral features for improving speaker diarization. We investigate here the combination of TDOA and spectral features in a fast diarization system based on the Information Bottleneck principle. We evaluate the algorithm on the NIST RT06 diarization task. Adding TDOA features to spectral features reduces the speaker error by 3\\% absolute. Results are comparable to those of conventional HMM/GMM based systems with consistent reduction in computational complexity. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
INTERSPEECH | 1 |
| 2007 | Agglomerative information bottleneck for speaker diarization of meetings dataabstractIn this paper, we investigate the use of agglomerative information bottleneck (aIB) clustering for the speaker diarization task of meetings data. In contrary to the state-of-the-art diarization systems that models individual speakers with Gaussian mixture models, the proposed algorithm is completely non parametric . Both clustering and model selection issues of non-parametric models are addressed in this work. The proposed algorithm is evaluated on meeting data on the RT06 evaluation data set. The system is able to achieve diarization error rates comparable to state-of-the-art systems at a much lower computational complexity. Deepu Vijayasenan, Fabio Valente, Hervé Bourlard |
ASRU | 1 |