EDBT 2026 Demo / reviewers in the wild / expert
Preeti Rao
dblp:54/790
· DBLP profile ↗
40ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-2842-8450ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Speaker Anonymization for Children's Oral Reading AssessmentabstractSpeaker anonymization aims to modify the speech signal in order to protect the identity of a speaker while preserving the linguistic content. Despite the increasing use of children's voices in educational applications, such as oral reading fluency (ORF) assessment, there is little work on the anonymization aspects. In this work, we investigate the effectiveness of available speaker anonymization methods drawing from traditional speech-production based approaches and a neural codec based method. We investigate the trade-off between privacy protection, measured as the degree of anonymity, and utility preservation, which in the current context of ORF assessment, includes the segmental and suprasegmental features of children’s read speech utterances. We report objective and subjective evaluations using two child-speaker datasets: MPS and SpeechOcean. Our objective evaluation results indicate that the speech-production based method of vocal tract length normalization coupled with pitch-transposition achieves the best balance between privacy and utility. Subjective listening results indicate that naturalness is achievable across methods while the neural method fails to preserve age characteristics, which are more easily controlled by the speech-production driven methods. Sandipan Dhar, Srikanth Raj Chetupalli, Preeti Rao |
AAAI | 3 |
| 2025 | Predicting Prosodic Boundaries for Children's TextsabstractReading fluency in any language requires accurate word decoding but also natural prosodic phrasing i.e the grouping of words into rhythmically and syntactically coherent units.This holds for, both, reading aloud and silent reading.While adults pause meaningfully at clause or punctuation boundaries, children aged 8-13 often insert inappropriate pauses due to limited breath control and underdeveloped prosodic awareness.We present a text-based model to predict cognitively appropriate pause locations in children's reading material.Using a curated dataset of 54 leveled English stories annotated for potential pauses, or prosodic boundaries, by 21 fluent speakers, we find that nearly 30% of pauses occur at non-punctuation locations of the text, highlighting the limitations of using only punctuation-based cues.Our model combines lexical, syntactic, and contextual features with a novel breath duration feature that captures syllable load since the last major boundary.This cognitively motivated approach can model both allowed and "forbidden" pauses.The proposed framework supports applications such as child-directed TTS and oral reading fluency assessment where the proper grouping of words is considered critical to reading comprehension. Mansi Dhamne, Sneha Raman, Preeti Rao |
EMNLP | 3 |
| 2025 | Psycholinguistic Features Predict Word Duration in Hindi Read Aloud SpeechabstractReliable assessment of oral reading fluency (ORF) is of great importance in foundational literacy missions globally. For the design of level appropriate testing passages, text difficulty has traditionally been based on coarse-grained measures of readability like the Flesch–Kincaid score. We present a novel study where we deploy psycholinguistic measures of reading difficulty from Natural Language Processing to predict the duration of words in Hindi read-aloud speech. We test the hypotheses that expectation-based measures of linguistic complexity are significant predictors of word duration in Hindi read-aloud speech. We validate the stated hypotheses by estimating surprisal measures inspired from Surprisal Theory of sentence comprehension and introduce a novel measure of orthographic complexity to model the intricacies of the Hindi script. Cognitive modelling experiments were conducted on a dataset of six Hindi short stories read aloud by 5 expert readers, containing 2 measures of word duration. Our results show that both surprisal as well as the orthographic complexity measures are significant predictors of word duration. In contrast with long words, we find duration reducing with increased orthographic complexity in the case of short words. The variation between individual speakers in terms of word duration is very low and the variance in the data is caused by the properties of the words used in the text. Finally, we reflect on the implications of our work for cognitive models of language production and for ORF assessment. Rajakrishnan Rajkumar, Sneha Raman, Aadya Ranjan, Mildred Pereira, Nagesh Nayak, Preeti Rao |
ICASSP | 6 |
| 2025 | Oral Reading Errors by Grade 3 Children in Indian Schools: A Hindi-English Perspective
Sneha Raman, Preeti Rao |
INTERSPEECH | 2 |
| 2024 | A Dataset and Two-pass System for Reading Miscue Detection
Raj Gothi, Mildred Pereira, Nagesh Nayak, Preeti Rao |
INTERSPEECH | 5 |
| 2024 | Emotion Arithmetic: Emotional Speech Synthesis via Weight Space Interpolation
Pavan Kalyan, Preeti Rao, Preethi Jyothi, Pushpak Bhattacharyya |
INTERSPEECH | 2 |
| 2024 | Predicting children's perceived reading proficiency with prosody modeling
Kamini Sabu, Preeti Rao |
Comput. Speech Lang. | 2 |
| 2023 | Narrator or Character: Voice Modulation in an Expressive Multi-speaker TTS
Tankala Pavan Kalyan, Preeti Rao, Preethi Jyothi, Pushpak Bhattacharyya |
INTERSPEECH | 2 |
| 2022 | Deep Learning for Prominence Detection In Children's Read SpeechabstractThe detection of perceived prominence in speech has attracted approaches ranging from the design of knowledge-based linguistic and acoustic features to the automatic feature learning from suprasegmental attributes such as pitch and intensity contours. We present here, in contrast, a system that operates directly on segmented speech waveforms to learn features relevant to prominent word detection for children’s oral fluency assessment. The chosen CRNN (convolutional recurrent neural network) framework, incorporating both word-level features and sequence information, is found to benefit from the perceptually motivated SincNet filters as the first convolutional layer. We further explore the benefits of the linguistic association between the prosodic events of phrase boundary and prominence with different multi-task architectures. Surpassing the previously reported performance on the same dataset of a random forest ensemble predictor trained on carefully chosen hand-crafted acoustic features, we evaluate further the possibly complementary information from hand-crafted acoustic and pre-trained lexical features. Mithilesh Vaidya, Kamini Sabu, Preeti Rao |
ICASSP | 3 |
| 2021 | The Four-Way Classification of Stops with Voicing and Aspiration for Non-Native Speech Evaluation
Titas Chakraborty, Vaishali Patil, Preeti Rao |
Interspeech | 3 |
| 2021 | Prosodic event detection in children's read speech
Kamini Sabu, Preeti Rao |
Comput. Speech Lang. | 2 |
| 2020 | Speech-To-Singing Conversion in an Encoder-Decoder FrameworkabstractIn this paper our goal is to convert a set of spoken lines into sung ones. Unlike previous signal processing based methods, we take a learning based approach to the problem. This allows us to automatically model various aspects of this transformation, thus overcoming dependence on specific inputs such as high quality singing templates or phoneme-score synchronization information. Specifically, we propose an encoder-decoder framework for our task. Given time-frequency representations of speech and a target melody contour, we learn encodings that enable us to synthesize singing that preserves the linguistic content and timbre of the speaker while adhering to the target melody. We also propose a multi-task learning based objective to improve lyric intelligibility. We present a quantitative and qualitative analysis of our framework. Jayneel Parekh, Preeti Rao, Yi-Hsuan Yang |
ICASSP | 2 |
| 2020 | Vapar Synth - A Variational Parametric Model for Audio SynthesisabstractWith the advent of data-driven statistical modeling and abundant computing power, researchers are turning increasingly to deep learning for audio synthesis. These methods try to model audio signals directly in the time or frequency domain. In the interest of more flexible control over the generated sound, it could be more useful to work with a parametric representation of the signal which corresponds more directly to the musical attributes such as pitch, dynamics and timbre. We present Va-Par Synth - a Variational Parametric Synthesizer which utilizes a conditional variational autoencoder (CVAE) trained on a suitable parametric representation. We demonstrate1our proposed model's capabilities via the reconstruction and generation of instrumental tones with flexible control over their pitch. Krishna Subramani, Preeti Rao, Alexandre D'Hooge |
ICASSP | 2 |
| 2020 | Automatic Prediction of Confidence Level from Children's Oral Reading Recordings
Kamini Sabu, Preeti Rao |
INTERSPEECH | 2 |
| 2018 | Acoustic-Prosodic Features of Tabla Bol Recitation and Correspondence with the Tabla Imitation
Rohit M. A., Preeti Rao |
INTERSPEECH | 2 |
| 2018 | A Non-convolutive NMF Model for Speech Dereverberation
Nikhil Mohanan, Rajbabu Velmurugan, Preeti Rao |
INTERSPEECH | 3 |
| 2018 | A Study of Lexical and Prosodic Cues to Segmentation in a Hindi-English Code-switched Discourse
Preeti Rao, Mugdha Pandya, Kamini Sabu, Kanhaiya Kumar, Nandini Bondale |
INTERSPEECH | 1 |
| 2018 | Automatic Detection of Expressiveness in Oral Reading
Kamini Sabu, Kanhaiya Kumar, Preeti Rao |
INTERSPEECH | 3 |
| 2017 | Speech dereverberation using NMF with regularized room impulse responseabstractIn this paper, various regularizations on the room impulse response (RIR) are proposed to obtain better single-channel speech dereverberation in the non-negative matrix factorization (NMF) framework. The regularizations on the RIR are motivated by the spectral domain representation of the RIR. To obtain better estimates of the RIR and clean speech, we propose three modifications (i) to obtain a sparse RIR (ii) a frequency envelop constrained RIR and (iii) to include the early part of the RIR. The performance of the proposed regularizers are evaluated by considering speech enhancement measures. While it is observed that the regularizers lead to an improved estimate of the RIR, they do not necessarily lead to speech enhancement in all the cases. For the experiments conducted using the RIRs from the REVERB 2014 challenge and sentences from the TIMIT database, the regularizer that includes the early part of the RIR shows reasonable improvement in speech enhancement measures. Nikhil Mohanan, Rajbabu Velmurugan, Preeti Rao |
ICASSP | 3 |
| 2016 | Automatic Assessment of Reading with Speech RecognitionTechnology
Preeti Rao, Prakhar Swarup, Ankita Pasad, Hitesh Tulsiani, Gargi Ghosh Das |
ICCE | 1 |
| 2015 | Structural segmentation of Hindustani concert audio with posterior featuresabstractStructural segmentation of music involves identifying boundaries between homogenous regions where the homogeneity involves one or more musical dimensions, and therefore depends on the musical genre. In this work, we address the segmentation of Hindustani instrumental concert recordings at the highest time-scale, that is, concert sections marked by prominent changes in rhythmic structure. Tempo features are effectively combined with energy and chroma features motivated by musicological knowledge and acoustic observations. Posterior probability features from unsupervised model fitting of the frame-level acoustic features are shown to significantly improve robustness to local acoustic variations. Finally, two diverse change detection criteria are combined to obtain a superior segmentation system. Vinutha T. P., Parthe Pandit, Preeti Rao |
ICASSP | 4 |
| 2015 | Vowel mispronunciation detection using DNN acoustic models with cross-lingual training
Shrikant Joshi, Nachiket Deo, Preeti Rao |
INTERSPEECH | 3 |
| 2014 | Acoustic characteristics of critical message utterances in noise applied to speech intelligibility enhancement
Neehar Jathar, Preeti Rao |
INTERSPEECH | 2 |
| 2013 | Acoustic features for detection of phonemic aspiration in voiced plosivesabstractPlosives in Indo-Aryan languages such as Hindi and Marathi display a 4-way contrast involving the two dimensions of voicing and aspiration. While many studies are available on the acoustics of aspiration in unvoiced stops due to their more universal presence in the world’s languages, voiced aspirated plosives have been less studied. Rather than the release duration cue of aspiration in unvoiced stops, the acoustic realization of aspiration in voiced plosives is marked by the coarticulatory breathiness of the following vowel. We consider the automatic detection of aspiration in Marathi word-initial voiced stops and affricates via several features relating to extent and timing of breathiness of the following vowel. The effectiveness of the features is evaluated by classification performance on a database of Marathi words. A practical application of this work to the detection of non-native pronunciation of voiced obstruents is presented. Index Terms: voiced obstruents, phonemic aspiration, breathy vowels, manner classification, acoustic-phonetic features Vaishali Patil, Preeti Rao |
INTERSPEECH | 2 |
| 2012 | Classification of place of articulation in unvoiced stops with spectro-temporal surface modeling
Veena Karjigi, Preeti Rao |
Speech Commun. | 2 |
| 2012 | Signal-Driven Window-Length Adaptation for Sinusoid Detection in Polyphonic MusicabstractAudio processing applications that use short-time signal analysis techniques typically utilize fixed window duration single- or multi-resolution analyses. However, different real-world signal conditions such as polyphony and non-stationarity, manifested as musical accompaniment and pitch-modulations, respectively, in the context of music content analysis, require varying data window lengths for reliable processing. In this paper, we investigate the use of signal sparsity for adapting analysis window lengths. Adaptive-window analysis driven by different measures of sparsity applied to the local spectrum, such as kurtosis and Gini index, is evaluated and shown to be superior to fixed-window analysis in terms of sinusoid detection and frequency estimation for simulated and real signals. A window main-lobe matching method for sinusoid detection is also shown to be more robust to signal conditions such as polyphony and frequency modulation relative to other methods. Vishweshwara Rao, Pradeep Gaddipati, Preeti Rao |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | Vocal Melody Extraction in the Presence of Pitched Accompaniment in Polyphonic MusicabstractMelody extraction algorithms for single-channel polyphonic music typically rely on the salience of the lead melodic instrument, considered here to be the singing voice. However the simultaneous presence of one or more pitched instruments in the polyphony can cause such a predominant-F0 tracker to switch between tracking the pitch of the voice and that of an instrument of comparable strength, resulting in reduced voice-pitch detection accuracy. We propose a system that, in addition to biasing the salience measure in favor of singing voice characteristics, acknowledges that the voice may not dominate the polyphony at all instants and therefore tracks an additional pitch to better deal with the potential presence of locally dominant pitched accompaniment. A feature based on the temporal instability of voice harmonics is used to finally identify the voice pitch. The proposed system is evaluated on test data that is representative of polyphonic music with strong pitched accompaniment. Results show that the proposed system is indeed able to recover melodic information lost to its single-pitch tracking counterpart, and also outperforms another state-of-the-art melody extraction system designed for polyphonic music. Vishweshwara Rao, Preeti Rao |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Improving the robustness of phonetic segmentation to accent and style variation with a two-staged approachabstractCorrect and temporally accurate phonetic segmentation of speech utterances is important in applications ranging from transcription alignment to pronunciation error detection. Automatic speech recognizers used in these tasks provide insufficient temporal alignment accuracy apart from a recognition performance that is sensitive to accent and style variations from the training data. A two-staged approach combining HMM broad-class recognition with acoustic-phonetic knowledge based refinement is evaluated for phonetic segmentation accuracy in the context of accent and style mismatches with training data. Vaishali Patil, Shrikant Joshi, Preeti Rao |
INTERSPEECH | 3 |
| 2009 | Singing voice detection in polyphonic music using predominant pitchabstractThis paper demonstrates the superiority of energy-based features derived from the knowledge of predominant-pitch, for singing voice detection in polyphonic music over commonly used spectral features. However, such energy-based features tend to misclassify loud, pitched instruments. To provide robustness to such accompaniment we exploit the relative instability of the pitch contour of the singing voice by attenuating harmonic spectral content belonging to stable-pitch instruments, using sinusoidal modeling. The obtained feature shows high classification accuracy when applied to north Indian classical music data and is also found suitable for automatic detection of vocal-instrumental boundaries required for smoothing the frame-level classifier decisions. Index Terms: audio segmentation, voice detection 1. Vishweshwara Rao, Preeti Rao |
INTERSPEECH | 3 |
| 2008 | Landmark based recognition of stops: acoustic attributes versus smoothed spectraabstractLandmark based recognition of unvoiced word-initial stops is investigated. The relative effectiveness of acoustic-phonetic attributes versus more global spectral shape features is experimentally evaluated for four-way place classification of unvoiced, unaspirated stops. Various feature sets derived from the burst and vocalic transition regions of word initial consonants are compared via GMM based classification under speaker, gender, and vowel-context variability. While a set of acoustic attributes derived from the burst shows the best invariance to vowel context, it is found that global spectral shape features provide the most robust representation of the vocalic transition region by overcoming the problem of errors in explicit formant tracking. A combination of features from the burst and vocalic regions was superior to burst-only cues, but still far from the near perfect identification achieved in human perception. Index Terms: Landmark based recognition, unvoiced stops, acoustic attributes, burst, vocalic transition Veena Karjigi, Preeti Rao |
INTERSPEECH | 2 |
| 2006 | Speech enhancement in nonstationary noise environments using noise properties
Kotta Manohar, Preeti Rao |
Speech Commun. | 2 |
| 2006 | Effect of voice quality on frequency-warped modeling of vowel spectra
Pushkar Patwardhan, Preeti Rao |
Speech Commun. | 2 |
| 2005 | Frequency warped modeling of vowel spectra: Dependence on vowel quality
Preeti Rao, Pushkar Patwardhan |
Speech Commun. | 1 |
| 2002 | Spectral enhancement preprocessing for the HNM coding of noisy speechabstractLow rate coders based on the harmonic-noise model are sensitive to acoustic background noise at low SNRs due to the increase in parameter errors from the analysis of noisy speech. We investi-gate the use of spectral subtraction enhancement preprocessing on the performance of the sinusoidal model based codec both by ob-jective assessment of parameter errors and the subjective testing of output speech quality and intelligibility. We find that while for noisy speech enhancement, improving speech quality is often ac-companied by a decrease in intelligibility, in the context of coding, significant combined improvements are obtained when the speech coder is combined with a speech enhancement preprocessor. 1. Gautam Moharir, Pushkar Patwardhan, Preeti Rao |
INTERSPEECH | 3 |
| 2002 | Controlling perceived degradation in spectrum envelope modeling via predistortionabstractThe compact representation of the discrete amplitude spectrum of voiced speech by an all-pole model of the spectral envelope is considered. Based on the properties of the all-pole modeling error, the use of spectrum predistortion for improving the perceptual fit at low model orders is motivated. Warping of the frequency scale before modeling of the spectral envelope of narrowband voiced sounds is investigated by subjective listening and objective measures. It is found that, contrary to what is generally accepted, the improvement in perceived quality brought about by frequency warping actually depends to a large extent on the underlying signal spectrum distribution. An objective distance measure based on partial noise loudness is found to show high correlation with subjective judgements of degradation, indicating that auditory frequency masking plays an important role in determining the perceptual accuracy of the spectrum envelope model. Pushkar Patwardhan, Preeti Rao |
INTERSPEECH | 2 |
| 2000 | Speech formant frequency estimation: evaluating a nonstationary analysis method
Preeti Rao, A. Das Barman |
Signal Process. | 1 |
| 1996 | A robust method for the estimation of formant frequency modulation in speech signalsabstractThe presence of amplitude and frequency modulations in individual formants of the speech signal was demonstrated using the discrete-time energy separation algorithm (DESA). Formant modulation estimates are valuable to the understanding of speech production, and have found several applications. While the DESA has been successfully applied in tracking the amplitude envelope of formants, the DESA frequency estimator has been relatively unexploited for tracking formant frequency modulation (FM) due to its lack of robustness to conditions commonly occurring in practice. We consider an alternate method of frequency estimation based on the Wigner distribution (WD) and investigate its application to the estimation of formant FM in speech signals. It is shown using simulated and speech signals that the WD estimator can track formant FM with high time resolution under a wider range of conditions and is more robust to additive noise present in the signal than the DESA estimator. Finally the computational complexities of the two methods are compared. Preeti Rao |
ICASSP | 1 |
| 1994 | 8 kb/s low-delay speech coding with 4 ms frame size
Preeti Rao, Yoshiaki Asakawa, Hidetoshi Sekine |
ICSLP | 1 |
| 1991 | A robust narrowband transient signal detectorabstractTh Wigner distribution (WD) is applied to the detection and localization in time of narrowband transient signals of unknown waveform in additive noise comprise of quasi-harmonic and random components. A traditional method of processing such signals is the spectrogram or short-time power spectrum. Based on the relationship between the spectrogram and the WD, the application of the WD to the detection problem is investigated. The smoothed-pseudo-WD is shown to provide the superior time and frequency resolutions of the received signals. By monitoring the received power in localized regions of the WD spectrum, a detector that achieves good time localization of the transient waveform and is relatively insensitive to changes in signal duration or frequency is obtained.> Preeti Rao |
ICASSP | 1 |
| 1989 | Roundoff error analysis of the discrete Wigner distribution using fixed-point arithmeticabstractThe authors investigate the roundoff noise in a fixed-point implementation of a discrete Wigner distribution (DWD) algorithm. The output noise-to-signal ratio (NSR) is derived using a statistical approach based on certain assumptions about the input signal. The expression for noise-to-signal ratio indicates that for each added stage in the DWD implementation, the fixed-point wordlength should be increased by one bit to keep the noise-to-signal ratio unchanged. The theoretical result has been validated by computer simulations using similar input data. It is found that doubling the length of the DWD increases the NSR by approximately 6 dB.> Chintana Griffin, Preeti Rao, Fred J. Taylor |
ICASSP | 2 |