VLDB 2026 Research / reviewers in the wild / expert
Samik Sadhu
dblp:207/9624
· DBLP profile ↗
9ranked-venue papers
7as first author
5since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Importance of Different Temporal Modulations of Speech: a Tale of two PerspectivesabstractHow important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic information in the modulation spectrum of speech by computing the mutual information between temporal modulations with frame-wise phoneme labels. Looking from another perspective, we ask - which speech modulations an Automatic Speech Recognition (ASR) system prefers for its operation. Data-driven weights are learned over the modulation spectrum and optimized for an end-to-end ASR task. Both methods unanimously agree that speech information is mostly contained in slow modulation. Maximum mutual information occurs around 3-6 Hz which also happens to be the range of modulations most preferred by the ASR. In addition, we show that the incorporation of this knowledge into ASRs significantly reduces their dependency on the amount of training data. Samik Sadhu, Hynek Hermansky |
ICASSP | 1 |
| 2022 | Complex Frequency Domain Linear Prediction: A Tool to Compute Modulation Spectrum of Speech
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 1 |
| 2022 | Dealing with Unknowns in Continual Learning for End-to-end Automatic Speech Recognition
Martin Sustek, Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 2 |
| 2021 | Radically Old Way of Computing Spectra: Applications in End-to-End ASRabstractWe propose a technique to compute spectrograms using Frequency Domain Linear Prediction (FDLP) that uses all-pole models to fit the squared Hilbert envelope of speech in different frequency sub-bands. The spectrogram of a complete speech utterance is computed by overlap-add of contiguous all-pole model responses. A long context window of 1.5 seconds allows us to capture the low frequency temporal modulations of speech in the spectrogram. For an end-to-end automatic speech recognition task, the FDLP spectrogram performs on par with the standard mel spectrogram features for clean read speech training and test data. For more realistic speech data with train-test domain mismatches or reverberations, FDLP spectrogram shows up to 25% and 22% relative WER improvements over mel spectrogram respectively. Samik Sadhu, Hynek Hermansky |
Interspeech | 1 |
| 2021 | wav2vec-C: A Self-Supervised Model for Speech Representation LearningabstractWav2vec-C introduces a novel representation learning technique combining elements from wav2vec 2.0 and VQ-VAE. Our model learns to reproduce quantized representations from partially masked speech encoding using a contrastive loss in a way similar to Wav2vec 2.0. However, the quantization process is regularized by an additional consistency network that learns to reconstruct the input features to the wav2vec 2.0 network from the quantized representations in a way similar to a VQ-VAE model. The proposed self-supervised model is trained on 10k hours of unlabeled data and subsequently used as the speech encoder in a RNN-T ASR model and fine-tuned with 1k hours of labeled data. This work is one of only a few studies of self-supervised learning on speech tasks with a large volume of real far-field labeled data. The Wav2vec-C encoded representations achieves, on average, twice the error reduction over baseline and a higher codebook utilization in comparison to wav2vec 2.0 Samik Sadhu, Di He 0004, Che-Wei Huang, Sri Harish Reddy Mallidi, Minhua Wu, Ariya Rastrow, Andreas Stolcke, Jasha Droppo, Roland Maas |
Interspeech | 1 |
| 2020 | Continual Learning in Automatic Speech Recognition
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 1 |
| 2019 | M-vectors: Sub-band Based Energy Modulation Features for Multi-stream Automatic Speech RecognitionabstractIn this paper, we propose a novel method to capture energy modulations from different frequency bands in speech into frame-level feature vectors, called Modulation-vectors or M-vectors, for use in Automatic Speech Recognition (ASR) systems. We show that in different multi-stream setups, with parallel streams for M-vectors and the popular Mel-frequency Cepstral Coefficient (MFCC) features, we can realize a boost in word recognition performance of end-to-end systems by ≈ 5%, and that of a monophone and triphone HMM-GMM ASR system by ≈ 18% and ≈ 16% respectively over using the traditional MFCC features. Samik Sadhu, Hynek Hermansky |
ICASSP | 1 |
| 2019 | Modulation Vectors as Robust Feature Representation for ASR in Domain Mismatched Conditions
Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 1 |
| 2019 | Exploring Methods for the Automatic Detection of Errors in Manual TranscriptionabstractQuality of data plays an important role in most deep learning tasks.In the speech community, transcription of speech recording is indispensable.Since the transcription is usually generated artificially, automatically finding errors in manual transcriptions not only saves time and labors but benefits the performance of tasks that need the training process.Inspired by the success of hybrid automatic speech recognition using both language model and acoustic model, two approaches of automatic error detection in the transcriptions have been explored in this work.Previous study using a biased language model approach, relying on a strong transcription-dependent language model, has been reviewed.In this work, we propose a novel acoustic model based approach, focusing on the phonetic sequence of speech.Both methods have been evaluated on a completely real dataset, which was originally transcribed with errors and strictly corrected manually afterwards. Xiaofei Wang 0007, Jinyi Yang, Samik Sadhu, Hynek Hermansky |
INTERSPEECH | 4 |