Mehul Kumar

dblp:140/7239 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 5 since 2021Theory of computation · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 CMAEH: Contrastive Masked Autoencoder Based Hashing for Efficient Image Retrieval
Mehul Kumar, Prerana Mukherjee, Koteswar Rao Jerripothula
ICPR (20)1
2024 Combatting Disinformation and Deepfake: Interdisciplinary Insights and Global Strategies
abstract
In this era led by Generative AI, any content you see on the internet cannot be trusted for authenticity, a disinformation video spreading like wildfire amongst the masses can do much greater damage than we can imagine, manipulate public opinion, agitate violence, ruin livelihoods, and endorse communal hate. As a defense mechanism usage of remote Photoplethysmography (rPPG), which is capable of capturing cardiovascular data, extract heartbeat signals to temporal PPG maps which when trained on Convolutional Neural Network helps us derive features and patterns observed in different methodologies followed to create deepfakes. Helping not only to detect the particular fake but also to get into the roots of it, making it simpler to detect the source and take necessary actions to catch the perpetrator. New day and age demand strict Law Regulations to prevent disinformation and derogatory deepfakes. This model can help clearly Distinguish Real and Fake apart.
Anubhav Hooda, Mehul Kumar
TENCON2
2024 VTHSC-MIR: Vision Transformer Hashing with Supervised Contrastive learning based medical image retrieval
Mehul Kumar, Rhythumwinder Singh, Prerana Mukherjee
Pattern Recognit. Lett.1
2023 Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data
abstract
In this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensive task. Also, finding a pool of representative English speakers for any arbitrary accent to collect unlabeled data can be a difficult task. However, collecting unlabeled speech data for any native language is a much simpler task. It is important to note that the accents of most non-native English speakers are heavily biased by the co-articulation of sounds in their own native language. In view of this, we propose to use unlabeled native language data to learn self-supervised representations during the pre-training stage. The pre-trained model is then fine-tuned using limited labeled English data for the target accent. Experiments using native language data to pre-train an English recognizer followed by fine-tuning using target accented English show significant improvements in word error rates on four different accents (Great Britain, Korean, Chinese, Spanish).
Mehul Kumar, Jiyeon Kim, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ICASSP1
2021 HiTNet: Byte-to-BPE Hierarchical Transcription Network for End-to-End Speech Recognition
abstract
In this paper, we propose a new byte to byte-pair-encoding (BPE) Hierarchical Transcription Network (HiTNet) architecture for end-to-end (e2e) automatic speech recognition (ASR). The proposed HiTNet architecture simultaneously encodes as well as decodes information hierarchically at different levels of linguistic granularity such as bytes and BPE. In general this idea can be extended to any levels of granularity including phonemes or graphemes or bytes (character to sub-character in some languages), to sub-words or byte-pair encodings (BPE), to words, and so on. Existing hierarchical e2e ASR models primarily encode the acoustic information in an hierarchical manner governed by weaker linguistic constraints at each level. The language information at each level is neither embedded or used explicitly, nor is the information decoded at each level passed on to the next stage. The proposed architecture primarily decodes information in an hierarchical manner utilizing the linguistic information at each level explicitly, while at the same time utilizing the hierarchically encoded acoustic information at each level. Experiments with a two-level byte-to-BPE (b2B) hierarchical transcription show that the proposed architecture significantly reduces the word error rates of both the byte and BPE decoders compared to baseline byte and BPE based attention encoder-decoder models.
Dhananjaya Gowda, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Nauman Dawalatabad, Aman Maghan, Shatrughan Singh, Chanwoo Kim 0001
ASRU4
2021 Semi-Supervised Transfer Learning for Language Expansion of End-to-End Speech Recognition Models to Low-Resource Languages
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ASRU2
2021 A Comparison of Streaming Models and Data Augmentation Methods for Robust Speech Recognition
abstract
In this paper, we present a comparative study on the robustness of two different online streaming speech recognition models: Monotonic Chunkwise Attention (MoChA) and Recurrent Neural Network-Transducer (RNN-T). We explore three recently proposed data augmentation techniques, namely, multi-conditioned training using an acoustic simulator, Vocal Tract Length Perturbation (VTLP) for speaker variability, and SpecAugment. Experimental results show that unidirectional models are in general more sensitive to noisy examples in the training set. It is observed that the final performance of the model depends on the proportion of training examples processed by data augmentation techniques. MoChA models generally perform better than RNN-T models. However, we observe that training of MoChA models seems to be more sensitive to various factors such as the characteristics of training sets and the incorporation of additional augmentations techniques. On the other hand, RNN-T models perform better than MoChA models in terms of latency, inference time, and the stability of training. Additionally, RNN-T models are generally more robust against noise and reverberation. All these advantages make RNN-T models a better choice for streaming on-device speech recognition compared to MoChA models.
Jiyeon Kim, Mehul Kumar, Dhananjaya Gowda, Abhinav Garg, Chanwoo Kim 0001
ASRU2
2020 Utterance Invariant Training for Hybrid Two-Pass End-to-End Speech Recognition
Dhananjaya Gowda, Kwangyoun Kim, Hejung Yang, Abhinav Garg, Jiyeon Kim, Mehul Kumar, Sichen Jin, Shatrughan Singh, Chanwoo Kim 0001
INTERSPEECH8
2019 Improved Multi-Stage Training of Online Attention-Based Encoder-Decoder Models
abstract
In this paper, we propose a refined multi-stage multi-task training strategy to improve the performance of onlineattention-based encoder-decoder (AED) models. A three-stage training based on three levels of architectural granularity namely, character encoder, byte pair encoding (BPE) based encoder, and attention decoder, is proposed. Also, multi-task learning based on two-levels of linguistic granularity namely, character and BPE, is used. We explore different pre-training strategies for the encoders including transfer learning from a bidirectional encoder. Our encoder-decoder models with online attention show ~35% and ~10% relative improvement over their baselines for smaller and bigger models, respectively. Our models achieve a word error rate (WER) of 5.04% and 4.48% on the Librispeech test-clean data for the smaller and bigger models respectively after fusion with long short-term memory (LSTM) based external language model (LM).
Abhinav Garg, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001
ASRU5
2019 Power-Law Nonlinearity with Maximally Uniform Distribution Criterion for Improved Neural Network Training in Automatic Speech Recognition
abstract
In this paper, we describe the Maximum Uniformity of Distribution (MUD) algorithm with the power-law nonlinearity. In this approach, we hypothesize that neural network training will become more stable if feature distribution is not too much skewed. We propose two different types of MUD approaches: power function-based MUD and histogram-based MUD. In these approaches, we first obtain the mel filterbank coefficients and apply nonlinearity functions for each filterbank channel. With the power function-based MUD, we apply a power-function based nonlinearity where power function coefficients are chosen to maximize the likelihood assuming that nonlinearity outputs follow the uniform distribution. With the histogram-based MUD, the empirical Cumulative Density Function (CDF) from the training database is employed to transform the original distribution into a uniform distribution. In MUD processing, we do not use any prior knowledge (e.g. logarithmic relation) about the energy of the incoming signal and the perceived intensity by a human. Experimental results using an end-to-end speech recognition system demonstrate that power-function based MUD shows better result than the conventional Mel Filterbank Cepstral Coefficients (MFCCs). On the LibriSpeech database, we could achieve 4.02 % WER on test-clean and 13.34 % WER on test-other without using any Language Models (LMs). The major contribution of this work is that we developed a new algorithm for designing the compressive nonlinearity in a data-driven way, which is much more flexible than the previous approaches and may be extended to other domains as well.
Chanwoo Kim 0001, Mehul Kumar, Kwangyoun Kim, Dhananjaya Gowda
ASRU2
2019 End-to-End Training of a Large Vocabulary End-to-End Speech Recognition System
abstract
In this paper, we present an end-to-end training framework for building state-of-the-art end-to-end speech recognition systems. Our training system utilizes a cluster of Central Processing Units (CPUs) and Graphics Processing Units (GPUs). The entire data reading, large scale data augmentation, neural network parameter updates are all performed “on-the-fly”. We use vocal tract length perturbation [1] and an acoustic simulator [2] for data augmentation. The processed features and labels are sent to the GPU cluster. The Horovod allreduce approach is employed to train neural network parameters. We evaluated the effectiveness of our system on the standard Librispeech corpus [3] and the 10,000-hr anonymized Bixby English dataset. Our end-to-end speech recognition system built using this training infrastructure showed a 2.44 % WER on test-clean of the LibriSpeech test set after applying shallow fusion with a Transformer language model (LM). For the proprietary English Bixby open domain test set, we obtained a WER of 7.92 % using a Bidirectional Full Attention (BFA) end-to-end model after applying shallow fusion with an RNN-LM. When the monotonic chunckwise attention (MoCha) based approach is employed for streaming speech recognition, we obtained a WER of 9.95 % on the same Bixby open domain test set.
Chanwoo Kim 0001, Minkyoo Shin, Shatrughan Singh, Larry Heck, Dhananjaya Gowda, Kwangyoun Kim, Mehul Kumar, Jiyeon Kim, Kyungmin Lee, Abhinav Garg, Eunhyang Kim
ASRU8
2019 Reoptimization of Path Vertex Cover Problem
Mehul Kumar, Amit Kumar 0015, C. Pandu Rangan
COCOON1
2019 Multi-Task Multi-Resolution Char-to-BPE Cross-Attention Decoder for End-to-End Speech Recognition
Dhananjaya Gowda, Abhinav Garg, Kwangyoun Kim, Mehul Kumar, Chanwoo Kim 0001
INTERSPEECH4
2015 Improved analysis of D2-sampling based PTAS for k-means and other clustering problems
Ragesh Jaiswal, Mehul Kumar, Pulkit Yadav
Inf. Process. Lett.2