Anshuman Tripathi

dblp:198/2060 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
6since 2021 · last 2024
0000-0002-4902-3719ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3
YearPublicationVenuePosition
2024 Monte Carlo Self-Training for Speech Recognition
abstract
Self-training in the teacher-student framework generally suffers from the confirmation bias problem, where errors from the teacher are propagated to the student and hence get amplified with multiple iterations. In this paper, we present Monte Carlo Self-training where pseudo labels are generated by sampling from a teacher distribution, as a way to mitigate this problem. We show that Monte Carlo Self-training is an approximation to minimizing label sequence level cross entropy between student and teacher. In our experiments we find that Monte Carlo Self-training always outperforms beam decoder based self-training and is quite robust even when the initial teacher WER is high. We also show that label sampling allows formulating pseudo label confidence in a more natural way and we show that these confidence measures give further improvements in our unsupervised adaptation experiments especially when the initial teacher WER is very high.
Anshuman Tripathi, Soheil Khorram, Han Lu 0003, Hasim Sak
ICASSP1
2023 Cross-Training: A Semi-Supervised Training Scheme for Speech Recognition
abstract
Semi-supervised training can be performed by jointly optimizing supervised and unsupervised losses. In many settings, supervised and unsupervised losses are inconsistent, and this inconsistency creates instability in training. As a solution, we propose cross-training: instead of training one network with two losses, we train two separate networks, each with a different loss; we then tie the parameters of the networks by minimizing an additional L2 loss between the parameters. This L2 loss acts as a knowledge bridge between the networks. It forces the networks to be similar; therefore both can learn from each other. This paper introduces the cross-training scheme to develop a stable contrastive siamese (c-siam) network. Our experiments on LibriSpeech and Google’s Voice-Search/YouTube datasets show that (1) cross-training provides 20% relative WER improvement over the SOTA systems on the LibriSpeech dataset; (2) cross-training stabilizes c-siam training and significantly outperforms SOTA systems on small supervised datasets; (3) cross-training is effective for cascaded encoders, unlike the original c-siam which shows weak convergence characteristics.
Soheil Khorram, Anshuman Tripathi, Han Lu 0003, Rohit Prabhavalkar, Hasim Sak
ICASSP2
2022 Contrastive Siamese Network for Semi-Supervised Speech Recognition
abstract
This paper introduces contrastive siamese (c-siam) network, an architecture for leveraging unlabeled acoustic data in speech recognition. c-siam is the first network that extracts high-level linguistic information from speech by matching outputs of two identical transformer encoders. It contains augmented and target branches which are trained by: (1) masking inputs and matching outputs with a contrastive loss, (2) incorporating a stop gradient operation on the target branch, (3) using an extra learnable transformation on the augmented branch, (4) introducing new temporal augment functions to prevent the shortcut learning problem. We use the Libri-light 60k unsupervised data and the LibriSpeech 100hrs/960hrs supervised data to compare c-siam and other best-performing systems. Our experiments show that c-siam provides 20% relative word error rate improvement over wav2vec baselines. A c-siam network with 450M parameters achieves competitive results compared to the state-of-the-art networks with 600M parameters.
Soheil Khorram, Anshuman Tripathi, Han Lu 0003, Hasim Sak
ICASSP3
2022 Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection
abstract
In this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, represent each speaker turn by a speaker embedding, then cluster these embeddings with constraints from the detected speaker turns. Compared with conventional clustering-based diarization systems, our system largely reduces the computational cost of clustering due to the sparsity of speaker turns. Unlike other supervised speaker diarization systems which require annotations of time-stamped speaker labels for training, our system only requires including speaker turn tokens during the transcribing process, which largely reduces the human efforts involved in data collection.
Han Lu 0003, Anshuman Tripathi, Ignacio López-Moreno, Hasim Sak
ICASSP4
2021 Reducing Streaming ASR Model Delay with Self Alignment
abstract
Reducing prediction delay for streaming end-to-end ASR models with minimal performance regression is a challenging problem. Constrained alignment is a well-known existing approach that penalizes predicted word boundaries using external low-latency acoustic models. On the contrary, recently proposed FastEmit is a sequence-level delay regularization scheme encouraging vocabulary tokens over blanks without any reference alignments. Although all these schemes are successful in reducing delay, ASR word error rate (WER) often severely degrades after applying these delay constraining schemes. In this paper, we propose a novel delay constraining method, named self alignment. Self alignment does not require external alignment models. Instead, it utilizes Viterbi forced-alignments from the trained model to find the lower latency alignment direction. From LibriSpeech evaluation, self alignment outperformed existing schemes: 25% and 56% less delay compared to FastEmit and constrained alignment at the similar word error rate. For Voice Search evaluation,12% and 25% delay reductions were achieved compared to FastEmit and constrained alignment with more than 2% WER improvements.
Han Lu 0003, Anshuman Tripathi, Hasim Sak
Interspeech3
2021 End-to-End Audio-Visual Speech Recognition for Overlapping Speech
Richard Rose, Olivier Siohan, Anshuman Tripathi, Otavio Braga
Interspeech3
2020 End-To-End Multi-Talker Overlapping Speech Recognition
abstract
In this paper we present an end-to-end speech recognition system that can recognize single-channel speech where multiple talkers can speak at the same time (overlapping speech) by using a neural network model based on Recurrent Neural Network Transducer (RNN-T) architecture. We augment the conventional RNN-T architecture by including a masking model for separation of encoded audio features, and multiple label encoders to encode transcripts from different speakers. We use a masking L2 loss to prevent transcripts to align to wrong speakers' audio, and a speaker embedding loss to facilitate speaker tracking. We show that by using these additional training objectives, the proposed augmented RNN-T model can be trained with simulated overlapping speech data and can achieve a WER of 32% on words in overlapping speech segments from real-life telephone conversations. Our analysis of manual transcription task on the same test set shows that transcribing overlapping speech is hard even for humans who can get a WER of 37% compared to ground-truth.
Anshuman Tripathi, Han Lu 0003, Hasim Sak
ICASSP1
2020 Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
abstract
In this paper we present an end-to-end speech recognition model with Transformer encoders that can be used in a streaming speech recognition system. Transformer computation blocks based on self-attention are used to encode both audio and label sequences independently. The activations from both audio and label encoders are combined with a feed-forward layer to compute a probability distribution over the label space for every combination of acoustic frame position and label history. This is similar to the Recurrent Neural Network Transducer (RNN-T) model, which uses RNNs for information encoding instead of Transformer encoders. The model is trained with the RNN-T loss well-suited to streaming decoding. We present results on the LibriSpeech dataset showing that limiting the left context for self-attention in the Transformer layers makes decoding computationally tractable for streaming, with only a slight degradation in accuracy. We also show that the full attention version of our model beats the-state-of-the art accuracy on the LibriSpeech benchmarks. Our results also show that we can bridge the gap between full attention and limited attention versions of our model by attending to a limited number of future frames.
Han Lu 0003, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, Shankar Kumar
ICASSP4
2020 Multilingual Speech Recognition with Self-Attention Structured Parameterization
Parisa Haghani, Anshuman Tripathi, Bhuvana Ramabhadran, Brian Farris, Hainan Xu, Han Lu 0003, Hasim Sak, Isabel Leal, Neeraj Gaur, Pedro J. Moreno 0001
INTERSPEECH3
2019 Monotonic Recurrent Neural Network Transducer and Decoding Strategies
abstract
Recurrent Neural Network Transducer (RNNT) is an end-to-end model which transduces discrete input sequences to output sequences by learning alignments between the sequences. In speech recognition tasks we generally have a strictly monotonic alignment between time frames and label sequence. However, the standard RNNT loss does not enforce this constraint. This can cause some anomalies in alignments such as the model outputting a sequence of labels at a single time frame. There is also no bound on the decoding time steps. To address these problems, we introduce a monotonic version of the RNNT loss. Under the assumption that the output sequence is not longer than the input sequence, this loss can be used with forward-backward algorithm to learn strictly monotonic alignments between the sequences. We present experimental studies showing that speech recognition accuracy for monotonic RNNT is equivalent to standard RNNT. We also explore best-first and breadth-first decoding strategies for both monotonic and standard RNNT models. Our experiments show that breadth-first search is effective in exploring and combining alternative alignments. Additionally, it also allows batching of hypotheses during search label expansion, allowing better resource utilization, and resulting in decoding speedup.
Anshuman Tripathi, Han Lu 0003, Hasim Sak, Hagen Soltau
ASRU1
2018 Temporal Modeling Using Dilated Convolution and Gating for Voice-Activity-Detection
abstract
Voice activity detection (VAD) is the task of predicting which parts of an utterance contains speech versus background noise. It is an important first step to determine which samples to send to the decoder and when to close the microphone. The long short-term memory neural network (LSTM) is a popular architecture for sequential modeling of acoustic signals, and has been successfully used in several VAD applications. However, it has been observed that LSTMs suffer from state saturation problems when the utterance is long (i.e., for voice dictation tasks), and thus requires the LSTM state to be periodically reset. In this paper, we propose an alternative architecture that does not suffer from saturation problems by modeling temporal variations through a stateless dilated convolution neural network (CNN). The proposed architecture differs from conventional CNNs in three respects: it uses dilated causal convolution, gated activations and residual connections. Results on a Google Voice Typing task shows that the proposed architecture achieves 14% relative FA improvement at a FR of 1% over state-of-the-art LSTMs for VAD task. We also include detailed experiments investigating the factors that distinguish the proposed architecture from conventional convolution.
Shuo-Yiin Chang, Bo Li 0028, Gabor Simko, Tara N. Sainath, Anshuman Tripathi, Aäron van den Oord, Oriol Vinyals
ICASSP5
2018 Analysis of Brushless Wound Rotor Synchronous Generator with Unity Power Factor Rectifier for Series Offshore DC Wind Power Collection
abstract
Due to various advantages offshore wind farms are getting higher attention from industrial and research perspectives. The DC power collection topologies are preferred as they reduce both system and transmission cost in case of offshore power transmission. Offshore wind turbines are expected to be lighter, smaller and prone to lesser maintenance. Permanent magnet synchronous generator (PMSG), though most popular choice, they are getting more and more expensive due to the availability of raw earth permanent magnets (PM). Thus a wound rotor synchronous generator (WRSG), with brushless excitation is a probable alternative to PMSG. For front-end converters three-phase-diode bridge rectifiers are suitable for high power ratings but they are not controllable. It requires additional auxiliary circuitry or to be followed by a DC-DC converter, in order to be used for DC power collection. With these regards, this study is carried out for the scenario of series connected DC power collection of wind turbines with brushless excited WRSG. For the front-end converter, an auxiliary circuit based three phase diode bridge unity-power-factor rectifier is selected. The brushless excitation ensures the generator terminal voltage regulation. The rectifier controls DC link voltage, improves the power factor and reduces the generator output current total harmonic distortion (THD). The control of converter is done based on hysteresis current control. To verify its efficacy, the proposed system is simulated under various dynamic wind speed scenarios.
Md Shafquat Ullah Khan, Ali I. Maswood, Kuntal Satpathi, Mohammad Tauquir Iqbal, Anshuman Tripathi
IECON5
2018 Experimental Verification on Thermal Modeling of Medium Frequency Transformers
abstract
Nowadays, medium frequency transformers have gained growing attention in modern power system. Higher power density, which is realized by increasing the operating frequency, leads to the reduction of the magnetic components size. Along with it, cooling surface is consequently reduced which results in high thermal stress. Careful attention must be paid on the loss mechanisms and thermal analysis of a medium frequency transformer at the design stage. This paper develops an equivalent thermal circuit of an oil-immersed medium frequency transformer to predict its temperature profile. Thermal resistance, capacitance and the losses generated in the core and winding are carefully estimated. Two oil-immersed shell-type transformer prototypes with specifications as 5kW, 500/5000V, 5kHz, interleaved winding construction have been developed and built for verification purpose. Loss and temperature measurements have been performed to verify the presented framework. The accuracy of the proposed thermal model is benchmarked and corroborated through experimental measurements as well as FEM-CFD study with good agreement.
Haonan Tian, Zhongbao Wei, Madasamy Palavesha Thevar, Sriram Vaisambhayana, Anshuman Tripathi, Philip Carne Kjaer
IECON5
2018 Speech Recognition for Medical Conversations
abstract
In this paper we document our experiences with developing speech recognition for medical transcription -a system that automatically transcribes doctor-patient conversations.Towards this goal, we built a system along two different methodological lines -a Connectionist Temporal Classification (CTC) phoneme based model and a Listen Attend and Spell (LAS) grapheme based model.To train these models we used a corpus of anonymized conversations representing approximately 14,000 hours of speech.Because of noisy transcripts and alignments in the corpus, a significant amount of effort was invested in data cleaning issues.We describe a two-stage strategy we followed for segmenting the data.The data cleanup and development of a matched language model was essential to the success of the CTC based models.The LAS based models, however were found to be resilient to alignment and transcript noise and did not require the use of language models.CTC models were able to achieve a word error rate of 20.1%, and the LAS models were able to achieve 18.3%.Our analysis shows that both models perform well on important medical utterances and therefore can be practical for transcribing medical conversations.
Chung-Cheng Chiu, Anshuman Tripathi, Katherine Chou, Chris Co, Navdeep Jaitly, Diana Jaunzeikare, Anjuli Kannan, Patrick Nguyen, Hasim Sak, Ananth Sankar, Justin Tansuwan, Nathan Wan
INTERSPEECH2
2018 Domain Adaptation Using Factorized Hidden Layer for Robust Automatic Speech Recognition
Khe Chai Sim, Arun Narayanan, Ananya Misra, Anshuman Tripathi, Golan Pundak, Tara N. Sainath, Parisa Haghani, Bo Li 0028, Michiel Bacchiani
INTERSPEECH4
2018 Toward Domain-Invariant Speech Recognition via Large Scale Training
abstract
Current state-of-the-art automatic speech recognition systems are trained to work in specific `domains', defined based on factors like application, sampling rate and codec. When such recognizers are used in conditions that do not match the training domain, performance significantly drops. This work explores the idea of building a single domain-invariant model for varied use-cases by combining large scale training data from multiple application domains. Our final system is trained using 162,000 hours of speech. Additionally, each utterance is artificially distorted during training to simulate effects like background noise, codec distortion, and sampling rates. Our results show that, even at such a scale, a model thus trained works almost as well as those fine-tuned to specific subsets: A single model can be robust to multiple application domains, and variations like codecs and noise. More importantly, such models generalize better to unseen conditions and allow for rapid adaptation - we show that by using as little as 10 hours of data from a new domain, an adapted domain-invariant model can match performance of a domain-specific model trained from scratch using 70 times as much data. We also highlight some of the limitations of such models and areas that need addressing in future work.
Arun Narayanan, Ananya Misra, Khe Chai Sim, Golan Pundak, Anshuman Tripathi, Mohamed G. Elfeky, Parisa Haghani, Trevor Strohman, Michiel Bacchiani
SLT5
2017 Thermal modeling and transient behavior analysis of a medium-frequency high-power transformer
abstract
A medium/high-power conversion system, using power electronic (PE) converter in conjunction with a medium/high-frequency transformer, has many desirable effects suitably oriented for modern power system architecture. Switching at high frequency results in lesser volume of magnetics but induces higher loss density. Thus design and characterization of a medium-frequency (MF) high-power (HP) transformer has significant ramification on its performance and application. Thermal management of a MF HP transformer is one of key aspects for its characterization. In this paper, an equivalent thermal model of a multi-layer concentrated winding is derived. Core and copper losses are carefully estimated. Thermal resistance and capacitance are accurately calculated. Time-domain response of proposed thermal network is obtained using Heun's method (Modified Euler) and validated with PLECS. Effects of temperature change on thermal properties of material and coolant (transformer oil) are also discussed. Furthermore, accuracy of said thermal network is corroborated through FEM-CFD study of a 10kW, 0.5/2.5kV, 1kHz natural oil-cooled transformer. Close agreement between analytical and simulation results is observed which substantiates proposed thermal model in terms of accuracy and efficacy of computation.
Annoy Kumar Das, Zhongbao Wei, Sriram Vaisambhayana, Shuyu Cao, Haonan Tian, Anshuman Tripathi, Philip Came Kjar
IECON6