VLDB 2026 Research / reviewers in the wild / expert
Aarthi M. Reddy
dblp:21/9279
· DBLP profile ↗
9ranked-venue papers
6as first author
0since 2021 · last 2014
0000-0002-6898-6473ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Machine translation · 60% Speech recognition and synthesis · 20% Deep learning architectures and training · 20% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
speech translation |
0.1 | 1 | 2010 | Integration of Statistical Models for Dictation of Document Translations in a Machine-Aided Human Translation Task · IEEE Trans. Speech Audio Process. 2010 |
Machine learning › Deep learning architectures and training
soft masking |
0.1 | 1 | 2007 | Soft Mask Methods for Single-Channel Speaker Separation · IEEE Trans. Speech Audio Process. 2007 |
Natural language and speech › Speech recognition and synthesis
speech separation |
0.1 | 1 | 2007 | Soft Mask Methods for Single-Channel Speaker Separation · IEEE Trans. Speech Audio Process. 2007 |
Methods — techniques the papers use, named apart from their topics
statistical machine translation · 0.1named entity recognition · 0.1factored approximation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | MVA: The Multimodal Virtual AssistantabstractMichael Johnston, John Chen, Patrick Ehlen, Hyuckchul Jung, Jay Lieske, Aarthi Reddy, Ethan Selfridge, Svetlana Stoyanchev, Brant Vasilieff, Jay Wilpon. Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL). 2014. Michael Johnston, John Chen 0001, Patrick Ehlen, Hyuckchul Jung, Jay Lieske, Aarthi M. Reddy, Ethan Selfridge, Svetlana Stoyanchev, Brant Vasilieff, Jay G. Wilpon |
SIGDIAL Conference | 6 |
| 2012 | Efficient integration of translation and speech models in dictation based machine aided human translationabstractThis paper is concerned with combining models for decoding an optimum translation for a dictation based machine aided human translation (MAHT) task. Statistical language model (SLM) probabilities in automatic speech recognition (ASR) are updated using statistical machine translation (SMT) model probabilities. The effect of this procedure is evaluated for utterances from human translators dictating translations of source language documents. It is shown that computational complexity is significantly reduced while at the same time word error rate is reduced by 30%. Luis Rodríguez, Aarthi M. Reddy, Richard C. Rose |
ICASSP | 2 |
| 2010 | Subword-based spoken term detection in audio course lecturesabstractThis paper investigates spoken term detection (STD) from audio recordings of course lectures obtained from an existing media repository. STD is performed from word lattices generated offline using an automatic speech recognition (ASR) system configured from a meetings domain. An efficient STD approach is presented where lattice paths which are likely to contain search terms are identified and an efficient phone based distance is used to detect the occurrence of search terms in phonetic expansions of promising lattice paths. STD and ASR results are reported for both in-vocabulary (IV) and out-of-vocabulary (OOV) search terms in this lecture speech domain. Richard C. Rose, Atta Norouzian, Aarthi M. Reddy, André Coy, Vishwa Gupta, Martin Karafiát |
ICASSP | 3 |
| 2010 | Integration of Statistical Models for Dictation of Document Translations in a Machine-Aided Human Translation TaskabstractThis paper presents a model for machine-aided human translation (MAHT) that integrates source language text and target language acoustic information to produce the text translation of source language document. It is evaluated on a scenario where a human translator dictates a first draft target language translation of a source language document. Information obtained from the source language document, including translation probabilities derived from statistical machine translation (SMT) and named entity tags derived from named entity recognition (NER), is incorporated with acoustic phonetic information obtained from an automatic speech recognition (ASR) system. One advantage of the system combination used here is that words that are not included in the ASR vocabulary can be correctly decoded by the combined system. The MAHT model and system implementation is presented. It is shown that a relative decrease in word error rate of 29% can be obtained by this combined system relative to the baseline ASR performance on a French to English document translation task in the Hansard domain. In addition, it is shown that transcriptions obtained by using the combined system show a relative increase in NIST score of 34% compared to transcriptions obtained from the baseline ASR system. Aarthi M. Reddy, Richard C. Rose |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Incorporating Knowledge of Source Language Text in a System for Dictation of Document Translations
Aarthi M. Reddy, Richard C. Rose, Hani Safadi, Samuel Larkin, Gilles Boulianne |
MTSummit | 1 |
| 2008 | Towards domain independence in machine aided human translationabstractThis paper presents an approach for integrating statistical ma-chine translation and automatic speech recognition for machine aided human translation (MAHT). It is applied to the problem of improving ASR performance for a human translator dictating translations in a target language while reading from a source language document. The approach addresses the issues asso-ciated with task independent ASR including out of vocabulary words and mismatched language models. We show in this paper that by obtaining domain information from the document in the form of labelled named entities from the source language text the accuracy of the ASR system can be improved by 34.5%. 1. Aarthi M. Reddy, Richard C. Rose |
INTERSPEECH | 1 |
| 2007 | Integration of ASR and machine translation models in a document translation taskabstractThis paper is concerned with the problem of machine aided human language translation. It addresses a translation scenario where a human translator dictates the spoken language translation of a source language text into an automatic speech dictation system. The source language text in this scenario is also presented to a statistical machine translation system (SMT). The techniques presented in the paper assume that the optimum target language word string which is produced by the dictation system is modeled using the combined SMT and ASR statistical models. These techniques were evaluated on a speech corpus involving human translators dictating English language translations of French language text obtained from transcriptions of the proceedings of the Canadian House of Commons. It will be shown in the paper that the combined ASR/SMT modeling techniques described in the paper were able to reduce ASR WER by 26.6 percent relative to the WER of an ASR system that did not incorporate SMT knowledge. 1. Aarthi M. Reddy, Richard C. Rose, Alain Désilets |
INTERSPEECH | 1 |
| 2007 | Soft Mask Methods for Single-Channel Speaker SeparationabstractThe problem of single-channel speaker separation attempts to extract a speech signal uttered by the speaker of interest from a signal containing a mixture of acoustic signals. Most algorithms that deal with this problem are based on masking, wherein unreliable frequency components from the mixed signal spectrogram are suppressed, and the reliable components are inverted to obtain the speech signal from speaker of interest. Most current techniques estimate this mask in a binary fashion, resulting in a hard mask. In this paper, we present two techniques to separate out the speech signal of the speaker of interest from a mixture of speech signals. One technique estimates all the spectral components of the desired speaker. The second technique estimates a soft mask that weights the frequency subbands of the mixed signal. In both cases, the speech signal of the speaker of interest is reconstructed from the complete spectral descriptions obtained. In their native form, these algorithms are computationally expensive. We also present fast factored approximations to the algorithms. Experiments reveal that the proposed algorithms can result in significant enhancement of individual speakers in mixed recordings, consistently achieving better performance than that obtained with hard binary masks. Aarthi M. Reddy, Bhiksha Raj |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | A minimum mean squared error estimator for single channel speaker separationabstractThe problem of separating out the signals for multiple speakers from a single mixed recording has received considerable atten- tio ni n recent times. Most current techniques are based on the principle of masking :i n order the separate out the signal for any speaker, frequency components that are not believed to be- long to that speaker are suppressed. The signals for the speaker is reconstructed fro mt hepartial spectral information that re- mains. In this paper we present a different kind of technique - one that attempts to estimate all spectral components for the desired speaker. Separated signals are derived from the com- plete spectral descriptions so obtained. Experiments show that this method results in superior reconstruction to masking based methods. form representations of the various speakers by hidden Markov models (HMMs). The parameters of the HMM for any speaker are learnt from training data recorded from the speaker. In addi- tion, Roweis assumes that the log energy in any frequency band of the mixed signal at any time can be attributed to only one of the speakers. This log-max assumption is justified by two observations. First, when two or more speakers speak simulta- neously, at any time, any given frequency band is usually domi- nated by a single speaker. Second, in any given frequency band the disparity in the energy levels of the dominant speaker and the other speakers is such that the logarithm of the sum of the energies of the individual speakers can be well approximated by the logarithm of the energy of the dominant speaker. In or- der to reconstruct the signal for any speaker, Roweis estimates the mask for that speaker, i.e. the identity of the time-frequency locations where the speaker dominates. The entire signal is re- constructed entirely from the masked spectrum for the speaker, i.e. fro mt he spectral components identified by the mask. The results achieved with this method are remarkably good. Hershey et. al. (6) augment audio recordings with visual features, such as lip and facial movement, in order to enhance the separation. Additionally, the ys eparate the signal into mul- tiple frequency bands, which are then processed independently. As in Roweis' algorithm, the signals for the individual speakers are reconstructed from masked spectra. Aarthi M. Reddy, Bhiksha Raj |
INTERSPEECH | 1 |