Ashish Vaswani

dblp:26/9012 · DBLP profile ↗
← Back
29ranked-venue papers
7as first author
6since 2021 · last 2023
0000-0002-7794-2085ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
Deep learning architectures and training · 33% Image recognition and object detection · 17% Machine translation · 14%
Computer graphics and multimedia
2 papers
Audio and music processing · 54% Image and video processing · 46%

Topics — the 30 heaviest of 38, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
image classification
1.442021
Scaling Local Self-Attention for Parameter Efficient Visual Backbones · CVPR 2021
Stand-Alone Self-Attention in Vision Models · NeurIPS 2019
Attention Augmented Convolutional Networks · ICCV 2019
Machine learning › Deep learning architectures and training
attention mechanism
1.032019
Stand-Alone Self-Attention in Vision Models · NeurIPS 2019
Attention Augmented Convolutional Networks · ICCV 2019
Attention is All you Need · NIPS 2017
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
1.032019
Stand-Alone Self-Attention in Vision Models · NeurIPS 2019
Attention Augmented Convolutional Networks · ICCV 2019
Attention is All you Need · NIPS 2017
Computer vision › Image recognition and object detection
object detection
1.032021
Bottleneck Transformers for Visual Recognition · CVPR 2021
Attention Augmented Convolutional Networks · ICCV 2019
Stand-Alone Self-Attention in Vision Models · NeurIPS 2019
Natural language and speech › Machine translation
neural machine translation
0.732018
The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation · ACL (1) 2018
Attention is All you Need · NIPS 2017
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Machine learning › Generative modeling
autoregressive model
0.722019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Image Transformer · ICML 2018
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pre-training and fine-tuning
0.612022
Scale Efficiently: Insights from Pretraining and Finetuning Transformers · ICLR 2022
Machine learning › Deep learning architectures and training
scaling laws
0.612022
Scale Efficiently: Insights from Pretraining and Finetuning Transformers · ICLR 2022
Computer vision › Segmentation and scene understanding
instance segmentation
0.512021
Bottleneck Transformers for Visual Recognition · CVPR 2021
Machine learning › Deep learning architectures and training › attention mechanism › attention network
self-attention network
0.512021
Bottleneck Transformers for Visual Recognition · CVPR 2021
Machine learning › Deep learning architectures and training
transformer
0.422019
Image Transformer · ICML 2018
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Natural language and speech › Machine translation
decipherment
0.422015
Unifying Bayesian Inference and Vector Space Models for Improved Decipherment · ACL (1) 2015
Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014
Computer vision › Vision and language
vision-and-language navigation
0.412019
Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation · ACL (1) 2019
Audio and music processing
music generation
0.412019
Music Transformer: Generating Music with Long-Term Structure · ICLR (Poster) 2019
Natural language and speech › Machine translation
statistical machine translation
0.422014
Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014
Decoding with Large-Scale Neural Language Models Improves Translation · EMNLP 2013
Natural language and speech › Machine translation › statistical machine translation
word alignment
0.322014
Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014
Smaller Alignment Models for Better Translations: Unsupervised Word Alignment with the l0-norm · ACL (1) 2012
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable
0.312018
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Machine learning › Efficient and distributed learning
distributed training
0.312018
Mesh-TensorFlow: Deep Learning for Supercomputers · NeurIPS 2018
Machine learning › Generative modeling
image generation
0.312018
Image Transformer · ICML 2018
Machine learning › Efficient and distributed learning › distributed training
model parallelism
0.312018
Mesh-TensorFlow: Deep Learning for Supercomputers · NeurIPS 2018
Natural language and speech › Language models and text generation › decoding › decoding strategy
parallel decoding
0.312018
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Machine learning › Deep learning architectures and training
sequence modeling
0.312018
Fast Decoding in Sequence Models Using Discrete Latent Variables · ICML 2018
Image and video processing
super-resolution
0.312018
Image Transformer · ICML 2018
Machine learning › Deep learning architectures and training › sequence modeling › sequence generation
sequence transduction
0.312017
Attention is All you Need · NIPS 2017
Natural language and speech › Language models and text generation › neural language model
neural language model representations
0.212014
Aligning context-based statistical models of language with brain activity during reading · EMNLP 2014
Natural language and speech › Language models and text generation
neural language model
0.212013
Decoding with Large-Scale Neural Language Models Improves Translation · EMNLP 2013
Natural language and speech › Machine translation › syntax-based machine translation
tree-to-string translation
0.112011
Rule Markov Models for Fast Tree-to-String Translation · ACL 2011
Natural language and speech › Language models and text generation › language modeling › language model architecture
sequence-to-sequence model
0.112018
Mesh-TensorFlow: Deep Learning for Supercomputers · NeurIPS 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference
0.112015
Unifying Bayesian Inference and Vector Space Models for Improved Decipherment · ACL (1) 2015
Natural language and speech › Machine translation
low-resource machine translation
0.112014
Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation · EMNLP 2014

Methods — techniques the papers use, named apart from their topics

self-attention · 2.8convolutional neural network · 1.3autoregressive modeling · 1.0relative positional encoding · 0.8scaling analysis · 0.6convolutional hybrid · 0.5reward shaping · 0.4relative self-attention · 0.4coverage weighted by length score · 0.4collective communication · 0.3attention · 0.3allreduce · 0.3SPMD programming · 0.3
YearPublicationVenuePosition
2023 Guest Editorial Introduction to the Special Section on Transformer Models in Vision
abstract
Transformer models have achieved outstanding results on a variety of language tasks, such as text classification, ma- chine translation, and question answering. This success in the field of Natural Language Processing (NLP) has sparked interest in the computer vision community to apply these models to vision and multi-modal learning tasks. However, visual data has a unique structure, requiring the need to rethink network designs and training methods. As a result, Transformer models and their variations have been suc- cessfully used for image recognition, object detection, seg- mentation, image super-resolution, video understanding, image generation, text-image synthesis, and visual question answering, among other applications.
Salman Khan 0001, Fahad Shahbaz Khan, Ashish Vaswani, Niki Parmar, Ming-Hsuan Yang 0001, Mubarak Shah
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 The Efficiency Misnomer
Mostafa Dehghani 0001, Yi Tay, Anurag Arnab, Lucas Beyer, Ashish Vaswani
ICLR5
2022 Scale Efficiently: Insights from Pretraining and Finetuning Transformers
Yi Tay, Mostafa Dehghani 0001, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, Donald Metzler
ICLR9
2021 Bottleneck Transformers for Visual Recognition
abstract
We present BoTNet, a conceptually simple yet powerful backbone architecture that incorporates self-attention for multiple computer vision tasks including image classification, object detection and instance segmentation. By just replacing the spatial convolutions with global self-attention in the final three bottleneck blocks of a ResNet and no other changes, our approach improves upon the baselines significantly on instance segmentation and object detection while also reducing the parameters, with minimal overhead in latency. Through the design of BoTNet, we also point out how ResNet bottleneck blocks with self-attention can be viewed as Transformer blocks. Without any bells and whistles, BoTNet achieves 44.4% Mask AP and 49.7% Box AP on the COCO Instance Segmentation benchmark using the Mask R-CNN framework; surpassing the previous best published single model and single scale results of ResNeSt [67] evaluated on the COCO validation set. Finally, we present a simple adaptation of the BoTNet design for image classification, resulting in models that achieve a strong performance of 84.7% top-1 accuracy on the ImageNet benchmark while being up to 1.64x faster in "compute"1time than the popular EfficientNet models on TPU-v3 hardware. We hope our simple and effective approach will serve as a strong baseline for future research in self-attention models for vision.2
Aravind Srinivas, Tsung-Yi Lin, Niki Parmar, Jonathon Shlens, Pieter Abbeel, Ashish Vaswani
CVPR6
2021 Scaling Local Self-Attention for Parameter Efficient Visual Backbones
abstract
Self-attention has the promise of improving computer vision systems due to parameter-independent scaling of receptive fields and content-dependent interactions, in contrast to parameter-dependent scaling and content-independent interactions of convolutions. Self-attention models have recently been shown to have encouraging improvements on accuracy-parameter trade-offs compared to baseline convolutional models such as ResNet-50. In this work, we develop self-attention models that can outperform not just the canonical baseline models, but even the high-performing convolutional models. We propose two extensions to self-attention that, in conjunction with a more efficient implementation of self-attention, improve the speed, memory usage, and accuracy of these models. We leverage these improvements to develop a new self-attention model family, HaloNets, which reach state-of-the-art accuracies on the parameter-limited setting of the ImageNet classification benchmark. In preliminary transfer learning experiments, we find that HaloNet models outperform much larger models and have better inference performance. On harder tasks such as object detection and instance segmentation, our simple local self-attention and convolutional hybrids show improvements over very strong baselines. These results mark another step in demonstrating the efficacy of self-attention models on settings traditionally dominated by convolutions.1
Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake A. Hechtman, Jonathon Shlens
CVPR1
2021 Efficient Content-Based Sparse Attention with Routing Transformers
abstract
Self-attention has recently been adopted for a wide range of sequence modeling problems. Despite its effectiveness, self-attention suffers from quadratic computation and memory requirements with respect to sequence length. Successful approaches to reduce this complexity focused on attending to local sliding windows or a small set of locations independent of content. Our work proposes to learn dynamic sparse attention patterns that avoid allocating computation and memory to attend to content unrelated to the query of interest. This work builds upon two lines of research: It combines the modeling flexibility of prior work on content-based sparse attention with the efficiency gains from approaches based on local, temporal sparse attention. Our model, the Routing Transformer, endows self-attention with a sparse routing module based on online k-means while reducing the overall complexity of attention to O( n 1.5 d) from O( n 2 d) for sequence length n and hidden dimension d. We show that our model outperforms comparable sparse attention models on language modeling on Wikitext-103 (15.8 vs 18.3 perplexity), as well as on image generation on ImageNet-64 (3.43 vs 3.44 bits/dim) while using fewer self-attention layers. Additionally, we set a new state-of-the-art on the newly released PG-19 data-set, obtaining a test perplexity of 33.2 with a 22 layer Routing Transformer model trained on sequences of length 8192. We open-source the code for Routing Transformer in Tensorflow. 1
Aurko Roy, Mohammad Saffar, Ashish Vaswani, David Grangier
Trans. Assoc. Comput. Linguistics3
2019 Stay on the Path: Instruction Fidelity in Vision-and-Language Navigation
abstract
Advances in learning and representations have reinvigorated work that connects language to other modalities.A particularly exciting direction is Vision-and-Language Navigation (VLN), in which agents interpret natural language instructions and visual scenes to move through environments and reach goals.Despite recent progress, current research leaves unclear how much of a role language understanding plays in this task, especially because dominant evaluation metrics have focused on goal completion rather than the sequence of actions corresponding to the instructions.Here, we highlight shortcomings of current metrics for the Room-to-Room dataset (Anderson et al., 2018b) and propose a new metric, Coverage weighted by Length Score (CLS).We also show that the existing paths in the dataset are not ideal for evaluating instruction following because they are direct-to-goal shortest paths.We join existing short paths to form more challenging extended paths to create a new data set, Room-for-Room (R4R).Using R4R and CLS, we show that agents that receive rewards for instruction fidelity outperform agents that focus on goal completion.
Vihan Jain, Gabriel Ilharco, Alexander Ku, Ashish Vaswani, Eugene Ie, Jason Baldridge
ACL (1)4
2019 Attention Augmented Convolutional Networks
abstract
Convolutional networks have enjoyed much success in many computer vision applications. The convolution operation however has a significant weakness in that it only operates on a local neighbourhood, thus missing global information. Self-attention, on the other hand, has emerged as a recent advance to capture long range interactions, but has mostly been applied to sequence modeling and generative modeling tasks. In this paper, we propose to augment convolutional networks with self-attention by concatenating convolutional feature maps with a set of feature maps produced via a novel relative self-attention mechanism. In particular, we extend previous work on relative self-attention over sequences to images and discuss a memory efficient implementation. Unlike Squeeze-and-Excitation, which performs attention over the channels and ignores spatial information, our self-attention mechanism attends jointly to both features and spatial locations while preserving translation equivariance. We find that Attention Augmentation leads to consistent improvements in image classification on ImageNet and object detection on COCO across many different models and scales, including ResNets and a state-of-the art mobile constrained network, while keeping the number of parameters similar. In particular, our method achieves a 1.3% top-1 accuracy improvement on ImageNet classification over a ResNet50 baseline and outperforms other attention mechanisms for images such as Squeeze-and-Excitation. It also achieves an improvement of 1.4 AP in COCO Object Detection on top of a RetinaNet baseline.
Irwan Bello, Barret Zoph, Quoc V. Le, Ashish Vaswani, Jonathon Shlens
ICCV4
2019 Music Transformer: Generating Music with Long-Term Structure
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Ian Simon, Curtis Hawthorne, Noam Shazeer, Andrew M. Dai, Matthew Hoffman 0001, Monica Dinculescu, Douglas Eck
ICLR (Poster)2
2019 Stand-Alone Self-Attention in Vision Models
abstract
Convolutions are a fundamental building block of modern computer vision systems. Recent approaches have argued for going beyond convolutions in order to capture long-range dependencies. These efforts focus on augmenting convolutional models with content-based interactions, such as self-attention and non-local means, to achieve gains on a number of vision tasks. The natural question that arises is whether attention can be a stand-alone primitive for vision models instead of serving as just an augmentation on top of convolutions. In developing and testing a pure self-attention vision model, we verify that self-attention can indeed be an effective stand-alone layer. A simple procedure of replacing all instances of spatial convolutions with a form of self-attention to ResNet-50 produces a fully self-attentional model that outperforms the baseline on ImageNet classification with 12% fewer FLOPS and 29% fewer parameters. On COCO object detection, a fully self-attention model matches the mAP of a baseline RetinaNet while having 39% fewer FLOPS and 34% fewer parameters. Detailed ablation studies demonstrate that self-attention is especially impactful when used in later layers. These results establish that stand-alone self-attention is an important addition to the vision practitioner's toolbox.
Niki Parmar, Prajit Ramachandran, Ashish Vaswani, Irwan Bello, Anselm Levskaya, Jonathon Shlens
NeurIPS3
2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
abstract
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes
ACL (1)11
2018 Fast Decoding in Sequence Models Using Discrete Latent Variables
abstract
Autoregressive sequence models based on deep neural networks, such as RNNs, Wavenet and Transformer are the state-of-the-art on many tasks. However, they lack parallelism and are thus slow for long sequences. RNNs lack parallelism both during training and decoding, while architectures like WaveNet and Transformer are much more parallel during training, but still lack parallelism during decoding. We present a method to extend sequence models using discrete latent variables that makes decoding much more parallel. The main idea behind this approach is to first autoencode the target sequence into a shorter discrete latent sequence, which is generated autoregressively, and finally decode the full sequence from this shorter latent sequence in a parallel manner. To this end, we introduce a new method for constructing discrete latent variables and compare it with previously introduced methods. Finally, we verify that our model works on the task of neural machine translation, where our models are an order of magnitude faster than comparable autoregressive models and, while lower in BLEU than purely autoregressive models, better than previously proposed non-autogregressive translation.
Lukasz Kaiser, Samy Bengio, Aurko Roy, Ashish Vaswani, Niki Parmar, Jakob Uszkoreit, Noam Shazeer
ICML4
2018 Image Transformer
abstract
Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran
ICML2
2018 Mesh-TensorFlow: Deep Learning for Supercomputers
abstract
Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to memory constraints), high latency, and inefficiency at small batch sizes. All of these can be solved by more general distribution strategies (model-parallelism). Unfortunately, efficient model-parallel algorithms tend to be complicated to discover, describe, and to implement, particularly on large clusters. We introduce Mesh-TensorFlow, a language for specifying a general class of distributed tensor computations. Where data-parallelism can be viewed as splitting tensors and operations along the "batch" dimension, in Mesh-TensorFlow, the user can specify any tensor-dimensions to be split across any dimensions of a multi-dimensional mesh of processors. A Mesh-TensorFlow graph compiles into a SPMD program consisting of parallel operations coupled with collective communication primitives such as Allreduce. We use Mesh-TensorFlow to implement an efficient data-parallel, model-parallel version of the Transformer sequence-to-sequence model. Using TPU meshes of up to 512 cores, we train Transformer models with up to 5 billion parameters, surpassing SOTA results on WMT'14 English-to-French translation task and the one-billion-word Language modeling benchmark. Mesh-Tensorflow is available at https://github.com/tensorflow/mesh
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, Blake A. Hechtman
NeurIPS5
2017 Attention is All you Need
abstract
The dominant sequence transduction models are based on complex recurrent orconvolutional neural networks in an encoder and decoder configuration. The best performing such models also connect the encoder and decoder through an attentionm echanisms. We propose a novel, simple network architecture based solely onan attention mechanism, dispensing with recurrence and convolutions entirely.Experiments on two machine translation tasks show these models to be superiorin quality while being more parallelizable and requiring significantly less timeto train. Our single model with 165 million parameters, achieves 27.5 BLEU onEnglish-to-German translation, improving over the existing best ensemble result by over 1 BLEU. On English-to-French translation, we outperform the previoussingle state-of-the-art with model by 0.7 BLEU, achieving a BLEU score of 41.1.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
NIPS1
2016 Supertagging With LSTMs
abstract
Ashish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Ashish Vaswani, Yonatan Bisk, Kenji Sagae, Ryan Musa
HLT-NAACL1
2016 Name Tagging for Low-resource Incident Languages based on Expectation-driven Learning
abstract
Boliang Zhang, Xiaoman Pan, Tianlu Wang, Ashish Vaswani, Heng Ji, Kevin Knight, Daniel Marcu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Boliang Zhang, Xiaoman Pan, Ashish Vaswani, Heng Ji 0001, Kevin Knight, Daniel Marcu
HLT-NAACL4
2016 Simple, Fast Noise-Contrastive Estimation for Large RNN Vocabularies
abstract
Barret Zoph, Ashish Vaswani, Jonathan May, Kevin Knight. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Barret Zoph, Ashish Vaswani, Jonathan May, Kevin Knight
HLT-NAACL2
2016 Efficient Structured Inference for Transition-Based Parsing with Neural Networks and Error States
abstract
Transition-based approaches based on local classification are attractive for dependency parsing due to their simplicity and speed, despite producing results slightly below the state-of-the-art. In this paper, we propose a new approach for approximate structured inference for transition-based parsing that produces scores suitable for global scoring using local models. This is accomplished with the introduction of error states in local training, which add information about incorrect derivation paths typically left out completely in locally-trained models. Using neural networks for our local classifiers, our approach achieves 93.61% accuracy for transition-based dependency parsing in English.
Ashish Vaswani, Kenji Sagae
Trans. Assoc. Comput. Linguistics1
2015 Unifying Bayesian Inference and Vector Space Models for Improved Decipherment
abstract
Qing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Qing Dou, Ashish Vaswani, Kevin Knight, Chris Dyer
ACL (1)2
2015 Model Invertibility Regularization: Sequence Alignment With or Without Parallel Data
abstract
We present Model Invertibility Regularization (MIR), a method that jointly trains two directional sequence alignment models, one in each direction, and takes into account the invertibility of the alignment task.By coupling the two models through their parameters (as opposed to through their inferences, as in Liang et al.'s Alignment by Agreement (ABA), and Ganchev et al.'s Posterior Regularization (PostCAT)), our method seamlessly extends to all IBMstyle word alignment models as well as to alignment without parallel data.Our proposed algorithm is mathematically sound and inherits convergence guarantees from EM.We evaluate MIR on two tasks: (1) On word alignment, applying MIR on fertility based models we attain higher F-scores than ABA and PostCAT.(2) On Japanese-to-English backtransliteration without parallel data, applied to the decipherment model of Ravi and Knight, MIR learns sparser models that close the gap in whole-name error rate by 33% relative to a model trained on parallel data, and further, beats a previous approach by Mylonakis et al.
Tomer Levinboim, Ashish Vaswani, David Chiang 0001
HLT-NAACL2
2014 Beyond Parallel Data: Joint Word Alignment and Decipherment Improves Machine Translation
abstract
Inspired by previous work, where decipherment is used to improve machine translation, we propose a new idea to combine word alignment and decipherment into a single learning process.We use EM to estimate the model parameters, not only to maximize the probability of parallel corpus, but also the monolingual corpus.We apply our approach to improve Malagasy-English machine translation, where only a small amount of parallel data is available.In our experiments, we observe gains of 0.9 to 2.1 Bleu over a strong baseline.
Qing Dou, Ashish Vaswani, Kevin Knight
EMNLP2
2014 Aligning context-based statistical models of language with brain activity during reading
abstract
Many statistical models for natural language processing exist, including context-based neural networks that (1) model the previously seen context as a latent feature vector, (2) integrate successive words into the context using some learned representation (embedding), and (3) compute output probabilities for incoming words given the context.On the other hand, brain imaging studies have suggested that during reading, the brain (a) continuously builds a context from the successive words and every time it encounters a word it (b) fetches its properties from memory and (c) integrates it with the previous context with a degree of effort that is inversely proportional to how probable the word is.This hints to a parallelism between the neural networks and the brain in modeling context (1 and a), representing the incoming words (2 and b) and integrating it (3 and c).We explore this parallelism to better understand the brain processes and the neural networks representations.We study the alignment between the latent vectors used by neural networks and brain activity observed via Magnetoencephalography (MEG) when subjects read a story.For that purpose we apply the neural network to the same text the subjects are reading, and explore the ability of these three vector representations to predict the observed word-by-word brain activity.Our novel results show that: before a new word i is read, brain activity is well predicted by the neural network latent representation of context and the predictability decreases as the brain integrates the word and changes its own representation of context.Secondly, the neural network embedding of word i can predict the MEG activity when word i is presented to the subject, revealing that it is correlated with the brain's own representation of word i.Moreover, we obtain that the activity is predicted in different regions of the brain with varying delay.The delay is consistent with the placement of each region on the processing pathway that starts in the visual cortex and moves to higher level regions.Finally, we show that the output probability computed by the neural networks agrees with the brain's own assessment of the probability of word i, as it can be used to predict the brain activity after the word i's properties have been fetched from memory and the brain is in the process of integrating it into the context.
Leila Wehbe, Ashish Vaswani, Kevin Knight, Tom M. Mitchell
EMNLP2
2013 Decoding with Large-Scale Neural Language Models Improves Translation
abstract
We explore the application of neural language models to machine translation.We develop a new model that combines the neural probabilistic language model of Bengio et al., rectified linear units, and noise-contrastive estimation, and we incorporate it into a machine translation system both by reranking k-best lists and by direct integration into the decoder.Our large-scale, large-vocabulary experiments across four language pairs show that our neural language model improves translation quality by up to 1.1 Bleu.
Ashish Vaswani, Yinggong Zhao, Victoria Fossum, David Chiang 0001
EMNLP1
2013 Learning Whom to Trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, Eduard H. Hovy
HLT-NAACL3
2012 Smaller Alignment Models for Better Translations: Unsupervised Word Alignment with the l0-norm
Ashish Vaswani, Liang Huang 0001, David Chiang 0001
ACL (1)1
2011 Rule Markov Models for Fast Tree-to-String Translation
Ashish Vaswani, Haitao Mi, Liang Huang 0001, David Chiang 0001
ACL1
2010 Fast, Greedy Model Minimization for Unsupervised Tagging
Sujith Ravi, Ashish Vaswani, Kevin Knight, David Chiang 0001
COLING2
2006 Radiobot-CFF: a spoken dialogue system for military training
abstract
We describe a spoken dialogue system which can engage in Call For Fire (CFF) radio dialogues to help train soldiers in proper procedures for requesting artillery fire missions. We describe the domain, an information-state dialogue manager with a novel system of interactive information components, and provide evaluation results. Index Terms: spoken dialogue systems. 1.
Antonio Roque, Anton Leuski, Vivek Kumar Rangarajan Sridhar, Susan Robinson, Ashish Vaswani, Shri Narayanan, David R. Traum
INTERSPEECH5