VLDB 2026 Research / reviewers in the wild / expert
Jacob Devlin
dblp:116/0575
· DBLP profile ↗
25ranked-venue papers
8as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Language models and text generation · 38% Transfer learning and domain adaptation · 14% Efficient and distributed learning · 12% | |
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
few-shot learning |
1.7 | 3 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 Neural Program Meta-Induction · NIPS 2017 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.4 | 2 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
instruction tuning |
0.8 | 1 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 |
Natural language and speech › Language models and text generation
instruction following |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Natural language and speech › Language models and text generation
large language model |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Machine learning › Deep learning architectures and training
scaling laws |
0.7 | 1 | 2023 | PaLM: Scaling Language Modeling with Pathways · J. Mach. Learn. Res. 2023 |
Program synthesis and code generation
neural program synthesis |
0.6 | 2 | 2018 | Leveraging Grammar and Reinforcement Learning for Neural Program Synthesis · ICLR (Poster) 2018 RobustFill: Neural Program Learning under Noisy I/O · ICML 2017 |
Program synthesis and code generation › inductive program synthesis
neural program induction |
0.6 | 2 | 2017 | Neural Program Meta-Induction · NIPS 2017 RobustFill: Neural Program Learning under Noisy I/O · ICML 2017 |
Natural language and speech › Machine translation
statistical machine translation |
0.6 | 3 | 2015 | Statistical Machine Translation Features with Multitask Tensor Networks · ACL (1) 2015 Fast and Robust Neural Network Joint Models for Statistical Machine Translation · ACL (1) 2014 Factored Soft Source Syntactic Constraints for Hierarchical Machine Translation · EMNLP 2013 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.5 | 2 | 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPU · EMNLP 2017 Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Machine learning › Efficient and distributed learning
model compression |
0.5 | 2 | 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPU · EMNLP 2017 Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Information retrieval › retrieval models › neural retrieval › dense retrieval
bi-encoder retrieval |
0.5 | 1 | 2021 | Multi-Vector Attention Models for Deep Re-ranking · EMNLP (1) 2021 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.5 | 1 | 2021 | Multi-Vector Attention Models for Deep Re-ranking · EMNLP (1) 2021 |
Information retrieval › reranking
document re-ranking |
0.5 | 1 | 2021 | Multi-Vector Attention Models for Deep Re-ranking · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training
data augmentation |
0.4 | 1 | 2019 | Synthetic QA Corpora Generation with Roundtrip Consistency · ACL (1) 2019 |
Natural language and speech › Language models and text generation › large language model training
domain-adaptive pre-training |
0.4 | 1 | 2019 | Zero-Shot Entity Linking by Reading Entity Descriptions · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
entity linking |
0.4 | 1 | 2019 | Zero-Shot Entity Linking by Reading Entity Descriptions · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › entity linking
zero-shot entity linking |
0.4 | 1 | 2019 | Zero-Shot Entity Linking by Reading Entity Descriptions · ACL (1) 2019 |
Machine learning › Reinforcement learning
policy learning |
0.3 | 1 | 2018 | Leveraging Grammar and Reinforcement Learning for Neural Program Synthesis · ICLR (Poster) 2018 |
Natural language and speech › Language models and text generation › code generation
program synthesis |
0.3 | 1 | 2018 | Leveraging Grammar and Reinforcement Learning for Neural Program Synthesis · ICLR (Poster) 2018 |
Program synthesis and code generation
syntax-guided synthesis |
0.3 | 1 | 2018 | Leveraging Grammar and Reinforcement Learning for Neural Program Synthesis · ICLR (Poster) 2018 |
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search |
0.3 | 1 | 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPU · EMNLP 2017 |
Natural language and speech › Language models and text generation
decoding |
0.3 | 1 | 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPU · EMNLP 2017 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.3 | 1 | 2017 | Neural Program Meta-Induction · NIPS 2017 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPU · EMNLP 2017 |
Program synthesis and code generation
inductive program synthesis |
0.3 | 1 | 2017 | Neural Program Meta-Induction · NIPS 2017 |
Computer vision › Vision and language › vision-language generation
visual question generation |
0.2 | 1 | 2016 | Generating Natural Questions About an Image · ACL (1) 2016 |
Machine learning › Efficient and distributed learning › large-scale learning
model scaling |
0.2 | 1 | 2024 | Scaling Instruction-Finetuned Language Models · J. Mach. Learn. Res. 2024 |
Natural language and speech › Language models and text generation
neural language model |
0.2 | 1 | 2015 | Pre-Computable Multi-Layer Neural Network Language Models · EMNLP 2015 |
Computer vision › Vision and language
vision-language dataset |
0.2 | 1 | 2015 | A Survey of Current Datasets for Vision and Language Research · EMNLP 2015 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.7pathways · 0.7grammar · 0.7multi-vector attention · 0.5learned pooling · 0.5cross-attention · 0.5sequence-to-sequence pre-training · 0.4reading comprehension model · 0.4question generation · 0.4pre-training · 0.4answer extraction · 0.4BERT fine-tuning · 0.4BERT · 0.4reinforcement learning · 0.3transfer learning · 0.3portfolio adaptation · 0.3attention RNN · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Scaling Instruction-Finetuned Language ModelsabstractFinetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation, RealToxicityPrompts). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PaLM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks (at time of release), such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints,1 which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models. Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang 0002, Mostafa Dehghani 0001, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu 0001, Slav Petrov, Ed H. Chi, Jeffrey Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei |
J. Mach. Learn. Res. | 31 |
| 2023 | PaLM: Scaling Language Modeling with PathwaysabstractLarge language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies. Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel |
J. Mach. Learn. Res. | 3 |
| 2021 | Multi-Vector Attention Models for Deep Re-rankingabstractLarge-scale document retrieval systems often utilize two styles of neural network models which live at two different ends of the joint computation vs. accuracy spectrum.The first style is dual encoder (or two-tower) models, where the query and document representations are computed completely independently and combined with a simple dot product operation.The second style is cross-attention models, where the query and document features are concatenated in the input layer and all computation is based on the joint querydocument representation.Dual encoder models are typically used for retrieval and deep re-ranking, while cross-attention models are typically used for shallow re-ranking.In this paper, we present a lightweight architecture that explores this joint cost vs.accuracy trade-off based on multi-vector attention (MVA).We thoroughly evaluate our method on the MS-MARCO passage retrieval dataset and show how to efficiently trade off retrieval accuracy with joint computation and offline document storage cost.We show that a highly compressed document representation and inexpensive joint computation can be achieved through a combination of learned pooling tokens and aggressive downprojection.Our code and model checkpoints are available on GitHub. Giulio Zhou, Jacob Devlin |
EMNLP (1) | 2 |
| 2019 | Synthetic QA Corpora Generation with Roundtrip ConsistencyabstractWe introduce a novel method of generating synthetic question answering corpora by combining models of question generation and answer extraction, and by filtering the results to ensure roundtrip consistency.By pretraining on the resulting corpora we obtain significant improvements on SQuAD2 (Rajpurkar et al., 2018) and NQ (Kwiatkowski et al., 2019), establishing a new state-of-the-art on the latter.Our synthetic data generation models, for both question generation and answer extraction, can be fully reproduced by finetuning a publicly available BERT model (Devlin et al., 2018) on the extractive subsets of SQuAD2 and NQ.We also describe a more powerful variant that does full sequence-to-sequence pretraining for question generation, obtaining exact match and F1 at less than 0.1% and 0.4% from human performance on SQuAD2. Christopher Alberti, Daniel Andor, Emily Pitler, Jacob Devlin, Michael Collins 0001 |
ACL (1) | 4 |
| 2019 | Zero-Shot Entity Linking by Reading Entity DescriptionsabstractWe present the zero-shot entity linking task, where mentions must be linked to unseen entities without in-domain labeled data.The goal is to enable robust transfer to highly specialized domains, and so no metadata or alias tables are assumed.In this setting, entities are only identified by text descriptions, and models must rely strictly on language understanding to resolve the new entities.First, we show that strong reading comprehension models pre-trained on large unlabeled data can be used to generalize to unseen entities.Second, we propose a simple and effective adaptive pre-training strategy, which we term domainadaptive pre-training (DAP), to address the domain shift problem associated with linking unseen entities in a new domain.We present experiments on a new dataset that we construct for this task and show that DAP improves over strong pre-training baselines, including BERT. Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, Honglak Lee |
ACL (1) | 5 |
| 2019 | Natural Questions: a Benchmark for Question Answering ResearchabstractWe present the Natural Questions corpus, a question answering data set. Questions consist of real anonymized, aggregated queries issued to the Google search engine. An annotator is presented with a question along with a Wikipedia page from the top 5 search results, and annotates a long answer (typically a paragraph) and a short answer (one or more entities) if present on the page, or marks null if no long/short answer is present. The public release consists of 307,373 training examples with single annotations; 7,830 examples with 5-way annotations for development data; and a further 7,842 examples with 5-way annotated sequestered as test data. We present experiments validating quality of the data. We also describe analysis of 25-way annotations on 302 examples, giving insights into human variability on the annotation task. We introduce robust metrics for the purposes of evaluating question answering systems; demonstrate high human upper bounds on these metrics; and establish baseline results using competitive methods drawn from related literature. Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins 0001, Ankur P. Parikh, Christopher Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, Slav Petrov |
Trans. Assoc. Comput. Linguistics | 9 |
| 2018 | Leveraging Grammar and Reinforcement Learning for Neural Program Synthesis
Rudy Bunel, Matthew J. Hausknecht, Jacob Devlin, Rishabh Singh, Pushmeet Kohli |
ICLR (Poster) | 3 |
| 2018 | Universal Neural Machine Translation for Extremely Low Resource LanguagesabstractJiatao Gu, Hany Hassan, Jacob Devlin, Victor O.K. Li. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jiatao Gu, Hany Hassan, Jacob Devlin, Victor O. K. Li |
NAACL-HLT | 3 |
| 2017 | Sharp Models on Dull Hardware: Fast and Accurate Neural Machine Translation Decoding on the CPUabstractAttentional sequence-to-sequence models have become the new standard for machine translation, but one challenge of such models is a significant increase in training and decoding cost compared to phrase-based systems.Here, we focus on efficient decoding, with a goal of achieving accuracy close the state-of-the-art in neural machine translation (NMT), while achieving CPU decoding speed/throughput close to that of a phrasal decoder.We approach this problem from two angles: First, we describe several techniques for speeding up an NMT beam search decoder, which obtain a 4.4x speedup over a very efficient baseline decoder without changing the decoder output.Second, we propose a simple but powerful network architecture which uses an RNN (GRU/LSTM) layer at bottom, followed by a series of stacked fully-connected layers applied at every timestep.This architecture achieves similar accuracy to a deep recurrent model, at a small fraction of the training and decoding cost.By combining these techniques, our best system achieves a very competitive accuracy of 38.3 BLEU on WMT English-French NewsTest2014, while decoding at 100 words/sec on single-threaded CPU.We believe this is the best published accuracy/speed trade-off of an NMT system. Jacob Devlin |
EMNLP | 1 |
| 2017 | RobustFill: Neural Program Learning under Noisy I/OabstractThe problem of automatically generating a computer program from some specification has been studied since the early days of AI. Recently, two competing approaches for `automatic program learning’ have received significant attention: (1) `neural program synthesis’, where a neural network is conditioned on input/output (I/O) examples and learns to generate a program, and (2) `neural program induction’, where a neural network generates new outputs directly using a latent program representation. Here, for the first time, we directly compare both approaches on a large-scale, real-world learning task and we additionally contrast to rule-based program synthesis, which uses hand-crafted semantics to guide the program generation. Our neural models use a modified attention RNN to allow encoding of variable-sized sets of I/O pairs, which achieve 92\% accuracy on a real-world test set, compared to the 34\% accuracy of the previous best neural synthesis approach. The synthesis model also outperforms a comparable induction model on this task, but we more importantly demonstrate that the strength of each approach is highly dependent on the evaluation metric and end-user application. Finally, we show that we can train our neural models to remain very robust to the type of noise expected in real-world data (e.g., typos), while a highly-engineered rule-based system fails entirely. Jacob Devlin, Jonathan Uesato, Surya Bhupatiraju, Rishabh Singh, Abdel-rahman Mohamed, Pushmeet Kohli |
ICML | 1 |
| 2017 | Neural Program Meta-InductionabstractMost recently proposed methods for Neural Program induction work under the assumption of having a large set of input/output (I/O) examples for learning any given input-output mapping. This paper aims to address the problem of data and computation efficiency of program induction by leveraging information from related tasks. Specifically, we propose two novel approaches for cross-task knowledge transfer to improve program induction in limited-data scenarios. In our first proposal, portfolio adaptation, a set of induction models is pretrained on a set of related tasks, and the best model is adapted towards the new task using transfer learning. In our second approach, meta program induction, a $k$-shot learning approach is used to make a model generalize to new tasks without additional training. To test the efficacy of our methods, we constructed a new benchmark of programs written in the Karel programming language. Using an extensive experimental evaluation on the Karel benchmark, we demonstrate that our proposals dramatically outperform the baseline induction method that does not use knowledge transfer. We also analyze the relative performance of the two approaches and study conditions in which they perform best. In particular, meta induction outperforms all existing approaches under extreme data sparsity (when a very small number of examples are available), i.e., fewer than ten. As the number of available I/O examples increase (i.e. a thousand or more), portfolio adapted program induction becomes the best approach. For intermediate data sizes, we demonstrate that the combined method of adapted meta program induction has the strongest performance. Jacob Devlin, Rudy Bunel, Rishabh Singh, Matthew J. Hausknecht, Pushmeet Kohli |
NIPS | 1 |
| 2016 | Generating Natural Questions About an ImageabstractNasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He, Lucy Vanderwende. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Margaret Mitchell, Xiaodong He 0001, Lucy Vanderwende |
ACL (1) | 3 |
| 2016 | Visual StorytellingabstractTing-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Ting-Hao 'Kenneth' Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross B. Girshick, Xiaodong He 0001, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, Margaret Mitchell |
HLT-NAACL | 6 |
| 2015 | Statistical Machine Translation Features with Multitask Tensor NetworksabstractWe present a three-pronged approach to improving Statistical Machine Translation (SMT), building on recent success in the application of neural networks to SMT. First, we propose new features based on neural networks to model various non-local translation phenomena. Second, we augment the architecture of the neural network with tensor layers that capture important higher-order interaction among the network units. Third, we apply multitask learning to estimate the neural network parameters jointly. Each of our proposed methods results in significant improvements that are complementary. The overall improvement is +2.7 and +1.8 BLEU points for Arabic-English and Chinese-English translation over a state-of-the-art system that already includes neural network features. Hendra Setiawan, Zhongqiang Huang, Jacob Devlin, Thomas Lamar, Rabih Zbib, Richard M. Schwartz, John Makhoul |
ACL (1) | 3 |
| 2015 | Pre-Computable Multi-Layer Neural Network Language ModelsabstractIn the last several years, neural network models have significantly improved accuracy in a number of NLP tasks.However, one serious drawback that has impeded their adoption in production systems is the slow runtime speed of neural network models compared to alternate models, such as maximum entropy classifiers.In Devlin et al. (2014), the authors presented a simple technique for speeding up feed-forward embedding-based neural network models, where the dot product between each word embedding and part of the first hidden layer are pre-computed offline.However, this technique cannot be used for hidden layers beyond the first.In this paper, we explore a neural network architecture where the embedding layer feeds into multiple hidden layers that are placed "next to" one another so that each can be pre-computed independently.On a large scale language modeling task, this architecture achieves a 10x speedup at runtime and a significant reduction in perplexity when compared to a standard multilayer network. Jacob Devlin, Chris Quirk, Arul Menezes |
EMNLP | 1 |
| 2015 | A Survey of Current Datasets for Vision and Language ResearchabstractFrancis Ferraro, Nasrin Mostafazadeh, Ting-Hao Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015. Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao 'Kenneth' Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, Margaret Mitchell |
EMNLP | 5 |
| 2014 | Fast and Robust Neural Network Joint Models for Statistical Machine TranslationabstractJacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard Schwartz, John Makhoul. Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2014. Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard M. Schwartz, John Makhoul |
ACL (1) | 1 |
| 2013 | Factored Soft Source Syntactic Constraints for Hierarchical Machine TranslationabstractThis paper describes a factored approach to incorporating soft source syntactic constraints into a hierarchical phrase-based translation system.In contrast to traditional approaches that directly introduce syntactic constraints to translation rules by explicitly decorating them with syntactic annotations, which often exacerbate the data sparsity problem and cause other problems, our approach keeps translation rules intact and factorizes the use of syntactic constraints through two separate models: 1) a syntax mismatch model that associates each nonterminal of a translation rule with a distribution of tags that is used to measure the degree of syntactic compatibility of the translation rule on source spans; 2) a syntax-based reordering model that predicts whether a pair of sibling constituents in the constituent parse tree of the source sentence should be reordered or not when translated to the target language.The features produced by both models are used as soft constraints to guide the translation process.Experiments on Chinese-English translation show that the proposed approach significantly improves a strong string-to-dependency translation system on multiple evaluation sets. Zhongqiang Huang, Jacob Devlin, Rabih Zbib |
EMNLP | 2 |
| 2013 | BBN TransTalk: Robust multilingual two-way speech-to-speech translation for mobile platforms
Rohit Prasad, Premkumar Natarajan, David Stallard, Shirin Saleem, Shankar Ananthakrishnan, Stavros Tsakalidis, Chia-Lin Kao, Fred Choi, Ralf Meermeier, Mark Rawls, Jacob Devlin, Kriste Krstovski, Aaron Challenner |
Comput. Speech Lang. | 11 |
| 2012 | Automatic Tune Set Generation for Machine Translation with Limited Indomain Data
Jinying Chen, Jacob Devlin, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
EAMT | 2 |
| 2012 | Statistical Machine Translation as a Language Model for Handwriting RecognitionabstractWhen performing handwriting recognition on natural language text, the use of a word-level language model (LM) is known to significantly improve recognition accuracy. The most common type of language model, the n-gram model, decomposes sentences into short, overlapping chunks. In this paper, we propose a new type of language model which we use in addition to the standard n-gram LM. Our new model uses the likelihood score from a statistical machine translation system as a reranking feature. In general terms, we automatically translate each OCR hypothesis into another language, and then create a feature score based on how "difficult" it was to perform the translation. Intuitively, the difficulty of translation correlates with how well-formed the input sentence is. In an Arabic handwriting recognition task, we were able to obtain an 0.4% absolute improvement to word error rate (WER) on top of a powerful 5-gram LM. Jacob Devlin, Matin Kamali, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICFHR | 1 |
| 2012 | Document recognition and translation system for unconstrained Arabic documents
Huaigu Cao, Jinying Chen, Jacob Devlin, Rohit Prasad, Premkumar Natarajan |
ICPR | 3 |
| 2012 | Trait-Based Hypothesis Selection For Machine Translation
Jacob Devlin, Spyridon Matsoukas |
HLT-NAACL | 1 |
| 2012 | Machine Translation of Arabic Dialects
Rabih Zbib, Erika Malchiodi, Jacob Devlin, David Stallard, Spyridon Matsoukas, Richard M. Schwartz, John Makhoul, Omar Zaidan, Chris Callison-Burch |
HLT-NAACL | 3 |
| 2011 | System Combination Using Discriminative Cross-Adaptation
Jacob Devlin, Antti-Veikko I. Rosti, Sankaranarayanan Ananthakrishnan, Spyridon Matsoukas |
IJCNLP | 1 |