VLDB 2026 Research / reviewers in the wild / expert
Tao Lei 0001
dblp:91/8024-1
· DBLP profile ↗
32ranked-venue papers
9as first author
10since 2021 · last 2025
0000-0002-3010-8085ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 9 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice RoutingabstractDiffusion transformers have been widely adopted for text-to-image synthesis. While scaling these models up to billions of parameters shows promise, the effectiveness of scaling beyond current sizes remains underexplored and challenging. By explicitly exploiting the computational heterogeneity of image generations, we develop a new family of Mixture-of-Experts (MoE) models (EC-DIT) for diffusion transformers with expert-choice routing. EC-DIT learns to adaptively optimize the compute allocated to understand the input texts and generate the respective image patches, enabling heterogeneous computation aligned with varying text-image complexities. This heterogeneity provides an efficient way of scaling EC-DIT up to 97 billion parameters and achieving significant improvements in training convergence, text-to-image alignment, and overall generation quality over dense models and conventional MoE models. Through extensive ablations, we show that EC-DIT demonstrates superior scalability and adaptive compute allocation by recognizing varying textual importance through end-to-end training. Notably, in text-to-image alignment evaluation, our largest models achieve a state-of-the-art GenEval score of 71.68% and still maintain competitive inference speed with intuitive interpretability. Tao Lei 0001, Bowen Zhang 0002, Yanghao Li, Haoshuo Huang, Ruoming Pang, Bo Dai 0001, Nan Du 0002 |
ICLR | 2 |
| 2025 | Instruction-Following Pruning for Large Language ModelsabstractWith the rapid scaling of large language models (LLMs), structured pruning has become a widely used technique to learn efficient, smaller models from larger ones, delivering superior performance compared to training similarly sized models from scratch. In this paper, we move beyond the traditional static pruning approach of determining a fixed pruning mask for a model, and propose a dynamic approach to structured pruning. In our method, the pruning mask is input-dependent and adapts dynamically based on the information described in a user instruction. Our approach, termed "instruction-following pruning'', introduces a sparse mask predictor that takes the user instruction as input and dynamically selects the most relevant model parameters for the given task. To identify and activate effective parameters, we jointly optimize the sparse mask predictor and the LLM, leveraging both instruction-following data and the pre-training corpus. Experimental results demonstrate the effectiveness of our approach on a wide range of evaluation benchmarks. For example, our 3B activated model improves over the 3B dense model by 5-8 points of absolute margin on domains such as math and coding, and rivals the performance of a 9B model. Bairu Hou, Guoli Yin, Nan Du 0002, Ruoming Pang, Shiyu Chang, Tao Lei 0001 |
ICML | 9 |
| 2024 | MM1: Methods, Analysis and Insights from Multimodal LLM Pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang 0002, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Anton Belyi, Haotian Zhang 0005, Karanjeet Singh 0003, Doug Kang, Hongyu Hè, Max Schwarzer, Tom Gunter, Xiang Kong, Aonan Zhang, Nan Du 0002, Tao Lei 0001, Sam Wiseman, Mark Lee 0003, Ruoming Pang, Peter Grasch, Alexander Toshev, Yinfei Yang |
ECCV (29) | 22 |
| 2023 | CoLT5: Faster Long-Range Transformers with Conditional ComputationabstractJoshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontanon, Siddhartha Brahma, Yury Zemlyanskiy, David Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, Yun-Hsuan Sung, Sumit Sanghai. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Joshua Ainslie, Tao Lei 0001, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskiy, David C. Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, Yun-Hsuan Sung, Sumit Sanghai |
EMNLP | 2 |
| 2023 | Rethinking the Role of Token Retrieval in Multi-Vector RetrievalabstractMulti-vector retrieval models such as ColBERT [Khattab et al., 2020] allow token-level interactions between queries and documents, and hence achieve state of the art on many information retrieval benchmarks. However, their non-linear scoring function cannot be scaled to millions of documents, necessitating a three-stage process for inference: retrieving initial candidates via token retrieval, accessing all token vectors, and scoring the initial candidate documents. The non-linear scoring function is applied over all token vectors of each candidate document, making the inference process complicated and slow. In this paper, we aim to simplify the multi-vector retrieval by rethinking the role of token retrieval. We present XTR, ConteXtualized Token Retriever, which introduces a simple, yet novel, objective function that encourages the model to retrieve the most important document tokens first. The improvement to token retrieval allows XTR to rank candidates only using the retrieved tokens rather than all tokens in the document, and enables a newly designed scoring stage that is two-to-three orders of magnitude cheaper than that of ColBERT. On the popular BEIR benchmark, XTR advances the state-of-the-art by 2.8 nDCG@10 without any distillation. Detailed analysis confirms our decision to revisit the token retrieval stage, as XTR demonstrates much better recall of the token retrieval stage compared to ColBERT. Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 0001, Iftekhar Naim, Ming-Wei Chang, Vincent Y. Zhao |
NeurIPS | 4 |
| 2023 | Conditional Adapters: Parameter-efficient Transfer Learning with Fast InferenceabstractWe propose Conditional Adapter (CoDA), a parameter-efficient transfer learning method that also improves inference efficiency. CoDA generalizes beyond standard adapter approaches to enable a new way of balancing speed and accuracy using conditional computation.
Starting with an existing dense pretrained model, CoDA adds sparse activation together with a small number of new parameters and a light-weight training phase.
Our experiments demonstrate that the CoDA approach provides an unexpectedly efficient way to transfer knowledge.
Across a variety of language, vision, and speech tasks, CoDA achieves a 2x to 8x inference speed-up compared to the state-of-the-art Adapter approaches with moderate to no accuracy loss and the same parameter efficiency. Tao Lei 0001, Junwen Bai, Siddhartha Brahma, Joshua Ainslie, Kenton Lee, Yanqi Zhou, Nan Du 0002, Vincent Y. Zhao, Yuexin Wu, Bo Li 0028, Yu Zhang 0033, Ming-Wei Chang |
NeurIPS | 1 |
| 2022 | SRU++: Pioneering Fast Recurrence with Attention for Speech RecognitionabstractThe Transformer architecture has been well adopted as a dominant architecture in most sequence transduction tasks including automatic speech recognition (ASR), since its attention mechanism excels in capturing long-range dependencies. While models built solely upon attention can be better parallelized than regular RNN, a novel network architecture, SRU++, was recently proposed. By combining the fast recurrence and attention mechanism, SRU++ exhibits strong capability in sequence modeling and achieves near-state-of-the-art results in various language modeling and machine translation tasks with improved compute efficiency. In this work, we present the advantages of applying SRU++ in ASR tasks by comparing with Conformer across multiple ASR benchmarks and study how the benefits can be generalized to long-form speech inputs. On the popular LibriSpeech benchmark, our SRU++ model achieves 2.0% / 4.7% WER on test-clean / test-other, showing competitive performances compared with the state-of-the-art Conformer encoder under the same set-up. Specifically, SRU++ can surpass Conformer on long-form speech input with a large margin, based on our analysis. Tao Lei 0001, Kwangyoun Kim, Kyu Jeong Han, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2022 | Mixture-of-Experts with Expert Choice RoutingabstractSparsely-activated Mixture-of-experts (MoE) models allow the number of parameters to greatly increase while keeping the amount of computation for a given token or a given sample unchanged. However, a poor expert routing strategy (e.g. one resulting in load imbalance) can cause certain experts to be under-trained, leading to an expert being under or over-specialized. Prior work allocates a fixed number of experts to each token using a top-k function regardless of the relative importance of different tokens. To address this, we propose a heterogeneous mixture-of-experts employing an expert choice method. Instead of letting tokens select the top-k experts, we have experts selecting the top-k tokens. As a result, each token can be routed to a variable number of experts and each expert can have a fixed bucket size. We systematically study pre-training speedups using the same computational resources of the Switch Transformer top-1 and GShard top-2 gating of prior work and find that our method improves training convergence time by more than 2×. For the same computational cost, our method demonstrates higher performance in fine-tuning 11 selected tasks in the GLUE and SuperGLUE benchmarks. For a smaller activation cost, our method outperforms the T5 dense model in 7 out of the 11 tasks. Yanqi Zhou, Tao Lei 0001, Hanxiao Liu, Nan Du 0002, Yanping Huang, Vincent Y. Zhao, Andrew M. Dai, Quoc V. Le, James Laudon |
NeurIPS | 2 |
| 2021 | Nutri-bullets: Summarizing Health Studies by Composing SegmentsabstractWe introduce Nutri-bullets, a multi-document summarization task for health and nutrition. First, we present two datasets of food and health summaries from multiple scientific studies. Furthermore, we propose a novel extract-compose model to solve the problem in the regime of limited parallel data. We explicitly select key spans from several abstracts using a policy network, followed by composing the selected spans to present a summary via a task specific language model. Compared to state-of-the-art methods, our approach leads to more faithful, relevant and diverse summarization -- properties imperative to this application. For instance, on the BreastCancer dataset our approach gets a more than 50% improvement on relevance and faithfulness. Darsh J. Shah, Lili Yu, Tao Lei 0001, Regina Barzilay |
AAAI | 3 |
| 2021 | Nutri-bullets Hybrid: Consensual Multi-document SummarizationabstractWe present a method for generating comparative summaries that highlights similarities and contradictions in input documents.The key challenge in creating such summaries is the lack of large parallel training data required for training typical summarization systems.To this end, we introduce a hybrid generation approach inspired by traditional concept-to-text systems.To enable accurate comparison between different sources, the model first learns to extract pertinent relations from input documents.The content planning component uses deterministic operators to aggregate these relations after identifying a subset for inclusion into a summary.The surface realization component lexicalizes this information using a text-infilling language model.By separately modeling content selection and realization, we can effectively train them with limited annotations.We implemented and tested the model in the domain of nutrition and health -rife with inconsistencies.Compared to conventional methods, our framework leads to more faithful, relevant and aggregation-sensitive summarization -while being equally fluent. 1 Darsh J. Shah, Lili Yu, Tao Lei 0001, Regina Barzilay |
NAACL-HLT | 3 |
| 2020 | Rationalizing Text Matching: Learning Sparse Alignments via Optimal TransportabstractSelecting input features of top relevance has become a popular method for building selfexplaining models.In this work, we extend this selective rationalization approach to text matching, where the goal is to jointly select and align text pieces, such as tokens or sentences, as a justification for the downstream prediction.Our approach employs optimal transport (OT) to find a minimal cost alignment between the inputs.However, directly applying OT often produces dense and therefore uninterpretable alignments.To overcome this limitation, we introduce novel constrained variants of the OT problem that result in highly sparse alignments with controllable sparsity.Our model is end-to-end differentiable using the Sinkhorn algorithm for OT and can be trained without any alignment annotations.We evaluate our model on the Stack-Exchange, MultiNews, e-SNLI, and MultiRC datasets.Our model achieves very sparse rationale selections with high fidelity while preserving prediction accuracy compared to strong attention baseline models.† * Denotes equal contribution.† Our code is publicly available at https://github. com/asappresearch/rationale-alignment.Can I find duplicate songs with different names?I have so many duplicate songs but they have different names.Is there an application I can use to find and delete the duplicates?How to find (and delete) duplicate files?I have a largish music collection and there are some duplicates in there.Is there any way to find duplicate files.At a minimum by doing a hash and seeing if two files have the same hash.… I'm happy using the command line if that is the easiest way. Kyle Swanson, Lili Yu, Tao Lei 0001 |
ACL | 3 |
| 2020 | Interactive Classification by Asking Informative QuestionsabstractWe study the potential for interaction in natural language classification.We add a limited form of interaction for intent classification, where users provide an initial query using natural language, and the system asks for additional information using binary or multichoice questions.At each turn, our system decides between asking the most informative question or making the final classification prediction.The simplicity of the model allows for bootstrapping of the system without interaction data, instead relying on simple crowdsourcing tasks.We evaluate our approach on two domains, showing the benefit of interaction and the advantage of learning to balance between asking additional questions and making the final prediction.What is the bill length of the bird: shorter, similar, or longer than head?Shorter than head.Is the bird underpart orange?Yes.The identified bird is: American Redstart FAQ Suggestion What data limits apply when roaming internationally?American Crow Bobolink … American Redstart How do I sign up for Sprint Global Roaming? . . .How do I purchase a High Speed Data Roaming Pass?Bird Identification Travel out of country.Do you need to activate global roaming service?Yes.Do you want high speed data roaming?No. Lili Yu, Howard Chen 0003, Sida I. Wang, Tao Lei 0001, Yoav Artzi |
ACL | 4 |
| 2020 | Autoregressive Knowledge Distillation through Imitation LearningabstractThe performance of autoregressive models on natural language generation tasks has dramatically improved due to the adoption of deep, self-attentive architectures.However, these gains have come at the cost of hindering inference speed, making state-of-the-art models cumbersome to deploy in real-world, timesensitive settings.We develop a compression technique for autoregressive models that is driven by an imitation learning perspective on knowledge distillation.The algorithm is designed to address the exposure bias problem.On prototypical language generation tasks such as translation and summarization, our method consistently outperforms other distillation algorithms, such as sequence-level knowledge distillation.Student models trained with our method attain 1.4 to 4.8 BLEU/ROUGE points higher than those trained from scratch, while increasing inference speed by up to 14 times in comparison to the teacher model. 1 Alexander Lin, Jeremy Wohlwend, Howard Chen 0003, Tao Lei 0001 |
EMNLP (1) | 4 |
| 2020 | Structured Pruning of Large Language ModelsabstractLarge language models have recently achieved state of the art performance across a wide variety of natural language tasks. Meanwhile, the size of these models and their latency have significantly increased, which makes their usage costly, and raises an interesting question: do language models need to be large? We study this question through the lens of model compression. We present a generic, structured pruning approach by parameterizing each weight matrix using its low-rank factorization, and adaptively removing rank-1 components during training. On language modeling tasks, our structured approach outperforms other unstructured and block-structured pruning baselines at various compression levels, while achieving significant speedups during both training and inference. We also demonstrate that our method can be applied to pruning adaptive word embeddings in large language models, and to pruning the BERT model on several downstream fine-tuning classification benchmarks. Jeremy Wohlwend, Tao Lei 0001 |
EMNLP (1) | 3 |
| 2020 | ASAPP-ASR: Multistream CNN and Self-Attentive SRU for SOTA Speech RecognitionabstractIn this paper we present state-of-the-art (SOTA) performance on the LibriSpeech corpus with two novel neural network architectures, a multistream CNN for acoustic modeling and a selfattentive simple recurrent unit (SRU) for language modeling.In the hybrid ASR framework, the multistream CNN acoustic model processes an input of speech frames in multiple parallel pipelines where each stream has a unique dilation rate for diversity.Trained with the SpecAugment data augmentation method, it achieves relative word error rate (WER) improvements of 4% on test-clean and 14% on test-other.We further improve the performance via N -best rescoring using a 24-layer self-attentive SRU language model, achieving WERs of 1.75% on test-clean and 4.46% on test-other. Joshua Shapiro, Jeremy Wohlwend, Kyu Jeong Han, Tao Lei 0001 |
INTERSPEECH | 5 |
| 2018 | Simple Recurrent Units for Highly Parallelizable RecurrenceabstractCommon recurrent neural architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations.In this work, we propose the Simple Recurrent Unit (SRU), a light recurrent unit that balances model capacity and scalability.SRU is designed to provide expressive recurrence, enable highly parallelized implementation, and comes with careful initialization to facilitate training of deep models.We demonstrate the effectiveness of SRU on multiple NLP tasks.SRU achieves 5-9x speed-up over cuDNN-optimized LSTM on classification and question answering datasets, and delivers stronger results than LSTM and convolutional models.We also obtain an average of 0.7 BLEU improvement over the Transformer model (Vaswani et al., 2017) on translation by incorporating SRU into the architecture.1 Tao Lei 0001, Yu Zhang 0033, Sida I. Wang, Hui Dai, Yoav Artzi |
EMNLP | 1 |
| 2018 | Adversarial Domain Adaptation for Duplicate Question DetectionabstractWe address the problem of detecting duplicate questions in forums, which is an important step towards automating the process of answering new questions.As finding and annotating such potential duplicates manually is very tedious and costly, automatic methods based on machine learning are a viable alternative.However, many forums do not have annotated data, i.e., questions labeled by experts as duplicates, and thus a promising solution is to use domain adaptation from another forum that has such annotations.Here we focus on adversarial domain adaptation, deriving important findings about when it performs well and what properties of the domains are important in this regard.Our experiments with StackExchange data show an average improvement of 5.6% over the best baseline across multiple pairs of domains. Darsh J. Shah, Tao Lei 0001, Alessandro Moschitti, Salvatore Romeo, Preslav Nakov |
EMNLP | 2 |
| 2017 | Deriving Neural Architectures from Sequence and Graph KernelsabstractThe design of neural architectures for structured objects is typically guided by experimental insights rather than a formal process. In this work, we appeal to kernels over combinatorial structures, such as sequences and graphs, to derive appropriate neural operations. We introduce a class of deep recurrent neural operations and formally characterize their associated kernel spaces. Our recurrent modules compare the input to virtual reference objects (cf. filters in CNN) via the kernels. Similar to traditional neural operations, these reference objects are parameterized and directly optimized in end-to-end training. We empirically evaluate the proposed class of neural architectures on standard applications such as language modeling and molecular graph regression, achieving state-of-the-art results across these applications. Tao Lei 0001, Wengong Jin, Regina Barzilay, Tommi S. Jaakkola |
ICML | 1 |
| 2017 | Style Transfer from Non-Parallel Text by Cross-AlignmentabstractThis paper focuses on style transfer on the basis of non-parallel text. This is an instance of a broad family of problems including machine translation, decipherment, and sentiment modification. The key challenge is to separate the content from other aspects such as style. We assume a shared latent content distribution across different text corpora, and propose a method that leverages refined alignment of latent representations to perform style transfer. The transferred sentences from one style should match example sentences from the other style as a population. We demonstrate the effectiveness of this cross-alignment method on three tasks: sentiment modification, decipherment of word substitution ciphers, and recovery of word order. Tianxiao Shen, Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola |
NIPS | 2 |
| 2016 | Learning to refine text based recommendations
Youyang Gu, Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola |
EMNLP | 2 |
| 2016 | Rationalizing Neural PredictionsabstractPrediction without justification has limited applicability.As a remedy, we learn to extract pieces of input text as justifications -rationales -that are tailored to be short and coherent, yet sufficient for making the same prediction.Our approach combines two modular components, generator and encoder, which are trained to operate well together.The generator specifies a distribution over text fragments as candidate rationales and these are passed through the encoder for prediction.Rationales are never given during training.Instead, the model is regularized by desiderata for rationales.We evaluate the approach on multi-aspect sentiment analysis against manually annotated test cases.Our approach outperforms attention-based baseline by a significant margin.We also successfully illustrate the method on the question retrieval task. 1 Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola |
EMNLP | 1 |
| 2016 | Semi-supervised Question Retrieval with Gated ConvolutionsabstractTao Lei, Hrishikesh Joshi, Regina Barzilay, Tommi Jaakkola, Kateryna Tymoshenko, Alessandro Moschitti, Lluís Màrquez. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Tao Lei 0001, Hrishikesh Joshi, Regina Barzilay, Tommi S. Jaakkola, Kateryna Tymoshenko, Alessandro Moschitti, Lluís Màrquez |
HLT-NAACL | 1 |
| 2016 | Making Dependency Labeling Simple, Fast and AccurateabstractThis work addresses the task of dependency labeling-assigning labels to an (unlabeled) dependency tree.We employ and extend a feature representation learning approach, optimizing it for both high speed and accuracy.We apply our labeling model on top of state-of-the-art parsers and evaluate its performance on standard benchmarks including the CoNLL-2009 and the English PTB datasets.Our model processes over 1,700 English sentences per second, which is 30 times faster than the sparse-feature method.It improves labeling accuracy over the outputs of top parsers, achieving the best LAS on 5 out of 7 datasets 1 . Tianxiao Shen, Tao Lei 0001, Regina Barzilay |
HLT-NAACL | 2 |
| 2015 | Molding CNNs for text: non-linear, non-consecutive convolutionsabstractThe success of deep learning often derives from well-chosen operational building blocks.In this work, we revise the temporal convolution operation in CNNs to better adapt it to text processing.Instead of concatenating word representations, we appeal to tensor algebra and use low-rank n-gram tensors to directly exploit interactions between words already at the convolution stage.Moreover, we extend the n-gram convolution to non-consecutive words to recognize patterns with intervening words.Through a combination of lowrank tensors, and pattern weighting, we can efficiently evaluate the resulting convolution operation via dynamic programming.We test the resulting architecture on standard sentiment classification and news categorization tasks.Our model achieves state-of-the-art performance both in terms of accuracy and training speed.For instance, we obtain 51.2% accuracy on the fine-grained sentiment classification task. 1 Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola |
EMNLP | 1 |
| 2015 | High-Order Low-Rank Tensors for Semantic Role LabelingabstractTao Lei, Yuan Zhang, Lluís Màrquez, Alessandro Moschitti, Regina Barzilay. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Tao Lei 0001, Yuan Zhang 0001, Lluís Màrquez, Alessandro Moschitti, Regina Barzilay |
HLT-NAACL | 1 |
| 2015 | Erratum: "Exploring Compositional Architectures and Word Vector Representations for Prepositional Phrase Attachment"abstractCorrection for the list of authors in the reference (Seddah et al., 2013). Yonatan Belinkov, Tao Lei 0001, Regina Barzilay, Amir Globerson |
Trans. Assoc. Comput. Linguistics | 2 |
| 2014 | Low-Rank Tensors for Scoring Dependency StructuresabstractAccurate scoring of syntactic structures such as head-modifier arcs in dependency parsing typically requires rich, highdimensional feature representations.A small subset of such features is often selected manually.This is problematic when features lack clear linguistic meaning as in embeddings or when the information is blended across features.In this paper, we use tensors to map high-dimensional feature vectors into low dimensional representations.We explicitly maintain the parameters as a low-rank tensor to obtain low dimensional representations of words in their syntactic roles, and to leverage modularity in the tensor for easy training with online algorithms.Our parser consistently outperforms the Turbo and MST parsers across 14 different languages.We also obtain the best published UAS results on 5 languages.1 Tao Lei 0001, Yuan Zhang 0001, Regina Barzilay, Tommi S. Jaakkola |
ACL (1) | 1 |
| 2014 | Steps to Excellence: Simple Inference with Refined Scoring of Dependency TreesabstractMuch of the recent work on depen-dency parsing has been focused on solv-ing inherent combinatorial problems as-sociated with rich scoring functions. In contrast, we demonstrate that highly ex-pressive scoring functions can be used with substantially simpler inference pro-cedures. Specifically, we introduce a sampling-based parser that can easily han-dle arbitrary global features. Inspired by SampleRank, we learn to take guided stochastic steps towards a high scoring parse. We introduce two samplers for traversing the space of trees, Gibbs and Metropolis-Hastings with Random Walk. The model outperforms state-of-the-art re-sults when evaluated on 14 languages of non-projective CoNLL datasets. Our sampling-based approach naturally ex-tends to joint prediction scenarios, such as joint parsing and POS correction. The resulting method outperforms the best re-ported results on the CATiB dataset, ap-proaching performance of parsing with gold tags.1 1 Yuan Zhang 0001, Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola, Amir Globerson |
ACL (1) | 2 |
| 2014 | Greed is Good if Randomized: New Inference for Dependency ParsingabstractDependency parsing with high-order features results in a provably hard decoding problem.A lot of work has gone into developing powerful optimization methods for solving these combinatorial problems.In contrast, we explore, analyze, and demonstrate that a substantially simpler randomized greedy inference algorithm already suffices for near optimal parsing: a) we analytically quantify the number of local optima that the greedy method has to overcome in the context of first-order parsing; b) we show that, as a decoding algorithm, the greedy method surpasses dual decomposition in second-order parsing; c) we empirically demonstrate that our approach with up to third-order and global features outperforms the state-of-the-art dual decomposition and MCMC sampling methods when evaluated on 14 languages of non-projective CoNLL datasets.1 Yuan Zhang 0001, Tao Lei 0001, Regina Barzilay, Tommi S. Jaakkola |
EMNLP | 2 |
| 2014 | Exploring Compositional Architectures and Word Vector Representations for Prepositional Phrase AttachmentabstractPrepositional phrase (PP) attachment disambiguation is a known challenge in syntactic parsing. The lexical sparsity associated with PP attachments motivates research in word representations that can capture pertinent syntactic and semantic features of the word. One promising solution is to use word vectors induced from large amounts of raw text. However, state-of-the-art systems that employ such representations yield modest gains in PP attachment accuracy. In this paper, we show that word vector representations can yield significant PP attachment performance gains. This is achieved via a non-linear architecture that is discriminatively trained to maximize PP attachment accuracy. The architecture is initialized with word vectors trained from unlabeled data, and relearns those to maximize attachment accuracy. We obtain additional performance gains with alternative representations such as dependency-based word vectors. When tested on both English and Arabic datasets, our method outperforms both a strong SVM classifier and state-of-the-art parsers. For instance, we achieve 82.6% PP attachment accuracy on Arabic, while the Turbo and Charniak self-trained parsers obtain 76.7% and 80.8% respectively. Yonatan Belinkov, Tao Lei 0001, Regina Barzilay, Amir Globerson |
Trans. Assoc. Comput. Linguistics | 2 |
| 2013 | From Natural Language Specifications to Program Input Parsers
Tao Lei 0001, Fan Long, Regina Barzilay, Martin C. Rinard |
ACL (1) | 1 |
| 2012 | Learning High-Level Planning from Text
S. R. K. Branavan, Nate Kushman, Tao Lei 0001, Regina Barzilay |
ACL (1) | 3 |