Ramesh Nallapati

dblp:59/4797 · DBLP profile ↗
← Back
45ranked-venue papers
12as first author
16since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 9 first-author · 15 since 2021Databases, data management, data science and information retrieval · 12 · 8 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-author
YearPublicationVenuePosition
2024 CoCoMIC: Code Completion by Jointly Modeling In-file and Cross-file Context
abstract
While pre-trained language models (LM) for code have achieved great success in code completion, they generate code conditioned only on the contents within the file, i.e., in-file context, but ignore the rich semantics in other files within the same project, i.e., project-level cross-file context, a critical source of information that is especially useful in modern modular software development. Such overlooking constrains code LMs’ capacity in code completion, leading to unexpected behaviors such as generating hallucinated class member functions or function calls with unexpected arguments. In this work, we propose CoCoMIC, a novel framework that jointly learns the in-file and cross-file context on top of code LMs. To empower CoCoMIC, we develop CCFinder, a static-analysis-based tool that locates and retrieves the most relevant project-level cross-file context for code completion. CoCoMIC successfully improves the existing code LM with a 33.94% relative increase in exact match and 28.69% in identifier matching for code completion when the cross-file context is provided. Finally, we perform a series of ablation studies and share valuable insights for future research on integrating cross-file context into code LMs.
Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang
LREC/COLING5
2024 Code Representation Learning at Scale
abstract
Recent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size.
Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang
ICLR5
2024 Bifurcated Attention for Single-Context Large-Batch Sampling
abstract
In our study, we present bifurcated attention, a method developed for language model inference in single-context batch sampling contexts. This approach aims to reduce redundant memory IO costs, a significant factor in latency for high batch sizes and long context lengths. Bifurcated attention achieves this by dividing the attention mechanism during incremental decoding into two distinct GEMM operations, focusing on the KV cache from prefill and the decoding process. This method ensures precise computation and maintains the usual computational load (FLOPs) of standard attention mechanisms, but with reduced memory IO. Bifurcated attention is also compatible with multi-query attention mechanism known for reduced memory IO for KV cache, further enabling higher batch size and context length. The resulting efficiency leads to lower latency, improving suitability for real-time applications, e.g., enabling massively-parallel answer generation without substantially increasing latency, enhancing performance when integrated with post-processing techniques such as reranking.
Ben Athiwaratkun, Sujan K. Gonugondla, Sanjay Krishna Gouda, Haifeng Qian, Hantian Ding, Qing Sun 0013, Jun Wang 0022, Jiacheng Guo, Liangfu Chen, Parminder Bhatia, Ramesh Nallapati, Sudipta Sengupta, Bing Xiang
ICML11
2023 ReCode: Robustness Evaluation of Code Generation Models
abstract
Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shiqi Wang 0002, Haifeng Qian, Chenghao Yang 0001, Zijian Wang 0002, Mingyue Shang, Samson Tan, Baishakhi Ray, Parminder Bhatia, Ramesh Nallapati, Murali Krishna Ramanathan, Dan Roth 0001, Bing Xiang
ACL (1)11
2023 Multitask Pretraining with Structured Knowledge for Text-to-SQL Generation
abstract
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Yang Li, Ming Tan, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Robert Giaquinto, Dejiao Zhang, Benjamin Kleiner, Parminder Bhatia, Ramesh Nallapati, Xiaofei Ma 0001
ACL (1)7
2023 ContraCLM: Contrastive Learning For Causal Language Model
abstract
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, Bing Xiang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang 0002, Feng Nan, Xiaopeng Li 0002, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma 0001, Bing Xiang
ACL (1)8
2023 Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati
ICLR20
2023 CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion
abstract
Code completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers.
Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang
NeurIPS8
2021 Answering Ambiguous Questions through Generative Evidence Fusion and Round-Trip Prediction
abstract
Yifan Gao, Henghui Zhu, Patrick Ng, Cicero Nogueira dos Santos, Zhiguo Wang, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yifan Gao 0001, Henghui Zhu, Patrick Ng, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Feng Nan, Dejiao Zhang, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)8
2021 Improving Factual Consistency of Abstractive Summarization via Question Answering
abstract
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, Bing Xiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Feng Nan, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathy McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang 0006, Andrew O. Arnold, Bing Xiang
ACL/IJCNLP (1)6
2021 Knowledge Graph Representation via Hierarchical Hyperbolic Neural Graph Embedding
abstract
Knowledge graph enhanced information retrieval systems have attracted considerable attention due to their ability to improve performance and provide additional explainability. As the knowledge graphs usually include fruitful facts, they are also good sources of side information. However, recent studies have shown that the usefulness of knowledge graphs depends highly on their representation, e.g., the embeddings of entities and relations. Embedding entities and relations in low-dimensional space is a successful knowledge graph representation solution. Most of the works lie in modeling symmetry/asymmetry/composition/inversion relations but pay less attention to the hierarchical relations. Recent studies have observed the fact that there exist rich semantic hierarchical relations in knowledge graphs such as Freebase (entities are connected in a taxonomic hierarchy) and WordNet (entities are synsets linked together in a hierarchy).To address the above problems, we propose Hierarchical Hyperbolic Neural Graph Embedding (H2E), a new knowledge graph representation approach, which is able to better preserve hierarchical relations. Specifically, the entities/relations representations are learned in a hyperbolic polar embedding space. In a hyperbolic polar embedding space, the entity and relation are modeled as a dual-embedding with modulus embedding part and phase embedding part, enabling the explicitly modeling of two types of hierarchies: inter-level hierarchy and intra-level hierarchy. As the polar embedding is defined i n hyperbolic space, the ability of modeling and inferring hierarchical relations are mutual enhanced. In addition, by noticing the existence of the rich relational context, we propose an attentional neural context aggregation to adaptively integrate the relational context for further enhancing the ability to preserve the hierarchical relations. The empirical study on three benchmark datasets for the link prediction task demonstrates significant performance gains compared to some existing state-of-the-art methods and verifies the effectiveness of the proposed method on hierarchical relations.
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Philip S. Yu
IEEE BigData5
2021 Entity-level Factual Consistency of Abstractive Text Summarization
abstract
Feng Nan, Ramesh Nallapati, Zhiguo Wang, Cicero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathleen McKeown, Bing Xiang. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Feng Nan, Ramesh Nallapati, Zhiguo Wang 0006, Cícero Nogueira dos Santos, Henghui Zhu, Dejiao Zhang, Kathy McKeown, Bing Xiang
EACL2
2021 Retrieval, Re-ranking and Multi-task Learning for Knowledge-Base Question Answering
abstract
Question answering over knowledge bases (KBQA) usually involves three sub-tasks, namely topic entity detection, entity linking and relation detection.Due to the large number of entities and relations inside knowledge bases (KB), previous work usually utilized sophisticated rules to narrow down the search space and managed only a subset of KBs in memory.In this work, we leverage a retrieveand-rerank framework to access KBs via traditional information retrieval (IR) method, and re-rank retrieved candidates with more powerful neural networks such as the pre-trained BERT model.Considering the fact that directly assigning a different BERT model for each sub-task may incur prohibitive costs, we propose to share a BERT encoder across all three sub-tasks and define task-specific layers on top of the shared layer.The unified model is then trained under a multi-task learning framework.Experiments show that: (1) Our IRbased retrieval method is able to collect highquality candidates efficiently, thus enables our method adapt to large-scale KBs easily; (2) the BERT model improves the accuracy across all three sub-tasks; and (3) benefiting from multitask learning, the unified model obtains further improvements with only 1/3 of the original parameters.Our final model achieves competitive results on the SimpleQuestions dataset and superior performance on the FreebaseQA dataset.
Zhiguo Wang 0006, Patrick Ng, Ramesh Nallapati, Bing Xiang
EACL3
2021 Pairwise Supervised Contrastive Learning of Sentence Representations
abstract
Many recent successes in sentence representation learning have been achieved by simply fine-tuning on the Natural Language Inference (NLI) datasets with triplet loss or siamese loss.Nevertheless, they share a common weakness: sentences in a contradiction pair are not necessarily from different semantic categories.Therefore, optimizing the semantic entailment and contradiction reasoning objective alone is inadequate to capture the high-level semantic structure.The drawback is compounded by the fact that the vanilla siamese or triplet losses only learn from individual sentence pairs or triplets, which often suffer from bad local optima.In this paper, we propose PairSupCon, an instance discrimination based approach aiming to bridge semantic entailment and contradiction understanding with high-level categorical concept encoding.We evaluate PairSupCon on various downstream tasks that involve understanding sentence semantics at different granularities.We outperform the previous state-of-theart method with 10%-13% averaged improvement on eight clustering tasks, and 5%-6% averaged improvement on seven semantic textual similarity (STS) tasks.
Dejiao Zhang, Shang-Wen Li 0001, Wei Xiao 0001, Henghui Zhu, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
EMNLP (1)5
2021 Supporting Clustering with Contrastive Learning
abstract
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li, Henghui Zhu, Kathleen McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Dejiao Zhang, Feng Nan, Xiaokai Wei, Shang-Wen Li 0001, Henghui Zhu, Kathy McKeown, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang
NAACL-HLT7
2021 Mixed-Curvature Multi-Relational Graph Neural Network for Knowledge Graph Completion
abstract
Knowledge graphs (KGs) have gradually become valuable assets for many AI applications. In a KG, a node denotes an entity, and an edge (or link) denotes a relationship between the entities represented by the nodes. Knowledge graph completion infers and predicts missing edges in a KG automatically. Knowledge graph embeddings have shed light on addressing this task. Recent research embeds KGs in hyperbolic (negatively curved) space instead of conventional Euclidean (zero curved) space and is effective in capturing hierarchical structures. However, as multi-relational graphs, KGs are not structured uniformly and display intrinsic heterogeneous structures. They usually contain rich types of structures, such as hierarchical and cyclic typed structures. Embedding KGs in single-curvature space, such as Euclidean or hyperbolic space, overlooks the intrinsic heterogeneous structures of KGs, and therefore cannot accurately capture their structures. To address this issue, we propose Mixed-Curvature Multi-Relational Graph Neural Network (M2GNN), a generic approach that embeds multi-relational KGs in a mixed-curvature space for knowledge graph completion. Specifically, we define and construct a mixed-curvature space through a product manifold combining multiple single-curvature spaces (e.g., spherical, hyperbolic, or Euclidean) with the purpose of modeling a variety of structures. However, constructing a mixed-curvature space typically requires manually defining the fixed curvatures, which needs domain knowledge and additional data analysis. Improperly defined curvature space also cannot capture the structures of KGs accurately. To address this problem, we set mixed-curvatures as trainable parameters to better capture the underlying structures of the KGs. Furthermore, we propose a Graph Neural Updater by leveraging the heterogeneous relational context in mixed-curvature space to improve the quality of the embedding. Experiments on three KG datasets demonstrate that the proposed M2GNN can outperform its single geometry counterpart as well as state-of-the-art embedding methods on the KG completion task.
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu, Isabel F. Cruz
WWW5
2020 Who Did They Respond to? Conversation Structure Modeling Using Masked Hierarchical Transformer
abstract
Conversation structure is useful for both understanding the nature of conversation dynamics and for providing features for many downstream applications such as summarization of conversations. In this work, we define the problem of conversation structure modeling as identifying the parent utterance(s) to which each utterance in the conversation responds to. Previous work usually took a pair of utterances to decide whether one utterance is the parent of the other. We believe the entire ancestral history is a very important information source to make accurate prediction. Therefore, we design a novel masking mechanism to guide the ancestor flow, and leverage the transformer model to aggregate all ancestors to predict parent utterances. Our experiments are performed on the Reddit dataset (Zhang, Culbertson, and Paritosh 2017) and the Ubuntu IRC dataset (Kummerfeld et al. 2019). In addition, we also report experiments on a new larger corpus from the Reddit platform and release this dataset. We show that the proposed model, that takes into account the ancestral history of the conversation, significantly outperforms several strong baselines including the BERT model on all datasets.
Henghui Zhu, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
AAAI4
2020 Template-Based Question Generation from Retrieved Sentences for Improved Unsupervised Question Answering
abstract
Question Answering (QA) is in increasing demand as the amount of information available online and the desire for quick access to this content grows.A common approach to QA has been to fine-tune a pretrained language model on a task-specific labeled dataset.This paradigm, however, relies on scarce, and costly to obtain, large-scale human-labeled data.We propose an unsupervised approach to training QA models with generated pseudotraining data.We show that generating questions for QA training by applying a simple template on a related, retrieved sentence rather than the original context sentence improves downstream QA performance by allowing the model to learn more complex context-question relationships.Training a QA model on this data gives a relative improvement over a previous unsupervised model in F1 score on the SQuAD dataset by about 14%, and 20% when the answer is a named entity, achieving stateof-the-art performance on SQuAD for unsupervised QA.
Alexander R. Fabbri, Patrick Ng, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
ACL4
2020 Beyond [CLS] through Ranking by Generation
abstract
Generative models for Information Retrieval, where ranking of documents is viewed as the task of generating a query from a document's language model, were very successful in various IR tasks in the past.However, with the advent of modern deep neural networks, attention has shifted to discriminative ranking functions that model the semantic similarity of documents and queries instead.Recently, deep generative models such as GPT2 and BART have been shown to be excellent text generators, but their effectiveness as rankers have not been demonstrated yet.In this work, we revisit the generative framework for information retrieval and show that our generative approaches are as effective as state-of-the-art semantic similarity-based discriminative models for the answer selection task.Additionally, we demonstrate the effectiveness of unlikelihood losses for IR.
Cícero Nogueira dos Santos, Xiaofei Ma 0001, Ramesh Nallapati, Zhiheng Huang, Bing Xiang
EMNLP (1)3
2020 End-to-End Synthetic Data Generation for Domain Adaptation of Question Answering Systems
abstract
Siamak Shakeri, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang, Ramesh Nallapati, Bing Xiang. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Siamak Shakeri, Cícero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Feng Nan, Zhiguo Wang 0006, Ramesh Nallapati, Bing Xiang
EMNLP (1)7
2020 H2KGAT: Hierarchical Hyperbolic Knowledge Graph Attention Network
Shen Wang 0005, Xiaokai Wei, Cícero Nogueira dos Santos, Zhiguo Wang 0006, Ramesh Nallapati, Andrew O. Arnold, Bing Xiang, Philip S. Yu
EMNLP (1)5
2020 Elastic Machine Learning Algorithms in Amazon SageMaker
abstract
There is a large body of research on scalable machine learning (ML). Nevertheless, training ML models on large, continuously evolving datasets is still a difficult and costly undertaking for many companies and institutions. We discuss such challenges and derive requirements for an industrial-scale ML platform. Next, we describe the computational model behind Amazon SageMaker, which is designed to meet such challenges. SageMaker is an ML platform provided as part of Amazon Web Services (AWS), and supports incremental training, resumable and elastic learning as well as automatic hyperparameter optimization. We detail how to adapt several popular ML algorithms to its computational model. Finally, we present an experimental evaluation on large datasets, comparing SageMaker to several scalable, JVM-based implementations of ML algorithms, which we significantly outperform with regard to computation time and cost.
Edo Liberty, Zohar S. Karnin, Bing Xiang, Laurence Rouesnel, Baris Coskun, Ramesh Nallapati, Julio Delgado, Amir Sadoughi, Yury Astashonok, Piali Das, Can Balioglu, Saswata Chakravarty, Madhav Jha, Philip Gautier, David Arpin, Tim Januschowski, Valentin Flunkert, Yuyang Wang 0001, Jan Gasthaus, Lorenzo Stella, Syama Sundar Rangapuram, David Salinas, Sebastian Schelter, Alexander J. Smola
SIGMOD Conference6
2019 Topic Modeling with Wasserstein Autoencoders
abstract
We propose a novel neural topic model in the Wasserstein autoencoders (WAE) framework.Unlike existing variational autoencoder based models, we directly enforce Dirichlet prior on the latent document-topic vectors.We exploit the structure of the latent space and apply a suitable kernel in minimizing the Maximum Mean Discrepancy (MMD) to perform distribution matching.We discover that MMD performs much better than the Generative Adversarial Network (GAN) in matching high dimensional Dirichlet distribution.We further discover that incorporating randomness in the encoder output during training leads to significantly more coherent topics.To measure the diversity of the produced topics, we propose a simple topic uniqueness metric.Together with the widely used coherence measure NPMI, we offer a more wholistic evaluation of topic quality.Experiments on several real datasets show that our model produces significantly better topics than existing topic models.
Feng Nan, Ramesh Nallapati, Bing Xiang
ACL (1)3
2019 OCGAN: One-Class Novelty Detection Using GANs With Constrained Latent Representations
abstract
We present a novel model called OCGAN for the classical problem of one-class novelty detection, where, given a set of examples from a particular class, the goal is to determine if a query example is from the same class. Our solution is based on learning latent representations of in-class examples using a de-noising auto-encoder network. The key contribution of our work is our proposal to explicitly constrain the latent space to exclusively represent the given class. In order to accomplish this goal, firstly, we force the latent space to have bounded support by introducing a tanh activation in the encoder's output layer. Secondly, using a discriminator in the latent space that is trained adversarially, we ensure that encoded representations of in-class examples resemble uniform random samples drawn from the same bounded space. Thirdly, using a second adversarial discriminator in the input space, we ensure all randomly drawn latent samples generate examples that look real. Finally, we introduce a gradient-descent based sampling technique that explores points in the latent space that generate potential out-of-class examples, which are fed back to the network to further train it to generate in-class examples from those points. The effectiveness of the proposed method is measured across four publicly available datasets using two one-class novelty detection protocols where we achieve state-of-the-art results.
Pramuditha Perera, Ramesh Nallapati, Bing Xiang
CVPR2
2019 Multi-passage BERT: A Globally Normalized BERT Model for Open-domain Question Answering
abstract
Zhiguo Wang, Patrick Ng, Xiaofei Ma, Ramesh Nallapati, Bing Xiang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhiguo Wang 0006, Patrick Ng, Xiaofei Ma 0001, Ramesh Nallapati, Bing Xiang
EMNLP/IJCNLP (1)4
2018 Coherence-Aware Neural Topic Modeling
abstract
Topic models are evaluated based on their ability to describe documents well (i.e.low perplexity) and to produce topics that carry coherent semantic meaning.In topic modeling so far, perplexity is a direct optimization target.However, topic coherence, owing to its challenging computation, is not optimized for and is only evaluated after training.In this work, under a neural variational inference framework, we propose methods to incorporate a topic coherence objective into the training process.We demonstrate that such a coherenceaware topic model exhibits a similar level of perplexity as baseline models but achieves substantially higher topic coherence.
Ramesh Nallapati, Bing Xiang
EMNLP2
2017 SummaRuNNer: A Recurrent Neural Network Based Sequence Model for Extractive Summarization of Documents
abstract
We present SummaRuNNer, a Recurrent Neural Network (RNN) based sequence model for extractive summarization of documents and show that it achieves performance better than or comparable to state-of-the-art. Our model has the additional advantage of being very interpretable, since it allows visualization of its predictions broken up by abstract features such as information content, salience and novelty. Another novel contribution of our work is abstractive training of our extractive model that can train on human generated reference summaries alone, eliminating the need for sentence-level extractive labels.
Ramesh Nallapati, Feifei Zhai
AAAI1
2016 Pointing the Unknown Words
abstract
The problem of rare and unknown words is an important issue that can potentially effect the performance of many NLP systems, including traditional count-based and deep learning models.We propose a novel way to deal with the rare and unseen words for the neural network models using attention.Our model uses two softmax layers in order to predict the next word in conditional language models: one predicts the location of a word in the source sentence, and the other predicts a word in the shortlist vocabulary.At each timestep, the decision of which softmax layer to use is adaptively made by an MLP which is conditioned on the context.We motivate this work from a psychological evidence that humans naturally have a tendency to point towards objects in the context or the environment when the name of an object is not known.Using our proposed model, we observe improvements on two tasks, neural machine translation on the Europarl English to French parallel corpora and text summarization on the Gigaword dataset.
Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou 0002, Yoshua Bengio
ACL (1)3
2016 Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond
abstract
In this work, we model abstractive text summarization using Attentional Encoder-Decoder Recurrent Neural Networks, and show that they achieve state-of-the-art performance on two different corpora.We propose several novel models that address critical problems in summarization that are not adequately modeled by the basic architecture, such as modeling key-words, capturing the hierarchy of sentence-toword structure, and emitting words that are rare or unseen at training time.Our work shows that many of our proposed models contribute to further improvement in performance.We also propose a new dataset consisting of multi-sentence summaries, and establish performance benchmarks for further research.
Ramesh Nallapati, Bowen Zhou 0002, Cícero Nogueira dos Santos, Caglar Gulcehre, Bing Xiang
CoNLL1
2014 Joint question clustering and relevance prediction for open domain non-factoid question answering
abstract
Web searches are increasingly formulated as natural language questions, rather than keyword queries. Retrieving answers to such questions requires a degree of understanding of user expectations. An important step in this direction is to automatically infer the type of answer implied by the question, e.g., factoids, statements on a topic, instructions, reviews, etc. Answer Type taxonomies currently exist for factoid-style questions, but not for open-domain questions. Building taxonomies for non-factoid questions is a harder problem since these questions can come from a very broad semantic space. A few attempts have been made to develop taxonomies for non-factoid questions, but these tend to be too narrow or domain specific. In this paper, we address this problem by modeling the Answer Type as a latent variable that is learned in a data-driven fashion, allowing the model to be more adaptive to new domains and data sets. We propose approaches that detect the relevance of candidate answers to a user question by jointly 'clustering' questions according to the hidden variable, and modeling relevance conditioned on this hidden variable.
Snigdha Chaturvedi, Vittorio Castelli, Radu Florian, Ramesh Nallapati, Hema Raghavan
WWW4
2014 Evaluating multimedia features and fusion for example-based event detection
abstract
Multimedia event detection (MED) is a challenging problem because of the heterogeneous content and variable quality found in large collections of Internet videos. To study the value of multimedia features and fusion for representing and learning events from a set of example video clips, we created SESAME, a system for video SEarch with Speed and Accuracy for Multimedia Events. SESAME includes multiple bag-of-words event classifiers based on single data types: low-level visual, motion, and audio features; high-level semantic visual concepts; and automatic speech recognition. Event detection performance was evaluated for each event classifier. The performance of low-level visual and motion features was improved by the use of difference coding. The accuracy of the visual concepts was nearly as strong as that of the low-level visual features. Experiments with a number of fusion methods for combining the event detection scores from these classifiers revealed that simple fusion methods, such as arithmetic mean, perform as well as or better than other, more complex fusion methods. SESAME’s performance in the 2012 TRECVID MED evaluation was one of the best reported.
Gregory K. Myers, Ramesh Nallapati, Julien van Hout, Stephanie Pancoast, Ramakant Nevatia, Chen Sun 0002, AmirHossein Habibian, Dennis C. Koelma, Koen E. A. van de Sande, Arnold W. M. Smeulders, Cees Snoek
Mach. Vis. Appl.2
2012 Multi-instance Multi-label Learning for Relation Extraction
Mihai Surdeanu, Julie Tibshirani, Ramesh Nallapati, Christopher D. Manning
EMNLP-CoNLL3
2011 Risk analysis for intellectual property litigation
abstract
We introduce the problem of risk analysis for Intellectual Property (IP) lawsuits. More specifically, we focus on estimating the risk for participating parties using solely prior factors, i. e., historical and concurrent behavior of the entities involved in the case. This work represents a first step towards building a comprehensive legal risk assessment system for parties involved in litigation. This technology will allow parties to optimize their case parameters to minimize their own risk, or to settle disputes out of court and thereby ease the burden on the judicial system. In addition, it will also help U.S. courts detect and fix any inherent biases in the system.
Mihai Surdeanu, Ramesh Nallapati, George Gregory, Joshua Walker, Christopher D. Manning
ICAIL2
2011 LeadLag LDA: Estimating Topic Specific Leads and Lags of Information Outlets
Ramesh Nallapati, Daniel A. McFarland, Jure Leskovec, Daniel Jurafsky
ICWSM1
2009 Labeled LDA: A supervised topic model for credit attribution in multi-labeled corpora
Daniel Ramage, David Hall 0006, Ramesh Nallapati, Christopher D. Manning
EMNLP3
2008 Exploiting Feature Hierarchy for Transfer Learning in Named Entity Recognition
Andrew O. Arnold, Ramesh Nallapati, William W. Cohen
ACL2
2008 Legal Docket Classification: Where Machine Learning Stumbles
Ramesh Nallapati, Christopher D. Manning
EMNLP1
2008 Link-PLSA-LDA: A New Unsupervised Model for Topics and Influence of Blogs
Ramesh Nallapati, William W. Cohen
ICWSM1
2008 Joint latent topic models for text and citations
abstract
In this work, we address the problem of joint modeling of text and citations in the topic modeling framework. We present two different models called the Pairwise-Link-LDA and the Link-PLSA-LDA models.
Ramesh Nallapati, Amr Ahmed 0001, Eric P. Xing, William W. Cohen
KDD1
2007 Multiscale topic tomography
abstract
Modeling the evolution of topics with time is of great value in automatic summarization and analysis of large document collections. In this work, we propose a new probabilistic graphical model to address this issue. The new model, which we call the Multiscale Topic Tomography Model (MTTM), employs non-homogeneous Poisson processes to model generation of word-counts. The evolution of topics is modeled through a multi-scale analysis using Haar wavelets. One of the new features of the model is its modeling the evolution of topics at various time-scales of resolution, allowing the user to zoom in and out of the time-scales. Our experiments on Science data using the new model uncovers some interesting patterns in topics. The new model is also comparable to LDA in predicting unseen data as demonstrated by our perplexity experiments.
Ramesh Nallapati, Susan Ditmore, John D. Lafferty, Kin Ung
KDD1
2004 Event threading within news topics
abstract
With the overwhelming volume of online news available today, there is an increasing need for automatic techniques to analyze and present news to the user in a meaningful and efficient manner. Previous research focused only on organizing news stories by their topics into a flat hierarchy. We believe viewing a news topic as a flat collection of stories is too restrictive and inefficient for a user to understand the topic quickly.
Ramesh Nallapati, Ao Feng, Fuchun Peng, James Allan 0001
CIKM1
2004 Discriminative models for information retrieval
abstract
Discriminative models have been preferred over generative models in many machine learning problems in the recent past owing to some of their attractive theoretical properties. In this paper, we explore the applicability of discriminative classifiers for IR. We have compared the performance of two popular discriminative models, namely the maximum entropy model and support vector machines with that of language modeling, the state-of-the-art generative model for IR. Our experiments on ad-hoc retrieval indicate that although maximum entropy is significantly worse than language models, support vector machines are on par with language models. We argue that the main reason to prefer SVMs over language models is their ability to learn arbitrary features automatically as demonstrated by our experiments on the home-page finding task of TREC-10.
Ramesh Nallapati
SIGIR1
2003 Relevant query feedback in statistical language modeling
abstract
In traditional relevance feedback, researchers have explored relevant document feedback, wherein, the query representation is updated based on a set of relevant documents returned by the user. In this work, we investigate relevant query feedback, in which we update a document's representation based on a set of relevant queries. We propose four statistical models to incorporate relevant query feedback.To validate our models, we considered anchor text of incoming links to a given document as feedback queries and performed experiments on the home-page retrieval task of TREC 2001. Our results show that three of our four models outperform the query-likelihood baseline by at least 35% in MRR score on a test set.
Ramesh Nallapati, W. Bruce Croft, James Allan 0001
CIKM1
2003 Semantic Language Models for Topic Detection and Tracking
Ramesh Nallapati
HLT-NAACL1
2002 Capturing term dependencies using a language model based on sentence trees
abstract
We describe a new probabilistic Sentence Tree Language Modeling approach that captures term dependency patterns in Topic Detection and Tracking's (TDT) Story Link Detection task. New features of the approach include modeling the syntactic structure of sentences in documents by a sentence-bin approach and a computationally efficient algorithm for capturing the most significant sentence-level term dependencies using a Maximum Spanning Tree approach, similar to Van Rijsbergen's modeling of document-level term dependencies.The new model is a good discriminator of on-topic and off-topic story pairs providing evidence that sentence-level term dependencies contain significant information about relevance. Although runs on a subset of the TDT2 corpus show that the model is outperformed by the unigram language model, a mixture of the unigram and the Sentence Tree models is shown to improve on the best performance especially in the regions of low false alarms.
Ramesh Nallapati, James Allan 0001
CIKM1