Vivek Srikumar

dblp:37/44 · DBLP profile ↗
← Back
62ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0003-0419-6568ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 6 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Systems, architecture and hardware · 1Security and privacy · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum
abstract
While dense retrieval models have been the standard for state-of-the-art information retrieval, their deployment is often constrained by high memory requirements and reliance on GPU accelerators for vector similarity search at scale. Learned sparse retrieval offers a compelling alternative by enabling efficient search via inverted indices, yet it has historically received less attention than dense approaches. In this paper, we introduce LACONIC, a family of learned sparse retrievers based on the Llama3 architecture (1B, 3B, and 8B). We propose a streamlined two-phase training curriculum consisting of (1) weakly supervised pre-finetuning to adapt causal LLMs for bidirectional contextualization and (2) high-signal finetuning using curated hard negatives. Our results demonstrate that LACONIC effectively bridges the performance gap with dense models: the 8B variant achieves a state-of-the-art 60.2 nDCG@10 on the MTEB Retrieval benchmark, ranking 15th on the leaderboard as of February 5th, 2026, while utilizing 74% less index memory than an equivalent dense model. By delivering high retrieval effectiveness on commodity CPU hardware with a fraction of the compute budget required by competing models, LACONIC provides a scalable and efficient solution for real-world search applications. We fully open source our code implementation and trained checkpoints to facilitate reproducibility.
Zhichao Xu 0001, Shengyao Zhuang, Xinyu Zhang 0018, Xueguang Ma, Yijun Tian 0001, Maitrey Mehta, Jimmy Lin, Vivek Srikumar
SIGIR8
2025 Understanding the Logic of Direct Preference Alignment through Logic
abstract
Recent direct preference alignment algorithms (DPA), such as DPO, have shown great promise in aligning large language models to human preferences. While this has motivated the development of many new variants of the original DPO loss, understanding the differences between these recent proposals, as well as developing new DPA loss functions, remains difficult given the lack of a technical and conceptual framework for reasoning about the underlying semantics of these algorithms. In this paper, we attempt to remedy this by formalizing DPA losses in terms of discrete reasoning problems. Specifically, we ask: Given an existing DPA loss, can we systematically derive a symbolic program that characterizes its semantics? We propose a novel formalism for characterizing preference losses for single model and reference model based approaches, and identify symbolic forms for a number of commonly used DPA variants. Further, we show how this formal view of preference learning sheds new light on both the size and structure of the DPA loss landscape, making it possible to not only rigorously characterize the relationships between recent loss proposals but also to systematically explore the landscape and derive new loss functions from first principles. We hope our framework and findings will help provide useful guidance to those working on human AI alignment.
Kyle Richardson 0001, Vivek Srikumar, Ashish Sabharwal
ICML2
2025 The Role of Transformer Architecture in the Logic-as-Loss Framework
Mattia Medina Grespan, Vivek Srikumar
ECML/PKDD (5)2
2024 Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness
abstract
Ashim Gupta, Rishanth Rajendhran, Nathan Stringham, Vivek Srikumar, Ana Marasovic. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Ashim Gupta, Rishanth Rajendhran, Nathan Stringham, Vivek Srikumar, Ana Marasovic
NAACL-HLT4
2024 Promptly Predicting Structures: The Return of Inference
abstract
Maitrey Mehta, Valentina Pyatkin, Vivek Srikumar. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Maitrey Mehta, Valentina Pyatkin, Vivek Srikumar
NAACL-HLT3
2024 An In-depth Investigation of User Response Simulation for Conversational Search
abstract
Conversational search has seen increased recent attention in both the IR and NLP communities. It seeks to clarify and solve users' search needs through multi-turn natural language interactions. However, most existing systems are trained and demonstrated with recorded or artificial conversation logs. Eventually, conversational search systems should be trained, evaluated, and deployed in an open-ended setting with unseen conversation trajectories. A key challenge is that training and evaluating such systems both require a human-in-the-loop, which is expensive and does not scale. One strategy is to simulate users, thereby reducing the scaling costs. However, current user simulators are either limited to only responding to yes-no questions from the conversational search system or unable to produce high-quality responses in general.
Zhenduo Wang, Zhichao Xu 0001, Vivek Srikumar, Qingyao Ai
WWW3
2024 VERB: Visualizing and Interpreting Bias Mitigation Techniques Geometrically for Word Representations
abstract
Word vector embeddings have been shown to contain and amplify biases in the data they are extracted from. Consequently, many techniques have been proposed to identify, mitigate, and attenuate these biases in word representations. In this article, we utilize interactive visualization to increase the interpretability and accessibility of a collection of state-of-the-art debiasing techniques. To aid this, we present the Visualization of Embedding Representations for deBiasing (VERB) system, an open-source web-based visualization tool that helps users gain a technical understanding and visual intuition of the inner workings of debiasing techniques, with a focus on their geometric properties. In particular, VERB offers easy-to-follow examples that explore the effects of these debiasing techniques on the geometry of high-dimensional word vectors. To help understand how various debiasing techniques change the underlying geometry, VERB decomposes each technique into interpretable sequences of primitive transformations and highlights their effect on the word vectors using dimensionality reduction and interactive visual exploration. VERB is designed to target natural language processing (NLP) practitioners who are designing decision-making systems on top of word embeddings and researchers working with the fairness and ethics of machine learning systems in NLP. It can also serve as a visual medium for education, which helps an NLP novice understand and mitigate biases in word embeddings.
Archit Rathore, Sunipa Dev, Jeff M. Phillips, Vivek Srikumar, Yan Zheng 0001, Chin-Chia Michael Yeh, Junpeng Wang 0001, Wei Zhang 0189, Bei Wang 0001
ACM Trans. Interact. Intell. Syst.4
2023 Logic-driven Indirect Supervision: An Application to Crisis Counseling
abstract
Ensuring the effectiveness of text-based crisis counseling requires observing ongoing conversations and providing feedback, both labor-intensive tasks. Automatic analysis of conversations-at the full chat and utterance levels-may help support counselors and provide better care. While some session-level training data (e.g., rating of patient risk) is often available from counselors, labeling utterances requires expensive post hoc annotation. But the latter can not only provide insights about conversation dynamics, but can also serve to support quality assurance efforts for counselors. In this paper, we examine if inexpensive-and potentially noisy-session-level annotation can help improve label utterances. To this end, we propose a logic-based indirect supervision approach that exploits declaratively stated structural dependencies between both levels of annotation to improve utterance modeling. We show that adding these rules gives an improvement of 3.5% f-score over a strong multi-task baseline for utterance-level predictions. We demonstrate via ablation studies how indirect supervision via logic rules also improves the consistency and robustness of the system.
Mattia Medina Grespan, Meghan Broadbent, Katherine Axford, Brent Kious, Zac E. Imel, Vivek Srikumar
ACL (1)7
2023 Don't Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting Text
abstract
Ashim Gupta, Carter Blum, Temma Choji, Yingjie Fei, Shalin Shah, Alakananda Vempala, Vivek Srikumar. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Ashim Gupta, Carter Wood Blum, Temma Choji, Yingjie Fei, Shalin Shah, Alakananda Vempala, Vivek Srikumar
ACL (1)7
2023 ClarifyDelphi: Reinforced Clarification Questions with Defeasibility Rewards for Social and Moral Situations
abstract
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Liwei Jiang, Yejin Choi, Chandra Bhagavatula. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Valentina Pyatkin, Jena D. Hwang, Vivek Srikumar, Ximing Lu, Yejin Choi 0001, Chandra Bhagavatula
ACL (1)3
2023 Elaboration-Generating Commonsense Question Answering at Scale
abstract
In question answering requiring common sense, language models (e.g., GPT-3) have been used to generate text expressing background knowledge that helps improve performance.Yet the cost of working with such models is very high; in this work, we finetune smaller language models to generate useful intermediate context, referred to here as elaborations.Our framework alternates between updating two language models-an elaboration generator and an answer predictor-allowing each to influence the other.Using less than 0.5% of the parameters of GPT-3, our model outperforms alternatives with similar sizes and closes the gap with GPT-3 on four commonsense question answering benchmarks.Human evaluations show that the quality of the generated elaborations is high. 1
Wenya Wang 0001, Vivek Srikumar, Hannaneh Hajishirzi, Noah A. Smith
ACL (1)2
2023 TempTabQA: Temporal Question Answering for Semi-Structured Tables
abstract
Semi-structured data, such as Infobox tables, often include temporal information about entities, either implicitly or explicitly.Can current NLP systems reason about such information in semi-structured tables?To tackle this question, we introduce the task of temporal question answering on semi-structured tables.We present a dataset, TEMPTABQA, which comprises 11,454 question-answer pairs extracted from 1,208 Wikipedia Infobox tables spanning more than 90 distinct domains.Using this dataset, we evaluate several state-ofthe-art models for temporal reasoning.We observe that even the top-performing LLMs lag behind human performance by more than 13.5 F1 points.Given these results, our dataset has the potential to serve as a challenging benchmark to improve the temporal reasoning capabilities of NLP models.
Vivek Gupta 0001, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang 0006, Yujie He 0003, Ridho Reinanda, Vivek Srikumar
EMNLP7
2023 AGRO: Adversarial discovery of error-prone Groups for Robust Optimization
Bhargavi Paranjape, Pradeep Dasigi, Vivek Srikumar, Luke Zettlemoyer, Hannaneh Hajishirzi
ICLR3
2022 PYLON: A PyTorch Framework for Learning with Constraints
abstract
Deep learning excels at learning task information from large amounts of data, but struggles with learning from declarative high-level knowledge that can be more succinctly expressed directly. In this work, we introduce PYLON, a neuro-symbolic training framework that builds on PyTorch to augment procedurally trained models with declaratively specified knowledge. PYLON lets users programmatically specify constraints as Python functions and compiles them into a differentiable loss, thus training predictive models that fit the data whilst satisfying the specified constraints. PYLON includes both exact as well as approximate compilers to efficiently compute the loss, employing fuzzy logic, sampling methods, and circuits, ensuring scalability even to complex models and constraints. Crucially, a guiding principle in designing PYLON is the ease with which any existing deep learning codebase can be extended to learn from constraints in a few lines code: a function that expresses the constraint, and a single line to compile it into a loss. Our demo comprises of models in NLP, computer vision, logical games, and knowledge graphs that can be interactively trained using constraints as supervision.
Kareem Ahmed, Tao Li 0039, Thy Ton, Quan Guo, Kai-Wei Chang 0001, Parisa Kordjamshidi, Vivek Srikumar, Guy Van den Broeck, Sameer Singh 0001
AAAI7
2022 Right for the Right Reason: Evidence Extraction for Trustworthy Tabular Reasoning
abstract
Vivek Gupta, Shuo Zhang, Alakananda Vempala, Yujie He, Temma Choji, Vivek Srikumar. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Vivek Gupta 0001, Shuo Zhang 0006, Alakananda Vempala, Yujie He 0003, Temma Choji, Vivek Srikumar
ACL (1)6
2022 A Closer Look at How Fine-tuning Changes BERT
abstract
Given the prevalence of pre-trained contextualized representations in today's NLP, there have been many efforts to understand what information they contain, and why they seem to be universally successful.The most common approach to use these representations involves fine-tuning them for an end task.Yet, how fine-tuning changes the underlying embedding space is less studied.In this work, we study the English BERT family and use two probing techniques to analyze how fine-tuning changes the space.We hypothesize that fine-tuning affects classification performance by increasing the distances between examples associated with different labels.We confirm this hypothesis with carefully designed experiments on five different NLP tasks.Via these experiments, we also discover an exception to the prevailing wisdom that "fine-tuning always improves performance".Finally, by comparing the representations before and after fine-tuning, we discover that fine-tuning does not introduce arbitrary changes to representations; instead, it adjusts the representations to downstream tasks while largely preserving the original spatial structure of the data points.
Yichu Zhou, Vivek Srikumar
ACL (1)2
2022 Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
abstract
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
Vivek Gupta 0001, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics6
2021 BERT & Family Eat Word Salad: Experiments with Text Understanding
abstract
In this paper, we study the response of large models from the BERT family to incoherent inputs that should confuse any model that claims to understand natural language. We define simple heuristics to construct such examples. Our experiments show that state-of-the-art models consistently fail to recognize them as ill-formed, and instead produce high confidence predictions on them. As a consequence of this phenomenon, models trained on sentences with randomly permuted word order perform close to state-of-the-art models. To alleviate these issues, we show that if models are explicitly trained to recognize invalid inputs, they can be robust to such attacks without a drop in performance.
Ashim Gupta, Giorgi Kvernadze, Vivek Srikumar
AAAI3
2021 OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings
abstract
Language representations are known to carry certain associations (e.g., gendered connotations) which may lead to invalid and harmful predictions in downstream tasks.While existing methods are effective at mitigating such unwanted associations by linear projection, we argue that they are too aggressive: not only do they remove such associations, they also erase information that should be retained.To address this issue, we propose OS-CAR (Orthogonal Subspace Correction and Rectification), a balanced approach of mitigation that focuses on disentangling associations between concepts that are deemed problematic, instead of removing concepts wholesale.We develop new measurements for evaluating information retention relevant to the debiasing goal.Our experiments on genderoccupation associations show that OSCAR is a well-balanced approach that ensures that semantic information is retained in the embeddings and unwanted associations are also effectively mitigated.
Sunipa Dev, Tao Li 0039, Jeff M. Phillips, Vivek Srikumar
EMNLP (1)4
2021 Putting Words in BERT's Mouth: Navigating Contextualized Vector Spaces with Pseudowords
abstract
We present a method for exploring regions around individual points in a contextualized vector space (particularly, BERT space), as a way to investigate how these regions correspond to word senses.By inducing a contextualized "pseudoword" as a stand-in for a static embedding in the input layer, and then performing masked prediction of a word in the sentence, we are able to investigate the geometry of the BERT-space in a controlled manner around individual instances.Using our method on a set of carefully constructed sentences targeting ambiguous English words, we find substantial regularity in the contextualized space, with regions that correspond to distinct word senses; but between these regions there are occasionally "sense voids"-regions that do not correspond to any intelligible sense. 1Learn pseudoword in place of that is customized to reconstruct .
Taelin Karidi, Yichu Zhou, Nathan Schneider 0001, Omri Abend, Vivek Srikumar
EMNLP (1)5
2021 Evaluating Relaxations of Logic for Neural Networks: A Comprehensive Study
abstract
Symbolic knowledge can provide crucial inductive bias for training neural models, especially in low data regimes. A successful strategy for incorporating such knowledge involves relaxing logical statements into sub-differentiable losses for optimization. In this paper, we study the question of how best to relax logical expressions that represent labeled examples and knowledge about a problem; we focus on sub-differentiable t-norm relaxations of logic. We present theoretical and empirical criteria for characterizing which relaxation would perform best in various scenarios. In our theoretical study driven by the goal of preserving tautologies, the Lukasiewicz t-norm performs best. However, in our empirical analysis on the text chunking and digit recognition tasks, the product t-norm achieves best predictive performance. We analyze this apparent discrepancy, and conclude with a list of best practices for defining loss functions via logic.
Mattia Medina Grespan, Ashim Gupta, Vivek Srikumar
IJCAI3
2021 A Visual Tour of Bias Mitigation Techniques for Word Representations
abstract
Word vector embeddings have been shown to contain and amplify biases in data they are extracted from. Consequently, many techniques have been proposed to identify, mitigate, and attenuate these biases in word representations. In this tutorial, we will review a collection of state-of-the-art debiasing techniques. To aid this, we provide an open source web-based visualization tool and offer hands-on experience in exploring the effects of these debiasing techniques on the geometry of high-dimensional word vectors. To help understand how various debiasing techniques change the underlying geometry, we decompose each technique into interpretable sequences of primitive operations, and study their effect on the word vectors using dimensionality reduction and interactive visual exploration.
Archit Rathore, Sunipa Dev, Jeff M. Phillips, Vivek Srikumar, Bei Wang 0001
KDD4
2021 Incorporating External Knowledge to Enhance Tabular Reasoning
abstract
Reasoning about tabular information presents unique challenges to modern NLP approaches which largely rely on pre-trained contextualized embeddings of text.In this paper, we study these challenges through the problem of tabular natural language inference.We propose easy and effective modifications to how information is presented to a model for this task.We show via systematic experiments that these strategies substantially improve tabular inference performance.
J. Neeraja, Vivek Gupta 0001, Vivek Srikumar
NAACL-HLT3
2021 DirectProbe: Studying Representations without Classifiers
abstract
Understanding how linguistic structure is encoded in contextualized embedding could help explain their impressive performance across NLP.Existing approaches for probing them usually call for training classifiers and use the accuracy, mutual information, or complexity as a proxy for the representation's goodness.In this work, we argue that doing so can be unreliable because different representations may need different classifiers.We develop a heuristic, DIRECTPROBE, that directly studies the geometry of a representation by building upon the notion of a version space for a task.Experiments with several linguistic tasks and contextualized embeddings show that, even without training classifiers, DIRECTPROBE can shine light into how an embedding space represents labels, and also anticipate classifier performance for the representation.
Yichu Zhou, Vivek Srikumar
NAACL-HLT2
2021 Database Workload Characterization with Query Plan Encoders
abstract
Smart databases are adopting artificial intelligence (AI) technologies to achieve instance optimality , and in the future, databases will come with prepackaged AI models within their core components. The reason is that every database runs on different workloads, demands specific resources, and settings to achieve optimal performance. It prompts the necessity to understand workloads running in the system along with their features comprehensively, which we dub as workload characterization. To address this workload characterization problem, we propose our query plan encoders that learn essential features and their correlations from query plans. Our pretrained encoders captures the structural and the computational performance of queries independently. We show that our pretrained encoders are adaptable to workloads that expedites the transfer learning process. We performed independent assessments of structural encoder and performance encoders with multiple downstream tasks. For the overall evaluation of our query plan encoders, we architect two downstream tasks (i) query latency prediction and (ii) query classification. These tasks show the importance of feature-based workload characterization. We also performed extensive experiments on individual encoders to verify the effectiveness of representation learning, and domain adaptability.
Debjyoti Paul, Jie Cao 0010, Feifei Li 0001, Vivek Srikumar
Proc. VLDB Endow.4
2021 Supertagging the Long Tail with Tree-Structured Decoding of Complex Categories
abstract
Abstract Although current CCG supertaggers achieve high accuracy on the standard WSJ test set, few systems make use of the categories’ internal structure that will drive the syntactic derivation during parsing. The tagset is traditionally truncated, discarding the many rare and complex category types in the long tail. However, supertags are themselves trees. Rather than give up on rare tags, we investigate constructive models that account for their internal structure, including novel methods for tree-structured prediction. Our best tagger is capable of recovering a sizeable fraction of the long-tail supertags and even generates CCG categories that have never been seen in training, while approximating the prior state of the art in overall tag accuracy with fewer parameters. We further investigate how well different approaches generalize to out-of-domain evaluation sets.
Jakob Prange, Nathan Schneider 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics3
2020 On Measuring and Mitigating Biased Inferences of Word Embeddings
abstract
Word embeddings carry stereotypical connotations from the text they are trained on, which can lead to invalid inferences in downstream models that rely on them. We use this observation to design a mechanism for measuring stereotypes using the task of natural language inference. We demonstrate a reduction in invalid inferences via bias mitigation strategies on static word embeddings (GloVe). Further, we show that for gender bias, these techniques extend to contextualized embeddings when applied selectively only to the static components of contextualized embeddings (ELMo, BERT).
Sunipa Dev, Tao Li 0039, Jeff M. Phillips, Vivek Srikumar
AAAI4
2020 INFOTABS: Inference on Tables as Semi-structured Data
abstract
In this paper, we observe that semi-structured tabulated text is ubiquitous; understanding them requires not only comprehending the meaning of text fragments, but also implicit relationships between them.We argue that such data can prove as a testing ground for understanding how we reason about information.To study this, we introduce a new dataset called INFOTABS, comprising of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.Our analysis shows that the semi-structured, multi-domain and heterogeneous nature of the premises admits complex, multi-faceted reasoning.Experiments reveal that, while human annotators agree on the relationships between a table-hypothesis pair, several standard modeling strategies are unsuccessful at the task, suggesting that reasoning about tables can pose a difficult modeling challenge.
Vivek Gupta 0001, Maitrey Mehta, Pegah Nokhiz, Vivek Srikumar
ACL4
2020 Structured Tuning for Semantic Role Labeling
abstract
Recent neural network-driven semantic role labeling (SRL) systems have shown impressive improvements in F1 scores.These improvements are due to expressive input representations, which, at least at the surface, are orthogonal to knowledge-rich constrained decoding mechanisms that helped linear SRL models.Introducing the benefits of structure to inform neural models presents a methodological challenge.In this paper, we present a structured tuning framework to improve models using softened constraints only at training time.Our framework leverages the expressiveness of neural networks and provides supervision with structured loss components.We start with a strong baseline (RoBERTa) to validate the impact of our approach, and show that our framework outperforms the baseline by learning to comply with declarative constraints.Additionally, our experiments with smaller training sizes show that we can achieve consistent improvements under low-resource scenarios.
Tao Li 0039, Parth Anand Jawale, Martha Palmer, Vivek Srikumar
ACL4
2020 Learning Constraints for Structured Prediction Using Rectifier Networks
abstract
Various natural language processing tasks are structured prediction problems where outputs are constructed with multiple interdependent decisions.Past work has shown that domain knowledge, framed as constraints over the output space, can help improve predictive accuracy.However, designing good constraints often relies on domain expertise.In this paper, we study the problem of learning such constraints.We frame the problem as that of training a two-layer rectifier network to identify valid structures or substructures, and show a construction for converting a trained network into a system of linear constraints over the inference variables.Our experiments on several NLP tasks show that the learned constraints can improve the prediction accuracy, especially when the number of training examples is small.
Xingyuan Pan, Maitrey Mehta, Vivek Srikumar
ACL3
2019 Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
abstract
Automatically analyzing dialogue can help understand and guide behavior in domains such as counseling, where interactions are largely mediated by conversation.In this paper, we study modeling behavioral codes used to asses a psychotherapy treatment style called Motivational Interviewing (MI), which is effective for addressing substance abuse and related problems.Specifically, we address the problem of providing real-time guidance to therapists with a dialogue observer that (1) categorizes therapist and client MI behavioral codes and, (2) forecasts codes for upcoming utterances to help guide the conversation and potentially alert the therapist.For both tasks, we define neural network models that build upon recent successes in dialogue modeling.Our experiments demonstrate that our models can outperform several baselines for both tasks.We also report the results of a careful analysis that reveals the impact of the various network design tradeoffs for modeling therapy dialogue.Code Count Description Examples Client Behavioral Codes FN 47715 Follow/ Neutral: unrelated to changing or sustaining behavior."You know, I didn't smoke for a while.""I have smoked for forty years now."CT 5099 Utterances about changing unhealthy behavior."I want to stop smoking."ST 4378 Utterances about sustaining unhealthy behavior."I really don't think I smoke too much."Therapist Behavioral Codes FA 17468 Facilitate conversation "Mm Hmm.", "OK.","Tell me more."GI 15271 Give information or feedback."I'm Steve.","Yes, alcohol is a depressant."RES 6246 Simple reflection about the clients most recent utterance.C: "I didn't smoke last week" T: "Cool, you avoided smoking last week."REC 4651 Complex reflection based on a client's history or the broader conversation.C: "I didn't smoke last week."T: "You mean things begin to change".QUC 5218 Closed question "Did you smoke this week?"QUO 4509 Open question "Tell me more about your week."MIA 3869 Other MI adherent,e.g., affirmation, advising with permission, etc. "You've accomplished a difficult task." "Is it OK if I suggested something?"MIN 1019 MI non-adherent, e.g., confrontation, advising without permission, etc. "You hurt the baby's health for cigarettes?" "You ask them not to drink at your house."
Jie Cao 0010, Michael Tanana, Zac E. Imel, Eric G. Poitras, David C. Atkins, Vivek Srikumar
ACL (1)6
2019 Augmenting Neural Networks with First-order Logic
abstract
Today, the dominant paradigm for training neural networks involves minimizing task loss on a large dataset.Using world knowledge to inform a model, and yet retain the ability to perform end-to-end training remains an open question.In this paper, we present a novel framework for introducing declarative knowledge to neural network architectures in order to guide training and prediction.Our framework systematically compiles logical statements into computation graphs that augment a neural network without extra learnable parameters or manual redesign.We evaluate our modeling strategy on three tasks: machine comprehension, natural language inference, and text chunking.Our experiments show that knowledge-augmented networks can strongly improve over baselines, especially in low-data regimes.Gaius Julius Caesar (July 100 BC -15 March 44 BC), Roman general, statesman, Consul and notable author of Latin prose, played a critical role in the events that led to the demise of the Roman Republic and the rise of the Roman Empire through his various military campaigns.
Tao Li 0039, Vivek Srikumar
ACL (1)2
2019 On the Limits of Learning to Actively Learn Semantic Representations
abstract
One of the goals of natural language understanding is to develop models that map sentences into meaning representations.However, training such models requires expensive annotation of complex structures, which hinders their adoption.Learning to actively-learn (LTAL) is a recent paradigm for reducing the amount of labeled data by learning a policy that selects which samples should be labeled.In this work, we examine LTAL for learning semantic representations, such as QA-SRL.We show that even an oracle policy that is allowed to pick examples that maximize performance on the test set (and constitutes an upper bound on the potential of LTAL), does not substantially improve performance compared to a random policy.We investigate factors that could explain this finding and show that a distinguishing characteristic of successful applications of LTAL is the interaction between optimization and the oracle policy selection process.In successful applications of LTAL, the examples selected by the oracle policy do not substantially depend on the optimization procedure, while in our setup the stochastic nature of optimization strongly affects the examples selected by the oracle.We conclude that the current applicability of LTAL for improving data efficiency in learning semantic meaning representations is limited.
Omri Koshorek, Gabriel Stanovsky, Yichu Zhou, Vivek Srikumar, Jonathan Berant
CoNLL4
2019 A Logic-Driven Framework for Consistency of Neural Models
abstract
Tao Li, Vivek Gupta, Maitrey Mehta, Vivek Srikumar. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Li 0039, Vivek Gupta 0001, Maitrey Mehta, Vivek Srikumar
EMNLP/IJCNLP (1)4
2019 NLIZE: A Perturbation-Driven Visual Interrogation Tool for Analyzing and Interpreting Natural Language Inference Models
abstract
With the recent advances in deep learning, neural network models have obtained state-of-the-art performances for many linguistic tasks in natural language processing. However, this rapid progress also brings enormous challenges. The opaque nature of a neural network model leads to hard-to-debug-systems and difficult-to-interpret mechanisms. Here, we introduce a visualization system that, through a tight yet flexible integration between visualization elements and the underlying model, allows a user to interrogate the model by perturbing the input, internal state, and prediction while observing changes in other parts of the pipeline. We use the natural language inference problem as an example to illustrate how a perturbation-driven paradigm can help domain experts assess the potential limitation of a model, probe its inner states, and interpret and form hypotheses about fundamental model mechanisms such as attention.
Shusen Liu 0001, Tao Li 0039, Vivek Srikumar, Valerio Pascucci, Peer-Timo Bremer
IEEE Trans. Vis. Comput. Graph.4
2018 Comprehensive Supersense Disambiguation of English Prepositions and Possessives
abstract
Nathan Schneider, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Nathan Schneider 0001, Jena D. Hwang, Vivek Srikumar, Jakob Prange, Austin Blodgett, Sarah R. Moeller, Aviram Stern, Adi Bitan, Omri Abend
ACL (1)3
2018 Learning to Speed Up Structured Output Prediction
abstract
Predicting structured outputs can be computationally onerous due to the combinatorially large output spaces. In this paper, we focus on reducing the prediction time of a trained black-box structured classifier without losing accuracy. To do so, we train a speedup classifier that learns to mimic a black-box classifier under the learning-to-search approach. As the structured classifier predicts more examples, the speedup classifier will operate as a learned heuristic to guide search to favorable regions of the output space. We present a mistake bound for the speedup classifier and identify inference situations where it can independently make correct judgments without input features. We evaluate our method on the task of entity and relation extraction and show that the speedup classifier outperforms even greedy search in terms of speed without loss of accuracy.
Xingyuan Pan, Vivek Srikumar
ICML2
2018 CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001
LREC6
2018 Visual Exploration of Semantic Relationships in Neural Word Embeddings
abstract
Constructing distributed representations for words through neural language models and using the resulting vector spaces for analysis has become a crucial component of natural language processing (NLP). However, despite their widespread application, little is known about the structure and properties of these spaces. To gain insights into the relationship between words, the NLP community has begun to adapt high-dimensional visualization techniques. In particular, researchers commonly use t-distributed stochastic neighbor embeddings (t-SNE) and principal component analysis (PCA) to create two-dimensional embeddings for assessing the overall structure and exploring linear relationships (e.g., word analogies), respectively. Unfortunately, these techniques often produce mediocre or even misleading results and cannot address domain-specific visualization challenges that are crucial for understanding semantic relationships in word embeddings. Here, we introduce new embedding techniques for visualizing semantic and syntactic analogies, and the corresponding tests to determine whether the resulting views capture salient structures. Additionally, we introduce two novel views for a comprehensive study of analogy relationships. Finally, we augment t-SNE embeddings to convey uncertainty information in order to allow a reliable interpretation. Combined, the different views address a number of domain-specific tasks difficult to solve with existing tools.
Shusen Liu 0001, Peer-Timo Bremer, Jayaraman J. Thiagarajan, Vivek Srikumar, Bei Wang 0001, Yarden Livnat, Valerio Pascucci
IEEE Trans. Vis. Comput. Graph.4
2017 An Algebra for Feature Extraction
abstract
Though feature extraction is a necessary first step in statistical NLP, it is often seen as a mere preprocessing step.Yet, it can dominate computation time, both during training, and especially at deployment.In this paper, we formalize feature extraction from an algebraic perspective.Our formalization allows us to define a message passing algorithm that can restructure feature templates to be more computationally efficient.We show via experiments on text chunking and relation extraction that this restructuring does indeed speed up feature extraction in practice by reducing redundant computation.
Vivek Srikumar
ACL (1)1
2017 DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning
abstract
Anomaly detection is a critical step towards building a secure and trustworthy system. The primary purpose of a system log is to record system states and significant events at various critical points to help debug system failures and perform root cause analysis. Such log data is universally available in nearly all computer systems. Log data is an important and valuable resource for understanding system status and performance issues; therefore, the various system logs are naturally excellent source of information for online monitoring and anomaly detection. We propose DeepLog, a deep neural network model utilizing Long Short-Term Memory (LSTM), to model a system log as a natural language sequence. This allows DeepLog to automatically learn log patterns from normal execution, and detect anomalies when log patterns deviate from the model trained from log data under normal execution. In addition, we demonstrate how to incrementally update the DeepLog model in an online fashion so that it can adapt to new log patterns over time. Furthermore, DeepLog constructs workflows from the underlying system log so that once an anomaly is detected, users can diagnose the detected anomaly and perform root cause analysis effectively. Extensive experimental evaluations over large log data have shown that DeepLog has outperformed other existing log-based anomaly detection methods based on traditional data mining methodologies.
Min Du 0003, Feifei Li 0001, Guineng Zheng, Vivek Srikumar
CCS4
2016 Exploiting Sentence Similarities for Better Alignments
abstract
We study the problem of jointly aligning sentence constituents and predicting their similarities.While extensive sentence similarity data exists, manually generating reference alignments and labeling the similarities of the aligned chunks is comparatively onerous.This prompts the natural question of whether we can exploit easy-to-create sentence level data to train better aligners.In this paper, we present a model that learns to jointly align constituents of two sentences and also predict their similarities.By taking advantage of both sentence and constituent level data, we show that our model achieves state-of-the-art performance at predicting alignments and constituent similarities.
Tao Li 0039, Vivek Srikumar
EMNLP2
2016 Expressiveness of Rectifier Networks
abstract
Rectified Linear Units (ReLUs) have been shown to ameliorate the vanishing gradient problem, allow for efficient backpropagation, and empirically promote sparsity in the learned parameters. They have led to state-of-the-art results in a variety of applications. However, unlike threshold and sigmoid networks, ReLU networks are less explored from the perspective of their expressiveness. This paper studies the expressiveness of ReLU networks. We characterize the decision boundary of two-layer ReLU networks by constructing functionally equivalent threshold networks. We show that while the decision boundary of a two-layer ReLU network can be captured by a threshold network, the latter may require an exponentially larger number of hidden units. We also formulate sufficient conditions for a corresponding logarithmic reduction in the number of hidden units to represent a sign network as a ReLU network. Finally, we experimentally compare threshold networks and their much smaller ReLU counterparts with respect to their ability to learn from synthetically generated data.
Xingyuan Pan, Vivek Srikumar
ICML2
2016 ISAAC: A Convolutional Neural Network Accelerator with In-Situ Analog Arithmetic in Crossbars
abstract
A number of recent efforts have attempted to design accelerators for popular machine learning algorithms, such as those involving convolutional and deep neural networks (CNNs and DNNs). These algorithms typically involve a large number of multiply-accumulate (dot-product) operations. A recent project, DaDianNao, adopts a near data processing approach, where a specialized neural functional unit performs all the digital arithmetic operations and receives input weights from adjacent eDRAM banks. This work explores an in-situ processing approach, where memristor crossbar arrays not only store input weights, but are also used to perform dot-product operations in an analog manner. While the use of crossbar memory as an analog dot-product engine is well known, no prior work has designed or characterized a full-fledged accelerator based on crossbars. In particular, our work makes the following contributions: (i) We design a pipelined architecture, with some crossbars dedicated for each neural network layer, and eDRAM buffers that aggregate data between pipeline stages. (ii) We define new data encoding techniques that are amenable to analog computations and that can reduce the high overheads of analog-to-digital conversion (ADC). (iii) We define the many supporting digital components required in an analog CNN accelerator and carry out a design space exploration to identify the best balance of memristor storage/compute, ADCs, and eDRAM storage on a chip. On a suite of CNN and DNN workloads, the proposed ISAAC architecture yields improvements of 14.8×, 5.5×, and 7.5× in throughput, energy, and computational density (respectively), relative to the state-of-the-art DaDianNao architecture.
Ali Shafiee, Anirban Nag, Naveen Muralimanohar, Rajeev Balasubramonian, John Paul Strachan, Miao Hu 0002, R. Stanley Williams, Vivek Srikumar
ISCA8
2016 EDISON: Feature Extraction for NLP, Simplified
Mark Sammons, Christos Christodoulopoulos 0001, Parisa Kordjamshidi, Daniel Khashabi, Vivek Srikumar, Dan Roth 0001
LREC5
2016 Continuous Kernel Learning
John Moeller, Vivek Srikumar, Sarathkrishna Swaminathan, Suresh Venkatasubramanian, Dustin Webb
ECML/PKDD (2)2
2014 Correcting Grammatical Verb Errors
abstract
Verb errors are some of the most common mistakes made by non-native writers of English but some of the least studied. The reason is that dealing with verb errors requires a new paradigm; essentially all research done on correcting grammatical errors assumes a closed set of triggers ‐ e.g., correcting the use of prepositions or articles ‐ but identifying mistakes in verbs necessitates identifying potentially ambiguous triggers first, and then determining the type of mistake made and correcting it. Moreover, once the verb is identified, modeling verb errors is challenging because verbs fulfill many grammatical functions, resulting in a variety of mistakes. Consequently, the little earlier work done on verb errors assumed that the error type is known in advance. We propose a linguistically-motivated approach to verb error correction that makes use of the notion of verb finiteness to identify triggers and types of mistakes, before using a statistical machine learning approach to correct these mistakes. We show that the linguistically-informed model significantly improves the accuracy of the verb correction approach.
Alla Rozovskaya, Dan Roth 0001, Vivek Srikumar
EACL3
2014 Modeling Biological Processes for Reading Comprehension
abstract
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Jonathan Berant, Vivek Srikumar, Pei-Chun Chen, Abby Vander Linden, Brittany Harding, Brad Huang, Peter Clark, Christopher D. Manning
EMNLP2
2014 Learning Distributed Representations for Structured Output Prediction
Vivek Srikumar, Christopher D. Manning
NIPS1
2013 Margin-based Decomposed Amortized Inference
Gourab Kundu, Vivek Srikumar, Dan Roth 0001
ACL (1)2
2013 Multi-core Structural SVM Training
Kai-Wei Chang 0001, Vivek Srikumar, Dan Roth 0001
ECML/PKDD (2)2
2013 Modeling Semantic Relations Expressed by Prepositions
abstract
This paper introduces the problem of predicting semantic relations expressed by prepositions and develops statistical learning models for predicting the relations, their arguments and the semantic types of the arguments. We define an inventory of 32 relations, building on the word sense disambiguation task for prepositions and collapsing related senses across prepositions. Given a preposition in a sentence, our computational task to jointly model the preposition relation and its arguments along with their semantic types, as a way to support the relation prediction. The annotated data, however, only provides labels for the relation label, and not the arguments and types. We address this by presenting two models for preposition relation labeling. Our generalization of latent structure SVM gives close to 90% accuracy on relation labeling. Further, by jointly predicting the relation, arguments, and their types along with preposition sense, we show that we can not only improve the relation accuracy, but also significantly improve sense prediction accuracy.
Vivek Srikumar, Dan Roth 0001
Trans. Assoc. Comput. Linguistics1
2012 Learning shared body plans
abstract
We cast the problem of recognizing related categories as a unified learning and structured prediction problem with shared body plans. When provided with detailed annotations of objects and their parts, these body plans model objects in terms of shared parts and layouts, simultaneously capturing a variety of categories in varied poses. We can use these body plans to jointly train many detectors in a shared framework with structured learning, leading to significant gains for each supervised task. Using our model, we can provide detailed predictions of objects and their parts for both familiar and unfamiliar categories.
Ian Endres, Vivek Srikumar, Ming-Wei Chang, Derek Hoiem
CVPR2
2012 On Amortizing Inference Cost for Structured Prediction
Vivek Srikumar, Gourab Kundu, Dan Roth 0001
EMNLP-CoNLL1
2012 An NLP Curator (or: How I Learned to Stop Worrying and Love NLP Pipelines)
James Clarke, Vivek Srikumar, Mark Sammons, Dan Roth 0001
LREC2
2012 Predicting Structures in NLP: Constrained Conditional Models and Integer Linear Programming in NLP
Dan Goldwasser, Vivek Srikumar, Dan Roth 0001
HLT-NAACL2
2011 A Joint Model for Extended Semantic Role Labeling
Vivek Srikumar, Dan Roth 0001
EMNLP1
2010 Structured Output Learning with Indirect Supervision
Ming-Wei Chang, Vivek Srikumar, Dan Goldwasser, Dan Roth 0001
ICML2
2010 Discriminative Learning over Constrained Latent Representations
Ming-Wei Chang, Dan Goldwasser, Dan Roth 0001, Vivek Srikumar
HLT-NAACL4
2008 Importance of Semantic Representation: Dataless Classification
Ming-Wei Chang, Lev-Arie Ratinov, Dan Roth 0001, Vivek Srikumar
AAAI4
2008 Proactive Intrusion Detection
Benjamin Liebald, Dan Roth 0001, Neelay Shah, Vivek Srikumar
AAAI4
2008 Extraction of Entailed Semantic Relations Through Syntax-Based Comma Resolution
Vivek Srikumar, Roi Reichart, Mark Sammons, Ari Rappoport, Dan Roth 0001
ACL1