Kanishka Misra

dblp:46/8138 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cross-Modal Taxonomic Generalization in (Vision-) Language Models
abstract
Tianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianyang Xu 0002, Marcelo Sandoval-Castañeda, Karen Livescu, Gregory Shakhnarovich, Kanishka Misra
ACL (1)5
2026 Language Models Learn Constructional Semantics, Not To Mention Syntax: Investigating LM Understanding of Paired-Focus Constructions
abstract
Grasping the semantics of rare constructions (form-meaning pairings) has been shown to be a challenging problem that has currently only been solved by the largest LLMs.It remains an open question if open-source models have robust constructional understanding, and if so, what learning dynamics underlie the acquisition of this knowledge.Focusing on a set of rare PAIRED-FOCUS constructions in English (e.g."let alone", "much less"), we construct a novel dataset to test their meanings using both scalar adjectival semantics and general world knowledge.Testing a wide range of models differing in parameter count, architecture, and pretraining dataset size, we find that several modestly sized models are sensitive to both the forms and the meanings of PAIRED-FOCUS constructions, though models trained on human-scale data fail at all meaning evaluations.Turning to training dynamics for a set of open-checkpoint models, we find that PAIRED-FOCUS understanding emerges later in training than PAIRED-FOCUS syntactic knowledge, and that learning of PAIRED-FOCUS semantics is correlated with gains in some domains of world knowledge.Overall, our empirical results support the conclusion that modestly sized open-source models can grasp the rare PAIRED-FOCUS constructions, and demonstrate a connection between knowledge of PAIRED-FOCUS constructions and other meaning domains.
Wesley Scivetti, Ethan Wilcox, Nathan Schneider 0001, Kanishka Misra, Leonie Weissweiler
CoNLL4
2025 Characterizing the Role of Similarity in the Property Inferences of Language Models
abstract
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Juan Diego Rodriguez, Aaron Mueller, Kanishka Misra
NAACL (Long Papers)3
2025 Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
abstract
Does vision-and-language (VL) training change the linguistic representations of language models in meaningful ways? In terms of downstream task performance on text-only tasks, most results in the literature have shown marginal differences. In this work, we start from the hypothesis that the domain in which VL training could have a significant effect is lexical-conceptual knowledge, in particular its taxonomic organization. Through comparing minimal pairs of text-only LMs and their VL-trained counterparts, we first show that the VL models often outperform their text-only counterparts on a text-only question-answering task that requires taxonomic understanding of concepts mentioned in the questions. Using an array of targeted behavioral and representational analyses, we show that the LMs and VLMs do not differ significantly in terms of their taxonomic knowledge itself, but they differ in how they represent questions that contain concepts in a taxonomic relation vs. a non-taxonomic relation. This implies that the taxonomic knowledge itself does not change substantially through additional VL training, but VL training does improve the deployment of this knowledge in the context of a specific task, even when the presentation of the task is purely linguistic.
Yulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli, Kanishka Misra, Najoung Kim
NeurIPS5
2024 Experimental Contexts Can Facilitate Robust Semantic Property Inference in Language Models, but Inconsistently
abstract
Recent zero-shot evaluations have highlighted important limitations in the abilities of language models (LMs) to perform meaning extraction.However, it is now well known that LMs can demonstrate radical improvements in the presence of experimental contexts such as in-context examples and instructions.How well does this translate to previously studied meaning-sensitive tasks?We present a casestudy on the extent to which experimental contexts can improve LMs' robustness in performing property inheritance-predicting semantic properties of novel concepts, a task that they have been previously shown to fail on.Upon carefully controlling the nature of the in-context examples and the instructions, our work reveals that they can indeed lead to nontrivial property inheritance behavior in LMs.However, this ability is inconsistent: with a minimal reformulation of the task, some LMs were found to pick up on shallow, non-semantic heuristics from their inputs, suggesting that the computational principles of semantic property inference are yet to be mastered by LMs.
Kanishka Misra, Allyson Ettinger, Kyle Mahowald
EMNLP1
2024 Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
abstract
Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question.To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction ("a beautiful five days").We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed.We found that AANNs were still learned better than systematically perturbed variants of the construction.Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., "a few days").An additional experiment showed that this learning is enhanced when there is more variability in the input.Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena.
Kanishka Misra, Kyle Mahowald
EMNLP1
2023 Language model acceptability judgements are not always robust to context
abstract
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Koustuv Sinha, Jon Gauthier, Aaron Mueller, Kanishka Misra, Keren Fuentes, Roger Levy, Adina Williams
ACL (1)4
2023 COMPS: Conceptual Minimal Pair Sentences for testing Robust Property Knowledge and its Inheritance in Pre-trained Language Models
abstract
A characteristic feature of human semantic cognition is its ability to not only store and retrieve the properties of concepts observed through experience, but to also facilitate the inheritance of properties (can breathe) from superordinate concepts (ANIMAL) to their subordinates (DOG)-i.e.demonstrate property inheritance.In this paper, we present COMPS, a collection of English minimal pair sentences that jointly tests pre-trained language models (PLMs) on their ability to attribute properties to concepts and their ability to demonstrate property inheritance behavior.Analyses of 22 different PLMs on COMPS reveal that they can easily distinguish between concepts on the basis of a property when they are trivially different, but find it relatively difficult when concepts are related on the basis of nuanced knowledge representations.Furthermore, we find that PLMs can show behaviors suggesting successful property inheritance in simple contexts, but fail in the presence of distracting information, which decreases the performance of many models sometimes even below chance.This lack of robustness in demonstrating simple reasoning raises important questions about PLMs' capacity to make correct inferences even when they appear to possess the prerequisite knowledge.
Kanishka Misra, Julia Taylor Rayz, Allyson Ettinger
EACL1
2023 Large Language Models Can Be Easily Distracted by Irrelevant Context
abstract
Large language models have achieved impressive performance on various natural language processing tasks. However, so far they have been evaluated primarily on benchmarks where all information in the input context is relevant for solving the task. In this work, we investigate the *distractibility* of large language models, i.e., how the model prediction can be distracted by irrelevant context. In particular, we introduce Grade-School Math with Irrelevant Context (GSM-IC), an arithmetic reasoning dataset with irrelevant information in the problem description. We use this benchmark to measure the distractibility of different prompting techniques for large language models, and find that the model is easily distracted by irrelevant information. We also identify several approaches for mitigating this deficiency, such as decoding with self-consistency and adding to the prompt an instruction that tells the language model to ignore the irrelevant information.
Freda Shi, Kanishka Misra, Nathan Scales, David Dohan, Ed H. Chi, Nathanael Schärli, Denny Zhou
ICML3
2022 On Semantic Cognition, Inductive Generalization, and Language Models
abstract
My doctoral research focuses on understanding semantic knowledge in neural network models trained solely to predict natural language (referred to as language models, or LMs), by drawing on insights from the study of concepts and categories grounded in cognitive science. I propose a framework inspired by 'inductive reasoning,' a phenomenon that sheds light on how humans utilize background knowledge to make inductive leaps and generalize from new pieces of information about concepts and their properties. Drawing from experiments that study inductive reasoning, I propose to analyze semantic inductive generalization in LMs using phenomena observed in human-induction literature, investigate inductive behavior on tasks such as implicit reasoning and emergent feature recognition, and analyze and relate induction dynamics to the learned conceptual representation space.
Kanishka Misra
AAAI1
2022 A Property Induction Framework for Neural Language Models
Kanishka Misra, Julia Taylor Rayz, Allyson Ettinger
CogSci1
2021 Do language models learn typicality judgments from text?
Kanishka Misra, Allyson Ettinger, Julia Taylor Rayz
CogSci1
2020 Exploring Lexical Relations in BERT using Semantic Priming
Kanishka Misra, Allyson Ettinger, Julia Taylor Rayz
CogSci1
2019 L1 Influence on Content Word errors in Learner English Corpora: Insights from Distributed Representation of Words
Kanishka Misra, Hemanth Devarapalli, Julia Taylor Rayz
CogSci1
2019 Authorship Analysis of Online Predatory Conversations using Character Level Convolution Neural Networks
abstract
Authorship Attribution (AA) of written content presents several advantages within the digital forensics domain. While AA has been traditionally applied to long documents, recent works have shown improved performance of neural AA models on short texts such as tweets and online conversations Concurrently, the rise of social media as well as a plethora of chat messaging platforms have made it easier for teenagers to be vulnerable to online predators. In this work, we present an authorship attribution model that trains on a corpus of online conversations involving predators, and perform subsequent analysis of the message representations. Our results show comparable performance relative to prior work for Authorship Attribution and highlight differences between predatory and non-predatory message styles.
Kanishka Misra, Hemanth Devarapalli, Tatiana R. Ringenberg, Julia Taylor Rayz
SMC1
2019 Not So Cute but Fuzzy: Estimating Risk of Sexual Predation in Online Conversations
abstract
The sexual exploitation of minors is a known and persistent problem for law enforcement. Assistance in prioritizing cases of sexual exploitation of potentially risky conversations is crucial. While attempts to automatically triage conversations for the risk of sexual exploitation of minors have been attempted in the past, most computational models use features which are not representative of the grooming process that is used by investigators. Accurately annotating an offender corpus for use with machine learning algorithms is difficult because the stages of the grooming process feed into one another and are non-linear. In this paper we propose a method for labeling risk, tied to stages and themes of the grooming process, using fuzzy sets. We develop a neural network model that uses these fuzzy membership functions of each line in a chat as input and predict the risk of interaction.
Tatiana R. Ringenberg, Kanishka Misra, Julia Taylor Rayz
SMC2
2019 A Sentiment Based Non-Factoid Question-Answering Framework
abstract
With the rapid advances in Artificial Intelligence, a question of emotional intelligence of a system may become as important as its accuracy. This paper investigates whether emotions should be considered for non-factoid “how” Question-Answering systems with the eventual goal of enabling the system to retrieve answers in a more emotionally intelligent way. This study proposes an architecture that adds extended representation of sentiment information to questions and answers, and reports on to what extent a prediction of the best answer be improved by the proposed architecture.
Qiaofei Ye, Kanishka Misra, Hemanth Devarapalli, Julia Taylor Rayz
SMC2