EDBT 2026 Demo / reviewers in the wild / expert
Nikhil Krishnaswamy
dblp:184/8937
· DBLP profile ↗
31ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-7878-7227ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 6 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Now They See It, Now They Don't: Multimodal Reward Models Exhibit Unreliability in Physical World ConstraintsabstractGenerative AI systems, especially those driven by autoregressive and diffusion-based models, are known to struggle with spatial reasoning. As such, it becomes critical to understand how humans regard those failure modes. In this paper, we examine how humans judge different types of errors in images generated by a text-to-image model. We curated prompts that described common household objects with variance in number, spatial relations, and orientations, and generated a variety of images using each prompt. Humans observed pairs of images generated using the same prompt and answered a set of systematic questions about each image. Survey results showed that incorrect spatial orientation regularly emerges as a reason that the generated images do not accurately represent the prompt. We further investigated how RLHF-based multimodal reward models score prompt-image alignment over the same data, and whether they can reliably distinguish the better image in a pairwise setting, as humans do. We find that even though a general cross-task reward model may output alignment scores that accord with those of humans, its reasoning traces are flawed with respect to spatial orientational and relational indicators—the very factors that human annotators rated as the most consequential errors in generated images. Our results show that human annotators regard spatial reasoning errors as highly impactful on the correctness of generated images, and undermine the reliability of multimodal reward model scores as a baseline for evaluating image quality. Sadaf Ghaffari, Nikhil Krishnaswamy |
CoNLL | 2 |
| 2026 | Identifying Contexts of Distress in College Students' Reddit Posts: A Comparative Study of Classical NLP and Large Language Models
Carine Graff, Nikhil Krishnaswamy |
LREC | 2 |
| 2026 | Distributed Partial Information Puzzles: Examining Common Ground Construction under Epistemic Asymmetry
Yifan Zhu 0014, Mariah Bradford, Kenneth Lai, Timothy Obiso, Videep Venkatesha, James Pustejovsky, Nikhil Krishnaswamy |
LREC | 7 |
| 2025 | Bidirectional Human-AI Learning in Real-Time Disoriented BalancingabstractWe present a real-time system that enables bidirectional human-AI learning and teaching in a balancing task that is a realistic analogue of disorientation during piloting and spaceflight. A human subject and autonomous AI model of choice guide each other in maintaining balance using a visual inverted pendulum (VIP) display. We show how AI assistance changes human performance and vice versa. Sheikh Mannan, Nikhil Krishnaswamy |
AAAI | 2 |
| 2025 | Speech Is Not Enough: Interpreting Nonverbal Indicators of Common Knowledge and EngagementabstractOur goal is to develop an AI Partner that can provide support for group problem solving and social dynamics. In multi-party working group environments, multimodal analytics is crucial for identifying non-verbal interactions of group members. In conjunction with their verbal participation, this creates an holistic understanding of collaboration and engagement that provides necessary context for the AI Partner. In this demo, we illustrate our present capabilities at detecting and tracking nonverbal behavior in student task-oriented interactions in the classroom, and the implications for tracking common ground and engagement. Derek Palmer, Yifan Zhu 0014, Kenneth Lai, Hannah VanderHoeven, Mariah Bradford, Ibrahim Khebour, Carlos Mabrey, Jack Fitzgerald, Nikhil Krishnaswamy, Martha Palmer, James Pustejovsky |
AAAI | 9 |
| 2025 | Frictional Agent Alignment Framework: Slow Down and Don't Break ThingsabstractAI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs.Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the explicit signals of interlocutor beliefs are sparse and skewed.We propose the Frictional Agent Alignment Framework (FAAF), to generate precise, context-aware "friction" that prompts for deliberation and re-examination of existing evidence.FAAF's two-player objective decouples from data skew: a frictive-state policy identifies belief misalignments, while an intervention policy crafts collaborator-preferred responses.We derive an analytical solution to this objective, enabling training a single policy via a simple supervised loss.Experiments on three benchmarks show FAAF outperforms competitors in producing concise, interpretable friction and in OOD generalization.By aligning LLMs to act as adaptive "thought partners"-not passive responders-FAAF advances scalable, dynamic human-AI collaboration.Our code and data can be found at https: //github.com/csu-signal/FAAF_ACL. Abhijnan Nath, Carine Graff, Andrei Bachinin, Nikhil Krishnaswamy |
ACL (1) | 4 |
| 2025 | The Impact of Background Speech on Interruption Detection in Collaborative Groups
Mariah Bradford, Nikhil Krishnaswamy, Nathaniel Blanchard |
AIED (3) | 2 |
| 2025 | HuTCH: Human Teachable Concept Highlighter for Post-hoc Visual Explanations
Erfan Mirhaji, Nikhil Krishnaswamy, Jill Zarestky, Lisa Mason, Sarath Sreedharan, Nathaniel Blanchard |
AIED (5) | 2 |
| 2025 | DPL: Diverse Preference Learning Without A Reference ModelabstractAbhijnan Nath, Andrey Volozin, Saumajit Saha, Albert Aristotle Nanda, Galina Grunin, Rahul Bhotika, Nikhil Krishnaswamy. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Abhijnan Nath, Andrey Volozin, Saumajit Saha, Albert Nanda, Galina Grunin, Rahul Bhotika, Nikhil Krishnaswamy |
NAACL (Long Papers) | 7 |
| 2025 | Learning "Partner-Aware" Collaborators in Multi-Party CollaborationabstractLarge Language Models (LLMs) are increasingly being deployed in agentic settings where they act as collaborators with humans. Therefore, it is increasingly important to be able to evaluate their abilities to collaborate effectively in multi-turn, multi-party tasks. In this paper, we build on the AI alignment and “safe interruptability” literature to offer novel theoretical insights on collaborative behavior between LLM-driven *collaborator agents* and an *intervention agent*. Our goal is to learn an ideal “partner-aware” collaborator that increases the group’s common-ground (CG)—alignment on task-relevant propositions—by intelligently collecting information provided in *interventions* by a partner agent. We show how LLM agents trained using standard RLHF and related approaches are naturally inclined to ignore possibly well-meaning interventions, which makes increasing group common ground non-trivial in this setting. We employ a two-player Modified-Action MDP to examine this suboptimal behavior of standard AI agents, and propose **Interruptible Collaborative Roleplayer (ICR)**—a novel “partner-aware” learning algorithm to train CG-optimal collaborators. Experiments on multiple collaborative task environments show that ICR, on average, is more capable of promoting successful CG convergence and exploring more diverse solutions in such tasks. Abhijnan Nath, Nikhil Krishnaswamy |
NeurIPS | 2 |
| 2024 | Modeling the development of intuitive mechanics
Mengguo Jing, Zakir Makhani, Iris Oved, Nikhil Krishnaswamy, James Pustejovsky, Joshua K. Hartshorne |
CogSci | 5 |
| 2024 | Computational Thought Experiments for a More Rigorous Philosophy and Science of the Mind
Iris Oved, Nikhil Krishnaswamy, James Pustejovsky, Joshua K. Hartshorne |
CogSci | 2 |
| 2024 | Common Ground Tracking in Multimodal DialogueabstractWithin Dialogue Modeling research in AI and NLP, considerable attention has been spent on “dialogue state tracking” (DST), which is the ability to update the representations of the speaker’s needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is “common ground tracking” (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ”questions under discussion” (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task. Ibrahim Khebour, Kenneth Lai, Mariah Bradford, Yifan Zhu 0014, Richard Brutti, Christopher Tam, Jingxuan Tu, Benjamin Ibarra, Nathaniel Blanchard, Nikhil Krishnaswamy, James Pustejovsky |
LREC/COLING | 10 |
| 2024 | Cross-Lingual Transfer Robustness to Lower-Resource Languages on Adversarial DatasetsabstractMultilingual Language Models (MLLMs) exhibit robust cross-lingual transfer capabilities, or the ability to leverage information acquired in a source language and apply it to a target language. These capabilities find practical applications in well-established Natural Language Processing (NLP) tasks such as Named Entity Recognition (NER). This study aims to investigate the effectiveness of a source language when applied to a target language, particularly in the context of perturbing the input test set. We evaluate on 13 pairs of languages, each including one high-resource language (HRL) and one low-resource language (LRL) with a geographic, genetic, or borrowing relationship. We evaluate two well-known MLLMs—MBERT and XLM-R—on these pairs, in native LRL and cross-lingual transfer settings, in two tasks, under a set of different perturbations. Our findings indicate that NER cross-lingual transfer depends largely on the overlap of entity chunks. If a source and target language have more entities in common, the transfer ability is stronger. Models using cross-lingual transfer also appear to be somewhat more robust to certain perturbations of the input, perhaps indicating an ability to leverage stronger representations derived from the HRL. Our research provides valuable insights into cross-lingual transfer and its implications for NLP applications, and underscores the need to consider linguistic nuances and potential limitations when employing MLLMs across distinct languages. Shadi Manafi Avari, Nikhil Krishnaswamy |
LREC/COLING | 2 |
| 2024 | Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality EnsemblesabstractEvent coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facilitate resolution when language is ambiguous. Here, we propose a multimodal cross-document event coreference resolution method that integrates visual and textual cues with a simple linear map between vision and language models. As existing ECR benchmark datasets rarely provide images for all event mentions, we augment the popular ECB+ dataset with event-centric images scraped from the internet and generated using image diffusion models. We establish three methods that incorporate images and text for coreference: 1) a standard fused model with finetuning, 2) a novel linear mapping method without finetuning and 3) an ensembling approach based on splitting mention pairs by semantic and discourse-level difficulty. We evaluate on 2 datasets: the augmented ECB+, and AIDA Phase 1. Our ensemble systems using cross-modal linear mapping establish an upper limit (91.9 CoNLL F1) on ECB+ ECR performance given the preprocessing assumptions used, and establish a novel baseline on AIDA Phase 1. Our results demonstrate the utility of multimodal information in ECR for certain challenging coreference problems, and highlight a need for more multimodal resources in the coreference resolution space. Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Arthur Baker, Rahul Ghosh, James H. Martin, Nathaniel Blanchard, Nikhil Krishnaswamy |
LREC/COLING | 8 |
| 2024 | Propositional Extraction from Natural Speech in Small Group Collaborative Tasks
Videep Venkatesha, Abhijnan Nath, Ibrahim Khebour, Avyakta Chelle, Mariah Bradford, Jingxuan Tu, James Pustejovsky, Nathaniel Blanchard, Nikhil Krishnaswamy |
EDM | 9 |
| 2024 | Combating Spatial Disorientation in a Dynamic Self-Stabilization Task Using AI AssistantsabstractSpatial disorientation is a leading cause of fatal aircraft accidents. This paper explores the potential of AI agents to aid pilots in maintaining balance and preventing unrecoverable losses of control by offering cues and corrective measures that ameliorate spatial disorientation. A multi-axis rotation system (MARS) was used to gather data from human subjects self-balancing in a spaceflight analog condition. We trained models over this data to create “digital twins” that exemplified performance characteristics of humans with different proficiency levels. We then trained various reinforcement learning and deep learning models to offer corrective cues if loss of control is predicted. Digital twins and assistant models then co-performed a virtual inverted pendulum (VIP) programmed with identical physics. From these simulations, we picked the 5 best-performing assistants based on task metrics such as crash frequency and mean distance from the direction of balance. These were used in a co-performance study with 20 new human subjects performing a version of the VIP task with degraded spatial information. We show that certain AI assistants were able to improve human performance and that reinforcement-learning based assistants were objectively more effective but rated as less trusted and preferable by humans. Sheikh Mannan, Paige Hansen, Vivekanand Pandey Vimal, Hannah N. Davies, Paul DiZio, Nikhil Krishnaswamy |
HAI | 6 |
| 2024 | Okay, Let's Do This! Modeling Event Coreference with Generated Rationales and Knowledge DistillationabstractAbhijnan Nath, Shadi Manafi Avari, Avyakta Chelle, Nikhil Krishnaswamy. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Abhijnan Nath, Shadi Manafi Avari, Avyakta Chelle, Nikhil Krishnaswamy |
NAACL-HLT | 4 |
| 2024 | Metacognitive AI: Framework and the Case for a Neurosymbolic Approach
Hua Wei 0001, Paulo Shakarian, Christian Lebiere, Bruce A. Draper, Nikhil Krishnaswamy, Sergei Nirenburg |
NeSy (2) | 5 |
| 2023 | Automatic Detection of Collaborative States in Small Groups Using Multimodal Features
Mariah Bradford, Ibrahim Khebour, Nathaniel Blanchard, Nikhil Krishnaswamy |
AIED | 4 |
| 2022 | Exploiting Embodied Simulation to Detect Novel Object Classes Through Interaction
Nikhil Krishnaswamy, Sadaf Ghaffari |
CogSci | 1 |
| 2022 | A Generalized Method for Automated Multilingual Loanword DetectionabstractLoanwords are words incorporated from one language into another without translation. Suppose two words from distantly-related or unrelated languages sound similar and have a similar meaning. In that case, this is evidence of likely borrowing. This paper presents a method to automatically detect loanwords across various language pairs, accounting for differences in script, pronunciation and phonetic transformation by the borrowing language. We incorporate edit distance, semantic similarity measures, and phonetic alignment. We evaluate on 12 language pairs and achieve performance comparable to or exceeding state of the art methods on single-pair loanword detection tasks. We also demonstrate that multilingual models perform the same or often better than models trained on single language pairs and can potentially generalize to unseen language pairs with sufficient data, and that our method can exceed human performance on loanword detection. Abhijnan Nath, Sina Mahdipour Saravani, Ibrahim Khebour, Sheikh Mannan, Zihui Li, Nikhil Krishnaswamy |
COLING | 6 |
| 2022 | A deep dive into microphone hardware for recording collaborative group work
Mariah Bradford, Paige Hansen, J. Ross Beveridge, Nikhil Krishnaswamy, Nathaniel Blanchard |
EDM | 4 |
| 2022 | The VoxWorld Platform for Multimodal Embodied AgentsabstractWe present a five-year retrospective on the development of the VoxWorld platform, first introduced as a multimodal platform for modeling motion language, that has evolved into a platform for rapidly building and deploying embodied agents with contextual and situational awareness, capable of interacting with humans in multiple modalities, and exploring their environments. In particular, we discuss the evolution from the theoretical underpinnings of the VoxML modeling language to a platform that accommodates both neural and symbolic inputs to build agents capable of multimodal interaction and hybrid reasoning. We focus on three distinct agent implementations and the functionality needed to accommodate all of them: Diana, a virtual collaborative agent; Kirby, a mobile robot; and BabyBAW, an agent who self-guides its own exploration of the world. Nikhil Krishnaswamy, William Pickard, Brittany Cates, Nathaniel Blanchard, James Pustejovsky |
LREC | 1 |
| 2020 | Diana's World: A Situated Multimodal Interactive AgentabstractState of the art unimodal dialogue agents lack some core aspects of peer-to-peer communication—the nonverbal and visual cues that are a fundamental aspect of human interaction. To facilitate true peer-to-peer communication with a computer, we present Diana, a situated multimodal agent who exists in a mixed-reality environment with a human interlocutor, is situation- and context-aware, and responds to the human's language, gesture, and affect to complete collaborative tasks. Nikhil Krishnaswamy, Pradyumna Narayana, Rahul Bangar, Kyeongmin Rim, Dhruva Patil, David G. McNeely-White, Jaime Ruiz 0002, Bruce A. Draper, J. Ross Beveridge, James Pustejovsky |
AAAI | 1 |
| 2020 | Embodied Human-Computer Interactions through Situated GroundingabstractIn this paper, we introduce a simulation platform for modeling and building Embodied Human-Computer Interactions (EHCI). This system, VoxWorld, is a multimodal dialogue system enabling communication through language, gesture, action, facial expressions, and gaze tracking, in the context of task-oriented interactions. A multimodal simulation is an embodied 3D virtual realization of both the situational environment and the co-situated agents, as well as the most salient content denoted by communicative acts in a discourse. It is built on the modeling language VoxML [7], which encodes objects with rich semantic typing and action affordances, and actions themselves as multimodal programs, enabling contextu-ally salient inferences and decisions in the environment. VoxWorld enables an embodied HCI by situating both human and computational agents within the same virtual simulation environment, where they share perceptual and epistemic common ground. James Pustejovsky, Nikhil Krishnaswamy |
IVA | 2 |
| 2020 | A Formal Analysis of Multimodal Referring Strategies Under Common GroundabstractIn this paper, we present an analysis of computationally generated mixed-modality definite referring expressions using combinations of gesture and linguistic descriptions. In doing so, we expose some striking formal semantic properties of the interactions between gesture and language, conditioned on the introduction of content into the common ground between the (computational) speaker and (human) viewer, and demonstrate how these formal features can contribute to training better models to predict viewer judgment of referring expressions, and potentially to the generation of more natural and informative referring expressions. Nikhil Krishnaswamy, James Pustejovsky |
LREC | 1 |
| 2019 | Combining Deep Learning and Qualitative Spatial Reasoning to Learn Complex Structures from Sparse Examples with NoiseabstractMany modern machine learning approaches require vast amounts of training data to learn new concepts; conversely, human learning often requires few examples—sometimes only one—from which the learner can abstract structural concepts. We present a novel approach to introducing new spatial structures to an AI agent, combining deep learning over qualitative spatial relations with various heuristic search algorithms. The agent extracts spatial relations from a sparse set of noisy examples of block-based structures, and trains convolutional and sequential models of those relation sets. To create novel examples of similar structures, the agent begins placing blocks on a virtual table, uses a CNN to predict the most similar complete example structure after each placement, an LSTM to predict the most likely set of remaining moves needed to complete it, and recommends one using heuristic search. We verify that the agent learned the concept by observing its virtual block-building activities, wherein it ranks each potential subsequent action toward building its learned concept. We empirically assess this approach with human participants’ ratings of the block structures. Initial results and qualitative evaluations of structures generated by the trained agent show where it has generalized concepts from the training data, which heuristics perform best within the search space, and how we might improve learning and execution. Nikhil Krishnaswamy, Scott Friedman 0001, James Pustejovsky |
AAAI | 1 |
| 2018 | An Evaluation Framework for Multimodal Interaction
Nikhil Krishnaswamy, James Pustejovsky |
LREC | 1 |
| 2016 | Visualizing Events: Simulating Meaning in Language
James Pustejovsky, Nikhil Krishnaswamy |
CogSci | 2 |
| 2016 | VoxML: A Visualization Modeling Language
James Pustejovsky, Nikhil Krishnaswamy |
LREC | 2 |