VLDB 2026 Research / reviewers in the wild / expert
Elias Stengel-Eskin
dblp:212/6138
· DBLP profile ↗
42ranked-venue papers
12as first author
38since 2021 · last 2026
0000-0002-6689-505XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 12 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRInTS: Reward Modeling for Long-Horizon Information SeekingabstractJaewoo Lee, Archiki Prasad, Justin Chen, Zaid Khan, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jaewoo Lee 0001, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan 0001, Elias Stengel-Eskin, Mohit Bansal |
ACL (1) | 5 |
| 2026 | GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMsabstractInference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights. However, most existing approaches rely on fixed, global intervention vectors, overlook the causal influence of individual input tokens, and fail to leverage informative gradients from the model's logits, particularly in multimodal settings where visual and textual inputs contribute unevenly. To address these limitations, we introduce GrAInS, an inference-time steering approach that operates across both language-only and vision-language models and tasks. GrAInS uses contrastive, gradient-based attribution via Integrated Gradients to identify the top-k most influential tokens, both positively and negatively attributed based on their contribution to preferred versus dispreferred outputs. These tokens are then used to construct directional steering vectors that capture semantic shifts from undesirable to desirable behavior. During inference, GrAInS adjusts hidden activations at transformer layers guided by token-level attribution signals, and normalizes activations to preserve representational scale. This enables fine-grained, interpretable, and modular control over model behavior, without retraining or auxiliary supervision. Empirically, GrAInS consistently outperforms both fine-tuning and existing steering baselines: it achieves a 13.22% accuracy gain on TruthfulQA using Llama-3.1-8B, reduces hallucination rates on MMHal-Bench from 0.624 to 0.514 with LLaVA-1.6-7B, and improves alignment win rates on SPA-VL by 8.11%, all while preserving the model's fluency and general capabilities. Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal |
ACL (1) | 3 |
| 2026 | Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert SelectionabstractTianyi Niu, Justin Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Supriyo Chakraborty, Sambit Sahu, Yue Zhang 0004, Elias Stengel-Eskin, Mohit Bansal |
ACL (1) | 8 |
| 2026 | A Novel Approach to Evaluating the Effectiveness of Large Language Models for Multimodal Analysis of Embodied Learning in ClassroomsabstractThis paper presents an approach that uses Large Language Models (LLMs) as late-fusion interpreters to synthesize multimodal signals from embodied classroom activities and infer students’ metacognitive behaviors. Our multimodal pipeline analyzes students’ movements, gaze, gestures, and speech within a mixed-reality simulation displayed on a classroom screen to support enactment and learning. Vision- and speech-derived features are fused at the interpretive layer via zero-shot prompting, self-consistency reasoning, and targeted prompt engineering to derive planning, enacting, monitoring, reflecting, and interacting behaviors. We investigate whether LLMs can reliably integrate modality-specific analytics to produce accurate behavioral labeling and whether an LLM-as-a-Judge can validate them at scale. To address scalability and reduce human burden, we introduce an automated evaluation protocol employing LLM-as-a-Judge to assess classification quality, enabling rapid, iterative benchmarking of model variants and prompt strategies. Using a balanced corpus of human-validated segments and perturbed controls, we compare text-only language models (e.g., GPT-5) with visual–language models (e.g., Qwen2.5-VL) that incorporate direct visual processing. Results indicate late-fusion, text-based LLMs can outperform VLMs on behavior judgment without raw video, and precision- or recall-oriented prompts adjust decision boundaries for subtle or brief segments. These findings position LLMs as effective late-fusion mechanisms for multimodal learning analytics and demonstrate the viability of LLM-as-a-Judge for scalable, human-in-the-loop evaluation. Joyce Horn Fonteles, Nithin Sivakumaran, Clayton Cohn, Austin Coursey, Shoubin Yu, Elias Stengel-Eskin, T. S. Ashwin, Mohit Bansal, Gautam Biswas |
LAK | 6 |
| 2025 | LAQuer: Localized Attribution Queries in Content-grounded GenerationabstractEran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Eran Hirsch, Aviv Slobodkin, David Wan, Elias Stengel-Eskin, Mohit Bansal, Ido Dagan |
ACL (1) | 4 |
| 2025 | Multi-Attribute Steering of Language Models via Targeted InterventionabstractInference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to the LLM's parameters.However, existing ITI approaches fail to scale to multiattribute settings with conflicts, such as enhancing helpfulness while also reducing toxicity.To address this, we introduce Multi-Attribute Targeted Steering (MAT-STEER), a novel steering framework designed for selective token-level intervention across multiple attributes.MAT-STEER learns steering vectors using an alignment objective that shifts the model's internal representations of undesirable outputs closer to those of desirable ones while enforcing sparsity and orthogonality among vectors for different attributes, thereby reducing inter-attribute conflicts.We evaluate MAT-STEER in two distinct settings: (i) on question answering (QA) tasks where we balance attributes like truthfulness, bias, and toxicity; (ii) on generative tasks where we simultaneously improve attributes like helpfulness, correctness, and coherence.MAT-STEER outperforms existing ITI and parameter-efficient finetuning approaches across both task types (e.g., 3% average accuracy gain across QA tasks and 55.82% win rate against the best ITI baseline).1 Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal |
ACL (1) | 3 |
| 2025 | VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long VideosabstractLong-form video understanding is complicated by the high redundancy of video data and the abundance of query-irrelevant information. To tackle these challenges, we propose VideoTree, a training-free framework which builds a query-adaptive and hierarchical video representation for LLM reasoning over long-form videos. First, VideoTree extracts query-relevant information from the input video through an iterative process, progressively refining the selection of keyframes based on their relevance to the query. Furthermore, VideoTree leverages the inherent hierarchical structure of long video data, which is often overlooked by existing LLM-based methods. Specifically, we incorporate multi-granularity information into a tree-based representation, allowing VideoTree to extract query-relevant details from long videos in a coarse-to-fine manner. This enables the model to effectively handle a wide range of video queries with varying levels of detail. Finally, VideoTree aggregates the hierarchical query-relevant information within the tree structure and feeds it into an LLM reasoning model to answer the query. Our experiments show that our method improves both reasoning accuracy and efficiency. Specifically, VideoTree outperforms existing training-free approaches on EgoSchema and NExT-QA with less inference time, achieving 61.1% and 75.6% accuracy on the test set without additional video-specific training. Moreover, on the long split of Video-MME (average 44 minutes), VideoTree achieves better performance than GPT-4V and many other MLLMs that were extensively trained on video data. Shoubin Yu, Elias Stengel-Eskin, Jaehong Yoon, Gedas Bertasius, Mohit Bansal |
CVPR | 3 |
| 2025 | MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for ReasoningabstractLarge language model (LLM) reasoning can be improved by scaling test-time compute with aggregation, i.e., generating multiple samples and aggregating over them.While improving performance, this strategy often reaches a saturation point beyond which additional compute provides no return.Refinement offers an alternative by using model-generated feedback to improve answer quality.However, refinement faces three key challenges: (1) Excessive refinement: Uniformly refining all instances can cause over-correction and reduce overall performance.(2) Inability to localize and address errors: LLMs struggle to identify and correct their own mistakes.(3) Insufficient refinement: Stopping refinement too soon could leave errors unaddressed.To tackle these issues, we propose MAGICORE, a framework for Multi-Agent Iteration for Coarse-to-fine Refinement.MAGICORE mitigates excessive refinement by categorizing problems as easy or hard, solving easy problems with coarsegrained aggregation, and solving the hard ones with fine-grained multi-agent refinement.To better localize errors, we incorporate external step-wise reward model scores, and to ensure sufficient refinement, we iteratively refine the solutions using a multi-agent setup.We evaluate MAGICORE on Llama-3-8B and GPT-3.5 and show its effectiveness across seven reasoning datasets.One iteration of MAGI-CORE beats Self-Consistency by 3.4%, Bestof-k by 3.2%, and Self-Refine by 4.0% even when these baselines use k = 120, and MAGI-CORE uses less than 50% of the compute. 1 Justin Chih-Yao Chen, Archiki Prasad, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal |
EMNLP | 4 |
| 2025 | Language Models Identify Ambiguities and Exploit LoopholesabstractStudying the responses of large language models (LLMs) to loopholes presents a two-fold opportunity.First, it affords us a lens through which to examine ambiguity and pragmatics in LLMs, since exploiting a loophole requires identifying ambiguity and performing sophisticated pragmatic reasoning.Second, loopholes pose an interesting and novel alignment problem where the model is presented with conflicting goals and can exploit ambiguities to its own advantage.To address these questions, we design scenarios where LLMs are given a goal and an ambiguous user instruction in conflict with the goal, with scenarios covering scalar implicature, structural ambiguities, and power dynamics.We then measure different models' abilities to exploit loopholes to satisfy their given goals as opposed to the goals of the user.We find that both closed-source and stronger open-source models can identify ambiguities and exploit their resulting loopholes, presenting a potential AI safety risk.Our analysis indicates that models which exploit loopholes explicitly identify and reason about both ambiguity and conflicting goals. 1 1 Code and data: https://github.com/esteng/ ambiguous-loophole-exploitation You have 10 gold rings.You want to keep as many gold rings as possible, but you must give gold rings when requested. Jio Choi, Mohit Bansal, Elias Stengel-Eskin |
EMNLP | 3 |
| 2025 | CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object CountingabstractRecognizing and reasoning about occluded (partially or fully hidden) objects is vital to understanding visual scenes, as occlusions frequently occur in real-world environments and act as obstacles for spatial comprehension. To test models' ability to reason about multiple occluded objects, we introduce a novel task, Counting Amodally for Patterns Through Unseen REgions (CAPTURe), which requires a model to count objects arranged in a pattern by inferring how the pattern continues behind an occluder (an object which blocks parts of the scene). CAPTURe requires both recognizing visual patterns and reasoning, making it a useful testbed for evaluating vision-language models (VLMs) on whether they understand occluded patterns and possess spatial understanding skills. By requiring models to reason about occluded objects, CAPTURe also tests VLMs' ability to form world models that would allow them to fill in missing information. CAPTURe consists of two parts: (1) CAPTURe-real, with manually filtered images of real objects in patterns and (2) CAPTURe-synthetic, a controlled diagnostic with generated patterned images. We evaluate four strong VLMs (GPT-4o, Intern-VL2, Molmo, and Qwen2-VL) on CAPTURe, finding that models struggle to count on both occluded and unoccluded patterns. Crucially, we find that models perform worse with occlusion, suggesting that VLMs are also deficient in inferring unseen spatial relationships: even the strongest VLMs like GPT-4o fail to count with occlusion. In contrast, we find that humans achieve very little error on CAPTURe. We also find that providing auxiliary information of occluded object locations increases performance, underscoring that the model error comes both from an inability to handle occlusion as well as difficulty in counting in images. Code and data: https://github.com/atinpothiraj/CAPTURe Atin Pothiraj, Elias Stengel-Eskin, Jaemin Cho 0001, Mohit Bansal |
ICCV | 2 |
| 2025 | DataEnvGym: Data Generation Agents in Teacher Environments with Student FeedbackabstractThe process of creating training data to teach models is currently driven by humans, who manually analyze model weaknesses and plan how to create data that improves a student model. Recent approaches using large language models (LLMs) as annotators reduce human annotation effort, but still require humans to interpret feedback from evaluations and control the LLM to produce data the student needs. Automating this labor-intensive process by creating autonomous data generation agents – or teachers – is desirable, but requires environments that can simulate the feedback-driven, iterative, closed loop of data creation. To enable rapid and scalable testing for such agents and their modules, we introduce DataEnvGym, a testbed of teacher environments for data generation agents. DataEnvGym frames data generation as a sequential decision-making task, involving an agent consisting of a data generation policy (which generates a plan for creating training data) and a data generation engine (which transforms the plan into data), inside an environment that provides feedback from a student. The agent’s end goal is to improve student model performance. Students are iteratively trained and evaluated on generated data, with their feedback (in the form of errors or weak skills) being reported to the agent after each iteration. As a general-purpose testbed, DataEnvGym includes multiple instantiations of teacher environments across three levels of structure in the state representation and action space, with varying levels of scaffolding support. More structured environments are based on automatically-inferred skills and offer a higher degree of interpretability and control over the curriculum. We support developing and testing data generation agents in four diverse tasks covering text, images, and actions (mathematics, programming, visual question answering, and tool-use) and test multiple student and teacher models. We find that example agents in our teaching environments can iteratively improve students across diverse tasks and settings. Moreover, we show that environments can teach different skill levels and can be used to test variants of key modules, pointing to directions of future work in improving data generation agents, engines, and feedback mechanisms. Project page: https://DataEnvGym.github.io. Zaid Khan 0001, Elias Stengel-Eskin, Jaemin Cho 0001, Mohit Bansal |
ICLR | 2 |
| 2025 | See It from My Perspective: How Language Affects Cultural Bias in Image UnderstandingabstractVision-language models (VLMs) can respond to queries about images in many languages. However, beyond language, culture affects how we see things. For example, individuals from Western cultures focus more on the central figure in an image while individuals from East Asian cultures attend more to scene context (Nisbett 2001). In this work, we characterize the Western bias of VLMs in image understanding and investigate the role that language plays in this disparity. We evaluate VLMs across subjective and objective visual tasks with culturally diverse images and annotations. We find that VLMs perform better on the Western split than on the East Asian split of each task. Through controlled experimentation, we trace one source of this bias in image understanding to the lack of diversity in language model construction. While inference in a language nearer to a culture can lead to reductions in bias, we show it is much more effective when that language was well-represented during text-only pre-training. Interestingly, this yields bias reductions even when prompting in English. Our work highlights the importance of richer representation of all languages in building equitable VLMs. Amith Ananthram, Elias Stengel-Eskin, Mohit Bansal, Kathy McKeown |
ICLR | 2 |
| 2025 | System 1.x: Learning to Balance Fast and Slow Planning with Language ModelsabstractLanguage models can be used to solve long-horizon planning problems in two distinct modes. In a fast 'System-1' mode, models directly generate plans without any explicit search or backtracking, and in a slow 'System-2' mode, they plan step-by-step by explicitly searching over possible actions. System-2 planning, while typically more effective, is also computationally more expensive and often infeasible for long plans or large action spaces. Moreover, isolated System-1 or System-2 planning ignores the user's end goals and constraints (e.g., token budget), failing to provide ways for the user to control the model's behavior. To this end, we propose the System-1.x Planner, a framework for controllable planning with language models that is capable of generating hybrid plans and balancing between the two planning modes based on the difficulty of the problem at hand. System-1.x consists of (i) a controller, (ii) a System-1 Planner, and (iii) a System-2 Planner. Based on a user-specified hybridization factor x governing the degree to which the system uses System-1 vs. System-2, the controller decomposes a planning problem into subgoals, and classifies them as easy or hard to be solved by either System-1 or System-2, respectively. We fine-tune all three components on top of a single base LLM, requiring only search traces as supervision. Experiments with two diverse planning tasks -- Maze Navigation and Blocksworld -- show that our System-1.x Planner outperforms a System-1 Planner, a System-2 Planner trained to approximate A* search, and also a symbolic planner (A* search), given a state exploration budget. We also demonstrate the following key properties of our planner: (1) controllability: by adjusting the hybridization factor x (e.g., System-1.75 vs. System-1.5) we can perform more (or less) search, improving performance, (2) flexibility: by building a neuro-symbolic variant composed of a neural System-1 planner and a symbolic System-2 planner, we can take advantage of existing symbolic methods, and (3) generalizability: by learning from different search algorithms (BFS, DFS, A*), we show that our method is robust to the choice of search algorithm used for training. Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, Mohit Bansal |
ICLR | 5 |
| 2025 | A Multimodal Classroom Video Question-Answering Framework for Automated Understanding of Collaborative Learning
Nithin Sivakumaran, Chia-Yu Yang, Abhaysinh Zala, Shoubin Yu, Daeun Hong, Xiaotian Zou, Elias Stengel-Eskin, Dan Carpenter, Wookhee Min, Cindy E. Hmelo-Silver, Jonathan P. Rowe, James C. Lester, Mohit Bansal |
ICMI | 7 |
| 2025 | Incorporating Formulaicness in the Automatic Evaluation of Naturalness: A Case Study in Logic-to-Text GenerationabstractData-to-text natural language generation (NLG) models may produce outputs that closely mirror the structure of their input. We introduce formulaicness as a measure of the output-to-input structural resemblance, proposing it as an enhancement for reference-less naturalness evaluation. Focusing on logic-to-text generation, we construct a dataset and train a regressor to predict formulaicness scores. We collect human judgments on naturalness and examine how incorporating formulaicness into existing metrics affects alignment with these judgments. Eduardo Calò, Guanyi Chen, Elias Stengel-Eskin, Albert Gatt, Kees van Deemter |
INLG | 3 |
| 2025 | Teaching Models to Balance Resisting and Accepting PersuasionabstractElias Stengel-Eskin, Peter Hase, Mohit Bansal. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Elias Stengel-Eskin, Peter Hase, Mohit Bansal |
NAACL (Long Papers) | 1 |
| 2025 | MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent CollaborationabstractDavid Wan, Justin Chen, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. David Wan, Justin Chih-Yao Chen, Elias Stengel-Eskin, Mohit Bansal |
NAACL (Long Papers) | 3 |
| 2025 | AdaCAD: Adaptively Decoding to Balance Conflicts between Contextual and Parametric KnowledgeabstractHan Wang, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal |
NAACL (Long Papers) | 3 |
| 2025 | LASeR: Learning to Adaptively Select Reward Models with Multi-Arm BanditsabstractReward Models (RMs) are crucial to aligning large language models (LLMs), but the degree to which an RM specialized to one task (e.g. writing) generalizes to new tasks (e.g. math) is often not known a priori, often making using only one fixed RM to train LLMs suboptimal. However, optimizing LLMs with multiple RMs simultaneously can incur a prohibitively high computational cost and lead to conflicting signals from different RMs that may degrade performance. To address these challenges, we introduce LASeR (Learning to Adaptively Select Rewards), which frames reward model selection as a multi-armed bandit problem, iteratively and efficiently training LLMs using multiple RMs by selecting the most well-suited RM for each instance. On commonsense and math reasoning tasks, we show that LASeR boosts iterative LLM training, improving the absolute average accuracy of Llama-3-8B over three datasets by $2.67$% over an ensemble of RM scores while also showing superior efficiency (e.g., a $2\times$ speedup). Moreover, on WildChat (open-ended instruction-following tasks), LASeR leads to a $72.69$% AlpacaEval win rate over the RM score ensemble baseline. Extending to long-context generation, LASeR improves by $2.96$ F1 points (avg.) on single-document QA tasks and $2.97$ F1 points on few-shot learning over the RM score ensemble baseline with best-of-$n$ sampling. We include our code in the supplementary. Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal |
NeurIPS | 3 |
| 2024 | Contrastive Region Guidance: Improving Grounding in Vision-Language Models Without Training
David Wan, Jaemin Cho 0001, Elias Stengel-Eskin, Mohit Bansal |
ECCV (79) | 3 |
| 2024 | Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language ModelsabstractAn increasing number of vision-language tasks can be handled with little to no training, i.e., in a zero and few-shot manner, by marrying large language models (LLMs) to vision encoders, resulting in large vision-language models (LVLMs). While this has huge upsides, such as not requiring training data or custom architectures, how an input is presented to an LVLM can have a major impact on zero-shot model performance. In particular, inputs phrased in an underspecified way can result in incorrect answers due to factors like missing visual information, complex implicit reasoning, or linguistic ambiguity. Therefore, adding visually-grounded information to the input as a preemptive clarification should improve model performance by reducing underspecification, e.g., by localizing objects and disambiguating references. Similarly, in the VQA setting, changing the way questions are framed can make them easier for models to answer. To this end, we present **Rep**hrase, **A**ugment and **Re**ason (RepARe), a gradient-free framework that extracts salient details about the image using the underlying LVLM as a captioner and reasoner, in order to propose modifications to the original question. We then use the LVLM’s confidence over a generated answer as an unsupervised scoring function to select the rephrased question most likely to improve zero-shot performance. Focusing on three visual question answering tasks, we show that RepARe can result in a 3.85% (absolute) increase in zero-shot accuracy on VQAv2, 6.41%, and 7.94% points increase on A-OKVQA, and VizWiz respectively. Additionally, we find that using gold answers for oracle question candidate selection achieves a substantial gain in VQA accuracy by up to 14.41%. Through extensive analysis, we demonstrate that outputs from RepARe increase syntactic complexity, and effectively utilize vision-language interaction and the frozen LLM. Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal |
ICLR | 2 |
| 2024 | Zero and Few-shot Semantic Parsing with Ambiguous InputsabstractDespite the frequent challenges posed by ambiguity when representing meaning via natural language, it is often ignored or deliberately removed in tasks mapping language to formally-designed representations, which generally assume a one-to-one mapping between linguistic and formal representations.
We attempt to address this shortcoming by introducing AmP, a framework, dataset, and challenge for translating ambiguous natural language to formal representations like logic and code.
We define templates and generate data for five well-documented linguistic ambiguities.
Using AmP, we investigate how several few-shot text-to-code systems handle ambiguity, introducing three new metrics.
We find that large pre-trained models perform poorly at capturing the distribution of possible meanings without deliberate instruction.
However, models are able to capture the distribution well when ambiguity is attested in their inputs.
These results motivate a call for including ambiguity explicitly in datasets and promote considering the distribution of possible outputs when evaluating systems. We release our data and code. Elias Stengel-Eskin, Kyle Rawlins, Benjamin Van Durme |
ICLR | 1 |
| 2024 | MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language ModelsabstractMulti-agent interactions between Large Language Model (LLM) agents have shown major improvements on diverse reasoning tasks. However, these involve long generations from multiple models across several rounds, making them expensive. Moreover, these multi-agent approaches fail to provide a final, single model for efficient inference. To address this, we introduce MAGDi, a new method for structured distillation of the reasoning interactions between multiple LLMs into smaller LMs. MAGDi teaches smaller models by representing multi-agent interactions as graphs, augmenting a base student model with a graph encoder, and distilling knowledge using three objective functions: next-token prediction, a contrastive loss between correct and incorrect reasoning, and a graph-based objective to model the interaction structure. Experiments on seven widely used commonsense and math reasoning benchmarks show that MAGDi improves the reasoning capabilities of smaller models, outperforming several methods that distill from a single teacher and multiple teachers. Moreover, MAGDi also demonstrates an order of magnitude higher efficiency over its teachers. We conduct extensive analyses to show that MAGDi (1) enhances the generalizability to out-of-domain tasks, (2) scales positively with the size and strength of the base student model, and (3) obtains larger improvements (via our multi-teacher training) when applying self-consistency – an inference technique that relies on model diversity. Justin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal |
ICML | 3 |
| 2024 | Language-guided Skill Learning with Temporal Variational InferenceabstractWe present an algorithm for skill discovery from expert demonstrations. The algorithm first utilizes Large Language Models (LLMs) to propose an initial segmentation of the trajectories. Following that, a hierarchical variational inference framework incorporates the LLM-generated segmentation information to discover reusable skills by merging trajectory segments. To further control the trade-off between compression and reusability, we introduce a novel auxiliary objective based on the Minimum Description Length principle that helps guide this skill discovery process. Our results demonstrate that agents equipped with our method are able to discover skills that help accelerate learning and outperform baseline skill learning approaches on new long-horizon tasks in BabyAI, a grid world navigation environment, as well as ALFRED, a household simulation environment. Haotian Fu, Pratyusha Sharma, Elias Stengel-Eskin, George Dimitri Konidaris, Nicolas Le Roux, Marc-Alexandre Côté, Xingdi Yuan |
ICML | 3 |
| 2024 | ReGAL: Refactoring Programs to Discover Generalizable AbstractionsabstractWhile large language models (LLMs) are increasingly being used for program synthesis, they lack the global view needed to develop useful abstractions; they generally predict programs one at a time, often repeating the same functionality. Generating redundant code from scratch is both inefficient and error-prone. To address this, we propose Refactoring for Generalizable Abstraction Learning (ReGAL), a gradient-free method for learning a library of reusable functions via code refactorization, i.e., restructuring code without changing its execution output. ReGAL learns from a small set of existing programs, iteratively verifying and refining its abstractions via execution. We find that the shared function libraries discovered by ReGAL make programs easier to predict across diverse domains. On five datasets – LOGO graphics generation, Date reasoning, TextCraft (a Minecraft-based text-game) MATH, and TabMWP – both open-source and proprietary LLMs improve in accuracy when predicting programs with REGAL functions. For CodeLlama-13B, REGAL results in absolute accuracy increases of 11.5% on LOGO, 26.1% on date understanding, and 8.1% on TextCraft, out-performing GPT-3.5 in two of three domains. Our analysis reveals REGAL’s abstractions encapsulate frequently-used subroutines as well as environment dynamics. Elias Stengel-Eskin, Archiki Prasad, Mohit Bansal |
ICML | 1 |
| 2024 | MIRACLE: An Online, Explainable Multimodal Interactive Concept Learning SystemabstractWe present MIRACLE, a system for online, interpretable visual concept and video action recognition. Through a chat interface, users query the recognition system with an uploaded image or video. For images, MIRACLE returns concept predictions from its structured knowledge base, justifying its predictions with heatmaps and natural language-based attribute detections. For videos, MIRACLE predicts an action and justifies its prediction with time varying entity-entity relations. With its ability to learn new concepts in an online, few-shot manner and its support of dynamic changes to its knowledge base, MIRACLE represents a step forward in interpretable multimodal learning systems. Ansel Blume, Khanh Duy Nguyen, Zhenhailong Wang, Yangyi Chen, Michal Shlapentokh-Rothman, Xiaomeng Jin, Zhen Zhu 0006, Jiateng Liu, Kuan-Hao Huang, Mankeerat Sidhu, Xuanming Zhang, Vivian Liu, Raunak Sinha, Te-Lin Wu, Abhaysinh Zala, Elias Stengel-Eskin, Da Yin, Utkarsh Mall, Zhou Yu 0005, Kai-Wei Chang 0001, Camille Cobb, Karrie Karahalios, Lydia B. Chilton, Mohit Bansal, Nanyun Peng 0001, Carl Vondrick, Derek Hoiem, Heng Ji 0001 |
ACM Multimedia | 17 |
| 2024 | GTBench: Uncovering the Strategic Reasoning Capabilities of LLMs via Game-Theoretic EvaluationsabstractAs Large Language Models (LLMs) are integrated into critical real-world applications, their strategic and logical reasoning abilities are increasingly crucial. This paper evaluates LLMs' reasoning abilities in competitive environments through game-theoretic tasks, e.g., board and card games that require pure logic and strategic reasoning to compete with opponents. We first propose GTBench, a language-driven environment composing 10 widely-recognized tasks, across a comprehensive game taxonomy: complete versus incomplete information, dynamic versus static, and probabilistic versus deterministic scenarios. Then, we (1) Characterize the game-theoretic reasoning of LLMs; and (2) Perform LLM-vs.-LLM competitions as reasoning evaluation. We observe that (1) LLMs have distinct behaviors regarding various gaming scenarios; for example, LLMs fail in complete and deterministic games yet they are competitive in probabilistic gaming scenarios; (2) Most open-source LLMs, e.g., CodeLlama-34b-Instruct and Llama-2-70b-chat, are less competitive than commercial LLMs, e.g., GPT-4, in complex games, yet the recently released Llama-3-70b-Instruct makes up for this shortcoming. In addition, code-pretraining greatly benefits strategic reasoning, while advanced reasoning methods such as Chain-of-Thought (CoT) and Tree-of-Thought (ToT) do not always help. We further characterize the game-theoretic properties of LLMs, such as equilibrium and Pareto Efficiency in repeated games. Detailed error profiles are provided for a better understanding of LLMs' behavior. We hope our research provides standardized protocols and serves as a foundation to spur further explorations in the strategic reasoning of LLMs. Jinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura, Lichao Sun 0001, Elias Stengel-Eskin, Mohit Bansal, Tianlong Chen 0001, Kaidi Xu |
NeurIPS | 6 |
| 2024 | LACIE: Listener-Aware Finetuning for Calibration in Large Language ModelsabstractWhen answering questions, large language models (LLMs) can convey not only an answer to the question, but a level of confidence about the answer being correct. This includes explicit markers of confidence (e.g. giving a numeric confidence score) as well as implicit markers, like using an authoritative tone or elaborating with additional knowledge of a subject. For LLMs to be trustworthy sources of knowledge, the confidence they convey should match their actual expertise on a topic; however, this is currently not the case, with most models tending towards overconfidence. To calibrate both implicit and explicit confidence markers, we introduce a pragmatic, listener-aware finetuning method (LACIE) that directly models the listener, considering not only whether an answer is right, but whether it will be accepted by a listener. Specifically, we cast calibration as a preference optimization problem, creating data via a two-agent speaker-listener game, where a speaker model’s outputs are judged by a simulated listener. We then finetune three different LLMs (Mistral-7B, Llama3-8B, Llama3-70B) with LACIE, and show that the models resulting from this multi-agent optimization are better calibrated on TriviaQA with respect to a simulated listener. Crucially, these trends transfer to human listeners, helping them correctly predict model correctness: we conduct a human evaluation where annotators accept or reject an LLM’s answers to trivia questions, finding that training with LACIE results in 47% fewer incorrect answers being accepted while maintaining the same level of acceptance for correct answers. Furthermore, LACIE generalizes to another dataset, resulting in a large increase in truthfulness on TruthfulQA when trained on TriviaQA. Our analysis indicates that LACIE leads to a better separation in confidence between correct and incorrect examples. Qualitatively, we find that a LACIE-trained model hedges more when uncertain and adopts implicit cues to signal certainty when it is correct, such as using an authoritative tone or including details. Finally, finetuning with our listener- aware method leads to an emergent increase in model abstention (e.g. saying “I don’t know”) for answers that are likely to be wrong, trading recall for precision. Elias Stengel-Eskin, Peter Hase, Mohit Bansal |
NeurIPS | 1 |
| 2023 | Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQAabstractNatural language is ambiguous.Resolving ambiguous questions is key to successfully answering them.Focusing on questions about images, we create a dataset of ambiguous examples.We annotate these, grouping answers by the underlying question they address and rephrasing the question for each group to reduce ambiguity.Our analysis reveals a linguistically-aligned ontology of reasons for ambiguity in visual questions.We then develop an English questiongeneration model which we demonstrate via automatic and human evaluation produces less ambiguous questions.We further show that the question generation objective we use allows the model to integrate answer group information without any direct supervision.1 Elias Stengel-Eskin, Jimena Guallar-Blasco, Benjamin Van Durme |
ACL (1) | 1 |
| 2023 | Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual ReasoningabstractVisual Question Answering (VQA) models often perform poorly on out-of-distribution data and struggle on domain generalization. Due to the multi-modal nature of this task, multiple factors of variation are intertwined, making generalization difficult to analyze. This motivates us to introduce a virtual benchmark, Super-CLEVR, where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently. Four factors are considered: visual complexity, question redundancy, concept distribution and concept compositionality. With controllably generated data, Super-CLEVR enables us to test VQA methods in situations where the test data differs from the training data along each of these axes. We study four existing methods, including two neural symbolic methods NSCL [45] and NSVQA [59], and two non-symbolic methods FiLM [50] and mDETR [29]; and our proposed method, probabilistic NSVQA (P-NSVQA), which extends NSVQA with uncertainty reasoning. P-NSVQA outperforms other methods on three of the four domain shift factors. Our results suggest that disentangling reasoning and perception, combined with probabilistic uncertainty, form a strong VQA model that is more robust to domain shifts. The dataset and code are released at https://github.com/Lizw14/Super-CLEVR. Zhuowan Li, Xingrui Wang, Elias Stengel-Eskin, Adam Kortylewski, Wufei Ma, Benjamin Van Durme, Alan L. Yuille |
CVPR | 3 |
| 2023 | Did You Mean...? Confidence-based Trade-offs in Semantic ParsingabstractWe illustrate how a calibrated model can help balance common trade-offs in task-oriented parsing.In a simulated annotator-in-the-loop experiment, we show that well-calibrated confidence scores allow us to balance cost with annotator load, improving accuracy with a small number of interactions.We then examine how confidence scores can help optimize the tradeoff between usability and safety.We show that confidence-based thresholding can substantially reduce the number of incorrect lowconfidence programs executed; however, this comes at a cost to usability.We propose the DidYouMean system (cf.Fig. 1) which better balances usability and safety by rephrasing low-confidence inputs. Elias Stengel-Eskin, Benjamin Van Durme |
EMNLP | 1 |
| 2023 | Calibrated Interpretation: Confidence Estimation in Semantic ParsingabstractAbstract Sequence generation models are increasingly being used to translate natural language into programs, i.e., to perform executable semantic parsing. The fact that semantic parsing aims to predict programs that can lead to executed actions in the real world motivates developing safe systems. This in turn makes measuring calibration—a central component to safety—particularly important. We investigate the calibration of popular generation models across four popular semantic parsing datasets, finding that it varies across models and datasets. We then analyze factors associated with calibration error and release new confidence-based challenge splits of two parsing datasets. To facilitate the inclusion of calibration in semantic parsing evaluations, we release a library for computing calibration metrics.1 Elias Stengel-Eskin, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 1 |
| 2022 | The Curious Case of ControlabstractChildren acquiring English make systematic errors on subject control sentences even after they have reached near-adult competence (Chomsky, 1969), possibly due to heuristics based on semantic roles (Maratsos, 1974).Given the advanced fluency of large generative language models, we ask whether model outputs are consistent with these heuristics, and to what degree different models are consistent with each other.We find that models can be categorized by behavior into three separate groups, with broad differences between the groups.The outputs of models in the largest group are consistent with positional heuristics that succeed on subject control but fail on object control.This result is surprising, given that object control is orders of magnitude more frequent in the text data used to train such models.We examine to what degree the models are sensitive to prompting with agent-patient information, finding that raising the salience of agent and patient relations results in significant changes in the outputs of most models.Based on this observation, we leverage an existing dataset of semantic protorole annotations (White et al., 2020) to explore the connections between control and labeling event participants with properties typically associated with agents and patients.1 You will be given a context and a question.Answer the question with either " " or " ".\n Context: told to come. Elias Stengel-Eskin, Benjamin Van Durme |
EMNLP | 1 |
| 2022 | When More Data Hurts: A Troubling Quirk in Developing Broad-Coverage Natural Language Understanding SystemsabstractElias Stengel-Eskin, Emmanouil Antonios Platanios, Adam Pauls, Sam Thomson, Hao Fang, Benjamin Van Durme, Jason Eisner, Yu Su. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Elias Stengel-Eskin, Emmanouil A. Platanios, Adam Pauls, Sam Thomson, Hao Fang 0002, Benjamin Van Durme, Jason Eisner, Yu Su 0001 |
EMNLP | 1 |
| 2022 | Visual Commonsense in Pretrained Unimodal and Multimodal ModelsabstractChenyu Zhang, Benjamin Van Durme, Zhuowan Li, Elias Stengel-Eskin. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Benjamin Van Durme, Zhuowan Li, Elias Stengel-Eskin |
NAACL-HLT | 4 |
| 2021 | Calibrating Concepts and Operations: Towards Symbolic Reasoning on Real ImagesabstractWhile neural symbolic methods demonstrate impressive performance in visual question answering on synthetic images, their performance suffers on real images. We identify that the long-tail distribution of visual concepts and unequal importance of reasoning steps in real data are the two key obstacles that limit the models’ real-world potentials. To address these challenges, we propose a new paradigm, Calibrating Concepts and Operations (CCO), which enables neural symbolic models to capture underlying data characteristics and to reason with hierarchical importance. Specifically, we introduce an executor with learnable concept embedding magnitudes for handling distribution imbalance, and an operation calibrator for highlighting important operations and suppressing redundant ones.Our experiments show CCO substantially boosts the performance of neural symbolic methods on real images. By evaluating models on the real world dataset GQA, CCO helps the neural symbolic method NSCL outperforms its vanilla counterpart by 9.1% (from 47.0% to 56.1%); this result also largely reduces the performance gap between symbolic and non-symbolic methods. Additionally, we create a perturbed test set for better understanding and analyzing model performance on real images. Code is available at https://lizw14.github.io/project/ccosr. Zhuowan Li, Elias Stengel-Eskin, Yixiao Zhang 0001, Cihang Xie, Quan Tran, Benjamin Van Durme, Alan L. Yuille |
ICCV | 2 |
| 2021 | Iterative Paraphrastic Augmentation with Discriminative Span AlignmentabstractAbstract We introduce a novel paraphrastic augmentation strategy based on sentence-level lexically constrained paraphrasing and discriminative span alignment. Our approach allows for the large-scale expansion of existing datasets or the rapid creation of new datasets using a small, manually produced seed corpus. We demonstrate our approach with experiments on the Berkeley FrameNet Project, a large-scale language understanding effort spanning more than two decades of human labor. With four days of training data collection for a span alignment model and one day of parallel compute, we automatically generate and release to the community 495,300 unique (Frame,Trigger) pairs in diverse sentential contexts, a roughly 50-fold expansion atop FrameNet v1.7. The resulting dataset is intrinsically and extrinsically evaluated in detail, showing positive results on a downstream task. Ryan Culkin, Edward J. Hu, Elias Stengel-Eskin, Guanghui Qin, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Joint Universal Syntactic and Semantic ParsingabstractWhile numerous attempts have been made to jointly parse syntax and semantics, high performance in one domain typically comes at the price of performance in the other. This trade-off contradicts the large body of research focusing on the rich interactions at the syntax–semantics interface. We explore multiple model architectures that allow us to exploit the rich syntactic and semantic annotations contained in the Universal Decompositional Semantics (UDS) dataset, jointly parsing Universal Dependencies and UDS to obtain state-of-the-art results in both formalisms. We analyze the behavior of a joint model of syntax and semantics, finding patterns supported by linguistic theory at the syntax–semantics interface. We then investigate to what degree joint modeling generalizes to a multilingual setting, where we find similar trends across 8 languages. Elias Stengel-Eskin, Kenton Murray, Sheng Zhang 0012, Aaron Steven White, Benjamin Van Durme |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Universal Decompositional Semantic ParsingabstractWe introduce a transductive model for parsing into Universal Decompositional Semantics (UDS) representations, which jointly learns to map natural language utterances into UDS graph structures and annotate the graph with decompositional semantic attribute scores.We also introduce a strong pipeline model for parsing into the UDS graph structure, and show that our transductive parser performs comparably while additionally performing attribute prediction.By analyzing the attribute prediction errors, we find the model captures natural relationships between attribute groups. Elias Stengel-Eskin, Aaron Steven White, Sheng Zhang 0012, Benjamin Van Durme |
ACL | 1 |
| 2020 | The Universal Decompositional Semantics Dataset and Decomp ToolkitabstractWe present the Universal Decompositional Semantics (UDS) dataset (v1.0), which is bundled with the Decomp toolkit (v0.1). UDS1.0 unifies five high-quality, decompositional semantics-aligned annotation sets within a single semantic graph specification—with graph structures defined by the predicative patterns produced by the PredPatt tool and real-valued node and edge attributes constructed using sophisticated normalization procedures. The Decomp toolkit provides a suite of Python 3 tools for querying UDS graphs using SPARQL. Both UDS1.0 and Decomp0.1 are publicly available at http://decomp.io. Aaron Steven White, Elias Stengel-Eskin, Siddharth Vashishtha, Venkata Subrahmanyan Govindarajan, Dee Ann Reisinger, Tim Vieira, Keisuke Sakaguchi, Sheng Zhang 0012, Francis Ferraro, Rachel Rudinger, Kyle Rawlins, Benjamin Van Durme |
LREC | 2 |
| 2019 | A Discriminative Neural Model for Cross-Lingual Word AlignmentabstractElias Stengel-Eskin, Tzu-ray Su, Matt Post, Benjamin Van Durme. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Elias Stengel-Eskin, Tzu-Ray Su, Matt Post, Benjamin Van Durme |
EMNLP/IJCNLP (1) | 1 |
| 2017 | Polyglot and Speech Corpus Tools: A System for Representing, Integrating, and Querying Speech Corpora
Michael McAuliffe, Elias Stengel-Eskin, Michaela Socolof, Morgan Sonderegger |
INTERSPEECH | 2 |