VLDB 2026 Research / reviewers in the wild / expert
Mrinmaya Sachan
dblp:86/10440
· DBLP profile ↗
93ranked-venue papers
13as first author
77since 2021 · last 2026
0000-0001-8787-8681ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 87 · 12 first-author · 72 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
Jingwei Ni, Ekaterina Fadeeva, Mubashara Akhtar, Jiaheng Zhang, Elliott Ash, Markus Leippold, Timothy Baldwin, See-Kiong Ng, Artem Shelmanov, Mrinmaya Sachan |
ACL (1) | 11 |
| 2026 | Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical ArgumentationabstractIdentifying logical fallacies in everyday discourse is challenging for many people.This challenge is amplified in the era of Large Language Models (LLMs), where malicious agents can deploy fallacious arguments to disseminate misinformation at scale.In this work, we explore the potential of LLMs as part of the solution.We introduce LFTutor, an intelligent tutoring system which uses LLMs to tutor laypeople and help them learn about logical fallacies.LFTutor integrates intent-driven Socratic questioning and critical argumentation principles to actively engage learners to reflect on their reasoning.Through both automatic and human evaluations, we demonstrate that LFTutor significantly outperforms baseline LLMs lacking these pedagogical strategies.This work highlights the promise of combining LLMs with pedagogical scaffolding to foster critical thinking and argument literacy in the age of AI.ing university students' ability to recognize argument structures and fallacies. Minjing Shi, Junling Wang 0001, Jingwei Ni, Sankalan Pal Chowdhury, Mrinmaya Sachan |
ACL (1) | 5 |
| 2026 | Test of Time: Rethinking Temporal Signal of Benchmark ContaminationabstractTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter, Punya Syon Pandey, Keenan Samway, Wenyuan Jiang, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Terry Jingchen Zhang, Gopal Dev, Max Obreiter, Wenyuan Jiang, Punya Syon Pandey, Keenan Samway, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin 0001 |
ACL (1) | 10 |
| 2026 | Misconception Acquisition Dynamics in Large Language Models
Naiming Liu, Xinghe Chen, Richard G. Baraniuk, Mrinmaya Sachan, Shashank Sonkar |
AIED (1) | 4 |
| 2026 | Bridging Instead of Replacing Online Coding Communities with AI through Community-Enriched Chatbot Designs CSCW008abstractLLM-based chatbots like ChatGPT have become popular tools for assisting with coding tasks. However, they often produce isolated responses and lack mechanisms for social learning or contextual grounding. In contrast, online coding communities like Kaggle offer socially mediated learning environments that foster critical thinking, engagement, and a sense of belonging. Yet, growing reliance on LLMs risks diminishing participation in these communities and weakening their collaborative value. To address this, we propose Community-Enriched AI, a design paradigm that embeds social learning dynamics into LLM-based chatbots by surfacing user-generated content and social design features from online coding communities. Using this paradigm, we implemented a RAG-based AI chatbot leveraging resources from Kaggle to validate our design. Across two empirical studies involving 28 and 12 data science learners, respectively, we found that Community-Enriched AI significantly enhances user trust, encourages engagement with community, and effectively supports learners in solving data science tasks. We conclude by discussing design implications for AI assistance systems that bridge—rather than replace—online coding communities. Junling Wang 0001, Lahari Goswami, Gustavo Umbelino, Kiara Chau, Mrinmaya Sachan, April Yi Wang |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2025 | Calibrating Large Language Models with Sample ConsistencyabstractAccurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional calibration techniques due to their proprietary nature and massive scale. In this work, we derive model confidence from the distribution of multiple randomly sampled generations, using three measures of consistency. We extensively evaluate eleven open and closed-source models on nine reasoning datasets. Results show that consistency-based calibration methods outperform existing post-hoc approaches in terms of calibration error. Meanwhile, we find that factors such as intermediate explanations, model scaling, and larger sample sizes enhance calibration, while instruction-tuning makes calibration more difficult. Moreover, confidence scores obtained from consistency can potentially enhance model performance. Finally, we offer guidance on choosing suitable consistency metrics for calibration, tailored to model characteristics such as the exposure to instruction-tuning and RLHF. Qing Lyu 0001, Kumar Shridhar, Chaitanya Malaviya, Li Zhang 0039, Yanai Elazar, Niket Tandon, Marianna Apidianaki, Mrinmaya Sachan, Chris Callison-Burch |
AAAI | 8 |
| 2025 | GPT-4 as a Homework Tutor Can Improve Student Engagement and Learning OutcomesabstractThis work contributes to the scarce empirical literature on LLM-based interactive homework in real-world educational settings and offers a practical, scalable solution to improve homework in schools.Homework is an important part of education in schools across the world, but to maximize benefit, it must be accompanied by feedback and follow-up questions.We developed a prompting strategy that enables GPT-4 to conduct interactive homework sessions for high school students learning English as a second language.Our strategy requires minimal effort in content preparation, one of the key challenges of alternatives such as home tutors or ITSs.We carried out a Randomized Controlled Trial (RCT) in four highschool classes, replacing traditional homework with GPT-4 homework sessions for the treatment group.We found that the treatment group had higher levels of satisfaction and desire to keep using the system among the students.This occurred without compromising learning outcomes, and one group even showed significantly better learning gains. Alessandro Vanzo, Sankalan Pal Chowdhury, Mrinmaya Sachan |
ACL (1) | 3 |
| 2025 | Can Vision-Language Models Solve Visual Math Equations?abstractDespite strong performance in visual understanding and language-based reasoning, Vision-Language Models (VLMs) struggle with tasks requiring integrated perception and symbolic computation.We study this limitation through visual equation solving, where mathematical equations are embedded in images, variables are represented by object icons, and coefficients must be inferred by counting.While VLMs perform well on textual equations, they fail on visually grounded counterparts.To understand this gap, we decompose the task into coefficient counting and variable recognition, and find that counting is the primary bottleneck, even when recognition is accurate.We also observe that composing recognition and reasoning introduces additional errors, highlighting challenges in multi-step visual reasoning.Finally, as equation complexity increases, symbolic reasoning itself becomes a limiting factor.These findings reveal key weaknesses in current VLMs and point toward future improvements in visually grounded mathematical reasoning.1 * Equal contribution. 1 Our code and data are publicly available. Monjoy Narayan Choudhury, Junling Wang 0001, Mrinmaya Sachan |
EMNLP | 4 |
| 2025 | From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement LearningabstractLarge language models (LLMs) can transform education, but their optimization for direct question-answering often undermines effective pedagogy which requires strategically withholding answers.To mitigate this, we propose an online reinforcement learning (RL)-based alignment framework that can quickly adapt LLMs into effective tutors using simulated student-tutor interactions by emphasizing pedagogical quality and guided problem-solving over simply giving away answers.We use our method to train a 7B parameter tutor model without human annotations which reaches similar performance to larger proprietary models like LearnLM.We introduce a controllable reward weighting to balance pedagogical support and student solving accuracy, allowing us to trace the Pareto frontier between these two objectives.Our models better preserve reasoning capabilities than single-turn SFT baselines and can optionally enhance interpretability through thinking tags that expose the model's instructional planning. David Dinucu-Jianu, Jakub Macina, Nico Daheim, Ido Hakimi, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 6 |
| 2025 | The Medium Is Not the Message: Deconfounding Document Embeddings via Linear Concept ErasureabstractEmbedding-based similarity metrics between text sequences can be influenced not just by the content dimensions we most care about, but can also be biased by spurious attributes like the text's source or language.These document confounders cause problems for many applications, but especially those that need to pool texts from different corpora.This paper shows that a debiasing algorithm that removes information about observed confounders from the encoder representations substantially reduces these biases at a minimal computational cost.Document similarity and clustering metrics improve across every embedding variant and task we evaluate-often dramatically.Interestingly, performance on out-of-distribution benchmarks is not impacted, indicating that the embeddings are not otherwise degraded.1 Yu Fan 0007, Shauli Ravfogel, Mrinmaya Sachan, Elliott Ash, Alexander Miserlis Hoyle |
EMNLP | 4 |
| 2025 | MathTutorBench: A Benchmark for Measuring Open-ended Pedagogical Capabilities of LLM TutorsabstractEvaluating the pedagogical capabilities of AIbased tutoring models is critical for making guided progress in the field.Yet, we lack a reliable, easy-to-use, and simple-to-run evaluation that reflects the pedagogical abilities of models.To fill this gap, we present MATH-TUTORBENCH, an open-source benchmark for holistic tutoring model evaluation.MATHTU-TORBENCH contains a collection of datasets and metrics that broadly cover tutor abilities as defined by learning sciences research in dialogbased teaching.To score the pedagogical quality of open-ended teacher responses, we train a reward model and show it can discriminate expert from novice teacher responses with high accuracy.We evaluate a wide set of closed-and open-weight models on MATHTUTORBENCH and find that subject expertise, indicated by solving ability, does not immediately translate to good teaching.Rather, pedagogy and subject expertise appear to form a trade-off that is navigated by the degree of tutoring specialization of the model.Furthermore, tutoring appears to become more challenging in longer dialogs, where simpler questioning strategies begin to fail.We release the benchmark, code, and leaderboard openly to enable rapid benchmarking of future models. 1 github.com/eth-lre/mathtutorbench Jakub Macina, Nico Daheim, Ido Hakimi, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 6 |
| 2025 | A Head to Predict and a Head to Question: Pre-trained Uncertainty Quantification Heads for Hallucination Detection in LLM OutputsabstractArtem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Artem Shelmanov, Ekaterina Fadeeva, Akim Tsvigun, Ivan Tsvigun, Zhuohan Xie, Igor Kiselev, Nico Daheim, Caiqi Zhang, Artem Vazhentsev, Mrinmaya Sachan, Preslav Nakov, Timothy Baldwin |
EMNLP | 10 |
| 2025 | Improving Large Language Model Safety with Contrastive Representation LearningabstractLarge Language Models (LLMs) are powerful tools with profound societal impacts, yet their ability to generate responses to diverse and uncontrolled inputs leaves them vulnerable to adversarial attacks.While existing defenses often struggle to generalize across varying attack types, recent advancements in representation engineering offer promising alternatives.In this work, we propose a defense framework that formulates model defense as a contrastive representation learning (CRL) problem.Our method finetunes a model using a triplet-based loss combined with adversarial hard negative mining to encourage separation between benign and harmful representations.Our experimental results across multiple models demonstrate that our approach outperforms prior representation engineering-based defenses, improving robustness against both input-space and embeddingspace attacks without compromising standard performance.1 Samuel Simko, Mrinmaya Sachan, Bernhard Schölkopf, Zhijing Jin 0001 |
EMNLP | 2 |
| 2025 | Probing for Arithmetic Errors in Language ModelsabstractWe investigate whether internal activations in language models can be used to detect arithmetic errors.Starting with a controlled setting of 3-digit addition, we show that simple probes can accurately decode both the model's predicted output and the correct answer from hidden states, regardless of whether the model's output is correct.Building on this, we train lightweight error detectors that predict model correctness with over 90% accuracy.We then extend our analysis to structured chain-ofthought traces on addition-only GSM8K problems and find that probes trained on simple arithmetic generalize well to this more complex setting, revealing consistent internal representations.Finally, we demonstrate that these probes can guide selective re-prompting of erroneous reasoning steps, improving task accuracy with minimal disruption to correct outputs.Our findings suggest that arithmetic errors can be anticipated from internal activations alone, and that simple probes offer a viable path toward lightweight model self-correction. 1 Alessandro Stolfo, Mrinmaya Sachan |
EMNLP | 3 |
| 2025 | Language Model Alignment in Multilingual Trolley ProblemsabstractWe evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems. Building on the Moral Machine experiment, which captures over 40 million human judgments across 200+ countries, we develop a cross-lingual corpus of moral dilemma vignettes in over 100 languages called MultiTP. This dataset enables the assessment of LLMs' decision-making processes in diverse linguistic contexts. Our analysis explores the alignment of 19 different LLMs with human judgments, capturing preferences across six moral dimensions: species, gender, fitness, status, age, and the number of lives involved. By correlating these preferences with the demographic distribution of language speakers and examining the consistency of LLM responses to various prompt paraphrasings, our findings provide insights into cross-lingual and ethical biases of LLMs and their intersection. We discover significant variance in alignment across languages, challenging the assumption of uniform moral reasoning in AI systems and highlighting the importance of incorporating diverse perspectives in AI ethics. The results underscore the need for further research on the integration of multilingual dimensions in responsible AI research to ensure fair and equitable AI interactions worldwide. Zhijing Jin 0001, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu 0004, Fernando Gonzalez Adauto, Francesco Ortu, András Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi 0001, Bernhard Schölkopf |
ICLR | 9 |
| 2025 | MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex ProofsabstractLarge language models (LLMs) can solve arithmetic word problems with high accuracy, but little is known about how well they generalize to more complex problems. This is difficult to study, as (i) much of the available evaluation data has already been seen by the most capable models during training, and (ii) existing benchmarks do not capture how problem proofs may be arbitrarily complex in various ways. In this paper, we present a data-generation framework for evaluating LLMs on problems with arbitrarily complex arithmetic proofs, called MathGAP. MathGAP generates problem statements and chain-of-thought reasoning traces according to specifications about their arithmetic proof structure, enabling systematic studies on easy-to-hard generalization with respect to complexity of proof trees. Using MathGAP, we find that LLMs show a significant decrease in performance as proofs get deeper and wider. This effect is more pronounced in complex, nonlinear proof structures, which are challenging even for the most capable models. The models are also sensitive to simple changes in sentence ordering. However, they remain capable of solving some complex problems, suggesting that reasoning generalization is noisy. Andreas Opedal, Haruki Shirakami, Bernhard Schölkopf, Abulhair Saparov, Mrinmaya Sachan |
ICLR | 5 |
| 2025 | Do Vision-Language Models Really Understand Visual Language?abstractVisual language is a system of communication that conveys information through symbols, shapes, and spatial arrangements. Diagrams are a typical example of a visual language depicting complex concepts and their relationships in the form of an image. The symbolic nature of diagrams presents significant challenges for building models capable of understanding them. Yet, recent studies seem to suggest that Large Vision-Language Models (LVLMs) can even tackle complex reasoning tasks involving diagrams. In this paper, we investigate this phenomenon by developing a comprehensive test suite to evaluate the diagram comprehension capability of LVLMs. Our test suite uses a variety of questions focused on concept entities and their relationships over a set of synthetic as well as real diagrams across several domains to evaluate the recognition and reasoning abilities of models. Our evaluation of six LVLMs shows that while these models can accurately identify and reason about entities, their ability to understand relationships is notably limited. Further testing reveals that the decent performance on diagram understanding largely stems from leveraging their background knowledge as shortcuts to identify and reason about the relational information. Thus, we conclude that LVLMs have a limited capability for genuine diagram understanding, and their impressive performance in diagram reasoning is an illusion emanating from other confounding factors, such as the background knowledge in the models. Buse Giledereli, Yilei Tu, Mrinmaya Sachan |
ICML | 4 |
| 2025 | Grammar Control in Dialogue Response Generation for Language Learning ChatbotsabstractDominik Glandorf, Peng Cui, Detmar Meurers, Mrinmaya Sachan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dominik Glandorf, Peng Cui 0006, Detmar Meurers, Mrinmaya Sachan |
NAACL (Long Papers) | 4 |
| 2025 | Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented GenerationabstractTianyu Liu, Jirui Qi, Paul He, Arianna Bisazza, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tianyu Liu 0004, Jirui Qi, Paul He 0001, Arianna Bisazza, Mrinmaya Sachan, Ryan Cotterell |
NAACL (Long Papers) | 5 |
| 2025 | DIRAS: Efficient LLM Annotation of Document Relevance for Retrieval Augmented GenerationabstractJingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
NAACL (Long Papers) | 4 |
| 2025 | AI-Assisted Human Evaluation of Machine TranslationabstractVilém Zouhar, Tom Kocmi, Mrinmaya Sachan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Vilém Zouhar, Tom Kocmi, Mrinmaya Sachan |
NAACL (Long Papers) | 3 |
| 2025 | Are Language Models Efficient Reasoners? A Perspective from Logic ProgrammingabstractModern language models (LMs) exhibit strong deductive reasoning capabilities, yet standard evaluations emphasize correctness while overlooking a key aspect of reasoning: *efficiency*. In real-world reasoning scenarios, much of the available information is irrelevant, and effective deductive inference requires identifying and ignoring such distractions. We propose a framework for assessing LM reasoning efficiency through the lens of logic programming, introducing a simple method to align proofs written in natural language---as generated by an LM---with shortest proofs found by executing the logic program. Efficiency is quantified by measuring how well a model avoids unnecessary inference. Empirically, we construct a dataset of math word problems injected with various number of irrelevant axioms that vary in semantic overlap with the goal theorem. We find that current LMs show marked accuracy declines under such conditions---even with minimal, domain-consistent distractions---and the proofs they generate frequently exhibit detours through irrelevant inferences. Andreas Opedal, Yanick Zengaffinen, Haruki Shirakami, Clemente Pasti, Mrinmaya Sachan, Abulhair Saparov, Ryan Cotterell, Bernhard Schölkopf |
NeurIPS | 5 |
| 2025 | Personalized Exercise Recommendation with Semantically-Grounded Knowledge TracingabstractWe introduce ExRec, a general framework for personalized exercise recommendation with semantically-grounded knowledge tracing. Our method builds on the observation that existing exercise recommendation approaches simulate student performance via knowledge tracing (KT) but they often overlook two key aspects: (a) the semantic content of questions and (b) the sequential, structured progression of student learning. To address this, our ExRec presents an end-to-end pipeline, from annotating the KCs of questions and learning their semantic representations to training KT models and optimizing several reinforcement learning (RL) methods. Moreover, we improve standard Q-learning-based continuous RL methods via a tailored model-based value estimation (MVE) approach that directly leverages the components of KT model in estimating cumulative knowledge improvement. We validate the effectiveness of our ExRec using various RL methods across four real-world tasks with different educational goals in online math learning. We further show that ExRec generalizes robustly to new, unseen questions and that it produces interpretable student learning trajectories. Together, our findings highlight the promise of KT-guided RL for effective personalization in education. Yilmazcan Özyurt, Tunaberk Almaci, Stefan Feuerriegel, Mrinmaya Sachan |
NeurIPS | 4 |
| 2025 | Dense SAE Latents Are Features, Not BugsabstractSparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that are both sparse and semantically meaningful. However, many SAE latents activate frequently (i.e., are *dense*), raising concerns that they may be undesirable artifacts of the training procedure. In this work, we systematically investigate the geometry, function, and origin of dense latents and show that they are not only persistent but often reflect meaningful model representations. We first demonstrate that dense latents tend to form antipodal pairs that reconstruct specific directions in the residual stream, and that ablating their subspace suppresses the emergence of new dense features in retrained SAEs---suggesting that high density features are an intrinsic property of the residual space. We then introduce a taxonomy of dense latents, identifying classes tied to position tracking, context binding, entropy regulation, letter-specific output signals, part-of-speech, and principal component reconstruction. Finally, we analyze how these features evolve across layers, revealing a shift from structural features in early layers, to semantic features in mid layers, and final to output-oriented signals in the last layers of the model. Our findings indicate that dense latents serve functional roles in language model computation and should not be dismissed as training noise. Xiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu 0001, Senthooran Rajamanoharan, Mrinmaya Sachan, Max Tegmark |
NeurIPS | 6 |
| 2025 | SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics ReasoningabstractWe present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts.Project Page: github.com/SeePhys/seephys-projectHugging Face: huggingface.co/datasets/SeePhys/SeePhys Kun Xiang, Terry Jingchen Zhang, Yinya Huang, Zirong Liu, Peixin Qu, Jixi He, Yu-Jie Yuan, Jianhua Han, Hang Xu 0004, Mrinmaya Sachan, Xiaodan Liang |
NeurIPS | 13 |
| 2025 | How to Select Datapoints for Efficient Human Evaluation of NLG Models?abstractAbstract Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However, randomly selected data may not accurately represent test performance, making this approach economically inefficient for model comparison. Thus, in this work, we develop and analyze a suite of selectors to get the most informative datapoints for human evaluation, taking the evaluation costs into account. We show that selectors based on variance in automated metric scores, diversity in model outputs, or Item Response Theory outperform random selection. We further develop an approach to distill these selectors to the scenario where the model outputs are not yet available. In particular, we introduce source-based estimators, which predict item usefulness for human evaluation just based on the source texts. We demonstrate the efficacy of our selectors in two common NLG tasks, machine translation and summarization, and show that only ∼70% of the test data is needed to produce the same evaluation result as the entire data. Vilém Zouhar, Peng Cui 0006, Mrinmaya Sachan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | How to Engage your Readers? Generating Guiding Questions to Promote Active ReadingabstractUsing questions in written text is an effective strategy to enhance readability.However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied.We introduce GUIDINGQ, a dataset of 10K in-text questions from textbooks and scientific articles.By analyzing the dataset, we present a comprehensive understanding of the use, distribution, and linguistic characteristics of these questions.Then, we explore various approaches to generate such questions using language models.Our results highlight the importance of capturing inter-question relationships and the challenge of question position identification in generating these questions.Finally, we conduct a human study to understand the implication of such questions on reading comprehension.We find that the generated questions are of high quality and are almost as effective as human-written questions in terms of improving readers' memorization and comprehension.github.com/eth-lre/engage-your-readers Questions in titles:How do Philosophers arrive at truth?Is there no quantum form of Einstein Gravity?Why do house-hunting ants recruit in both directions? Peng Cui 0006, Vilém Zouhar, Xiaoyu Zhang 0014, Mrinmaya Sachan |
ACL (1) | 4 |
| 2024 | What Do Language Models Learn in Context? The Structured Task HypothesisabstractLarge language models (LLMs) exhibit an intriguing ability to learn a novel task from incontext examples presented in a demonstration, termed in-context learning (ICL).Understandably, a swath of research has been dedicated to uncovering the theories underpinning ICL.One popular hypothesis explains ICL by task selection.LLMs identify the task based on the demonstration and generalize it to the prompt.Another popular hypothesis is that ICL is a form of meta-learning, i.e., the models learn a learning algorithm at pre-training time and apply it to the demonstration.Finally, a third hypothesis argues that LLMs use the demonstration to select a composition of tasks learned during pre-training to perform ICL.In this paper, we empirically explore these three hypotheses that explain LLMs' ability to learn in context with a suite of experiments derived from common text classification tasks.We invalidate the first two hypotheses with counterexamples and provide evidence in support of the last hypothesis.Our results suggest an LLM could learn a novel task in context via composing tasks learned during pre-training. Jiaoda Li, Mrinmaya Sachan, Ryan Cotterell |
ACL (1) | 3 |
| 2024 | AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsabstractJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
ACL (1) | 4 |
| 2024 | Competition of Mechanisms: Tracing How Language Models Handle Facts and CounterfactualsabstractFrancesco Ortu, Zhijing Jin, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, Bernhard Schölkopf. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Francesco Ortu, Zhijing Jin 0001, Diego Doimo, Mrinmaya Sachan, Alberto Cazzaniga, Bernhard Schölkopf |
ACL (1) | 4 |
| 2024 | RELIC: Investigating Large Language Model Responses using Self-ConsistencyabstractLarge Language Models (LLMs) are notorious for blending fact with fiction and generating non-factual content, known as hallucinations. To address this challenge, we propose an interactive system that helps users gain insight into the reliability of the generated text. Our approach is based on the idea that the self-consistency of multiple samples generated by the same LLM relates to its confidence in individual claims in the generated texts. Using this idea, we design RELIC, an interactive system that enables users to investigate and verify semantic-level variations in multiple long-form responses. This allows users to recognize potentially inaccurate information in the generated text and make necessary corrections. From a user study with ten participants, we demonstrate that our approach helps users better verify the reliability of the generated text. We further summarize the design implications and lessons learned from this research for future studies of reliable human-LLM interactions. Furui Cheng, Vilém Zouhar, Simran Arora, Mrinmaya Sachan, Hendrik Strobelt, Mennatallah El-Assady |
CHI | 4 |
| 2024 | PWESuite: Phonetic Word Embeddings and Tasks They FacilitateabstractMapping words into a fixed-dimensional vector space is the backbone of modern NLP. While most word embedding methods successfully encode semantic information, they overlook phonetic information that is crucial for many tasks. We develop three methods that use articulatory features to build phonetically informed word embeddings. To address the inconsistent evaluation of existing phonetic word embedding methods, we also contribute a task suite to fairly evaluate past, current, and future methods. We evaluate both (1) intrinsic aspects of phonetic word embeddings, such as word retrieval and correlation with sound similarity, and (2) extrinsic performance on tasks such as rhyme and cognate detection and sound analogies. We hope our task suite will promote reproducibility and inspire future phonetic embedding research. Vilém Zouhar, Kalvin Chang, Chenxuan Cui, Nate B. Carlson, Nathaniel R. Robinson, Mrinmaya Sachan, David R. Mortensen |
LREC/COLING | 6 |
| 2024 | RETRO-LI: Small-Scale Retrieval Augmented Generation Supporting Noisy Similarity Searches and Domain Shift GeneralizationabstractThe retrieval augmented generation (RAG) system such as RETRO has been shown to improve language modeling capabilities and reduce toxicity and hallucinations by retrieving from a database of non-parametric memory containing trillions of entries. We introduce RETRO-LI that shows retrieval can also help using a small scale database, but it demands more accurate and better neighbors when searching in a smaller hence sparser non-parametric memory. This can be met by using a proper semantic similarity search. We further propose adding a regularization to the non-parametric memory for the first time: it significantly reduces perplexity when the neighbor search operations are noisy during inference, and it improves generalization when a domain shift occurs. We also show that the RETRO-LI’s non-parametric memory can potentially be implemented on analog in-memory computing hardware, exhibiting O(1) search time while causing noise in retrieving neighbors, with minimal (<1%) performance loss. Our code is available at: https://github.com/IBM/Retrieval-Enhanced-Transformer-Little Gentiana Rashiti, Geethan Karunaratne, Mrinmaya Sachan, Abu Sebastian, Abbas Rahimi |
ECAI | 3 |
| 2024 | Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model TutorsabstractLarge language models (LLMs) present an opportunity to scale high-quality personalized education to all.A promising approach towards this means is to build dialog tutoring models that scaffold students' problem-solving.However, even though existing LLMs perform well in solving reasoning questions, they struggle to precisely detect student's errors and tailor their feedback to these errors.Inspired by realworld teaching practice where teachers identify student errors and customize their response based on them, we focus on verifying student solutions and show how grounding to such verification improves the overall quality of tutor response generation.We collect a dataset of 1K stepwise math reasoning chains with the first error step annotated by teachers.We show empirically that finding the mistake in a student solution is challenging for current models.We propose and evaluate several verifiers for detecting these errors.Using both automatic and human evaluation we show that the student solution verifiers steer the generation model towards highly targeted responses to student errors which are more often correct with less hallucinations compared to existing baselines.https://github.com/eth-lre/ verify-then-generate Teacher If the height is 6, what is the length of the box?Volume of a box is height * width * length.Student Multi-turn dialog tutoring task Goal: Generate next teacher utterance.Not quite.Is the length you computed 2-times more than height?targeted and correct A. Error reason (baseline): Student made a careless mistake.B. Correctness verification: incorrect C. Stepwise verification: Step 2 -We set an equation 2 * length = 6 ... D. Error Description: length is used as a label instead of a variable representing the number.E. Alignment: Missing student steps: We know height is 6,... Matching steps are: Next we know length...<=>We set an equation... Verification-based Conditional Generation ModelEquation is height * width * length.Volume of a box is height * width * length.We set an equation 2 * length = 6, so length is 3.Next we know length = 2 * height, so length is 12. Student Reasoning Chain-of-Thought (CoT) SolutionThe volume is 6 * 4 * 3 = 72.We know height is 6, width is 4, and we found the length is 12.So the volume is 6 * 4 * 12 = 288.I think the answer is 72.Great work, this is correct!factually incorrect Conditional Generation Model (baseline) Targeted Correct Actionable 1. Stepwise verification + - Nico Daheim, Jakub Macina, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EMNLP | 5 |
| 2024 | Towards Aligning Language Models with Textual FeedbackabstractWe present ALT (ALignment with Textual feedback), an approach that aligns language models with user preferences expressed in text.We argue that text offers greater expressiveness, enabling users to provide richer feedback than simple comparative preferences, leading to more efficient and effective alignment.ALT aligns the model by conditioning its generations on the textual feedback.Our method relies solely on language modeling techniques and requires minimal hyper-parameter tuning while retaining the main benefits of RL-based alignment algorithms.We demonstrate the efficacy and efficiency of textual feedback across different tasks, including toxicity reduction, summarization, and dialogue response generation.Notably, ALT outperforms PPO in toxicity reduction and matches its performance on summarization with only 20% of the samples.We also explore using ALT with feedback from an existing LLM, examining constrained and unconstrained feedback.Additionally, we outline future directions to align models with natural language feedback.1 Saüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan |
EMNLP | 4 |
| 2024 | Can Large Language Models Infer Causation from Correlation?abstractCausal inference is one of the hallmarks of human intelligence. While the field of CausalNLP has attracted much interest in the recent years, existing causal inference datasets in NLP primarily rely on discovering causality from empirical knowledge (e.g., commonsense knowledge). In this work, we propose the first benchmark dataset to test the pure causal inference skills of large language models (LLMs). Specifically, we formulate a novel task Corr2Cause, which takes a set of correlational statements and determines the causal relationship between the variables. We curate a large-scale dataset of more than 200K samples, on which we evaluate seventeen existing LLMs. Through our experiments, we identify a key shortcoming of LLMs in terms of their causal inference skills, and show that these models achieve almost close to random performance on the task. This shortcoming is somewhat mitigated when we try to re-purpose LLMs for this skill via finetuning, but we find that these models still fail to generalize – they can only perform causal inference in in-distribution settings when variable names and textual expressions used in the queries are similar to those in the training set, but fail in out-of-distribution settings generated by perturbing these queries. Corr2Cause is a challenging task for LLMs, and can be helpful in guiding future research on improving LLMs’ pure reasoning skills and generalizability. Our data is at https://huggingface.co/datasets/causalnlp/corr2cause. Our code is at https://github.com/causalNLP/corr2cause. Zhijing Jin 0001, Jiarui Liu 0004, Zhiheng Lyu, Spencer Poff, Mrinmaya Sachan, Rada Mihalcea, Mona T. Diab, Bernhard Schölkopf |
ICLR | 5 |
| 2024 | Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?abstractThere is increasing interest in employing large language models (LLMs) as cognitive models. For such purposes, it is central to understand which properties of human cognition are well-modeled by LLMs, and which are not. In this work, we study the biases of LLMs in relation to those known in children when solving arithmetic word problems. Surveying the learning science literature, we posit that the problem-solving process can be split into three distinct steps: text comprehension, solution planning and solution execution. We construct tests for each one in order to understand whether current LLMs display the same cognitive biases as children in these steps. We generate a novel set of word problems for each of these tests, using a neuro-symbolic approach that enables fine-grained control over the problem features. We find evidence that LLMs, with and without instruction-tuning, exhibit human-like biases in both the text-comprehension and the solution-planning steps of the solving process, but not in the final step, in which the arithmetic expressions are executed to obtain the answer. Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao, Ryan Cotterell, Bernhard Schölkopf, Abulhair Saparov, Mrinmaya Sachan |
ICML | 8 |
| 2024 | Slicing, Chatting, and Refining: A Concept-Based Approach for Machine Learning Model Validation with ConceptSlicerabstractAs machine learning (ML) gains wider adoption in real-world applications, the validation of ML models becomes fundamental for its productization, particularly in safety-critical applications. Recently, data slice finding has emerged as a popular method for validating ML models, but it requires additional metadata or cross-modal embeddings for the slices to be interpretable. We propose ConceptSlicer, an integrated workflow that facilitates the slicing of computer vision models using visual concepts. This approach breaks down the image dataset into interpretable visual concepts, serving as metadata in the slice finding process. Our system offers insights into model issues and enables a deeper understanding of computer vision models’ strengths and weaknesses. We evaluate ConceptSlicer through interviews with eight domain experts and machine learning practitioners, and fine-tune the ML models based on their feedback. Our study also highlights varied attitudes towards large foundational models, encouraging contemplation of the challenges and opportunities presented by this technological advancement. Xiaoyu Zhang 0014, Jorge Henrique Piazentin Ono, Liang Gou, Mrinmaya Sachan, Kwan-Liu Ma, Liu Ren 0001 |
IUI | 5 |
| 2024 | AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and GuardrailsabstractLarge Language Models (LLMs) have found several use cases in education, ranging from automatic question generation to essay evaluation. In this paper, we explore the potential of using LLMs to author Intelligent Tutoring Systems. A common pitfall of using LLMs as tutors is their straying from desired pedagogical strategies such as leaking the answer to the student, and in general, providing no guarantees on the validity or appropriateness of the tutor assistance. We argue that while LLMs with certain guardrails can take the place of subject experts, the overall pedagogical design still needs to be handcrafted for the best learning results. Based on this principle, we create a sample end-to-end tutoring system named MWPTutor, which uses LLMs to fill in the state space of a predefined finite state transducer. This approach retains the structure and the pedagogy of traditional tutoring systems that has been developed over the years by learning scientists but brings in additional flexibility of LLM-based approaches. Through a human evaluation study on two datasets with math word problems, we show that our hybrid approach achieves a better overall tutoring score than an instructed, but otherwise free-form, GPT-4. MWPTutor is completely modular and opens up the scope for the community to improve its performance by refining its individual modules or using different teaching strategies that it can follow. Sankalan Pal Chowdhury, Vilém Zouhar, Mrinmaya Sachan |
L@S | 3 |
| 2024 | Elastic Weight Removal for Faithful and Abstractive Dialogue GenerationabstractNico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Ponti. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, Edoardo Maria Ponti |
NAACL-HLT | 3 |
| 2024 | The ART of LLM Refinement: Ask, Refine, and TrustabstractKumar Shridhar, Koustuv Sinha, Andrew Cohen, Tianlu Wang, Ping Yu, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz |
NAACL-HLT | 7 |
| 2024 | On Affine Homotopy between Language EncodersabstractPre-trained language encoders---functions that represent text as vectors---are an integral component of many NLP tasks.
We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar?
We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity---the performance on downstream tasks.
It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation.
In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity.
We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity.
We confirm this on datasets of natural language representations.
Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them. Robin Shing Moon Chan, Reda Boumasmoud, Anej Svete, Qipeng Guo, Zhijing Jin 0001, Shauli Ravfogel, Mrinmaya Sachan, Bernhard Schölkopf, Mennatallah El-Assady, Ryan Cotterell |
NeurIPS | 8 |
| 2024 | Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM AgentsabstractAs AI systems pervade human life, ensuring that large language models (LLMs) make safe decisions remains a significant challenge. We introduce the Governance of the Commons Simulation (GovSim), a generative simulation platform designed to study strategic interactions and cooperative decision-making in LLMs. In GovSim, a society of AI agents must collectively balance exploiting a common resource with sustaining it for future use. This environment enables the study of how ethical considerations, strategic planning, and negotiation skills impact cooperative outcomes. We develop an LLM-based agent architecture and test it with the leading open and closed LLMs. We find that all but the most powerful LLM agents fail to achieve a sustainable equilibrium in GovSim, with the highest survival rate below 54%. Ablations reveal that successful multi-agent communication between agents is critical for achieving cooperation in these cases. Furthermore, our analyses show that the failure to achieve sustainable cooperation in most LLMs stems from their inability to formulate and analyze hypotheses about the long-term effects of their actions on the equilibrium of the group. Finally, we show that agents that leverage "Universalization"-based reasoning, a theory of moral thinking, are able to achieve significantly better sustainability. Taken together, GovSim enables us to study the mechanisms that underlie sustainable self-government with specificity and scale. We open source the full suite of our research results, including the simulation environment, agent prompts, and a comprehensive web interface. Giorgio Piatti, Zhijing Jin 0001, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, Rada Mihalcea |
NeurIPS | 5 |
| 2024 | Confidence Regulation Neurons in Language ModelsabstractDespite their widespread use, the mechanisms by which large language models (LLMs) represent and regulate uncertainty in next-token predictions remain largely unexplored. This study investigates two critical components believed to influence this uncertainty: the recently discovered entropy neurons and a new set of components that we term token frequency neurons. Entropy neurons are characterized by an unusually high weight norm and influence the final layer normalization (LayerNorm) scale to effectively scale down the logits. Our work shows that entropy neurons operate by writing onto an \textit{unembedding null space}, allowing them to impact the residual stream norm with minimal direct effect on the logits themselves. We observe the presence of entropy neurons across a range of models, up to 7 billion parameters. On the other hand, token frequency neurons, which we discover and describe here for the first time, boost or suppress each token’s logit proportionally to its log frequency, thereby shifting the output distribution towards or away from the unigram distribution. Finally, we present a detailed case study where entropy neurons actively manage confidence: the setting of induction, i.e. detecting and continuing repeated subsequences. Alessandro Stolfo, Ben Wu 0001, Wes Gurnee, Yonatan Belinkov, Xingyi Song, Mrinmaya Sachan, Neel Nanda |
NeurIPS | 6 |
| 2023 | Adaptive and Personalized Exercise Generation for Online Language LearningabstractAdaptive learning aims to provide customized educational activities (e.g., exercises) to address individual learning needs.However, manual construction and delivery of such activities is a laborious process.Thus, in this paper, we study a novel task of adaptive and personalized exercise generation for online language learning.To this end, we combine a knowledge tracing model that estimates each student's evolving knowledge states from their learning history and a controlled text generation model that generates exercise sentences based on the student's current estimated knowledge state and instructor requirements of desired properties (e.g., domain knowledge and difficulty).We train and evaluate our model on real-world learner interaction data from Duolingo and demonstrate that LMs guided by student states can generate superior exercises.Then, we discuss the potential use of our model in educational applications using various simulations.These simulations show that our model can adapt to students' individual abilities and can facilitate their learning efficiency by personalizing learning sequences.1 Peng Cui 0006, Mrinmaya Sachan |
ACL (1) | 2 |
| 2023 | Discourse-Centric Evaluation of Document-level Machine Translation with a New Densely Annotated Parallel Corpus of NovelsabstractYuchen Eleanor Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yuchen Eleanor Jiang, Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Mrinmaya Sachan, Ryan Cotterell |
ACL (1) | 5 |
| 2023 | XDailyDialog: A Multilingual Parallel Dialogue CorpusabstractZeming Liu, Ping Nie, Jie Cai, Haifeng Wang, Zheng-Yu Niu, Peng Zhang, Mrinmaya Sachan, Kaiping Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zeming Liu, Ping Nie, Haifeng Wang 0001, Zhengyu Niu, Mrinmaya Sachan, Kaiping Peng |
ACL (1) | 7 |
| 2023 | When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPabstractMulti-task learning (MTL) aims at achieving a better model by leveraging data and knowledge from multiple tasks.However, MTL does not always work -sometimes negative transfer occurs between tasks, especially when aggregating loosely related skills, leaving it an open question when MTL works.Previous studies show that MTL performance can be improved by algorithmic tricks.However, what tasks and skills should be included is less well explored.In this work, we conduct a case study in Financial NLP where multiple datasets exist for skills relevant to the domain, such as numeric reasoning and sentiment analysis.Due to the task difficulty and data scarcity in the Financial NLP domain, we explore when aggregating such diverse skills from multiple datasets with MTL can work.Our findings suggest that the key to MTL success lies in skill diversity, relatedness between tasks, and choice of aggregation size and shared capacity.Specifically, MTL works well when tasks are diverse but related, and when the size of the task aggregation and the shared capacity of the model are balanced to avoid overwhelming certain tasks. 1 Jingwei Ni, Zhijing Jin 0001, Mrinmaya Sachan, Markus Leippold |
ACL (1) | 4 |
| 2023 | A Causal Framework to Quantify the Robustness of Mathematical Reasoning with Language ModelsabstractAlessandro Stolfo, Zhijing Jin, Kumar Shridhar, Bernhard Schoelkopf, Mrinmaya Sachan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Alessandro Stolfo, Zhijing Jin 0001, Kumar Shridhar, Bernhard Schölkopf, Mrinmaya Sachan |
ACL (1) | 5 |
| 2023 | Tokenization and the Noiseless ChannelabstractVilém Zouhar, Clara Meister, Juan Gastaldi, Li Du, Mrinmaya Sachan, Ryan Cotterell. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Vilém Zouhar, Clara Meister, Juan Luis Gastaldi, Mrinmaya Sachan, Ryan Cotterell |
ACL (1) | 5 |
| 2023 | Automatic Educational Question Generation with Difficulty Level Controls
Ying Jiao, Kumar Shridhar, Peng Cui 0006, Wangchunshu Zhou, Mrinmaya Sachan |
AIED | 5 |
| 2023 | Opportunities and Challenges in Neural Dialog TutoringabstractJakub Macina, Nico Daheim, Lingzhi Wang, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jakub Macina, Nico Daheim, Lingzhi Wang 0001, Tanmay Sinha, Manu Kapur, Iryna Gurevych, Mrinmaya Sachan |
EACL | 7 |
| 2023 | Poor Man's Quality Estimation: Predicting Reference-Based MT Metrics Without the ReferenceabstractVilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Vilém Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, Mrinmaya Sachan |
EACL | 7 |
| 2023 | Linear-Time Modeling of Linguistic Structure: An Order-Theoretic PerspectiveabstractTasks that model the relation between pairs of tokens in a string are a vital part of understanding natural language.Such tasks, in general, require exhaustive pair-wise comparisons of tokens, thus having a quadratic runtime complexity in the length of the string.We show that these exhaustive comparisons can be avoided, and, moreover, the complexity of such tasks can be reduced to linear by casting the relation between tokens as a partial order over the string.Our method predicts real numbers for each token in a string in parallel and sorts the tokens accordingly, resulting in total orders of the tokens in the string.Each total order implies a set of arcs oriented from smaller to greater tokens, sorted by their predicted numbers.The intersection of total orders results in a partial order over the set of tokens in the string, which is then decoded into a directed graph representing the desired linguistic structure.Our experiments on dependency parsing and coreference resolution show that our method achieves state-of-the-art or comparable performance.Moreover, the linear complexity and parallelism of our method double the speed of graph-based coref- erence resolution models, and bring a 10-times speed-up over graph-based dependency parsers. Tianyu Liu 0004, Afra Amini, Mrinmaya Sachan, Ryan Cotterell |
EMNLP | 3 |
| 2023 | A Diachronic Perspective on User Trust in AI under UncertaintyabstractIn a human-AI collaboration, users build a mental model of the AI system based on its reliability and how it presents its decision, e.g. its presentation of system confidence and an explanation of the output.Modern NLP systems are often uncalibrated, resulting in confidently incorrect predictions that undermine user trust.In order to build trustworthy AI, we must understand how user trust is developed and how it can be regained after potential trust-eroding events.We study the evolution of user trust in response to these trust-eroding events using a betting game.We find that even a few incorrect instances with inaccurate confidence estimates damage user trust and performance, with very slow recovery.We also show that this degradation in trust reduces the success of human-AI collaboration and that different types of miscalibration-unconfidently correct and confidently incorrect-have different negative effects on user trust.Our findings highlight the importance of calibration in user-facing AI applications and shed light on what aspects help users decide whether to trust the AI system. Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, Mrinmaya Sachan |
EMNLP | 4 |
| 2023 | Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language ModelsabstractYifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jiaoda Li, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, Mrinmaya Sachan |
EMNLP | 8 |
| 2023 | Enhancing Textbooks with Visuals from the Web for Improved LearningabstractTextbooks are one of the main mediums for delivering high-quality education to students.In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge.However, many textbooks lack these interesting visuals to support student learning.In this paper, we investigate the effectiveness of vision-language models to automatically enhance textbooks with images from the web.We collect a dataset of e-textbooks in the math, science, social science and business domains.We then set up a text-image matching task that involves retrieving and appropriately assigning web images to textbooks, which we frame as a matching optimization problem.Through a crowd-sourced evaluation, we verify that (1) while the original textbook images are rated higher, automatically assigned ones are not far behind, and (2) the precise formulation of the optimization problem matters.We release the dataset of textbooks with an associated image bank to inspire further research in this intersectional area of computer vision and NLP for education. Janvijay Singh, Vilém Zouhar, Mrinmaya Sachan |
EMNLP | 3 |
| 2023 | Revisiting Automated Topic Model Evaluation with Large Language ModelsabstractTopic models help make sense of large text collections.Automatically evaluating their output and determining the optimal number of topics are both longstanding challenges, with no effective automated solutions to date.This paper evaluates the effectiveness of large language models (LLMs) for these tasks.We find that LLMs appropriately assess the resulting topics, correlating more strongly with human judgments than existing automated metrics.However, the type of evaluation task matters -LLMs correlate better with coherence ratings of word sets than on a word intrusion task.We find that LLMs can also guide users toward a reasonable number of topics.In actual applications, topic models are typically used to answer a research question related to a collection of texts.We can incorporate this research question in the prompt to the LLM, which helps estimate the optimal number of topics.github.com/dominiksinsaarland/ evaluating-topic-model-output Dominik Stammbach, Vilém Zouhar, Alexander Miserlis Hoyle, Mrinmaya Sachan, Elliott Ash |
EMNLP | 4 |
| 2023 | A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation AnalysisabstractMathematical reasoning in large language models (LMs) has garnered significant attention in recent work, but there is a limited understanding of how these models process and store information related to arithmetic tasks within their architecture.In order to improve our understanding of this aspect of language models, we present a mechanistic interpretation of Transformer-based LMs on arithmetic questions using a causal mediation analysis framework.By intervening on the activations of specific model components and measuring the resulting changes in predicted probabilities, we identify the subset of parameters responsible for specific predictions.This provides insights into how information related to arithmetic is processed by LMs.Our experimental results indicate that LMs process the input by transmitting the information relevant to the query from mid-sequence early layers to the final token using the attention mechanism.Then, this information is processed by a set of MLP modules, which generate result-related information that is incorporated into the residual stream.To assess the specificity of the observed activation dynamics, we compare the effects of different model components on arithmetic queries with other tasks, including number retrieval from prompts and factual knowledge questions. 1 Alessandro Stolfo, Yonatan Belinkov, Mrinmaya Sachan |
EMNLP | 3 |
| 2023 | Infusing Lattice Symmetry Priors in Attention Mechanisms for Sample-Efficient Abstract Geometric ReasoningabstractThe Abstraction and Reasoning Corpus (ARC) (Chollet, 2019) and its most recent language-complete instantiation (LARC) has been postulated as an important step towards general AI. Yet, even state-of-the-art machine learning models struggle to achieve meaningful performance on these problems, falling behind non-learning based approaches. We argue that solving these tasks requires extreme generalization that can only be achieved by proper accounting for core knowledge priors. As a step towards this goal, we focus on geometry priors and introduce LatFormer, a model that incorporates lattice symmetry priors in attention masks. We show that, for any transformation of the hypercubic lattice, there exists a binary attention mask that implements that group action. Hence, our study motivates a modification to the standard attention mechanism, where attention weights are scaled using soft masks generated by a convolutional network. Experiments on synthetic geometric reasoning show that LatFormer requires 2 orders of magnitude fewer data than standard attention and transformers. Moreover, our results on ARC and LARC tasks that incorporate geometric priors provide preliminary evidence that these complex datasets do not lie out of the reach of deep learning models. Mattia Atzeni, Mrinmaya Sachan, Andreas Loukas |
ICML | 2 |
| 2023 | Controlled Text Generation with Natural Language InstructionsabstractLarge language models can be prompted to pro- duce fluent output for a wide range of tasks without being specifically trained to do so. Nevertheless, it is notoriously difficult to control their generation in such a way that it satisfies user-specified constraints. In this paper, we present InstructCTG, a simple controlled text generation framework that incorporates different constraints by verbalizing them as natural language instructions. We annotate natural texts through a combination of off-the-shelf NLP tools and simple heuristics with the linguistic and extra-linguistic constraints they satisfy. Then, we verbalize the constraints into natural language instructions to form weakly supervised training data, i.e., we prepend the natural language verbalizations of the constraints in front of their corresponding natural language sentences. Next, we fine-tune a pre-trained language model on the augmented corpus. Compared to existing methods, InstructCTG is more flexible in terms of the types of constraints it allows the practitioner to use. It also does not require any modification of the decoding procedure. Finally, InstructCTG allows the model to adapt to new constraints without re-training through the use of in-context learning. Wangchunshu Zhou, Yuchen Eleanor Jiang, Ethan Wilcox, Ryan Cotterell, Mrinmaya Sachan |
ICML | 5 |
| 2023 | CLadder: A Benchmark to Assess Causal Reasoning Capabilities of Language Models
Zhijing Jin 0001, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez Adauto, Max Kleiman-Weiner, Mrinmaya Sachan, Bernhard Schölkopf |
NeurIPS | 10 |
| 2022 | Deep Clustering of Text Representations for Supervision-Free Probing of SyntaxabstractWe explore deep clustering of multilingual text representations for unsupervised model interpretation and induction of syntax. As these representations are high-dimensional, out-of-the-box methods like K-means do not work well. Thus, our approach jointly transforms the representations into a lower-dimensional cluster-friendly space and clusters them. We consider two notions of syntax: Part of Speech Induction (POSI) and Constituency Labelling (CoLab) in this work. Interestingly, we find that Multilingual BERT (mBERT) contains surprising amount of syntactic knowledge of English; possibly even as much as English BERT (E-BERT). Our model can be used as a supervision-free probe which is arguably a less-biased way of probing. We find that unsupervised probes show benefits from higher layers as compared to supervised probes. We further note that our unsupervised probe utilizes E-BERT and mBERT representations differently, especially for POSI. We validate the efficacy of our probe by demonstrating its capabilities as a unsupervised syntax induction technique. Our probe works well for both syntactic formalisms by simply adapting the input representations. We report competitive performance of our probe on 45-tag English POSI, state-of-the-art performance on 12-tag POSI across 10 languages, and competitive results on CoLab. We also perform zero-shot syntax induction on resource impoverished languages and report strong results. Vikram Gupta, Freda Shi, Kevin Gimpel, Mrinmaya Sachan |
AAAI | 4 |
| 2022 | Slangvolution: A Causal Analysis of Semantic Change and Frequency Dynamics in SlangabstractLanguages are continuously undergoing changes, and the mechanisms that underlie these changes are still a matter of debate.In this work, we approach language evolution through the lens of causality in order to model not only how various distributional factors associate with language change, but how they causally affect it.In particular, we study slang, which is an informal language that is typically restricted to a specific group or social setting.We analyze the semantic change and frequency shift of slang words and compare them to those of standard, nonslang words.With causal discovery and causal inference techniques, we measure the effect that word type (slang/nonslang) has on both semantic change and frequency shift, as well as its relationship to frequency, polysemy and part of speech.Our analysis provides some new insights in the study of language change, e.g., we show that slang words undergo less semantic change but tend to have larger frequency shifts over time. 1 Daphna Keidar, Andreas Opedal, Zhijing Jin 0001, Mrinmaya Sachan |
ACL (1) | 4 |
| 2022 | Beyond prompting: Making Pre-trained Language Models Better Zero-shot Learners by Clustering RepresentationsabstractRecent work has demonstrated that pre-trained language models (PLMs) are zero-shot learners.However, most existing zero-shot methods involve heavy human engineering or complicated self-training pipelines, hindering their application to new situations.In this work, we show that zero-shot text classification can be improved simply by clustering texts in the embedding spaces of PLMs.Specifically, we fit the unlabeled texts with a Bayesian Gaussian Mixture Model after initializing cluster positions and shapes using class names.Despite its simplicity, this approach achieves superior or comparable performance on both topic and sentiment classification datasets and outperforms prior works significantly on unbalanced datasets.We further explore the applicability of our clustering approach by evaluating it on 14 datasets with more diverse topics, text lengths, and numbers of classes.Our approach achieves an average of 20% absolute improvement over prompt-based zero-shot learning.Finally, we compare different PLM embedding spaces and find that texts are well-clustered by topics even if the PLM is not explicitly pre-trained to generate meaningful sentence embeddings.This work indicates that PLM embeddings can categorize texts without task-specific fine-tuning, thus providing a new way to analyze and utilize their knowledge and zero-shot learning ability 1 . Ping Nie, Roger Wattenhofer, Mrinmaya Sachan |
EMNLP | 5 |
| 2022 | Differentially Private Language Models for Secure Data SharingabstractTo protect the privacy of individuals whose data is being shared, it is of high importance to develop methods allowing researchers and companies to release textual data while providing formal privacy guarantees to its originators.In the field of NLP, substantial efforts have been directed at building mechanisms following the framework of local differential privacy, thereby anonymizing individual text samples before releasing them.In practice, these approaches are often dissatisfying in terms of the quality of their output language due to the strong noise required for local differential privacy.In this paper, we approach the problem at hand using global differential privacy, particularly by training a generative language model in a differentially private manner and consequently sampling data from it.Using natural language prompts and a new prompt-mismatch loss, we are able to create highly accurate and fluent textual datasets taking on specific desired attributes such as sentiment or topic and resembling statistical properties of the training data.We perform thorough experiments indicating that our synthetic datasets do not leak information from our original data and are of high language quality and highly suitable for training models for further analysis on real-world data.Notably, we also demonstrate that training classifiers on private synthetic data outperforms directly training classifiers on real data with DP-SGD. 1 Justus Mattern, Zhijing Jin 0001, Benjamin Weggenmann, Bernhard Schölkopf, Mrinmaya Sachan |
EMNLP | 5 |
| 2022 | Automatic Generation of Socratic Subquestions for Teaching Math Word ProblemsabstractSocratic questioning is an educational method that allows students to discover answers to complex problems by asking them a series of thoughtful questions.Generation of didactically sound questions is challenging, requiring understanding of the reasoning process involved in the problem.We hypothesize that such questioning strategy can not only enhance the human performance, but also assist the math word problem (MWP) solvers.In this work, we explore the ability of large language models (LMs) in generating sequential questions for guiding math word problem-solving.We propose various guided question generation schemes based on input conditioning and reinforcement learning.On both automatic and human quality evaluations, we find that LMs constrained with desirable question properties generate superior questions and improve the overall performance of a math word problem solver.We conduct a preliminary user study to examine the potential value of such question generation models in the education domain.Results suggest that the difficulty level of problems plays an important role in determining whether questioning improves or hinders human performance.We discuss the future of using such questioning strategies in education.https://github.com/eth-nlped/ scaffolding-generation Kumar Shridhar, Jakub Macina, Mennatallah El-Assady, Tanmay Sinha, Manu Kapur, Mrinmaya Sachan |
EMNLP | 6 |
| 2022 | Case-based reasoning for better generalization in textual reinforcement learning
Mattia Atzeni, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya Sachan |
ICLR | 4 |
| 2022 | BlonDe: An Automatic Evaluation Metric for Document-level Machine TranslationabstractYuchen Jiang, Tianyu Liu, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tianyu Liu 0004, Shuming Ma, Dongdong Zhang 0001, Jian Yang 0030, Haoyang Huang, Rico Sennrich, Ryan Cotterell, Mrinmaya Sachan, Ming Zhou 0001 |
NAACL-HLT | 9 |
| 2022 | Probing via PromptingabstractProbing is a popular method to discern what linguistic information is contained in the representations of pre-trained language models.However, the mechanism of selecting the probe model has recently been subject to intense debate, as it is not clear if the probes are merely extracting information or modeling the linguistic property themselves.To address this challenge, this paper introduces a novel model-free approach to probing, by formulating probing as a prompting task.We conduct experiments on five probing tasks and show that our approach is comparable or better at extracting information than diagnostic probes while learning much less on its own.We further combine the probing via prompting approach with attention head pruning to analyze where the model stores the linguistic information in its architecture.We then examine the usefulness of a specific linguistic property for pre-training by removing the heads that are essential to that property and evaluating the resulting model's performance on language modeling. Jiaoda Li, Ryan Cotterell, Mrinmaya Sachan |
NAACL-HLT | 3 |
| 2022 | A Structured Span SelectorabstractTianyu Liu, Yuchen Jiang, Ryan Cotterell, Mrinmaya Sachan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tianyu Liu 0004, Yuchen Eleanor Jiang, Ryan Cotterell, Mrinmaya Sachan |
NAACL-HLT | 4 |
| 2022 | Original or Translated? A Causal Analysis of the Impact of Translationese on Machine Translation PerformanceabstractJingwei Ni, Zhijing Jin, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jingwei Ni, Zhijing Jin 0001, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf |
NAACL-HLT | 4 |
| 2022 | When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentabstractAI systems are becoming increasingly intertwined with human life. In order to effectively collaborate with humans and ensure safety, AI systems need to be able to understand, interpret and predict human moral judgments and decisions. Human moral judgments are often guided by rules, but not always. A central challenge for AI safety is capturing the flexibility of the human moral mind — the ability to determine when a rule should be broken, especially in novel or unusual situations. In this paper, we present a novel challenge set consisting of moral exception question answering (MoralExceptQA) of cases that involve potentially permissible moral exceptions – inspired by recent moral psychology studies. Using a state-of-the-art large language model (LLM) as a basis, we propose a novel moral chain of thought (MoralCoT) prompting strategy that combines the strengths of LLMs with theories of moral reasoning developed in cognitive science to predict human moral judgments. MoralCoT outperforms seven existing LLMs by 6.2% F1, suggesting that modeling human reasoning might be necessary to capture the flexibility of the human moral mind. We also conduct a detailed error analysis to suggest directions for future work to improve AI safety using MoralExceptQA. Our data is open-sourced at https://huggingface.co/datasets/feradauto/MoralExceptQA and code at https://github.com/feradauto/MoralCoT. Zhijing Jin 0001, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, Bernhard Schölkopf |
NeurIPS | 6 |
| 2021 | Text-based RL Agents with Commonsense Knowledge: New Challenges, Environments and BaselinesabstractText-based games have emerged as an important test-bed for Reinforcement Learning (RL) research, requiring RL agents to combine grounded language understanding with sequential decision making. In this paper, we examine the problem of infusing RL agents with commonsense knowledge. Such knowledge would allow agents to efficiently act in the world by pruning out implausible actions, and to perform look-ahead planning to determine how current actions might affect future world states. We design a new text-based gaming environment called TextWorld Commonsense (TWC) for training and evaluating RL agents with a specific kind of commonsense knowledge about objects, their attributes, and affordances. We also introduce several baseline RL agents which track the sequential context and dynamically retrieve the relevant commonsense knowledge from ConceptNet. We show that agents which incorporate commonsense knowledge in TWC perform better, while acting more efficiently. We conduct user-studies to estimate human performance on TWC and show that there is ample room for future improvement. Keerthiram Murugesan, Mattia Atzeni, Pavan Kapanipathi, Pushkar Shukla, Sadhana Kumaravel, Gerald Tesauro, Kartik Talamadupula, Mrinmaya Sachan, Murray Campbell |
AAAI | 8 |
| 2021 | Bird's Eye: Probing for Linguistic Graph Structures with a Simple Information-Theoretic ApproachabstractYifan Hou, Mrinmaya Sachan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Mrinmaya Sachan |
ACL/IJCNLP (1) | 2 |
| 2021 | Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPabstractZhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schoelkopf. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhijing Jin 0001, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schölkopf |
EMNLP (1) | 6 |
| 2021 | Differentiable Subset Pruning of Transformer HeadsabstractAbstract Multi-head attention, a collection of several attention mechanisms that independently attend to different parts of the input, is the key ingredient in the Transformer. Recent work has shown, however, that a large proportion of the heads in a Transformer’s multi-head attention mechanism can be safely pruned away without significantly harming the performance of the model; such pruning leads to models that are noticeably smaller and faster in practice. Our work introduces a new head pruning technique that we term differentiable subset pruning. ntuitively, our method learns per- head importance variables and then enforces a user-specified hard constraint on the number of unpruned heads. he importance variables are learned via stochastic gradient descent. e conduct experiments on natural language inference and machine translation; we show that differentiable subset pruning performs comparably or better than previous works while offering precise control of the sparsity level.1 Jiaoda Li, Ryan Cotterell, Mrinmaya Sachan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Knowledge Graph Embedding CompressionabstractKnowledge graph (KG) representation learning techniques that learn continuous embeddings of entities and relations in the KG have become popular in many AI applications.With a large KG, the embeddings consume a large amount of storage and memory.This is problematic and prohibits the deployment of these techniques in many real world settings.Thus, we propose an approach that compresses the KG embedding layer by representing each entity in the KG as a vector of discrete codes and then composes the embeddings from these codes.The approach can be trained end-toend with simple modifications to any existing KG embedding technique.We evaluate the approach on various standard KG embedding evaluations and show that it achieves 50-1000x compression of embeddings with a minor loss in performance.The compressed embeddings also retain the ability to perform various reasoning tasks such as KG inference. Mrinmaya Sachan |
ACL | 1 |
| 2019 | Discourse in Multimedia: A Case Study in Extracting Geometry Knowledge from TextbooksabstractTo ensure readability, text is often written and presented with due formatting. These text formatting devices help the writer to effectively convey the narrative. At the same time, these help the readers pick up the structure of the discourse and comprehend the conveyed information. There have been a number of linguistic theories on discourse structure of text. However, these theories only consider unformatted text. Multimedia text contains rich formatting features that can be leveraged for various NLP tasks. In this article, we study some of these discourse features in multimedia text and what communicative function they fulfill in the context. As a case study, we use these features to harvest structured subject knowledge of geometry from textbooks. We conclude that the discourse and text layout features provide information that is complementary to lexical semantic information. Finally, we show that the harvested structured knowledge can be used to improve an existing solver for geometry problems, making it more accurate as well as more explainable. Mrinmaya Sachan, Avinava Dubey, Eduard H. Hovy, Tom M. Mitchell, Dan Roth 0001, Eric P. Xing |
Comput. Linguistics | 1 |
| 2018 | Contextual Parameter Generation for Universal Neural Machine TranslationabstractWe propose a simple modification to existing neural machine translation (NMT) models that enables using a single universal model to translate between multiple languages while allowing for language specific parameterization, and that can also be used for domain adaptation.Our approach requires no changes to the model architecture of a standard NMT system, but instead introduces a new component, the contextual parameter generator (CPG), that generates the parameters of the system (e.g., weights in a neural network).This parameter generator accepts source and target language embeddings as input, and generates the parameters for the encoder and the decoder, respectively.The rest of the model remains unchanged and is shared across all languages.We show how this simple modification enables the system to use monolingual data for training and also perform zero-shot translation.We further show it is able to surpass state-of-theart performance for both the IWSLT-15 and IWSLT-17 datasets and that the learned language embeddings are able to uncover interesting relationships between languages. Emmanouil A. Platanios, Mrinmaya Sachan, Graham Neubig, Tom M. Mitchell |
EMNLP | 2 |
| 2018 | Parsing to Programs: A Framework for Situated QAabstractThis paper introduces Parsing to Programs, a framework that combines ideas from parsing and probabilistic programming for situated question answering. As a case study, we build a system that solves pre-university level Newtonian physics questions. Our approach represents domain knowledge of Newtonian physics as programs. When presented with a novel question, the system learns a formal representation of the question by combining interpretations from the question text and any associated diagram. Finally, the system uses this formal representation to solve the questions using the domain knowledge. We collect a new dataset of Newtonian physics questions from a number of textbooks and use it to train our system. The system achieves near human performance on held-out textbook questions and section 1 of AP Physics C mechanics - both on practice questions as well as on freely available actual exams held in 1998 and 2012. Mrinmaya Sachan, Eric P. Xing |
KDD | 1 |
| 2018 | Self-Training for Jointly Learning to Ask and Answer QuestionsabstractMrinmaya Sachan, Eric Xing. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mrinmaya Sachan, Eric P. Xing |
NAACL-HLT | 1 |
| 2018 | Learning Pipelines with Limited Data and Domain Knowledge: A Study in Parsing Physics ProblemsabstractAs machine learning becomes more widely used in practice, we need new methods to build complex intelligent systems that integrate learning with existing software, and with domain knowledge encoded as rules. As a case study, we present such a system that learns to parse Newtonian physics problems in textbooks. This system, Nuts&Bolts, learns a pipeline process that incorporates existing code, pre-learned machine learning models, and human engineered rules. It jointly trains the entire pipeline to prevent propagation of errors, using a combination of labelled and unlabelled data. Our approach achieves a good performance on the parsing task, outperforming the simple pipeline and its variants. Finally, we also show how Nuts&Bolts can be used to achieve improvements on a relation extraction task and on the end task of answering Newtonian physics problems. Mrinmaya Sachan, Avinava Dubey, Tom M. Mitchell, Dan Roth 0001, Eric P. Xing |
NeurIPS | 1 |
| 2017 | From Textbooks to Knowledge: A Case Study in Harvesting Axiomatic Knowledge from Textbooks to Solve Geometry ProblemsabstractTextbooks are rich sources of knowledge.Harvesting knowledge from textbooks is a key challenge in many educational applications.In this paper, we present an approach to obtain axiomatic knowledge of geometry in the form of horn-clause rules from math textbooks.The approach uses rich contextual and typographical features extracted from the textbooks.It also leverages the redundancy and shared ordering of axioms across multiple textbooks to accurately harvest axioms.These axioms are then parsed into horn-clause rules that are used to improve the state-of-the-art in solving geometry problems. Mrinmaya Sachan, Avinava Dubey, Eric P. Xing |
EMNLP | 1 |
| 2016 | Easy Questions First? A Case Study on Curriculum Learning for Question AnsweringabstractCognitive science researchers have emphasized the importance of ordering a complex task into a sequence of easy to hard problems.Such an ordering provides an easier path to learning and increases the speed of acquisition of the task compared to conventional learning.Recent works in machine learning have explored a curriculum learning approach called selfpaced learning which orders data samples on the easiness scale so that easy samples can be introduced to the learning algorithm first and harder samples can be introduced successively.We introduce a number of heuristics that improve upon selfpaced learning.Then, we argue that incorporating easy, yet, a diverse set of samples can further improve learning.We compare these curriculum learning proposals in the context of four non-convex models for QA and show that they lead to real improvements in each of them. Mrinmaya Sachan, Eric P. Xing |
ACL (1) | 1 |
| 2016 | Learning Concept Taxonomies from Multi-modal DataabstractWe study the problem of automatically building hypernym taxonomies from textual and visual data.Previous works in taxonomy induction generally ignore the increasingly prominent visual data, which encode important perceptual semantics.Instead, we propose a probabilistic model for taxonomy induction by jointly leveraging text and images.To avoid hand-crafted feature engineering, we design end-to-end features based on distributed representations of images and words.The model is discriminatively trained given a small set of existing ontologies and is capable of building full taxonomies from scratch for a collection of unseen conceptual label items with associated images.We evaluate our model and features on the WordNet hierarchies, where our system outperforms previous approaches by a large gap. Hao Zhang 0025, Zhiting Hu, Yuntian Deng, Mrinmaya Sachan, Zhicheng Yan 0001, Eric P. Xing |
ACL (1) | 4 |
| 2016 | Grounding Topic Models with Knowledge Bases
Zhiting Hu, Mrinmaya Sachan, Eric P. Xing, Zaiqing Nie |
IJCAI | 3 |
| 2015 | Learning Answer-Entailing Structures for Machine ComprehensionabstractMrinmaya Sachan, Kumar Dubey, Eric Xing, Matthew Richardson. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mrinmaya Sachan, Avinava Dubey, Eric P. Xing, Matthew Richardson |
ACL (1) | 1 |
| 2015 | An Active Learning Approach to Coreference Resolution
Mrinmaya Sachan, Eduard H. Hovy, Eric P. Xing |
IJCAI | 1 |
| 2014 | Spatial compactness meets topical consistency: jointly modeling links and content for community detectionabstractIn this paper, we address the problem of discovering topically meaningful, yet compact (densely connected) communities in a social network. Assuming the social network to be an integer-weighted graph (where the weights can be intuitively defined as the number of common friends, followers, documents exchanged, etc.), we transform the social network to a more efficient representation. In this new representation, each user is a bag of her one-hop neighbors. We propose a mixed-membership model to identify compact communities using this transformation. Next, we augment the representation and the model to incorporate user-content information imposing topical consistency in the communities. In our model a user can belong to multiple communities and a community can participate in multiple topics. This allows us to discover community memberships as well as community and user interests. Our method outperforms other well known baselines on two real-world social networks. Finally, we also provide a fast, parallel approximation of the same. Mrinmaya Sachan, Avinava Dubey, Eric P. Xing, Eduard H. Hovy |
WSDM | 1 |
| 2012 | Using content and interactions for discovering communities in social networksabstractIn recent years, social networking sites have not only enabled people to connect with each other using social links but have also allowed them to share, communicate and interact over diverse geographical regions. Social network provide a rich source of heterogeneous data which can be exploited to discover previously unknown relationships and interests among groups of people. In this paper, we address the problem of discovering topically meaningful communities from a social network. We assume that a persons' membership in a community is conditioned on its social relationship, the type of interaction and the information communicated with other members of that community. We propose generative models that can discover communities based on the discussed topics, interaction types and the social connections among people. In our models a person can belong to multiple communities and a community can participate in multiple topics. This allows us to discover both community interests and user interests based on the information and linked associations. We demonstrate the effectiveness of our model on two real word data sets and show that it performs better than existing community discovery models. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
WWW | 1 |
| 2011 | Probabilistic model for discovering topic based communities in social networksabstractSocial graphs have received renewed interest as a research topic with the advent of social networking websites. These online networks provide a rich source of data to study user relationships and interaction patterns on a large scale. In this paper, we propose a generative Bayesian model for extracting latent communities from a social graph. We assume that community memberships depend on topics of interest between users and the link relationships between them in the social graph topology. In addition, we make use of the nature of interaction to gauge user interests. Our model allows communities to be related to multiple topics and each user in the graph can be a member of multiple communities. This gives an insight into user interests and topical distribution in communities. We show the effectiveness of our model using a real world data set and also compare our model with existing community discovery methods. Mrinmaya Sachan, Danish Contractor, Tanveer A. Faruquie, L. Venkata Subramaniam |
CIKM | 1 |
| 2011 | Using Text Reviews for Product Entity Completion
Mrinmaya Sachan, Tanveer A. Faruquie, L. Venkata Subramaniam, Mukesh K. Mohania |
IJCNLP | 1 |