Thomas Icard

dblp:32/8219 · also Thomas F. Icard III · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
16since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 5 since 2021Theory of computation · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Toward a Formal Pragmatics of Explanation
Jacqueline Harding, Tobias Gerstenberg, Thomas Icard
CogSci3
2025 Norms moderate causal judgments in cases of double prevention
Kevin O'Neill, Paul Henne, Tadeg Quillien, Thomas Icard, Felipe De Brigard
CogSci4
2025 Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
abstract
Interpretability research now offers a variety of techniques for identifying abstract internal mechanisms in neural networks. Can such techniques be used to predict how models will behave on out-of-distribution examples? In this work, we provide a positive answer to this question. Through a diverse set of language modeling tasks—including symbol manipulation, knowledge retrieval, and instruction following—we show that the most robust features for correctness prediction are those that play a distinctive causal role in the model’s behavior. Specifically, we propose two methods that leverage causal mechanisms to predict the correctness of model outputs: counterfactual simulation (checking whether key causal variables are realized) and value probing (using the values of those variables to make predictions). Both achieve high AUC-ROC in distribution and outperform methods that rely on causal-agnostic features in out-of-distribution settings, where predicting model behaviors is more crucial. Our work thus highlights a novel and significant application for internal causal analysis of language models.
Jing Huang 0014, Junyi Tao, Thomas Icard, Diyi Yang, Christopher Potts
ICML3
2025 Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
abstract
Causal abstraction provides a theoretical foundation for mechanistic interpretability, the field concerned with providing intelligible algorithms that are faithful simplifications of the known, but opaque low-level details of black box AI models. Our contributions are (1) generalizing the theory of causal abstraction from mechanism replacement (i.e., hard and soft interventions) to arbitrary mechanism transformation (i.e., functionals from old mechanisms to new mechanisms), (2) providing a flexible, yet precise formalization for the core concepts of polysemantic neurons, the linear representation hypothesis, modular features, and graded faithfulness, and (3) unifying a variety of mechanistic interpretability methods in the common language of causal abstraction, namely, activation and path patching, causal mediation analysis, causal scrubbing, causal tracing, circuit analysis, concept erasure, sparse autoencoders, differential binary masking, distributed alignment search, and steering.
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang 0014, Aryaman Arora, Zhengxuan Wu, Noah D. Goodman, Christopher Potts, Thomas Icard
J. Mach. Learn. Res.11
2024 Anticipating the Risks and Benefits of Counterfactual World Simulation Models (Extended Abstract)
abstract
This paper examines the transformative potential of Counterfactual World Simulation Models (CWSMs). CWSMs use pieces of multi-modal evidence, such as the CCTV footage or sound recordings of a road accident, to build a high-fidelity 3D reconstruction of the scene. They can also answer causal questions, such as whether the accident happened because the driver was speeding, by simulating what would have happened in relevant counterfactual situations. CWSMs will enhance our capacity to envision alternate realities and investigate the outcomes of counterfactual alterations to how events unfold. This also, however, raises questions about what alternative scenarios we should be considering and what to do with that knowledge. We present a normative and ethical framework that guides and constrains the simulation of counterfactuals. We address the challenge of ensuring fidelity in reconstructions while simultaneously preventing stereotype perpetuation during counterfactual simulations. We anticipate different modes of how users will interact with CWSMs and discuss how their outputs may be presented. Finally, we address the prospective applications of CWSMs in the legal domain, recognizing both their potential to revolutionize legal proceedings as well as the ethical concerns they engender. Anticipating a new type of AI, this paper seeks to illuminate a path forward for responsible and effective use of CWSMs.
Lara Kirfel, Rob MacCoun, Thomas Icard, Tobias Gerstenberg
AIES (1)3
2024 Do as I explain: Explanations communicate optimal interventions
Lara Kirfel, Jacqueline Harding, Jeong Yeon Shin, Cindy Xin, Thomas Icard, Tobias Gerstenberg
CogSci5
2024 Probing the quantitative-qualitative divide in probabilistic reasoning
abstract
This paper explores the space of (propositional) probabilistic logical languages, ranging from a purely ‘qualitative’ comparative language to a highly ‘quantitative’ language involving arbitrary polynomials over probability terms. While talk of qualitative vs. quantitative may be suggestive, we identify a robust and meaningful boundary in the space by distinguishing systems that encode (at most) additive reasoning from those that encode additive and multiplicative reasoning. The latter includes not only languages with explicit multiplication but also languages expressing notions of dependence and conditionality. We show that the distinction tracks a divide in computational complexity: additive systems remain complete for NP, while multiplicative systems are robustly complete for ∃R. We also address axiomatic questions, offering several new completeness results as well as a proof of non-finite-axiomatizability for comparative probability. Repercussions of our results for conceptual and empirical questions are addressed, and open problems are discussed.
Duligur Ibeling, Thomas Icard, Krzysztof Mierzewski, Milan Mossé
Ann. Pure Appl. Log.2
2023 A Semantics for Causing, Enabling, and Preventing Verbs Using Structural Causal Models
Angela Cao, Atticus Geiger, Elisa Kreiss, Thomas Icard, Tobias Gerstenberg
CogSci4
2023 Show and tell: Learning causal structures from observations and explanations
Andrew Nam, Christopher Hughes, Thomas Icard, Tobias Gerstenberg
CogSci3
2023 Comparing Causal Frameworks: Potential Outcomes, Structural Models, Graphs, and Abstractions
abstract
The aim of this paper is to make clear and precise the relationship between the Rubin causal model (RCM) and structural causal model (SCM) frameworks for causal inference. Adopting a neutral logical perspective, and drawing on previous work, we show what is required for an RCM to be representable by an SCM. A key result then shows that every RCM---including those that violate algebraic principles implied by the SCM framework---emerges as an abstraction of some representable RCM. Finally, we illustrate the power of this ameliorative perspective by pinpointing an important role for SCM principles in classic applications of RCMs; conversely, we offer a characterization of the algebraic constraints implied by a graph, helping to substantiate further comparisons between the two frameworks.
Duligur Ibeling, Thomas Icard
NeurIPS2
2023 Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
abstract
Obtaining human-interpretable explanations of large, general-purpose language models is an urgent goal for AI safety. However, it is just as important that our interpretability methods are faithful to the causal dynamics underlying model behavior and able to robustly generalize to unseen inputs. Distributed Alignment Search (DAS) is a powerful gradient descent method grounded in a theory of causal abstraction that uncovered perfect alignments between interpretable symbolic algorithms and small deep learning models fine-tuned for specific tasks. In the present paper, we scale DAS significantly by replacing the remaining brute-force search steps with learned parameters -- an approach we call Boundless DAS. This enables us to efficiently search for interpretable causal structure in large language models while they follow instructions. We apply Boundless DAS to the Alpaca model (7B parameters), which, off the shelf, solves a simple numerical reasoning problem. With Boundless DAS, we discover that Alpaca does this by implementing a causal model with two interpretable boolean variables. Furthermore, we find that the alignment of neural representations with these variables is robust to changes in inputs and instructions. These findings mark a first step toward deeply understanding the inner-workings of our largest and most widely deployed language models.
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, Noah D. Goodman
NeurIPS3
2022 Inducing Causal Structure for Interpretable Neural Networks
abstract
In many areas, we have well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g., a deterministic program or Bayesian network) with representations in a neural model and (2) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model.
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, Christopher Potts
ICML6
2022 Causal Distillation for Language Models
abstract
Zhengxuan Wu, Atticus Geiger, Joshua Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah Goodman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Zhengxuan Wu, Atticus Geiger, Josh Rozner, Elisa Kreiss, Hanson Lu, Thomas Icard, Christopher Potts, Noah D. Goodman
NAACL-HLT6
2021 Causal Abstractions of Neural Networks
abstract
Structural analysis methods (e.g., probing and feature attribution) are increasingly important tools for neural network analysis. We propose a new structural analysis method grounded in a formal theory of causal abstraction that provides rich characterizations of model-internal representations and their roles in input/output behavior. In this method, neural representations are aligned with variables in interpretable causal models, and then interchange interventions are used to experimentally verify that the neural representations have the causal properties of their aligned variables. We apply this method in a case study to analyze neural models trained on Multiply Quantified Natural Language Inference (MQNLI) corpus, a highly complex NLI dataset that was constructed with a tree-structured natural logic causal model. We discover that a BERT-based model with state-of-the-art performance successfully realizes parts of the natural logic model’s causal structure, whereas a simpler baseline model fails to show any such structure, demonstrating that neural representations encode the compositional structure of MQNLI examples.
Atticus Geiger, Hanson Lu, Thomas Icard, Christopher Potts
NeurIPS3
2021 A Topological Perspective on Causal Inference
abstract
This paper presents a topological learning-theoretic perspective on causal inference by introducing a series of topologies defined on general spaces of structural causal models (SCMs). As an illustration of the framework we prove a topological causal hierarchy theorem, showing that substantive assumption-free causal inference is possible only in a meager set of SCMs. Thanks to a known correspondence between open sets in the weak topology and statistically verifiable hypotheses, our results show that inductive assumptions sufficient to license valid causal inferences are statistically unverifiable in principle. Similar to no-free-lunch theorems for statistical inference, the present results clarify the inevitability of substantial assumptions for causal inference. An additional benefit of our topological approach is that it easily accommodates SCMs with infinitely many variables. We finally suggest that our framework may be helpful for the positive project of exploring and assessing alternative causal-inductive assumptions.
Duligur Ibeling, Thomas Icard
NeurIPS2
2021 Logics of imprecise comparative probability
Wesley H. Holliday, Thomas Icard
Int. J. Approx. Reason.3
2020 Probabilistic Reasoning Across the Causal Hierarchy
abstract
We propose a formalization of the three-tier causal hierarchy of association, intervention, and counterfactuals as a series of probabilistic logical languages. Our languages are of strictly increasing expressivity, the first capable of expressing quantitative probabilistic reasoning—including conditional independence and Bayesian inference—the second encoding do-calculus reasoning for causal effects, and the third capturing a fully expressive do-calculus for arbitrary counterfactual queries. We give a corresponding series of finitary axiomatizations complete over both structural causal models and probabilistic programs, and show that satisfiability and validity for each language are decidable in polynomial space.
Duligur Ibeling, Thomas Icard
AAAI2
2020 Learning from explanations
Lara Kirfel, Thomas Icard, Tobias Gerstenberg
CogSci2
2020 Intention as commitment toward time
Marc van Zee, Dragan Doder, Leon van der Torre, Mehdi Dastani, Thomas Icard, Eric Pacuit
Artif. Intell.5
2019 Inflated inflation and superseded supersession: testing counterfactual sampling accounts of causal strength judgments
Maureen Gill, Jonathan F. Kominsky, Joshua Knobe, Thomas Icard
CogSci4
2019 On Open-Universe Causal Reasoning
Duligur Ibeling, Thomas Icard
UAI2
2018 On the instrumental value of hypothetical and counterfactual thought
Thomas Icard, Fiery Cushman, Joshua Knobe
CogSci1
2018 On the Conditional Logic of Simulation Models
abstract
We propose analyzing conditional reasoning by appeal to a notion of intervention on a simulation program, formalizing and subsuming a number of approaches to conditional thinking in the recent AI literature. Our main results include a series of axiomatizations, allowing comparison between this framework and existing frameworks (normality-ordering models, causal structural equation models), and a complexity result establishing NP-completeness of the satisfiability problem. Perhaps surprisingly, some of the basic logical principles common to all existing approaches are invalidated in our causal simulation approach. We suggest that this additional flexibility is important in modeling some intuitive examples.
Duligur Ibeling, Thomas Icard
IJCAI2
2017 Preferential Structures for Comparative Probabilistic Reasoning
abstract
Qualitative and quantitative approaches to reasoning about uncertainty can lead to different logical systems for formalizing such reasoning, even when the language for expressing uncertainty is the same. In the case of reasoning about relative likelihood, with statements of the form φ
Matthew Harrison-Trainor, Wesley H. Holliday, Thomas Icard
AAAI3
2017 Beyond Almost-Sure Termination
Thomas Icard
CogSci1
2016 Causality, Normality, and Sampling Propensity
Thomas Icard, Joshua Knobe
CogSci1
2015 A Resource-Rational Approach to the Causal Frame Problem
Thomas Icard, Noah D. Goodman
CogSci1
2014 Toward Boundedly Rational Analysis
Thomas Icard
CogSci1
2012 A Uniform Logic of Information Dynamics
Wesley H. Holliday, Tomohiro Hoshi, Thomas Icard
Advances in Modal Logic3
2011 A Topological Study of the Closed Fragment of GLP
abstract
In this article, we study the canonical model for the closed fragment of GLP and establish its precise relationship with a universal model constructed by Ignatiev. In particular, we effectively characterize the canonical model in terms of a coordinate system based on sequences of ordinals up to ϵ0.We then define a simple topological model of this logic by defining a natural polytopology on the ordinal ϵ0 itself.
Thomas Icard
J. Log. Comput.1
2010 Moorean Phenomena in Epistemic Logic
Wesley H. Holliday, Thomas Icard
Advances in Modal Logic2
2010 Joint Revision of Beliefs and Intention
Thomas Icard, Eric Pacuit, Yoav Shoham
KR1