Felix J. Binder

dblp:294/8679 · also Felix Jedidja Binder · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Looking Inward: Language Models Can Learn About Themselves by Introspection
abstract
Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g. thoughts and feelings) that are not accessible to external observers. Do LLMs have this introspective capability of privileged access? If they do, this would show that LLMs can acquire knowledge not contained in or inferable from training data. We investigate LLMs predicting properties of their own behavior in hypothetical situations. If a model M1 has this capability, it should outperform a different model M2 in predicting M1's behavior—even if M2 is trained on M1's ground-truth behavior. The idea is that M1 has privileged access to its own behavioral tendencies, and this enables it to predict itself better than M2 (even if M2 is generally stronger). In experiments with GPT-4, GPT-4o, and Llama-3 models, we find that the model M1 outperforms M2 in predicting itself, providing evidence for privileged access. Further experiments and ablations provide additional evidence. Our results show that LLMs can offer reliable self-information independent of external data in certain domains. By demonstrating this, we pave the way for further work on introspection in more practical domains, which would have significant implications for model transparency and explainability. However, while we successfully show introspective capabilities in simple tasks, we are unsuccessful on more complex tasks or those requiring out-of-distribution generalization.
Felix J. Binder, James Chua, Tomek Korbak, Henry Sleight, Robert Long, Ethan Perez, Miles Turpin, Owain Evans
ICLR1
2024 Probabilistic simulation supports generalizable intuitive physics
Khaled Jedoui, Rahul M. V., Felix J. Binder, Josh Tenenbaum, Judith E. Fan, Dan Yamins, Kevin A. Smith 0001
CogSci4
2024 Understanding Physical Dynamics with Counterfactual World Modeling
Rahul M. V., Kevin T. Feigelis, Daniel Bear, Khaled Jedoui, Klemen Kotar, Felix J. Binder, Wanhee Lee, Sherry Liu, Kevin A. Smith 0001, Judith E. Fan, Dan Yamins
ECCV (24)7
2023 Advancing Cognitive Science and AI with Cognitive-AI Benchmarking
Felix J. Binder, Logan Matthew Cross, Yoni Friedman, Robert D. Hawkins, Dan Yamins, Judith E. Fan
CogSci1
2023 Humans choose visual subgoals to reduce cognitive cost
Felix J. Binder, Marcelo G. Mattar, David Kirsh, Judith E. Fan
CogSci1
2023 Measuring and Modeling Physical Intrinsic Motivation
Julio Martinez, Felix J. Binder, Nick Haber, Judith E. Fan, Dan Yamins
CogSci2
2021 Cognitive cost and information gain trade off in a large-scale number guessing game
Felix J. Binder, Cameron Jones, Robert Kaufman 0001, Naomi T. Lin, Crystal R. Poole, Ed Vul
CogSci1
2021 Visual scoping operations for physical assembly
Felix J. Binder, Marcelo G. Mattar, David Kirsh, Judith E. Fan
CogSci1