Ilker Yildirim

dblp:29/840 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-9072-0938ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 8 first-author · 10 since 2021
YearPublicationVenuePosition
2025 New Perspectives in Computational Modeling of Human Attention
Ilker Yildirim
CogSci1
2024 Generative Semantic Transformation Process: A Case Study in Goal Prediction via Online Bayesian Language Inference
Lorenss Martinsons, John Muchovej, Ilker Yildirim
CogSci3
2024 An Intuitive Physics Approach to Modeling Melodic Expectation
Breanna K. Nguyen, Ilker Yildirim
CogSci2
2023 Perception of Mooney Faces: Extreme Generalization through Inverse Rendering?
Shreya Kapoor, Maximilian Weiherer, Max H. Siegel, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum, Bernhard Egger 0001
CogSci5
2023 Where does the flow go? Humans automatically predict liquid pathing with coarse-grained simulation
Mario Belledonne, Tristan Yates, Ilker Yildirim
CogSci4
2023 When are Lemons Purple? The Concept Association Bias of Vision-Language Models
abstract
Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-totext retrieval.However, such performance does not realize in tasks that require a finergrained correspondence between vision and language, such as Visual Question Answering (VQA).As a potential cause of the difficulty of applying these models to VQA and similar tasks, we report an interesting phenomenon of vision-language models, which we call the Concept Association Bias (CAB).We find that models with CAB tend to treat input as a bag of concepts and attempt to fill in the other missing concept crossmodally, leading to an unexpected zero-shot prediction.We demonstrate CAB by showing that CLIP's zeroshot classification performance greatly suffers when there is a strong concept association between an object (e.g.eggplant) and an attribute (e.g.color purple).We also show that the strength of CAB predicts the performance on VQA.We observe that CAB is prevalent in vision-language models trained with contrastive losses, even when autoregressive losses are jointly employed.However, a model that solely relies on autoregressive loss seems to exhibit minimal or no signs of CAB. * Equal contribution.CLIP: "In this picture, the color of the lemon is purple."
Yingtian Tang, Yutaro Yamada, Yoyo Zhang, Ilker Yildirim
EMNLP4
2022 Perception of liquids relies on generalizable, physics-based representations
Wenyan Bi, Ilker Yildirim
CogSci3
2021 Automatic computation of navigational affordances explains selective processing of geometry in scene perception: behavioral and computational evidence
Mario Belledonne, Ilker Yildirim
CogSci2
2021 Perception of soft materials relies on physics-based object representations: Behavioral and computational evidence
Wenyan Bi, Aalap D. Shah, Kimberly W. Wong, Brian J. Scholl, Ilker Yildirim
CogSci5
2021 Detecting the involvement of agents through physical reasoning
Michael Lopez-Brau, Joseph Kwon, Breanna McBean, Ilker Yildirim, Julian Jara-Ettinger
CogSci4
2021 Seeing in the dark: Testing deep neural network and analysis-by-synthesis accounts of 3D shape perception with highly degraded images
Hakan Yilmaz, Gargi Singh, Bernhard Egger 0001, Josh Tenenbaum, Ilker Yildirim
CogSci5
2020 Modeling temporal attention in dynamic scenes: Hypothesis-driven resource allocation using adaptive computation explains both objective tracking performance and subjective effort judgments
Eivinas Butkus, Mario Belledonne, Brian J. Scholl, Ilker Yildirim
CogSci4
2020 Inverse Rendering Best Explains Face Perception Under Extreme Illuminations
Bernhard Egger 0001, Max H. Siegel, Riya Arora, Amir Arsalan Soltani, Ilker Yildirim, Josh Tenenbaum
CogSci5
2019 Real-time inference of physical properties in dynamic scenes
Kevin A. Smith 0001, Mario Belledonne, Ilker Yildirim, Jiajun Wu 0001, Josh Tenenbaum
CogSci3
2019 Draping an Elephant: Uncovering Children's Reasoning About Cloth-Covered Objects
Tomer D. Ullman, Eliza Kosoy, Ilker Yildirim, Amir Arsalan Soltani, Max H. Siegel, Josh Tenenbaum, Elizabeth S. Spelke
CogSci3
2019 Explaining intuitive difficulty judgments by modeling physical effort and risk
Ilker Yildirim, Basil Saeed, Grace Bennett-Pierre, Tobias Gerstenberg, Josh Tenenbaum, Hyowon Gweon
CogSci1
2019 Modeling human intuitions about liquid flow with particle-based simulation
abstract
Humans can easily describe, imagine, and, crucially, predict a wide variety of behaviors of liquids-splashing, squirting, gushing, sloshing, soaking, dripping, draining, trickling, pooling, and pouring-despite tremendous variability in their material and dynamical properties. Here we propose and test a computational model of how people perceive and predict these liquid dynamics, based on coarse approximate simulations of fluids as collections of interacting particles. Our model is analogous to a "game engine in the head", drawing on techniques for interactive simulations (as in video games) that optimize for efficiency and natural appearance rather than physical accuracy. In two behavioral experiments, we found that the model accurately captured people's predictions about how liquids flow among complex solid obstacles, and was significantly better than several alternatives based on simple heuristics and deep neural networks. Our model was also able to explain how people's predictions varied as a function of the liquids' properties (e.g., viscosity and stickiness). Together, the model and empirical results extend the recent proposal that human physical scene understanding for the dynamics of rigid, solid objects can be supported by approximate probabilistic simulation, to the more complex and unexplored domain of fluid dynamics.
Christopher Bates, Ilker Yildirim, Josh Tenenbaum, Peter W. Battaglia
PLoS Comput. Biol.2
2017 Physical problem solving: Joint planning with symbolic, geometric, and dynamic constraints
Ilker Yildirim, Tobias Gerstenberg, Basil Saeed, Marc Toussaint, Josh Tenenbaum
CogSci1
2017 Causal and compositional generative models in online perception
Ilker Yildirim, Michael Janner, Mario Belledonne, Christian Wallraven, Winrich Freiwald, Josh Tenenbaum
CogSci1
2017 Workshop proposal: Deep Learning in Computational Cognitive Science
Ilker Yildirim, Josh Tenenbaum
CogSci1
2017 Self-Supervised Intrinsic Image Decomposition
abstract
Intrinsic decomposition from a single image is a highly challenging task, due to its inherent ambiguity and the scarcity of training data. In contrast to traditional fully supervised learning approaches, in this paper we propose learning intrinsic image decomposition by explaining the input image. Our model, the Rendered Intrinsics Network (RIN), joins together an image decomposition pipeline, which predicts reflectance, shape, and lighting conditions given a single image, with a recombination function, a learned shading model used to recompose the original input based off of intrinsic image predictions. Our network can then use unsupervised reconstruction error as an additional signal to improve its intermediate representations. This allows large-scale unlabeled data to be useful during training, and also enables transferring learned knowledge to images of unseen object categories, lighting conditions, and shapes. Extensive experiments demonstrate that our method performs well on both intrinsic image decomposition and knowledge transfer.
Michael Janner, Jiajun Wu 0001, Tejas D. Kulkarni, Ilker Yildirim, Josh Tenenbaum
NIPS4
2016 Integrating identification and perception: A case study of familiar and unfamiliar face processing
Kelsey R. Allen, Ilker Yildirim, Josh Tenenbaum
CogSci2
2016 Integrating physical reasoning and visual object recognition for fully occluded scene interpretation
Ilker Yildirim, Max H. Siegel, Josh Tenenbaum
CogSci1
2015 Humans predict liquid dynamics using probabilistic simulation
Christopher Bates, Peter W. Battaglia, Ilker Yildirim, Josh Tenenbaum
CogSci3
2015 Efficient analysis-by-synthesis in vision: A computational framework, behavioral tests, and modeling neuronal representations
Ilker Yildirim, Tejas D. Kulkarni, Winrich Freiwald, Josh Tenenbaum
CogSci1
2015 Galileo: Perceiving Physical Object Properties by Integrating a Physics Engine with Deep Learning
abstract
Humans demonstrate remarkable abilities to predict physical events in dynamic scenes, and to infer the physical properties of objects from static images. We propose a generative model for solving these problems of physical scene understanding from real-world videos and images. At the core of our generative model is a 3D physics engine, operating on an object-based representation of physical properties, including mass, position, 3D shape, and friction. We can infer these latent properties using relatively brief runs of MCMC, which drive simulations in the physics engine to fit key features of visual observations. We further explore directly mapping visual inputs to physical properties, inverting a part of the generative process using deep learning. We name our model Galileo, and evaluate it on a video dataset with simple yet physically rich scenarios. Results show that Galileo is able to infer the physical properties of objects and predict the outcome of a variety of physical events, with an accuracy comparable to human subjects. Our study points towards an account of human vision with generative physical knowledge at its core, and various recognition models as helpers leading to efficient inference.
Jiajun Wu 0001, Ilker Yildirim, Joseph J. Lim, William T. Freeman, Josh Tenenbaum
NIPS2
2015 From Sensory Signals to Modality-Independent Conceptual Representations: A Probabilistic Language of Thought Approach
abstract
People learn modality-independent, conceptual representations from modality-specific sensory signals. Here, we hypothesize that any system that accomplishes this feat will include three components: a representational language for characterizing modality-independent representations, a set of sensory-specific forward models for mapping from modality-independent representations to sensory signals, and an inference algorithm for inverting forward models-that is, an algorithm for using sensory signals to infer modality-independent representations. To evaluate this hypothesis, we instantiate it in the form of a computational model that learns object shape representations from visual and/or haptic signals. The model uses a probabilistic grammar to characterize modality-independent representations of object shape, uses a computer graphics toolkit and a human hand simulator to map from object representations to visual and haptic features, respectively, and uses a Bayesian inference algorithm to infer modality-independent object representations from visual and/or haptic signals. Simulation results show that the model infers identical object representations when an object is viewed, grasped, or both. That is, the model's percepts are modality invariant. We also report the results of an experiment in which different subjects rated the similarity of pairs of objects in different sensory conditions, and show that the model provides a very accurate account of subjects' ratings. Conceptually, this research significantly contributes to our understanding of modality invariance, an important type of perceptual constancy, by demonstrating how modality-independent representations can be acquired and used. Methodologically, it provides an important contribution to cognitive modeling, particularly an emerging probabilistic language-of-thought approach, by showing how symbolic and statistical approaches can be combined in order to understand aspects of human perception.
Goker Erdogan, Ilker Yildirim, Robert A. Jacobs
PLoS Comput. Biol.2
2014 Transfer of object shape knowledge across visual and haptic modalities
Goker Erdogan, Ilker Yildirim, Robert A. Jacobs
CogSci2
2013 Linguistic Variability and Adaptation in Quantifier Meanings
Ilker Yildirim, Judith Degen, Michael K. Tanenhaus, T. Florian Jaeger
CogSci1