Tao Gao 0004

dblp:08/17-4 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
20since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 2 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 17 since 2021Systems, architecture and hardware · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Communicating through Acting: The Role of Contextual Affordance in Intuitive Pantomimetic Gestural Communication
Siyi Gong, Jessica G. Li, Mireille Karadanaian, Ziyi Meng 0004, Tao Gao 0004
CogSci6
2025 Territorial Gestalt in the Strategy of Conflicts
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci6
2025 Beyond Emotion: Unraveling the Limited Role of Sentiment in Extended-Format Communication
Shuhao Fu, Rick Dale, Tao Gao 0004, Junying Liang
CogSci4
2024 Intentional commitment as a spontaneous presentation of self
Shaozhe Cheng, Jingyin Zhu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci5
2023 Simplifying Group Communication: A Shared Agency Modeling Approach
Max Potter, Tao Gao 0004, Stephanie Stacy
CogSci2
2022 Intentional commitment through an internalized theory of mind: Acting in the eyes of an imagined observer
Shaozhe Cheng, Minglu Zhao, Jingyin Zhu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci6
2022 What Is the point? a Theory of Mind Model of Relevance
Stephanie Stacy, Annya L. Dahmani, Boxuan Jiang, Federico Rossano, Yixin Zhu 0001, Tao Gao 0004
CogSci7
2022 Perceptual Grouping for War and Peace
Siyi Gong, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci7
2022 Overloaded Communication as Paternalistic Helping
Stephanie Stacy, Aishni Parab, Max Kleiman-Weiner, Tao Gao 0004
CogSci4
2022 No Such Thing as the Average Listener: Belief-driven versus Action-driven Strategies in Signaling
Stephanie Stacy, Yiling Yun, Max Potter, Naomi Moskowitz, Tao Gao 0004
CogSci5
2022 Exploring an Imagined "We" in Human Collective Hunting: Joint Commitment within Shared Intentionality
Siyi Gong, Minglu Zhao, Chenya Gu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci7
2022 The development of commitment: Attention for intention
Shuyi Zhai, Shaozhe Cheng, Naomi Moskowitz, Mowei Shen, Tao Gao 0004
CogSci5
2022 Emergent Graphical Conventions in a Visual Communication Game
abstract
Humans communicate with graphical sketches apart from symbolic languages. Primarily focusing on the latter, recent studies of emergent communication overlook the sketches; they do not account for the evolution process through which symbolic sign systems emerge in the trade-off between iconicity and symbolicity. In this work, we take the very first step to model and simulate this process via two neural agents playing a visual communication game; the sender communicates with the receiver by sketching on a canvas. We devise a novel reinforcement learning method such that agents are evolved jointly towards successful communication and abstract graphical conventions. To inspect the emerged conventions, we define three key properties -- iconicity, symbolicity, and semanticity -- and design evaluation methods accordingly. Our experimental results under different controls are consistent with the observation in studies of human graphical conventions. Of note, we find that evolved sketches can preserve the continuum of semantics under proper environmental pressures. More interestingly, co-evolved agents can switch between conventionalized and iconic communication based on their familiarity with referents. We hope the present research can pave the path for studying emergent communication with the modality of sketches.
Shuwen Qiu, Sirui Xie, Lifeng Fan, Tao Gao 0004, Jungseock Joo, Song-Chun Zhu, Yixin Zhu 0001
NeurIPS4
2021 Intention beyond Desire: Humans Spontaneously Commit to Future Actions
Shaozhe Cheng, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci7
2021 Individual vs. Joint Perception: a Pragmatic Model of Pointing as Smithian Helping
Stephanie Stacy, Adelpha Chan, Chuyu Wei, Federico Rossano, Yixin Zhu 0001, Tao Gao 0004
CogSci7
2021 Modeling Communication to Coordinate Perspectives in Cooperation
Stephanie Stacy, Chenfei Li, Minglu Zhao, Yiling Yun, Qingyi Zhao, Max Kleiman-Weiner, Tao Gao 0004
CogSci7
2021 Jointly Perceiving Physics and Mind: Motion, force and intention
Siyi Gong, Ziqian Liao, Haokui Xu, Jifan Zhou, Mowei Shen, Tao Gao 0004
CogSci7
2021 Sharing is not Needed: Modeling Animal Coordinated Hunting with Reinforcement Learning
Minglu Zhao, Annya L. Dahmani, Ross Richard Perry, Yixin Zhu 0001, Federico Rossano, Tao Gao 0004
CogSci7
2021 Learning Triadic Belief Dynamics in Nonverbal Communication From Videos
abstract
Humans possess a unique social cognition capability [43], [20]; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene understanding literature. In this paper, we incorporate different nonverbal communication cues (e.g., gaze, human poses, and gestures) to represent, model, learn, and infer agents’ mental states from pure visual inputs. Crucially, such a mental representation takes the agent’s belief into account so that it represents what the true world state is and infers the beliefs in each agent’s mental state, which may differ from the true world states. By aggregating different beliefs and true world states, our model essentially forms "five minds" during the interactions between two agents. This "five minds" model differs from prior works that infer beliefs in an infinite recursion; instead, agents’ beliefs are converged into a "common mind" [31], [47]. Based on this representation, we further devise a hierarchical energy-based model that jointly tracks and predicts all five minds. From this new perspective, a social event is interpreted by a series of nonverbal communication and belief dynamics, which transcends the classic keyframe video summary. In the experiments, we demonstrate that using such a social account provides a better video summary on videos with rich social interactions compared with state-of-the-art keyframe video summary methods.
Lifeng Fan, Shuwen Qiu, Zilong Zheng, Tao Gao 0004, Song-Chun Zhu, Yixin Zhu 0001
CVPR4
2021 YouRefIt: Embodied Reference Understanding with Language and Gesture
abstract
We study the machine’s understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with perspective-taking to identify which object is being referred to. To tackle this problem, we introduce YouRefIt, a new crowd-sourced dataset of embodied reference collected in various physical scenes; the dataset contains 4,195 unique reference clips in 432 indoor scenes. To the best of our knowledge, this is the first embodied reference dataset that allows us to study referring expressions in daily physical scenes to understand referential behavior, human communication, and human-robot interaction. We further devise two benchmarks for image-based and video-based embodied reference understanding. Comprehensive baselines and extensive experiments provide the very first result of machine perception on how the referring expressions and gestures affect the embodied reference understanding. Our results provide essential evidence that gestural cues are as critical as language cues in understanding the embodied reference.
Yixin Chen 0003, Qing Li 0003, Deqian Kong, Yik Lun Kei, Song-Chun Zhu, Tao Gao 0004, Yixin Zhu 0001, Siyuan Huang 0001
ICCV6
2020 Intuitive Signaling Through an "Imagined We'"
Stephanie Stacy, Qingyi Zhao, Minglu Zhao, Max Kleiman-Weiner, Tao Gao 0004
CogSci5
2020 Bootstrapping an Imagined We for Cooperation
Stephanie Stacy, Minglu Zhao, Gabriel Marquez, Tao Gao 0004
CogSci5
2020 Joint Inference of States, Robot Knowledge, and Human (False-)Beliefs
abstract
Aiming to understand how human (false-)belief- a core socio-cognitive ability-would affect human interactions with robots, this paper proposes to adopt a graphical model to unify the representation of object states, robot knowledge, and human (false-)beliefs. Specifically, a parse graph (pg) is learned from a single-view spatiotemporal parsing by aggregating various object states along the time; such a learned representation is accumulated as the robot's knowledge. An inference algorithm is derived to fuse individual pg from all robots across multi-views into a joint pg, which affords more effective reasoning and inference capability to overcome the errors originated from a single view. In the experiments, through the joint inference over pgs, the system correctly recognizes human (false-)belief in various settings and achieves better cross-view accuracy on a challenging small object tracking dataset.
Hangxin Liu, Lifeng Fan, Zilong Zheng, Tao Gao 0004, Yixin Zhu 0001, Song-Chun Zhu
ICRA5
2017 Perception Meets Examination: Studying Deceptive Behaviors in VR
Carla Aravena, Mark Vo, Tao Gao 0004, Takaaki Shiratori, Lap-Fai Yu
CogSci3
2016 Proposal of the Second Workshop on Physical and Social Scene Understanding
Tao Gao 0004, Chenfanfu Jiang, Yixin Zhu 0001, Yibiao Zhao, Lap-Fai Yu
CogSci1
2016 Inferring human intent from video by sampling hierarchical plans
abstract
This paper presents a method which allows robots to infer a human's hierarchical intent from partially observed RGBD videos by imagining how the human will behave in the future. This capability is critical for creating robots which can interact socially or collaboratively with humans. We represent intent as a novel hierarchical, compositional, and probabilistic And-Or graph structure which describes a relationship between actions and plans. We infer human intent by reverse-engineering a human's decision-making and action planning processes under a Bayesian probabilistic programming framework. We present experiments from a 3D environment which demonstrate that the inferred human intent (1) matches well with human judgment, and (2) provides useful contextual cues for object tracking and action recognition.
Steven Holtzen, Yibiao Zhao, Tao Gao 0004, Josh Tenenbaum, Song-Chun Zhu
IROS3
2015 Physical and Social Scene Understanding
Tao Gao 0004, Yibiao Zhao, Lap-Fai Yu
CogSci1
2012 Modeling the Perception of Intentions
Barbara Tversky, Shimon Ullman, Dare A. Baldwin, Frank E. Pollick, Josh Tenenbaum, Tao Gao 0004, Peter C. Pantelis, David Pautler
CogSci6