EDBT 2026 Demo / reviewers in the wild / expert
Chen Yu 0001
dblp:98/4839-1
· DBLP profile ↗
93ranked-venue papers
13as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 88 · 11 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 70 · 3 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Children with ASD show diminished input statistics for word learning during caregiver-child interaction
Catherine Bianco, Chen Yu 0001 |
CogSci | 2 |
| 2025 | Near-Zipfian Distribution is Prevalent in Infant Input
Brianna E. Kaplan, Chen Yu 0001 |
CogSci | 2 |
| 2025 | Using Head-Mounted Eye Tracking to Examine Infant Face Looking During Naturalistic Freeplay
Brianna E. Kaplan, Elton Martinez, Chen Yu 0001 |
CogSci | 3 |
| 2025 | Zero-Shot Cross-Situational Learning for Building Word-Referent Mappings
Melina Knabe, Chen Yu 0001 |
CogSci | 2 |
| 2025 | Parents and Children Create Semantic Regularities During Naturalistic Toy Play
Melina Knabe, Chen Yu 0001 |
CogSci | 2 |
| 2025 | The Roles of Speech Complexity and Pointing Gesture in Guiding Children's Attention During Shared Book Reading
Yayun Zhang, Jennifer Sander, Thalassia Kontino, Caroline F. Rowland, Chen Yu 0001 |
CogSci | 5 |
| 2024 | Learning semantic knowledge based on infant real-time attention and parent in-situ speech
Jane Yang, Yayun Zhang, Chen Yu 0001 |
CogSci | 3 |
| 2023 | The categorization and utility of ambiguity for cross-situational verb learning
Spencer Caplan, Chen Yu 0001 |
CogSci | 2 |
| 2023 | Using an Egocentric Human Simulation Paradigm to quantify referential and semantic ambiguity in early word learning
Spencer Caplan, Misty Z. Peng, Yayun Zhang, Chen Yu 0001 |
CogSci | 4 |
| 2023 | Embodied attention resolves visual ambiguity to support infants' real-time word learning
Sara E. Schroer, Chen Yu 0001 |
CogSci | 2 |
| 2023 | Using manual actions to create visual saliency: an outside-in solution to sustained attention and joint attention
Jane Yang, Linda B. Smith, David Crandall, Chen Yu 0001 |
CogSci | 4 |
| 2022 | Quantifying cross-situational statistics during parent-child toy play
Ellis Cain, Rachel Ryskin, Chen Yu 0001 |
CogSci | 3 |
| 2022 | Infant Action Prediction of Everyday Food Preparation
Claire D. Monroy, Sara E. Schroer, Mary M. Hayhoe, Chen Yu 0001 |
CogSci | 4 |
| 2022 | Visual attention and language exposure during everyday activities: an at-home study of early word learning using wearable eye trackers
Sara E. Schroer, Ryan E. Peters, Alyssa Yarbrough, Chen Yu 0001 |
CogSci | 4 |
| 2022 | Examining Real-time Attention Dynamics in Parent-infant Picture Book Reading
Yayun Zhang, Chen Yu 0001 |
CogSci | 2 |
| 2022 | Grounding Action Verbs in Egocentric Visual Perception
Yayun Zhang, Ellis Cain, David Crandall, Chen Yu 0001 |
CogSci | 4 |
| 2021 | In-the-Moment Visual Information from the Infant's Egocentric View Determines the Success of Infant Word Learning: A Computational Study
Andrei Amatuni, Sara E. Schroer, Yayun Zhang, Ryan E. Peters, Md. Alimoor Reza, David Crandall, Chen Yu 0001 |
CogSci | 7 |
| 2021 | Parents Adaptively Use Anaphora During Parent-child Social Interaction
Jasmine J. Falk, Yayun Zhang, Matthias Scheutz, Chen Yu 0001 |
CogSci | 4 |
| 2021 | Joint Action in Deaf and Hearing Toddlers: A Mobile Eye-Tracking Study
Claire D. Monroy, Derek Houston, Chen Yu 0001 |
CogSci | 3 |
| 2021 | Modeling joint attention from egocentric vision
Ryan E. Peters, Andrei Amatuni, Sara E. Schroer, Shujon Naha, David Crandall, Chen Yu 0001 |
CogSci | 6 |
| 2021 | The Sensorimotor Dynamics of Joint Attention
Sara E. Schroer, Chen Yu 0001 |
CogSci | 2 |
| 2021 | Multimodal Behaviors from Children Elicit Parent Responses in Real-Time Social Interaction
Julia Yurkovic, Daniel P. Kennedy, Chen Yu 0001 |
CogSci | 3 |
| 2021 | Human Learners Integrate Visual and Linguistic Information Cross-Situational Verb Learning
Yayun Zhang, Andrei Amatuni, Ellis Cain, David Crandall, Chen Yu 0001 |
CogSci | 6 |
| 2020 | Localizing Novel Attended Objects in Egocentric Views
Shujon Naha, Md. Alimoor Reza, Chen Yu 0001, David Crandall |
BMVC | 3 |
| 2020 | Decoding Eye Movements in Cross-Situational Word Learning via Tensor Component Analysis
Andrei Amatuni, Chen Yu 0001 |
CogSci | 2 |
| 2020 | Examining a developmental pathway of early word learning: From Qualitative Characteristics of Parent Speech, to Sustained Attention, to Vocabulary Size
Ryan E. Peters, Chen Yu 0001 |
CogSci | 2 |
| 2020 | Active Vision in the Perception of Actions: An Eye Tracking Study in Naturalistic Contexts
Ryan E. Peters, Dian Zhi, Matthew Petersen, Chen Yu 0001 |
CogSci | 4 |
| 2020 | A Computational Model of Early Word Learning from the Infant's Point of View
Satoshi Tsutsui, Arjun Chandrasekaran, Md. Alimoor Reza, David Crandall, Chen Yu 0001 |
CogSci | 5 |
| 2020 | Examining Sustained Attention in Child-Parent Interaction: A Comparative Study of Typically Developing Children and Children with Autism Spectrum Disorder
Julia Yurkovic, Grace Lisandrelli, Rebecca C. Schaffer, Kelli Dominick, Ernest V. Pedapati, Craig A. Erickson, Daniel P. Kennedy, Chen Yu 0001 |
CogSci | 8 |
| 2020 | Seeking Meaning: Examining a Cross-situational Solution to Learn Action Verbs Using Human Simulation Paradigm
Yayun Zhang, Andrei Amatuni, Ellis Cain, Chen Yu 0001 |
CogSci | 4 |
| 2019 | Modeling Gaze Distribution in Cross-situational Learning
Andrei Amatuni, Chen Yu 0001 |
CogSci | 2 |
| 2019 | Why Some Verbs are Harder to Learn than Others - A Micro-Level Analysis of Everyday Learning Contexts for Early Verb Learning
Siyun Liu, Yayun Zhang, Chen Yu 0001 |
CogSci | 3 |
| 2019 | Action prediction during real-time social interactions in infancy
Claire D. Monroy, Chi-hsin Chen, Derek Houston, Chen Yu 0001 |
CogSci | 4 |
| 2019 | How do infants start learning object names in a sea of clutter?
Hadar Karmazyn Raz, Drew H. Abney, David Crandall, Chen Yu 0001, Linda B. Smith |
CogSci | 4 |
| 2019 | Examining the multimodal effects of parent speech in parent-infant interactions
Sara E. Schroer, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2019 | Incremental Object Learning From Contiguous ViewsabstractIn this work, we present CRIB (Continual Recognition Inspired by Babies), a synthetic incremental object learning environment that can produce data that models visual imagery produced by object exploration in early infancy. CRIB is coupled with a new 3D object dataset, Toys-200, that contains 200 unique toy-like object instances, and is also compatible with existing 3D datasets. Through extensive empirical evaluation of state-of-the-art incremental learning algorithms, we find the novel empirical result that repetition can significantly ameliorate the effects of catastrophic forgetting. Furthermore, we find that in certain cases repetition allows for performance approaching that of batch learning algorithms. Finally, we propose an unsupervised incremental learning task with intriguing baseline results. Stefan Stojanov, Samarth Mishra, Ngoc Anh Thai, Nikhil Dhanda, Ahmad Humayun, Chen Yu 0001, Linda B. Smith, James M. Rehg |
CVPR | 6 |
| 2019 | A Self Validation Network for Object-Level Human Attention EstimationabstractDue to the foveated nature of the human vision system, people can focus their visual attention on a small region of their visual field at a time, which usually contains only a single object. Estimating this object of attention in first-person (egocentric) videos is useful for many human-centered real-world applications such as augmented reality applications and driver assistance systems. A straightforward solution for this problem is to pick the object whose bounding box is hit by the gaze, where eye gaze point estimation is obtained from a traditional eye gaze estimator and object candidates are generated from an off-the-shelf object detector. However, such an approach can fail because it addresses the where and the what problems separately, despite that they are highly related, chicken-and-egg problems. In this paper, we propose a novel unified model that incorporates both spatial and temporal evidence in identifying as well as locating the attended object in firstperson videos. It introduces a novel Self Validation Module that enforces and leverages consistency of the where and the what concepts. We evaluate on two public datasets, demonstrating that Self Validation Module significantly benefits both training and testing and that our model outperforms the state-of-the-art. Chen Yu 0001, David Crandall |
NeurIPS | 2 |
| 2018 | From Coarse Attention to Fine-Grained Gaze: A Two-stage 3D Fully Convolutional Network for Predicting Eye Gaze in First Person Video
David Crandall, Chen Yu 0001, Sven Bambach |
BMVC | 3 |
| 2018 | Hand-Eye Coordination and Visual Attention in Infancy
Drew H. Abney, Hadar Karmazyn Raz, Linda B. Smith, Chen Yu 0001 |
CogSci | 4 |
| 2018 | The Roles of Gesture and Statistical Cues on Infants' Word Learning in Shared Storybook Reading
Yayun Zhang, Chen Yu 0001 |
CogSci | 2 |
| 2018 | Estimating Head Motion from Egocentric VisionabstractThe recent availability of lightweight, wearable cameras allows for collecting video data from a "first-person' perspective, capturing the visual world of the wearer in everyday interactive contexts. In this paper, we investigate how to exploit egocentric vision to infer multimodal behaviors from people wearing head-mounted cameras. More specifically, we estimate head (camera) motion from egocentric video, which can be further used to infer non-verbal behaviors such as head turns and nodding in multimodal interactions. We propose several approaches based on Convolutional Neural Networks (CNNs) that combine raw images and optical flow fields to learn to distinguish regions with optical flow caused by global ego-motion from those caused by other motion in a scene. Our results suggest that CNNs do not directly learn useful visual features with end-to-end training from raw images alone; instead, a better approach is to first extract optical flow explicitly and then train CNNs to integrate optical flow and visual information. Satoshi Tsutsui, Sven Bambach, David Crandall, Chen Yu 0001 |
ICMI | 4 |
| 2018 | Toddler-Inspired Visual Object LearningabstractReal-world learning systems have practical limitations on the quality and quantity of the training datasets that they can collect and consider. How should a system go about choosing a subset of the possible training examples that still allows for learning accurate, generalizable models? To help address this question, we draw inspiration from a highly efficient practical learning system: the human child. Using head-mounted cameras, eye gaze trackers, and a model of foveated vision, we collected first-person (egocentric) images that represents a highly accurate approximation of the "training data" that toddlers' visual systems collect in everyday, naturalistic learning contexts. We used state-of-the-art computer vision learning models (convolutional neural networks) to help characterize the structure of these data, and found that child data produce significantly better object models than egocentric data experienced by adults in exactly the same environment. By using the CNNs as a modeling tool to investigate the properties of the child data that may enable this rapid learning, we found that child data exhibit a unique combination of quality and diversity, with not only many similar large, high-quality object views but also a greater number and diversity of rare views. This novel methodology of analyzing the visual "training data" used by children may not only reveal insights to improve machine learning, but also may suggest new experimental tools to better understand infant learning in developmental psychology. Sven Bambach, David Crandall, Linda B. Smith, Chen Yu 0001 |
NeurIPS | 4 |
| 2017 | It's Time: Quantifying the Relevant Timescales for Joint Attention
Drew H. Abney, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2017 | Learning Object Names from Visual Pervasiveness: the Visual Statistics Predict
Elizabeth M. Clerkin, Chen Yu 0001, Linda B. Smith |
CogSci | 2 |
| 2017 | Information Signatures in Children's Language Environment
Steven L. Elmlinger, Drew H. Abney, David W. Vinson, Linda B. Smith, Chen Yu 0001 |
CogSci | 5 |
| 2017 | Slow Change: The Visual Context for Real World Learning
Charlene Tay, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2017 | Discovering Multicausality in the Development of Coordinated Behavior
Tian Xu 0001, Drew H. Abney, Chen Yu 0001 |
CogSci | 3 |
| 2017 | Seeing Is Not Enough for Sustained Visual Attention
Tian Xu 0001, Chen Yu 0001, Linda B. Smith |
CogSci | 3 |
| 2017 | Iterative Machine TeachingabstractIn this paper, we consider the problem of machine teaching, the inverse problem of machine learning. Different from traditional machine teaching which views the learners as batch algorithms, we study a new paradigm where the learner uses an iterative algorithm and a teacher can feed examples sequentially and intelligently based on the current performance of the learner. We show that the teaching complexity in the iterative case is very different from that in the batch case. Instead of constructing a minimal training set for learners, our iterative machine teaching focuses on achieving fast convergence in the learner model. Depending on the level of information the teacher has from the learner model, we design teaching algorithms which can provably reduce the number of teaching examples and achieve faster convergence than learning without teachers. We also validate our theoretical findings with extensive experiments on different data distribution and real image datasets. Weiyang Liu, Bo Dai 0001, Ahmad Humayun, Charlene Tay, Chen Yu 0001, Linda B. Smith, James M. Rehg |
ICML | 5 |
| 2016 | Active Viewing in Toddlers Facilitates Visual Object Learning: An Egocentric Vision Approach
Sven Bambach, David Crandall, Linda B. Smith, Chen Yu 0001 |
CogSci | 4 |
| 2016 | Which Statistic Matters? Effects of Category Size and Distribution on Statistical Category Learning
Chi-hsin Chen, Paulo Carvalho 0004, Chen Yu 0001 |
CogSci | 3 |
| 2016 | Infants' Developing Coordinated Visual-Manual Object Exploration and Links with Vocabulary Development
Lauren Slone, Chen Yu 0001, Linda B. Smith |
CogSci | 2 |
| 2016 | More than Words: The Many Ways Extended Discourse Facilitates Word Learning
Sumarga H. Suanda, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2016 | Finding Clarity Amidst the Clutter: How Parents Name Objects
Charlene Tay, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2016 | Quantifying Joint Activities using Cross-Recurrence Block Representation
Tian Xu 0001, Chen Yu 0001 |
CogSci | 2 |
| 2016 | Learning Hierarchical Labels through Cross-situational Learning
Chen Yu 0001, Chi-hsin Chen |
CogSci | 1 |
| 2016 | Examining Referential Uncertainty in Naturalistic Contexts from the Child's View: Evidence from an Eye-Tracking Study with Infants
Yayun Zhang, Chen Yu 0001 |
CogSci | 2 |
| 2016 | See You See Me: The Role of Eye Contact in Multimodal Human-Robot InteractionabstractWe focus on a fundamental looking behavior in human-robot interactions - gazing at each other's face. Eye contact and mutual gaze between two social partners are critical in smooth human-human interactions. Therefore, investigating at what moments and in what ways a robot should look at a human user's face as a response to the human's gaze behavior is an important topic. Toward this goal, we developed a gaze-contingent human-robot interaction system, which relied on momentary gaze behaviors from a human user to control an interacting robot in real time. Using this system, we conducted an experiment in which human participants interacted with the robot in a joint attention task. In the experiment, we systematically manipulated the robot's gaze toward the human partner's face in real time and then analyzed the human's gaze behavior as a response to the robot's gaze behavior. We found that more face looks from the robot led to more look-backs (to the robot's face) from human participants and consequently created more mutual gaze and eye contact between the two. Moreover, participants demonstrated more coordinated and synchronized multimodal behaviors between speech and gaze when more eye contact was successfully established and maintained. Tian Xu 0001, Hui Zhang 0006, Chen Yu 0001 |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Visual-motor coordination in natural reaching of young children and adults
John M. Franchak, Chen Yu 0001 |
CogSci | 2 |
| 2015 | Linking Joint Attention with Hand-Eye Coordination - A Sensorimotor Approach to Understanding Child-Parent Social Interaction
Chen Yu 0001, Linda B. Smith |
CogSci | 1 |
| 2015 | Statistical Word Learning is a Continuous Process: Evidence from the Human Simulation Paradigm
Yayun Zhang, Daniel Yurovsky, Chen Yu 0001 |
CogSci | 3 |
| 2015 | Lending A Hand: Detecting Hands and Recognizing Activities in Complex Egocentric InteractionsabstractHands appear very often in egocentric video, and their appearance and pose give important cues about what people are doing and what they are paying attention to. But existing work in hand detection has made strong assumptions that work well in only simple scenarios, such as with limited interaction with other people or in lab settings. We develop methods to locate and distinguish between hands in egocentric video using strong appearance models with Convolutional Neural Networks, and introduce a simple candidate region generation approach that outperforms existing techniques at a fraction of the computational cost. We show how these high-quality bounding boxes can be used to create accurate pixelwise hand regions, and as an application, we investigate the extent to which hand segmentation alone can distinguish between different activities. We evaluate these techniques on a new dataset of 48 first-person videos of people interacting in realistic environments, with pixel-level ground truth for over 15,000 hand instances. Sven Bambach, Stefan Lee, David Crandall, Chen Yu 0001 |
ICCV | 4 |
| 2015 | Viewpoint Integration for Hand-Based Recognition of Social Interactions from a First-Person ViewabstractWearable devices are becoming part of everyday life, from first-person cameras (GoPro, Google Glass), to smart watches (Apple Watch), to activity trackers (FitBit). These devices are often equipped with advanced sensors that gather data about the wearer and the environment. These sensors enable new ways of recognizing and analyzing the wearer's everyday personal activities, which could be used for intelligent human-computer interfaces and other applications. We explore one possible application by investigating how egocentric video data collected from head-mounted cameras can be used to recognize social activities between two interacting partners (e.g. playing chess or cards). In particular, we demonstrate that just the positions and poses of hands within the first-person view are highly informative for activity recognition, and present a computer vision approach that detects hands to automatically estimate activities. While hand pose detection is imperfect, we show that combining evidence across first-person views from the two social partners significantly improves activity recognition accuracy. This result highlights how integrating weak but complimentary sources of evidence from social partners engaged in the same task can help to recognize the nature of their interaction. Sven Bambach, David Crandall, Chen Yu 0001 |
ICMI | 3 |
| 2014 | Detecting Hands in Children's Egocentric Views to Understand Embodied Attention during Social Interaction
Sven Bambach, John M. Franchak, David Crandall, Chen Yu 0001 |
CogSci | 4 |
| 2014 | Developing Semantic Knowledge through Cross-situational Word Learning
George Kachergis, Chen Yu 0001, Richard M. Shiffrin |
CogSci | 2 |
| 2014 | Interactions between statistical aggregation and hypothesis testing mechanisms during word learning
Alexa R. Romberg, Chen Yu 0001 |
CogSci | 2 |
| 2014 | It's in the Hands: Developmental Changes in the Quality of Naming Events for Two-Year-Olds
Sumarga H. Suanda, Danika Geisler, Linda B. Smith, Chen Yu 0001 |
CogSci | 4 |
| 2014 | Learning to interact and interacting to learn: Active statistical learning in human-robot interactionabstractLearning and interaction are viewed as two related but distinct topics in developmental robotics. Many studies focus solely on either building a robot that can acquire new knowledge and learn to perform new tasks, or designing smooth human-robot interactions with pre-acquired knowledge and skills. The present paper focuses on linking language learning with human-robot interaction, showing how better human-robot interaction can lead to better language learning by robot. Toward this goal, we developed a real-time human-robot interaction paradigm in which a robot learner acquired lexical knowledge from a human teacher through free-flowing interaction. With the same statistical learning mechanism in the robot's system, we systematically manipulated the degree of activity in human-robot interaction in three experimental conditions: the robot learner was either highly active with lots of speaking and looking acts, or moderately active with a few acts, or passive without actions. Our results show that more talking and looking acts from the robot, including those immature behaviors such as saying non-sense words or looking at random targets, motivated human teachers to be more engaged in the interaction. In addition, more activities from the robot revealed its robot's internal learning states in real time, which allowed human teachers to provide more useful and "on-demand" teaching signals to facilitate learning. Thus, compared with passive and batch-mode training, an active robot learner can create more and better training data through smooth and effective social interactions that consequentially lead to more successful language learning. Chen Yu 0001, Tian Xu 0001, Yiwen Zhong, Seth B. Foster, Hui Zhang 0006 |
IJCNN | 1 |
| 2013 | Embodied Approaches to Interpersonal Coordination: Infants, Adults, Robots, and Agents
Rick Dale, Chen Yu 0001, Yukie Nagai, Moreno I. Coco, Stefan Kopp |
CogSci | 2 |
| 2013 | More Naturalistic Cross-situational Word Learning
George Kachergis, Chen Yu 0001 |
CogSci | 2 |
| 2013 | Informavores: Active information foraging and human cognition
Douglas Markant, Todd M. Gureckis, Björn Meder, Jonathan D. Nelson, Peter Pirolli, Chen Yu 0001 |
CogSci | 6 |
| 2013 | Integration and Inference: Cross-Situational Word Learning Involves More than Simple Co-occurrences
Alexa R. Romberg, Chen Yu 0001 |
CogSci | 2 |
| 2013 | Mechanistic Developmental Process: Rumelhart Prize Symposium in Honor of Linda Smith
Larissa K. Samuelson, Anthony F. Morse, Chen Yu 0001, Eliana Colunga, Thomas T. Hills |
CogSci | 3 |
| 2013 | An Attentionally Constrained Model of Statistical Word Learning
Sumarga H. Suanda, Seth B. Foster, Linda B. Smith, Chen Yu 0001 |
CogSci | 4 |
| 2012 | Actively Learning Nouns Across Ambiguous Situations
George Kachergis, Chen Yu 0001, Richard M. Shiffrin |
CogSci | 2 |
| 2012 | Does Statistical Word Learning Scale? It's a Matter of Perspective
Daniel Yurovsky, Linda B. Smith, Chen Yu 0001 |
CogSci | 3 |
| 2012 | Adaptive eye gaze patterns in interactions with human and artificial agentsabstractEfficient collaborations between interacting agents, be they humans, virtual or embodied agents, require mutual recognition of the goal, appropriate sequencing and coordination of each agent's behavior with others, and making predictions from and about the likely behavior of others. Moment-by-moment eye gaze plays an important role in such interaction and collaboration. In light of this, we used a novel experimental paradigm to systematically investigate gaze patterns in both human-human and human-agent interactions. Participants in the study were asked to interact with either another human or an embodied agent in a joint attention task. Fine-grained multimodal behavioral data were recorded including eye movement data, speech, first-person view video, which were then analyzed to discover various behavioral patterns. Those patterns show that human participants are highly sensitive to momentary multimodal behaviors generated by the social partner (either another human or an artificial agent) and they rapidly adapt their gaze behaviors accordingly. Our results from this data-driven approach provide new findings for understanding micro-behaviors in human-human communication which will be critical for the design of artificial agents that can generate human-like gaze behaviors and engage in multimodal interactions with humans. Chen Yu 0001, Paul W. Schermerhorn, Matthias Scheutz |
ACM Trans. Interact. Intell. Syst. | 1 |
| 2011 | Visual Attention and Change Detection
Ty W. Boyer, Thomas G. Smith, Chen Yu 0001, Bennett I. Bertenthal |
CogSci | 3 |
| 2011 | From Data Streams to Information Flow: Information Exchange in Child-Parent Interaction
Heeyoul Choi, Chen Yu 0001, Linda B. Smith, Olaf Sporns |
CogSci | 2 |
| 2011 | Active Cross-situational Learning
George Kachergis, Chen Yu 0001, Richard M. Shiffrin |
CogSci | 2 |
| 2011 | Spatial and Temporal Cues in Statistical Cross Situdational Learning
Roy Seo, Damian Fricker, Chen Yu 0001 |
CogSci | 3 |
| 2011 | Word Learning through Sensorimotor Child-Parent Interaction: A Feature Selection Approach
Chen Yu 0001, Jun-Ming Xu 0002, Xiaojin Zhu 0001 |
CogSci | 1 |
| 2010 | Investigating multimodal real-time patterns of joint attention in an hri word learning task
Chen Yu 0001, Matthias Scheutz, Paul W. Schermerhorn |
HRI | 1 |
| 2010 | A Data-Driven Paradigm to Understand Multimodal Communication in Human-Human and Human-Robot Interaction
Chen Yu 0001, Thomas G. Smith, Shohei Hidaka, Matthias Scheutz, Linda B. Smith |
IDA | 1 |
| 2010 | A Multimodal Real-Time Platform for Studying Human-Avatar Interactions
Hui Zhang 0006, Damian Fricker, Chen Yu 0001 |
IVA | 3 |
| 2008 | Social coordination in toddler's word learning: interacting systems of perception and actionabstractWe measured turn-taking in terms of hand and head movements and asked if the global rhythm of the participants' body activity relates to word learning. Six dyads composed of parents and toddlers (M = 18 months) interacted in a tabletop task wearing motion-tracking sensors on their hands and head. Parents were instructed to teach the labels of 10 novel objects and the child was later tested on a name-comprehension task. Using dynamic time warping, we compared the motion data of all body-part pairs, within and between partners. For every dyad, we also computed an overall measure of the quality of the interaction, that takes into consideration the state of interaction when the parent uttered an object label and the overall smoothness of the turn-taking. The overall interaction quality measure was correlated with the total number of words learned.In particular, head movements were inversely related to other partner's hand movements, and the degree of bodily coupling of parent and toddler predicted the words that children learned during the interaction. The implications of joint body dynamics to understanding joint coordination of activity in a social interaction, its scaffolding effect on the child's learning and its use in the development of artificial systems are discussed. Alfredo F. Pereira, Linda B. Smith, Chen Yu 0001 |
Connect. Sci. | 3 |
| 2007 | A unified model of early word learning: Integrating statistical and social cues
Chen Yu 0001, Dana H. Ballard |
Neurocomputing | 1 |
| 2005 | The emergence of links between lexical acquisition and object categorization: a computational studyabstractLanguage is about symbols, and those symbols must be grounded in the physical world. Children learn to associate language with sensorimotor experiences during their development. In light of this, we first provide a computational account of how words are mapped to their perceptually grounded meanings. Moreover, the main part of this work proposes and implements a computational model of how word learning influences the formation of object categories to which those words refer. This model simulates the bi-directional relationship between word and object category learning: (1) object categorization provides mental representations of meanings that are mapped to words to form lexical items; (2) linguistic labels help object categorization by providing additional teaching signals; and (3) these two learning processes interplay with each other and form a developmental feedback loop. Compared with the method that performs these two tasks separately, our model shows promising improvements in both word-to-world mapping and perceptual categorization, suggesting a unified view of lexical and category learning in an integrative framework. Most importantly, this work provides a cognitively plausible explanation of the mechanistic nature of early word learning and object learning from co-occurring multisensory data. Chen Yu 0001 |
Connect. Sci. | 1 |
| 2004 | On the Integration of Grounding Language and Learning Objects
Chen Yu 0001, Dana H. Ballard |
AAAI | 1 |
| 2004 | A multimodal learning interface for grounding spoken language in sensory perceptionsabstractWe present a multimodal interface that learns words from natural interactions with users. In light of studies of human language development, the learning system is trained in an unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behaviors. The system collects acoustic signals in concert with user-centric multisensory information from nonspeech modalities, such as user's perspective video, gaze positions, head directions, and hand movements. A multimodal learning algorithm uses this data to first spot words from continuous speech and then associate action verbs and object names with their perceptually grounded meanings. The central ideas are to make use of nonspeech contextual information to facilitate word spotting, and utilize body movements as deictic references to associate temporally cooccurring data from different modalities and build hypothesized lexical items. From those items, an EM-based method is developed to select correct word--meaning pairs. Successful learning is demonstrated in the experiments of three natural tasks: "unscrewing a jar," "stapling a letter," and "pouring water." Chen Yu 0001, Dana H. Ballard |
ACM Trans. Appl. Percept. | 1 |
| 2003 | A multimodal learning interface for word acquisitionabstractWe present a multimodal interface that learns words from natural interactions with users. The system can be trained in an unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behavior. We collect acoustic signals in concert with user-centric multisensory information from non-speech modalities, such as user's perspective video, gaze positions, head directions and hand movements. A multimodal learning algorithm is developed that firstly spots words from continuous speech and then associates action verbs and object names with their grounded meanings. The central idea is to make use of non-speech contextual information to facilitate word spotting, and utilize temporal correlations of data from different modalities to build hypothesized lexical items. From those items, an EM-based method selects correct word-meaning pairs. Successful learning has been demonstrated in the experiment of the natural task of "stapling papers". Dana H. Ballard, Chen Yu 0001 |
ICASSP (5) | 2 |
| 2003 | A multimodal learning interface for grounding spoken language in sensory perceptionsabstractMost speech interfaces are based on natural language processing techniques that use pre-defined symbolic representations of word meanings and process only linguistic information. To understand and use language like their human counterparts in multimodal human-computer interaction, computers need to acquire spoken language and map it to other sensory perceptions. This paper presents a multimodal interface that learns to associate spoken language with perceptual features by being situated in users' everyday environments and sharing user-centric multisensory information. The learning interface is trained in unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behaviors. We collect acoustic signals in concert with multisensory information from non-speech modalities, such as user's perspective video, gaze positions, head directions and hand movements. The system firstly estimates users' focus of attention from eye and head cues. Attention, as represented by gaze fixation, is used for spotting the target object of user interest. Attention switches are calculated and used to segment an action sequence into action units which are then categorized by mixture hidden Markov models. A multimodal learning algorithm is developed to spot words from continuous speech and then associate them with perceptually grounded meanings extracted from visual perception and action. Successful learning has been demonstrated in the experiments of three natural tasks: "unscrewing a jar", "stapling a letter" and "pouring water". Chen Yu 0001, Dana H. Ballard |
ICMI | 1 |
| 2002 | Attentional Object Spotting by Integrating Multimodal InputabstractAn intelligent human-computer interface is expected to allow computers to work with users in a cooperative manner. To achieve this goal, computers need to be aware of user attention and provide assistance without explicit user requests. Cognitive studies of eye movements suggest that in accomplishing well-learned tasks, the performer's focus of attention is locked onto ongoing work and more than 90% of eye movements are closely related to the objects being manipulated in the tasks. In light of this, we have developed an attentional object spotting system that integrates multimodal data consisting of eye position, head position and video from the "first-person" perspective. To detect the user's focus of attention, we modeled eye gaze and head movements using a hidden Markov model (HMM) representation. For each attentional point in time, the object of user interest is automatically extracted and recognized. We report the results of experiments on finding attentional objects in the natural task of "making a peanut-butter sandwich". Chen Yu 0001, Dana H. Ballard, Shenghuo Zhu |
ICMI | 1 |