VLDB 2026 Research / reviewers in the wild / expert
Joyce Y. Chai
dblp:c/JoyceYChai · also Joyce Chai, Joyce Yue Chai
· DBLP profile ↗
106ranked-venue papers
18as first author
35since 2021 · last 2026
0000-0002-9658-2230ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 12 first-author · 33 since 2021Human-computer interaction and ubiquitous computing · 16 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 1 since 2021Systems, architecture and hardware · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language ModelsabstractRuixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Mike Angstadt, Chandra Sripada, Joyce Chai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks, Aman Taxali, Michael Angstadt, Chandra Sekhar Sripada, Joyce Y. Chai |
ACL (1) | 8 |
| 2025 | LinkGPT: Leveraging Large Language Models for Enhanced Link Prediction in Text-Attributed GraphsabstractInspired by the success of Large Language Models (LLMs) in language and vision tasks, there has been growing interest in applying LLMs to graph tasks, particularly on Text-Attributed Graphs (TAGs). However, most prior work tackles the node classification task. In this work, we evaluate an LLM's ability to reason over structured data and infer new facts based on learned patterns by focusing on link prediction (LP)-the task of predicting missing links between nodes-that is understudied in the literature. This task poses two key challenges: (1) How to effectively integrate pairwise structural information, which is crucial for LP performance, into LLMs, and (2) how to address the computational bottleneck during inference. To tackle these challenges, we propose LinkGPT, the first LLM-based training and inference framework specifically designed for LP on homogeneous TAGs. To enhance the LLM's ability to understand the underlying structure, we carefully design a node encoder and pairwise encoder, and leverage a two-stage instruction tuning to effectively incorporate the nodewise and pairwise information into LLMs. For inference efficiency, we introduce a retrieval-reranking scheme. Extensive experiments show that LinkGPT achieves state-of-the-art performance on real-world graphs and demonstrates superior zero-shot and few-shot generalization. At inference time, it achieves a 10× speedup while maintaining high LP accuracy. Zhongmou He, Jing Zhu 0005, Shengyi Qian 0001, Joyce Y. Chai, Danai Koutra |
CIKM | 4 |
| 2025 | Flexible Physical Problem Solving with Strategy Acquisition and Composition
Jung-Chun Liu, Jiayuan Mao, Joyce Y. Chai, Josh Tenenbaum |
CogSci | 3 |
| 2025 | 3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less HallucinationabstractThe integration of language and 3D perception is crucial for embodied agents and robots that comprehend and interact with the physical world. While large language models (LLMs) have demonstrated impressive language understanding and generation capabilities, their adaptation to 3D environments (3D-LLMs) remains in its early stages. A primary challenge is a lack of large-scale datasets with dense grounding between language and 3D scenes. We introduce 3D-GRAND, a pioneering large-scale dataset comprising 40,087 household scenes paired with 6.2 million densely-grounded scene-language instructions. Our results show that instruction tuning with 3D-GRAND significantly enhances grounding capabilities and reduces hallucinations in 3D-LLMs. As part of our contributions, we propose a comprehensive benchmark 3D-POPE to systematically evaluate hallucination in 3D-LLMs, enabling fair comparisons of models. Our experiments highlight a scaling effect between dataset size and 3D-LLM performance, emphasizing the importance of large-scale 3D-text datasets for embodied AI research. Our results demonstrate early signals for effective sim-to-real transfer, indicating that models trained on large synthetic data can perform well on real-world 3D scans. Through 3D-GRAND and 3D-POPE, we aim to equip the embodied AI community with resources and insights to lead to more reliable and better-grounded 3D-LLMs. Xuweiyi Chen, Nikhil Madaan, Madhavan Iyengar, Shengyi Qian 0001, David F. Fouhey, Joyce Y. Chai |
CVPR | 7 |
| 2025 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassabstractMulti-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives. Current leading methods such as DUSt3R employ a fundamentally pairwise approach, processing images in pairs and necessitating costly global alignment procedures to reconstruct from multiple views. In this work, we propose Fast 3D Reconstruction (Fast3R), a novel multi-view generalization to DUSt3R that achieves efficient and scalable 3D reconstruction by processing many views in parallel. Fast3R’s Transformer-based architecture forwards N images in a single forward pass, bypassing the need for iterative alignment. Through extensive experiments on camera pose estimation and 3D reconstruction, Fast3R demonstrates state-of-the-art performance, with significant improvements in inference speed and reduced error accumulation. These results establish Fast3R as a robust alternative for multi-view applications, offering enhanced scalability without compromising reconstruction accuracy. Alexander Sax, Kevin J. Liang, Mikael Henaff, Ang Cao, Joyce Y. Chai, Franziska Meier, Matt Feiszli |
CVPR | 7 |
| 2025 | Transparent and Coherent Procedural Mistake DetectionabstractProcedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a procedural text).Despite significant recent efforts, machine performance in the wild remains nonviable, and the reasoning processes underlying this performance are opaque.As such, we extend PMD to require generating visual self-dialog rationales to inform decisions.Given the impressive, mature image understanding capabilities observed in recent visionand-language models (VLMs), we curate a suitable benchmark dataset for PMD based on individual frames.As our reformulation enables unprecedented transparency, we leverage a natural language inference (NLI) model to formulate two automated metrics for the coherence of generated rationales.We establish baselines for this reframed task, showing that VLMs struggle off-the-shelf, but with some trade-offs, their accuracy, coherence, and efficiency can be improved by incorporating these metrics into common inference and finetuning methods.Lastly, our multi-faceted metrics visualize common outcomes, highlighting areas for further improvement. Shane Storks, Itamar Bar-Yossef, Yayuan Li, Jason J. Corso, Joyce Y. Chai |
EMNLP | 6 |
| 2025 | Proactive Assistant Dialogue Generation from Streaming Egocentric VideosabstractYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, Seungwhan Moon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichi Zhang 0001, Xin Dong 0001, Zhaojiang Lin, Andrea Madotto, Babak Damavandi, Joyce Y. Chai, Seungwhan Moon |
EMNLP | 7 |
| 2025 | VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
Shoubin Yu, Difan Liu, Ziqiao Ma 0001, Yicong Hong, Yang Zhou 0009, Hao Tan 0002, Joyce Y. Chai, Mohit Bansal |
ICCV | 7 |
| 2025 | Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference under AmbiguitiesabstractSpatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning. Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Y. Chai, Ziqiao Ma 0001 |
ICLR | 6 |
| 2025 | RACER: Rich Language-Guided Failure Recovery Policies for Imitation LearningabstractDeveloping robust and correctable visuomotor policies for robotic manipulation is challenging due to the lack of self-recovery mechanisms from failures and the limitations of simple language instructions in guiding robot actions. To address these issues, we propose a scalable data generation pipeline that automatically augments expert demonstrations with failure recovery trajectories and fine-grained language annotations for training. We then introduce Rich languAge-guided failure reCovERy (RACER), a supervisor-actor frame-work, which combines failure recovery data with rich language descriptions to enhance robot control. RACER features a vision-language model (VLM) that acts as an online supervisor, providing detailed language guidance for error correction and task execution, and a language-conditioned visuomotor policy as an actor to predict the next actions. Our experimental results show that RACER outperforms the state-of-the-art Robotic View Transformer (RVT) on RLbench across various evaluation settings, including standard long-horizon tasks, dynamic goal-change tasks and zero-shot unseen tasks, achieving superior performance in both simulated and real world environments. Videos and code are available at: https://rich-language-failure-recovery.github.io. Yinpei Dai, Jayjun Lee, Nima Fazeli, Joyce Y. Chai |
ICRA | 4 |
| 2025 | Babysit A Language Model From Scratch: Interactive Language Learning by Trials and DemonstrationsabstractZiqiao Ma, Zekun Wang, Joyce Chai. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ziqiao Ma 0001, Zekun Wang 0002, Joyce Y. Chai |
NAACL (Long Papers) | 3 |
| 2025 | 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any TimeabstractCan we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction model that takes input from unconstrained views and timestamps and renders arbitrary novel view-time combinations. Unlike prior 4D approaches, e.g., optimization-based, geometry-based, or generative, that struggle with efficiency, generalization, or faithfulness, 4D-LRM learns a unified space-time representation and directly predicts per-pixel 4D Gaussian primitives from posed image tokens across time, enabling fast, high-quality rendering at, in principle, infinite frame rate. Our results demonstrate that scaling spatiotemporal pretraining enables accurate and efficient 4D reconstruction. We show that 4D-LRM generalizes to novel objects, interpolates across time, and handles diverse camera setups. It reconstructs 24-frame sequences in one forward pass with less than 1.5 seconds on a single A100 GPU. Ziqiao Ma 0001, Xuweiyi Chen, Shoubin Yu, Sai Bi, Kai Zhang 0045, Sihan Xu, Zexiang Xu, Kalyan Sunkavalli, Mohit Bansal, Joyce Y. Chai, Hao Tan 0002 |
NeurIPS | 12 |
| 2025 | Position: Towards Bidirectional Human-AI AlignmentabstractRecent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values. Yet, the lack of a clear, shared understanding of what constitutes "alignment" limits meaningful progress and cross-disciplinary collaboration. In this position paper, we argue that the research community should explicitly define and critically reflect on "alignment" to account for the bidirectional and dynamic relationship between humans and AI. Through a systematic review of over 400 papers spanning HCI, NLP, ML, and more, we examine how alignment is currently defined and operationalized. Building on this analysis, we introduce the Bidirectional Human-AI Alignment framework, which not only incorporates traditional efforts to align AI with human values but also introduces the critical, underexplored dimension of aligning humans with AI – supporting cognitive, behavioral, and societal adaptation to rapidly advancing AI technologies. Our findings reveal significant gaps in current literature, especially in long-term interaction design, human value modeling, and mutual understanding. We conclude with three central challenges and actionable recommendations to guide future research toward more nuanced, reciprocal, and human-AI alignment approaches. Hua Shen 0005, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Savvas Petridis, Yi-Hao Peng, Li Qiwei, Chenglei Si, Yutong Xie 0007, Jeffrey P. Bigham, Frank Bentley, Joyce Y. Chai, Zachary C. Lipton, Qiaozhu Mei, Michael Terry, Diyi Yang, Meredith Ringel Morris, Paul Resnick, David Jurgens |
NeurIPS | 14 |
| 2024 | Inversion-Free Image Editing with Language-Guided Diffusion ModelsabstractDespite recent advances in inversion-based editing, text-guided image manipulation remains challenging for diffusion models. The primary bottlenecks include 1) the time-consuming nature of the inversion process; 2) the struggle to balance consistency with accuracy; 3) the lack of compatibility with efficient consistency sampling methods used in consistency models. To address the above issues, we start by asking ourselves if the inversion process can be eliminated for editing. We show that when the initial sample is known, a special variance schedule reduces the denoising step to the same form as the multi-step consistency sam- pling. We name this Denoising Diffusion Consistent Model (DDCM), and note that it implies a virtual inversion strat-egy without explicit inversion in sampling. We further unify the attention control mechanisms in a tuning-free framework for text-guided editing. Combining them, we present inversion-free editing (InfEdit), which allows for consistent and faithful editing for both rigid and non-rigid semantic changes, catering to intricate modifications without compromising on the image's integrity and explicit inversion. Through extensive experiments, InfEdit shows strong performance in various editing tasks and also maintains a seamless workflow (less than 3 seconds on one single A40), demonstrating the potential for real-time applications. Sihan Xu, Yidong Huang, Jiayi Pan 0002, Ziqiao Ma 0001, Joyce Y. Chai |
CVPR | 5 |
| 2024 | Groundhog Grounding Large Language Models to Holistic SegmentationabstractMost multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level representations that are important for fine-grained visual understanding and diagnosis. In this work, we introduce Groundhog, an MLLM developed by grounding Large Language Models to holistic segmentation. GROUNDHOG incorporates a masked feature extractor and converts extracted features into visual entity tokens for the MLLM backbone, which then connects groundable phrases to unified grounding masks by retrieving and merging the entity masks. To train GROUNDHOG, we carefully curated M3G2, a grounded visual instruction tuning dataset with Multi-Modal Multi-Grained Grounding, by harvesting a collection of segmentation-grounded datasets with rich annotations. Our experimental results show that GROUNDHOG achieves superior performance on various language grounding tasks without task-specific fine-tuning, and significantly reduces object hallucination. GROUNDHOG also demonstrates better grounding towards complex forms of visual input and provides easy-to-understand diagnosis in failure cases. Yichi Zhang 0001, Ziqiao Ma 0001, Xiaofeng Gao 0002, Suhaila M. Shakiah, Qiaozi Gao, Joyce Y. Chai |
CVPR | 6 |
| 2024 | Teaching Embodied Reinforcement Learning Agents: Informativeness and Diversity of Language UseabstractIn real-world scenarios, it is desirable for embodied agents to have the ability to leverage human language to gain explicit or implicit knowledge for learning tasks.Despite recent progress, most previous approaches adopt simple low-level instructions as language inputs, which may not reflect natural human communication.It's not clear how to incorporate rich language use to facilitate task learning.To address this question, this paper studies different types of language inputs in facilitating reinforcement learning (RL) embodied agents.More specifically, we examine how different levels of language informativeness (i.e., feedback on past behaviors and future guidance) and diversity (i.e., variation of language expressions) impact agent learning and inference.Our empirical results based on four RL benchmarks demonstrate that agents trained with diverse and informative language feedback can achieve enhanced generalization and fast adaptation to new tasks.These findings highlight the pivotal role of language use in teaching embodied agents new tasks in an open world. 1 Jiajun Xi, Yinong He, Yinpei Dai, Joyce Y. Chai |
EMNLP | 5 |
| 2024 | Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional PropertiesabstractA major reason behind the recent success of large language models (LLMs) is their incontext learning capability, which makes it possible to rapidly adapt them to downstream textbased tasks by prompting them with a small number of relevant demonstrations.While large vision-language models (VLMs) have recently been developed for tasks requiring both text and images, they largely lack in-context learning over visual information, especially in understanding and generating text about videos.In this work, we implement Emergent In-context Learning on Videos (EILeV), a novel training paradigm that induces in-context learning over video and text by capturing key properties of pre-training data found by prior work to be essential for in-context learning in transformers.In our experiments, we show that EILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.Furthermore, we demonstrate that these key properties of bursty distributions, skewed marginal distributions, and dynamic meaning each contribute to varying degrees to VLMs' in-context learning capability in narrating procedural videos.Our results, analysis, and EILeV-trained models yield numerous insights about the emergence of in-context learning over video and text, creating a foundation for future work to optimize and scale VLMs for open-domain video understanding and reasoning.1 Keunwoo Peter Yu, Fengyuan Hu, Shane Storks, Joyce Y. Chai |
EMNLP | 5 |
| 2024 | Think, Act, and Ask: Open-World Interactive Personalized Robot NavigationabstractZero-Shot Object Navigation (ZSON) enables agents to navigate towards open-vocabulary objects in unknown environments. The existing works of ZSON mainly focus on following individual instructions to find generic object classes, neglecting the utilization of natural language interaction and the complexities of identifying user-specific objects. To address these limitations, we introduce Zero-shot Interactive Personalized Object Navigation (ZIPON), where robots need to navigate to personalized goal objects while engaging in conversations with users. To solve ZIPON, we propose a new framework termed Open-woRld Interactive persOnalized Navigation (ORION)1, which uses Large Language Models (LLMs) to make sequential decisions to manipulate different modules for perception, navigation and communication. Experimental results show that the performance of interactive agents that can leverage user feedback exhibits significant improvement. However, obtaining a good balance between task completion and the efficiency of navigation and interaction remains challenging for all methods. We further provide more findings on the impact of diverse user feedback forms on the agents’ performance. Yinpei Dai, Run Peng, Sikai Li, Joyce Y. Chai |
ICRA | 4 |
| 2024 | LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agentabstract3D visual grounding is a critical skill for household robots, enabling them to navigate, manipulate objects, and answer questions based on their environment. While existing approaches often rely on extensive labeled data or exhibit limitations in handling complex language queries, we propose LLM-Grounder, a novel zero-shot, open-vocabulary, Large Language Model (LLM)-based 3D visual grounding pipeline. LLM-Grounder utilizes an LLM to decompose complex natural language queries into semantic constituents and employs a visual grounding tool, such as OpenScene or LERF, to identify objects in a 3D scene. The LLM then evaluates the spatial and commonsense relations among the proposed objects to make a final grounding decision. Our method does not require any labeled training data and can generalize to novel 3D scenes and arbitrary text queries. We evaluate LLM-Grounder on the ScanRefer benchmark and demonstrate state-of-the-art zero-shot grounding accuracy. Our findings indicate that LLMs significantly improve the grounding capability, especially for complex language queries, making LLM-Grounder an effective approach for 3D vision-language tasks in robotics. Xuweiyi Chen, Shengyi Qian 0001, Nikhil Madaan, Madhavan Iyengar, David F. Fouhey, Joyce Y. Chai |
ICRA | 7 |
| 2024 | DriVLMe: Enhancing LLM-based Autonomous Driving Agents with Embodied and Social ExperiencesabstractRecent advancements in foundation models (FMs) have unlocked new prospects in autonomous driving, yet the experimental settings of these studies are preliminary, oversimplified, and fail to capture the complexity of real-world driving scenarios in human environments. It remains under-explored whether FM agents can handle long-horizon navigation tasks with free-from dialogue and deal with unexpected situations caused by environmental dynamics or task changes. To explore the capabilities and boundaries of FMs faced with the challenges above, we introduce DriVLMe, a video-language-model-based agent to facilitate natural and effective communication between humans and autonomous vehicles that perceive the environment and navigate. We develop DriVLMe from both embodied experiences in a simulated environment and social experiences from real human dialogue. While DriVLMe demonstrates competitive performance in both open-loop benchmarks and closed-loop human studies, we reveal several limitations and challenges, including unacceptable inference time, imbalanced training data, limited visual understanding, challenges with multi-turn interactions, simplified language generation from robotic experiences, and difficulties in handling on-the-fly unexpected situations like environmental dynamics and task changes. Nevertheless, DriVLMe offers a promising new direction for autonomous driving agents that need to navigate not just complex environments but also complex social interactions. Yidong Huang, Jacob Sansom, Ziqiao Ma 0001, Felix Gervits, Joyce Y. Chai |
IROS | 5 |
| 2024 | Multi-Object Hallucination in Vision Language ModelsabstractLarge vision language models (LVLMs) often suffer from object hallucination, producing objects not present in the given images.
While current benchmarks for object hallucination primarily concentrate on the presence of a single object class rather than individual entities, this work systematically investigates multi-object hallucination, examining how models misperceive (e.g., invent nonexistent objects or become distracted) when tasked with focusing on multiple objects simultaneously.
We introduce Recognition-based Object Probing Evaluation (ROPE), an automated evaluation protocol that considers the distribution of object classes within a single image during testing and uses visual referring prompts to eliminate ambiguity.
With comprehensive empirical studies and analysis of potential factors leading to multi-object hallucination, we found that (1) LVLMs suffer more hallucinations when focusing on multiple objects compared to a single object.
(2) The tested object class distribution affects hallucination behaviors, indicating that LVLMs may follow shortcuts and spurious correlations.
(3) Hallucinatory behaviors are influenced by data-specific factors, salience and frequency, and model intrinsic behaviors.
We hope to enable LVLMs to recognize and reason about multiple objects that often occur in realistic visual scenes, provide insights, and quantify our progress towards mitigating the issues. Xuweiyi Chen, Ziqiao Ma 0001, Xuejun Zhang 0003, Sihan Xu, Shengyi Qian 0001, David F. Fouhey, Joyce Y. Chai |
NeurIPS | 8 |
| 2024 | GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot LearningabstractPre-trained vision-language models (VLMs) have achieved promising success in many fields, especially with prompt learning paradigm. In this work, we propose GIPCOL (Graph-Injected Soft Prompting for Compositional Learning) to better explore the compositional zero-shot learning (CZSL) ability of VLMs within the prompt-based learning framework. The soft prompt in GIPCOL is structured and consists of the prefix learnable vectors, attribute label and object label. In addition, the attribute and object labels in the soft prompt are designated as nodes in a compositional graph. The compositional graph is constructed based on the compositional structure of the objects and attributes extracted from the training data and consequently feeds the updated concept representation into the soft prompt to capture this compositional structure for a better prompting for CZSL. With the new prompting strategy, GIPCOL achieves state-of-the-art AUC results on all three CZSL benchmarks, including MIT-States, UT-Zappos, and C-GQA datasets in both closed and open settings compared to previous non-CLIP as well as CLIP-based methods. We analyze when and why GIPCOL operates well given the CLIP backbone and its training data limitations, and our findings shed light on designing more effective prompts for CZSL. Guangyue Xu, Joyce Y. Chai, Parisa Kordjamshidi |
WACV | 2 |
| 2023 | Human Inspired Progressive Alignment and Comparative Learning for Grounded Word AcquisitionabstractHuman language acquisition is an efficient, supervised, and continual process.In this work, we took inspiration from how human babies acquire their first language, and developed a computational process for word acquisition through comparative learning.Motivated by cognitive findings, we generated a small dataset that enables the computation models to compare the similarities and differences of various attributes, learn to filter out and extract the common information for each shared linguistic label.We frame the acquisition of words as not only the information filtration process, but also as representation-symbol mapping.This procedure does not involve a fixed vocabulary size, nor a discriminative objective, and allows the models to continually learn more concepts efficiently.Our results in controlled experiments have shown the potential of this approach for efficient continual learning of grounded words. Yuwei Bao, Barrett Martin Lattimer, Joyce Y. Chai |
ACL (1) | 3 |
| 2023 | In-Context Analogical Reasoning with Pre-Trained Language ModelsabstractAnalogical reasoning is a fundamental capacity of human cognition that allows us to reason abstractly about novel situations by relating them to past experiences.While it is thought to be essential for robust reasoning in AI systems, conventional approaches require significant training and/or hard-coding of domain knowledge to be applied to benchmark tasks.Inspired by cognitive science research that has found connections between human language and analogy-making, we explore the use of intuitive language-based abstractions to support analogy in AI systems.Specifically, we apply large pre-trained language models (PLMs) to visual Raven's Progressive Matrices (RPM), a common relational reasoning test.By simply encoding the perceptual features of the problem into language form, we find that PLMs exhibit a striking capacity for zero-shot relational reasoning, exceeding human performance and nearing supervised vision-based methods.We explore different encodings that vary the level of abstraction over task features, finding that higherlevel abstractions further strengthen PLMs' analogical reasoning.Our detailed analysis reveals insights on the role of model complexity, incontext learning, and prior knowledge in solving RPM tasks. Xiaoyang Hu, Shane Storks, Richard L. Lewis, Joyce Y. Chai |
ACL (1) | 4 |
| 2023 | World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language ModelsabstractThe ability to connect language units to their referents in the physical world, referred to as grounding, is crucial to learning and understanding grounded meanings of words.While humans demonstrate fast mapping in new word learning, it remains unclear whether modern vision-language models can truly represent language with their grounded meanings, and how grounding may further bootstrap new word learning.To this end, we introduce Grounded Open Vocabulary Acquisition (GOVA) to examine grounding and bootstrapping in openworld language learning.As an initial attempt, we propose World-to-Words (W2W), a novel visually-grounded language model by pre-training on image-text pairs highlighting grounding as an objective.Through extensive experiments and analysis, we demonstrate that W2W is a more coherent and fast grounded word learner, and that the grounding ability acquired during pre-training helps the model to learn unseen words more rapidly and robustly.1 Ziqiao Ma 0001, Jiayi Pan 0002, Joyce Y. Chai |
ACL (1) | 3 |
| 2023 | NLP Reproducibility For All: Understanding Experiences of BeginnersabstractAs natural language processing (NLP) has recently seen an unprecedented level of excitement, and more people are eager to enter the field, it is unclear whether current research reproducibility efforts are sufficient for this group of beginners to apply the latest developments.To understand their needs, we conducted a study with 93 students in an introductory NLP course, where students reproduced the results of recent NLP papers.Surprisingly, we find that their programming skill and comprehension of research papers have a limited impact on their effort spent completing the exercise.Instead, we find accessibility efforts by research authors to be the key to success, including complete documentation, better coding practice, and easier access to data files.Going forward, we recommend that NLP researchers pay close attention to these simple aspects of open-sourcing their work, and use insights from beginners' feedback to provide actionable ideas on how to better support them. Shane Storks, Keunwoo Peter Yu, Ziqiao Ma 0001, Joyce Y. Chai |
ACL (1) | 4 |
| 2023 | Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?abstractVision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world.However, known as visual illusions, human's perception of reality isn't always faithful to the physical world.This raises a key question: do VLMs have the similar kind of illusions as humans do, or do they faithfully learn to represent reality?To investigate this question, we build a dataset containing five types of visual illusions and formulate four tasks to examine visual illusions in state-of-the-art VLMs.Our findings have shown that although the overall alignment is low, larger models are closer to human perception and more susceptible to visual illusions.Our dataset and initial findings will promote a better understanding of visual illusions in humans and machines and provide a stepping stone for future computational models that can better align humans and machines in perceiving and communicating about the shared visual world.The code and data are available at github.com/vl-illusion/dataset. Yichi Zhang 0001, Jiayi Pan 0002, Joyce Y. Chai |
EMNLP | 5 |
| 2023 | From Heuristic to Analytic: Cognitively Motivated Strategies for Coherent Physical Commonsense ReasoningabstractPre-trained language models (PLMs) have shown impressive performance in various language tasks.However, they are prone to spurious correlations, and often generate illusory information.In real-world applications, PLMs should justify decisions with formalized, coherent reasoning chains, but this challenge remains under-explored.Cognitive psychology theorizes that humans are capable of utilizing fast and intuitive heuristic thinking to make decisions based on past experience, then rationalizing the decisions through slower and deliberative analytic reasoning.We incorporate these interlinked dual processes in fine-tuning and in-context learning with PLMs, applying them to two language understanding tasks that require coherent physical commonsense reasoning.We show that our proposed Heuristic-Analytic Reasoning (HAR) strategies drastically improve the coherence of rationalizations for model decisions, yielding state-of-the-art results on Tiered Reasoning for Intuitive Physics (TRIP).We also find that this improved coherence is a direct result of more faithful attention to relevant language context in each step of reasoning.Our findings suggest that human-like reasoning strategies can effectively improve the coherence and reliability of PLM reasoning. Shane Storks, Fengyuan Hu, Sungryull Sohn, Moontae Lee, Honglak Lee, Joyce Y. Chai |
EMNLP | 7 |
| 2023 | Towards Collaborative Plan Acquisition through Theory of Mind Modeling in Situated DialogueabstractCollaborative tasks often begin with partial task knowledge and incomplete plans from each partner. To complete these tasks, partners need to engage in situated communication with their partners and coordinate their partial plans towards a complete plan to achieve a joint task goal. While such collaboration seems effortless in a human-human team, it is highly challenging for human-AI collaboration. To address this limitation, this paper takes a step towards Collaborative Plan Acquisition, where humans and agents strive to learn and communicate with each other to acquire a complete plan for joint tasks. Specifically, we formulate a novel problem for agents to predict the missing task knowledge for themselves and for their partners based on rich perceptual and dialogue history. We extend a situated dialogue benchmark for symmetric collaborative tasks in a 3D blocks world and investigate computational strategies for plan acquisition. Our empirical results suggest that predicting the partner's missing knowledge is a more viable approach than predicting one's own. We show that explicit modeling of the partner's dialogue moves and mental states produces improved and more stable results than without. These results provide insight for future AI agents that can predict what knowledge their partner is missing and, therefore, can proactively communicate such information to help the partner acquire such missing knowledge toward a common understanding of joint tasks. Cristian-Paul Bara, Ziqiao Ma 0001, Yingzhuo Yu, Julie A. Shah, Joyce Y. Chai |
IJCAI | 5 |
| 2023 | Pragmatic Communication with Embodied AgentsabstractWith the emergence of a new generation of embodied AI agents (e.g., cognitive robots), it has become increasingly important to empower these agents with the ability to learn and collaborate with humans through language communication. Despite recent advances, language communication in embodied AI still faces many challenges. Human language not only needs to ground to agents’ perception and action but also needs to facilitate collaboration between humans and agents. To address these challenges, I will introduce several efforts in my lab that study pragmatic communication with embodied agents. I will talk about how language use is shaped by shared experience and knowledge (i.e., common ground) and how collaborative effort is important to mediate perceptual differences and handle exceptions. I will discuss task learning by following language instructions and highlight the need for neuro-symbolic representations for situation awareness and transparency. I will further present explicit modeling of partners’ goals, beliefs, and abilities (i.e., theory of mind) and discuss its role in language communication for situated collaborative tasks. Joyce Y. Chai |
IUI | 1 |
| 2023 | CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image ManipulationabstractDiffusion models (DMs) have enabled breakthroughs in image synthesis tasks but lack an intuitive interface for consistent image-to-image (I2I) translation. Various methods have been explored to address this issue, including mask-based methods, attention-based methods, and image-conditioning. However, it remains a critical challenge to enable unpaired I2I translation with pre-trained DMs while maintaining satisfying consistency. This paper introduces Cyclenet, a novel but simple method that incorporates cycle consistency into DMs to regularize image manipulation. We validate Cyclenet on unpaired I2I tasks of different granularities. Besides the scene and object level translation, we additionally contribute a multi-domain I2I translation dataset to study the physical state changes of objects. Our empirical studies show that Cyclenet is superior in translation consistency and quality, and can generate high-quality images for out-of-domain distributions with a simple change of the textual prompt. Cyclenet is a practical framework, which is robust even with very limited training data (around 2k) and requires minimal computational resources (1 GPU) to train. Project homepage: https://cyclenetweb.github.io/ Sihan Xu, Ziqiao Ma 0001, Yidong Huang, Honglak Lee, Joyce Y. Chai |
NeurIPS | 5 |
| 2022 | Learning to Mediate Disparities Towards Pragmatic CommunicationabstractHuman communication is a collaborative process.Speakers, on top of conveying their own intent, adjust the content and language expressions by taking the listeners into account, including their knowledge background, personalities, and physical capabilities.Towards building AI agents with similar abilities in language communication, we propose Pragmatic Rational Speaker (PRS), a framework extending Rational Speech Act (RSA).The PRS attempts to learn the speaker-listener disparity and adjust the speech accordingly, by adding a light-weighted disparity adjustment layer into working memory on top of speaker's long-term memory system.By fixing the long-term memory, the PRS only needs to update its working memory to learn and adapt to different types of listeners.To validate our framework, we create a dataset that simulates different types of speaker-listener disparities in the context of referential games.Our empirical results demonstrate that the PRS is able to shift its output towards the language that listeners are able to understand, significantly improve the collaborative task outcome. Yuwei Bao, Joyce Y. Chai |
ACL (1) | 3 |
| 2022 | DANLI: Deliberative Agent for Following Natural Language InstructionsabstractYichi Zhang, Jianing Yang, Jiayi Pan, Shane Storks, Nikhil Devraj, Ziqiao Ma, Keunwoo Yu, Yuwei Bao, Joyce Chai. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Yichi Zhang 0001, Jiayi Pan 0002, Shane Storks, Nikhil Devraj, Ziqiao Ma 0001, Keunwoo Peter Yu, Yuwei Bao, Joyce Y. Chai |
EMNLP | 9 |
| 2022 | Spoken language interaction with robots: Recommendations for future researchabstractWith robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with. Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005 |
Comput. Speech Lang. | 8 |
| 2021 | MindCraft: Theory of Mind Modeling for Situated Dialogue in Collaborative TasksabstractAn ideal integration of autonomous agents in a human world implies that they are able to collaborate on human terms.In particular, theory of mind plays an important role in maintaining common ground during human collaboration and communication.To enable theory of mind modeling in situated interactions, we introduce a fine-grained dataset of collaborative tasks performed by pairs of human subjects in the 3D virtual blocks world of Minecraft.It provides information that captures partners' beliefs of the world and of each other as an interaction unfolds, bringing abundant opportunities to study human collaborative behaviors in situated language communication.As a first step towards our goal of developing embodied AI agents able to infer belief states of collaborative partners in situ, we build and present results on computational models for several theory of mind tasks. Cristian-Paul Bara, Sky CH-Wang, Joyce Y. Chai |
EMNLP (1) | 3 |
| 2020 | Experience Grounds LanguageabstractYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph Turian. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Y. Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph P. Turian |
EMNLP (1) | 6 |
| 2018 | What Action Causes This? Towards Naive Physical Action-Effect PredictionabstractDespite recent advances in knowledge representation, automated reasoning, and machine learning, artificial agents still lack the ability to understand basic actioneffect relations regarding the physical world, for example, the action of cutting a cucumber most likely leads to the state where the cucumber is broken apart into smaller pieces.If artificial agents (e.g., robots) ever become our partners in joint tasks, it is critical to empower them with such action-effect understanding so that they can reason about the state of the world and plan for actions.Towards this goal, this paper introduces a new task on naive physical action-effect prediction, which addresses the relations between concrete actions (expressed in the form of verbnoun pairs) and their effects on the state of the physical world as depicted by images.We collected a dataset for this task and developed an approach that harnesses web image data through distant supervision to facilitate learning for action-effect prediction.Our empirical results have shown that web data can be used to complement a small number of seed examples (e.g., three examples for each action) for model learning.This opens up possibilities for agents to learn physical action-effect relations for tasks at hand through communication with humans with a few examples. Qiaozi Gao, Joyce Y. Chai, Lucy Vanderwende |
ACL (1) | 3 |
| 2018 | Commonsense Justification for Action ExplanationabstractTo enable collaboration and communication between humans and agents, this paper investigates learning to acquire commonsense evidence for action justification.In particular, we have developed an approach based on the generative Conditional Variational Autoencoder (CVAE) that models object relations/attributes of the world as latent variables and jointly learns a performer that predicts actions and an explainer that gathers commonsense evidence to justify the action.Our empirical results have shown that, compared to a typical attention-based model, CVAE achieves significantly higher performance in both action prediction and justification.A human subject study further shows that the commonsense evidence gathered by CVAE can be communicated to humans to achieve a significantly higher common ground between humans and agents. Qiaozi Gao, Sari Saba-Sadiya, Joyce Y. Chai |
EMNLP | 4 |
| 2018 | Language to Action: Towards Interactive Task Learning with Physical AgentsabstractLanguage communication plays an important role in human learning and knowledge acquisition. With the emergence of a new generation of cognitive robots, empowering these robots to learn directly from human partners becomes increasingly important. This paper gives a brief introduction to interactive task learning where humans can teach physical agents new tasks through natural language communication and action demonstration. It discusses research challenges and opportunities in language and communication grounding that are critical in this process. It further highlights the importance of commonsense knowledge, particularly the very basic physical causality knowledge, in grounding language to perception and action. Joyce Y. Chai, Qiaozi Gao, Lanbo She, Sari Saba-Sadiya, Guangyue Xu |
IJCAI | 1 |
| 2017 | Interactive Learning of Grounded Verb Semantics towards Human-Robot CommunicationabstractTo enable human-robot communication and collaboration, previous works represent grounded verb semantics as the potential change of state to the physical world caused by these verbs.Grounded verb semantics are acquired mainly based on the parallel data of the use of a verb phrase and its corresponding sequences of primitive actions demonstrated by humans.The rich interaction between teachers and students that is considered important in learning new skills has not yet been explored.To address this limitation, this paper presents a new interactive learning approach that allows robots to proactively engage in interaction with human partners by asking good questions to learn models for grounded verb semantics.The proposed approach uses reinforcement learning to allow the robot to acquire an optimal policy for its question-asking behaviors by maximizing the long-term reward.Our empirical results have shown that the interactive learning approach leads to more reliable models for grounded verb semantics, especially in the noisy environment which is full of uncertainties.Compared to previous work, the models acquired from interactive learning result in a 48% to 145% performance gain when applied in new situations. Lanbo She, Joyce Y. Chai |
ACL (1) | 2 |
| 2017 | Detecting clinically related content in online patient posts
Courtland VanDam, Shaheen Kanthawala, Wanda Pratt, Joyce Y. Chai, Jina Huh |
J. Biomed. Informatics | 4 |
| 2016 | What's Hot in Human Language Technology: Highlights from NAACL HLT 2015abstractThis paper shows a few examples to highlight the trends observed at the NAACL HLT 2015 conference. Joyce Y. Chai, Anoop Sarkar, Rada Mihalcea |
AAAI | 1 |
| 2016 | Physical Causality of Action Verbs in Grounded Language UnderstandingabstractLinguistics studies have shown that action verbs often denote some Change of State (CoS) as the result of an action.However, the causality of action verbs and its potential connection with the physical world has not been systematically explored.To address this limitation, this paper presents a study on physical causality of action verbs and their implied changes in the physical world.We first conducted a crowdsourcing experiment and identified eighteen categories of physical causality for action verbs.For a subset of these categories, we then defined a set of detectors that detect the corresponding change from visual perception of the physical environment.We further incorporated physical causality modeling and state detection in grounded language understanding.Our empirical studies have demonstrated the effectiveness of causality modeling in grounding language to perception. Qiaozi Gao, Malcolm Doering, Joyce Y. Chai |
ACL (1) | 4 |
| 2016 | Incremental Acquisition of Verb Hypothesis Space towards Physical World Interaction
Lanbo She, Joyce Y. Chai |
ACL (1) | 2 |
| 2016 | Jointly Learning Grounded Task Structures from Language Instruction and Visual DemonstrationabstractTo enable language-based communication and collaboration with cognitive robots, this paper presents an approach where an agent can learn task models jointly from language instruction and visual demonstration using an And-Or Graph (AoG) representation.The learned AoG captures a hierarchical task structure where linguistic labels (for language communication) are grounded to corresponding state changes from the physical environment (for perception and action).Our empirical results on a cloth-folding domain have shown that, although state detection through visual processing is full of uncertainties and error prone, by a tight integration with language the agent is able to learn an effective AoG for task representation.The learned AoG can be further applied to infer and interpret on-going actions from new visual demonstration using linguistic labels at different levels of granularity. Changsong Liu, Sari Saba-Sadiya, Nishant Shukla, Yunzhong He, Song-Chun Zhu, Joyce Y. Chai |
EMNLP | 7 |
| 2016 | Grounded Semantic Role LabelingabstractShaohua Yang, Qiaozi Gao, Changsong Liu, Caiming Xiong, Song-Chun Zhu, Joyce Y. Chai. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Qiaozi Gao, Changsong Liu, Caiming Xiong, Song-Chun Zhu, Joyce Y. Chai |
HLT-NAACL | 6 |
| 2015 | Learning to Mediate Perceptual Differences in Situated Human-Robot DialogueabstractIn human-robot dialogue, although a robot and its human partner are co-present in a shared environment, they have significantly mismatched perceptual capabilities (e.g., recognizing objects in the surroundings). When a shared perceptual basis is missing, it becomes difficult for the robot to identify referents in the physical world that are referred to by the human (i.e., a problem of referential grounding). To overcome this problem, we have developed an optimization based approach that allows the robot to detect and adapt to perceptual differences. Through online interaction with the human, the robot can learn a set of weights indicating how reliably/unreliably each dimension (e.g., object type, object color, etc.) of its perception of the environment maps to the human's linguistic descriptors and thus adjust its word models accordingly. Our empirical evaluation has shown that this weight-learning approach can successfully adjust the weights to reflect the robot's perceptual limitations. The learned weights, together with updated word models, can lead to a significant improvement for referential grounding in future dialogues. Changsong Liu, Joyce Y. Chai |
AAAI | 2 |
| 2015 | Question Types in Online Health Communities
Shaheen Kanthawala, Courtland VanDam, Barbara Given, Joyce Y. Chai, Amber Vermeesch, Marianne Huebner, John Crowley, Jina Huh |
AMIA | 4 |
| 2015 | Embodied Collaborative Referring Expression Generation in Situated Human-Robot InteractionabstractTo facilitate referential communication between humans and robots and mediate their differences in representing the shared environment, we are exploring embodied collaborative models for referring expression generation (REG). Instead of a single minimum description to describe a target object, episodes of expressions are generated based on human feedback during human-robot interaction. We particularly investigate the role of embodiment such as robot gesture behaviors (i.e., pointing to an object) and human's gaze feedback (i.e., looking at a particular object) in the collaborative process. This paper examines different strategies of incorporating embodiment and collaboration in REG and discusses their possibilities and challenges in enabling human-robot referential communication. Malcolm Doering, Joyce Y. Chai |
HRI | 3 |
| 2014 | Collaborative Models for Referring Expression Generation in Situated DialogueabstractIn situated dialogue with artificial agents (e.g., robots), although a human and an agent are co-present, the agent's representation and the human's representation of the shared environment are significantly mismatched. Because of this misalignment, our previous work has shown that when the agent applies traditional approaches to generate referring expressions for describing target objects with minimum descriptions, the intended objects often cannot be correctly identified by the human. To address this problem, motivated by collaborative behaviors in human referential communication, we have developed two collaborative models - an episodic model and an installment model - for referring expression generation. Both models, instead of generating a single referring expression to describe a target object as in the previous work, generate multiple small expressions that lead to the target object with the goal of minimizing the collaborative effort. In particular, our installment model incorporates human feedback in a reinforcement learning framework to learn the optimal generation strategies. Our empirical results have shown that the episodic model and the installment model outperform previous non-collaborative models with an absolute gain of 6% and 21% respectively. Malcolm Doering, Joyce Y. Chai |
AAAI | 3 |
| 2014 | Collaborative effort towards common ground in situated human-robot dialogueabstractIn situated human-robot dialogue, although humans and robots are co-present in a shared environment, they have significantly mismatched capabilities in perceiving the shared environment. Their representations of the shared world are misaligned. In order for humans and robots to communicate with each other successfully using language, it is important for them to mediate such differences and to establish common ground. To address this issue, this paper describes a dialogue system that aims to mediate a shared perceptual basis during human-robot dialogue. In particular, we present an empirical study that examines the role of the robot's collaborative effort and the performance of natural language processing modules in dialogue grounding. Our empirical results indicate that in situated human-robot dialogue, a low collaborative effort from the robot may lead its human partner to believe a common ground is established. However, such beliefs may not reflect true mutual understanding. To support truly grounded dialogues, the robot should make an extra effort by making its partner aware of its internal representation of the shared world. Joyce Y. Chai, Lanbo She, Spencer Ottarson, Cody Littley, Changsong Liu, Kenneth Hanson |
HRI | 1 |
| 2014 | Perceptive feedback for natural language control of robotic operationsabstractA new planning and control scheme for natural language control of robotic operations using the perceptive feedback is presented. Different from the traditional open-loop natural language control, the scheme incorporates the high-level planning and low-level control of the robotic systems and makes the high-level planning become a closed-loop process such that it is able to handle some unexpected events in the robotics system and the environment. The experimental results on a natural language controlled mobile manipulator clearly demonstrate the advantages of the proposed method. Yunyi Jia, Ning Xi 0001, Joyce Y. Chai, Yu Cheng 0006, Lanbo She |
ICRA | 3 |
| 2014 | Teaching Robots New Actions through Natural Language InstructionsabstractRobots often have limited knowledge and need to continuously acquire new knowledge and skills in order to collaborate with its human partners. To address this issue, this paper describes an approach which allows human partners to teach a robot (i.e., a robotic arm) new high-level actions through natural language instructions. In particular, built upon the traditional planning framework, we propose a representation of high-level actions that only consists of the desired goal states rather than step-by-step operations (although these operations may be specified by the human in their instructions). Our empirical results have shown that, given this representation, the robot can reply on automated planning and immediately apply the newly learned action knowledge to perform actions under novel situations. Lanbo She, Yu Cheng 0006, Joyce Y. Chai, Yunyi Jia, Ning Xi 0001 |
RO-MAN | 3 |
| 2014 | Back to the Blocks World: Learning New Actions through Situated Human-Robot DialogueabstractThis paper describes an approach for a robotic arm to learn new actions through dialogue in a simplified blocks world. In particular, we have developed a three-tier action knowledge representation that on one hand, supports the connection be-tween symbolic representations of lan-guage and continuous sensorimotor repre-sentations of the robot; and on the other hand, supports the application of existing planning algorithms to address novel situ-ations. Our empirical studies have shown that, based on this representation the robot was able to learn and execute basic actions in the blocks world. When a human is engaged in a dialogue to teach the robot new actions, step-by-step instructions lead to better learning performance compared to one-shot instructions. 1 Lanbo She, Yu Cheng 0006, Yunyi Jia, Joyce Y. Chai, Ning Xi 0001 |
SIGDIAL Conference | 5 |
| 2013 | Towards Situated Dialogue: Revisiting Referring Expression GenerationabstractIn situated dialogue, humans and agents have mismatched capabilities of perceiving the shared environment.Their representations of the shared world are misaligned.Thus referring expression generation (REG) will need to take this discrepancy into consideration.To address this issue, we developed a hypergraph-based approach to account for group-based spatial relations and uncertainties in perceiving the environment.Our empirical results have shown that this approach outperforms a previous graph-based approach with an absolute gain of 9%.However, while these graph-based approaches perform effectively when the agent has perfect knowledge or perception of the environment (e.g., 84%), they perform rather poorly when the agent has imperfect perception of the environment (e.g., 45%).This big performance gap calls for new solutions to REG that can mediate a shared perceptual basis in situated dialogue. Changsong Liu, Lanbo She, Joyce Y. Chai |
EMNLP | 4 |
| 2013 | Modeling Collaborative Referring for Situated Referential Grounding
Changsong Liu, Lanbo She, Joyce Y. Chai |
SIGDIAL Conference | 4 |
| 2013 | Introduction to the special section on eye gaze and conversationabstractThis editorial introduction first explains the origin of this special section. It then outlines how each of the two articles included sheds light on possibilities for conversational dialog systems to use eye gaze as a signal that reflects aspects of participation in the dialog: degree of engagement and turn taking behavior, respectively. Elisabeth André, Joyce Y. Chai |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2012 | Integrating word acquisition and referential grounding towards physical world interactionabstractIn language-based interaction between a human and an artificial agent (e.g., robot) in a physical world, because the human and the agent have different knowledge and capabilities in perceiving the shared environment, referential grounding is very difficult. To facilitate such interaction, it is important for the agent to continuously learn and acquire knowledge about the environment through interactions with humans and incorporate the learned knowledge in grounding references from human utterances. To address this issue, this paper presents a graph-based approach for referential grounding and examines how referential grounding and word acquisition influence each other in physical world interaction. Our empirical results have shown that for most words, automated word acquisition through interaction improves referential grounding performance. However, this is not the case for words describing object types, where human supervision is important. Nevertheless, better referential grounding enables more accurate acquisition of word meanings, which in turn further improves grounding performance for references in subsequent utterances. Changsong Liu, Joyce Y. Chai |
ICMI | 3 |
| 2012 | Towards online adaptation and personalization of key-target resizing for mobile devicesabstractSoftware (soft) keyboards are becoming increasingly popular on mobile devices. To attempt to improve soft keyboard input accuracy, key-target resizing algorithms that dynamically change the size of each key's target area have been developed. Although methods that employ personalized touch models have been shown to outperform general models, previous work has relied upon laboratory-based offline calibration to collect the data necessary to build these models. Such approaches are unrealistic and interuptive, and it is unlikely that offline calibration can be applied in a realistic usage setting, as hundreds or thousands of touch points are necessary to build the models. To combat this problem, this paper explores the possibility of online adaptation of key-target resizing algorithms. In particular, we propose and examine three online data collection methods that can be used to build and dynamically update personalized key-target resizing models. Our results suggest that a data collection methodology that makes inference based on vocabulary and error correction behavior is able to perform on par with gold standard personalized models, while reducing relative error rate by 10.4% over general models. This approach is simple, computationally inexpensive, and calculable via information that the system already has access to. Additionally, we show that these models can be built quickly, requiring less than one week's worth of text input by an average mobile device user. Tyler Baldwin, Joyce Y. Chai |
IUI | 2 |
| 2012 | Autonomous Self-Assessment of Autocorrections: Exploring Text Message Dialogues
Tyler Baldwin, Joyce Y. Chai |
HLT-NAACL | 2 |
| 2012 | Towards Mediating Shared Perceptual Basis in Situated Dialogue
Changsong Liu, Joyce Y. Chai |
SIGDIAL Conference | 3 |
| 2012 | Semantic Role Labeling of Implicit Arguments for Nominal PredicatesabstractNominal predicates often carry implicit arguments. Recent work on semantic role labeling has focused on identifying arguments within the local context of a predicate; implicit arguments, however, have not been systematically examined. To address this limitation, we have manually annotated a corpus of implicit arguments for ten predicates from NomBank. Through analysis of this corpus, we find that implicit arguments add 71% to the argument structures that are present in NomBank. Using the corpus, we train a discriminative model that is able to identify implicit arguments with an F1 score of 50%, significantly outperforming an informed baseline model. This article describes our investigation, explores a wide variety of features important for the task, and discusses future directions for work on implicit argument identification. Matthew Gerber, Joyce Y. Chai |
Comput. Linguistics | 2 |
| 2012 | Introduction to the special issue on eye gaze in intelligent human-machine interactionabstractGiven the recent advances in eye tracking technology and the availability of nonintrusive and high-performance eye tracking devices, there has never been a better time to explore new opportunities to incorporate eye gaze in intelligent and natural human-machine communication. In this special issue, we present six articles that cover various aspects of eye gaze in human-machine interaction, including applications of gaze tracking in human-machine interaction, techniques that recognize gaze gestures and render gaze behaviors, and the analysis of gaze behaviors in social interactions. Elisabeth André, Joyce Y. Chai |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2011 | Beyond Normalization: Pragmatics of Word Form in Text Messages
Tyler Baldwin, Joyce Y. Chai |
IJCNLP | 2 |
| 2010 | Beyond NomBank: A Study of Implicit Arguments for Nominal Predicates
Matthew Gerber, Joyce Y. Chai |
ACL | 2 |
| 2010 | Fusing Eye Gaze with Speech Recognition Hypotheses to Resolve Exophoric References in Situated Dialogue
Zahar Prasov, Joyce Y. Chai |
EMNLP | 2 |
| 2010 | Towards Conversation Entailment: An Empirical Investigation
Joyce Y. Chai |
EMNLP | 2 |
| 2010 | Workshop: eye gaze in intelligent human machine interactionabstractThis workshop brought researchers from academia and industry together to share recent advances and discuss research directions and opportunities for next generation of intelligent human machine interaction that incorporate eye gaze. Elisabeth André, Joyce Y. Chai |
IUI | 2 |
| 2010 | Hand Gestures in Disambiguating Types of You Expressions in Multiparty Meetings
Tyler Baldwin, Joyce Y. Chai, Katrin Kirchhoff |
SIGDIAL Conference | 2 |
| 2010 | Context-based Word Acquisition for Situated Dialogue in a Virtual WorldabstractTo tackle the vocabulary problem in conversational systems, previous work has applied unsupervised learning approaches on co-occurring speech and eye gaze during interaction to automatically acquire new words. Although these approaches have shown promise, several issues related to human language behavior and human-machine conversation have not been addressed. First, psycholinguistic studies have shown certain temporal regularities between human eye movement and language production. While these regularities can potentially guide the acquisition process, they have not been incorporated in the previous unsupervised approaches. Second, conversational systems generally have an existing knowledge base about the domain and vocabulary. While the existing knowledge can potentially help bootstrap and constrain the acquired new words, it has not been incorporated in the previous models. Third, eye gaze could serve different functions in human-machine conversation. Some gaze streams may not be closely coupled with speech stream, and thus are potentially detrimental to word acquisition. Automated recognition of closely-coupled speech-gaze streams based on conversation context is important. To address these issues, we developed new approaches that incorporate user language behavior, domain knowledge, and conversation context in word acquisition. We evaluated these approaches in the context of situated dialogue in a virtual world. Our experimental results have shown that incorporating the above three types of contextual information significantly improves word acquisition performance. Shaolin Qu, Joyce Y. Chai |
J. Artif. Intell. Res. | 2 |
| 2009 | Communicative gestures in coreference identification in multiparty meetingsabstractDuring multiparty meetings, participants can use non-verbal modalities such as hand gestures to make reference to the shared environment. Therefore, one hypothesis is that incorporating hand gestures can improve coreference identification, a task that automatically identifies what participants refer to with their linguistic expressions. To evaluate this hypothesis, this paper examines the role of hand gestures in coreference identification, in particular, focusing on two questions: (1) what signals can distinguish communicative gestures that can potentially help coreference identification from non-communicative gestures; and (2) in what ways can communicative gestures help coreference identification. Based on the AMI data, our empirical results have shown that the length of gesture production is highly indicative of whether a gesture is communicative and potentially helpful in language understanding. Our experiments on the automated identification of coreferring expressions indicate that while the incorporation of simple gesture features does not improve overall performance, it does show potential on expressions referring to participants, an important and unique component of the meeting domain. A further analysis suggests that communicative gestures provide both redundant and complementary information, but further domain modeling and world knowledge incorporation is required to take full advantage of information that is complementary. Tyler Baldwin, Joyce Y. Chai, Katrin Kirchhoff |
ICMI | 2 |
| 2009 | Between linguistic attention and gaze fixations inmultimodal conversational interfacesabstractIn multimodal human machine conversation, successfully interpreting human attention is critical. While attention has been studied extensively in linguistic processing and visual processing, it is not clear how linguistic attention is aligned with visual attention in multimodal conversational interfaces. To address this issue, we conducted a preliminary investigation on how attention reflected by linguistic discourse aligns with attention indicated by gaze fixations during human machine conversation. Our empirical findings have shown that more attended entities based on linguistic discourse correspond to higher intensity of gaze fixations. The smoother a linguistic transition is, the less distance between corresponding fixation distributions. These findings provide insight into how language and gaze can be combined to predict attention, which have important implications in many tasks such as word acquisition and object recognition. Joyce Y. Chai, Fernanda Ferreira |
ICMI | 2 |
| 2009 | The Role of Implicit Argumentation in Nominal SRL
Matthew Gerber, Joyce Y. Chai, Adam Meyers 0001 |
HLT-NAACL | 2 |
| 2009 | The Role of Interactivity in Human-Machine Conversation for Automatic Word Acquisition
Shaolin Qu, Joyce Y. Chai |
SIGDIAL Conference | 2 |
| 2009 | What do We Know about Conversation Participants: Experiments on Conversation Entailment
Joyce Y. Chai |
SIGDIAL Conference | 2 |
| 2008 | Incorporating Temporal and Semantic Information with Eye Gaze for Automatic Word Acquisition in Multimodal Conversational Systems
Shaolin Qu, Joyce Y. Chai |
EMNLP | 2 |
| 2008 | What's in a gaze?: the role of eye-gaze in reference resolution in multimodal conversational interfacesabstractMultimodal conversational interfaces allow users to carry a dialog with a graphical display using speech to accomplish a particular task. Motivated by previous psycholinguistic findings, we examine how eye-gaze contributes to reference resolution in such a setting. Specifically, we present an integrated probabilistic framework that combines speech and eye-gaze for reference resolution. We further examine the relationship between eye-gaze and increased domain modeling with corresponding linguistic processing. Our empirical results show that the incorporation of eye-gaze significantly improves reference resolution performance. This improvement is most dramatic when a simple domain model is used. Our results also show that minimal domain modeling combined with eye-gaze significantly outperforms complex domain modeling without eye-gaze, which indicates that eye-gaze can be used to potentially compensate a lack of domain modeling for reference resolution. Zahar Prasov, Joyce Y. Chai |
IUI | 2 |
| 2008 | Beyond attention: the role of deictic gesture in intention recognition in multimodal conversational interfacesabstractIn a multimodal conversational interface supporting speech and deictic gesture, deictic gestures on the graphical display have been traditionally used to identify user attention, for example, through reference resolution. Since the context of the identified attention can potentially constrain the associated intention, our hypothesis is that deictic gestures can go beyond attention and apply to intention recognition. Driven by this assumption, this paper systematically investigates the role of deictic gestures in intention recognition. We experiment with different model-based methods and instancebased methods to incorporate gestural information for intention recognition. We examine the effects of utilizing gestural information in two different processing stages: speech recognition stage and language understanding stage. Our empirical results have shown that utilizing gestural information improves intention recognition. The performance is further improved when gestures are incorporated in both speech recognition and language understanding stages compared to either stage alone. Shaolin Qu, Joyce Y. Chai |
IUI | 2 |
| 2007 | Automated Vocabulary Acquisition and Interpretation in Multimodal Conversational Systems
Yi Liu 0054, Joyce Y. Chai, Rong Jin 0001 |
ACL | 2 |
| 2007 | An Exploration of Eye Gaze in Spoken Language Processing for Multimodal Conversational Interfaces
Shaolin Qu, Joyce Y. Chai |
HLT-NAACL | 2 |
| 2007 | Discourse processing for context question answering based on linguistic knowledge
Mingyu Sun, Joyce Y. Chai |
Knowl. Based Syst. | 2 |
| 2007 | An empirical investigation of user term feedback in text-based targeted image searchabstractText queries are natural and intuitive for users to describe their information needs. However, text-based image retrieval faces many challenges. Traditional text retrieval techniques on image descriptions have not been very successful. This is mainly due to the inconsistent textual descriptions and the discrepancies between user queries and terms in the descriptions. To investigate strategies to alleviate this vocabulary problem, this article examines the role of user term feedback in targeted image search that is based on text-based image retrieval. Term feedback refers to the feedback from a user on specific terms regarding their relevance to a target image. Previous studies have indicated the effectiveness of term feedback in interactive text retrieval. However, in our experiments on text-based image retrieval, the term feedback has not been shown to be effective. Our results indicate that, although term feedback has a positive effect by allowing users to identify more relevant terms, it also has a strong negative effect by providing more opportunities for users to specify irrelevant terms. To understand these different effects and their implications, this article further analyzes important factors that contribute to the utility of term feedback and discusses the outlook of term feedback in interactive text-based image retrieval. Joyce Y. Chai, Rong Jin 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2006 | Towards Conversational QA: Automatic Identification of Problematic Situations and User Intent
Joyce Y. Chai, Tyler Baldwin |
ACL | 1 |
| 2006 | Salience modeling based on non-verbal modalities for spoken language understandingabstractPrevious studies have shown that, in multimodal conversational systems, fusing information from multiple modalities together can improve the overall input interpretation through mutual disambiguation. Inspired by these findings, this paper investigates non-verbal modalities, in particular deictic gesture, in spoken language processing. Our assumption is that during multimodal conversation, user's deictic gestures on the graphic display can signal the underlying domain model that is salient at that particular point of interaction. This salient domain model can be used to constrain hypotheses for spoken language processing. Based on this assumption, this paper examines different configurations of salience driven language models (e.g., n-gram and probabilistic context free grammar) for spoken language processing across different stages. Our empirical results have shown the potential of integrating salience models based on non-verbal modalities in spoken language understanding. Shaolin Qu, Joyce Y. Chai |
ICMI | 2 |
| 2006 | Towards intelligent QA interfaces: discourse processing for context questionsabstractQuestion answering (QA) systems take users' natural language questions and retrieve relevant answers from large repositories of free texts. Despite recent progress in QA research, most work on question answering is still focused on isolated questions. In a real-world information seeking scenario, questions are not asked in isolation, but rather in a coherent manner that involves a sequence of related questions to meet users' information needs. Therefore, to support coherent information seeking, intelligent QA interfaces will inevitably require techniques to support context question answering. To address this problem, this paper investigates approaches to discourse processing of a sequence of coherent questions and their implications on query expansion. In particular, we examine three models for query expansion that are motivated by Centering Theory. Our empirical results indicate that more sophisticated processing based on discourse transitions and centers can significantly improve the performance of document retrieval compared to models that only resolve references. Mingyu Sun, Joyce Y. Chai |
IUI | 2 |
| 2006 | Automated performance assessment in interactive QAabstractIn interactive question answering (QA), users and systems take turns to ask questions and provide answers. In such an interactive setting, user questions largely depend on the answers provided by the system. One question is whether user follow-up questions can provide feedback for the system to automatically assess its performance (e.g., assess whether a correct answer is delivered). This self-awareness can make QA systems more intelligent for information seeking, for example, by adapting better strategies to cope with problematic situations. Therefore, this paper describes our initial investigation in addressing this problem. Our results indicate that interaction context can provide useful cues for automated performance assessment in interactive QA. Joyce Y. Chai, Tyler Baldwin |
SIGIR | 1 |
| 2006 | Cognitive Principles in Robust Multimodal InterpretationabstractMultimodal conversational interfaces provide a natural means for users to communicate with computer systems through multiple modalities such as speech and gesture. To build effective multimodal interfaces, automated interpretation of user multimodal inputs is important. Inspired by the previous investigation on cognitive status in multimodal human machine interaction, we have developed a greedy algorithm for interpreting user referring expressions (i.e., multimodal reference resolution). This algorithm incorporates the cognitive principles of Conversational Implicature and Givenness Hierarchy and applies constraints from various sources (e.g., temporal, semantic, and contextual) to resolve references. Our empirical results have shown the advantage of this algorithm in efficiently resolving a variety of user references. Because of its simplicity and generality, this approach has the potential to improve the robustness of multimodal input interpretation. Joyce Y. Chai, Zahar Prasov, Shaolin Qu |
J. Artif. Intell. Res. | 1 |
| 2006 | A statistical framework for query translation disambiguationabstractResolving ambiguity in the process of query translation is crucial to cross-language information retrieval (CLIR), given the short length of queries. This problem is even more challenging when only a bilingual dictionary is available, which is the focus of our work described here. In this paper, we will present a statistical framework for dictionary-based CLIR that estimates the translation probabilities of query words based on the monolingual word co-occurrence statistics. In addition, we will present two realizations of the proposed framework, i.e., the “maximum coherence model” and the “spectral query-translation model,” that exploit different metrics for the coherence measurement between a translation of a query word and the theme of the entire query. Compared to previous work on dictionary-based CLIR, the proposed framework is advantageous in three aspects: (1) Translation probabilities are calculated explicitly to capture the uncertainty in translating queries; (2) translations of all query words are estimated simultaneously rather than independently; and (3) the formulated problem can be solved efficiently with a unique optimal solution. Empirical studies with Chinese--English cross-language information retrieval using TREC datasets have shown that the proposed models achieve a relative 10%--50% improvement, compared to other approaches that also exploit word co-occurrence statistics for query translation disambiguation. Yi Liu 0054, Rong Jin 0001, Joyce Y. Chai |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2005 | Learn to weight terms in information retrieval using category informationabstractHow to assign appropriate weights to terms is one of the critical issues in information retrieval. Many term weighting schemes are unsupervised. They are either based on the empirical observation in information retrieval, or based on generative approaches for language modeling. As a result, the existing term weighting schemes are usually insufficient in distinguishing informative words from the uninformative ones, which is crucial to the performance of information retrieval. In this paper, we present supervised term weighting schemes that automatically learn term weights based on the correlation between word frequency and category information of documents. Empirical studies with the ImageCLEF dataset have indicated that the proposed methods perform substantially better than the state-of-the-art approaches for term weighting and other alternatives that exploit category information for information retrieval. Rong Jin 0001, Joyce Y. Chai, Luo Si |
ICML | 2 |
| 2005 | Linguistic theories in efficient multimodal reference resolution: an empirical investigationabstractMultimodal conversational interfaces provide a natural means for users to communicate with computer systems through multiple modalities such as speech, gesture, and gaze. To build effective multimodal interfaces, understanding user multimodal inputs is important. Previous linguistic and cognitive studies indicate that user language behavior does not occur randomly, but rather follows certain linguistic and cognitive principles. Therefore, this paper investigates the use of linguistic theories in multimodal interpretation. In particular, we present a greedy algorithm that incorporates Conversation Implicature and Givenness Hierarchy for efficient multimodal reference resolution. Empirical studies indicate that this algorithm significantly reduces the complexity in multimodal reference resolution compared to a previous graph-matching approach. One major advantage of this greedy algorithm is that the prior linguistic and cognitive knowledge can be used to guide the search and significantly prune the search space. Because of its simplicity and generality, this approach has the potential to improve the robustness of interpretation and provide a more practical solution to multimodal input interpretation. Joyce Y. Chai, Zahar Prasov, Joseph Blaim, Rong Jin 0001 |
IUI | 1 |
| 2005 | Study of cross lingual information retrieval using on-line translation systemsabstractTypical cross language retrieval requires special linguistic resources, such as bilingual dictionaries and parallel corpus. In this study, we focus on the cross lingual retrieval problem that only uses online translation systems. We compare two approaches: a translation-based approach that directly translates queries into the language of documents and then applies traditional information retrieval techniques; and a model-based approach that first learns a statistical translation model from the translations acquired from an online translation system and then applies the learned statistical model to cross lingual information retrieval. Our empirical study with ImageCLEF has shown the model-based approach performs significantly better than the translation-based approach. Rong Jin 0001, Joyce Y. Chai |
SIGIR | 2 |
| 2005 | A maximum coherence model for dictionary-based cross-language information retrievalabstractOne key to cross-language information retrieval is how to efficiently resolve the translation ambiguity of queries given their short length. This problem is even more challenging when only bilingual dictionaries are available, which is the focus of this paper. In the previous research of cross-language information retrieval using bilingual dictionaries, the word co-occurrence statistics is used to determine the most likely translations of queries. In this paper, we propose a novel statistical model, named ``maximum coherence model'', which estimates the translation probabilities of query words that are consistent with the word co-occurrence statistics. Unlike the previous work, where a binary decision is made for the selection of translations, the new model maintains the uncertainty in translating query words when their sense ambiguity is difficult to resolve. Furthermore, this new model is able to estimate translations of multiple query words simultaneously. This is in contrast to many previous approaches where translations of individual query words are determined independently. Empirical studies with TREC datasets have shown that the maximum coherence model achieves a relative 10% - 40% improvement in cross-language information retrieval, comparing to other approaches that also use word co-occurrence statistics for sense disambiguation. Yi Liu 0054, Rong Jin 0001, Joyce Y. Chai |
SIGIR | 3 |
| 2005 | User term feedback in interactive text-based image retrievalabstractTo alleviate the vocabulary problem, this paper investigates the role of user term feedback in interactive text-based image retrieval. Term feedback refers to the feedback from a user on specific terms regarding their relevance to a target image. Previous studies have indicated the effectiveness of term feedback in interactive text retrieval [14]. However, the term feedback has not shown to be effective in our experiments on text-based image retrieval. Our results indicate that, although term feedback has a positive effect by allowing users to identify more relevant terms, it also has a strong negative effect by providing more opportunities for users to specify irrelevant terms. To understand these different effects and their implications on the potential of term feedback, this paper further presents analysis of important factors that contribute to the utility of term feedback and discusses the outlook of term feedback in interactive text-based image retrieval. Joyce Y. Chai, Rong Jin 0001 |
SIGIR | 2 |
| 2004 | Optimization in Multimodal InterpretationabstractIn a multimodal conversation, the way users communicate with a system depends on the available interaction channels and the situated context (e.g., conversation focus, visual feedback). These dependencies form a rich set of constraints from various perspectives such as temporal alignments between different modalities, coherence of conversation, and the domain semantics. There is strong evidence that competition and ranking of these constraints is important to achieve an optimal interpretation. Thus, we have developed an optimization approach for multimodal interpretation, particularly for interpreting multimodal references. A preliminary evaluation indicates the effectiveness of this approach, especially for complex user inputs that involve multiple referring expressions in a speech utterance and multiple gestures. Joyce Y. Chai, Pengyu Hong, Michelle X. Zhou, Zahar Prasov |
ACL | 1 |
| 2004 | Regularizing translation models for better automatic image annotationabstractThe goal of automatic image annotation is to automatically generate annotations for images to describe their content. In the past, statistical machine translation models have been successfully applied to automatic image annotation task [8]. It views the process of annotating images as a process of translating the content from a 'visual language' to textual words. One problem with the existing translation models is that common words are usually associated with too many different image regions. As a result, uncommon words have little chance to be used for annotating images. Uncommon words are important for automatic image annotation because they are often used in the queries. In this paper, we propose two modified translation models for automatic image annotation, namely the normalized translation model and the regularized translation model, that specifically address the problem of common annotated words. The basic idea is to raise the number of blobs that are associated with uncommon words. The normalized translation model realizes this by scaling translation probabilities of different words with different factors. The same goal is achieved in the regularized translation model through the introduction of a special Dirichlet prior. Empirical study with the Corel dataset has shown that both two modified translation models outperform the original translation model and several existing approaches for automatic image annotation substantially. Feng Kang, Rong Jin 0001, Joyce Y. Chai |
CIKM | 3 |
| 2004 | A probabilistic approach to reference resolution in multimodal user interfacesabstractMultimodal user interfaces allow users to interact with computers through multiple modalities, such as speech, gesture, and gaze. To be effective, multimodal user interfaces must correctly identify all objects which users refer to in their inputs. To systematically resolve different types of references, we have developed a probabilistic approach that uses a graph-matching algorithm. Our approach identifies the most probable referents by optimizing the satisfaction of semantic, temporal, and contextual constraints simultaneously. Our preliminary user study results indicate that our approach can successfully resolve a wide variety of referring expressions, ranging from simple to complex and from precise to ambiguous ones. Joyce Y. Chai, Pengyu Hong, Michelle X. Zhou |
IUI | 1 |
| 2004 | Effective automatic image annotation via a coherent language model and active learningabstractImage annotations allow users to access a large image database with textual queries. There have been several studies on automatic image annotation utilizing machine learning techniques, which automatically learn statistical models from annotated images and apply them to generate annotations for unseen images. One common problem shared by most previous learning approaches for automatic image annotation is that each annotated word is predicated for an image independently from other annotated words. In this paper, we proposed a coherent language model for automatic image annotation that takes into account the word-to-word correlation by estimating a coherent language model for an image. This new approach has two important advantages: 1) it is able to automatically determine the annotation length to improve the accuracy of retrieval results, and 2) it can be used with active learning to significantly reduce the required number of annotated image examples. Empirical studies with Corel dataset are presented to show the effectiveness of the coherent language model for automatic image annotation. Rong Jin 0001, Joyce Y. Chai, Luo Si |
ACM Multimedia | 2 |
| 2004 | An automatic weighting scheme for collaborative filteringabstractCollaborative filtering identifies information interest of a particular user based on the information provided by other similar users. The memory-based approaches for collaborative filtering (e.g., Pearson correlation coefficient approach) identify the similarity between two users by comparing their ratings on a set of items. In these approaches, different items are weighted either equally or by some predefined functions. The impact of rating discrepancies among different users has not been taken into consideration. For example, an item that is highly favored by most users should have a smaller impact on the user-similarity than an item for which different types of users tend to give different ratings. Even though simple weighting methods such as variance weighting try to address this problem, empirical studies have shown that they are ineffective in improving the performance of collaborative filtering. In this paper, we present an optimization algorithm to automatically compute the weights for different items based on their ratings from training users. More specifically, the new weighting scheme will create a clustered distribution for user vectors in the item space by bringing users of similar interests closer and separating users of different interests more distant. Empirical studies over two datasets have shown that our new weighting scheme substantially improves the performance of the Pearson correlation coefficient method for collaborative filtering. Rong Jin 0001, Joyce Y. Chai, Luo Si |
SIGIR | 2 |
| 2002 | Semantics-based Representation for Multimodal Interpretation in Conversational Systems
Joyce Y. Chai |
COLING | 1 |
| 2002 | Context-Based Multimodal Input Understanding in Conversational SystemsabstractIn a multimodal human-machine conversation, user inputs are often abbreviated or imprecise. Sometimes, merely fusing multimodal inputs together cannot derive a complete understanding. To address these inadequacies, we are building a semantics-based multimodal interpretation framework called MIND (Multimodal Interpretation for Natural Dialog). The unique feature of MIND is the use of a variety of contexts (e.g., domain context and conversation context) to enhance multimodal fusion. In this paper we present a semantically rich modeling scheme and a context-based approach that enable MIND to gain a full understanding of user inputs, including ambiguous and incomplete ones. Joyce Y. Chai, Shimei Pan, Michelle X. Zhou, Keith Houck |
ICMI | 1 |
| 2002 | Operations for context-based multimodal interpretation in conversational systems
Joyce Y. Chai |
INTERSPEECH | 1 |
| 2001 | Natural Language Sales Assistant - A Web-Based Dialog System for Online Sales
Joyce Y. Chai, Malgorzata Budzikowska, Veronika Horvath, Nicolas Nicolov, Nanda Kambhatla, Wlodek Zadrozny |
IAAI | 1 |
| 2000 | A multi-modal dialog system for business transactions
Joyce Y. Chai, Sylvie Levesque, Malgorzata Budzikowska, Veronika Horvath, Nanda Kambhatla, Nicolas Nicolov, Wlodek Zadrozny |
INTERSPEECH | 1 |
| 2000 | Evaluation of a Generic Lexical Semantic Resource in Information Extraction
Joyce Y. Chai |
LREC | 1 |
| 1997 | A Trainable Message Understanding System
Amit Bagga, Joyce Y. Chai |
CoNLL | 2 |
| 1997 | A WordNet Based Rule Generalization Engine for Meaning Extraction System
Joyce Y. Chai, Alan W. Biermann |
ISMIS | 1 |