VLDB 2026 Research / reviewers in the wild / expert
Seungwhan Moon
dblp:120/4131
· DBLP profile ↗
35ranked-venue papers
9as first author
20since 2021 · last 2025
0009-0003-2507-1884ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proactive Assistant Dialogue Generation from Streaming Egocentric VideosabstractYichi Zhang, Xin Luna Dong, Zhaojiang Lin, Andrea Madotto, Anuj Kumar, Babak Damavandi, Joyce Chai, Seungwhan Moon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yichi Zhang 0001, Xin Dong 0001, Zhaojiang Lin, Andrea Madotto, Babak Damavandi, Joyce Y. Chai, Seungwhan Moon |
EMNLP | 8 |
| 2025 | WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenariosabstractWe introduce WearVQA, the first benchmark specifically designed to evaluate the visual questionanswering (VQA) capabilities of multi-modal AI assistant on wearable devices like smart glasses. Unlikeprior benchmarks that focus on high-quality, third-person imagery, WearVQA reflects the unique chal-lenges of ego-centric interaction—where visual inputs may be occluded, poorly lit, unzoomed, or blurry,and questions are grounded in realistic wearable use cases. The benchmark comprises 2,500 carefullycurated image-question-answer triplets, spanning 7 diverse image domains including both text-centricand general scenes, 10 cognitive task types ranging from basic recognition to various forms of reasoning,and 6 common wearables-specific image quality issues. All questions are designed to be answerable usingonly the visual input and common senses. WearVQA is paired with a rigorous LLM-as-a-judge evaluationframework with 96% labeling accuracy. Open-source and proprietary multi-modal LLMs achieved a QAaccuracy as low as 24–52% on WearVQA, with substantial drops on lower-quality images and reasoning-heavy tasks. These observations position WearVQA as a comprehensive and challenging benchmark forguiding technicial advancement towards robust, real-world multi-modal wearables AI systems. Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, Adel Ahmadyan, Akil Iyer, Elissa Li, Nicolas Scheffer, Ahmed Kirmani, Babak Damavandi, Rakesh Wanga, Rohit Patel, Seungwhan Moon, Xin Dong 0001 |
NeurIPS | 21 |
| 2025 | PerceptionLM: Open-Access Data and Models for Detailed Visual UnderstandingabstractVision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The research community has responded by using distillation from black-box models to label training data, achieving strong benchmark results, at the cost of measurable scientific progress. However, without knowing the details of the teacher model and its data sources, scientific progress remains difficult to measure. In this paper, we study building a Perception Language Model (PLM) in a fully open and reproducible framework for transparent research in image and video understanding. We analyze standard training pipelines without distillation from proprietary models and explore large-scale synthetic data to identify critical data gaps, particularly in detailed video understanding. To bridge these gaps, we release 2.8M human-labeled instances of fine-grained video question-answer pairs and spatio-temporally grounded video captions. Additionally, we introduce PLM–VideoBench, a suite for evaluating challenging video understanding tasks focusing on the ability to reason about ''what'', ''where'', ''when'', and ''how'' of a video. We make our work fully reproducible by providing data, training recipes, code & models. Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi, Triantafyllos Afouras, Tushar Nagarajan, Muhammad Maaz 0001, Yale Song, Tengyu Ma 0005, Shuming Hu, Suyog Dutt Jain, Hanoona Abdul Rasheed, Peize Sun, Po-Yao Huang 0001, Daniel Bolya, Nikhila Ravi, Shashank Jain, Tammy Stark, Seungwhan Moon, Babak Damavandi, Vivian Lee, Andrew Westbury, Salman Khan 0001, Philipp Krähenbühl, Piotr Dollár, Lorenzo Torresani, Kristen Grauman, Christoph Feichtenhofer |
NeurIPS | 20 |
| 2025 | VisualLens: Personalization through Task-Agnostic Visual HistoryabstractExisting recommendation systems either rely on user interaction logs, such as online shopping history for shopping recommendations, or focus on text signals.
However, item-based histories are not always accessible and generalizable for multimodal recommendation.
We hypothesize that a user's visual history --- comprising images from daily life --- can offer rich, task-agnostic insights into their interests and preferences, and thus be leveraged for effective personalization.
To this end, we propose VisualLens, a novel framework that leverages multimodal large language models (MLLMs) to enable personalization using task-agnostic visual history.
VisualLens extracts, filters, and refines a spectrum user profile from the visual history to support personalized recommendation.
We created two new benchmarks, Google-Review-V and Yelp-V, with task-agnostic visual histories, and show that VisualLens improves over state-of-the-art item-based multimodal recommendations by 5-10\% on Hit@3, and outperforms GPT-4o by 2-5\%.
Further analysis shows that VisualLens is robust across varying history lengths and excels at adapting to both longer histories and unseen content categories. Wang Zhu 0001, Deqing Fu, Kai Sun 0006, Zhaojiang Lin, Seungwhan Moon, Kanika Narang, Mustafa Canim, Xin Dong 0001 |
NeurIPS | 6 |
| 2025 | Corgi: Cached Memory Guided Video GenerationabstractText-to-Video generation has achieved remarkable progress with the rise of diffusion models. In this work, we introduce Cached Memory-Guided Video Generation (Corgi), aiming to generate multi-scene videos with arbi-trary number of video clips, conditioned on input images and instruction prompts. This is a challenging task, as tra-ditional T2V methods often struggle to maintain the quality of longer videos due to the difficulties in preserving visual context from earlier scenes. We address this by introducing a cached memory mechanism that stores the key frames. Our multi-scene video generation process is explicitly con-ditioned on the cached memories to avoid forgetting the vi-sual appearance of target subjects. Corgi shows significant improvement in multi-scene video generation compared to the prior art, with up to 59.2% in long-term consistency and 7.6% in diversity. Xindi Wu, Uriel Singer, Zhaojiang Lin, Andrea Madotto, Xide Xia, Paul A. Crook, Xin Dong 0001, Seungwhan Moon |
WACV | 9 |
| 2024 | Large Language Models as Zero-shot Dialogue State Tracker through Function CallingabstractZekun Li, Zhiyu Zoey Chen, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Luna Dong, Adithya Sagar, Xifeng Yan, Paul A. Crook. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zekun Li 0001, Zhiyu Chen 0002, Mike Ross, Patrick Huber, Seungwhan Moon, Zhaojiang Lin, Xin Dong 0001, Adithya Sagar, Xifeng Yan, Paul A. Crook |
ACL (1) | 5 |
| 2024 | Embodied Executable Policy Learning with Language-based Scene SummarizationabstractJielin Qiu, Mengdi Xu, William Han, Seungwhan Moon, Ding Zhao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jielin Qiu, Mengdi Xu, William Jongwon Han, Seungwhan Moon, Ding Zhao |
NAACL-HLT | 4 |
| 2024 | Overview of the Ninth Dialog System Technology Challenge: DSTC9abstractThis paper introduces the Ninth Dialog System Technology Challenge (DSTC-9). This edition of the DSTC focuses on applying end-to-end dialog technologies for four distinct tasks in dialog systems, namely, 1. Task-oriented dialog Modeling with Unstructured Knowledge Access, 2. Multi-domain task-oriented dialog, 3. Interactive evaluation of dialog and 4. Situated interactive multimodal dialog. This paper describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. R. Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D'Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan 0001, Yang Liu 0004, Chao-Wei Huang, Dilek Hakkani-Tür, Jinchao Li, Qi Zhu 0007, Lingxiao Luo, Lars Liden, Kaili Huang, Shahin Shayandeh, Runze Liang, Baolin Peng, Zheng Zhang 0020, Swadheen Shukla, Minlie Huang, Jianfeng Gao 0001, Shikib Mehri, Yulan Feng, Carla Gordon, Seyed Hossein Alavi, David R. Traum, Maxine Eskénazi, Ahmad Beirami, Eunjoon Cho, Paul A. Crook, Ankita De, Alborz Geramifard, Satwik Kottur, Seungwhan Moon, Shivani Poddar, Rajen Subba |
IEEE ACM Trans. Audio Speech Lang. Process. | 36 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2023 | SIMMC-VR: A Task-oriented Multimodal Dialog Dataset with Situated and Immersive VR StreamsabstractTe-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodriguez, Babak Damavandi, Nanyun Peng, Seungwhan Moon. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Te-Lin Wu, Satwik Kottur, Andrea Madotto, Mahmoud Azab, Pedro Rodríguez 0001, Babak Damavandi, Nanyun Peng 0001, Seungwhan Moon |
ACL (1) | 8 |
| 2023 | Towards Next-Generation Intelligent Assistants Leveraging LLM TechniquesabstractVirtual Intelligent Assistants take user requests in the voice form, perform actions such as setting an alarm, turning on a light, and answering a question, and provide answers or confirmations in the voice form or through other channels such as a screen. Assistants have become prevalent in the past decade, and users have been taking services from assistants like Amazon Alexa, Apple Siri, Google Assistant, and Microsoft Cortana. Xin Dong 0001, Seungwhan Moon, Yifan Ethan Xu, Kshitiz Malik, Zhou Yu 0005 |
KDD | 2 |
| 2022 | Navigating Connected Memories with a Task-oriented Dialog SystemabstractRecent years have seen an increasing trend in the volume of personal media captured by users, thanks to the advent of smartphones and smart glasses, resulting in large media collections.Despite conversation being an intuitive humancomputer interface, current efforts focus mostly on single-shot natural language based media retrieval to aid users query their media and re-live their memories.This severely limits the search functionality as users can neither ask followup queries nor obtain information without first formulating a single-turn query.In this work, we propose dialogs for connected memories as a powerful tool to empower users to search their media collection through a multiturn, interactive conversation.Towards this, we collect a new task-oriented dialog dataset COMET, which contains 11.5k user↔assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.We employ a resource-efficient, two-phase data collection pipeline that uses: (1) a novel multimodal dialog simulator that generates synthetic dialog flows grounded in memory graphs, and, (2) manual paraphrasing to obtain natural language utterances.We analyze COMET, formulate four main tasks to benchmark meaningful progress, and adopt state-of-the-art language models as strong baselines, in order to highlight the multimodal challenges captured by our dataset 1 . Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi |
EMNLP | 2 |
| 2022 | Normalized Contrastive Learning for Text-Video RetrievalabstractCross-modal contrastive learning has led the recent advances in multimodal retrieval with its simplicity and effectiveness.In this work, however, we reveal that cross-modal contrastive learning suffers from incorrect normalization of the sum retrieval probabilities of each text or video instance.Specifically, we show that many test instances are either overor under-represented during retrieval, significantly hurting the retrieval performance.To address this problem, we propose Normalized Contrastive Learning (NCL) which utilizes the Sinkhorn-Knopp algorithm to compute the instance-wise biases that properly normalize the sum retrieval probabilities of each instance so that every text and video instance is fairly represented during cross-modal retrieval.Empirical study shows that NCL brings consistent and significant gains in text-video retrieval on different model architectures, with new stateof-the-art multimodal retrieval metrics on the ActivityNet, MSVD, and MSR-VTT datasets without any architecture engineering. Yookoon Park, Mahmoud Azab, Seungwhan Moon, Florian Metze, Gourab Kundu, Ahmed Kirmani |
EMNLP | 3 |
| 2021 | DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded DialogueabstractHung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hung Le 0003, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami, Alborz Geramifard, Satwik Kottur |
ACL/IJCNLP (1) | 3 |
| 2021 | SIMMC 2.0: A Task-oriented Dialog Dataset for Immersive Multimodal ConversationsabstractNext generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment.Existing task-oriented dialog datasets aimed towards virtual assistance fall short and do not situate the dialog in the user's multimodal context.To overcome, we present a new dataset for Situated and Interactive Multimodal Conversations, SIMMC 2.0, which includes 11K taskoriented user$assistant dialogs (117K utterances) in the shopping domain, grounded in immersive and photo-realistic scenes.The dialogs are collected using a two-phase pipeline: (1) A novel multimodal dialog simulator generates simulated dialog flows, with an emphasis on diversity and richness of interactions, (2) Manual paraphrasing of the generated utterances to collect diverse referring expressions.We provide an in-depth analysis of the collected dataset, and describe in detail the four main benchmark tasks we propose.Our baseline model, powered by the state-of-theart language model, shows promising results, and highlights new challenges and directions for the community to study 1 . Satwik Kottur, Seungwhan Moon, Alborz Geramifard, Babak Damavandi |
EMNLP (1) | 2 |
| 2021 | Zero-Shot Dialogue State Tracking via Cross-Task TransferabstractZhaojiang Lin, Bing Liu, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul Crook, Zhiguang Wang, Zhou Yu, Eunjoon Cho, Rajen Subba, Pascale Fung. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhaojiang Lin, Andrea Madotto, Seungwhan Moon, Zhenpeng Zhou, Paul A. Crook, Zhiguang Wang, Zhou Yu 0005, Eunjoon Cho, Rajen Subba, Pascale Fung |
EMNLP (1) | 4 |
| 2021 | Continual Learning in Task-Oriented Dialogue SystemsabstractAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul Crook, Bing Liu, Zhou Yu, Eunjoon Cho, Pascale Fung, Zhiguang Wang. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Andrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon, Paul A. Crook, Zhou Yu 0005, Eunjoon Cho, Pascale Fung, Zhiguang Wang |
EMNLP (1) | 4 |
| 2021 | Leveraging Slot Descriptions for Zero-Shot Cross-Domain Dialogue StateTrackingabstractZhaojiang Lin, Bing Liu, Seungwhan Moon, Paul Crook, Zhenpeng Zhou, Zhiguang Wang, Zhou Yu, Andrea Madotto, Eunjoon Cho, Rajen Subba. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Zhaojiang Lin, Seungwhan Moon, Paul A. Crook, Zhenpeng Zhou, Zhiguang Wang, Andrea Madotto, Eunjoon Cho, Rajen Subba |
NAACL-HLT | 3 |
| 2021 | Adding Chit-Chat to Enhance Task-Oriented DialoguesabstractKai Sun, Seungwhan Moon, Paul Crook, Stephen Roller, Becka Silvert, Bing Liu, Zhiguang Wang, Honglei Liu, Eunjoon Cho, Claire Cardie. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Kai Sun 0006, Seungwhan Moon, Paul A. Crook, Stephen Roller, Becka Silvert, Zhiguang Wang, Eunjoon Cho, Claire Cardie |
NAACL-HLT | 2 |
| 2021 | An Analysis of State-of-the-Art Models for Situated Interactive MultiModal Conversations (SIMMC)abstractSatwik Kottur, Paul Crook, Seungwhan Moon, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard. Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2021. Satwik Kottur, Paul A. Crook, Seungwhan Moon, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard |
SIGDIAL | 3 |
| 2020 | Situated and Interactive Multimodal ConversationsabstractSeungwhan Moon, Satwik Kottur, Paul Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard. Proceedings of the 28th International Conference on Computational Linguistics. 2020. Seungwhan Moon, Satwik Kottur, Paul A. Crook, Ankita De, Shivani Poddar, Theodore Levin, David Whitney, Daniel Difranco, Ahmad Beirami, Eunjoon Cho, Rajen Subba, Alborz Geramifard |
COLING | 1 |
| 2020 | User Memory Reasoning for Conversational RecommendationabstractWe study an end-to-end approach for conversational recommendation that dynamically manages and reasons over users' past (offline) preferences and current (online) requests through a structured and cumulative user memory knowledge graph. This formulation extends existing state tracking beyond the boundary of a single dialog to user state tracking (UST). For this study, we create a new Memory Graph (MG) <-> Conversational Recommendation parallel corpus called MGConvRex with 7K+ human-to-human role-playing dialogs, grounded on a large-scale user memory bootstrapped from real-world user scenarios. MGConvRex captures human-level reasoning over user memory and has disjoint training/testing sets of users for zero-shot (cold-start) reasoning for recommendation. We propose a simple yet expandable formulation for constructing and updating the MG, and an end-to-end graph-based reasoning model that updates MG from unstructured utterances and predicts optimal dialog policies (eg recommendation) based on updated MG. The prediction of our proposed model inherits the graph structure, providing a natural way to explain policies. Experiments are conducted for both offline metrics and online simulation, showing competitive results. Hu Xu 0001, Seungwhan Moon, Honglei Liu 0001, Bing Liu 0024, Pararth Shah, Philip S. Yu |
COLING | 2 |
| 2020 | Information Seeking in the Spirit of Learning: A Dataset for Conversational CuriosityabstractOpen-ended human learning and information-seeking are increasingly mediated by digital assistants. However, such systems often ignore the user's pre-existing knowledge. Assuming a correlation between engagement and user responses such as "liking" messages or asking followup questions, we design a Wizard-of-Oz dialog task that tests the hypothesis that engagement increases when users are presented with facts related to what they know. Through crowd-sourcing of this experiment, we collect and release 14K dialogs (181K utterances) where users and assistants converse about geographic topics like geopolitical entities and locations. This dataset is annotated with pre-existing user knowledge, message-level dialog acts, grounding to Wikipedia, and user reactions to messages. Responses using a user's prior knowledge increase engagement. We incorporate this knowledge into a multi-task model that reproduces human assistant policies and improves over a BERT content model by 13 mean reciprocal rank points. Pedro Rodríguez 0001, Paul A. Crook, Seungwhan Moon, Zhiguang Wang |
EMNLP (1) | 3 |
| 2019 | OpenDialKG: Explainable Conversational Reasoning with Attention-based Walks over Knowledge GraphsabstractWe study a conversational reasoning model that strategically traverses through a largescale common fact knowledge graph (KG) to introduce engaging and contextually diverse entities and attributes.For this study, we collect a new Open-ended Dialog ↔ KG parallel corpus called OpenDialKG, where each utterance from 15K human-to-human roleplaying dialogs is manually annotated with ground-truth reference to corresponding entities and paths from a large-scale KG with 1M+ facts.We then propose the DialKG Walker model that learns the symbolic transitions of dialog contexts as structured traversals over KG, and predicts natural entities to introduce given previous dialog contexts via a novel domain-agnostic, attention-based graph path decoder.Automatic and human evaluations show that our model can retrieve more natural and human-like responses than the state-ofthe-art baselines or rule-based models, in both in-domain and cross-domain tasks.The proposed model also generates a KG walk path for each entity retrieved, providing a natural way to explain conversational reasoning. Seungwhan Moon, Pararth Shah, Rajen Subba |
ACL (1) | 1 |
| 2019 | Memory Graph Networks for Explainable Memory-grounded Question AnsweringabstractWe introduce Episodic Memory QA, the task of answering personal user questions grounded on memory graph (MG), where episodic memories and related entity nodes are connected via relational edges.We create a new benchmark dataset first by generating synthetic memory graphs with simulated attributes, and by composing 100K QA pairs for the generated MG with bootstrapped scripts.To address the unique challenges for the proposed task, we propose Memory Graph Networks (MGN), a novel extension of memory networks to enable dynamic expansion of memory slots through graph traversals, thus able to answer queries in which contexts from multiple linked episodes and external knowledge are required.We then propose the Episodic Memory QA Net with multiple module networks to effectively handle various question types.Empirical results show improvement over the QA baselines in top-k answer prediction accuracy in the proposed task.The proposed model also generates a graph walk path and attention vectors for each predicted answer, providing a natural way to explain its QA reasoning. Seungwhan Moon, Pararth Shah, Rajen Subba |
CoNLL | 1 |
| 2018 | Multimodal Named Entity Disambiguation for Noisy Social Media PostsabstractWe introduce the new Multimodal Named Entity Disambiguation (MNED) task for multimodal social media posts such as Snapchat or Instagram captions, which are composed of short captions with accompanying images.Social media posts bring significant challenges for disambiguation tasks because 1) ambiguity not only comes from polysemous entities, but also from inconsistent or incomplete notations, 2) very limited context is provided with surrounding words, and 3) there are many emerging entities often unseen during training.To this end, we build a new dataset called SnapCaptionsKB, a collection of Snapchat image captions submitted to public and crowd-sourced stories, with named entity mentions fully annotated and linked to entities in an external knowledge base.We then build a deep zeroshot multimodal network for MNED that 1) extracts contexts from both text and image, and 2) predicts correct entity in the knowledge graph embeddings space, allowing for zeroshot disambiguation of entities unseen in training set as well.The proposed model significantly outperforms the stateof-the-art text-only NED models, showing efficacy and potentials of the MNED task. Seungwhan Moon, Leonardo Neves, Vitor Carvalho |
ACL (1) | 1 |
| 2018 | Multimodal Named Entity Recognition for Short Social Media PostsabstractSeungwhan Moon, Leonardo Neves, Vitor Carvalho. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Seungwhan Moon, Leonardo Neves, Vitor Carvalho |
NAACL-HLT | 1 |
| 2017 | Completely Heterogeneous Transfer Learning with Attention - What And What Not To TransferabstractWe study a transfer learning framework where source and target datasets are heterogeneous in both feature and label spaces. Specifically, we do not assume explicit relations between source and target tasks a priori, and thus it is crucial to determine what and what not to transfer from source knowledge. Towards this goal, we define a new heterogeneous transfer learning approach that (1) selects and attends to an optimized subset of source samples to transfer knowledge from, and (2) builds a unified transfer network that learns from both source and target knowledge. This method, termed "Attentional Heterogeneous Transfer", along with a newly proposed unsupervised transfer loss, improve upon the previous state-of-the-art approaches on extensive simulations as well as a challenging hetero-lingual text classification task. Seungwhan Moon, Jaime G. Carbonell |
IJCAI | 1 |
| 2016 | Metaphor Detection with Topic Transition, Emotion and Cognition in ContextabstractMetaphor is a common linguistic tool in communication, making its detection in discourse a crucial task for natural language understanding.One popular approach to this challenge is to capture semantic incohesion between a metaphor and the dominant topic of the surrounding text.While these methods are effective, they tend to overclassify target words as metaphorical when they deviate in meaning from its context.We present a new approach that (1) distinguishes literal and non-literal use of target words by examining sentence-level topic transitions and (2) captures the motivation of speakers to express emotions and abstract concepts metaphorically.Experiments on an online breast cancer discussion forum dataset demonstrate a significant improvement in metaphor detection over the state-of-theart.These experimental results also reveal a tendency toward metaphor usage in personal topics and certain emotional contexts. Hyeju Jang, Yohan Jo, Qinlan Shen, Michael Miller Yoder, Seungwhan Moon, Carolyn P. Rosé |
ACL (1) | 5 |
| 2016 | Proactive Transfer Learning for Heterogeneous Feature and Label Spaces
Seungwhan Moon, Jaime G. Carbonell |
ECML/PKDD (2) | 1 |
| 2015 | Ranking and retrieval of image sequences from multiple paragraph queriesabstractWe propose a method to rank and retrieve image sequences from a natural language text query, consisting of multiple sentences or paragraphs. One of the method's key applications is to visualize visitors' text-only reviews on TRIPADVISOR or YELP, by automatically retrieving the most illustrative image sequences. While most previous work has dealt with the relations between a natural language sentence and an image or a video, our work extends to the relations between paragraphs and image sequences. Our approach leverages the vast user-generated resource of blog posts and photo streams on the Web. We use blog posts as text-image parallel training data that co-locate informative text with representative images that are carefully selected by users. We exploit large-scale photo streams to augment the image samples for retrieval. We design a latent structural SVM framework to learn the semantic relevance relations between text and image sequences. We present both quantitative and qualitative results on the newly created DISNEYLAND dataset. Gunhee Kim, Seungwhan Moon, Leonid Sigal |
CVPR | 2 |
| 2015 | Joint photo stream and blog post summarization and explorationabstractWe propose an approach that utilizes large collections of photo streams and blog posts, two of the most prevalent sources of data on the Web, for joint story-based summarization and exploration. Blogs consist of sequences of images and associated text; they portray events and experiences with concise sentences and representative images. We leverage blogs to help achieve story-based semantic summarization of collections of photo streams. In the opposite direction, blog posts can be enhanced with sets of photo streams by showing interpolations between consecutive images in the blogs. We formulate the problem of joint alignment from blogs to photo streams and photo stream summarization in a unified latent ranking SVM framework. We alternate between solving the two coupled latent SVM problems, by first fixing the summarization and solving for the alignment from blog images to photo streams and vice versa. On a newly collected large-scale Disneyland dataset of 10K blogs (120K associated images) and 6K photo streams (540K images), we demonstrate that blog posts and photo streams are mutually beneficial for summarization, exploration, semantic knowledge transfer, and photo interpolation. Gunhee Kim, Seungwhan Moon, Leonid Sigal |
CVPR | 2 |
| 2015 | Metaphor Detection in DiscourseabstractUnderstanding contextual information is key to detecting metaphors in discourse.Most current work aims at detecting metaphors given a single sentence, thus focusing mostly on local contextual cues within a short text.In this paper, we present a novel approach that explicitly leverages global context of a discourse to detect metaphors.In addition, we show that syntactic information such as dependency structures can help better describe local contextual information, thus improving detection results when combined.We apply our methods on a newly annotated online discussion forum, and show that our approach outperforms the state-of-the-art baselines in previous literature. Hyeju Jang, Seungwhan Moon, Yohan Jo, Carolyn P. Rosé |
SIGDIAL Conference | 2 |
| 2014 | Proactive learning with multiple class-sensitive labelersabstractProactive learning extends active learning by considering multiple labelers with different accuracies and costs, thus optimizing labeler selection as well as instance selection. In this paper, we propose a novel method to estimate labeler accuracy per class and to select labelers based on both cost and estimated accuracy, combined with an ensemble approach called multi-class information density (MCID) as a selection criterion. Our approach relaxes the common assumption found in past work that labeler accuracy is independent of class for multi-class learning, and by estimating the class-conditional accuracy better assigns instances to labelers. Results on several datasets with both real and simulated experts strongly demonstrate the efficacy of these methods. Seungwhan Moon, Jaime G. Carbonell |
DSAA | 1 |
| 2012 | Design and implementation of the note-taking style haptic voice recognition for mobile devicesabstractThis research proposes the "note-taking style" Haptic Voice Recognition (HVR) technology which incorporates speech and touch sensory inputs in a note-like form to enhance the performance of speech recognition. A note is taken from a user via two different haptic input methods - handwriting and a keyboard. A note consists of some of the keywords in the given utterance, either partially spelled or fully spelled. In order to facilitate fast input, the interface allows a shorthand writing system such as Gregg Shorthand. Using this haptic note sequence as an additional knowledge source, the algorithm re-ranks the n-best list generated by a speech engine. The simulation and experimental results show that the proposed HVR method improves the Word Error Rate (WER) and Keyword Error Rate (KER) performance in comparison to an Automatic Speech Recognition (ASR) system. Although it generates an inevitable increase in speech duration due to disfluency and occasional mistakes in haptic input, the compensation is shown to be less than conventional HVR methods. As such, this new note-taking style HVR interaction has the potential to be both natural and effective in increasing the recognition performance by choosing the most likely utterance among multiple hypotheses. This paper discusses the algorithm for the proposed system, the results from the simulation and the experiments, and the possible applications of this new technology such as aiding spoken document retrieval with haptic notes. Seungwhan Moon, Khe Chai Sim |
ICMI | 1 |