VLDB 2026 Research / reviewers in the wild / expert
Mo Yu
dblp:32/7445
· DBLP profile ↗
85ranked-venue papers
15as first author
36since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 14 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ComoRAG: A Cognitive-Inspired Memory-Organized RAG for Stateful Long Narrative ReasoningabstractNarrative comprehension on long stories and novels has been a challenging domain attributed to their intricate plotlines and entangled, often evolving relations among characters and entities. Given the LLM's diminished reasoning over extended context and its high computational cost, retrieval-based approaches remain a pivotal role in practice. However, traditional RAG methods could fall short due to their stateless, single-step retrieval process, which often overlooks the dynamic nature of capturing interconnected relations within long-range context. In this work, we propose ComoRAG, holding the principle that narrative reasoning is not a one-shot process, but a dynamic, evolving interplay between new evidence acquisition and past knowledge consolidation, analogous to human cognition on reasoning with memory-related signals in the brain. Specifically, when encountering a reasoning impasse, ComoRAG undergoes iterative reasoning cycles while interacting with a dynamic memory workspace. In each cycle, it generates probing queries to devise new exploratory paths, then integrates the retrieved evidence of new aspects into a global memory pool, thereby supporting the emergence of a coherent context for the query resolution. Across four challenging long-context narrative benchmarks (200K+ tokens), ComoRAG outperforms strong RAG baselines with consistent relative gains up to 11% compared to the strongest baseline. Further analysis reveals that ComoRAG is particularly advantageous for complex queries requiring global comprehension, offering a principled, cognitively motivated paradigm for retrieval-based stateful reasoning. Juyuan Wang, Rongchen Zhao, Mo Yu, Jie Zhou 0016, Liyan Xu |
AAAI | 5 |
| 2026 | DHCom-NAS: Dynamic Heterogeneous Community Detection via Neural Architecture Search
Mo Yu, Zhengyang Wu 0001, Chaobo He |
ICIC (4) | 1 |
| 2025 | Coherency Improved Explainable Recommendation via Large Language ModelabstractExplainable recommender systems are designed to elucidate the explanation behind each recommendation, enabling users to comprehend the underlying logic. Previous works perform rating prediction and explanation generation in a multi-task manner. However, these works suffer from incoherence between predicted ratings and explanations. To address the issue, we propose a novel framework that employs a large language model (LLM) to generate a rating, transforms it into a rating vector, and finally generates an explanation based on the rating vector and user-item information. Moreover, we propose utilizing publicly available LLMs and pre-trained sentiment analysis models to automatically evaluate the coherence without human annotations. Extensive experimental results on three datasets of explainable recommendation show that the proposed framework is effective, outperforming state-of-the-art baselines with improvements of 7.3% in explainability and 4.4% in text quality. Ruixin Ding, Weihai Lu, Jun Wang 0006, Mo Yu, Wei Zhang 0056 |
AAAI | 5 |
| 2025 | The Essence of Contextual Understanding in Theory of Mind: A Study on Question Answering with Story CharactersabstractTheory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others’ thoughts by integrating causal cues and indirect clues from broad contextual information, often derived from past interactions. In other words, human ToM heavily relies on the understanding about the backgrounds and life stories of others. Unfortunately, this aspect is largely overlooked in existing benchmarks for evaluating machines’ ToM capabilities, due to their usage of short narratives without global context, especially personal background of characters. In this paper, we verify the importance of comprehensive contextual understanding about personal backgrounds in ToM and assess the performance of LLMs in such complex scenarios. To achieve this, we introduce CharToM-QA benchmark, comprising 1,035 ToM questions based on characters from classic novels. Our human study reveals a significant disparity in performance: the same group of educated participants performs dramatically better when they have read the novels compared to when they have not. In parallel, our experiments on state-of-the-art LLMs, including the very recent o1 and DeepSeek-R1 models, show that LLMs still perform notably worse than humans, despite that they have seen these stories during pre-training. This highlights the limitations of current LLMs in capturing the nuanced contextual information required for ToM reasoning. Chulun Zhou, Qiujing Wang, Mo Yu, Xiaoqian Yue, Shunchi Zhang, Jie Zhou 0016, Wai Lam |
ACL (1) | 3 |
| 2025 | Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC TaskabstractJunjie Wu, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Junjie Wu 0007, Mo Yu, Lemao Liu, Dit-Yan Yeung, Jie Zhou 0016 |
NAACL (Long Papers) | 2 |
| 2025 | The Stochastic Parrot on LLM's Shoulder: A Summative Assessment of Physical Concept UnderstandingabstractMo Yu, Lemao Liu, Junjie Wu, Tsz Ting Chung, Shunchi Zhang, Jiangnan Li, Dit-Yan Yeung, Jie Zhou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Mo Yu, Lemao Liu, Junjie Wu 0007, Tsz Ting Chung, Shunchi Zhang, Dit-Yan Yeung, Jie Zhou 0016 |
NAACL (Long Papers) | 1 |
| 2025 | Rebalancing Discriminative Responses for Knowledge TracingabstractKnowledge Tracing (KT) is a crucial task in computer-aided education and intelligent tutoring systems, predicting students’ performance on new questions from their responses to prior ones. An accurate KT model can capture a student’s mastery level of different knowledge topics, as reflected in their predicted performance on different questions. This helps improve the learning efficiency by suggesting appropriate new questions that complement students’ knowledge states. However, current KT models have significant drawbacks that they neglect the imbalanced discrimination of historical responses. A significant proportion of question responses provide limited information for discerning students’ knowledge mastery, such as those that demonstrate uniform performance across different students. Optimizing the prediction of these cases may increase overall KT accuracy, but also negatively impact the model’s ability to trace personalized knowledge states, especially causing a deceptive surge of performance. Towards this end, we propose a framework to reweight the contribution of different responses based on their discrimination in training. Additionally, we introduce an adaptive predictive score fusion technique to maintain accuracy on less discriminative responses, achieving proper balance between student knowledge mastery and question difficulty. Experimental results demonstrate that our framework enhances the performance of three mainstream KT methods on three widely used datasets. Jiajun Cui, Hong Qian, Chanjin Zheng, Lu Wang 0029, Mo Yu, Wei Zhang 0056 |
ACM Trans. Inf. Syst. | 5 |
| 2024 | Fine-Grained Modeling of Narrative Context: A Coherence Perspective via Retrospective QuestionsabstractThis work introduces an original and practical paradigm for narrative comprehension, stemming from the characteristics that individual passages within narratives tend to be more cohesively related than isolated.Complementary to the common end-to-end paradigm, we propose a fine-grained modeling of narrative context, by formulating a graph dubbed NARCO, which explicitly depicts task-agnostic coherence dependencies that are ready to be consumed by various downstream tasks.In particular, edges in NARCO encompass free-form retrospective questions between context snippets, inspired by human cognitive perception that constantly reinstates relevant events from prior context.Importantly, our graph formalism is practically instantiated by LLMs without human annotations, through our designed two-stage prompting scheme.To examine the graph properties and its utility, we conduct three studies in narratives, each from a unique angle: edge relation efficacy, local context enrichment, and broader application in QA.All tasks could benefit from the explicit coherence captured by NARCO. Liyan Xu, Mo Yu, Jie Zhou 0016 |
ACL (1) | 3 |
| 2024 | Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating additional information from retrieval.However, studies have shown that LLMs still face challenges in effectively using the retrieved information, even ignoring it or being misled by it.The key reason is that the training of LLMs does not clearly make LLMs learn how to utilize input retrieved texts with varied quality.In this paper, we propose a novel perspective that considers the role of LLMs in RAG as "Information Refiner", which means that regardless of correctness, completeness, or usefulness of retrieved texts, LLMs can consistently integrate knowledge within the retrieved texts and model parameters to generate the texts that are more concise, accurate, and complete than the retrieved texts.To this end, we propose an information refinement training method named INFO-RAG that optimizes LLMs for RAG in an unsupervised manner.INFO-RAG is low-cost and general across various tasks.Extensive experiments on zero-shot prediction of 11 datasets in diverse tasks including Question Answering, Slot-Filling, Language Modeling, Dialogue, and Code Generation show that INFO-RAG improves the performance of LLaMA2 by an average of 9.39% relative points.INFO-RAG also shows advantages in in-context learning and robustness of RAG.LLMs Liang Pang 0001, Mo Yu, Fandong Meng, Huawei Shen, Xueqi Cheng 0001, Jie Zhou 0016 |
ACL (1) | 3 |
| 2024 | Rethinking the Evaluation of In-Context Learning for LLMsabstractIn-context learning (ICL) has demonstrated excellent performance across various downstream NLP tasks, especially when synergized with powerful large language models (LLMs).Existing studies evaluate ICL methods primarily based on downstream task performance.This evaluation protocol overlooks the significant cost associated with the demonstration configuration process, i.e., tuning the demonstration as the ICL prompt.However, in this work, we point out that the evaluation protocol leads to unfair comparisons and potentially biased evaluation, because we surprisingly find the correlation between the configuration costs and task performance.Then we call for a twodimensional evaluation paradigm that considers both of these aspects, facilitating a fairer comparison.Finally, based on our empirical finding that the optimized demonstration on one language model generalizes across language models of different sizes, we introduce a simple yet efficient strategy that can be applied to any ICL method as a plugin, yielding a better trade-off between the two dimensions according to the proposed evaluation paradigm. Guoxin Yu, Lemao Liu, Mo Yu, Xiang Ao 0001 |
EMNLP | 3 |
| 2024 | Few-Shot Character Understanding in Movies as an Assessment to Meta-Learning of Theory-of-MindabstractWhen reading a story, humans can quickly understand new fictional characters with a few observations, mainly by drawing analogies to fictional and real people they already know. This reflects the few-shot and meta-learning essence of humans' inference of characters' mental states, *i.e.*, theory-of-mind (ToM), which is largely ignored in existing research. We fill this gap with a novel NLP dataset in a realistic narrative understanding scenario, ToM-in-AMC. Our dataset consists of $\sim$1,000 parsed movie scripts, each corresponding to a few-shot character understanding task that requires models to mimic humans' ability of fast digesting characters with a few starting scenes in a new movie. We further propose a novel ToM prompting approach designed to explicitly assess the influence of multiple ToM dimensions. It surpasses existing baseline models, underscoring the significance of modeling multiple ToM dimensions for our task. Our extensive human study verifies that humans are capable of solving our problem by inferring characters' mental states based on their previously seen movies. In comparison, all the AI systems lag $>20\%$ behind humans, highlighting a notable limitation in existing approaches' ToM capabilities. Code and data are available at https://github.com/ShunchiZhang/ToM-in-AMC Mo Yu, Qiujing Wang, Shunchi Zhang, Yisi Sang, Kangsheng Pu, Zekai Wei, Liyan Xu, Jie Zhou 0016 |
ICML | 1 |
| 2024 | On Large Language Models' Hallucination with Regard to Known FactsabstractChe Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Yang Cheng, Fandong Meng, Mo Yu, Bowen Zhou, Jie Zhou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Che Jiang, Biqing Qi, Xiangyu Hong, Dayuan Fu, Fandong Meng, Mo Yu, Bowen Zhou 0002, Jie Zhou 0016 |
NAACL-HLT | 7 |
| 2024 | An Energy-based Model for Word-level AutoCompletion in Computer-aided TranslationabstractAbstract Word-level AutoCompletion (WLAC) is a rewarding yet challenging task in Computer-aided Translation. Existing work addresses this task through a classification model based on a neural network that maps the hidden vector of the input context into its corresponding label (i.e., the candidate target word is treated as a label). Since the context hidden vector itself does not take the label into account and it is projected to the label through a linear classifier, the model cannot sufficiently leverage valuable information from the source sentence as verified in our experiments, which eventually hinders its overall performance. To alleviate this issue, this work proposes an energy-based model for WLAC, which enables the context hidden vector to capture crucial information from the source sentence. Unfortunately, training and inference suffer from efficiency and effectiveness challenges, therefore we employ three simple yet effective strategies to put our model into practice. Experiments on four standard benchmarks demonstrate that our reranking-based approach achieves substantial improvements (about 6.07%) over the previous state-of-the-art model. Further analyses show that each strategy of our approach contributes to the final performance.1 Cheng Yang 0007, Guoping Huang, Mo Yu, Zhirui Zhang, Siheng Li, Shuming Shi 0001, Yujiu Yang 0001, Lemao Liu |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Personality Understanding of Fictional Characters during Book ReadingabstractMo Yu, Jiangnan Li, Shunyu Yao, Wenjie Pang, Xiaochen Zhou, Zhou Xiao, Fandong Meng, Jie Zhou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Mo Yu, Wenjie Pang, Xiaochen Zhou, Xiao Zhou 0004, Fandong Meng, Jie Zhou 0016 |
ACL (1) | 1 |
| 2023 | Explicit Planning Helps Language Models in Logical ReasoningabstractLanguage models have been shown to perform remarkably well on a wide range of natural language processing tasks.In this paper, we propose LEAP, a novel system that uses language models to perform multi-step logical reasoning and incorporates explicit planning into the inference procedure.Explicit planning enables the system to make more informed reasoning decisions at each step by looking ahead into their future effects.Moreover, we propose a training strategy that safeguards the planning process from being led astray by spurious features.Our full system significantly outperforms other competing methods on multiple standard datasets.When using small T5 models as its core selection and deduction components, our system performs competitively compared to GPT-3 despite having only about 1B parameters (i.e., 175 times smaller than GPT-3).When using GPT-3.5, it significantly outperforms chain-of-thought prompting on the challenging PrOntoQA dataset.We have conducted extensive empirical studies to demonstrate that explicit planning plays a crucial role in the system's performance. Hongyu Zhao 0006, Kangrui Wang, Mo Yu, Hongyuan Mei |
EMNLP | 3 |
| 2023 | Towards Coherent Image Inpainting Using Denoising Diffusion Implicit ModelsabstractImage inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. Jiabao Ji, Yang Zhang 0001, Mo Yu, Tommi S. Jaakkola, Shiyu Chang |
ICML | 4 |
| 2023 | Generalizable Reinforcement Learning-Based Coarsening Model for Resource Allocation over Large and Diverse Stream Processing GraphsabstractResource allocation for stream processing graphs on computing devices is critical to the performance of stream processing. Efficient allocations need to balance workload distribution and minimize communication simultaneously and globally. Since this problem is known to be NP-complete, recent machine learning solutions were proposed based on an encoder-decoder framework, which predicts the device assignment of computing nodes sequentially as an approximation. However, for large graphs, these solutions suffer from the deficiency in handling long-distance dependency and global information, resulting in suboptimal predictions. This work proposes a new paradigm to deal with this challenge, which first coarsens the graph and conducts assignments on the smaller graph with existing graph partitioning methods. Unlike existing graph coarsening works, we leverage the theoretical insights in this resource allocation problem, formulate the coarsening of stream graphs as edge-collapsing predictions, and propose an edge-aware coarsening model. Extensive experiments on various datasets show that our framework significantly improves over existing learning-based and heuristic-based baselines with up to 56% relative improvement on large graphs. Lanshun Nie, Yuqi Qiu, Mo Yu, Jing Li 0025 |
IPDPS | 4 |
| 2022 | Fantastic Questions and Where to Find Them: FairytaleQA - An Authentic Dataset for Narrative ComprehensionabstractYing Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Bingsheng Yao, Tongshuang Wu, Zheng Zhang, Toby Li, Nora Bradford, Branda Sun, Tran Hoang, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu, Mark Warschauer. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Dakuo Wang, Mo Yu, Daniel Ritchie 0002, Bingsheng Yao, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Yisi Sang, Yufang Hou 0001, Xiaojuan Ma, Diyi Yang, Nanyun Peng 0001, Zhou Yu 0005, Mark Warschauer |
ACL (1) | 3 |
| 2022 | It is AI's Turn to Ask Humans a Question: Question-Answer Pair Generation for Children's Story BooksabstractExisting question answering (QA) techniques are created mainly to answer questions asked by humans.But in educational applications, teachers often need to decide what questions they should ask, in order to help students to improve their narrative understanding capabilities.We design an automated question-answer generation (QAG) system for this education scenario: given a story book at the kindergarten to eighth-grade level as input, our system can automatically generate QA pairs that are capable of testing a variety of dimensions of a student's comprehension skills.Our proposed QAG model architecture is demonstrated using a new expert-annotated FairytaleQA dataset, which has 278 child-friendly storybooks with 10,580 QA pairs.Automatic and human evaluations show that our model outperforms stateof-the-art QAG baseline systems.On top of our QAG system, we also start to build an interactive story-telling application for the future real-world deployment in this educational scenario. Bingsheng Yao, Dakuo Wang, Sherry Tongshuang Wu, Zheng Zhang 0043, Toby Jia-Jun Li, Mo Yu |
ACL (1) | 6 |
| 2022 | Educational Question Generation of Children Storybooks via Question Type Distribution Learning and Event-centric SummarizationabstractGenerating educational questions of fairytales or storybooks is vital for improving children's literacy ability.However, it is challenging to generate questions that capture the interesting aspects of a fairytale story with educational meaningfulness.In this paper, we propose a novel question generation method that first learns the question type distribution of an input story paragraph, and then summarizes salient events which can be used to generate high-cognitive-demand questions.To train the event-centric summarizer, we finetune a pre-trained transformer-based sequenceto-sequence model using silver samples composed by educational question-answer pairs.On a newly proposed educational questionanswering dataset FairytaleQA, we show good performance of our method on both automatic and human evaluation metrics.Our work indicates the necessity of decomposing question type distribution learning and event-centric summary generation for educational question generation. Zhenjie Zhao, Yufang Hou 0001, Dakuo Wang, Mo Yu, Chengzhong Liu, Xiaojuan Ma |
ACL (1) | 4 |
| 2022 | StoryBuddy: A Human-AI Collaborative Chatbot for Parent-Child Interactive Storytelling with Flexible Parental InvolvementabstractDespite its benefits for children’s skill development and parent-child bonding, many parents do not often engage in interactive storytelling by having story-related dialogues with their child due to limited availability or challenges in coming up with appropriate questions. While recent advances made AI generation of questions from stories possible, the fully-automated approach excludes parent involvement, disregards educational goals, and underoptimizes for child engagement. Informed by need-finding interviews and participatory design (PD) results, we developed StoryBuddy, an AI-enabled system for parents to create interactive storytelling experiences. StoryBuddy’s design highlighted the need for accommodating dynamic user needs between the desire for parent involvement and parent-child bonding and the goal of minimizing parent intervention when busy. The PD revealed varied assessment and educational goals of parents, which StoryBuddy addressed by supporting configuring question types and tracking child progress. A user study validated StoryBuddy’s usability and suggested design insights for future parent-AI collaboration systems. Zheng Zhang 0043, Bingsheng Yao, Daniel Ritchie 0002, Sherry Tongshuang Wu, Mo Yu, Dakuo Wang, Toby Jia-Jun Li |
CHI | 7 |
| 2022 | Linking Emergent and Natural Languages via Corpus Transfer
Shunyu Yao 0006, Mo Yu, Yang Zhang 0001, Karthik Narasimhan, Josh Tenenbaum, Chuang Gan 0001 |
ICLR | 2 |
| 2022 | A Survey of Machine Narrative Reading Comprehension AssessmentsabstractAs the body of research on machine narrative comprehension grows, there is a critical need for consideration of performance assessment strategies as well as the depth and scope of different benchmark tasks. Based on narrative theories, reading comprehension theories, as well as existing machine narrative reading comprehension tasks and datasets, we propose a typology that captures the main similarities and differences among assessment tasks; and discuss the implications of our typology for new task design and the challenges of narrative reading comprehension. Yisi Sang, Xiangyang Mou, Jing Li 0025, Jeffrey M. Stanton, Mo Yu |
IJCAI | 5 |
| 2022 | Learning as Conversation: Dialogue Systems Reinforced for Information AcquisitionabstractPengshan Cai, Hui Wan, Fei Liu, Mo Yu, Hong Yu, Sachindra Joshi. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Pengshan Cai, Hui Wan 0001, Fei Liu 0004, Mo Yu, Hong Yu 0001, Sachindra Joshi |
NAACL-HLT | 4 |
| 2022 | On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models?abstractNouha Dziri, Sivan Milton, Mo Yu, Osmar Zaiane, Siva Reddy. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nouha Dziri, Sivan Milton, Mo Yu, Osmar R. Zaïane, Siva Reddy |
NAACL-HLT | 3 |
| 2022 | TVShowGuess: Character Comprehension in Stories as Speaker GuessingabstractYisi Sang, Xiangyang Mou, Mo Yu, Shunyu Yao, Jing Li, Jeffrey Stanton. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yisi Sang, Xiangyang Mou, Mo Yu, Jing Li 0025, Jeffrey M. Stanton |
NAACL-HLT | 3 |
| 2022 | Group Chat Ecology in Enterprise Instant Messaging: How Employees Collaborate Through Multi-User Chat Channels on SlackabstractDespite the long history of studying instant messaging usage, we know very little about how today's people participate in group chat channels and interact with others inside a real-world organization. In this short paper, we aim to update the existing knowledge on how group chat is used in the context of today's organizations. The knowledge is particularly important for the new norm of remote works under the COVID-19 pandemic. We have the privilege of collecting two valuable datasets: a total of 4,300 group chat channels in Slack from an R&D department in a multinational IT company; and a total of 117 groups' performance data. Through qualitative coding of 100 randomly sampled group channels from the 4,300 channels dataset, we identified and reported 9 categories such as Project channels, IT-Support channels, and Event channels. We further defined a feature metric with 21 meta-features (and their derived features) without looking at the message content to depict the group communication style for these group chat channels, with which we successfully trained a machine learning model that can automatically classify a given group channel into one of the 9 categories. In addition to the descriptive data analysis, we illustrated how these communication metrics can be used to analyze team performance. We cross-referenced 117 project teams and their team-based Slack channels and identified 57 teams that appeared in both datasets, then we built a regression model to reveal the relationship between these group communication styles and the project team performance. This work contributes an updated empirical understanding of human-human communication practices within the enterprise setting, and suggests design opportunities for the future of human-AI communication experience. Dakuo Wang, Haoyu Wang 0002, Mo Yu, Zahra Ashktorab |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2022 | FaithDial: A Faithful Benchmark for Information-Seeking DialogueabstractAbstract The goal of information-seeking dialogue is to respond to seeker queries with natural language utterances that are grounded on knowledge sources. However, dialogue systems often produce unsupported utterances, a phenomenon known as hallucination. To mitigate this behavior, we adopt a data-centric solution and create FaithDial, a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark. We observe that FaithDial is more faithful than WoW while also maintaining engaging conversations. We show that FaithDial can serve as training signal for: i) a hallucination critic, which discriminates whether an utterance is faithful or not, and boosts the performance by 12.8 F1 score on the BEGIN benchmark compared to existing datasets for dialogue coherence; ii) high-quality dialogue generation. We benchmark a series of state-of-the-art models and propose an auxiliary contrastive objective that achieves the highest level of faithfulness and abstractiveness based on several automated metrics. Further, we find that the benefits of FaithDial generalize to zero-shot transfer on other datasets, such as CMU-Dog and TopicalChat. Finally, human evaluation reveals that responses generated by models trained on FaithDial are perceived as more interpretable, cooperative, and engaging. Nouha Dziri, Ehsan Kamalloo, Sivan Milton, Osmar R. Zaïane, Mo Yu, Edoardo Maria Ponti, Siva Reddy |
Trans. Assoc. Comput. Linguistics | 5 |
| 2022 | Evidence Integration for Multi-Hop Reading Comprehension With Graph Neural NetworksabstractMulti-hop reading comprehension focuses on one type of factoid question, where a system needs to properly integrate multiple pieces of evidence to correctly answer a question. Previous work approximates global evidence with local coreference information, encoding coreference chains with DAG-styled GRU layers within a gated-attention reader. However, coreference is limited in providing information for rich inference. We introduce a new method for better connecting global evidence, which forms more complex graphs compared to DAGs. To perform evidence integration on our graphs, we investigate two recent graph neural networks, namely graph convolutional network (GCN) and graph recurrent network (GRN). Experiments on two standard datasets show that richer global information leads to better answers. Our approach shows highly competitive performances on these datasets without deep language models (such as ELMo). Linfeng Song, Zhiguo Wang 0006, Mo Yu, Yue Zhang 0004, Radu Florian, Daniel Gildea |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | Complementary Evidence Identification in Open-Domain Question AnsweringabstractThis paper proposes a new problem of complementary evidence identification for opendomain question answering (QA).The problem aims to efficiently find a small set of passages that covers full evidence from multiple aspects as to answer a complex question.To this end, we proposes a method that learns vector representations of passages and models the sufficiency and diversity within the selected set, in addition to the relevance between the question and passages.Our experiments demonstrate that our method considers the dependence within the supporting evidence and significantly improves the accuracy of complementary evidence selection in QA domain. Xiangyang Mou, Mo Yu, Shiyu Chang, Hui Su |
EACL | 2 |
| 2021 | Timeline Summarization based on Event Graph Compression via Time-Aware Optimal TransportabstractTimeline Summarization identifies major events from a news collection and describes them following temporal order, with key dates tagged.Previous methods generally generate summaries separately for each date after they determine the key dates of events.These methods overlook the events' intra-structures (arguments) and inter-structures (event-event connections).Following a different route, we propose to represent the news articles as an event-graph, thus the summarization task becomes compressing the whole graph to its salient sub-graph.The key hypothesis is that the events connected through shared arguments and temporal order depict the skeleton of a timeline, containing events that are semantically related, structurally salient, and temporally coherent in the global event graph.A time-aware optimal transport distance is then introduced for learning the compression model in an unsupervised manner.We show that our approach significantly improves the state of the art on three real-world datasets, including two public standard benchmarks and our newly collected Timeline 100 dataset. 1 Manling Li, Tengfei Ma 0001, Mo Yu, Lingfei Wu 0001, Tian Gao 0007, Heng Ji 0001, Kathy McKeown |
EMNLP (1) | 3 |
| 2021 | Interpretable Visual Reasoning via Induced Symbolic SpaceabstractWe study the problem of concept induction in visual reasoning, i.e., identifying concepts and their hierarchical relationships from question-answer pairs associated with images; and achieve an interpretable model via working on the induced symbolic concept space. To this end, we first design a new framework named object-centric compositional attention model (OCCAM) to perform the visual reasoning task with object-level visual features. Then, we come up with a method to induce concepts of objects and relations using clues from the attention patterns between objects’ visual features and question words. Finally, we achieve a higher level of interpretability by imposing OCCAM on the objects represented in the induced symbolic concept space. Experiments on the CLEVR and GQA datasets demonstrate: 1) our OCCAM achieves a new state of the art without human-annotated functional programs; 2) our induced concepts are both accurate and sufficient as OCCAM achieves an on-par performance on objects represented either in visual features or in the induced symbolic concept space. Zhonghao Wang 0001, Kai Wang 0058, Mo Yu, Jinjun Xiong, Wen-Mei W. Hwu, Mark Hasegawa-Johnson, Humphrey Shi |
ICCV | 3 |
| 2021 | Multilingual BERT Post-Pretraining AlignmentabstractLin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, Mo Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Lin Pan 0003, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, Mo Yu |
NAACL-HLT | 6 |
| 2021 | Understanding Interlocking Dynamics of Cooperative RationalizationabstractSelective rationalization explains the prediction of complex neural networks by finding a small subset of the input that is sufficient to predict the neural model output. The selection mechanism is commonly integrated into the model itself by specifying a two-component cascaded system consisting of a rationale generator, which makes a binary selection of the input features (which is the rationale), and a predictor, which predicts the output based only on the selected features. The components are trained jointly to optimize prediction performance. In this paper, we reveal a major problem with such cooperative rationalization paradigm --- model interlocking. Inter-locking arises when the predictor overfits to the features selected by the generator thus reinforcing the generator's selection even if the selected rationales are sub-optimal. The fundamental cause of the interlocking problem is that the rationalization objective to be minimized is concave with respect to the generator’s selection policy. We propose a new rationalization framework, called A2R, which introduces a third component into the architecture, a predictor driven by soft attention as opposed to selection. The generator now realizes both soft and hard attention over the features and these are fed into the two different predictors. While the generator still seeks to support the original predictor performance, it also minimizes a gap between the two predictors. As we will show theoretically, since the attention-based predictor exhibits a better convexity property, A2R can overcome the concavity barrier. Our experiments on two synthetic benchmarks and two real datasets demonstrate that A2R can significantly alleviate the interlock problem and find explanations that better align with human judgments. Mo Yu, Yang Zhang 0001, Shiyu Chang, Tommi S. Jaakkola |
NeurIPS | 1 |
| 2021 | CASS: Towards Building a Social-Support Chatbot for Online Health CommunityabstractChatbots systems, despite their popularity in today's HCI and CSCW research, fall short for one of the two reasons: 1) many of the systems use a rule-based dialog flow, thus they can only respond to a limited number of pre-defined inputs with pre-scripted responses; or 2) they are designed with a focus on single-user scenarios, thus it is unclear how these systems may affect other users or the community. In this paper, we develop a generalizable chatbot architecture (CASS) to provide social support for community members in an online health community. The CASS architecture is based on advanced neural network algorithms, thus it can handle new inputs from users and generate a variety of responses to them. CASS is also generalizable as it can be easily migrate to other online communities. With a follow-up field experiment, CASS is proven useful in supporting individual members who seek emotional support. Our work also contributes to fill the research gap on how a chatbot may influence the whole community's engagement. Liuping Wang, Dakuo Wang, Feng Tian 0001, Zhenhui Peng, Xiangmin Fan, Zhan Zhang 0008, Mo Yu, Xiaojuan Ma, Hongan Wang |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2021 | Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive StudyabstractAbstract Recent advancements in open-domain question answering (ODQA), that is, finding answers from large open-domain corpus like Wikipedia, have led to human-level performance on many datasets. However, progress in QA over book stories (Book QA) lags despite its similar task formulation to ODQA. This work provides a comprehensive and quantitative analysis about the difficulty of Book QA: (1) We benchmark the research on the NarrativeQA dataset with extensive experiments with cutting-edge ODQA techniques. This quantifies the challenges Book QA poses, as well as advances the published state-of-the-art with a ∼7% absolute improvement on ROUGE-L. (2) We further analyze the detailed challenges in Book QA through human studies.1 Our findings indicate that the event-centric questions dominate this task, which exemplifies the inability of existing QA models to handle event-oriented scenarios. Xiangyang Mou, Chenghao Yang 0001, Mo Yu, Bingsheng Yao, Saloni Potdar, Hui Su |
Trans. Assoc. Comput. Linguistics | 3 |
| 2020 | Generalizable Resource Allocation in Stream Processing via Deep Reinforcement LearningabstractThis paper considers the problem of resource allocation in stream processing, where continuous data flows must be processed in real time in a large distributed system. To maximize system throughput, the resource allocation strategy that partitions the computation tasks of a stream processing graph onto computing devices must simultaneously balance workload distribution and minimize communication. Since this problem of graph partitioning is known to be NP-complete yet crucial to practical streaming systems, many heuristic-based algorithms have been developed to find reasonably good solutions. In this paper, we present a graph-aware encoder-decoder framework to learn a generalizable resource allocation strategy that can properly distribute computation tasks of stream processing graphs unobserved from training data. We, for the first time, propose to leverage graph embedding to learn the structural information of the stream processing graphs. Jointly trained with the graph-aware decoder using deep reinforcement learning, our approach can effectively find optimized solutions for unseen graphs. Our experiments show that the proposed model outperforms both METIS, a state-of-the-art graph partitioning algorithm, and an LSTM-based encoder-decoder model, in about 70% of the test cases. Xiang Ni, Jing Li 0025, Mo Yu, Kun-Lung Wu |
AAAI | 3 |
| 2020 | Differential Treatment for Stuff and Things: A Simple Unsupervised Domain Adaptation Method for Semantic SegmentationabstractWe consider the problem of unsupervised domain adaptation for semantic segmentation by easing the domain shift between the source domain (synthetic data) and the target domain (real data) in this work. State-of-the-art approaches prove that performing semantic-level alignment is helpful in tackling the domain shift issue. Based on the observation that stuff categories usually share similar appearances across images of different domains while things (i.e. object instances) have much larger differences, we propose to improve the semantic-level alignment with different strategies for stuff regions and for things: 1) for the stuff categories, we generate feature representation for each class and conduct the alignment operation from the target domain to the source domain; 2) for the thing categories, we generate feature representation for each individual instance and encourage the instance in the target domain to align with the most similar one in the source domain. In this way, the individual differences within thing categories will also be considered to alleviate over-alignment. In addition to our proposed method, we further reveal the reason why the current adversarial loss is often unstable in minimizing the distribution discrepancy and show that our method can help ease this issue by minimizing the most similar stuff and instance features between the source and the target domains. We conduct extensive experiments in two unsupervised domain adaptation tasks, i.e. GTA5 → Cityscapes and SYNTHIA → Cityscapes, and achieve the new state-of-the-art segmentation accuracy. Zhonghao Wang 0001, Mo Yu, Yunchao Wei, Rogério Feris, Jinjun Xiong, Wen-Mei W. Hwu, Thomas S. Huang, Humphrey Shi |
CVPR | 2 |
| 2020 | Interactive Fiction Game Playing as Multi-Paragraph Reading Comprehension with Reinforcement LearningabstractInteractive Fiction (IF) games with real humanwritten natural language texts provide a new natural evaluation for language understanding techniques.In contrast to previous text games with mostly synthetic texts, IF games pose language understanding challenges on the humanwritten textual descriptions of diverse and sophisticated game worlds and language generation challenges on the action command generation from less restricted combinatorial space.We take a novel perspective of IF game solving and re-formulate it as Multi-Passage Reading Comprehension (MPRC) tasks.Our approaches utilize the context-query attention mechanisms and the structured prediction in MPRC to efficiently generate and evaluate action outputs and apply an object-centric historical observation retrieval strategy to mitigate the partial observability of the textual observations.Extensive experiments on the recent IF benchmark (Jericho) demonstrate clear advantages of our approaches achieving high winning rates and low data requirements compared to all previous approaches. 1 Mo Yu, Yupeng Gao, Chuang Gan 0001, Murray Campbell, Shiyu Chang |
EMNLP (1) | 2 |
| 2020 | Invariant RationalizationabstractSelective rationalization improves neural network interpretability by identifying a small subset of input features {—} the rationale {—} that best explains or supports the prediction. A typical rationalization criterion, i.e. maximum mutual information (MMI), finds the rationale that maximizes the prediction performance based only on the rationale. However, MMI can be problematic because it picks up spurious correlations between the input features and the output. Instead, we introduce a game-theoretic invariant rationalization criterion where the rationales are constrained to enable the same predictor to be optimal across different environments. We show both theoretically and empirically that the proposed rationales can rule out spurious correlations and generalize better to different test scenarios. The resulting explanations also align better with human judgments. Our implementations are publicly available at https://github.com/code-terminator/invariant_rationalization. Shiyu Chang, Yang Zhang 0001, Mo Yu, Tommi S. Jaakkola |
ICML | 3 |
| 2020 | Leveraging Semantic Parsing for Relation Linking over Knowledge Bases
Nandana Mihindukulasooriya, Gaetano Rossiello, Pavan Kapanipathi, Ibrahim Abdelaziz, Srinivas Ravishankar, Mo Yu, Alfio Massimiliano Gliozzo, Salim Roukos, Alexander G. Gray |
ISWC (1) | 6 |
| 2019 | Hybrid Reinforcement Learning with Expert State SequencesabstractExisting imitation learning approaches often require that the complete demonstration data, including sequences of actions and states, are available. In this paper, we consider a more realistic and difficult scenario where a reinforcement learning agent only has access to the state sequences of an expert, while the expert actions are unobserved. We propose a novel tensor-based model to infer the unobserved actions of the expert state sequences. The policy of the agent is then optimized via a hybrid objective combining reinforcement learning and imitation learning. We evaluated our hybrid approach on an illustrative domain and Atari games. The empirical results show that (1) the agents are able to leverage state expert sequences to learn faster than pure reinforcement learning baselines, (2) our tensor-based action inference model is advantageous compared to standard deep neural networks in inferring expert actions, and (3) the hybrid policy optimization objective is robust against noise in expert state sequences. Shiyu Chang, Mo Yu, Gerald Tesauro, Murray Campbell |
AAAI | 3 |
| 2019 | Improving Natural Language Inference Using External Knowledge in the Science Questions DomainabstractNatural Language Inference (NLI) is fundamental to many Natural Language Processing (NLP) applications including semantic search and question answering. The NLI problem has gained significant attention due to the release of large scale, challenging datasets. Present approaches to the problem largely focus on learning-based methods that use only textual information in order to classify whether a given premise entails, contradicts, or is neutral with respect to a given hypothesis. Surprisingly, the use of methods based on structured knowledge – a central topic in artificial intelligence – has not received much attention vis-a-vis the NLI problem. While there are many open knowledge bases that contain various types of reasoning information, their use for NLI has not been well explored. To address this, we present a combination of techniques that harness external knowledge to improve performance on the NLI problem in the science questions domain. We present the results of applying our techniques on text, graph, and text-and-graph based models; and discuss the implications of using external knowledge to solve the NLI problem. Our model achieves close to state-of-the-art performance for NLI on the SciTail science questions dataset. Pavan Kapanipathi, Ryan Musa, Mo Yu, Kartik Talamadupula, Ibrahim Abdelaziz, Maria Chang 0001, Achille Fokoue, Bassem Makni, Nicholas Mattei, Michael Witbrock |
AAAI | 4 |
| 2019 | Extracting Multiple-Relations in One-Pass with Pre-Trained TransformersabstractThe state-of-the-art solutions for extracting multiple entity-relations from an input paragraph always require a multiple-pass encoding on the input.This paper proposes a new solution that can complete the multiple entityrelations extraction task with only one-pass encoding on the input corpus, and achieve a new state-of-the-art accuracy performance, as demonstrated in the ACE 2005 benchmark.Our solution is built on top of the pre-trained self-attentive models (Transformer).Since our method uses a single-pass to compute all relations at once, it scales to larger datasets easily; which makes it more usable in real-world applications.1 * Equal contributions from the corresponding authors: {wanghaoy,mingtan,yum}@us.ibm.com.Part of Haoyu Wang 0002, Mo Yu, Shiyu Chang, Dakuo Wang, Saloni Potdar |
ACL (1) | 3 |
| 2019 | Self-Supervised Learning for Contextualized Extractive SummarizationabstractExisting models for extractive summarization are usually trained from scratch with a crossentropy loss, which does not explicitly capture the global context at the document level.In this paper, we aim to improve this task by introducing three auxiliary pre-training tasks that learn to capture the document-level context in a self-supervised fashion.Experiments on the widely-used CNN/DM dataset validate the effectiveness of the proposed auxiliary tasks.Furthermore, we show that after pretraining, a clean model with simple building blocks is able to outperform previous state-ofthe-art that are carefully designed.1 Hong Wang 0023, Xin Wang 0061, Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang |
ACL (1) | 4 |
| 2019 | TWEETQA: A Social Media Focused Question Answering DatasetabstractWith social media becoming increasingly popular on which lots of news and real-time events are reported, developing automated question answering systems is critical to the effective-ness of many applications that rely on real-time knowledge. While previous datasets have concentrated on question answering (QA) for formal text like news and Wikipedia, we present the first large-scale dataset for QA over social media data. To ensure that the tweets we collected are useful, we only gather tweets used by journalists to write news articles. We then ask human annotators to write questions and answers upon these tweets. Unlike otherQA datasets like SQuAD in which the answers are extractive, we allow the answers to be abstractive. We show that two recently proposed neural models that perform well on formal texts are limited in their performance when applied to our dataset. In addition, even the fine-tuned BERT model is still lagging behind human performance with a large margin. Our results thus point to the need of improved QA systems targeting social media text. Wenhan Xiong, Jiawei Wu 0003, Hong Wang 0023, Vivek Kulkarni, Mo Yu, Shiyu Chang, William Yang Wang |
ACL (1) | 5 |
| 2019 | Improving Question Answering over Incomplete KBs with Knowledge-Aware ReaderabstractWe propose a new end-to-end question answering model, which learns to aggregate answer evidence from an incomplete knowledge base (KB) and a set of retrieved text snippets.Under the assumptions that the structured KB is easier to query and the acquired knowledge can help the understanding of unstructured text, our model first accumulates knowledge of entities from a question-related KB subgraph; then reformulates the question in the latent space and reads the texts with the accumulated entity knowledge at hand.The evidence from KB and texts are finally aggregated to predict answers.On the widely-used KBQA benchmark WebQSP, our model achieves consistent improvements across settings with different extents of KB incompleteness. 1 Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang |
ACL (1) | 2 |
| 2019 | Cross-lingual Knowledge Graph Alignment via Graph Matching Neural NetworkabstractPrevious cross-lingual knowledge graph (KG) alignment studies rely on entity embeddings derived only from monolingual KG structural information, which may fail at matching entities that have different facts in two KGs. In this paper, we introduce the topic entity graph, a local sub-graph of an entity, to represent entities with their contextual information in KG. From this view, the KB-alignment task can be formulated as a graph matching problem; and we further propose a graph-attention based solution, which first matches all entities in two topic entity graphs, and then jointly model the local matching information to derive a graph-level matching vector. Experiments show that our model outperforms previous state-of-the-art methods by a large margin. Kun Xu 0005, Liwei Wang 0009, Mo Yu, Yansong Feng 0002, Yan Song 0003, Zhiguo Wang 0006, Dong Yu 0001 |
ACL (1) | 3 |
| 2019 | Selection Bias Explorations and Debias Methods for Natural Language Sentence Matching DatasetsabstractNatural Language Sentence Matching (NLSM) has gained substantial attention from both academics and the industry, and rich public datasets contribute a lot to this process.However, biased datasets can also hurt the generalization performance of trained models and give untrustworthy evaluation results.For many NLSM datasets, the providers select some pairs of sentences into the datasets, and this sampling procedure can easily bring unintended pattern, i.e., selection bias.One example is the QuoraQP dataset, where some content-independent naïve features are unreasonably predictive.Such features are the reflection of the selection bias and termed as the "leakage features."In this paper, we investigate the problem of selection bias on six NLSM datasets and find that four out of them are significantly biased.We further propose a training and evaluation framework to alleviate the bias.Experimental results on QuoraQP suggest that the proposed framework can improve the generalization ability of trained models, and give more trustworthy evaluation results for real-world adoptions. Jian Liang 0002, Shiyu Chang, Mo Yu, Conghui Zhu, Tiejun Zhao |
ACL (1) | 6 |
| 2019 | Leveraging Dependency Forest for Neural Medical Relation ExtractionabstractLinfeng Song, Yue Zhang, Daniel Gildea, Mo Yu, Zhiguo Wang, Jinsong Su. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Linfeng Song, Yue Zhang 0004, Daniel Gildea, Mo Yu, Zhiguo Wang 0006, Jinsong Su |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Context-Aware Conversation Thread Detection in Multi-Party ChatabstractMing Tan, Dakuo Wang, Yupeng Gao, Haoyu Wang, Saloni Potdar, Xiaoxiao Guo, Shiyu Chang, Mo Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Dakuo Wang, Yupeng Gao, Haoyu Wang 0002, Saloni Potdar, Shiyu Chang, Mo Yu |
EMNLP/IJCNLP (1) | 8 |
| 2019 | Out-of-Domain Detection for Low-Resource Text Classification TasksabstractMing Tan, Yang Yu, Haoyu Wang, Dakuo Wang, Saloni Potdar, Shiyu Chang, Mo Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yang Yu 0029, Haoyu Wang 0002, Dakuo Wang, Saloni Potdar, Shiyu Chang, Mo Yu |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Rethinking Cooperative Rationalization: Introspective Extraction and Complement ControlabstractMo Yu, Shiyu Chang, Yang Zhang, Tommi Jaakkola. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Mo Yu, Shiyu Chang, Yang Zhang 0001, Tommi S. Jaakkola |
EMNLP/IJCNLP (1) | 1 |
| 2019 | DAG-GNN: DAG Structure Learning with Graph Neural NetworksabstractLearning a faithful directed acyclic graph (DAG) from samples of a joint distribution is a challenging combinatorial problem, owing to the intractable search space superexponential in the number of graph nodes. A recent breakthrough formulates the problem as a continuous optimization with a structural constraint that ensures acyclicity (Zheng et al., 2018). The authors apply the approach to the linear structural equation model (SEM) and the least-squares loss function that are statistically well justified but nevertheless limited. Motivated by the widespread success of deep learning that is capable of capturing complex nonlinear mappings, in this work we propose a deep generative model and apply a variant of the structural constraint to learn the DAG. At the heart of the generative model is a variational autoencoder parameterized by a novel graph neural network architecture, which we coin DAG-GNN. In addition to the richer capacity, an advantage of the proposed model is that it naturally handles discrete variables as well as vector-valued ones. We demonstrate that on synthetic data sets, the proposed method learns more accurate graphs for nonlinearly generated samples; and on benchmark data sets with discrete variables, the learned graphs are reasonably close to the global optima. The code is available at \url{https://github.com/fishmoon1234/DAG-GNN}. Yue Yu 0011, Jie Chen 0007, Tian Gao 0007, Mo Yu |
ICML | 4 |
| 2019 | A Game Theoretic Approach to Class-wise Selective RationalizationabstractSelection of input features such as relevant pieces of text has become a common technique of highlighting how complex neural predictors operate. The selection can be optimized post-hoc for trained models or incorporated directly into the method itself (self-explaining). However, an overall selection does not properly capture the multi-faceted nature of useful rationales such as pros and cons for decisions. To this end, we propose a new game theoretic approach to class-dependent rationalization, where the method is specifically trained to highlight evidence supporting alternative conclusions. Each class involves three players set up competitively to find evidence for factual and counterfactual scenarios. We show theoretically in a simplified scenario how the game drives the solution towards meaningful class-dependent rationales. We evaluate the method in single- and multi-aspect sentiment classification tasks and demonstrate that the proposed method is able to identify both factual (justifying the ground truth label) and counterfactual (countering the ground truth label) rationales consistent with human rationalization. The code for our method is publicly available. Shiyu Chang, Yang Zhang 0001, Mo Yu, Tommi S. Jaakkola |
NeurIPS | 3 |
| 2019 | A hybrid approach with optimization-based and metric-based meta-learner for few-shot learning
Yu Cheng 0001, Mo Yu, Tao Zhang 0006 |
Neurocomputing | 3 |
| 2018 | R3: Reinforced Ranker-Reader for Open-Domain Question AnsweringabstractIn recent years researchers have achieved considerable success applying neural network methods to question answering (QA). These approaches have achieved state of the art results in simplified closed-domain settings such as the SQuAD (Rajpurkar et al. 2016) dataset, which provides a pre-selected passage, from which the answer to a given question may be extracted. More recently, researchers have begun to tackle open-domain QA, in which the model is given a question and access to a large corpus (e.g., wikipedia) instead of a pre-selected passage (Chen et al. 2017a). This setting is more complex as it requires large-scale search for relevant passages by an information retrieval component, combined with a reading comprehension model that “reads” the passages to generate an answer to the question. Performance in this setting lags well behind closed-domain performance. In this paper, we present a novel open-domain QA system called Reinforced Ranker-Reader (R3), based on two algorithmic innovations. First, we propose a new pipeline for open-domain QA with a Ranker component, which learns to rank retrieved passages in terms of likelihood of extracting the ground-truth answer to a given question. Second, we propose a novel method that jointly trains the Ranker along with an answer-extraction Reader model, based on reinforcement learning. We report extensive experimental results showing that our method significantly improves on the state of the art for multiple open-domain QA datasets. Shuohang Wang, Mo Yu, Tim Klinger, Wei Zhang 0057, Shiyu Chang, Gerald Tesauro, Bowen Zhou 0002, Jing Jiang 0001 |
AAAI | 2 |
| 2018 | Image Super-Resolution via Dual-State Recurrent NetworksabstractAdvances in image super-resolution (SR) have recently benefited significantly from rapid developments in deep neural networks. Inspired by these recent discoveries, we note that many state-of-the-art deep SR architectures can be reformulated as a single-state recurrent neural network (RNN) with finite unfoldings. In this paper, we explore new structures for SR based on this compact RNN view, leading us to a dual-state design, the Dual-State Recurrent Network (DSRN). Compared to its single-state counterparts that operate at a fixed spatial resolution, DSRN exploits both low-resolution (LR) and high-resolution (HR) signals jointly. Recurrent signals are exchanged between these states in both directions (both LR to HR and HR to LR) via delayed feedback. Extensive quantitative and qualitative evaluations on benchmark datasets and on a recent challenge demonstrate that the proposed DSRN performs favorably against state-of-the-art algorithms in terms of both memory consumption and predictive accuracy. The code for our method is publicly available1. Wei Han 0002, Shiyu Chang, Ding Liu 0001, Mo Yu, Michael Witbrock, Thomas S. Huang |
CVPR | 4 |
| 2018 | Deriving Machine Attention from Human RationalesabstractAttention-based models are successful when trained on large amounts of data.In this paper, we demonstrate that even in the low-resource scenario, attention can be learned effectively.To this end, we start with discrete humanannotated rationales and map them into continuous attention.Our central hypothesis is that this mapping is general across domains, and thus can be transferred from resource-rich domains to low-resource ones.Our model jointly learns a domain-invariant representation and induces the desired mapping between rationales and attention.Our empirical results validate this hypothesis and show that our approach delivers significant gains over state-ofthe-art baselines, yielding over 15% average error reduction on benchmark datasets. Yujia Bao, Shiyu Chang, Mo Yu, Regina Barzilay |
EMNLP | 3 |
| 2018 | Improving Reinforcement Learning Based Image Captioning with Natural Language PriorabstractRecently, Reinforcement Learning (RL) approaches have demonstrated advanced performance in image captioning by directly optimizing the metric used for testing.However, this shaped reward introduces learning biases, which reduces the readability of generated text.In addition, the large sample space makes training unstable and slow.To alleviate these issues, we propose a simple coherent solution that constrains the action space using an n-gram language prior.Quantitative and qualitative evaluations on benchmarks show that RL with the simple add-on module performs favorably against its counterpart in terms of both readability and speed of convergence.Human evaluation results show that our model is more human readable and graceful.The implementation will become publicly available upon the acceptance of the paper 1 . Tszhang Guo, Shiyu Chang, Mo Yu |
EMNLP | 3 |
| 2018 | One-Shot Relational Learning for Knowledge GraphsabstractKnowledge graphs (KGs) are the key components of various natural language processing applications.To further expand KGs' coverage, previous studies on knowledge graph completion usually require a large number of training instances for each relation.However, we observe that long-tail relations are actually more common in KGs and those newly added relations often do not have many known triples for training.In this work, we aim at predicting new facts under a challenging setting where only one training instance is available.We propose a one-shot relational learning framework, which utilizes the knowledge extracted by embedding models and learns a matching metric by considering both the learned embeddings and one-hop graph structures.Empirically, our model yields considerable performance improvements over existing embedding models, and also eliminates the need of retraining the embedding models when dealing with newly added relations. 1 Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang |
EMNLP | 2 |
| 2018 | Exploiting Rich Syntactic Information for Semantic Parsing with Graph-to-Sequence ModelabstractExisting neural semantic parsers mainly utilize a sequence encoder, i.e., a sequential LSTM, to extract word order features while neglecting other valuable syntactic information such as dependency or constituent trees.In this paper, we first propose to use the syntactic graph to represent three types of syntactic information, i.e., word order, dependency and constituency features; then employ a graph-tosequence model to encode the syntactic graph and decode a logical form.Experimental results on benchmark datasets show that our model is comparable to the state-of-the-art on Jobs640, ATIS, and Geo880.Experimental results on adversarial examples demonstrate the robustness of the model is also improved by encoding more syntactic information. Kun Xu 0005, Lingfei Wu 0001, Zhiguo Wang 0006, Mo Yu, Vadim Sheinin |
EMNLP | 4 |
| 2018 | Evidence Aggregation for Answer Re-Ranking in Open-Domain Question Answering
Shuohang Wang, Mo Yu, Jing Jiang 0001, Wei Zhang 0057, Shiyu Chang, Tim Klinger, Gerald Tesauro, Murray Campbell |
ICLR (Poster) | 2 |
| 2018 | Scheduled Policy Optimization for Natural Language Communication with Intelligent AgentsabstractWe investigate the task of learning to interpret natural language instructions by jointly reasoning with visual observations and language inputs. Unlike current methods which start with learning from demonstrations (LfD) and then use reinforcement learning (RL) to fine-tune the model parameters, we propose a novel policy optimization algorithm which can dynamically schedule demonstration learning and RL. The proposed training paradigm provides efficient exploration and generalization beyond existing methods. Comparing to existing ensemble models, the best single model based on our proposed method tremendously decreases the execution error by 55% on a block-world environment. To further illustrate the exploration strategy of our RL algorithm, our paper includes systematic studies on the evolution of policy entropy during training. Wenhan Xiong, Mo Yu, Shiyu Chang, William Yang Wang |
IJCAI | 3 |
| 2018 | Diverse Few-Shot Text Classification with Multiple MetricsabstractMo Yu, Xiaoxiao Guo, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng, Gerald Tesauro, Haoyu Wang, Bowen Zhou. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Mo Yu, Jinfeng Yi, Shiyu Chang, Saloni Potdar, Yu Cheng 0001, Gerald Tesauro, Haoyu Wang 0002 |
NAACL-HLT | 1 |
| 2017 | Improved Neural Relation Detection for Knowledge Base Question AnsweringabstractRelation detection is a core component of many NLP applications including Knowledge Base Question Answering (KBQA).In this paper, we propose a hierarchical recurrent neural network enhanced by residual learning which detects KB relations given an input question.Our method uses deep residual bidirectional LSTMs to compare questions and relation names via different levels of abstraction.Additionally, we propose a simple KBQA system that integrates entity linking and our proposed relation detector to make the two components enhance each other.Our experimental results show that our approach not only achieves outstanding relation detection performance, but more importantly, it helps our KBQA system achieve state-of-the-art accuracy for both single-relation (SimpleQuestions) and multi-relation (WebQSP) QA benchmarks. Mo Yu, Wenpeng Yin 0001, Kazi Saidul Hasan, Cícero Nogueira dos Santos, Bing Xiang, Bowen Zhou 0002 |
ACL (1) | 1 |
| 2017 | A Structured Self-Attentive Sentence Embedding
Zhouhan Lin, Minwei Feng, Cícero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou 0002, Yoshua Bengio |
ICLR (Poster) | 4 |
| 2017 | Dilated Recurrent Neural NetworksabstractLearning with recurrent neural networks (RNNs) on long sequences is a notoriously difficult task. There are three major challenges: 1) complex dependencies, 2) vanishing and exploding gradients, and 3) efficient parallelization. In this paper, we introduce a simple yet effective RNN connection structure, the DilatedRNN, which simultaneously tackles all of these challenges. The proposed architecture is characterized by multi-resolution dilated recurrent skip connections and can be combined flexibly with diverse RNN cells. Moreover, the DilatedRNN reduces the number of parameters needed and enhances training efficiency significantly, while matching state-of-the-art performance (even with standard RNN cells) in tasks involving very long-term dependencies. To provide a theory-based quantification of the architecture's advantages, we introduce a memory capacity measure, the mean recurrent length, which is more suitable for RNNs with long skip connections than existing measures. We rigorously prove the advantages of the DilatedRNN over other recurrent neural architectures. The code for our method is publicly available at https://github.com/code-terminator/DilatedRNN. Shiyu Chang, Yang Zhang 0001, Wei Han 0002, Mo Yu, Wei Tan 0001, Michael Witbrock, Mark Hasegawa-Johnson, Thomas S. Huang |
NIPS | 4 |
| 2016 | New to online dating? Learning from experienced users for a successful matchabstractOnline dating arises as a popular venue for finding romantic partners in recent years. Many online dating sites adopt recommender systems to help their users. However, few of current research provides solutions to cold start problem, i.e., providing recommendations to new users. In this research, we propose a new approach of providing reciprocal online dating recommendations to new users. Specifically, we detect communities from existing users, match new users to these communities, and take advantage of reciprocal activities of those community members to provide recommendations to new users. Using data from a popular U.S. online dating site, experiments show that our approach greatly outperforms existing methods. Mo Yu, Xiaolong Zhang 0001, Derek Kreager |
ASONAM | 1 |
| 2016 | Simple Question Answering by Attentive Convolutional Neural NetworkabstractThis work focuses on answering single-relation factoid questions over Freebase. Each question can acquire the answer from a single fact of form (subject, predicate, object) in Freebase. This task, simple question answering (SimpleQA), can be addressed via a two-step pipeline: entity linking and fact selection. In fact selection, we match the subject entity in a fact candidate with the entity mention in the question by a character-level convolutional neural network (char-CNN), and match the predicate in that fact with the question by a word-level CNN (word-CNN). This work makes two main contributions. (i) A simple and effective entity linker over Freebase is proposed. Our entity linker outperforms the state-of-the-art entity linker over SimpleQA task. (ii) A novel attentive maxpooling is stacked over word-CNN, so that the predicate representation can be matched with the predicate-focused question representation more effectively. Experiments show that our system sets new state-of-the-art in this task. Wenpeng Yin 0001, Mo Yu, Bing Xiang, Bowen Zhou 0002, Hinrich Schütze |
COLING | 2 |
| 2016 | Leveraging Sentence-level Information with Encoder LSTM for Semantic Slot FillingabstractRecurrent Neural Network (RNN) and one of its specific architectures, Long Short-Term Memory (LSTM), have been widely used for sequence labeling.Explicitly modeling output label dependencies on top of RNN/LSTM is a widely-studied and effective extension.We propose another extension to incorporate the global information spanning over the whole input sequence.The proposed method, encoder-labeler LSTM, first encodes the whole input sequence into a fixed length vector with the encoder LSTM, and then uses this encoded vector as the initial state of another LSTM for sequence labeling.With this method, we can predict the label sequence while taking the whole input sequence information into consideration.In the experiments of a slot filling task, which is an essential component of natural language understanding, with using the standard ATIS corpus, we achieved the state-of-the-art F 1 -score of 95.66%. Gakuto Kurata, Bing Xiang, Bowen Zhou 0006, Mo Yu |
EMNLP | 4 |
| 2016 | Embedding Lexical Features via Low-Rank TensorsabstractMo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Mo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley |
HLT-NAACL | 1 |
| 2015 | Improved Relation Extraction with Feature-Rich Compositional Embedding ModelsabstractCompositional embedding models build a representation (or embedding) for a linguistic structure based on its component word embeddings.We propose a Feature-rich Compositional Embedding Model (FCM) for relation extraction that is expressive, generalizes to new domains, and is easy-to-implement.The key idea is to combine both (unlexicalized) handcrafted features with learned word embeddings.The model is able to directly tackle the difficulties met by traditional compositional embeddings models, such as handling arbitrary types of sentence annotations and utilizing global information for composition.We test the proposed model on two relation extraction tasks, and demonstrate that our model outperforms both previous compositional models and traditional feature rich models on the ACE 2005 relation extraction task, and the SemEval 2010 relation classification task.The combination of our model and a loglinear classifier with hand-crafted features gives state-of-the-art results.We made our implementation available for general use 1 . Matthew R. Gormley, Mo Yu, Mark Dredze |
EMNLP | 2 |
| 2015 | A Concrete Chinese NLP PipelineabstractNanyun Peng, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015. Nanyun Peng 0001, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze |
HLT-NAACL | 3 |
| 2015 | Combining Word Embeddings and Feature Embeddings for Fine-grained Relation ExtractionabstractCompositional embedding models build a rep-resentation for a linguistic structure based on its component word embeddings. While re-cent work has combined these word embed-dings with hand crafted features for improved performance, it was restricted to a small num-ber of features due to model complexity, thus limiting its applicability. We propose a new model that conjoins features and word em-beddings while maintaing a small number of parameters by learning feature embeddings jointly with the parameters of a compositional model. The result is a method that can scale to more features and more labels, while avoiding overfitting. We demonstrate that our model at-tains state-of-the-art results on ACE and ERE fine-grained relation extraction. 1 Mo Yu, Matthew R. Gormley, Mark Dredze |
HLT-NAACL | 1 |
| 2015 | Learning Composition Models for Phrase EmbeddingsabstractLexical embeddings can serve as useful representations for words for a variety of NLP tasks, but learning embeddings for phrases can be challenging. While separate embeddings are learned for each word, this is infeasible for every phrase. We construct phrase embeddings by learning how to compose word embeddings using features that capture phrase structure and context. We propose efficient unsupervised and task-specific learning objectives that scale our model to large datasets. We demonstrate improvements on both language modeling and several phrase semantic similarity tasks with various phrase lengths. We make the implementation of our model and the datasets available for general use. Mo Yu, Mark Dredze |
Trans. Assoc. Comput. Linguistics | 1 |
| 2014 | Who were you talking to - Mining interpersonal relationships from cellphone network dataabstractPeople play different roles in various social networks. Even in a single network, people may interact with others based on different roles, and there are various relationships among them. However, current research usually treats all relationships homogeneously (i.e. friendship). In this paper, we try to identify different types of relationship (family, colleague, and social) within social networks. By analyzing a large-scale cellphone network, we gain insights about human mobility patterns. We design three metrics to capture colocation behaviors for cellphone users, taking spatial-temporal factors into consideration. These metrics show that users with different relationships demonstrate significantly different co-locating patterns. With these metrics as features, we adopt supervised approach to classify cellphone user pairs into different relationship categories. Comparing to using network and communication features, co-location metrics demonstrate better performance to fulfill the task of relationship identification. Mo Yu, Wenjun Si, Guojie Song, Zhenhui Li, John Yen |
ASONAM | 1 |
| 2014 | Accelerated Mini-batch Randomized Block Coordinate Descent Method
Tuo Zhao, Mo Yu, Raman Arora, Han Liu 0001 |
NIPS | 2 |
| 2013 | Learning Domain Differences Automatically for Dependency Parsing Adaptation
Mo Yu, Tiejun Zhao, Yalong Bai |
IJCAI | 1 |
| 2013 | Compound Embedding Features for Semi-supervised Learning
Mo Yu, Tiejun Zhao, Daxiang Dong, Dianhai Yu |
HLT-NAACL | 1 |
| 2012 | Locally Training the Log-Linear Model for SMT
Lemao Liu, Hailong Cao, Taro Watanabe, Tiejun Zhao, Mo Yu, Conghui Zhu |
EMNLP-CoNLL | 5 |
| 2011 | Target-dependent Twitter Sentiment Classification
Long Jiang, Mo Yu, Ming Zhou 0001, Tiejun Zhao |
ACL | 2 |
| 2010 | Diversifying landmark image search results by learning interested views from community photosabstractIn this paper, we demonstrate a novel landmark photo search and browsing system: Agate, which ranks landmark image search results considering their relevance, diversity and quality. Agate learns from community photos the most interested aspects and related activities of a landmark, and generates adaptively a Table of Content (TOC) as a summary of the attractions to facilitate the user browsing. Image search results are thus re-ranked with the TOC so as to ensure a quick overview of the attractions of the landmarks. A novel non-parametric TOC generation and set-based ranking algorithm, MoM-DPM Sets, is proposed as the key technology of Agate. Experimental results based on human evaluation show the effectiveness of our model and users' preference for Agate. Yuheng Ren, Mo Yu, Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
WWW | 2 |
| 2009 | Advertising based on users' photosabstractIn this paper, we tackle the problem of learning a user's interest from his photo collections and suggesting relevant ads. We address two key challenges in this work: 1) understanding a user's photos to detect his interest, and 2) bridging the lexical and semantic gap between the vocabulary of ads and that of general users' photos. We solve the first problem by employing a data-driven image annotation approach to annotate each photo and modeling a group of photos, and tackle the second problem by learning and matching the topics of users' photos and ads. The experiments based on real flicker data showed the effectiveness of the approach. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ICME | 2 |
| 2009 | Argo: intelligent advertising made possible from users' photosabstractThough monetizing user-generated photos has a great potential in image business, this topic is seldom touched due to the difficulties of both image understanding and ads-to-images vocabulary matching. In this technical demonstration, we show case the Argo system, which attempts to monetize UGC (user-generated content) photos by mining a user's interest from a group of his photos and advertising the photos accordingly. Given a page of photos, it first auto-tags each photo by a large-scale search-based image annotation method, then maps both image annotations and the textual descriptions of ads onto an ODP-based topic hierarchy. The mapping produces semantic features which are statistical distributions on ODP topics. Ads are ranked by their similarities to such topic distributions of the photos and the top-ranked ones are output. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 2 |