EDBT 2026 Demo / reviewers in the wild / expert
Xiaodan Zhu 0001
dblp:93/310
· DBLP profile ↗
81ranked-venue papers
15as first author
30since 2021 · last 2026
0000-0003-3856-3696ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 65 · 10 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward ModellingabstractProcess or step-wise supervision has played a crucial role in advancing complex multi-step reasoning capabilities of Large Language Models (LLMs). However, efficient, high-quality automated process annotation remains a significant challenge. To address this, we introduce Single-Pass Annotation with Reference-Guided Evaluation (SPARE), a novel structured framework that enables efficient per-step annotation by jointly aligning solution steps to reference solutions and determine its accuracy with explicit reasoning in single generation. We demonstrate SPARE's effectiveness across four diverse datasets spanning mathematical reasoning (GSM8K, MATH), multi-hop question answering (MuSiQue-Ans), and spatial reasoning (SpaRP), showing consistent improvements in two applications: (1) training Process Reward Models (PRMs) for ranking and aggregating multiple generations, and (2) fine-tuning models via offline reinforcement learning for greedy decoding. On PROCESSBENCH, SPARE demonstrates data-efficient out-of-distribution generalization, using only ~16% of training samples compared to human-labeled and other synthetically trained baselines. Additionally, it achieves competitive performance with MCTS-based methods while offering 2.3x speedup in terms of total token count. Manual analysis reveals complementary precision-recall characteristics with MCTS approaches, suggesting potential for ensemble methods. These results establish SPARE as a practical and scalable solution for automatic process supervision in LLM reasoning. Md Imbesat Hassan Rizvi, Xiaodan Zhu 0001, Iryna Gurevych |
AAAI | 2 |
| 2026 | CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow AlignmentabstractBuilding Task-Oriented Dialogue (TOD) systems that generalize across different tasks remains a challenging problem.Data-driven approaches often struggle to transfer effectively to unseen tasks.While recent schema-based TOD frameworks improve generalization by decoupling task logic from language understanding, their reliance on neural or generative models often obscures how task schemas influence behaviour and hence impair interpretability.In this work, we introduce a novel framework, CoDial (Code for Dialogue), at the core of which is converting a predefined task schema to a structured heterogeneous graph and then to programmatic LLM guardrailing code, such as NVIDIA's Colang.The pipeline enables efficient and interpretable alignment of dialogue policies during inference.We introduce two paradigms for LLM guardrailing code generation, CoDial free and CoDial structured , and propose a mechanism that integrates human feedback to iteratively improve the generated code.Empirically, CoDial achieves state-ofthe-art (SOTA) performance on the widely used benchmark datasets, while providing inherent interpretability in the design.We additionally demonstrate CoDial's iterative improvement via manual and LLM-aided feedback, making it a practical tool for human-guided alignment of LLMs in unseen domains. 1 Radin Shayanfar, Chu Fei Luo, Rohan Bhambhoria, Samuel Dahan, Xiaodan Zhu 0001 |
ACL (1) | 5 |
| 2025 | Error Diversity Matters: An Error-Resistant Ensemble Method for Unsupervised Dependency ParsingabstractWe address unsupervised dependency parsing by building an ensemble of diverse existing models through post hoc aggregation of their output dependency parse structures. We observe that these ensembles often suffer from low robustness against weak ensemble components due to error accumulation. To tackle this problem, we propose an efficient ensemble-selection approach that considers error diversity and avoids error accumulation. Results demonstrate that our approach outperforms each individual model as well as previous ensemble techniques. Additionally, our experiments show that the proposed ensemble-selection method significantly enhances the performance and robustness of our ensemble, surpassing previously proposed strategies, which have not accounted for error diversity. Behzad Shayegh, Hobie H.-B. Lee, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou |
AAAI | 3 |
| 2025 | Robust Utility-Preserving Text Anonymization Based on Large Language ModelsabstractAnonymizing text that contains sensitive information is crucial for a wide range of applications. Existing techniques face the emerging challenges of the re-identification ability of large language models (LLMs), which have shown advanced capability in memorizing detailed information and reasoning over dispersed pieces of patterns to draw conclusions. When defending against LLM-based re-identification, anonymization could jeopardize the utility of the resulting anonymized data in downstream tasks. In general, the interaction between anonymization and data utility requires a deeper understanding within the context of LLMs. In this paper, we propose a framework composed of three key LLM-based components: \textit{a privacy evaluator}, \textit{a utility evaluator} and \textit{an optimization component}, which work collaboratively to perform anonymization. Extensive experiments demonstrate that the proposed model outperforms existing baselines, showing robustness in reducing the risk of re-identification while preserving greater data utility in downstream tasks. We provide detailed studies on these core modules. To consider large-scale and real-time applications, we investigate the distillation of the anonymization capabilities into lightweight models. All of our code and datasets will be made publicly available at \texttt{[Github URL]}. Tianyu Yang 0004, Xiaodan Zhu 0001, Iryna Gurevych |
ACL (1) | 2 |
| 2025 | Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMsabstractHaritz Puerto, Tilek Chubakov, Xiaodan Zhu, Harish Tayyar Madabushi, Iryna Gurevych. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Haritz Puerto, Tilek Chubakov, Xiaodan Zhu 0001, Harish Tayyar Madabushi, Iryna Gurevych |
ACL (1) | 3 |
| 2025 | Preemptive Detection and Correction of Misaligned Actions in LLM AgentsabstractDeploying LLM-based agents in real-life applications often faces a critical challenge: the misalignment between agents' behavior and user intent.Such misalignment may lead agents to unintentionally execute some critical actions that carry negative outcomes (e.g., accidentally triggering a 'buy-now' in web shopping), resulting in undesirable or even irreversible consequences.Although addressing these issues is crucial, the preemptive detection and correction of misaligned actions remains relatively underexplored.To fill this gap, we introduce InferAct, a novel approach that leverages the belief reasoning ability of LLMs, grounded in Theory-of-Mind, to detect misaligned actions before execution.Once the misalignment is detected, InferAct alerts users for timely correction, preventing adverse outcomes and enhancing the reliability of LLM agents' decision-making processes.Experiments on three widely used tasks demonstrate InferAct achieves up to 20% improvements on Marco-F1 against baselines in misaligned action detection.An in-depth evaluation of misalignment correction further highlights InferAct's effectiveness in improving agent alignment.1 I want to find white blackout shades that are 66 inches in width and 66 inches in height.they need to be easy to install, and price lower than 90.00 dollars.Thought: I think I should use the search bar to look for white blackout shades that are 66x66… Action: search[white blackout shades 66x66 inches easy to install price < 90.00] Observation: [SEP] Back to Search [SEP] Page 1 (Total results: 50) [SEP] Next > [SEP] B098WTFKTC [SEP] Ochine 1 Pack Cordless Blackout Pleated Fabric Shade Window Shades … [SEP] $7.99 [SEP] B0126EWYNI [SEP] Redi Shade Easy Lift Trim-At-Home Cordless ...Thought: The product B09LS7KQMC seems to offer custom cut to size cellular shades, which could potentially meet the size requirement of 66 inches in width and 66 inches in height, and it is also within the price range specified.... Haishuo Fang, Xiaodan Zhu 0001, Iryna Gurevych |
EMNLP | 2 |
| 2025 | On the Robustness of Verbal Confidence of LLMs in Adversarial AttacksabstractRobust verbal confidence generated by large language models (LLMs) is crucial for the deployment of LLMs to help ensure transparency, trust, and safety in many applications, including those involving human-AI interactions. In this paper, we present the first comprehensive study on the robustness of verbal confidence under adversarial attacks. We introduce attack frameworks targeting verbal confidence scores through both perturbation and jailbreak-based methods, and demonstrate that these attacks can significantly impair verbal confidence estimates and lead to frequent answer changes. We examine a variety of prompting strategies, model sizes, and application domains, revealing that current verbal confidence is vulnerable and that commonly used defence techniques are largely ineffective or counterproductive. Our findings underscore the need to design robust mechanisms for confidence expression in LLMs, as even subtle semantic-preserving modifications can lead to misleading confidence in responses. Stephen Obadinma, Xiaodan Zhu 0001 |
NeurIPS | 2 |
| 2024 | SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language ModelsabstractSpatial reasoning is a crucial component of both biological and artificial intelligence.In this work, we present a comprehensive study of the capability of current state-of-the-art large language models (LLMs) on spatial reasoning.To support our study, we created and contribute a novel Spatial Reasoning Characterization (SpaRC) framework and Spatial Reasoning Paths (SpaRP) 1 datasets, to enable an in-depth understanding of the spatial relations and compositions as well as the usefulness of spatial reasoning chains.We found that all the stateof-the-art LLMs do not perform well on the datasets-their performances are consistently low across different setups.The spatial reasoning capability improves substantially as model sizes scale up.Finetuning both large language models (e.g., Llama-2-70B) and smaller ones (e.g., Llama-2-13B) can significantly improve their F1-scores by 7-32 absolute points.We also found that the top proprietary LLMs still significantly outperform their open-source counterparts in topological spatial understanding and reasoning. Md Imbesat Hassan Rizvi, Xiaodan Zhu 0001, Iryna Gurevych |
ACL (1) | 2 |
| 2024 | Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMsabstractReasoning is a fundamental component of language understanding.Recent prompting techniques, such as chain of thought, have consistently improved LLMs' performance on various reasoning tasks.Nevertheless, there is still little understanding of what triggers reasoning abilities in LLMs in the inference stage.In this paper, we investigate the effect of the input representation on the reasoning abilities of LLMs.We hypothesize that representing natural language tasks as code can enhance specific reasoning abilities such as entity tracking or logical reasoning.To study this, we propose code prompting, a methodology we operationalize as a chain of prompts that transforms a natural language problem into code and directly prompts the LLM using the generated code without resorting to external code execution.We find that code prompting exhibits a high-performance boost for multiple LLMs (up to 22.52 percentage points on GPT 3.5, 7.75 on Mixtral, and 16.78 on Mistral) across multiple conditional reasoning datasets.We then conduct comprehensive experiments to understand how the code representation triggers reasoning abilities and which capabilities are elicited in the underlying models.Our analysis on GPT 3.5 reveals that the code formatting of the input problem is essential for performance improvement.Furthermore, the code representation improves sample efficiency of in-context learning and facilitates state tracking of entities.1 Haritz Puerto, Martin Tutek, Somak Aditya, Xiaodan Zhu 0001, Iryna Gurevych |
EMNLP | 4 |
| 2024 | Exploring the Role of Reasoning Structures for Constructing Proofs in Multi-Step Natural Language Reasoning with Large Language ModelsabstractWhen performing complex multi-step reasoning tasks, the ability of Large Language Models (LLMs) to derive structured intermediate proof steps is important for ensuring that the models truly perform the desired reasoning and for improving models' explainability.This paper is centred around a focused study: whether the current state-of-the-art generalist LLMs can leverage the structures in a few examples to better construct the proof structures with incontext learning.Our study specifically focuses on structure-aware demonstration and structureaware pruning.We demonstrate that they both help improve performance.A detailed analysis is provided to help understand the results. 1 Zi'ou Zheng, Christopher Malon, Martin Renqiang Min, Xiaodan Zhu 0001 |
EMNLP | 4 |
| 2024 | Ensemble Distillation for Unsupervised Constituency ParsingabstractWe investigate the unsupervised constituency parsing task, which organizes words and phrases of a sentence into a hierarchical structure without using linguistically annotated data. We observe that existing unsupervised parsers capture different aspects of parsing structures, which can be leveraged to enhance unsupervised parsing performance.
To this end, we propose a notion of "tree averaging," based on which we further propose a novel ensemble method for unsupervised parsing.
To improve inference efficiency, we further distill the ensemble knowledge into a student model; such an ensemble-then-distill process is an effective approach to mitigate the over-smoothing problem existing in common multi-teacher distilling methods.
Experiments show that our method surpasses all previous approaches, consistently demonstrating its effectiveness and robustness across various runs, with different ensemble components, and under domain-shift conditions. Behzad Shayegh, Yanshuai Cao, Xiaodan Zhu 0001, Jackie Chi Kit Cheung, Lili Mou |
ICLR | 3 |
| 2023 | NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural LogicabstractReasoning has been a central topic in artificial intelligence from the beginning.The recent progress made on distributed representation and neural networks continues to improve the state-of-the-art performance of natural language inference.However, it remains an open question whether the models perform real reasoning to reach their conclusions or rely on spurious correlations.Adversarial attacks have proven to be an important tool to help evaluate the Achilles' heel of the victim models.In this study, we explore the fundamental problem of developing attack models based on logic formalism.We propose NatLogAttack to perform systematic attacks centring around natural logic, a classical logic formalism that is traceable back to Aristotle's syllogism and has been closely developed for natural language inference.The proposed framework renders both label-preserving and label-flipping attacks.We show that compared to the existing attack models, NatLogAttack generates better adversarial examples with fewer visits to the victim models.The victim models are found to be more vulnerable under the label-flipping setting.NatLogAttack provides a tool to probe the existing and future NLI models' capacity from a key viewpoint and we hope more logicbased attacks will be further explored for understanding the desired property of reasoning.1 Zi'ou Zheng, Xiaodan Zhu 0001 |
ACL (1) | 2 |
| 2023 | OssCSE: Overcoming Surface Structure Bias in Contrastive Learning for Unsupervised Sentence EmbeddingabstractZhan Shi, Guoyin Wang, Ke Bai, Jiwei Li, Xiang Li, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Guoyin Wang 0002, Ke Bai 0001, Jiwei Li 0001, Qingjun Cui, Belinda Zeng, Trishul Chilimbi, Xiaodan Zhu 0001 |
EMNLP | 9 |
| 2023 | OpenJustice.ai: A Global Open-Source Legal Language ModelabstractGeneralized AI like ChatGPT cannot and should not be used for legal tasks. It presents significant risks for both the legal professions as well as litigants. However, domain-specific AI should not be ruled out. It has the potential for legal research as well as access to justice. In this paper, we call for the development of an open-source and distributed legal AI accessible to the entire legal community. We believe it has the potential to address some of the limitations related to the use of general AI for legal problems and resolving disputes – shortcomings that include legal misinformation or hallucinations, lack of transparency and precision, and inability to offer diverse and multiple narratives. Samuel Dahan, Rohan Bhambhoria, David Liang, Xiaodan Zhu 0001 |
JURIX | 4 |
| 2023 | Robust NLP for Finance (RobustFin)abstractNatural language processing (NLP) technologies have been widely applied in business domains such as e-commerce and customer service, but their adoption in the financial sector has been constrained by industry-specific performance standards and regulatory restrictions. This challenge has created new opportunities for core research in related areas. Recent advancements in NLP, such as the advent of large language models, has encouraged adoption in the finance sector. However, compared to other domains, finance has stricter requirements for robustness, explainability, and generalizability. Given this background, we propose to organize the first Robust NLP for Finance (RobustFin) workshop at KDD '23 to encourage the study of and research on robustness and explainability technologies with regard to financial NLP. The goal of the workshop is to extend the applications of NLP in finance, while motivating further research in robust NLP. Sameena Shah, Xiaodan Zhu 0001, Gerard de Melo, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Zhiyu Chen 0002 |
KDD | 2 |
| 2023 | Knowledge Discovery from Unstructured Data in Financial Services (KDF) WorkshopabstractKnowledge discovery from unstructured data, including business documents, web content, and news articles, has been a key AI challenge for the financial services industry. Comprehending these corpora and discovering knowledge from them, which could be textual, tabular, or graphic, are the cornerstone of supporting business decisions in the financial services domain, where information retrieval and content analysis techniques are of fundamental importance. We propose a workshop on knowledge discovery from unstructured data in financial services at SIGIR 2023 to highlight the current and emerging opportunities, invite original research, and prompt success sharing between researchers. Sameena Shah, Xiaodan Zhu 0001, Wenhu Chen, Manling Li, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Yulong Pei, Akshat Gupta |
SIGIR | 2 |
| 2022 | Interpretable Low-Resource Legal Decision MakingabstractOver the past several years, legal applications of deep learning have been on the rise. However, as with other high-stakes decision making areas, the requirement for interpretability is of crucial importance. Current models utilized by legal practitioners are more of the conventional machine learning type, wherein they are inherently interpretable, yet unable to harness the performance capabilities of data-driven deep learning models. In this work, we utilize deep learning models in the area of trademark law to shed light on the issue of likelihood of confusion between trademarks. Specifically, we introduce a model-agnostic interpretable intermediate layer, a technique which proves to be effective for legal documents. Furthermore, we utilize weakly supervised learning by means of a curriculum learning strategy, effectively demonstrating the improved performance of a deep learning model. This is in contrast to the conventional models which are only able to utilize the limited number of expensive manually-annotated samples by legal experts. Although the methods presented in this work tackles the task of risk of confusion for trademarks, it is straightforward to extend them to other fields of law, or more generally, to other similar high-stakes application scenarios. Rohan Bhambhoria, Hui Liu 0033, Samuel Dahan, Xiaodan Zhu 0001 |
AAAI | 4 |
| 2022 | Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial RelationsabstractPhrase grounding is a multi-modal problem that localizes a particular noun phrase in an image referred to by a text query. In the challenging zero-shot phrase grounding setting, the existing state-of-the-art grounding models have limited capacity in handling the unseen phrases. Humans, however, can ground novel types of objects in images with little effort, significantly benefiting from reasoning with commonsense. In this paper, we design a novel phrase grounding architecture that builds multi-modal knowledge graphs using external knowledge and then performs graph reasoning and spatial relation reasoning to localize the referred nouns phrases. We perform extensive experiments on different zero-shot grounding splits sub-sampled from the Flickr30K Entity and Visual Genome dataset, demonstrating that the proposed framework is orthogonal to backbone image encoders and outperforms the baselines by 2~3% in accuracy, resulting in a significant improvement under the standard evaluation metrics. Yilin Shen, Hongxia Jin, Xiaodan Zhu 0001 |
AAAI | 4 |
| 2022 | Neuro-symbolic Natural Logic with Introspective Revision for Natural Language InferenceabstractAbstract We introduce a neuro-symbolic natural logic framework based on reinforcement learning with introspective revision. The model samples and rewards specific reasoning paths through policy gradient, in which the introspective revision algorithm modifies intermediate symbolic reasoning steps to discover reward-earning operations as well as leverages external knowledge to alleviate spurious reasoning and training inefficiency. The framework is supported by properly designed local relation models to avoid input entangling, which helps ensure the interpretability of the proof paths. The proposed model has built-in interpretability and shows superior capability in monotonicity inference, systematic generalization, and interpretability, compared with previous models on the existing datasets. Xiaoyu Yang 0002, Xiaodan Zhu 0001, Michael A. Greenspan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | Dynamic Hybrid Relation Exploration Network for Cross-Domain Context-Dependent Semantic ParsingabstractSemantic parsing has long been a fundamental problem in natural language processing. Recently, cross-domain context-dependent semantic parsing has become a new focus of research. Central to the problem is the challenge of leveraging contextual information of both natural language queries and database schemas in the interaction history. In this paper, we present a dynamic graph framework that is capable of effectively modelling contextual utterances, tokens, database schemas, and their complicated interaction as the conversation proceeds. The framework employs a dynamic memory decay mechanism that incorporates inductive bias to integrate enriched contextual relation representation, which is further enhanced with a powerful reranking model. At the time of writing, we demonstrate that the proposed framework outperforms all existing models by large margins, achieving new state-of-the-art performance on two large-scale benchmarks, the SParC and CoSQL datasets. Specifically, the model attains a 55.8% question-match and 30.8% interaction-match accuracy on SParC, and a 46.8% question-match and 17.0% interaction-match accuracy on CoSQL. Binyuan Hui, Ruiying Geng, Qiyu Ren, Binhua Li, Yongbin Li 0001, Jian Sun 0021, Fei Huang 0002, Luo Si, Pengfei Zhu 0001, Xiaodan Zhu 0001 |
AAAI | 10 |
| 2021 | Identifying Untrustworthy Samples: Data Filtering for Open-domain Dialogues with Bayesian OptimizationabstractBeing able to reply with a related, fluent, and informative response is an indispensable requirement for building high-quality conversational agents. In order to generate better responses, some approaches have been proposed, such as feeding extra information by collecting large-scale datasets with human annotations, designing neural conversational models (NCMs) with complex architecture and loss functions, or filtering out untrustworthy samples based on a dialogue attribute, e.g., Relatedness or Genericness. In this paper, we follow the third research branch and present a data filtering method for open-domain dialogues, which identifies untrustworthy samples from training data with a quality measure that linearly combines seven dialogue attributes. The attribute weights are obtained via Bayesian Optimization (BayesOpt) that aims to optimize an objective function for dialogue generation iteratively on the validation set. Then we score training samples with the quality measure, sort them in descending order, and filter out those at the bottom. Furthermore, to accelerate the "filter-train-evaluate'' iterations involved in BayesOpt on large-scale datasets, we propose a training framework that integrates maximum likelihood estimation (MLE) and negative training method (NEG). The training method updates parameters of a trained NCMs on two small sets with newly maintained and removed samples, respectively. Specifically, MLE is applied to maximize the log-likelihood of newly maintained samples, while NEG is used to minimize the log-likelihood of newly removed ones. Experimental results on two datasets show that our method can effectively identify untrustworthy samples, and NCMs trained on the filtered datasets achieve better performance. Lei Shen 0001, Haolan Zhan, Hongshen Chen, Xiaodan Zhu 0001 |
CIKM | 6 |
| 2021 | Unsupervised Conversation Disentanglement through Co-TrainingabstractConversation disentanglement aims to separate intermingled messages into detached sessions, which is a fundamental task in understanding multi-party conversations.Existing work on conversation disentanglement relies heavily upon human-annotated datasets, which are expensive to obtain in practice.In this work, we explore to train a conversation disentanglement model without referencing any human annotations.Our method is built upon a deep co-training algorithm, which consists of two neural networks: a messagepair classifier and a session classifier.The former is responsible for retrieving local relations between two messages while the latter categorizes a message to a session by capturing context-aware information.Both networks are initialized respectively with pseudo data built from an unannotated corpus.During the deep co-training process, we use the session classifier as a reinforcement learning component to learn a session assigning policy by maximizing the local rewards given by the messagepair classifier.For the message-pair classifier, we enrich its training data by retrieving message pairs with high confidence from the disentangled sessions predicted by the session classifier.Experimental results on the large Movie Dialogue Dataset demonstrate that our proposed approach achieves competitive performance compared to the previous supervised methods.Further experiments show that the predicted disentangled conversations can promote the performance on the downstream task of multi-party response selection. Hui Liu 0033, Xiaodan Zhu 0001 |
EMNLP (1) | 3 |
| 2021 | Detecting Speaker Personas from Conversational TextsabstractPersonas are useful for dialogue response prediction.However, the personas used in current studies are pre-defined and hard to obtain before a conversation.To tackle this issue, we study a new task, named Speaker Persona Detection (SPD), which aims to detect speaker personas based on the plain conversational text.In this task, a best-matched persona is searched out from candidates given the conversational text.This is a many-to-many semantic matching task because both contexts and personas in SPD are composed of multiple sentences.The long-term dependency and the dynamic redundancy among these sentences increase the difficulty of this task.We build a dataset for SPD, dubbed as Persona Match on Persona-Chat (PMPC).Furthermore, we evaluate several baseline models and propose utterance-to-profile (U2P) matching networks for this task.The U2P models operate at a fine granularity which treat both contexts and personas as sets of multiple sequences.Then, each sequence pair is scored and an interpretable overall score is obtained for a context-persona pair through aggregation.Evaluation results show that the U2P models outperform their baseline counterparts significantly. Jia-Chen Gu, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Xiaodan Zhu 0001 |
EMNLP (1) | 6 |
| 2021 | WinoLogic: A Zero-Shot Logic-based Diagnostic Dataset for Winograd Schema ChallengeabstractThe recent success of neural language models (NLMs) on the Winograd Schema Challenge has called for further investigation of the commonsense reasoning ability of these models.Previous diagnostic datasets rely on crowd-sourcing which fails to provide coherent commonsense crucial for solving WSC problems.To better evaluate NLMs, we propose a logic-based framework that focuses on highquality commonsense knowledge.Specifically, we identify and collect formal knowledge formulas verified by theorem provers and translate such formulas into natural language sentences.Based on these true knowledge sentences, adversarial false ones are generated.We propose a new dataset named WINOLOGIC with these sentences.Given a problem in WINOLOGIC, NLMs need to decide whether the plausible knowledge sentences could correctly solve the corresponding WSC problems in a zero-shot setting.We also ask human annotators to validate WINOLOGIC to ensure it is humanagreeable.Experiments show that NLMs still struggle to comprehend commonsense knowledge as humans do, indicating that their reasoning ability could have been overestimated. Weinan He 0001, Canming Huang, Yongmei Liu 0001, Xiaodan Zhu 0001 |
EMNLP (1) | 4 |
| 2021 | Emotion Inference in Multi-Turn Conversations with Addressee-Aware Module and Ensemble StrategyabstractEmotion inference in multi-turn conversations aims to predict the participant's emotion in the next upcoming turn without knowing the participant's response yet, and is a necessary step for applications such as dialogue planning.However, it is a severe challenge to perceive and reason about the future feelings of participants, due to the lack of utterance information from the future.Moreover, it is crucial for emotion inference to capture the characteristics of emotional propagation in conversations, such as persistence and contagiousness.In this study, we focus on investigating the task of emotion inference in multi-turn conversations by modeling the propagation of emotional states among participants in the conversation history, and propose an addresseeaware module to automatically learn whether the participant keeps the historical emotional state or is affected by others in the next upcoming turn.In addition, we propose an ensemble strategy to further enhance the model performance.Empirical studies on three different benchmark conversation datasets demonstrate the effectiveness of the proposed model over several strong baselines. Dayu Li, Xiaodan Zhu 0001, Yang Li 0074, Suge Wang, Deyu Li 0001, Jian Liao 0005, Jianxing Zheng |
EMNLP (1) | 2 |
| 2021 | Have You Made a Decision? Where? A Pilot Study on Interpretability of Polarity Analysis Based on Advising Problem
Tianda Li, Jia-Chen Gu, Hui Liu 0033, Quan Liu 0003, Zhen-Hua Ling, Zhiming Su, Xiaodan Zhu 0001 |
ICASSP | 7 |
| 2021 | Improving Pretrained Models for Zero-shot Multi-label Text Classification through Reinforced Label Hierarchy ReasoningabstractExploiting label hierarchies has become a promising approach to tackling the zero-shot multi-label text classification (ZS-MTC) problem.Conventional methods aim to learn a matching model between text and labels, using a graph encoder to incorporate label hierarchies to obtain effective label representations (Rios and Kavuluru, 2018).More recently, pretrained models like BERT (Devlin et al., 2018) have been used to convert classification tasks into a textual entailment task (Yin et al., 2019).This approach is naturally suitable for the ZS-MTC task.However, pretrained models are underexplored in the existing work because they do not generate individual vector representations for text or labels, making it unintuitive to combine them with conventional graph encoding methods.In this paper, we explore to improve pretrained models with label hierarchies on the ZS-MTC task.We propose a Reinforced Label Hierarchy Reasoning (RLHR) approach to encourage interdependence among labels in the hierarchies during training.Meanwhile, to overcome the weakness of flat predictions, we design a rollback algorithm that can remove logical errors from predictions during inference.Experimental results on three reallife datasets show that our approach achieves better performance and outperforms previous non-pretrained methods on the ZS-MTC task. Hui Liu 0033, Danqing Zhang, Xiaodan Zhu 0001 |
NAACL-HLT | 4 |
| 2021 | Partner Matters! An Empirical Study on Fusing Personas for Personalized Response Selection in Retrieval-Based ChatbotsabstractPersona can function as the prior knowledge for maintaining the consistency of dialogue systems. Most of previous studies adopted the self persona in dialogue whose response was about to be selected from a set of candidates or directly generated, but few have noticed the role of partner in dialogue. This paper makes an attempt to thoroughly explore the impact of utilizing personas that describe either self or partner speakers on the task of response selection in retrieval-based chatbots. Four persona fusion strategies are designed, which assume personas interact with contexts or responses in different ways. These strategies are implemented into three representative models for response selection, which are based on the Hierarchical Recurrent Encoder (HRE), Interactive Matching Network (IMN) and Bidirectional Encoder Representations from Transformers (BERT) respectively. Empirical studies on the Persona-Chat dataset show that the partner personas neglected in previous studies can improve the accuracy of response selection in the IMN- and BERT-based models. Besides, our BERT-based model implemented with the context-response-aware persona fusion strategy outperforms previous methods by margins larger than 2.7% on original personas and 4.6% on revised personas in terms of [email protected] (top-1 accuracy), achieving a new state-of-the-art performance on the Persona-Chat dataset. Jia-Chen Gu, Hui Liu 0033, Zhen-Hua Ling, Quan Liu 0003, Zhigang Chen 0003, Xiaodan Zhu 0001 |
SIGIR | 6 |
| 2021 | Enhancing emotion inference in conversations with commonsense knowledge
Dayu Li, Xiaodan Zhu 0001, Yang Li 0074, Suge Wang, Deyu Li 0001, Jian Liao 0005, Jianxing Zheng |
Knowl. Based Syst. | 2 |
| 2021 | Deep Contextualized Utterance Representations for Response Selection and Dialogue AnalysisabstractThe NOESIS II challenge, as the Track 2 in the Eighth Dialogue System Technology Challenge (DSTC 8), is the extension of Track 1 in DSTC 7. Three new elements are incorporated into the extended track, i.e., dialogue with multiple participants, dialogue success, and dialogue disentanglement. These are vital for the creation of a deployed task-oriented dialogue system. This track is divided into four subtasks, the first two of which are evaluated in the form of response selection and the last two focus on dialogue analysis. This paper describes our methods developed for these four subtasks, which all employ deep contextualized utterance representations to make models aware of contextual information and to keep the intrinsic property of multi-turn dialogue systems. In the released evaluation results of Track 2 in DSTC 8, our proposed methods ranked fourth in subtask 1, third in subtask 2, and first in subtask 3 and subtask 4 respectively. In addition to the challenge tasks, we also compare our proposed methods with previous ones on public benchmark datasets. Experimental results show that our proposed methods outperform existing ones by large margins and achieve new state-of-the-art performances on multi-turn response selection and dialogue disentanglement. Jia-Chen Gu, Tianda Li, Zhen-Hua Ling, Quan Liu 0003, Zhiming Su, Yu-Ping Ruan, Xiaodan Zhu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2020 | Learning Cross-Modal Context Graph for Visual GroundingabstractVisual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic ambiguities. Prior works typically focus on learning representations of individual phrases with limited context information. To address their limitations, this paper proposes a language-guided graph representation to capture the global context of grounding entities and their relations, and develop a cross-modal graph matching strategy for the multiple-phrase visual grounding task. In particular, we introduce a modular graph neural network to compute context-aware representations of phrases and object proposals respectively via message propagation, followed by a graph-based matching module to generate globally consistent localization of grounding phrases. We train the entire graph neural network jointly in a two-stage strategy and evaluate it on the Flickr30K Entities benchmark. Extensive experiments show that our method outperforms the prior state of the arts by a sizable margin, evidencing the efficacy of our grounding framework. Code is available at https://github.com/youngfly11/LCMCG-PyTorch. Yongfei Liu, Xiaodan Zhu 0001, Xuming He 0001 |
AAAI | 3 |
| 2020 | Learning Low-Resource End-To-End Goal-Oriented Dialog for Fast and Reliable System DeploymentabstractExisting end-to-end dialog systems perform less effectively when data is scarce. To obtain an acceptable success in real-life online services with only a handful of training examples, both fast adaptability and reliable performance are highly desirable for dialog systems. In this paper, we propose the Meta-Dialog System (MDS), which combines the advantages of both meta-learning approaches and human-machine collaboration. We evaluate our methods on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning. Experimental results show that MDS significantly outperforms non-meta-learning baselines and can achieve more than 90% per-turn accuracies with only 10 dialogs on the extended-bAbI dataset. Yinpei Dai, Hangyu Li 0003, Chengguang Tang, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001 |
ACL | 6 |
| 2020 | Dynamic Memory Induction Networks for Few-Shot Text ClassificationabstractThis paper proposes Dynamic Memory Induction Networks (DMIN) for few-shot text classification.The model utilizes dynamic routing to provide more flexibility to memory-based few-shot learning in order to better adapt the support sets, which is a critical capacity of fewshot classification models.Based on that, we further develop induction models with query information, aiming to enhance the generalization ability of meta-learning.The proposed model achieves new state-of-the-art results on the miniRCV1 and ODIC dataset, improving the best performance (accuracy) by 2∼4%.Detailed analysis is further performed to show the effectiveness of each component. Ruiying Geng, Binhua Li, Yongbin Li 0001, Jian Sun 0021, Xiaodan Zhu 0001 |
ACL | 5 |
| 2020 | Improving Image Captioning with Better Use of CaptionabstractImage captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community.In this paper, we present a novel image captioning architecture to better explore semantics available in captions and leverage that to enhance both image representation and caption generation.Our models first construct caption-guided visual relationship graphs that introduce beneficial inductive bias using weakly supervised multi-instance learning.The representation is then enhanced with neighbouring and contextual nodes with their textual and visual features.During generation, the model further incorporates visual relationships using multi-task learning for jointly predicting word and object/predicate tag sequences.We perform extensive experiments on the MSCOCO dataset, showing that the proposed framework significantly outperforms the baselines, resulting in the state-of-the-art performance under a wide range of evaluation metrics.The code of our paper has been made publicly available.1 Xipeng Qiu, Xiaodan Zhu 0001 |
ACL | 4 |
| 2020 | Speaker-Aware BERT for Multi-Turn Response Selection in Retrieval-Based ChatbotsabstractIn this paper, we study the problem of employing pre-trained language models for multi-turn response selection in retrieval-based chatbots. A new model, named Speaker-Aware BERT (SA-BERT), is proposed in order to make the model aware of the speaker change information, which is an important and intrinsic property of multi-turn dialogues. Furthermore, a speaker-aware disentanglement strategy is proposed to tackle the entangled dialogues. This strategy selects a small number of most important utterances as the filtered context according to the speakers' information in them. Finally, domain adaptation is performed to incorporate the in-domain knowledge into pre-trained language models. Experiments on five public datasets show that our proposed model outperforms the present models on all metrics by large margins and achieves new state-of-the-art performances for multi-turn response selection. Jia-Chen Gu, Tianda Li, Quan Liu 0003, Zhen-Hua Ling, Zhiming Su, Si Wei, Xiaodan Zhu 0001 |
CIKM | 7 |
| 2020 | Exploring End-to-End Differentiable Natural Logic ModelingabstractWe explore end-to-end trained differentiable models that integrate natural logic with neural networks, aiming to keep the backbone of natural language reasoning based on the natural logic formalism while introducing subsymbolic vector representations and neural components.The proposed model adapts module networks to model natural logic operations, which is enhanced with a memory component to model contextual information.Experiments show that the proposed framework can effectively model monotonicity-based reasoning, compared to the baseline neural network models without built-in inductive bias for monotonicity-based reasoning.Our proposed model shows to be robust when transferred from upward to downward inference.We perform further analyses on the performance of the proposed model on aggregation, showing the effectiveness of the proposed subcomponents on helping achieve better intermediate aggregation performance. Zi'ou Zheng, Quan Liu 0003, Michael A. Greenspan, Xiaodan Zhu 0001 |
COLING | 5 |
| 2020 | Program Enhanced Fact Verification with Verbalization and Graph Attention NetworkabstractPerforming fact verification based on structured data is important for many real-life applications and is a challenging research problem, particularly when it involves both symbolic operations and informal inference based on language understanding.In this paper, we present a Program-enhanced Verbalization and Graph ATtention Network (ProgVGAT) to integrate programs and execution into textual inference models.Specifically, a verbalization with program execution model is proposed to accumulate evidences that are embedded in operations over the tables.Built on that, we construct the graph attention verification networks, which are designed to fuse different sources of evidences from verbalized program execution, program structures, and the original statements and tables, to make the final verification decision.To support the above framework, we propose a program selection module optimized with a new training strategy based on margin loss, to produce more accurate programs, which is shown to be effective in enhancing the final verification results.Experimental results show that the proposed framework achieves the new state-of-the-art performance, a 74.4% accuracy, on the benchmark dataset TABFACT.Our code is available at https://github.com/arielsho/Program-Enhanced-Table-Fact-Checking. Xiaoyu Yang 0002, Feng Nie, Quan Liu 0003, Zhigang Chen 0003, Xiaodan Zhu 0001 |
EMNLP (1) | 6 |
| 2020 | Multi-Modal Fusion With Observation Points For Skeleton Action RecognitionabstractCurrent methods for skeleton-based action recognition compute features based on the given skeleton joint information. We show that introducing new observation points in skeleton motion sequences and using them to create fused representations from multiple modalities such as joints and bones, can enhance the discriminative power of the original modalities. Moreover, such representations can be used to create new streams in multi-stream networks that fuse constructively with other streams trained on the original modalities, effectively exhibiting a dual behaviour and collectively boosting the performance of the network even further. We present one possible multi-modal fusion system with a single observation point that can easily be incorporated in existing networks and improves state-of-the-art results on the two popular J-HMDB and Kinetics-Skeleton action recognition datasets. Iqbal Singh, Xiaodan Zhu 0001, Michael A. Greenspan |
ICIP | 2 |
| 2020 | End-to-End Transition-Based Online Dialogue DisentanglementabstractDialogue disentanglement aims to separate intermingled messages into detached sessions. The existing research focuses on two-step architectures, in which a model first retrieves the relationships between two messages and then divides the message stream into separate clusters. Almost all existing work puts significant efforts on selecting features for message-pair classification and clustering, while ignoring the semantic coherence within each session. In this paper, we introduce the first end-to- end transition-based model for online dialogue disentanglement. Our model captures the sequential information of each session as the online algorithm proceeds on processing a dialogue. The coherence in a session is hence modeled when messages are sequentially added into their best-matching sessions. Meanwhile, the research field still lacks data for studying end-to-end dialogue disentanglement, so we construct a large-scale dataset by extracting coherent dialogues from online movie scripts. We evaluate our model on both the dataset we developed and the publicly available Ubuntu IRC dataset [Kummerfeld et al., 2019]. The results show that our model significantly outperforms the existing algorithms. Further experiments demonstrate that our model better captures the sequential semantics and obtains more coherent disentangled sessions. Hui Liu 0033, Jia-Chen Gu, Quan Liu 0003, Si Wei, Xiaodan Zhu 0001 |
IJCAI | 6 |
| 2020 | Anomaly Detection Based on Unsupervised Disentangled Representation Learning in Combination with Manifold LearningabstractIdentifying anomalous samples from highly complex and unstructured data is a crucial but challenging task in a variety of intelligent systems. In this paper, we present a novel deep anomaly detection framework named AnoDM (standing for Anomaly detection based on unsupervised Disentangled representation learning and Manifold learning). The disentanglement learning is currently implemented by β-VAE for automatically discovering interpretable factorized latent representations in a completely unsupervised manner. The manifold learning is realized by t-SNE for projecting the latent representations to a 2D map. We define a new anomaly score function by combining β-VAE's reconstruction error in the raw feature space and local density estimation in the t-SNE space. AnoDM was evaluated on both image and time-series data and achieved better results than models that use just one of the two measures and other deep learning methods. Iluju Kiringa, Tet Hin Yeap, Xiaodan Zhu 0001, Yifeng Li 0001 |
IJCNN | 4 |
| 2020 | Capsule Deep Generative Model That Forms Parse TreesabstractSupervised capsule networks are theoretically advantageous over convolutional neural networks, because they aim to model a range of transformations of local physical or abstract objects and part-whole relationships among them. However, it remains unclear how to use the concept of capsules in deep generative models. In this study, to address this challenge, we present a statistical modelling of capsules in deep generative models where distributions are formulated in the exponential family. The major contribution of this unsupervised method is that parse trees as representations of part-whole relationships can be dynamically learned from the data. Yifeng Li 0001, Xiaodan Zhu 0001, Richard Naud, Pengcheng Xi |
IJCNN | 2 |
| 2020 | Generating diverse conversation responses by creating and ranking multiple candidates
Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003, Jia-Chen Gu |
Comput. Speech Lang. | 3 |
| 2020 | Condition-Transforming Variational Autoencoder for Generating Diverse Short Text ConversationsabstractIn this article, conditional-transforming variational autoencoders (CTVAEs) are proposed for generating diverse short text conversations. In conditional variational autoencoders (CVAEs), the prior distribution of latent variable z follows a multivariate Gaussian distribution with mean and variance modulated by the input conditions. Previous work found that this distribution tended to become condition-independent in practical applications. Thus, this article designs CTVAEs to enhance the influence of conditions in CVAEs. In a CTVAE model, the latent variable z is sampled by performing a non-linear transformation on the combination of the input conditions and the samples from a condition-independent prior distribution N (0, I). In our experiments using a Chinese Sina Weibo dataset, the CTVAE model derives z samples for decoding with better condition-dependency than that of the CVAE model. The earth mover’s distance (EMD) between the distributions of the latent variable z at the training stage, and the testing stage is also reduced by using the CTVAE model. In subjective preference tests, our proposed CTVAE model performs significantly better than CVAE and sequence-to-sequence (Seq2Seq) models on generating diverse, informative, and topic-relevant responses. Yu-Ping Ruan, Zhen-Hua Ling, Xiaodan Zhu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Induction Networks for Few-Shot Text ClassificationabstractRuiying Geng, Binhua Li, Yongbin Li, Xiaodan Zhu, Ping Jian, Jian Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ruiying Geng, Binhua Li, Yongbin Li 0001, Xiaodan Zhu 0001, Ping Jian, Jian Sun 0021 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Dually Interactive Matching Network for Personalized Response Selection in Retrieval-Based ChatbotsabstractJia-Chen Gu, Zhen-Hua Ling, Xiaodan Zhu, Quan Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jia-Chen Gu, Zhen-Hua Ling, Xiaodan Zhu 0001, Quan Liu 0003 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Capsule Generative Models
Yifeng Li 0001, Xiaodan Zhu 0001 |
ICANN (1) | 2 |
| 2019 | Exploring deep neural networks for multitarget stance detectionabstractAbstract Detecting subjectivity expressed toward concerned targets is an interesting problem and has received intensive study. Previous work often treated each target independently, ignoring the potential (sometimes very strong) dependency that could exist among targets (eg, the subjectivity expressed toward two products or two political candidates in an election). In this paper, we relieve such an independence assumption in order to jointly model the subjectivity expressed toward multiple targets. We propose and show that an attention‐based encoder‐decoder framework is very effective for this problem, outperforming several alternatives that jointly learn dependent subjectivity through cascading classification or multitask learning, as well as models that independently predict subjectivity toward individual targets. Parinaz Sobhani, Diana Inkpen, Xiaodan Zhu 0001 |
Comput. Intell. | 3 |
| 2018 | Neural Natural Language Inference Models Enhanced with External KnowledgeabstractModeling natural language inference is a very challenging task.With the availability of large annotated data, it has recently become feasible to train complex models such as neural-network-based inference models, which have shown to achieve the state-of-the-art performance.Although there exist relatively large annotated data, can machines learn all knowledge needed to perform natural language inference (NLI) from these data?If not, how can neural-network-based NLI models benefit from external knowledge and how to build NLI models to leverage it?In this paper, we enrich the state-of-the-art neural natural language inference models with external knowledge.We demonstrate that the proposed models improve neural NLI models to achieve the state-of-the-art performance on the SNLI and MultiNLI datasets. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Diana Inkpen, Si Wei |
ACL (1) | 2 |
| 2018 | Enhancing Sentence Embedding with Generalized PoolingabstractPooling is an essential component of a wide variety of sentence representation and embedding models. This paper explores generalized pooling methods to enhance sentence embedding. We propose vector-based multi-head attention that includes the widely used max pooling, mean pooling, and scalar self-attention as special cases. The model benefits from properly designed penalization terms to reduce redundancy in multi-head attention. We evaluate the proposed model on three different tasks: natural language inference (NLI), author profiling, and sentiment classification. The experiments show that the proposed model achieves significant improvement over strong sentence-encoding-based methods, resulting in state-of-the-art performances on four datasets. The proposed approach can be easily implemented for more problems than we discuss in this paper. Qian Chen 0003, Zhen-Hua Ling, Xiaodan Zhu 0001 |
COLING | 3 |
| 2018 | Exponential Family Restricted Boltzmann Machines and Annealed Importance SamplingabstractIn this paper, we investigate restricted Boltzmann machines (RBMs) from the exponential family perspective, en-abling the visible units to follow any suitable distributions from the exponential family. We derive a unified view to compute the free energy function for exponential family RBMs (exp-RBMs). Based on that, annealed important sampling (AIS) is generalized to the entire exponential family, allowing for estimating the log-partition function and log-likelihood. Our experiments on a document processing task demonstrate that the generalized free energy functions and AIS estimation perform well in helping capture useful knowledge from the data; the estimated log-partition functions are stable. The appropriate instances of exp-RBMs can generate novel and meaningful samples and can be applied to classification tasks. Yifeng Li 0001, Xiaodan Zhu 0001 |
IJCNN | 2 |
| 2017 | Enhanced LSTM for Natural Language InferenceabstractReasoning and inference are central to human and artificial intelligence.Modeling inference in human language is very challenging.With the availability of large annotated data (Bowman et al., 2015), it has recently become feasible to train neural network based inference models, which have shown to be very effective.In this paper, we present a new state-of-the-art result, achieving the accuracy of 88.6% on the Stanford Natural Language Inference Dataset.Unlike the previous top models that use very complicated network architectures, we first demonstrate that carefully designing sequential inference models based on chain LSTMs can outperform all previous models.Based on this, we further show that by explicitly considering recursive architectures in both local inference modeling and inference composition, we achieve additional improvement.Particularly, incorporating syntactic parsing information contributes to our best result-it further improves the performance even when added to the already very strong model. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001, Diana Inkpen |
ACL (1) | 2 |
| 2017 | Cause-Effect Knowledge Acquisition and Neural Association Model for Solving A Set of Winograd Schema ProblemsabstractThis paper focuses on the investigations in Winograd Schema (WS), a challenging problem which has been proposed for measuring progress in commonsense reasoning.Due to the lack of commonsense knowledge and training data, very little work has been found on the WS problems in recent years.Actually, there is no shortcut to solve this problem except to collect more commonsense knowledge and design suitable models.Therefore, this paper addresses a set of WS problems by proposing a knowledge acquisition method and a general neural association model.To avoid the sparseness issue, the knowledge we aim to collect is the cause-effect relationships between thousands of commonly used words.The knowledge acquisition method supports us to extract hundreds of thousands of cause-effect pairs from large text corpus automatically.Meanwhile, a neural association model (NAM) is proposed to encode the association relationships between any two discrete events.Based on the extracted knowledge and the NAM models, in this paper, we successfully build a system for solving WS problems from scratch and achieve 70.0% accuracy.Most importantly, this paper provides a flexible framework to solve WS problems based on event association and neural network methods. Quan Liu 0003, Hui Jiang 0001, Andrew Evdokimov, Zhen-Hua Ling, Xiaodan Zhu 0001, Si Wei, Yu Hu 0003 |
IJCAI | 5 |
| 2016 | Extracting Discriminative Keyphrases with Learned Semantic HierarchiesabstractThe goal of keyphrase extraction is to automatically identify the most salient phrases from documents. The technique has a wide range of applications such as rendering a quick glimpse of a document, or extracting key content for further use. While previous work often assumes keyphrases are a static property of a given documents, in many applications, the appropriate set of keyphrases that should be extracted depends on the set of documents that are being considered together. In particular, good keyphrases should not only accurately describe the content of a document, but also reveal what discriminates it from the other documents. In this paper, we study this problem of extracting discriminative keyphrases. In particularly, we propose to use the hierarchical semantic structure between candidate keyphrases to promote keyphrases that have the right level of specificity to clearly distinguish the target document from others. We show that such knowledge can be used to construct better discriminative keyphrase extraction systems that do not assume a static, fixed set of keyphrases for a document. We show how this helps identify key expertise of authors from their papers, as well as competencies covered by online courses within different domains. Yunli Wang, Xiaodan Zhu 0001, Cyril Goutte |
COLING | 3 |
| 2016 | Distraction-Based Neural Networks for Modeling Document
Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001 |
IJCAI | 2 |
| 2016 | A Dataset for Detecting Stance in Tweets
Saif M. Mohammad, Svetlana Kiritchenko, Parinaz Sobhani, Xiaodan Zhu 0001, Colin Cherry |
LREC | 4 |
| 2016 | DAG-Structured Long Short-Term Memory for Semantic CompositionalityabstractRecurrent neural networks, particularly long short-term memory (LSTM), have recently shown to be very effective in a wide range of sequence modeling problems, core to which is effective learning of distributed representation for subsequences as well as the sequences they form.An assumption in almost all the previous models, however, posits that the learned representation (e.g., a distributed representation for a sentence), is fully compositional from the atomic components (e.g., representations for words), while non-compositionality is a basic phenomenon in human languages.In this paper, we relieve the assumption by extending the chain-structured LSTM to directed acyclic graphs (DAGs), with the aim to endow linear-chain LSTMs with the capability of considering compositionality together with non-compositionality in the same semantic composition framework.From a more general viewpoint, the proposed models incorporate additional prior knowledge into recurrent neural networks, which is interesting to us, considering most NLP tasks have relatively small training data and appropriate prior knowledge could be beneficial to help cover missing semantics.Our experiments on sentiment composition demonstrate that the proposed models achieve the state-of-the-art performance, outperforming models that lack this ability. Xiaodan Zhu 0001, Parinaz Sobhani |
HLT-NAACL | 1 |
| 2015 | Revisiting Word Embedding for Contrasting MeaningabstractZhigang Chen, Wei Lin, Qian Chen, Xiaoping Chen, Si Wei, Hui Jiang, Xiaodan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Zhigang Chen 0003, Qian Chen 0003, Si Wei, Hui Jiang 0001, Xiaodan Zhu 0001 |
ACL (1) | 7 |
| 2015 | Long Short-Term Memory Over Recursive StructuresabstractThe chain-structured long short-term memory (LSTM) has showed to be effective in a wide range of problems such as speech recognition and machine translation. In this paper, we propose to extend it to tree structures, in which a memory cell can reflect the history memories of multiple child cells or multiple descendant cells in a recursive process. We call the model S-LSTM, which provides a principled way of considering long-distance interaction over hierarchies, e.g., language or image parse structures. We leverage the models for semantic composition to understand the meaning of text, a fundamental problem in natural language understanding, and show that it outperforms a state-of-the-art recursive model by replacing its composition layers with the S-LSTM memory blocks. We also show that utilizing the given structures is helpful in achieving a performance better than that without considering the structures. Xiaodan Zhu 0001, Parinaz Sobhani |
ICML | 1 |
| 2015 | Sentiment, emotion, purpose, and style in electoral tweets
Saif M. Mohammad, Xiaodan Zhu 0001, Svetlana Kiritchenko, Joel D. Martin |
Inf. Process. Manag. | 2 |
| 2015 | Measuring academic influence: Not all citations are equalabstractThe importance of a research article is routinely measured by counting how many times it has been cited. However, treating all citations with equal weight ignores the wide variety of functions that citations perform. We want to automatically identify the subset of references in a bibliography that have a central academic influence on the citing paper. For this purpose, we examine the effectiveness of a variety of features for determining the academic influence of a citation. By asking authors to identify the key references in their own work, we created a data set in which citations were labeled according to their academic influence. Using automatic feature selection with supervised machine learning, we found a model for predicting academic influence that achieves good performance on this data set using only four features. The best features, among those we evaluated, were those based on the number of times a reference is mentioned in the body of a citing paper. The performance of these features inspired us to design an influence‐primed h‐index (the hip‐index). Unlike the conventional h‐index, it weights citations by how many times a reference is mentioned. According to our experiments, the hip‐index is a better indicator of researcher performance than the conventional h‐index. Xiaodan Zhu 0001, Peter D. Turney, Daniel Lemire, André Vellino |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2014 | An Empirical Study on the Effect of Negation Words on SentimentabstractNegation words, such as no and not, play a fundamental role in modifying sentiment of textual expressions. We will refer to a negation word as the negator and the text span within the scope of the negator as the argument. Commonly used heuristics to estimate the sentiment of negated expressions rely simply on the sentiment of argument (and not on the negator or the argument itself). We use a sentiment treebank to show that these existing heuristics are poor estimators of sentiment. We then modify these heuristics to be dependent on the negators and show that this improves prediction. Next, we evaluate a recently proposed composition model (Socher et al., 2013) that relies on both the negator and the argument. This model learns the syntax and semantics of the negator's argument with a recursive neural network. We show that this approach performs better than those mentioned above. In addition, we explicitly incorporate the prior sentiment of the argument and observe that this information can help reduce fitting errors. Xiaodan Zhu 0001, Saif M. Mohammad, Svetlana Kiritchenko |
ACL (1) | 1 |
| 2014 | Bilingual Sentiment Consistency for Statistical Machine TranslationabstractIn this paper, we explore bilingual sentiment knowledge for statistical machine translation (SMT). We propose to explicitly model the consistency of sentiment between the source and target side with a lexicon-based approach. The experiments show that the proposed model significantly improves Chinese-to-English NIST translation over a competitive baseline. © 2014 Association for Computational Linguistics. Boxing Chen, Xiaodan Zhu 0001 |
EACL | 2 |
| 2014 | Sentiment Analysis of Short Informal TextsabstractWe describe a state-of-the-art sentiment analysis system that detects (a) the sentiment of short informal textual messages such as tweets and SMS (message-level task) and (b) the sentiment of a word or a phrase within a message (term-level task). The system is based on a supervised statistical text classification approach leveraging a variety of surface-form, semantic, and sentiment features. The sentiment features are primarily derived from novel high-coverage tweet-specific sentiment lexicons. These lexicons are automatically generated from tweets with sentiment-word hashtags and from tweets with emoticons. To adequately capture the sentiment of words in negated contexts, a separate sentiment lexicon is generated for negated words. The system ranked first in the SemEval-2013 shared task `Sentiment Analysis in Twitter' (Task 2), obtaining an F-score of 69.02 in the message-level task and 88.93 in the term-level task. Post-competition improvements boost the performance to an F-score of 70.45 (message-level task) and 89.50 (term-level task). The system also obtains state-of-the-art performance on two additional datasets: the SemEval-2013 SMS test set and a corpus of movie review excerpts. The ablation experiments demonstrate that the use of the automatically generated lexicons results in performance gains of up to 6.5 absolute percentage points. Svetlana Kiritchenko, Xiaodan Zhu 0001, Saif M. Mohammad |
J. Artif. Intell. Res. | 2 |
| 2013 | À la Recherche du Temps Perdu: extracting temporal relations from medical text in the 2012 i2b2 NLP challengeabstractOBJECTIVE: An analysis of the timing of events is critical for a deeper understanding of the course of events within a patient record. The 2012 i2b2 NLP challenge focused on the extraction of temporal relationships between concepts within textual hospital discharge summaries. MATERIALS AND METHODS: The team from the National Research Council Canada (NRC) submitted three system runs to the second track of the challenge: typifying the time-relationship between pre-annotated entities. The NRC system was designed around four specialist modules containing statistical machine learning classifiers. Each specialist targeted distinct sets of relationships: local relationships, 'sectime'-type relationships, non-local overlap-type relationships, and non-local causal relationships. RESULTS: The best NRC submission achieved a precision of 0.7499, a recall of 0.6431, and an F1 score of 0.6924, resulting in a statistical tie for first place. Post hoc improvements led to a precision of 0.7537, a recall of 0.6455, and an F1 score of 0.6954, giving the highest scores reported on this task to date. DISCUSSION AND CONCLUSIONS: Methods for general relation extraction extended well to temporal relations, and gave top-ranked state-of-the-art results. Careful ordering of predictions within result sets proved critical to this success. Colin Cherry, Xiaodan Zhu 0001, Joel D. Martin, Berry de Bruijn |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Detecting concept relations in clinical text: Insights from a state-of-the-art model
Xiaodan Zhu 0001, Colin Cherry, Svetlana Kiritchenko, Joel D. Martin, Berry de Bruijn |
J. Biomed. Informatics | 1 |
| 2013 | A Graph-Partitioning Framework for Aligning Hierarchical Topic Structures to PresentationsabstractThis paper studies the problem of imposing an existing hierarchical semantic structure onto a corresponding spoken document in which the structures are embedded, with the goal of indexing such documents for easier access. We propose a graph-partitioning framework to solve a semantic tree-to-string alignment problem through optimizing a normalized-cut criterion. We present models with different modeling capabilities and time complexities in this framework and provide experimental evidence of their performance. We relate graph partitioning to conventional dynamic time warping (DTW) as it applies to this problem, and show that the proposed framework can naturally include topic segmentation to accommodate cohesion constraints. Xiaodan Zhu 0001, Colin Cherry, Gerald Penn |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Spotting keywords and sensing topic changes in speechabstractSecurity concerns involved in dealing with sensitive information conveyed in human languages must be able to handle speech, which is the most basic, natural form of human communication and a huge amount of data are being generated daily. Dealing with such data is naturally associated with typical big-data problems in terms of both computational complexity and storage space. Unfortunately, compared with written texts, speech is inherently more difficult to browse, if no technical support is provided. In this paper we are interested in spotting keywords, which could reflect a security agent's information needs, and study its usefulness in helping automatically disclose topic changes (boundaries) in speech data under concern. Our results show that keyword spotting can help identify topics with a competitive performance. Xiaodan Zhu 0001 |
CISDA | 1 |
| 2012 | Ecological validity and the evaluation of speech summarization qualityabstractThere is little evidence of widespread adoption of speech summarization systems. This may be due in part to the fact that the natural language heuristics used to generate summaries are often optimized with respect to a class of evaluation measures that, while computationally and experimentally inexpensive, rely on subjectively selected gold standards against which automatically generated summaries are scored. This evaluation protocol does not take into account the usefulness of a summary in assisting the listener in achieving his or her goal. In this paper we study how current measures and methods for evaluating summarization systems compare to human-centric evaluation criteria. For this, we have designed and conducted an ecologically valid evaluation that determines the value of a summary when embedded in a task, rather than how closely a summary resembles a gold standard. The results of our evaluation demonstrate that in the domain of lecture summarization, the well-known baseline of maximal marginal relevance [1] is statistically significantly worse than human-generated extractive summaries, and even worse than having no summary at all in a simple quiz-taking task. Priming seems to have no statistically significant effect on the usefulness of the human summaries. This is interesting because priming had been proposed as a technique for increasing kappa scores and/or maintaining goal orientation among summary authors. In addition, our results suggest that ROUGE scores, regardless of whether they are derived from numerically-ranked reference data or ecologically valid human-extracted summaries, may not always be reliable as inexpensive proxies for task-embedded evaluations. In fact, under some conditions, relying exclusively on ROUGE may lead to scoring human-generated summaries very favourably even when a task-embedded score calls their usefulness into question relative to using no summaries at all. Anthony McCallum, Gerald Penn, Cosmin Munteanu, Xiaodan Zhu 0001 |
SLT | 4 |
| 2011 | A Normalized-Cut Alignment Model for Mapping Hierarchical Semantic Structures onto Spoken Documents
Xiaodan Zhu 0001 |
CoNLL | 1 |
| 2011 | Indexing Spoken Documents with Hierarchical Semantic Structures: Semantic Tree-to-string Alignment Models
Xiaodan Zhu 0001, Colin Cherry, Gerald Penn |
IJCNLP | 1 |
| 2011 | Machine-learned solutions for three stages of clinical information extraction: the state of the art at i2b2 2010abstractOBJECTIVE: As clinical text mining continues to mature, its potential as an enabling technology for innovations in patient care and clinical research is becoming a reality. A critical part of that process is rigid benchmark testing of natural language processing methods on realistic clinical narrative. In this paper, the authors describe the design and performance of three state-of-the-art text-mining applications from the National Research Council of Canada on evaluations within the 2010 i2b2 challenge. DESIGN: The three systems perform three key steps in clinical information extraction: (1) extraction of medical problems, tests, and treatments, from discharge summaries and progress notes; (2) classification of assertions made on the medical problems; (3) classification of relations between medical concepts. Machine learning systems performed these tasks using large-dimensional bags of features, as derived from both the text itself and from external sources: UMLS, cTAKES, and Medline. MEASUREMENTS: Performance was measured per subtask, using micro-averaged F-scores, as calculated by comparing system annotations with ground-truth annotations on a test set. RESULTS: The systems ranked high among all submitted systems in the competition, with the following F-scores: concept extraction 0.8523 (ranked first); assertion detection 0.9362 (ranked first); relationship detection 0.7313 (ranked second). CONCLUSION: For all tasks, we found that the introduction of a wide range of features was crucial to success. Importantly, our choice of machine learning algorithms allowed us to be versatile in our feature design, and to introduce a large number of features without overfitting and without encountering computing-resource bottlenecks. Berry de Bruijn, Colin Cherry, Svetlana Kiritchenko, Joel D. Martin, Xiaodan Zhu 0001 |
J. Am. Medical Informatics Assoc. | 5 |
| 2009 | Improving Automatic Speech Recognition for Lectures through Transformation-based Rules Learned from Minimal Data
Cosmin Munteanu, Gerald Penn, Xiaodan Zhu 0001 |
ACL/IJCNLP | 3 |
| 2009 | Summarizing multiple spoken documents: finding evidence from untranscribed audio
Xiaodan Zhu 0001, Gerald Penn, Frank Rudzicz |
ACL/IJCNLP | 1 |
| 2008 | A Critical Reassessment of Evaluation Baselines for Speech Summarization
Gerald Penn, Xiaodan Zhu 0001 |
ACL | 2 |
| 2008 | Using latent Dirichlet allocation to incorporate domain knowledge for topic transition detectionabstractThis paper studies automatic detection of topic transitions for recorded presentations. This can be achieved by matching slide content with presentation transcripts directly with some similarity metrics. Such literal matching, however, misses domain-specific knowledge and is sensitive to speech recognition errors. In this paper, we incorporate relevant written materials, e.g., textbooks for lectures, which convey semantic relationships, in particular domain-specific relationships, between words. To this end, we train latent Dirichlet allocation (LDA) models on these materials and measure the similarity between slides and transcripts in the acquired hidden-topic space. This similarity is then combined with literal matchings. Experiments show that the proposed approach reduces the errors in slide transition detection by 17-41 % on manual transcripts and 27-37% on automatic transcripts. Index Terms: slides transition detection, boundary detection. 1. Xiaodan Zhu 0001, Xuming He 0001, Cosmin Munteanu, Gerald Penn |
INTERSPEECH | 1 |
| 2008 | Identifying salient utterances of online spoken documents using descriptive hypertextabstractThe Internet has become an important supply channel of spoken documents. Efficient ways of navigating their content are highly desirable. This paper aims to identify the most salient utterances from online spoken documents using relevant hypertext that encapsulates key information. Experimental results show that hypertext features are helpful when properly utilized and if the bit rates used to compress the spoken documents are reasonable. Xiaodan Zhu 0001, Siavash Kazemian, Gerald Penn |
SLT | 1 |
| 2006 | Using Outcome Polarity in Sentence Extraction for Medical Question-Answering
Yun Niu, Xiaodan Zhu 0001, Graeme Hirst |
AMIA | 2 |
| 2006 | Utterance-Level Extractive Summarization of Open-Domain Spontaneous Conversations with Rich FeaturesabstractTo identify important utterances from open-domain spontaneous conversations, previous work has focused on using textual features that are extracted from transcripts, e.g., word frequencies and noun senses. In this paper, we summarize spontaneous conversations with features of a wide variety that have not been explored before. Experiments show that the use of speech-related features improves summarization performance. In addition, the effectiveness of individual features is examined and compared Xiaodan Zhu 0001, Gerald Penn |
ICME | 1 |
| 2006 | Summarization of spontaneous conversationsabstractSpontaneous conversations are an integral element in many CSCW environments. Although speech is often regarded as the most natural and effective way of communication between human beings, speech data are not efficient for quick review. One solution to help people access speech data efficiently in CSCW environments is to conduct speech summarization. Up till now, most speech summarization research has focused on broadcast news; nevertheless summarizing spontaneous conversations is more valuable for CSCW. The task is also more challenging, for example, spontaneous conversations often contain more speech disfluencies, which need to be coped with properly; they are also more vulnerable to speech recognition errors. This demonstration is built to show the prototype of our summarization system. Compared with previous work, our summarizer addresses the problem further in several important respects. First, the system summarizes spontaneous conversations with a wide variety of information/features that have not been explored before, which improve summarization performance according to our experiments. Second, our summarizer handles speech disfluencies, which in all previous work was either not explicitly handled or removed as noise. Xiaodan Zhu 0001, Gerald Penn |
INTERSPEECH | 1 |
| 2006 | Comparing the roles of textual, acoustic and spoken-language features on spontaneous-conversation summarization
Xiaodan Zhu 0001, Gerald Penn |
HLT-NAACL | 1 |
| 2005 | Analysis of Polarity Information in Medical Text
Yun Niu, Xiaodan Zhu 0001, Graeme Hirst |
AMIA | 2 |