Xiaoyan Zhu 0001

dblp:50/1222-1 · DBLP profile ↗
← Back
114ranked-venue papers
1as first author
11since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 77 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 29Applied, interdisciplinary, general and emerging computing · 25Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Theory of computation · 1
YearPublicationVenuePosition
2024 Maximum Gaussianality training for deep speaker vector normalization
Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013
Pattern Recognit.4
2023 KPT: Keyword-Guided Pre-training for Grounded Dialog Generation
abstract
Incorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.r.t. different types of knowledge. In this work, we propose KPT (Keyword-guided Pre-Training), a novel self-supervised pre-training method for grounded dialog generation without relying on extra knowledge annotation. Specifically, we use a pre-trained language model to extract the most uncertain tokens in the dialog as keywords. With these keywords, we construct two kinds of knowledge and pre-train a knowledge-grounded response generation model, aiming at handling two different scenarios: (1) the knowledge should be faithfully grounded; (2) it can be selectively used. For the former, the grounding knowledge consists of keywords extracted from the response. For the latter, the grounding knowledge is additionally augmented with keywords extracted from other utterances in the same dialog. Since the knowledge is extracted from the dialog itself, KPT can be easily performed on a large volume and variety of dialogue data. We considered three data sources (open-domain, task-oriented, conversational QA) with a total of 2.5M dialogues. We conduct extensive experiments on various few-shot knowledge-grounded generation tasks, including grounding on dialog acts, knowledge graphs, persona descriptions, and Wikipedia passages. Our comprehensive experiments and analyses demonstrate that KPT consistently outperforms state-of-the-art methods on these tasks with diverse grounding knowledge.
Qi Zhu 0007, Fei Mi, Zheng Zhang 0020, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang
AAAI8
2023 DecompEval: Evaluating Generated Texts as Unsupervised Decomposed Question Answering
abstract
Existing evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability.Specifically, most of the wellperformed metrics are required to train on evaluation datasets of specific NLG tasks and evaluation dimensions, which may cause over-fitting to task-specific datasets.Furthermore, existing metrics only provide an evaluation score for each dimension without revealing the evidence to interpret how this score is obtained.To deal with these challenges, we propose a simple yet effective metric called DecompEval.This metric formulates NLG evaluation as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models (PLMs) without training on evaluation datasets, aiming to enhance the generalization ability.To make the evaluation process more interpretable, we decompose our devised instruction-style question about the quality of generated texts into the subquestions that measure the quality of each sentence.The subquestions with their answers generated by PLMs are then recomposed as evidence to obtain the evaluation result.Experimental results show that DecompEval achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which also exhibits strong dimension-level / task-level generalization ability and interpretability 1 .
Pei Ke, Fei Huang 0005, Fei Mi, Yasheng Wang, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang
ACL (1)6
2023 Building Multi-domain Dialog State Trackers from Single-domain Dialogs
abstract
Existing multi-domain dialog state tracking (DST) models are developed based on multidomain dialogs, which require significant manual effort to define domain relations and collect data.This process can be challenging and expensive, particularly when numerous domains are involved.In this paper, we propose a divide-and-conquer (DAC) DST paradigm and a multi-domain dialog synthesis framework, which makes building multi-domain DST models from single-domain dialogs possible.The DAC paradigm segments a multi-domain dialog into multiple single-domain dialogs for DST, which makes models generalize better on dialogs involving unseen domain combinations.The multi-domain dialog synthesis framework merges several potentially related single-domain dialogs into one multi-domain dialog and modifies the dialog to simulate domain relations.The synthesized dialogs can help DST models capture the value transfer between domains.Experiments with three representative DST models on two datasets demonstrate the effectiveness of our proposed DAC paradigm and data synthesis framework.
Qi Zhu 0007, Zheng Zhang 0020, Xiaoyan Zhu 0001, Minlie Huang
EMNLP3
2022 Continual Prompt Tuning for Dialog State Tracking
abstract
A desirable dialog system should be able to continually learn new skills without forgetting old ones, and thereby adapt to new domains or tasks in its life cycle.However, continually training a model often leads to a well-known catastrophic forgetting issue.In this paper, we present Continual Prompt Tuning, a parameterefficient framework that not only avoids forgetting but also enables knowledge transfer between tasks.To avoid forgetting, we only learn and store a few prompt tokens' embeddings for each task while freezing the backbone pre-trained model.To achieve bi-directional knowledge transfer among tasks, we propose several techniques (continual prompt initialization, query fusion, and memory replay) to transfer knowledge from preceding tasks and a memory-guided technique to transfer knowledge from subsequent tasks.Extensive experiments demonstrate the effectiveness and efficiency of our proposed method on continual learning for dialog state tracking, compared with state-of-the-art baselines.
Qi Zhu 0007, Fei Mi, Xiaoyan Zhu 0001, Minlie Huang
ACL (1)4
2022 CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation
abstract
Existing reference-free metrics have obvious limitations for evaluating controlled text generation models.Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgments, whereas supervised ones may overfit task-specific data with poor generalization ability to other datasets.In this paper, we propose an unsupervised reference-free metric called CTRLEval, which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks.On top of these tasks, the metric assembles the generation probabilities from a pre-trained language model without any model training.Experimental results show that our metric has higher correlations with human judgments than other baselines, while obtaining better generalization of evaluating generated texts from different models and with different qualities 1 .
Pei Ke, Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xiaoyan Zhu 0001, Minlie Huang
ACL (1)6
2022 Learning Instructions with Unlabeled Data for Zero-Shot Cross-Task Generalization
abstract
Training language models to learn from human instructions for zero-shot cross-task generalization has attracted much attention in NLP communities.Recently, instruction tuning (IT), which fine-tunes a pre-trained language model on a massive collection of tasks described via human-craft instructions, has been shown effective in instruction learning for unseen tasks.However, IT relies on a large amount of humanannotated samples, which restricts its generalization.Unlike labeled data, unlabeled data are often massive and cheap to obtain.In this work, we study how IT can be improved with unlabeled data.We first empirically explore the IT performance trends versus the number of labeled data, instructions, and training tasks.We find it critical to enlarge the number of training instructions, and the instructions can be underutilized due to the scarcity of labeled data.Then, we propose Unlabeled Data Augmented Instruction Tuning (UDIT) to take better advantage of the instructions during IT by constructing pseudo-labeled data from unlabeled plain texts.We conduct extensive experiments to show UDIT's effectiveness in various scenarios of tasks and datasets.We also comprehensively analyze the key factors of UDIT to investigate how to better improve IT with unlabeled data.
Yuxian Gu, Pei Ke, Xiaoyan Zhu 0001, Minlie Huang
EMNLP3
2022 Curriculum-Based Self-Training Makes Better Few-Shot Learners for Data-to-Text Generation
abstract
Despite the success of text-to-text pre-trained models in various natural language generation (NLG) tasks, the generation performance is largely restricted by the number of labeled data in downstream tasks, particularly in data-to-text generation tasks. Existing works mostly utilize abundant unlabeled structured data to conduct unsupervised pre-training for task adaption, which fail to model the complex relationship between source structured data and target texts. Thus, we introduce self-training as a better few-shot learner than task-adaptive pre-training, which explicitly captures this relationship via pseudo-labeled data generated by the pre-trained model. To alleviate the side-effect of low-quality pseudo-labeled data during self-training, we propose a novel method called Curriculum-Based Self-Training (CBST) to effectively leverage unlabeled data in a rearranged order determined by the difficulty of text generation. Experimental results show that our method can outperform fine-tuning and task-adaptive pre-training methods, and achieve state-of-the-art performance in the few-shot setting of data-to-text generation.
Pei Ke, Haozhe Ji, Yi Huang 0017, Junlan Feng, Xiaoyan Zhu 0001, Minlie Huang
IJCAI6
2021 A Semantic-based Method for Unsupervised Commonsense Question Answering
abstract
Yilin Niu, Fei Huang, Jiaming Liang, Wenkai Chen, Xiaoyan Zhu, Minlie Huang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yilin Niu, Fei Huang 0005, Xiaoyan Zhu 0001, Minlie Huang
ACL/IJCNLP (1)5
2021 EARL: Informative Knowledge-Grounded Conversation Generation with Entity-Agnostic Representation Learning
abstract
Generating informative and appropriate responses is challenging but important for building human-like dialogue systems.Although various knowledge-grounded conversation models have been proposed, these models have limitations in utilizing knowledge that infrequently occurs in the training data, not to mention integrating unseen knowledge into conversation generation.In this paper, we propose an Entity-Agnostic Representation Learning (EARL) method to introduce knowledge graphs to informative conversation generation.Unlike traditional approaches that parameterize the specific representation for each entity, EARL utilizes the context of conversations and the relational structure of knowledge graphs to learn the category representation for entities, which is generalized to incorporating unseen entities in knowledge graphs into conversation generation.Automatic and manual evaluations demonstrate that our model can generate more informative, coherent, and natural responses than baseline models.
Hao Zhou 0012, Minlie Huang, Wei Chen 0034, Xiaoyan Zhu 0001
EMNLP (1)5
2021 Deep Normalization for Speaker Vectors
abstract
Deep speaker embedding has demonstrated state-of-the-art performance in speaker recognition tasks. However, one potential issue with this approach is that the speaker vectors derived from deep embedding models tend to be non-Gaussian for each individual speaker, and non-homogeneous for distributions of different speakers. These irregular distributions can seriously impact speaker recognition performance, especially with the popular PLDA scoring method, which assumes homogeneous Gaussian distribution. In this article, we argue that deep speaker vectors require deep normalization, and propose a deep normalization approach based on a novel discriminative normalization flow (DNF) model. We demonstrate the effectiveness of the proposed approach with experiments using the widely used SITW and CNCeleb corpora. In these experiments, the DNF-based normalization delivered substantial performance gains and also showed strong generalization capability in out-of-domain tests.
Yunqi Cai, Lantian Li, Andrew Abel, Xiaoyan Zhu 0001, Dong Wang 0013
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven Conversation
abstract
The research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consists of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese multi-domain knowledge-driven conversation dataset, KdConv, which grounds the topics in multi-turn conversations to knowledge graphs. Our corpus contains 4.5K conversations from three domains (film, music, and travel), and 86K utterances with an average turn number of 19.0. These conversations contain in-depth discussions on related topics and natural transition between multiple topics. To facilitate the following research on this corpus, we provide several benchmark models. Comparative results show that the models can be enhanced by introducing background knowledge, yet there is still a large space for leveraging knowledge to model multi-turn conversations for further research. Results also show that there are obvious performance differences between different domains, indicating that it is worth further explore transfer learning and domain adaptation. The corpus and benchmark models are publicly available.
Hao Zhou 0012, Chujie Zheng, Kaili Huang, Minlie Huang, Xiaoyan Zhu 0001
ACL5
2020 Language Generation with Multi-Hop Reasoning on Commonsense Knowledge Graph
abstract
Despite the success of generative pre-trained language models on a series of text generation tasks, they still suffer in cases where reasoning over underlying commonsense knowledge is required during generation.Existing approaches that integrate commonsense knowledge into generative pre-trained language models simply transfer relational knowledge by post-training on individual knowledge triples while ignoring rich connections within the knowledge graph.We argue that exploiting both the structural and semantic information of the knowledge graph facilitates commonsenseaware text generation.In this paper, we propose Generation with Multi-Hop Reasoning Flow (GRF) that enables pre-trained models with dynamic multi-hop reasoning on multirelational paths extracted from the external commonsense knowledge graph.We empirically show that our model outperforms existing baselines on three text generation tasks that require reasoning over commonsense knowledge.We also demonstrate the effectiveness of the dynamic multi-hop reasoning module with reasoning paths inferred by the model that provide rationale to the generation. 1
Haozhe Ji, Pei Ke, Shaohan Huang, Furu Wei, Xiaoyan Zhu 0001, Minlie Huang
EMNLP (1)5
2020 SentiLARE: Sentiment-Aware Language Representation Learning with Linguistic Knowledge
abstract
Most of the existing pre-trained language representation models neglect to consider the linguistic knowledge of texts, which can promote language understanding in NLP tasks.To benefit the downstream tasks in sentiment analysis, we propose a novel language representation model called SentiLARE, which introduces word-level linguistic knowledge including part-of-speech tag and sentiment polarity (inferred from SentiWordNet) into pretrained models.We first propose a contextaware sentiment attention mechanism to acquire the sentiment polarity of each word with its part-of-speech tag by querying SentiWord-Net.Then, we devise a new pre-training task called label-aware masked language model to construct knowledge-aware language representation.Experiments show that SentiLARE obtains new state-of-the-art performance on a variety of sentiment analysis tasks 1 .
Pei Ke, Haozhe Ji, Siyang Liu 0003, Xiaoyan Zhu 0001, Minlie Huang
EMNLP (1)4
2020 A Large-Scale Chinese Short-Text Conversation Dataset
Yida Wang 0009, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang 0001, Xiaoyan Zhu 0001, Minlie Huang
NLPCC (1)6
2020 A Knowledge-Enhanced Pretraining Model for Commonsense Story Generation
abstract
Story generation, namely, generating a reasonable story from a leading context, is an important but challenging task. In spite of the success in modeling fluency and local coherence, existing neural language generation models (e.g., GPT-2) still suffer from repetition, logic conflicts, and lack of long-range coherence in generated stories. We conjecture that this is because of the difficulty of associating relevant commonsense knowledge, understanding the causal relationships, and planning entities and events with proper temporal order. In this paper, we devise a knowledge-enhanced pretraining model for commonsense story generation. We propose to utilize commonsense knowledge from external knowledge bases to generate reasonable stories. To further capture the causal and temporal dependencies between the sentences in a reasonable story, we use multi-task learning, which combines a discriminative objective to distinguish true and fake stories during fine-tuning. Automatic and manual evaluation shows that our model can generate more reasonable stories than state-of-the-art baselines, particularly in terms of logic and global coherence.
Jian Guan 0002, Fei Huang 0005, Minlie Huang, Xiaoyan Zhu 0001
Trans. Assoc. Comput. Linguistics5
2020 CrossWOZ: A Large-Scale Chinese Cross-Domain Task-Oriented Dialogue Dataset
abstract
To advance multi-domain (cross-domain) dialogue modeling as well as alleviate the shortage of Chinese task-oriented datasets, we propose CrossWOZ, the first large-scale Chinese Cross-Domain Wizard-of-Oz task-oriented dataset. It contains 6K dialogue sessions and 102K utterances for 5 domains, including hotel, restaurant, attraction, metro, and taxi. Moreover, the corpus contains rich annotation of dialogue states and dialogue acts on both user and system sides. About 60% of the dialogues have cross-domain user goals that favor inter-domain dependency and encourage natural transition across domains in conversation. We also provide a user simulator and several benchmark models for pipelined task-oriented dialogue systems, which will facilitate researchers to compare and evaluate their models on this corpus. The large size and rich annotation of CrossWOZ make it suitable to investigate a variety of tasks in cross-domain dialogue modeling, such as dialogue state tracking, policy learning, user simulation, etc.
Qi Zhu 0007, Kaili Huang, Zheng Zhang 0020, Xiaoyan Zhu 0001, Minlie Huang
Trans. Assoc. Comput. Linguistics4
2020 Robust Reading Comprehension With Linguistic Constraints via Posterior Regularization
abstract
In spite of the great advancements of machine reading comprehension (RC), existing RC models are still vulnerable and not robust to different types of adversarial examples. Neural models over-confidently predict wrong answers to semantic different adversarial examples, while over-sensitively predict wrong answers to semantic equivalent adversarial examples. Existing methods which improve the robustness of such neural models merely mitigate one of the two issues but ignore the other. In this article, we address the over-confidence issue and the over-sensitivity issue existing in current RC models simultaneously with the help of external linguistic knowledge. We first incorporate external knowledge to impose different linguistic constraints (entity constraint, lexical constraint, and predicate constraint), and then regularize RC models through posterior regularization. Linguistic constraints induce more reasonable predictions for both semantic different and semantic equivalent adversarial examples, and posterior regularization provides an effective mechanism to incorporate these constraints. Our method can be applied to any existing neural RC models including state-of-the-art BERT models. Extensive experiments show that our method remarkably improves the robustness of base RC models, and is better to cope with these two issues simultaneously.
Mantong Zhou, Minlie Huang, Xiaoyan Zhu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Challenges in Building Intelligent Open-domain Dialog Systems
abstract
There is a resurgent interest in developing intelligent open-domain dialog systems due to the availability of large amounts of conversational data and the recent progress on neural approaches to conversational AI [33]. Unlike traditional task-oriented bots, an open-domain dialog system aims to establish long-term connections with users by satisfying the human need for communication, affection, and social belonging. This article reviews the recent work on neural approaches that are devoted to addressing three challenges in developing such systems: semantics , consistency , and interactiveness . Semantics requires a dialog system to not only understand the content of the dialog but also identify users’ emotional and social needs during the conversation. Consistency requires the system to demonstrate a consistent personality to win users’ trust and gain their long-term confidence. Interactiveness refers to the system’s ability to generate interpersonal responses to achieve particular social goals such as entertainment and conforming. The studies we select to present in this survey are based on our unique views and are by no means complete. Nevertheless, we hope that the discussion will inspire new research in developing more intelligent open-domain dialog systems.
Minlie Huang, Xiaoyan Zhu 0001, Jianfeng Gao 0001
ACM Trans. Inf. Syst.2
2019 ARAML: A Stable Adversarial Training Framework for Text Generation
abstract
Pei Ke, Fei Huang, Minlie Huang, Xiaoyan Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pei Ke, Fei Huang 0005, Minlie Huang, Xiaoyan Zhu 0001
EMNLP/IJCNLP (1)4
2019 Long and Diverse Text Generation with Planning-based Hierarchical Variational Model
abstract
Zhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu, Xiaoyan Zhu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhihong Shao, Minlie Huang, Jiangtao Wen, Wenfei Xu, Xiaoyan Zhu 0001
EMNLP/IJCNLP (1)5
2019 Aspect-level Sentiment Analysis using AS-Capsules
abstract
Aspect-level sentiment analysis aims to provide complete and detailed view of sentiment analysis from different aspects. Existing solutions usually adopt a two-staged approach: first detecting aspect category in a document, then categorizing the polarity of opinion expressions for detected aspect(s). Inevitably, such methods lead to error accumulation. Moreover, aspect detection and aspect-level sentiment classification are highly correlated with each other. The key issue here is how to perform aspect detection and aspect-level sentiment classification jointly, and effectively. In this paper, we propose the aspect-level sentiment capsules model (AS-Capsules), which is capable of performing aspect detection and sentiment classification simultaneously, in a joint manner. AS-Capsules utilizes the correlation between aspect and sentiment through shared components including capsule embedding, shared encoders, and shared attentions. AS-Capsules is also capable of communicating with different capsules through a shared Recurrent Neural Network (RNN). More importantly, AS-Capsules model does not require any linguistic knowledge as additional input. Instead, through the attention mechanism, this model is able to attend aspect related words and sentiment words corresponding to different aspect(s). Experiments show that the AS-Capsules model achieves state-of-the-art performances on a benchmark dataset for aspect-level sentiment analysis.
Yequan Wang, Aixin Sun, Minlie Huang, Xiaoyan Zhu 0001
WWW4
2019 Neural Multimodal Belief Tracker with Adaptive Attention for Dialogue Systems
abstract
Multimodal dialogue systems are attracting increasing attention with a more natural and informative way for human-computer interaction. As one of its core components, the belief tracker estimates the user's goal at each step of the dialogue and provides a direct way to validate the ability of dialogue understanding. However, existing studies on belief trackers are largely limited to textual modality, which cannot be easily extended to capture the rich semantics in multimodal systems such as those with product images. For example, in fashion domain, the visual appearance of clothes play a crucial role in understanding the user's intention. In this case, the existing belief trackers may fail to generate accurate belief states for a multimodal dialogue system.
Zheng Zhang 0020, Lizi Liao, Minlie Huang, Xiaoyan Zhu 0001, Tat-Seng Chua
WWW4
2019 Domain-Constrained Advertising Keyword Generation
abstract
Advertising (ad for short) keyword suggestion is important for sponsored search to improve online advertising and increase search revenue. There are two common challenges in this task. First, the keyword bidding problem: hot ad keywords are very expensive for most of the advertisers because more advertisers are bidding on more popular keywords, while unpopular keywords are difficult to discover. As a result, most ads have few chances to be presented to the users. Second, the inefficient ad impression issue: a large proportion of search queries, which are unpopular yet relevant to many ad keywords, have no ads presented on their search result pages. Existing retrieval-based or matching-based methods either deteriorate the bidding competition or are unable to suggest novel keywords to cover more queries, which leads to inefficient ad impressions.
Hao Zhou 0012, Minlie Huang, Yishun Mao, Changlei Zhu, Peng Shu, Xiaoyan Zhu 0001
WWW6
2019 Story Ending Selection by Finding Hints From Pairwise Candidate Endings
abstract
The ability of story comprehension is a strong indicator of natural language understanding. Recently, Story Cloze Test has been introduced as a new task of machine reading comprehension, i.e., selecting a correct ending from two candidate endings given a four-sentence story context. Most existing methods for Story Cloze Test are essentially matching-based that operate by comparing an individual ending with a given context, therefore suffering from the evidence bias issue: both candidate endings can obtain supporting evidence from the story context, which misleads the classifier to choose an incorrect ending. To address this issue, we present a novel idea to improve story comprehension by utilizing the hints that are obtained through comparing two candidate endings. The proposed model firstly anticipates a feature vector for a possible ending solely based on the context, and then refines the feature prediction using the hints which encode the difference between two candidates. The candidate ending whose feature vector is more similar to the predicted ending vector is regarded as correct. Experimental results demonstrate that our approach can alleviate the evidence bias issue and improve story comprehension.
Mantong Zhou, Minlie Huang, Xiaoyan Zhu 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Memory-Augmented Dialogue Management for Task-Oriented Dialogue Systems
abstract
Dialogue management (DM) is responsible for predicting the next action of a dialogue system according to the current dialogue state and thus plays a central role in task-oriented dialogue systems. Since DM requires having access not only to local utterances but also to the global semantics of the entire dialogue session, modeling the long-range history information is a critical issue. To this end, we propose MAD, a novel memory-augmented dialogue management model that employs a memory controller and two additional memory structures (i.e., a slot-value memory and an external memory). The slot-value memory tracks the dialogue state by memorizing and updating the values of semantic slots (i.e., cuisine, price, and location), and the external memory augments the representation of hidden states of traditional recurrent neural networks by storing more context information. To update the dialogue state efficiently, we also propose slot-level attention on user utterances to extract specific semantic information for each slot. Experiments show that our model can obtain state-of-the-art performance and outperforms existing baselines.
Zheng Zhang 0020, Minlie Huang, Zhongzhou Zhao, Haiqing Chen, Xiaoyan Zhu 0001
ACM Trans. Inf. Syst.6
2018 Reinforcement Learning for Relation Classification From Noisy Data
abstract
Existing relation classification methods that rely on distant supervision assume that a bag of sentences mentioning an entity pair are all describing a relation for the entity pair. Such methods, performing classification at the bag level, cannot identify the mapping between a relation and a sentence, and largely suffers from the noisy labeling problem. In this paper, we propose a novel model for relation classification at the sentence level from noisy data. The model has two modules: an instance selector and a relation classifier. The instance selector chooses high-quality sentences with reinforcement learning and feeds the selected sentences into the relation classifier, and the relation classifier makes sentence-level prediction and provides rewards to the instance selector. The two modules are trained jointly to optimize the instance selection and relation classification processes.Experiment results show that our model can deal with the noise of data effectively and obtains better performance for relation classification at the sentence level.
Minlie Huang, Li Zhao 0007, Yang Yang 0009, Xiaoyan Zhu 0001
AAAI5
2018 Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory
abstract
Perception and expression of emotion are key factors to the success of dialogue systems or conversational agents. However, this problem has not been studied in large-scale conversation generation so far. In this paper, we propose Emotional Chatting Machine (ECM) that can generate appropriate responses not only in content (relevant and grammatical) but also in emotion (emotionally consistent). To the best of our knowledge, this is the first work that addresses the emotion factor in large-scale conversation generation. ECM addresses the factor using three new mechanisms that respectively (1) models the high-level abstraction of emotion expressions by embedding emotion categories, (2) captures the change of implicit internal emotion states, and (3) uses explicit emotion expressions with an external emotion vocabulary. Experiments show that the proposed model can generate responses appropriate not only in content but also in emotion.
Hao Zhou 0012, Minlie Huang, Xiaoyan Zhu 0001, Bing Liu 0001
AAAI4
2018 Generating Informative Responses with Controlled Sentence Function
abstract
Sentence function is a significant factor to achieve the purpose of the speaker, which, however, has not been touched in largescale conversation generation so far.In this paper, we present a model to generate informative responses with controlled sentence function.Our model utilizes a continuous latent variable to capture various word patterns that realize the expected sentence function, and introduces a type controller to deal with the compatibility of controlling sentence function and generating informative content.Conditioned on the latent variable, the type controller determines the type (i.e., function-related, topic, and ordinary word) of a word to be generated at each decoding position.Experiments show that our model outperforms state-of-the-art baselines, and it has the ability to generate responses with both controlled sentence function and informative content.
Pei Ke, Jian Guan 0002, Minlie Huang, Xiaoyan Zhu 0001
ACL (1)4
2018 An Operation Network for Abstractive Sentence Compression
abstract
Sentence compression condenses a sentence while preserving its most important contents. Delete-based models have the strong ability to delete undesired words, while generate-based models are able to reorder or rephrase the words, which are more coherent to human sentence compression. In this paper, we propose Operation Network, a neural network approach for abstractive sentence compression, which combines the advantages of both delete-based and generate-based sentence compression models. The central idea of Operation Network is to model the sentence compression process as an editing procedure. First, unnecessary words are deleted from the source sentence, then new words are either generated from a large vocabulary or copied directly from the source sentence. A compressed sentence can be obtained by a series of such edit operations (delete, copy and generate). Experiments show that Operation Network outperforms state-of-the-art baselines.
Naitong Yu, Jie Zhang 0002, Minlie Huang, Xiaoyan Zhu 0001
COLING4
2018 An Interpretable Reasoning Network for Multi-Relation Question Answering
abstract
Multi-relation Question Answering is a challenging task, due to the requirement of elaborated analysis on questions and reasoning over multiple fact triples in knowledge base. In this paper, we present a novel model called Interpretable Reasoning Network that employs an interpretable, hop-by-hop reasoning process for question answering. The model dynamically decides which part of an input question should be analyzed at each hop; predicts a relation that corresponds to the current parsed results; utilizes the predicted relation to update the question representation and the state of the reasoning process; and then drives the next-hop reasoning. Experiments show that our model yields state-of-the-art results on two datasets. More interestingly, the model can offer traceable and observable intermediate predictions for reasoning analysis and failure diagnosis, thereby allowing manual manipulation in predicting the final answer.
Mantong Zhou, Minlie Huang, Xiaoyan Zhu 0001
COLING3
2018 Assigning Personality/Profile to a Chatting Machine for Coherent Conversation Generation
abstract
Endowing a chatbot with personality is challenging but significant to deliver more realistic and natural conversations. In this paper, we address the issue of generating responses that are coherent to a pre-specified personality or profile. We present a method that uses generic conversation data from social media (without speaker identities) to generate profile-coherent responses. The central idea is to detect whether a profile should be used when responding to a user post (by a profile detector), and if necessary, select a key-value pair from the profile to generate a response forward and backward (by a bidirectional decoder) so that a personality-coherent response can be generated. Furthermore, in order to train the bidirectional decoder with generic dialogue data, a position detector is designed to predict a word position from which decoding should start given a profile value. Manual and automatic evaluation shows that our model can deliver more coherent, natural, and diversified responses.
Qiao Qian, Minlie Huang, Haizhou Zhao, Jingfang Xu, Xiaoyan Zhu 0001
IJCAI5
2018 A Weakly Supervised Method for Topic Segmentation and Labeling in Goal-oriented Dialogues via Reinforcement Learning
abstract
Topic structure analysis plays a pivotal role in dialogue understanding. We propose a reinforcement learning (RL) method for topic segmentation and labeling in goal-oriented dialogues, which aims to detect topic boundaries among dialogue utterances and assign topic labels to the utterances. We address three common issues in the goal-oriented customer service dialogues: informality, local topic continuity, and global topic structure. We explore the task in a weakly supervised setting and formulate it as a sequential decision problem. The proposed method consists of a state representation network to address the informality issue, and a policy network with rewards to model local topic continuity and global topic structure. To train the two networks and offer a warm-start to the policy, we firstly use some keywords to annotate the data automatically. We then pre-train the networks on noisy data. Henceforth, the method continues to refine the data labels using the current policy to learn better state representations on the refined data for obtaining a better policy. Results demonstrate that this weakly supervised method obtains substantial improvements over state-of-the-art baselines.
Ryuichi Takanobu, Minlie Huang, Zhongzhou Zhao, Feng-Lin Li, Haiqing Chen, Xiaoyan Zhu 0001, Liqiang Nie
IJCAI6
2018 Commonsense Knowledge Aware Conversation Generation with Graph Attention
abstract
Commonsense knowledge is vital to many natural language processing tasks. In this paper, we present a novel open-domain conversation generation model to demonstrate how large-scale commonsense knowledge can facilitate language understanding and generation. Given a user post, the model retrieves relevant knowledge graphs from a knowledge base and then encodes the graphs with a static graph attention mechanism, which augments the semantic information of the post and thus supports better understanding of the post. Then, during word generation, the model attentively reads the retrieved knowledge graphs and the knowledge triples within each graph to facilitate better generation through a dynamic graph attention mechanism. This is the first attempt that uses large-scale commonsense knowledge in conversation generation. Furthermore, unlike existing models that use knowledge triples (entities) separately and independently, our model treats each knowledge graph as a whole, which encodes more structured, connected semantic information in the graphs. Experiments show that the proposed model can generate more appropriate and informative responses than state-of-the-art baselines.
Hao Zhou 0012, Tom Young, Minlie Huang, Haizhou Zhao, Jingfang Xu, Xiaoyan Zhu 0001
IJCAI6
2018 Learning to Collaborate: Multi-Scenario Ranking via Multi-Agent Reinforcement Learning
abstract
Ranking is a fundamental and widely studied problem in scenarios such as search, advertising, and recommendation. However, joint optimization for multi-scenario ranking, which aims to improve the overall performance of several ranking strategies in different scenarios, is rather untouched. Separately optimizing each individual strategy has two limitations. The first one is lack of collaboration between scenarios meaning that each strategy maximizes its own objective but ignores the goals of other strategies, leading to a sub-optimal overall performance. The second limitation is the inability of modeling the correlation between scenarios meaning that independent optimization in one scenario only uses its own user data but ignores the context in other scenarios. In this paper, we formulate multi-scenario ranking as a fully cooperative, partially observable, multi-agent sequential decision problem. We propose a novel model named Multi-Agent Recurrent Deterministic Policy Gradient (MA-RDPG) which has a communication component for passing messages, several private actors (agents) for making actions for ranking, and a centralized critic for evaluating the overall performance of the co-working actors. Each scenario is treated as an agent (actor). Agents collaborate with each other by sharing a global action-value function (the critic) and passing messages that encodes historical information across scenarios. The model is evaluated with online settings on a large E-commerce platform. Results show that the proposed model exhibits significant improvements against baselines in terms of the overall performance.
Minlie Huang, Shichen Liu, Wenwu Ou, Zhirong Wang, Xiaoyan Zhu 0001
WWW7
2018 Sentiment Analysis by Capsules
abstract
In this paper, we propose RNN-Capsule, a capsule model based on Recurrent Neural Network (RNN) for sentiment analysis. For a given problem, one capsule is built for each sentiment category e.g., 'positive' and 'negative'. Each capsule has an attribute, a state, and three modules: representation module, probability module, and reconstruction module. The attribute of a capsule is the assigned sentiment category. Given an instance encoded in hidden vectors by a typical RNN, the representation module builds capsule representation by the attention mechanism. Based on capsule representation, the probability module computes the capsule's state probability. A capsule's state is active if its state probability is the largest among all capsules for the given instance, and inactive otherwise. On two benchmark datasets (i.e., Movie Review and Stanford Sentiment Treebank) and one proprietary dataset (i.e., Hospital Feedback), we show that RNN-Capsule achieves state-of-the-art performance on sentiment classification. More importantly, without using any linguistic knowledge, RNN-Capsule is capable of outputting words with sentiment tendencies reflecting capsules' attributes. The words well reflect the domain specificity of the dataset.
Yequan Wang, Aixin Sun, Jialong Han, Ying Liu 0004, Xiaoyan Zhu 0001
WWW5
2017 SSP: Semantic Space Projection for Knowledge Graph Embedding with Text Descriptions
abstract
Knowledge graph embedding represents entities and relations in knowledge graph as low-dimensional, continuous vectors, and thus enables knowledge graph compatible with machine learning models. Though there have been a variety of models for knowledge graph embedding, most methods merely concentrate on the fact triples, while supplementary textual descriptions of entities and relations have not been fully employed. To this end, this paper proposes the semantic space projection (SSP) model which jointly learns from the symbolic triples and textual descriptions. Our model builds interaction between the two information sources, and employs textual descriptions to discover semantic relevance and offer precise semantic embedding. Extensive experiments show that our method achieves substantial improvements against baselines on the tasks of knowledge graph completion and entity classification.
Han Xiao 0005, Minlie Huang, Lian Meng, Xiaoyan Zhu 0001
AAAI4
2017 Linguistically Regularized LSTM for Sentiment Classification
abstract
This paper deals with sentence-level sentiment classification.Though a variety of neural network models have been proposed recently, however, previous models either depend on expensive phrase-level annotation, most of which has remarkably degraded performance when trained with only sentence-level annotation; or do not fully employ linguistic resources (e.g., sentiment lexicons, negation words, intensity words).In this paper, we propose simple models trained with sentence-level annotation, but also attempt to model the linguistic role of sentiment lexicons, negation words, and intensity words.Results show that our models are able to capture the linguistic role of sentiment words, negation words, and intensity words in sentiment expression.
Qiao Qian, Minlie Huang, Jinhao Lei, Xiaoyan Zhu 0001
ACL (1)4
2017 Give me Something Unknown: Incorporate Exploration Preference in Cognition into Recommender System
abstract
Recommender systems (RSs) make recommendation for users by capturing their preference from historical behaviors. However, emphasizing too much on the perspective of prediction precision leads to the problem of preference overfitting, which blockades the user from touching a wide range of new but potentially interesting items. Usually, it would be beneficial for the system to attempt suitable extensions beyond users' existing preferences already expressed in the recommender system. Traditional recommender systems concentrate much more on what a user had explored than what she/he is yet to explore. Too much attention to the precision of recommendation can lead to a preference overfitting problem. Users have the need for certain amount of extension based on the existing preference. In this paper, we investigate how novel items should be recommended given the user's preference which is modeled from her/his historical behaviors. Inspired by Wundt Curve, a famous law in cognition psychology field, we propose an idea of Personal Exploration Preference (PEP) model to capture the user preference from another perspective, the strength of aspiration to explore. Then we integrate PEP model into traditional recommendation algorithms. Finally,we experimentally evaluate the improvement of PEP model to traditional algorithms on HetRec2011 dataset. Experiment results show that, on the Top-N recommendation task, PEP model can effectively capture user preference of exploration and improve recommendation quality significantly.
Daoyi Li, Minlie Huang, Xiaoyan Zhu 0001
ICTAI4
2017 Encoding Syntactic Knowledge in Neural Networks for Sentiment Classification
abstract
Phrase/Sentence representation is one of the most important problems in natural language processing. Many neural network models such as Convolutional Neural Network (CNN), Recursive Neural Network (RNN), and Long Short-Term Memory (LSTM) have been proposed to learn representations of phrase/sentence, however, rich syntactic knowledge has not been fully explored when composing a longer text from its shorter constituent words. In most traditional models, only word embeddings are utilized to compose phrase/sentence representations, while the syntactic information of words is yet to be explored. In this article, we discover that encoding syntactic knowledge (part-of-speech tag) in neural networks can enhance sentence/phrase representation. Specifically, we propose to learn tag-specific composition functions and tag embeddings in recursive neural networks, and propose to utilize POS tags to control the gates of tree-structured LSTM networks. We evaluate these models on two benchmark datasets for sentiment classification, and demonstrate that improvements can be obtained with such syntactic knowledge encoded.
Minlie Huang, Qiao Qian, Xiaoyan Zhu 0001
ACM Trans. Inf. Syst.3
2016 Semi-Supervised Multinomial Naive Bayes for Text Classification by Leveraging Word-Level Statistical Constraint
abstract
Multinomial Naive Bayes with Expectation Maximization (MNB-EM) is a standard semi-supervised learning method to augment Multinomial Naive Bayes (MNB) for text classification. Despite its success, MNB-EM is not stable, and may succeed or fail to improve MNB. We believe that this is because MNB-EM lacks the ability to preserve the class distribution on words. In this paper, we propose a novel method to augment MNB-EM by leveraging the word-level statistical constraint to preserve the class distribution on words. The word-level statistical constraints are further converted to constraints on document posteriors generated by MNB-EM. Experiments demonstrate that our method can consistently improve MNB-EM, and outperforms state-of-art baselines remarkably.
Li Zhao 0007, Minlie Huang, Ziyu Yao 0002, Rongwei Su, Yingying Jiang 0001, Xiaoyan Zhu 0001
AAAI6
2016 TransG : A Generative Model for Knowledge Graph Embedding
abstract
Recently, knowledge graph embedding, which projects symbolic entities and relations into continuous vector space, has become a new, hot topic in artificial intelligence.This paper proposes a novel generative model (TransG) to address the issue of multiple relation semantics that a relation may have multiple meanings revealed by the entity pairs associated with the corresponding triples.The new model can discover latent semantics for a relation and leverage a mixture of relationspecific component vectors to embed a fact triple.To the best of our knowledge, this is the first generative model for knowledge graph embedding, and at the first time, the issue of multiple relation semantics is formally discussed.Extensive experiments show that the proposed model achieves substantial improvements against the state-of-the-art baselines.
Han Xiao 0005, Minlie Huang, Xiaoyan Zhu 0001
ACL (1)3
2016 GAKE: Graph Aware Knowledge Embedding
abstract
Knowledge embedding, which projects triples in a given knowledge base to d-dimensional vectors, has attracted considerable research efforts recently. Most existing approaches treat the given knowledge base as a set of triplets, each of whose representation is then learned separately. However, as a fact, triples are connected and depend on each other. In this paper, we propose a graph aware knowledge embedding method (GAKE), which formulates knowledge base as a directed graph, and learns representations for any vertices or edges by leveraging the graph’s structural information. We introduce three types of graph context for embedding: neighbor context, path context, and edge context, each reflects properties of knowledge from different perspectives. We also design an attention mechanism to learn representative power of different vertices or edges. To validate our method, we conduct several experiments on two tasks. Experimental results suggest that our method outperforms several state-of-art knowledge embedding models.
Minlie Huang, Yang Yang 0009, Xiaoyan Zhu 0001
COLING4
2016 Product Review Summarization by Exploiting Phrase Properties
abstract
We propose a phrase-based approach for generating product review summaries. The main idea of our method is to leverage phrase properties to choose a subset of optimal phrases for generating the final summary. Specifically, we exploit two phrase properties, popularity and specificity. Popularity describes how popular the phrase is in the original reviews. Specificity describes how descriptive a phrase is in comparison to generic comments. We formalize the phrase selection procedure as an optimization problem and solve it using integer linear programming (ILP). An aspect-based bigram language model is used for generating the final summary with the selected phrases. Experiments show that our summarizer outperforms the other baselines.
Naitong Yu, Minlie Huang, Xiaoyan Zhu 0001
COLING4
2016 Context-aware Natural Language Generation for Spoken Dialogue Systems
abstract
Natural language generation (NLG) is an important component of question answering(QA) systems which has a significant impact on system quality. Most tranditional QA systems based on templates or rules tend to generate rigid and stylised responses without the natural variation of human language. Furthermore, such methods need an amount of work to generate the templates or rules. To address this problem, we propose a Context-Aware LSTM model for NLG. The model is completely driven by data without manual designed templates or rules. In addition, the context information, including the question to be answered, semantic values to be addressed in the response, and the dialogue act type during interaction, are well approached in the neural network model, which enables the model to produce variant and informative responses. The quantitative evaluation and human evaluation show that CA-LSTM obtains state-of-the-art performance.
Hao Zhou 0012, Minlie Huang, Xiaoyan Zhu 0001
COLING3
2016 Attention-based LSTM for Aspect-level Sentiment Classification
abstract
Aspect-level sentiment classification is a finegrained task in sentiment analysis.Since it provides more complete and in-depth results, aspect-level sentiment analysis has received much attention these years.In this paper, we reveal that the sentiment polarity of a sentence is not only determined by the content but is also highly related to the concerned aspect.For instance, "The appetizers are ok, but the service is slow.",for aspect taste, the polarity is positive while for service, the polarity is negative.Therefore, it is worthwhile to explore the connection between an aspect and the content of a sentence.To this end, we propose an Attention-based Long Short-Term Memory Network for aspect-level sentiment classification.The attention mechanism can concentrate on different parts of a sentence when different aspects are taken as input.We experiment on the SemEval 2014 dataset and results show that our model achieves state-ofthe-art performance on aspect-level sentiment classification.
Yequan Wang, Minlie Huang, Xiaoyan Zhu 0001, Li Zhao 0007
EMNLP3
2016 From One Point to a Manifold: Knowledge Graph Embedding for Precise Link Prediction
Han Xiao 0005, Minlie Huang, Xiaoyan Zhu 0001
IJCAI3
2016 Knowledge Graph Embedding by Flexible Translation
Minlie Huang, Mingdong Wang, Mantong Zhou, Yu Hao 0001, Xiaoyan Zhu 0001
KR6
2015 Learning Tag Embeddings and Tag-specific Composition Functions in Recursive Neural Network
abstract
Qiao Qian, Bo Tian, Minlie Huang, Yang Liu, Xuan Zhu, Xiaoyan Zhu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Qiao Qian, Minlie Huang, Yang Liu 0107, Xuan Zhu 0006, Xiaoyan Zhu 0001
ACL (1)6
2015 Optimizing the Bayesian Inference of Phylogeny on Graphic Processors
abstract
Searching for the evolutionary relationships between groups of organism has become a routine procedure in molecular biology. MrBayes is a popular model based phylogenetic inference tool using Bayesian statistics. Unfortunately, the computational cost is very high, resulting in undesirably long execution time. In this paper, we present what we believe the fastest solution of the MrBayes MC3 algorithm running on off-the-shelf graphic processors. The performance benefits are offered by the multi-granularity parallelism model, coarse-grained GPU kernel system, efficient thread arrangement strategy and GPU code level optimizations. MrBayes goMC3 (proposed herein) provides a significant performance improvement over the sequential MrBayes MC3 by a speedup of up to 48× when using single Tesla C2075 GPU card, whereas a speedup factor of 77× can be achieved when using dual GPUs. In comparison to the state-of-the-art version of other publicly available GPU implementations of MrBayes MC3, the cumulative optimizations adopted in goMC3 resulted in a speedup of up 2.5× over oMC3 (v1.0), 1.75× over tgMC3 (v1.0) and 1.46× over nMC3(v2.1.1) for realistic empirical biological datasets. Besides, experimental results indicated that goMC3 outstrips these GPU implementations on the analysis of simulated datasets composed of ultra-large-scale sequences. As a consequence, the reported performance improvement of goMC3 is significant and appears to scale well with increasing dataset sizes.
Cheng Ling, Chunbao Zhou, Arong Luo, Guoguang Zhao, Tsuyoshi Hamada, Xiaoyan Zhu 0001
CCGRID6
2015 Sentiment Extraction by Leveraging Aspect-Opinion Association Structure
abstract
Sentiment extraction aims to extract and group the task of extracting and grouping aspect and opinion words from online reviews. Previous works usually extract aspect and opinion words by leveraging association between a single pair of aspect and opinion word[5] [14] [9] [4][11], but the structure of aspect and opinion word clusters has not been fully exploited.
Li Zhao 0007, Minlie Huang, Jiashen Sun, Hengliang Luo, Xiankai Yang, Xiaoyan Zhu 0001
CIKM6
2015 Tackling Data Sparseness in Recommendation using Social Media based Topic Hierarchy Modeling
Xingwei Zhu, Zhaoyan Ming, Yu Hao 0001, Xiaoyan Zhu 0001
IJCAI4
2015 Clustering Sentiment Phrases in Product Reviews by Constrained Co-clustering
abstract
Clustering sentiment phrases in product reviews is convenient for us to get the most important information about one product directly through thousands of reviews. There are mainly two components in a sentiment phrase, the aspect word and the opinion word. We need to cluster these two parts simultaneously. Although several methods have been proposed to cluster words or phrases, limited work has been done on clustering two-dimensional sentiment phrases. In this paper, we apply a two-sided hidden Markov random field (HMRF) model on this task. We use the approach of constrained co-clustering with some priori knowledge, in a semi-supervised setting. Experimental results on sentiment phrases extracted from about 0.7 million mobile phone reviews show that this method is promising for this task and our method outperforms baselines remarkably.
Minlie Huang, Xiaoyan Zhu 0001
NLPCC3
2015 A Question Answering System Built on Domain Knowledge Base
Yu Hao 0001, Xiaoyan Zhu 0001, Jiao Li 0001
WAIM3
2015 Preface
Jie Tang 0001, Xiaoyan Zhu 0001
J. Comput. Sci. Technol.2
2015 Preface
Jie Tang 0001, Xiaoyan Zhu 0001
J. Comput. Sci. Technol.2
2014 New Word Detection for Sentiment Analysis
abstract
Automatic extraction of new words is an indispensable precursor to many NLP tasks such as Chinese word segmentation, named entity extraction, and sentiment analysis.This paper aims at extracting new sentiment words from large-scale user-generated content.We propose a fully unsupervised, purely data-driven framework for this purpose.We design statistical measures respectively to quantify the utility of a lexical pattern and to measure the possibility of a word being a new word.The method is almost free of linguistic resources (except POS tags), and requires no elaborated linguistic rules.We also demonstrate how new sentiment word will benefit sentiment analysis.Experiment results demonstrate the effectiveness of the proposed method.
Minlie Huang, Borui Ye, Haiqiang Chen, Junjun Cheng, Xiaoyan Zhu 0001
ACL (1)6
2014 Ranking Sentiment Explanations for Review Summarization Using Dual Decomposition
abstract
For online reviews, sentiment explanations refer to the sentences that may suggest detailed reasons of sentiment, which are very important for applications in review mining like opinion summarization. In this paper, we address the problem of ranking sentiment explanations by formulating the process as two subproblems: sentence informativeness ranking and structural sentiment analysis. Tractable inference in joint prediction is performed through dual decomposition. Preliminary experiments on publicly available data demonstrate that our approach obtains promising performance.
Lei Fang 0004, Qiao Qian, Minlie Huang, Xiaoyan Zhu 0001
CIKM4
2014 Customized Organization of Social Media Contents using Focused Topic Hierarchy
abstract
With the popularity of social media platforms such as Facebook and Twitter, the amount of useful data in these sources is rapidly increasing, making them promising places for information acquisition. This research aims at the customized organization of a social media corpus using focused topic hierarchy. It organizes the contents into different structures to meet with users' different information needs (e.g., "iPhone 5 problem" or "iPhone 5 camera"). To this end, we introduce a novel function to measure the likelihood of a topic hierarchy, by which the users' information need can be incorporated into the process of topic hierarchy construction. Using the structure information within the generated topic hierarchy, we then develop a probability based model to identify the representative contents for topics to assist users in document retrieval on the hierarchy. Experimental results on real world data illustrate the effectiveness of our method and its superiority over state-of-the-art methods for both information organization and retrieval tasks.
Xingwei Zhu, Zhaoyan Ming, Yu Hao 0001, Xiaoyan Zhu 0001, Tat-Seng Chua
CIKM4
2014 Clustering Aspect-related Phrases by Leveraging Sentiment Distribution Consistency
abstract
Clustering aspect-related phrases in terms of product’s property is a precursor pro-cess to aspect-level sentiment analysis which is a central task in sentiment analy-sis. Most of existing methods for address-ing this problem are context-based models which assume that domain synonymous phrases share similar co-occurrence con-texts. In this paper, we explore a novel idea, sentiment distribution consistency, which states that different phrases (e.g. “price”, “money”, “worth”, and “cost”) of the same aspect tend to have consistent sentiment distribution. Through formal-izing sentiment distribution consistency as soft constraint, we propose a novel unsu-pervised model in the framework of Poste-rior Regularization (PR) to cluster aspect-related phrases. Experiments demonstrate that our approach outperforms baselines remarkably. 1
Li Zhao 0007, Minlie Huang, Haiqiang Chen, Junjun Cheng, Xiaoyan Zhu 0001
EMNLP5
2014 Contextual Combinatorial Bandit and its Application on Diversified Online Recommendation
abstract
Recommender systems are faced with new challenges that are beyond traditional techniques. For example, most traditional techniques are based on similarity or overlap among existing data, however, there may not exist sufficient historical records for some new users to predict their preference, or users can hold diverse interest, but the similarity based methods may probably over-narrow it. To address the above challenges, we develop a principled approach called contextual combinatorial bandit in which a learning algorithm can dynamically identify diverse items that interest a new user. Specifically, each item is represented as a feature vector, and each user is represented as an unknown preference vector. On each of n rounds, the bandit algorithm sequentially selects a set of items according to the item-selection strategy that balances exploration and exploitation, and collects the user feedback on these selected items. A reward function is further designed to measure the quality (e.g. relevance or diversity) of the selected set based on observed feedback, and the goal of the algorithm is to maximize the total rewards of n rounds. The reward function only needs to satisfy two mild assumptions that is general enough to accommodate a large class of nonlinear functions. To solve this bandit problem, we provide algorithm that achieves Õ(√n) regret after playing n rounds. Experiments conducted on real-wold movie recommendation dataset demonstrate that our approach can effectively address the above challenges and hence improve the performance of recommendation task.
Lijing Qin, Shouyuan Chen, Xiaoyan Zhu 0001
SDM3
2014 A Chinese Question Answering System for Specific Domain
Tanche Li, Yu Hao 0001, Xiaoyan Zhu 0001, Xian Zhang 0006
WAIM3
2014 Sampling dilemma: towards effective data sampling for click prediction in sponsored search
abstract
Precise prediction of the probability that users click on ads plays a key role in sponsored search. State-of-the-art sponsored search systems typically employ a machine learning approach to conduct click prediction. While paying much attention to extracting useful features and building effective models, previous studies have overshadowed seemingly less obvious but essentially important challenges in terms of data sampling. To fulfill the learning objective of click prediction, it is not only necessary to ensure that the sampled training data implies the similar input distribution compared with the real world one, but also to guarantee that the sampled training data yield the consistent conditional output distribution, i.e. click-through rate (CTR), with the real world data. However, due to the sparseness of clicks in sponsored search, it is a bit contradictory to address these two challenges simultaneously. In this paper, we first take a theoretical analysis to reveal this sampling dilemma, followed by a thorough data analysis which demonstrates that the straightforward random sampling method may not be effective to balance these two kinds of consistency in sampling dilemma simultaneously. To address this problem, we propose a new sampling algorithm which can succeed in retaining the consistency between the sampled data and real world in terms of both input distribution and conditional output distribution. Large scale evaluations on the click-through logs from a commercial search engine demonstrate that this new sampling algorithm can effectively address the sampling dilemma. Further experiments illustrate that, by using the training data obtained by our new sampling algorithm, we can learn the model with much higher accuracy in click prediction.
Jiang Bian 0002, Taifeng Wang, Wei Chen 0034, Xiaoyan Zhu 0001, Tie-Yan Liu
WSDM5
2014 Exploring the Interactions of Storylines from Informative News Events
Minlie Huang, Xiaoyan Zhu 0001
J. Comput. Sci. Technol.3
2014 Estimating feature ratings through an effective review selection approach
Chong Long, Jie Zhang 0002, Minlie Huang, Xiaoyan Zhu 0001, Ming Li 0001, Bin Ma 0002
Knowl. Inf. Syst.4
2013 Exploring weakly supervised latent sentiment explanations for aspect-level review analysis
abstract
In sentiment analysis, aspect-level review analysis has been an important task because it can catalogue, aggregate, or summarize various opinions according to a product's properties. In this paper, we explore a new concept for aspect-level review analysis, latent sentiment explanations, which are defined as a set of informative aspect-specific sentences whose polarities are consistent with that of the review. In other words, sentiment explanations best represent a review in terms of both aspect and polarity. We formulate the problem as a structure learning problem, and sentiment explanations are modeled with latent variables. Training samples are automatically identified through a set of pre-defined aspect signature terms (i.e., without manual annotation on samples), which we term the way weakly supervised.
Lei Fang 0004, Minlie Huang, Xiaoyan Zhu 0001
CIKM3
2013 Functional dirichlet process
abstract
Dirichlet process mixture (DPM) model is one of the most important Bayesian nonparametric models owing to its efficiency of inference and flexibility for various applications. A fundamental assumption made by DPM model is that all data items are generated from a single, shared DP. This assumption, however, is restrictive in many practical settings where samples are generated from a collection of dependent DPs, each associated with a point in some covariate space. For example, documents in the proceedings of a conference are organized by year, or photos may be tagged and recorded with GPS locations. We present a general method for constructing dependent Dirichlet processes (DP) on arbitrary covariate space. The approach is based on restricting and projecting a DP defined on a space of continuous functions with different domains, which results in a collection of dependent random measures, each associated with a point in covariate space and is marginally DP distributed. The constructed collection of dependent DPs can be used as a nonparametric prior of infinite dynamic mixture models, which allow each mixture component to appear/disappear and vary in a subspace of covariate space. Furthermore, we discuss choices of base distributions of functions in a variety of settings as a flexible method to control dependencies. In addition, we develop an efficient Gibbs sampler for model inference where all underlying random measures are integrated out. Finally, experiment results on temporal modeling and spatial modeling datasets demonstrate the effectiveness of the method in modeling dynamic mixture models on different types of covariates.
Lijing Qin, Xiaoyan Zhu 0001
CIKM2
2013 Promoting Diversity in Recommendation by Entropy Regularizer
Lijing Qin, Xiaoyan Zhu 0001
IJCAI2
2013 Topic hierarchy construction for the organization of multi-source user generated contents
abstract
User generated contents (UGCs) carry a huge amount of high quality information. However, the information overload and diversity of UGC sources limit their potential uses. In this research, we propose a framework to organize information from multiple UGC sources by a topic hierarchy which is automatically generated and updated using the UGCs. We explore the unique characteristics of UGCs like blogs, cQAs, microblogs, etc., and introduce a novel scheme to combine them. We also propose a graph-based method to enable incremental update of the generated topic hierarchy. Using the hierarchy, users can easily obtain a comprehensive, in-depth and up-to-date picture of their topics of interests. The experiment results demonstrate how information from multiple heterogeneous sources improves the resultant topic hierarchies. It also shows that the proposed method achieves better F1 scores in hierarchy generation as compared to the state-of-the-art methods.
Xingwei Zhu, Zhaoyan Ming, Xiaoyan Zhu 0001, Tat-Seng Chua
SIGIR3
2012 Using First-Order Logic to Compress Sentences
abstract
Sentence compression is one of the most challenging tasks in natural language processing,which may be of increasing interest to many applicationssuch as abstractive summarization and text simplification for mobile devices.In this paper, we present a novel sentence compression model based on first-order logic, using Markov Logic Network.Sentence compression is formulated as a word/phrase deletion problem in this model.By taking advantage of first-order logic, the proposed method is able to incorporate local linguistic features and to capture global dependencies between word deletion operations. Experiments on both written and spoken corpora show that our approach produces competitive performance against the state-of-the-art methods in terms of manual evaluation measures such as importance, grammaticality, and overall quality.
Minlie Huang, Xiaoyan Zhu 0001
AAAI4
2012 Finding nuggets in IP portfolios: core patent mining through textual temporal analysis
abstract
Patents are critical for a company to protect its core technologies. Effective patent mining in massive patent databases can provide companies with valuable insights to develop strategies for IP management and marketing. In this paper, we study a novel patent mining problem of automatically discovering core patents (i.e., patents with high novelty and influence in a domain). We address the unique patent vocabulary usage problem, which is not considered in traditional word-based statistical methods, and propose a topic-based temporal mining approach to quantify a patent's novelty and influence. Comprehensive experimental results on real-world patent portfolios show the effectiveness of our method.
Minlie Huang, Peng Xu 0002, Weichang Li, Adam K. Usadi, Xiaoyan Zhu 0001
CIKM6
2012 Sentiment Analysis with Multi-source Product Reviews
Minlie Huang, Xiaoyan Zhu 0001
ICIC (1)3
2012 Entity Disambiguation with Freebase
abstract
Entity disambiguation with a knowledge base becomes increasingly popular in the NLP community. In this paper, we employ Freebase as the knowledge base, which contains significantly more entities than Wikipedia and others. While huge in size, Freebase lacks context for most entities, such as the descriptive text and hyperlinks in Wikipedia, which are useful for disambiguation. Instead, we leverage two features of Freebase, namely the naturally disambiguated mention phrases (aka aliases) and the rich taxonomy, to perform disambiguation in an iterative manner. Specifically, we explore both generative and discriminative models for each iteration. Experiments on 2, 430, 707 English sentences and 33, 743 Freebase entities show the effectiveness of the two features, where 90% accuracy can be reached without any labeled data. We also show that discriminative models with proposed split training strategy is robust against over fitting problem, and constantly outperforms the generative ones.
Zhicheng Zheng, Xiance Si, Fangtao Li, Edward Y. Chang, Xiaoyan Zhu 0001
Web Intelligence5
2012 A Unified Active Learning Framework for Biomedical Relation Extraction
Minlie Huang, Xiaoyan Zhu 0001
J. Comput. Sci. Technol.3
2011 Generating Breakpoint-based Timeline Overview for News Topic Retrospection
abstract
Though news readers can easily access a large number of news articles from the Internet, they can be overwhelmed by the quantity of information available, making it hard to get a concise, global picture of a news topic. In this paper we propose a novel method to address this problem. Given a set of articles for a given news topic, the proposed method models theme variation through time and identifies the breakpoints, which are time points when decisive changes occur. For each breakpoint, a brief summary is automatically constructed based on articles associated with the particular time point. Summaries are then ordered chronologically to form a timeline overview of the news topic. In this fashion, readers can easily track various news topics efficiently. We have conducted experiments on 15 popular topics in 2010. Empirical experiments show the effectiveness of our approach and its advantages over other approaches.
Minlie Huang, Peng Xu 0002, Weichang Li, Adam K. Usadi, Xiaoyan Zhu 0001
ICDM6
2011 Semantic Relationship Discovery with Wikipedia Structure
abstract
Thanks to the idea of social collaboration, Wikipedia has accumulated vast amount of semi-structured knowledge in which the link structure reflects human's cognition on semantic relationship to some extent. In this paper, we proposed a novel method RCRank to jointly compute concept-concept relatedness and concept-category relatedness base on the assumption that information carried in concept-concept links and concept-category links can mutually reinforce each other. Different from previous work, RCRank can not only find semantically related concepts but also interpret their relations by categories. Experimental results on concept recommendation and relation interpretation show that our method substantially outperforms classical methods.
Yu Hao 0001, Xiaoyan Zhu 0001
IJCAI3
2011 Learning to Identify Review Spam
abstract
In the past few years, sentiment analysis and opinion mining becomes a popular and important task. These studies all assume that their opinion resources are real and trustful. However, they may encounter the faked opinion or opinion spam problem. In this paper, we study this issue in the context of our product review mining system. On product review site, people may write faked reviews, called review spam, to promote their products, or defame their competitors’ products. It is important to identify and filter out the review spam. Previous work only focuses on some heuristic rules, such as helpfulness voting, or rating deviation, which limits the performance of this task. In this paper, we exploit machine learning methods to identify review spam. Toward the end, we manually build a spam collection from our crawled reviews. We first analyze the effect of various features in spam identification. We also observe that the review spammer consistently writes spam. This provides us another view to identify review spam: we can identify if the author of the review is spammer. Based on this observation, we provide a twoview semi-supervised method, co-training, to exploit the large amount of unlabeled data. The experiment results show that our proposed method is effective. Our designed machine learning methods achieve significant improvements in comparison to the heuristic baselines.
Fangtao Li, Minlie Huang, Yi Yang 0038, Xiaoyan Zhu 0001
IJCAI4
2011 Quality-biased Ranking of Short Texts in Microblogging Services
Minlie Huang, Yi Yang 0038, Xiaoyan Zhu 0001
IJCNLP3
2011 K2Q: Generating Natural Language Questions from Keywords with User Refinements
Zhicheng Zheng, Xiance Si, Edward Y. Chang, Xiaoyan Zhu 0001
IJCNLP4
2011 GeneTUKit: a software for document-level gene normalization
abstract
MOTIVATION: Linking gene mentions in an article to entries of biological databases can facilitate indexing and querying biological literature greatly. Due to the high ambiguity of gene names, this task is particularly challenging. Manual annotation for this task is cost expensive, time consuming and labor intensive. Therefore, providing assistive tools to facilitate the task is of high value. RESULTS: We developed GeneTUKit, a document-level gene normalization software for full-text articles. This software employs both local context surrounding gene mentions and global context from the whole full-text document. It can normalize genes of different species simultaneously. When participating in BioCreAtIvE III, the system obtained good results among 37 runs: the system was ranked first, fourth and seventh in terms of TAP-20, TAP-10 and TAP-5, respectively on the 507 full-text test articles. AVAILABILITY AND IMPLEMENTATION: The software is available at http://www.qanswers.net/GeneTUKit/.
Minlie Huang, Jingchen Liu, Xiaoyan Zhu 0001
Bioinform.3
2011 A New Multiword Expression Metric and Its Applications
Xiaoyan Zhu 0001, Ming Li 0001
J. Comput. Sci. Technol.2
2011 Guided Structure-Aware Review Summarization
Minlie Huang, Xiaoyan Zhu 0001
J. Comput. Sci. Technol.3
2010 Sentiment Analysis with Global Topics and Local Dependency
abstract
With the development of Web 2.0, sentiment analysis has now become a popular research problem to tackle. Recently, topic models have been introduced for the simultaneous analysis for topics and the sentiment in a document. These studies, which jointly model topic and sentiment, take the advantage of the relationship between topics and sentiment, and are shown to be superior to traditional sentiment analysis tools. However, most of them make the assumption that, given the parameters, the sentiments of the words in the document are all independent. In our observation, in contrast, sentiments are expressed in a coherent way. The local conjunctive words, such as “and” or “but”, are often indicative of sentiment transitions. In this paper, we propose a major departure from the previous approaches by making two linked contributions. First, we assume that the sentiments are related to the topic in the document, and put forward a joint sentiment and topic model, i.e. Sentiment-LDA. Second, we observe that sentiments are dependent on local context. Thus, we further extend the Sentiment-LDA model to Dependency-Sentiment-LDA model by relaxing the sentiment independent assumption in Sentiment-LDA. The sentiments of words are viewed as a Markov chain in Dependency-Sentiment-LDA. Through experiments, we show that exploiting the sentiment dependency is clearly advantageous, and that the Dependency-Sentiment-LDA is an effective approach for sentiment analysis.
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
AAAI3
2010 Measuring the Non-compositionality of Multiword Expressions
Xiaoyan Zhu 0001, Ming Li 0001
COLING2
2010 Structure-Aware Review Mining and Summarization
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001, Yingju Xia, Shu Zhang 0004, Hao Yu 0005
COLING4
2010 Function-Based Question Classification for General QA
Xingwei Zhu, Yu Hao 0001, Xiaoyan Zhu 0001
EMNLP4
2010 Learning to Link Entities with Knowledge Base
Zhicheng Zheng, Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
HLT-NAACL4
2010 A New Approach for Multi-Document Update Summarization
Chong Long, Minlie Huang, Xiaoyan Zhu 0001, Ming Li 0001
J. Comput. Sci. Technol.3
2009 Answering Opinion Questions with Random Walks on Graphs
Fangtao Li, Minlie Huang, Xiaoyan Zhu 0001
ACL/IJCNLP4
2009 Multi-document Summarization by Information Distance
abstract
Fast changing knowledge on the Internet can be acquired more efficiently with the help of automatic document summarization and updating techniques. This paper described a novel approach for multi-document update summarization. The best summary is defined to be the one which has the minimum information distance to the entire document set. The best update summary has the minimum conditional information distance to a document cluster given that a prior document cluster has already been read. Experiments on the DUC 2007 dataset and the TAC 2008 dataset have proved that our method closely correlates with the human summaries and outperforms other programs such as LexRank in many categories under the ROUGE evaluation criterion.
Chong Long, Minlie Huang, Xiaoyan Zhu 0001, Ming Li 0001
ICDM3
2009 Specialized Review Selection for Feature Rating Estimation
abstract
On participatory Websites, users provide opinions about products, with both overall ratings and textual reviews. In this paper, we propose an approach to accurately estimate feature ratings of the products. This approach selects user reviews that extensively discuss specific features of the products (called specialized reviews), using information distance of reviews on the features. Experiments on real data show that overall ratings of the specialized reviews can be used to represent their feature ratings. The average of these overall ratings can be used by recommender systems to provide feature specific recommendations that better help users make purchasing decisions.
Chong Long, Jie Zhang 0002, Minlie Huang, Xiaoyan Zhu 0001, Ming Li 0001, Bin Ma 0002
Web Intelligence4
2009 Extract interaction detection methods from the biological literature
abstract
BACKGROUND: Considerable efforts have been made to extract protein-protein interactions from the biological literature, but little work has been done on the extraction of interaction detection methods. It is crucial to annotate the detection methods in the literature, since different detection methods shed different degrees of reliability on the reported interactions. However, the diversity of method mentions in the literature makes the automatic extraction quite challenging. RESULTS: In this article, we develop a generative topic model, the Correlated Method-Word model (CMW model) to extract the detection methods from the literature. In the CMW model, we formulate the correlation between the different methods and related words in a probabilistic framework in order to infer the potential methods from the given document. By applying the model on a corpus of 5319 full text documents annotated by the MINT and IntAct databases, we observe promising results, which outperform the best result reported in the BioCreative II challenge evaluation. CONCLUSION: From the promising experiment results, we can see that the CMW model overcomes the issues caused by the diversity in the method mentions and properly captures the in-depth correlations between the detection methods and related words. The performance outperforming the baseline methods confirms that the dependence assumptions of the model are reasonable and the model is competent for the practical processing.
Hongning Wang, Minlie Huang, Xiaoyan Zhu 0001
BMC Bioinform.3
2009 Building Disease-Specific Drug-Protein Connectivity Maps from Molecular Interaction Networks and PubMed Abstracts
abstract
The recently proposed concept of molecular connectivity maps enables researchers to integrate experimental measurements of genes, proteins, metabolites, and drug compounds under similar biological conditions. The study of these maps provides opportunities for future toxicogenomics and drug discovery applications. We developed a computational framework to build disease-specific drug-protein connectivity maps. We integrated gene/protein and drug connectivity information based on protein interaction networks and literature mining, without requiring gene expression profile information derived from drug perturbation experiments on disease samples. We described the development and application of this computational framework using Alzheimer's Disease (AD) as a primary example in three steps. First, molecular interaction networks were incorporated to reduce bias and improve relevance of AD seed proteins. Second, PubMed abstracts were used to retrieve enriched drug terms that are indirectly associated with AD through molecular mechanistic studies. Third and lastly, a comprehensive AD connectivity map was created by relating enriched drugs and related proteins in literature. We showed that this molecular connectivity map development approach outperformed both curated drug target databases and conventional information retrieval systems. Our initial explorations of the AD connectivity map yielded a new hypothesis that diltiazem and quinidine may be investigated as candidate drugs for AD treatment. Molecular connectivity maps derived computationally can help study molecular signature differences between different classes of drugs in specific disease contexts. To achieve overall good data coverage and quality, a series of statistical methods have been developed to overcome high levels of data noise in biological networks and literature mining results. Further development of computational molecular connectivity maps to cover major disease areas will likely set up a new model for drug development, in which therapeutic/toxicological profiles of candidate drugs can be checked computationally before costly clinical trials begin.
Jiao Li 0001, Xiaoyan Zhu 0001, Jake Yue Chen
PLoS Comput. Biol.2
2008 Using Conditional Random Fields to Extract Contexts and Answers of Questions from Online Forums
Shilin Ding, Gao Cong, Chin-Yew Lin, Xiaoyan Zhu 0001
ACL4
2008 Information shared by many objects
abstract
If Kolmogorov complexity [25] measures information in one object and Information Distance measures information shared by two objects, how do we measure information shared by many objects? This paper provides an initial pragmatic study of this fundamental data mining question. Firstly, Em(x1,x2,...,xn) is defined to be the minimum amount of thermodynamic energy needed to convert from any xi to any xj. With this definition several theoretical problems have been solved. Second, our newly proposed theory is applied to select a comprehensive review and a specialized review from many reviews: (1) Core feature words, expanded words and dependent words are extracted respectively. (2) Comprehensive and specialized reviews are selected according to the information among them. This method of selecting a single review can be extended to select multiple reviews as well. Finally, experiments show that this comprehensive and specialized review mining method based on our new theory can do the job efficiently.
Chong Long, Xiaoyan Zhu 0001, Ming Li 0001, Bin Ma 0002
CIKM2
2008 Classifying What-Type Questions by Head Noun Tagging
Fangtao Li, Xian Zhang 0006, Jinhui Yuan, Xiaoyan Zhu 0001
COLING4
2008 A Generative Probabilistic Model for Multi-label Classification
abstract
Traditional discriminative classification method makes little attempt to reveal the probabilistic structure and the correlation within both input and output spaces. In the scenario of multi-label classification, most of the classifiers simply assume the predefined classes are independently distributed, which would definitely hinder the classification performance when there are intrinsic correlations between the classes. In this article, we propose a generative probabilistic model, the Correlated Labeling Model (CoL Model), to formulate the correlation between different classes. The CoL model is presented to capture the correlation between classes and the underlying structures via the latent random variables in a supervised manner. We develop a variational procedure to approximate the posterior distribution and employ the EM algorithm for the empirical Bayes parameter estimation. In our evaluations, the proposed model achieved promising results on various data sets.
Hongning Wang, Minlie Huang, Xiaoyan Zhu 0001
ICDM3
2008 Exploiting and integrating rich features for biological literature classification
abstract
BACKGROUND: Efficient features play an important role in automated text classification, which definitely facilitates the access of large-scale data. In the bioscience field, biological structures and terminologies are described by a large number of features; domain dependent features would significantly improve the classification performance. How to effectively select and integrate different types of features to improve the biological literature classification performance is the major issue studied in this paper. RESULTS: To efficiently classify the biological literatures, we propose a novel feature value schema TF*ML, features covering from lower level domain independent "string feature" to higher level domain dependent "semantic template feature", and proper integrations among the features. Compared to our previous approaches, the performance is improved in terms of AUC and F-Score by 11.5% and 8.8% respectively, and outperforms the best performance achieved in BioCreAtIvE 2006. CONCLUSIONS: Different types of features possess different discriminative capabilities in literature classification; proper integration of domain independent and dependent features would significantly improve the performance and overcome the over-fitting on data distribution.
Hongning Wang, Minlie Huang, Shilin Ding, Xiaoyan Zhu 0001
BMC Bioinform.4
2008 New Information Distance Measure and Its Application in Question Answering System
Xian Zhang 0006, Yu Hao 0001, Xiaoyan Zhu 0001, Ming Li 0001
J. Comput. Sci. Technol.3
2007 Semi-supervised Pattern Learning for Extracting Relations from Bioscience Texts
Shilin Ding, Minlie Huang, Xiaoyan Zhu 0001
APBC3
2007 A Novel Kernel-Based Approach for Predicting Binding Peptides for HLA Class II Molecules
Hao Yu 0005, Minlie Huang, Xiaoyan Zhu 0001, Yabin Guo
ISBRA3
2007 Information distance from a question to an answer
abstract
We provide three key missing pieces of a general theory of information distance [3, 23, 24]. We take bold steps in formulating a revised theory to avoid some pitfalls in practical applications. The new theory is then used to construct a question answering system. Extensive experiments are conducted to justify the new theory.
Xian Zhang 0006, Yu Hao 0001, Xiaoyan Zhu 0001, Ming Li 0001, David R. Cheriton
KDD3
2007 Combining Convolution Kernels Defined on Heterogeneous Sub-structures
Minlie Huang, Xiaoyan Zhu 0001
PAKDD2
2006 ONBRIRES: Ontology-Based Biological Relation Extraction System
Minlie Huang, Xiaoyan Zhu 0001, Shilin Ding, Hao Yu 0005, Ming Li 0001
APBC2
2006 Movie review mining and summarization
abstract
With the flourish of the Web, online review is becoming a more and more useful and important information resource for people. As a result, automatic review mining and summarization has become a hot research topic recently. Different from traditional text summarization, review mining and summarization aims at extracting the features on which the reviewers express their opinions and determining whether the opinions are positive or negative. In this paper, we focus on a specific domain - movie review. A multi-knowledge based approach is proposed, which integrates WordNet, statistical analysis and movie knowledge. The experimental results show the effectiveness of the proposed approach in movie review mining and summarization.
Li Zhuang, Xiaoyan Zhu 0001
CIKM3
2006 A Hybrid Handwritten Chinese Address Recognition Approach
Kaizhu Huang, Jun Sun 0004, Yoshinobu Hotta, Katsuhito Fujimoto, Satoshi Naoi, Chong Long, Li Zhuang, Xiaoyan Zhu 0001
ICONIP (2)8
2005 An OCR Post-processing Approach Based on Multi-knowledge
Li Zhuang, Xiaoyan Zhu 0001
KES (1)2
2005 Discovering patterns to extract protein-protein interactions from the literature: Part II
abstract
MOTIVATION: An enormous number of protein-protein interaction relationships are buried in millions of research articles published over the years, and the number is growing. Rediscovering them automatically is a challenging bioinformatics task. Solutions to this problem also reach far beyond bioinformatics. RESULTS: We study a new approach that involves automatically discovering English expression patterns, optimizing them and using them to extract protein-protein interactions. In a sister paper, we described how to generate English expression patterns related to protein-protein interactions, and this approach alone has already achieved precision and recall rates significantly higher than those of other automatic systems. This paper continues to present our theory, focusing on how to improve the patterns. A minimum description length (MDL)-based pattern-optimization algorithm is designed to reduce and merge patterns. This has significantly increased generalization power, and hence the recall and precision rates, as confirmed by our experiments. AVAILABILITY: http://spies.cs.tsinghua.edu.cn.
Hao Yu 0005, Xiaoyan Zhu 0001, Minlie Huang, Ming Li 0001
Bioinform.2
2005 Programming Style Based Program Partition
abstract
Program partitioning is a task of splitting a large, complex software system into functionally independent program modules. It is a key step in program understanding, software maintenance and software reuse. Traditional program partitioning methods are nonlinear. In most cases, the computational efforts needed for partitioning a source program will increase exponentially with the size of the source program. The NP-hard complexity constitutes a computational barrier for partitioning legacy software systems properly and efficiently. In this paper, we propose a new method that can partition a source program into program modules within a timescale that is linear with the size of the program. Our method uses special heuristic knowledge, based on psychological analysis on human programming styles, to partition a source program into domain-oriented program modules. A case study on a legacy C program that consists of 92 functions is reported to demonstrate the efficiency and effectiveness of this method.
Yang Li 0028, Xiaochun Cheng, Xiaoyan Zhu 0001
Int. J. Softw. Eng. Knowl. Eng.4
2004 PathwayFinder: Paving the Way Towards Automatic Pathway Extraction
Daming Yao, Yanmei Lu, Nathan Noble, Huandong Sun, Xiaoyan Zhu 0001, Donald G. Payan, Ming Li 0001, Kunbin Qu
APBC6
2004 Discovering patterns to extract protein-protein interactions from full texts
abstract
MOTIVATION: Although there are several databases storing protein-protein interactions, most such data still exist only in the scientific literature. They are scattered in scientific literature written in natural languages, defying data mining efforts. Much time and labor have to be spent on extracting protein pathways from literature. Our aim is to develop a robust and powerful methodology to mine protein-protein interactions from biomedical texts. RESULTS: We present a novel and robust approach for extracting protein-protein interactions from literature. Our method uses a dynamic programming algorithm to compute distinguishing patterns by aligning relevant sentences and key verbs that describe protein interactions. A matching algorithm is designed to extract the interactions between proteins. Equipped only with a dictionary of protein names, our system achieves a recall rate of 80.0% and precision rate of 80.5%. AVAILABILITY: The program is available on request from the authors.
Minlie Huang, Xiaoyan Zhu 0001, Hao Yu 0005, Donald G. Payan, Kunbin Qu, Ming Li 0001
Bioinform.2
2001 A Chinese spoken dialog system for blind men
abstract
This paper introduces a Chinese spoken dialog system providing services for blind people through which. they can use computers. A description of the architecture of the dialog system is presented briefly and the way in which each component works is also explained. The key factor of such a dialog system is extraction of the intention of a user's utterance so as to make an appropriate response. To achieve this, a case grammar formalism was applied for semantic description and a robust spoken language parsing method based on case-frames was adopted to obtain the semantic interpretation of the input. It shows that this parsing method can tolerate errors of speech recognition and grammatical deviation of spoken language to some extent.
Xiaoyan Zhu 0001, Yu Hao 0001
SMC2
2000 A criterion based on Fourier transform for segmentation of connected digits
Xiaoyan Zhu 0001, Yu Hao 0001
Int. J. Document Anal. Recognit.1
2000 A novel text-independent speaker verification method based on the global speaker model
abstract
This correspondence introduces a new text-independent speaker verification method, which is derived from the basic idea of pattern recognition that the discriminating ability of a classifier can be improved by removing the common information between classes. In looking for the common speech characteristics between a group of speakers, a global speaker model can be established. By subtracting the score acquired from this model, the conventional likelihood score is normalized with the consequence of more compact score distribution and lower equal error rates. Several experiments are carried out to demonstrate the effectiveness of the proposed method.
David Zhang 0001, Xiaoyan Zhu 0001
IEEE Trans. Syst. Man Cybern. Part A3