VLDB 2026 Research / reviewers in the wild / expert
Zhengyu Niu
dblp:31/9311 · also Zheng-Yu Niu
· DBLP profile ↗
47ranked-venue papers
10as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Thinking in Character: Advancing Role-Playing Agents with Role-Aware ReasoningabstractThe advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i.e., RPAs forget their role) and style drift (i.e., overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift. Yihong Tang, Kehai Chen, Muyun Yang, Zhengyu Niu, Tiejun Zhao, Min Zhang 0005 |
NeurIPS | 4 |
| 2025 | Towards few-shot mixed-type dialogue generation
Zeming Liu, Haifeng Wang 0001, Zeyang Lei, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
Sci. China Inf. Sci. | 4 |
| 2024 | Learning to Select External Knowledge With Multi-Scale Negative SamplingabstractThe Track-1 of DSTC9 aims to effectively answer user requests or questions during task-oriented dialogues, which are out of the scope of APIs/DB. By leveraging external knowledge resources, relevant information can be retrieved and encoded into the response generation for these out-of-API-coverage queries. In this work, we have explored several advanced techniques to enhance the utilization of external knowledge and boost the quality of response generation, includingschema guided knowledge decision,negatives enhanced knowledge selection, andknowledge grounded response generation. To evaluate the performance of our proposed method, comprehensive experiments have been carried out on the publicly available dataset. Our approach was ranked as the best in human evaluation of DSTC9 Track-1. Huang He, Hua Lu 0014, Siqi Bao, Fan Wang 0021, Hua Wu 0003, Zhengyu Niu, Haifeng Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | XDailyDialog: A Multilingual Parallel Dialogue CorpusabstractZeming Liu, Ping Nie, Jie Cai, Haifeng Wang, Zheng-Yu Niu, Peng Zhang, Mrinmaya Sachan, Kaiping Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zeming Liu, Ping Nie, Haifeng Wang 0001, Zhengyu Niu, Mrinmaya Sachan, Kaiping Peng |
ACL (1) | 5 |
| 2023 | Graph-Grounded Goal Planning for Conversational RecommendationabstractConversational recommendation casts the recommendation problem as a dialog-based interactive task, which could acquire user interest more efficiently and effectively by allowing users to express what they like. In this work, we move a step towards a new conversational recommendation task that is more suitable for real-world applications. In this task, the recommender proactively and naturally lead a dialog from non-recommendation content to approach an item being of interest to users, and allow users to ask questions for better support of user decisions. The challenge of this task lies in how to effectively control the dialog flow to complete the recommendation while appropriately responding to user utterances. To address this challenge, we first construct a Chinese recommendation dialog dataset DuRecDial. We then propose a two-stage Multi-Goal driven Conversation Generation framework, MGCG. In particular, the goal planning module leverages the global graph structure information and local goal-sequence information to effectively control the dialog flow step by step. The goal-guided responding module can produce an in-depth dialog about each goal by fully exploiting hierarchical goal information for response retrieval or generation. Results on DuRecDial demonstrate that MGCG can lead the dialog more proactively and naturally, and complete the recommendation task more effectively. Zeming Liu, Hao Liu 0026, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User GoalsabstractMost dialog systems posit that users have figured out clear and specific goals before starting an interaction.For example, users have determined the departure, the destination, and the travel time for booking a flight.However, in many scenarios, limited by experience and knowledge, users may know what they need, but still struggle to figure out clear and specific goals by determining all the necessary slots. Zeming Liu, Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003 |
ACL (1) | 5 |
| 2022 | CDConv: A Benchmark for Contradiction Detection in Chinese ConversationsabstractChujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Zhen Guo, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Minlie Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Wenquan Wu, Zhengyu Niu, Hua Wu 0003, Minlie Huang |
EMNLP | 7 |
| 2021 | Discovering Dialog Structure Graph for Coherent Dialog GenerationabstractJun Xu, Zeyang Lei, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
ACL/IJCNLP (1) | 4 |
| 2021 | DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational RecommendationabstractIn this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendation.The difference between DuRecDial 2.0 and existing conversational recommendation datasets is that the data item (Profile, Goal, Knowledge, Context, Response) in DuRecDial 2.0 is annotated in two languages, both English and Chinese, while other datasets are built with the setting of a single language.We collect 8.2k dialogs aligned across English and Chinese languages (16.5k dialogs and 255k utterances in total) that are annotated by crowdsourced workers with strict quality control procedure.We then build monolingual, multilingual, and cross-lingual conversational recommendation baselines on DuRecDial 2.0.Experiment results show that the use of additional English data can bring performance improvement for Chinese conversational recommendation, indicating the benefits of DuRecDial 2.0.Finally, this dataset provides a challenging testbed for future studies of monolingual, multilingual, and cross-lingual conversational recommendation. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
EMNLP (1) | 3 |
| 2021 | Multi-modal visual adversarial Bayesian personalized ranking model for recommendation
Guangli Li, Jianwu Zhuo, Chuanxiu Li, Jin Hua, Zhengyu Niu, Donghong Ji, Renzhong Wu, Hongbin Zhang 0004 |
Inf. Sci. | 6 |
| 2021 | Knowledge graph embedding with shared latent semantic units
Zhao Zhang 0011, Fuzhen Zhuang, Meng Qu, Zhengyu Niu, Hui Xiong 0001, Qing He 0003 |
Neural Networks | 4 |
| 2021 | Coherent Dialog Generation with Query GraphabstractLearning to generate coherent and informative dialogs is an enduring challenge for open-domain conversation generation. Previous work leverage knowledge graph or documents to facilitate informative dialog generation, with little attention on dialog coherence. In this article, to enhance multi-turn open-domain dialog coherence, we propose to leverage a new knowledge source, web search session data, to facilitate hierarchical knowledge sequence planning, which determines a sketch of a multi-turn dialog. Specifically, we formulate knowledge sequence planning or dialog policy learning as a graph grounded Reinforcement Learning (RL) problem. To this end, we first build a two-level query graph with queries as utterance-level vertices and their topics (entities in queries) as topic-level vertices. We then present a two-level dialog policy model that plans a high-level topic sequence and a low-level query sequence over the query graph to guide a knowledge aware response generator. In particular, to foster forward-looking knowledge planning decisions for better dialog coherence, we devise a heterogeneous graph neural network to incorporate neighbouring vertex information, or possible future RL action information, into each vertex (as an RL action) representation. Experiment results on two benchmark dialog datasets demonstrate that our framework can outperform strong baselines in terms of dialog coherence, informativeness, and engagingness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Jizhou Huang, Ting Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2020 | Knowledge Graph Grounded Goal Planning for Open-Domain Conversation GenerationabstractPrevious neural models on open-domain conversation generation have no effective mechanisms to manage chatting topics, and tend to produce less coherent dialogs. Inspired by the strategies in human-human dialogs, we divide the task of multi-turn open-domain conversation generation into two sub-tasks: explicit goal (chatting about a topic) sequence planning and goal completion by topic elaboration. To this end, we propose a three-layer Knowledge aware Hierarchical Reinforcement Learning based Model (KnowHRL). Specifically, for the first sub-task, the upper-layer policy learns to traverse a knowledge graph (KG) in order to plan a high-level goal sequence towards a good balance between dialog coherence and topic consistency with user interests. For the second sub-task, the middle-layer policy and the lower-layer one work together to produce an in-depth multi-turn conversation about a single topic with a goal-driven generation mechanism. The capability of goal-sequence planning enables chatbots to conduct proactive open-domain conversations towards recommended topics, which has many practical applications. Experiments demonstrate that our model outperforms state of the art baselines in terms of user-interest consistency, dialog coherence, and knowledge accuracy. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
AAAI | 3 |
| 2020 | Towards Conversational Recommendation over Multi-Type DialogsabstractWe focus on the study of conversational recommendation in the context of multi-type dialogs, where the bots can proactively and naturally lead a conversation from a nonrecommendation dialog (e.g., QA) to a recommendation dialog, taking into account user's interests and feedback.To facilitate the study of this task, we create a human-to-human Chinese dialog dataset DuRecDial (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).In each dialog, the recommender proactively leads a multi-type dialog to approach recommendation targets and then makes multiple recommendations with rich interaction behavior.This dataset allows us to systematically investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how to interact with users for recommendation.Finally we establish baseline results on DuRecDial for future studies. 1 Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 3 |
| 2020 | Conversational Graph Grounded Policy Learning for Open-Domain Conversation GenerationabstractTo address the challenge of policy learning in open-domain multi-turn conversation, we propose to represent prior information about dialog transitions as a graph and learn a graph grounded dialog policy, aimed at fostering a more coherent and controllable dialog.To this end, we first construct a conversational graph (CG) from dialog corpora, in which there are vertices to represent "what to say" and "how to say", and edges to represent natural transition between a message (the last utterance in a dialog context) and its response.We then present a novel CG grounded policy learning framework that conducts dialog flow planning by graph traversal, which learns to identify a what-vertex and a how-vertex from the CG at each turn to guide response generation.In this way, we effectively leverage the CG to facilitate policy learning as follows: (1) it enables more effective long-term reward design, (2) it provides high-quality candidate actions, and (3) it gives us more control over the policy.Results on two benchmark corpora demonstrate the effectiveness of this framework. Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001 |
ACL | 3 |
| 2020 | Enhancing Dialog Coherence with Event Graph Grounded Content PlanningabstractHow to generate informative, coherent and sustainable open-domain conversations is a non-trivial task. Previous work on knowledge grounded conversation generation focus on improving dialog informativeness with little attention on dialog coherence. In this paper, to enhance multi-turn dialog coherence, we propose to leverage event chains to help determine a sketch of a multi-turn dialog. We first extract event chains from narrative texts and connect them as a graph. We then present a novel event graph grounded Reinforcement Learning (RL) framework. It conducts high-level response content (simply an event) planning by learning to walk over the graph, and then produces a response conditioned on the planned content. In particular, we devise a novel multi-policy decision making mechanism to foster a coherent dialog with both appropriate content ordering and high contextual relevance. Experimental results indicate the effectiveness of this framework in terms of dialog coherence and informativeness. Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
IJCAI | 4 |
| 2020 | DE-Ada*: A novel model for breast mass classification using cross-modal pathological semantic mining and organic integration of multi-feature fusions
Hongbin Zhang 0004, Renzhong Wu, Ziliang Jiang, Jinpeng Wu, Jin Hua, Zhengyu Niu, Donghong Ji |
Inf. Sci. | 8 |
| 2019 | Knowledge Aware Conversation Generation with Explainable Reasoning over Augmented GraphsabstractZhibin Liu, Zheng-Yu Niu, Hua Wu, Haifeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Zhengyu Niu, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | A Key-Phrase Aware End2end Neural Response Generation Model
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che |
NLPCC (2) | 3 |
| 2019 | Knowledge triple mining via multi-task learning
Zhao Zhang 0011, Fuzhen Zhuang, Xuebing Li, Zhengyu Niu, Jia He 0001, Qing He 0003, Hui Xiong 0001 |
Inf. Syst. | 4 |
| 2019 | Supervised representation learning for multi-label classification
Fuzhen Zhuang, Xiao Zhang 0015, Xiang Ao 0001, Zhengyu Niu, Min-Ling Zhang, Qing He 0003 |
Mach. Learn. | 5 |
| 2018 | MultiE: Multi-Task Embedding for Knowledge Base CompletionabstractCompleting knowledge bases (KBs) with missing facts is of great importance, since most existing KBs are far from complete. To this end, many knowledge base completion (KBC) methods have been proposed. However, most existing methods embed each relation into a vector separately, while ignoring the correlations among different relations. Actually, in large-scale KBs, there always exist some relations that are semantically related, and we believe this can help to facilitate the knowledge sharing when learning the embedding of related relations simultaneously. Along this line, we propose a novel KBC model by Multi -Task E mbedding, named MultiE. In this model, semantically related relations are first clustered into the same group, and then learning the embedding of each relation can leverage the knowledge among different relations. Moreover, we propose a three-layer network to predict the missing values of incomplete knowledge triples. Finally, experiments on three popular benchmarks FB15k, FB15k-237 and WN18 are conducted to demonstrate the effectiveness of MultiE against some state-of-the-art baseline competitors. Zhao Zhang 0011, Fuzhen Zhuang, Zhengyu Niu, Deqing Wang 0001, Qing He 0003 |
CIKM | 3 |
| 2018 | Learning a unified embedding space of web search from large-scale query log
Lidong Bing, Zhengyu Niu, Piji Li, Wai Lam, Haifeng Wang 0001 |
Knowl. Based Syst. | 2 |
| 2017 | Multi-task Attention-based Neural Networks for Implicit Discourse Relationship Representation and IdentificationabstractWe present a novel multi-task attentionbased neural network model to address implicit discourse relationship representation and identification through two types of representation learning, an attentionbased neural network for learning discourse relationship representation with two arguments and a multi-task framework for learning knowledge from annotated and unannotated corpora.The extensive experiments have been performed on two benchmark corpora (i.e., PDTB and CoNLL-2016 datasets).Experimental results show that our proposed model outperforms the state-of-the-art systems on benchmark corpora. Man Lan, Jianxiang Wang, Yuanbin Wu, Zhengyu Niu, Haifeng Wang 0001 |
EMNLP | 4 |
| 2015 | Integrating word embeddings and traditional NLP features to measure textual entailment and semantic relatedness of sentence pairsabstractRecent years the distributed representations of words (i.e., word embeddings) have been shown to be able to significantly improve performance in many natural language processing tasks, such as pos-of-tag tagging, chunking, named entity recognition and sentiment polarity judgement, etc. However, previous tasks only involve a single sentence. In contrast, this paper evaluates the effectiveness of word embeddings in sentence pair classification or regression problems. Specifically, we propose novel simple yet effective features based on word embeddings and extract many traditional linguistic features. Then these features serve as input of a classification/regression algorithm in isolation and in combination. Evaluations are conducted on three sentence pair classification/regression tasks, i.e., textual entailment, cross-lingual textual entailment and semantic relatedness estimation. Experiments on benchmark datasets provided by Semantic Evaluation 2013 and 2014 showed that using word embeddings is able to significantly improve the performance and our results outperform the best achieved results so far. Jiang Zhao, Man Lan, Zhengyu Niu, Yue Lu 0001 |
IJCNN | 3 |
| 2014 | Recognizing cross-lingual textual entailment with co-training using similarity and difference viewsabstractCross-lingual textual entailment is a relatively new problem that detects the entailment relationship between two text fragments written in different languages. Previous work adopted machine learning algorithms and similarity measures as features to address this task. In order to overcome the high cost of human annotation and further improve the recognition performance, we present a novel co-training approach to solve this problem. We first use an off-the-shelf machine translation tool to eliminate the language gap between two texts. Then we measure the similarities and differences between two texts and regard them as sufficient and redundant views. We use those two views to conduct the co-training procedure to perform classification. Besides, a new effective Kullback-Leibler (KL) based criterion is proposed to select the results from all possible iterations. Experiments on cross-lingual datasets provided by SemEval 2013 show that our method significantly outperforms the baseline systems and previous work. Jiang Zhao, Man Lan, Zhengyu Niu, Donghong Ji |
IJCNN | 3 |
| 2014 | Web page segmentation with structured prediction and its application in web page classificationabstractWe propose a framework which can perform Web page segmentation with a structured prediction approach. It formulates the segmentation task as a structured labeling problem on a transformed Web page segmentation graph (WPS-graph). WPS-graph models the candidate segmentation boundaries of a page and the dependency relation among the adjacent segmentation boundaries. Each labeling scheme on the WPS-graph corresponds to a possible segmentation of the page. The task of finding the optimal labeling of the WPS-graph is transformed into a binary Integer Linear Programming problem, which considers the entire WPS-graph as a whole to conduct structured prediction. A learning algorithm based on the structured output Support Vector Machine framework is developed to determine the feature weights, which is capable to consider the inter-dependency among candidate segmentation boundaries. Furthermore, we investigate its efficacy in supporting the development of automatic Web page classification. Lidong Bing, Wai Lam, Zhengyu Niu, Haifeng Wang 0001 |
SIGIR | 4 |
| 2013 | Leveraging Synthetic Discourse Data via Multi-task Learning for Implicit Discourse Relation Recognition
Man Lan, Zhengyu Niu |
ACL (1) | 3 |
| 2012 | Connective prediction using machine learning for implicit discourse relation classificationabstractImplicit discourse relation classification is a challenge task due to missing discourse connective. Some work directly adopted machine learning algorithms and linguistically informed features to address this task. However, one interesting solution is to automatically predict implicit discourse connective. In this paper, we present a novel two-step machine learning-based approach to implicit discourse relation classification. We first use machine learning method to automatically predict the discourse connective that can best express the implicit discourse relation. Then the predicted implicit discourse connective is used to classify the implicit discourse relation. Experiments on Penn Discourse Treebank 2.0 (PDTB) and Biomedical Discourse Relation Bank (BioDRB) show that our method performs better than the baseline system and previous work. Man Lan, Yue Lu 0001, Zhengyu Niu, Chew Lim Tan |
IJCNN | 4 |
| 2010 | The Effects of Discourse Connectives Prediction on Implicit Discourse Relation Recognition
Zhi-Min Zhou, Man Lan, Zhengyu Niu, Jian Su 0002 |
SIGDIAL Conference | 3 |
| 2009 | Exploiting Heterogeneous Treebanks for Parsing
Zhengyu Niu, Haifeng Wang 0001, Hua Wu 0003 |
ACL/IJCNLP | 1 |
| 2007 | Learning model order from labeled and unlabeled data for partially supervised classification, with application to word sense disambiguation
Zhengyu Niu, Donghong Ji, Chew Lim Tan |
Comput. Speech Lang. | 1 |
| 2007 | Using cluster validation criterion to identify optimal feature subset and cluster number for document clustering
Zhengyu Niu, Donghong Ji, Chew Lim Tan |
Inf. Process. Manag. | 1 |
| 2006 | Relation Extraction Using Label Propagation Based Semi-Supervised LearningabstractShortage of manually labeled data is an obstacle to supervised relation extraction methods. In this paper we investigate a graph based semi-supervised learning algorithm, a label propagation (LP) algorithm, for relation extraction. It represents labeled and unlabeled examples and their distances as the nodes and the weights of edges of a graph, and tries to obtain a labeling function to satisfy two constraints: 1) it should be fixed on the labeled nodes, 2) it should be smooth on the whole graph. Experiment results on the ACE corpus showed that this LP algorithm achieves better performance than SVM when only very few labeled examples are available, and it also performs better than bootstrapping for the relation extraction task. Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu |
ACL | 4 |
| 2006 | Unsupervised Relation Disambiguation Using Spectral Clustering
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu |
ACL | 4 |
| 2006 | Unsupervised Relation Disambiguation with Order Identification Capabilities
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu |
EMNLP | 4 |
| 2006 | Partially Supervised Sense Disambiguation by Learning Sense Number from Tagged and Untagged Corpora
Zhengyu Niu, Donghong Ji, Chew Lim Tan |
EMNLP | 1 |
| 2006 | Semi-supervised Relation Extraction with Label Propagation
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu |
HLT-NAACL | 4 |
| 2005 | Word Sense Disambiguation Using Label Propagation Based Semi-Supervised LearningabstractShortage of manually sense-tagged data is an obstacle to supervised word sense disambiguation methods. In this paper we investigate a label propagation based semi-supervised learning algorithm for WSD, which combines labeled and unlabeled data in learning process to fully realize a global consistency assumption: similar examples should have similar labels. Our experimental results on benchmark corpora indicate that it consistently outperforms SVM when only very few labeled examples are available, and its performance is also better than monolingual bootstrapping, and comparable to bilingual bootstrapping. Zhengyu Niu, Donghong Ji, Chew Lim Tan |
ACL | 1 |
| 2005 | Word Sense Disambiguation by Semi-supervised Learning
Zhengyu Niu, Donghong Ji, Chew Lim Tan, Lingpeng Yang |
CICLing | 1 |
| 2005 | Automatic Relation Extraction with Model Order Selection and Discriminative Label Identification
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu |
IJCNLP | 4 |
| 2005 | Chinese information retrieval based on terms and relevant termsabstractIn this article we describe our approach to Chinese information retrieval, where a query is a short natural language description. First, we use automatically extracted short terms from document sets to build indexes and use the short terms in both the query and documents to do initial retrieval. Next, we use long terms extracted from the document collection to reorder the top N retrieved documents to improve precision. Finally, we acquire the relevant terms of the short terms from the Internet and the top retrieved documents and use them to do query expansion. Experiments on the NTCIR-4 CLIR Chinese SLIR sub-collection show that document reranking can both improve the retrieval performance on its own and make a significant contribution to query expansion. The experiments also show that the extended query expansion proposed in this article is more effective than the standard Rocchio query expansion. Lingpeng Yang, Donghong Ji, Zhengyu Niu |
ACM Trans. Asian Lang. Inf. Process. | 4 |
| 2004 | Learning Word Sense With Feature Selection and Order Identification CapabilitiesabstractThis paper presents an unsupervised word sense learning algorithm, which induces senses of target word by grouping its occurrences into a "natural" number of clusters based on the similarity of their contexts. For removing noisy words in feature set, feature selection is conducted by optimizing a cluster validation criterion subject to some constraint in an unsupervised manner. Gaussian mixture model and Minimum Description Length criterion are used to estimate cluster structure and cluster number. Experimental results show that our algorithm can find important feature subset, estimate model order (cluster number) and achieve better performance than another algorithm which requires cluster number to be provided. Zhengyu Niu, Donghong Ji, Chew Lim Tan |
ACL | 1 |
| 2004 | Feature Selection for Chinese Character Sense Discrimination
Zhengyu Niu, Donghong Ji |
CICLing | 1 |
| 2004 | Document clustering based on cluster validationabstractInternational Conference on Information and Knowledge Management, Proceedings Zhengyu Niu, Donghong Ji, Chew Lim Tan |
CIKM | 1 |
| 2003 | Microsoft Mulan - a bilingual TTS systemabstractThis paper describes a bilingual text-to-speech (TTS) system, Microsoft Mulan, which switches between Mandarin and English smoothly and which maintains the sentence level intonation even for mixed-lingual texts. Mulan is constructed on the basis of the Soft Prediction Only prosodic strategy and the Prosodic-Constraint Orient unit-selection strategy. The unit-selection module of Mulan is shared across languages. It is insensitive to language identity, even though the syllable is used as the smallest unit in Mandarin, and the phoneme in English. Mulan has a unique module, the language-dispatching module, which dispatches texts to the language-specific front-ends and merges the outputs of the two front-ends together. The mixed texts are "uttered" out with the same voice. According to our informal listening test, the speech synthesized with Mulan sounds quite natural. Sample waves can be heard at: http://research.microsoft.com/-echang/proiects/tts/mulan.htm. Min Chu, Hu Peng, Yong Zhao 0008, Zhengyu Niu, Eric Chang |
ICASSP (1) | 4 |
| 2000 | Segmentation of prosodic phrases for improving the naturalness of synthesized Mandarin Chinese speech
Zhengyu Niu, Peiqi Chai |
INTERSPEECH | 1 |