Zhengyu Niu

dblp:31/9311 · also Zheng-Yu Niu · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
abstract
The advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i.e., RPAs forget their role) and style drift (i.e., overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift.
Yihong Tang, Kehai Chen, Muyun Yang, Zhengyu Niu, Tiejun Zhao, Min Zhang 0005
NeurIPS4
2025 Towards few-shot mixed-type dialogue generation
Zeming Liu, Haifeng Wang 0001, Zeyang Lei, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
Sci. China Inf. Sci.4
2024 Learning to Select External Knowledge With Multi-Scale Negative Sampling
abstract
The Track-1 of DSTC9 aims to effectively answer user requests or questions during task-oriented dialogues, which are out of the scope of APIs/DB. By leveraging external knowledge resources, relevant information can be retrieved and encoded into the response generation for these out-of-API-coverage queries. In this work, we have explored several advanced techniques to enhance the utilization of external knowledge and boost the quality of response generation, includingschema guided knowledge decision,negatives enhanced knowledge selection, andknowledge grounded response generation. To evaluate the performance of our proposed method, comprehensive experiments have been carried out on the publicly available dataset. Our approach was ranked as the best in human evaluation of DSTC9 Track-1.
Huang He, Hua Lu 0014, Siqi Bao, Fan Wang 0021, Hua Wu 0003, Zhengyu Niu, Haifeng Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 XDailyDialog: A Multilingual Parallel Dialogue Corpus
abstract
Zeming Liu, Ping Nie, Jie Cai, Haifeng Wang, Zheng-Yu Niu, Peng Zhang, Mrinmaya Sachan, Kaiping Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zeming Liu, Ping Nie, Haifeng Wang 0001, Zhengyu Niu, Mrinmaya Sachan, Kaiping Peng
ACL (1)5
2023 Graph-Grounded Goal Planning for Conversational Recommendation
abstract
Conversational recommendation casts the recommendation problem as a dialog-based interactive task, which could acquire user interest more efficiently and effectively by allowing users to express what they like. In this work, we move a step towards a new conversational recommendation task that is more suitable for real-world applications. In this task, the recommender proactively and naturally lead a dialog from non-recommendation content to approach an item being of interest to users, and allow users to ask questions for better support of user decisions. The challenge of this task lies in how to effectively control the dialog flow to complete the recommendation while appropriately responding to user utterances. To address this challenge, we first construct a Chinese recommendation dialog dataset DuRecDial. We then propose a two-stage Multi-Goal driven Conversation Generation framework, MGCG. In particular, the goal planning module leverages the global graph structure information and local goal-sequence information to effectively control the dialog flow step by step. The goal-guided responding module can produce an in-depth dialog about each goal by fully exploiting hierarchical goal information for response retrieval or generation. Results on DuRecDial demonstrate that MGCG can lead the dialog more proactively and naturally, and complete the recommendation task more effectively.
Zeming Liu, Hao Liu 0026, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.5
2022 Where to Go for the Holidays: Towards Mixed-Type Dialogs for Clarification of User Goals
abstract
Most dialog systems posit that users have figured out clear and specific goals before starting an interaction.For example, users have determined the departure, the destination, and the travel time for booking a flight.However, in many scenarios, limited by experience and knowledge, users may know what they need, but still struggle to figure out clear and specific goals by determining all the necessary slots.
Zeming Liu, Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003
ACL (1)5
2022 CDConv: A Benchmark for Contradiction Detection in Chinese Conversations
abstract
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Zhen Guo, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Minlie Huang. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Chujie Zheng, Jinfeng Zhou, Yinhe Zheng, Libiao Peng, Wenquan Wu, Zhengyu Niu, Hua Wu 0003, Minlie Huang
EMNLP7
2021 Discovering Dialog Structure Graph for Coherent Dialog Generation
abstract
Jun Xu, Zeyang Lei, Haifeng Wang, Zheng-Yu Niu, Hua Wu, Wanxiang Che. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
ACL/IJCNLP (1)4
2021 DuRecDial 2.0: A Bilingual Parallel Corpus for Conversational Recommendation
abstract
In this paper, we provide a bilingual parallel human-to-human recommendation dialog dataset (DuRecDial 2.0) to enable researchers to explore a challenging task of multilingual and cross-lingual conversational recommendation.The difference between DuRecDial 2.0 and existing conversational recommendation datasets is that the data item (Profile, Goal, Knowledge, Context, Response) in DuRecDial 2.0 is annotated in two languages, both English and Chinese, while other datasets are built with the setting of a single language.We collect 8.2k dialogs aligned across English and Chinese languages (16.5k dialogs and 255k utterances in total) that are annotated by crowdsourced workers with strict quality control procedure.We then build monolingual, multilingual, and cross-lingual conversational recommendation baselines on DuRecDial 2.0.Experiment results show that the use of additional English data can bring performance improvement for Chinese conversational recommendation, indicating the benefits of DuRecDial 2.0.Finally, this dataset provides a challenging testbed for future studies of monolingual, multilingual, and cross-lingual conversational recommendation. 1
Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
EMNLP (1)3
2021 Multi-modal visual adversarial Bayesian personalized ranking model for recommendation
Guangli Li, Jianwu Zhuo, Chuanxiu Li, Jin Hua, Zhengyu Niu, Donghong Ji, Renzhong Wu, Hongbin Zhang 0004
Inf. Sci.6
2021 Knowledge graph embedding with shared latent semantic units
Zhao Zhang 0011, Fuzhen Zhuang, Meng Qu, Zhengyu Niu, Hui Xiong 0001, Qing He 0003
Neural Networks4
2021 Coherent Dialog Generation with Query Graph
abstract
Learning to generate coherent and informative dialogs is an enduring challenge for open-domain conversation generation. Previous work leverage knowledge graph or documents to facilitate informative dialog generation, with little attention on dialog coherence. In this article, to enhance multi-turn open-domain dialog coherence, we propose to leverage a new knowledge source, web search session data, to facilitate hierarchical knowledge sequence planning, which determines a sketch of a multi-turn dialog. Specifically, we formulate knowledge sequence planning or dialog policy learning as a graph grounded Reinforcement Learning (RL) problem. To this end, we first build a two-level query graph with queries as utterance-level vertices and their topics (entities in queries) as topic-level vertices. We then present a two-level dialog policy model that plans a high-level topic sequence and a low-level query sequence over the query graph to guide a knowledge aware response generator. In particular, to foster forward-looking knowledge planning decisions for better dialog coherence, we devise a heterogeneous graph neural network to incorporate neighbouring vertex information, or possible future RL action information, into each vertex (as an RL action) representation. Experiment results on two benchmark dialog datasets demonstrate that our framework can outperform strong baselines in terms of dialog coherence, informativeness, and engagingness.
Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Jizhou Huang, Ting Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2020 Knowledge Graph Grounded Goal Planning for Open-Domain Conversation Generation
abstract
Previous neural models on open-domain conversation generation have no effective mechanisms to manage chatting topics, and tend to produce less coherent dialogs. Inspired by the strategies in human-human dialogs, we divide the task of multi-turn open-domain conversation generation into two sub-tasks: explicit goal (chatting about a topic) sequence planning and goal completion by topic elaboration. To this end, we propose a three-layer Knowledge aware Hierarchical Reinforcement Learning based Model (KnowHRL). Specifically, for the first sub-task, the upper-layer policy learns to traverse a knowledge graph (KG) in order to plan a high-level goal sequence towards a good balance between dialog coherence and topic consistency with user interests. For the second sub-task, the middle-layer policy and the lower-layer one work together to produce an in-depth multi-turn conversation about a single topic with a goal-driven generation mechanism. The capability of goal-sequence planning enables chatbots to conduct proactive open-domain conversations towards recommended topics, which has many practical applications. Experiments demonstrate that our model outperforms state of the art baselines in terms of user-interest consistency, dialog coherence, and knowledge accuracy.
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
AAAI3
2020 Towards Conversational Recommendation over Multi-Type Dialogs
abstract
We focus on the study of conversational recommendation in the context of multi-type dialogs, where the bots can proactively and naturally lead a conversation from a nonrecommendation dialog (e.g., QA) to a recommendation dialog, taking into account user's interests and feedback.To facilitate the study of this task, we create a human-to-human Chinese dialog dataset DuRecDial (about 10k dialogs, 156k utterances), which contains multiple sequential dialogs for every pair of a recommendation seeker (user) and a recommender (bot).In each dialog, the recommender proactively leads a multi-type dialog to approach recommendation targets and then makes multiple recommendations with rich interaction behavior.This dataset allows us to systematically investigate different parts of the overall problem, e.g., how to naturally lead a dialog, how to interact with users for recommendation.Finally we establish baseline results on DuRecDial for future studies. 1
Zeming Liu, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001
ACL3
2020 Conversational Graph Grounded Policy Learning for Open-Domain Conversation Generation
abstract
To address the challenge of policy learning in open-domain multi-turn conversation, we propose to represent prior information about dialog transitions as a graph and learn a graph grounded dialog policy, aimed at fostering a more coherent and controllable dialog.To this end, we first construct a conversational graph (CG) from dialog corpora, in which there are vertices to represent "what to say" and "how to say", and edges to represent natural transition between a message (the last utterance in a dialog context) and its response.We then present a novel CG grounded policy learning framework that conducts dialog flow planning by graph traversal, which learns to identify a what-vertex and a how-vertex from the CG at each turn to guide response generation.In this way, we effectively leverage the CG to facilitate policy learning as follows: (1) it enables more effective long-term reward design, (2) it provides high-quality candidate actions, and (3) it gives us more control over the policy.Results on two benchmark corpora demonstrate the effectiveness of this framework.
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che, Ting Liu 0001
ACL3
2020 Enhancing Dialog Coherence with Event Graph Grounded Content Planning
abstract
How to generate informative, coherent and sustainable open-domain conversations is a non-trivial task. Previous work on knowledge grounded conversation generation focus on improving dialog informativeness with little attention on dialog coherence. In this paper, to enhance multi-turn dialog coherence, we propose to leverage event chains to help determine a sketch of a multi-turn dialog. We first extract event chains from narrative texts and connect them as a graph. We then present a novel event graph grounded Reinforcement Learning (RL) framework. It conducts high-level response content (simply an event) planning by learning to walk over the graph, and then produces a response conditioned on the planned content. In particular, we devise a novel multi-policy decision making mechanism to foster a coherent dialog with both appropriate content ordering and high contextual relevance. Experimental results indicate the effectiveness of this framework in terms of dialog coherence and informativeness.
Jun Xu 0027, Zeyang Lei, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
IJCAI4
2020 DE-Ada*: A novel model for breast mass classification using cross-modal pathological semantic mining and organic integration of multi-feature fusions
Hongbin Zhang 0004, Renzhong Wu, Ziliang Jiang, Jinpeng Wu, Jin Hua, Zhengyu Niu, Donghong Ji
Inf. Sci.8
2019 Knowledge Aware Conversation Generation with Explainable Reasoning over Augmented Graphs
abstract
Zhibin Liu, Zheng-Yu Niu, Hua Wu, Haifeng Wang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Zhengyu Niu, Hua Wu 0003, Haifeng Wang 0001
EMNLP/IJCNLP (1)2
2019 A Key-Phrase Aware End2end Neural Response Generation Model
Jun Xu 0027, Haifeng Wang 0001, Zhengyu Niu, Hua Wu 0003, Wanxiang Che
NLPCC (2)3
2019 Knowledge triple mining via multi-task learning
Zhao Zhang 0011, Fuzhen Zhuang, Xuebing Li, Zhengyu Niu, Jia He 0001, Qing He 0003, Hui Xiong 0001
Inf. Syst.4
2019 Supervised representation learning for multi-label classification
Fuzhen Zhuang, Xiao Zhang 0015, Xiang Ao 0001, Zhengyu Niu, Min-Ling Zhang, Qing He 0003
Mach. Learn.5
2018 MultiE: Multi-Task Embedding for Knowledge Base Completion
abstract
Completing knowledge bases (KBs) with missing facts is of great importance, since most existing KBs are far from complete. To this end, many knowledge base completion (KBC) methods have been proposed. However, most existing methods embed each relation into a vector separately, while ignoring the correlations among different relations. Actually, in large-scale KBs, there always exist some relations that are semantically related, and we believe this can help to facilitate the knowledge sharing when learning the embedding of related relations simultaneously. Along this line, we propose a novel KBC model by Multi -Task E mbedding, named MultiE. In this model, semantically related relations are first clustered into the same group, and then learning the embedding of each relation can leverage the knowledge among different relations. Moreover, we propose a three-layer network to predict the missing values of incomplete knowledge triples. Finally, experiments on three popular benchmarks FB15k, FB15k-237 and WN18 are conducted to demonstrate the effectiveness of MultiE against some state-of-the-art baseline competitors.
Zhao Zhang 0011, Fuzhen Zhuang, Zhengyu Niu, Deqing Wang 0001, Qing He 0003
CIKM3
2018 Learning a unified embedding space of web search from large-scale query log
Lidong Bing, Zhengyu Niu, Piji Li, Wai Lam, Haifeng Wang 0001
Knowl. Based Syst.2
2017 Multi-task Attention-based Neural Networks for Implicit Discourse Relationship Representation and Identification
abstract
We present a novel multi-task attentionbased neural network model to address implicit discourse relationship representation and identification through two types of representation learning, an attentionbased neural network for learning discourse relationship representation with two arguments and a multi-task framework for learning knowledge from annotated and unannotated corpora.The extensive experiments have been performed on two benchmark corpora (i.e., PDTB and CoNLL-2016 datasets).Experimental results show that our proposed model outperforms the state-of-the-art systems on benchmark corpora.
Man Lan, Jianxiang Wang, Yuanbin Wu, Zhengyu Niu, Haifeng Wang 0001
EMNLP4
2015 Integrating word embeddings and traditional NLP features to measure textual entailment and semantic relatedness of sentence pairs
abstract
Recent years the distributed representations of words (i.e., word embeddings) have been shown to be able to significantly improve performance in many natural language processing tasks, such as pos-of-tag tagging, chunking, named entity recognition and sentiment polarity judgement, etc. However, previous tasks only involve a single sentence. In contrast, this paper evaluates the effectiveness of word embeddings in sentence pair classification or regression problems. Specifically, we propose novel simple yet effective features based on word embeddings and extract many traditional linguistic features. Then these features serve as input of a classification/regression algorithm in isolation and in combination. Evaluations are conducted on three sentence pair classification/regression tasks, i.e., textual entailment, cross-lingual textual entailment and semantic relatedness estimation. Experiments on benchmark datasets provided by Semantic Evaluation 2013 and 2014 showed that using word embeddings is able to significantly improve the performance and our results outperform the best achieved results so far.
Jiang Zhao, Man Lan, Zhengyu Niu, Yue Lu 0001
IJCNN3
2014 Recognizing cross-lingual textual entailment with co-training using similarity and difference views
abstract
Cross-lingual textual entailment is a relatively new problem that detects the entailment relationship between two text fragments written in different languages. Previous work adopted machine learning algorithms and similarity measures as features to address this task. In order to overcome the high cost of human annotation and further improve the recognition performance, we present a novel co-training approach to solve this problem. We first use an off-the-shelf machine translation tool to eliminate the language gap between two texts. Then we measure the similarities and differences between two texts and regard them as sufficient and redundant views. We use those two views to conduct the co-training procedure to perform classification. Besides, a new effective Kullback-Leibler (KL) based criterion is proposed to select the results from all possible iterations. Experiments on cross-lingual datasets provided by SemEval 2013 show that our method significantly outperforms the baseline systems and previous work.
Jiang Zhao, Man Lan, Zhengyu Niu, Donghong Ji
IJCNN3
2014 Web page segmentation with structured prediction and its application in web page classification
abstract
We propose a framework which can perform Web page segmentation with a structured prediction approach. It formulates the segmentation task as a structured labeling problem on a transformed Web page segmentation graph (WPS-graph). WPS-graph models the candidate segmentation boundaries of a page and the dependency relation among the adjacent segmentation boundaries. Each labeling scheme on the WPS-graph corresponds to a possible segmentation of the page. The task of finding the optimal labeling of the WPS-graph is transformed into a binary Integer Linear Programming problem, which considers the entire WPS-graph as a whole to conduct structured prediction. A learning algorithm based on the structured output Support Vector Machine framework is developed to determine the feature weights, which is capable to consider the inter-dependency among candidate segmentation boundaries. Furthermore, we investigate its efficacy in supporting the development of automatic Web page classification.
Lidong Bing, Wai Lam, Zhengyu Niu, Haifeng Wang 0001
SIGIR4
2013 Leveraging Synthetic Discourse Data via Multi-task Learning for Implicit Discourse Relation Recognition
Man Lan, Zhengyu Niu
ACL (1)3
2012 Connective prediction using machine learning for implicit discourse relation classification
abstract
Implicit discourse relation classification is a challenge task due to missing discourse connective. Some work directly adopted machine learning algorithms and linguistically informed features to address this task. However, one interesting solution is to automatically predict implicit discourse connective. In this paper, we present a novel two-step machine learning-based approach to implicit discourse relation classification. We first use machine learning method to automatically predict the discourse connective that can best express the implicit discourse relation. Then the predicted implicit discourse connective is used to classify the implicit discourse relation. Experiments on Penn Discourse Treebank 2.0 (PDTB) and Biomedical Discourse Relation Bank (BioDRB) show that our method performs better than the baseline system and previous work.
Man Lan, Yue Lu 0001, Zhengyu Niu, Chew Lim Tan
IJCNN4
2010 The Effects of Discourse Connectives Prediction on Implicit Discourse Relation Recognition
Zhi-Min Zhou, Man Lan, Zhengyu Niu, Jian Su 0002
SIGDIAL Conference3
2009 Exploiting Heterogeneous Treebanks for Parsing
Zhengyu Niu, Haifeng Wang 0001, Hua Wu 0003
ACL/IJCNLP1
2007 Learning model order from labeled and unlabeled data for partially supervised classification, with application to word sense disambiguation
Zhengyu Niu, Donghong Ji, Chew Lim Tan
Comput. Speech Lang.1
2007 Using cluster validation criterion to identify optimal feature subset and cluster number for document clustering
Zhengyu Niu, Donghong Ji, Chew Lim Tan
Inf. Process. Manag.1
2006 Relation Extraction Using Label Propagation Based Semi-Supervised Learning
abstract
Shortage of manually labeled data is an obstacle to supervised relation extraction methods. In this paper we investigate a graph based semi-supervised learning algorithm, a label propagation (LP) algorithm, for relation extraction. It represents labeled and unlabeled examples and their distances as the nodes and the weights of edges of a graph, and tries to obtain a labeling function to satisfy two constraints: 1) it should be fixed on the labeled nodes, 2) it should be smooth on the whole graph. Experiment results on the ACE corpus showed that this LP algorithm achieves better performance than SVM when only very few labeled examples are available, and it also performs better than bootstrapping for the relation extraction task.
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu
ACL4
2006 Unsupervised Relation Disambiguation Using Spectral Clustering
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu
ACL4
2006 Unsupervised Relation Disambiguation with Order Identification Capabilities
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu
EMNLP4
2006 Partially Supervised Sense Disambiguation by Learning Sense Number from Tagged and Untagged Corpora
Zhengyu Niu, Donghong Ji, Chew Lim Tan
EMNLP1
2006 Semi-supervised Relation Extraction with Label Propagation
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu
HLT-NAACL4
2005 Word Sense Disambiguation Using Label Propagation Based Semi-Supervised Learning
abstract
Shortage of manually sense-tagged data is an obstacle to supervised word sense disambiguation methods. In this paper we investigate a label propagation based semi-supervised learning algorithm for WSD, which combines labeled and unlabeled data in learning process to fully realize a global consistency assumption: similar examples should have similar labels. Our experimental results on benchmark corpora indicate that it consistently outperforms SVM when only very few labeled examples are available, and its performance is also better than monolingual bootstrapping, and comparable to bilingual bootstrapping.
Zhengyu Niu, Donghong Ji, Chew Lim Tan
ACL1
2005 Word Sense Disambiguation by Semi-supervised Learning
Zhengyu Niu, Donghong Ji, Chew Lim Tan, Lingpeng Yang
CICLing1
2005 Automatic Relation Extraction with Model Order Selection and Discriminative Label Identification
Jinxiu Chen, Donghong Ji, Chew Lim Tan, Zhengyu Niu
IJCNLP4
2005 Chinese information retrieval based on terms and relevant terms
abstract
In this article we describe our approach to Chinese information retrieval, where a query is a short natural language description. First, we use automatically extracted short terms from document sets to build indexes and use the short terms in both the query and documents to do initial retrieval. Next, we use long terms extracted from the document collection to reorder the top N retrieved documents to improve precision. Finally, we acquire the relevant terms of the short terms from the Internet and the top retrieved documents and use them to do query expansion. Experiments on the NTCIR-4 CLIR Chinese SLIR sub-collection show that document reranking can both improve the retrieval performance on its own and make a significant contribution to query expansion. The experiments also show that the extended query expansion proposed in this article is more effective than the standard Rocchio query expansion.
Lingpeng Yang, Donghong Ji, Zhengyu Niu
ACM Trans. Asian Lang. Inf. Process.4
2004 Learning Word Sense With Feature Selection and Order Identification Capabilities
abstract
This paper presents an unsupervised word sense learning algorithm, which induces senses of target word by grouping its occurrences into a "natural" number of clusters based on the similarity of their contexts. For removing noisy words in feature set, feature selection is conducted by optimizing a cluster validation criterion subject to some constraint in an unsupervised manner. Gaussian mixture model and Minimum Description Length criterion are used to estimate cluster structure and cluster number. Experimental results show that our algorithm can find important feature subset, estimate model order (cluster number) and achieve better performance than another algorithm which requires cluster number to be provided.
Zhengyu Niu, Donghong Ji, Chew Lim Tan
ACL1
2004 Feature Selection for Chinese Character Sense Discrimination
Zhengyu Niu, Donghong Ji
CICLing1
2004 Document clustering based on cluster validation
abstract
International Conference on Information and Knowledge Management, Proceedings
Zhengyu Niu, Donghong Ji, Chew Lim Tan
CIKM1
2003 Microsoft Mulan - a bilingual TTS system
abstract
This paper describes a bilingual text-to-speech (TTS) system, Microsoft Mulan, which switches between Mandarin and English smoothly and which maintains the sentence level intonation even for mixed-lingual texts. Mulan is constructed on the basis of the Soft Prediction Only prosodic strategy and the Prosodic-Constraint Orient unit-selection strategy. The unit-selection module of Mulan is shared across languages. It is insensitive to language identity, even though the syllable is used as the smallest unit in Mandarin, and the phoneme in English. Mulan has a unique module, the language-dispatching module, which dispatches texts to the language-specific front-ends and merges the outputs of the two front-ends together. The mixed texts are "uttered" out with the same voice. According to our informal listening test, the speech synthesized with Mulan sounds quite natural. Sample waves can be heard at: http://research.microsoft.com/-echang/proiects/tts/mulan.htm.
Min Chu, Hu Peng, Yong Zhao 0008, Zhengyu Niu, Eric Chang
ICASSP (1)4
2000 Segmentation of prosodic phrases for improving the naturalness of synthesized Mandarin Chinese speech
Zhengyu Niu, Peiqi Chai
INTERSPEECH1