Cheng Niu

dblp:48/2002 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 6 first-author · 6 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
Question answering and dialogue systems · 29% Language models and text generation · 22% Speech recognition and synthesis · 19%
Databases, data mining, and information retrieval
6 papers
Information retrieval · 49% Web and social media mining · 27% Query processing and optimization · 12%

Topics — the 30 heaviest of 59, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
dialogue generation
1.232020
MovieChats: Chat like Humans in a Closed Domain · EMNLP (1) 2020
Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020
Adaboost with Auto-Evaluation for Conversational Models · IJCAI 2018
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
1.222024
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation · ACL (1) 2024
A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State Tracking · ACL 2020
Natural language and speech › Speech recognition and synthesis › speech synthesis
diffusion-based speech synthesis
1.012026
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis · AAAI 2026
Natural language and speech › Speech recognition and synthesis
speech synthesis
1.012026
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis · AAAI 2026
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
1.012026
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis · AAAI 2026
Query processing and optimization
adaptive sampling
1.012026
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking · ACL (1) 2026
Information retrieval › ranking › search relevance › relevance modeling
contextual relevance
1.012026
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking · ACL (1) 2026
Web and social media mining › misinformation detection › rumor detection
early rumor detection
1.012026
LLM-based Few-Shot Early Rumor Detection with Imitation Agent · KDD (1) 2026
Information retrieval › reranking
LLM-based reranking
1.012026
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking · ACL (1) 2026
Web and social media mining › misinformation detection
rumor detection
1.012026
LLM-based Few-Shot Early Rumor Detection with Imitation Agent · KDD (1) 2026
Information retrieval › reranking
search result re-ranking
1.012026
Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking · ACL (1) 2026
Machine learning › Trustworthy machine learning
hallucination
0.812024
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models · ACL (1) 2024
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation
0.812024
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models · ACL (1) 2024
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.812024
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models · ACL (1) 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
0.812024
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation · ACL (1) 2024
Natural language and speech › Language models and text generation › text generation
data-to-text generation
0.412020
Neural Data-to-Text Generation via Jointly Learning the Segmentation and Correspondence · ACL 2020
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
response diversity
0.412020
Diversifying Dialogue Generation with Non-Conversational Text · ACL 2020
Natural language and speech › Language models and text generation › text generation › poetry generation
chinese poetry generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Language models and text generation
controllable text generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Question answering and dialogue systems › knowledge-grounded dialogue
document-grounded dialogue
0.412019
Incremental Transformer with Deliberation Decoder for Document Grounded Conversations · ACL (1) 2019
Natural language and speech › Question answering and dialogue systems
multi-turn dialogue
0.412019
Improving Multi-turn Dialogue Modelling with Utterance ReWriter · ACL (1) 2019
Natural language and speech › Language models and text generation › text generation
poetry generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Natural language and speech › Language models and text generation › text generation › text rewriting
sentence rewriting
0.412019
Improving Multi-turn Dialogue Modelling with Utterance ReWriter · ACL (1) 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation · ACL (1) 2019
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.312018
Adaboost with Auto-Evaluation for Conversational Models · IJCAI 2018
Natural language and speech › Question answering and dialogue systems
conversational modeling
0.312018
Adaboost with Auto-Evaluation for Conversational Models · IJCAI 2018
Machine learning › Transfer learning and domain adaptation › cross-domain learning
cross-domain text classification
0.312018
Cross-Domain Labeled LDA for Cross-Domain Text Classification · ICDM 2018
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.312018
Adaboost with Auto-Evaluation for Conversational Models · IJCAI 2018
Natural language and speech › Information extraction and text analysis
text classification
0.312018
Cross-Domain Labeled LDA for Cross-Domain Text Classification · ICDM 2018

Methods — techniques the papers use, named apart from their topics

large language model · 3.8imitation learning · 2.0few-shot learning · 2.0fine-tuning · 1.2uncertainty-aware sampling · 1.0thompson sampling · 1.0teacher-guided sampling · 1.0group relative preference optimization · 1.0diffusion model · 1.0latent variable model · 0.8synthetic dialogue generation · 0.8corpus annotation · 0.8gaussian weight distributions · 0.3bayesian neural network · 0.3discriminative model · 0.3query log mining · 0.1pseudo-relevance feedback · 0.1word co-occurrence statistics · 0.1
YearPublicationVenuePosition
2026 DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
abstract
Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration prediction remained unoptimized. This paper presents DMOSpeech 2, which extends metric optimization to the duration predictor through a reinforcement learning approach. The proposed system implements a novel duration policy framework using group relative preference optimization (GRPO) with speaker similarity and word error rate as reward signals. By optimizing this previously unoptimized component, DMOSpeech 2 creates a more complete metric-optimized synthesis pipeline. Additionally, this paper introduces teacher-guided sampling, a hybrid approach leveraging a teacher model for initial denoising steps before transitioning to the student model, significantly improving output diversity while maintaining efficiency. Comprehensive evaluations demonstrate superior performance across all metrics compared to previous systems, while reducing sampling steps by half without quality degradation. These advances represent a significant step toward speech synthesis systems with metric optimization across multiple components.
Yinghao Aaron Li, Xilin Jiang, Cheng Niu, Kaifeng Xu, Juntong Song, Nima Mesgarani
AAAI4
2026 Contextual Relevance and Adaptive Sampling for LLM-Based Document Reranking
abstract
Reranking algorithms have made progress in improving document retrieval quality by efficiently aggregating relevance judgments generated by large language models (LLMs).However, identifying relevant documents for queries that require in-depth reasoning remains a major challenge.Reasoning-intensive queries often exhibit multifaceted information needs and nuanced interpretations, rendering document relevance inherently context dependent and often noisy.To address this, we propose contextual relevance, which we define as the probability that a document is relevant to a given query, marginalized over the distribution of different reranking contexts it may appear in (i.e., the set of candidate documents it is ranked alongside and the order in which the documents are presented to a reranking model).While prior works have studied methods to mitigate the positional bias LLMs exhibit by accounting for the ordering of documents, we empirically show that batch composition also materially affects relevance judgments.To efficiently estimate contextual relevance, we propose TS-SetRank, a sampling-based, uncertainty-aware reranking algorithm.Empirically, TS-SetRank improves nDCG@10 over retrieval and reranking baselines by 15-25% on BRIGHT and 6-21% on BEIR, highlighting the importance of modeling relevance as context-dependent.
Jerry Huang, Siddarth Madala, Cheng Niu, Julia Hockenmaier, Tong Zhang 0001
ACL (1)3
2026 LLM-based Few-Shot Early Rumor Detection with Imitation Agent
Fengzhu Zeng, Qian Shao, Ling Cheng 0002, Wei Gao 0001, Shih-Fen Cheng, Jing Ma 0004, Cheng Niu
KDD (1)7
2024 Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation
abstract
Dialogue State Tracking (DST) is designed to monitor the evolving dialogue state in the conversations and plays a pivotal role in developing task-oriented dialogue systems.However, obtaining the annotated data for the DST task is usually a costly endeavor.In this paper, we focus on employing LLMs to generate dialogue data to reduce dialogue collection and annotation costs.Specifically, GPT-4 is used to simulate the user and agent interaction, generating thousands of dialogues annotated with DST labels.Then a two-stage fine-tuning on LLaMA 2 is performed on the generated data and the real data for the DST prediction.Experimental results on two public DST benchmarks show that with the generated dialogue data, our model performs better than the baseline trained solely on real data.In addition, our approach is also capable of adapting to the dynamic demands in real-world scenarios, generating dialogues in new domains swiftly.After replacing dialogue segments in any domain with the corresponding generated ones, the model achieves comparable performance to the model trained on real data 1 .
Cheng Niu, Xingguang Wang, Xuxin Cheng, Juntong Song, Tong Zhang 0001
ACL (1)1
2024 RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
abstract
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, Tong Zhang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, Tong Zhang 0001
ACL (1)1
2022 Global Graph Attention Embedding Network for Relation Prediction in Knowledge Graphs
abstract
The incompleteness of knowledge graphs triggers considerable research interest in relation prediction. As the key to predicting relations among entities, many efforts have been devoted to learning the embeddings of entities and relations by incorporating a variety of neighbors' information which includes not only the information from direct outgoing and incoming neighbors but also the ones from the indirect neighbors on the multihop paths. However, previous models usually consider entity paths of limited length or ignore sequential information of the paths. Either simplification will make the model lack a global understanding of knowledge graphs and may result in the loss of important and indispensable information. In this article, we propose a novel global graph attention embedding network (GGAE) for relation prediction by combining global information from both direct neighbors and multihop neighbors. Concretely, given a knowledge graph, we first introduce the path construction algorithms to obtain meaningful paths, then design path modeling methods to capture the potential long-distance sequential information in the multihop paths, final propose an entity graph attention and a relation graph attention mechanisms to obtain entity embeddings and relation embeddings. Moreover, an entity graph attention mechanism is proposed to calculate the entity embeddings by aggregating direct incoming and outgoing neighbors from: 1) an original knowledge graph with the original entity and relation embeddings and 2) a new knowledge graph constructed by the paths whose embeddings are updated by path modeling methods. for each relation, we construct a new graph with related entities and present a relation graph attention to learn the features. Therefore, our model can encapsulate the information from different distance neighbors, and enable the embeddings of entities and relations to better capture all-sided semantic information. The experimental results on benchmark datasets verify the superiority of our model over the state-of-the-art ones.
Qian Li 0043, Daling Wang, Shi Feng 0001, Cheng Niu, Yifei Zhang 0003
IEEE Trans. Neural Networks Learn. Syst.4
2020 A Contextual Hierarchical Attention Network with Adaptive Objective for Dialogue State Tracking
abstract
Recent studies in dialogue state tracking (DST) leverage historical information to determine states which are generally represented as slot-value pairs. However, most of them have limitations to efficiently exploit relevant context due to the lack of a powerful mechanism for modeling interactions between the slot and the dialogue history. Besides, existing methods usually ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots and eventually hurts overall performance. In this paper, we propose to enhance the DST through employing a contextual hierarchical attention network to not only discern relevant information at both word level and turn level but also learn contextual representations. We further propose an adaptive objective to alleviate the slot imbalance problem by dynamically adjust weights of different slots during training. Experimental results show that our approach reaches 52.68% and 58.55% joint accuracy on MultiWOZ 2.0 and MultiWOZ 2.1 datasets respectively and achieves new state-of-the-art performance with considerable improvements (+1.24% and +5.98%).
Yong Shan, Zekang Li, Jinchao Zhang 0001, Fandong Meng, Yang Feng 0004, Cheng Niu, Jie Zhou 0016
ACL6
2020 Neural Data-to-Text Generation via Jointly Learning the Segmentation and Correspondence
abstract
The neural attention model has achieved great success in data-to-text generation tasks. Though usually excelling at producing fluent text, it suffers from the problem of information missing, repetition and "hallucination". Due to the black-box nature of the neural attention architecture, avoiding these problems in a systematic way is non-trivial. To address this concern, we propose to explicitly segment target text into fragment units and align them with their data correspondences. The segmentation and correspondence are jointly learned as latent variables without any human annotations. We further impose a soft statistical constraint to regularize the segmental granularity. The resulting architecture maintains the same expressive power as neural attention models, while being able to generate fully interpretable outputs with several times less computational cost. On both E2E and WebNLG benchmarks, we show the proposed model consistently outperforms its neural attention counterparts.
Xiaoyu Shen 0001, Ernie Chang, Hui Su, Cheng Niu, Dietrich Klakow
ACL4
2020 Diversifying Dialogue Generation with Non-Conversational Text
abstract
Neural network-based sequence-to-sequence (seq2seq) models strongly suffer from the lowdiversity problem when it comes to opendomain dialogue generation.As bland and generic utterances usually dominate the frequency distribution in our daily chitchat, avoiding them to generate more interesting responses requires complex data filtering, sampling techniques or modifying the training objective.In this paper, we propose a new perspective to diversify dialogue generation by leveraging non-conversational text.Compared with bilateral conversations, nonconversational text are easier to obtain, more diverse and cover a much broader range of topics.We collect a large-scale nonconversational corpus from multi sources including forum comments, idioms and book snippets.We further present a training paradigm to effectively incorporate these text via iterative back translation.The resulting model is tested on two conversational datasets and is shown to produce significantly more diverse responses without sacrificing the relevance with context.
Hui Su, Xiaoyu Shen 0001, Sanqiang Zhao, Xiao Zhou 0004, Pengwei Hu 0001, Randy Zhong, Cheng Niu, Jie Zhou 0016
ACL7
2020 Contrastive Zero-Shot Learning for Cross-Domain Slot Filling with Adversarial Attack
abstract
Zero-shot slot filling has widely arisen to cope with data scarcity in target domains.However, previous approaches often ignore constraints between slot value representation and related slot description representation in the latent space and lack enough model robustness.In this paper, we propose a Contrastive Zero-Shot Learning with Adversarial Attack (CZSL-Adv) method for the cross-domain slot filling.The contrastive loss aims to map slot value contextual representations to the corresponding slot description representations.And we introduce an adversarial attack training strategy to improve model robustness.Experimental results show that our model significantly outperforms state-of-the-art baselines under both zero-shot and few-shot settings.
Keqing He 0001, Jinchao Zhang 0001, Yuanmeng Yan, Weiran Xu, Cheng Niu, Jie Zhou 0016
COLING5
2020 MovieChats: Chat like Humans in a Closed Domain
abstract
Being able to perform in-depth chat with humans in a closed domain is a precondition before an open-domain chatbot can ever be claimed.In this work, we take a close look at the movie domain and present a large-scale high-quality corpus with fine-grained annotations in hope of pushing the limit of moviedomain chatbots.We propose a unified, readily scalable neural approach which reconciles all subtasks like intent prediction and knowledge retrieval.The model is first pretrained on the huge general-domain data, then finetuned on our corpus.We show this simple neural approach trained on high-quality data is able to outperform commercial systems replying on complex rules.On both the static and interactive tests, we find responses generated by our system exhibits remarkably good engagement and sensibleness close to human-written ones.We further analyze the limits of our work and point out potential directions for future work 1 .
Hui Su, Xiaoyu Shen 0001, Xiao Zhou 0004, Ernie Chang, Cheng Niu, Jie Zhou 0016
EMNLP (1)7
2019 Incremental Transformer with Deliberation Decoder for Document Grounded Conversations
abstract
Document Grounded Conversations is a task to generate dialogue responses when chatting about the content of a given document. Obviously, document knowledge plays a critical role in Document Grounded Conversations, while existing dialogue models do not exploit this kind of knowledge effectively enough. In this paper, we propose a novel Transformer-based architecture for multi-turn document grounded conversations. In particular, we devise an Incremental Transformer to encode multi-turn utterances along with knowledge in related documents. Motivated by the human cognitive process, we design a two-pass decoder (Deliberation Decoder) to improve context coherence and knowledge correctness. Our empirical study on a real-world Document Grounded Dataset proves that responses generated by our model significantly outperform competitive baselines on both context coherence and knowledge relevance.
Zekang Li, Cheng Niu, Fandong Meng, Yang Feng 0004, Jie Zhou 0016
ACL (1)2
2019 Rhetorically Controlled Encoder-Decoder for Modern Chinese Poetry Generation
abstract
Rhetoric is a vital element in modern poetry, and plays an essential role in improving its aesthetics.However, to date, it has not been considered in research on automatic poetry generation.In this paper, we propose a rhetorically controlled encoder-decoder for modern Chinese poetry generation.Our model relies on a continuous latent variable as a rhetoric controller to capture various rhetorical patterns in an encoder, and then incorporates rhetoricbased mixtures while generating modern Chinese poetry.For metaphor and personification, an automated evaluation shows that our model outperforms state-of-the-art baselines by a substantial margin, while a human evaluation shows that our model generates better poems than baseline methods in terms of fluency, coherence, meaningfulness, and rhetorical aesthetics.
Zuohui Fu, Jie Cao 0010, Gerard de Melo, Yik-Cheung Tam, Cheng Niu, Jie Zhou 0016
ACL (1)6
2019 Improving Multi-turn Dialogue Modelling with Utterance ReWriter
abstract
Recent research has achieved impressive results in single-turn dialogue modelling. In the multi-turn setting, however, current models are still far from satisfactory. One major challenge is the frequently occurred coreference and information omission in our daily conversation, making it hard for machines to understand the real intention. In this paper, we propose rewriting the human utterance as a pre-process to help multi-turn dialgoue modelling. Each utterance is first rewritten to recover all coreferred and omitted information. The next processing steps are then performed based on the rewritten utterance. To properly train the utterance rewriter, we collect a new dataset with human annotations and introduce a Transformer-based utterance rewriting architecture using the pointer network. We show the proposed architecture achieves remarkably good performance on the utterance rewriting task. The trained utterance rewriter can be easily integrated into online chatbots and brings general improvement over different domains.
Hui Su, Xiaoyu Shen 0001, Rongzhi Zhang, Fei Sun 0001, Pengwei Hu 0001, Cheng Niu, Jie Zhou 0016
ACL (1)6
2018 A General Cross-Domain Recommendation Framework via Bayesian Neural Network
abstract
Collaborative filtering is an effective and widely used recommendation approach by applying the user-item rating matrix for recommendations, however, which usually suffers from cold-start and sparsity problems. To address these problems, hybrid methods are proposed to incorporate auxiliary information such as user/item profiles to collaborative filtering models; Cross-domain recommendation systems add a new dimension to solve these problems by leveraging ratings from other domains to improve recommendation performance. Among these methods, deep neural network based recommendation systems achieve excellent performance due to their excellent ability in learning powerful representations. However, these cross-domain recommendation systems based on deep neural network rarely consider the uncertainty of weights. Therefore, they maybe lack of calibrated probabilistic predictions and make overly confident decisions. Along this line, we propose a general cross-domain recommendation framework via Bayesian neural network to incorporate auxiliary information, which takes advantage of both the hybrid recommendation methods and the cross-domain recommendation systems. Specifically, our framework consists of two kinds of neural networks, one to learn the low dimensional representation from the one-hot codings of users/items, while the other one is to project the auxiliary information of users/items into another latent space. The final rating is produced by integrating the latent representations of the one-hot codings of users/items and the auxiliary information of users/items. The latent representations of users learnt from ratings and auxiliary information are shared across different domains for knowledge transfer. Moreover, we capture the uncertainty in all weights by representing weights with Gaussian distributions to make calibrated probabilistic predictions. We have done extensive experiments on real-world data sets to verify the effectiveness of our framework.
Jia He 0001, Rui Liu 0007, Fuzhen Zhuang, Cheng Niu, Qing He 0003
ICDM5
2018 Cross-Domain Labeled LDA for Cross-Domain Text Classification
abstract
Cross-domain text classification aims at building a classifier for a target domain which leverages data from both source and target domain. One promising idea is to minimize the feature distribution differences of the two domains. Most existing studies explicitly minimize such differences by an exact alignment mechanism (aligning features by one-to-one feature alignment, projection matrix etc.). Such exact alignment, however, will restrict models' learning ability and will further impair models' performance on classification tasks when the semantic distributions of different domains are very different. To address this problem, we propose a novel group alignment which aligns the semantics at group level. In addition, to help the model learn better semantic groups and semantics within these groups, we also propose a partial supervision for model's learning in source domain. To this end, we embed the group alignment and a partial supervision into a cross-domain topic model, and propose a Cross-Domain Labeled LDA (CDL-LDA). On the standard 20Newsgroup and Reuters dataset, extensive quantitative (classification, perplexity etc.) and qualitative (topic detection) experiments are conducted to show the effectiveness of the proposed group alignment and partial supervision.
Baoyu Jing, Chenwei Lu, Deqing Wang 0001, Fuzhen Zhuang, Cheng Niu
ICDM5
2018 MiRNN: An Improved Prediction Model of MicroRNA Precursors Using Gated Recurrent Units
Dancheng Li, Zhitao Lin, Cheng Niu, Chen Ding 0005
ICIC (2)4
2018 Adaboost with Auto-Evaluation for Conversational Models
abstract
We propose a boosting method for conversational models to encourage them to generate more human-like dialogs. In our method, we consider existing conversational models as weak generators and apply Adaboost to update those models. However, conventional Adaboost cannot be directly applied on conversational models. Because for conversational models, conventional Adaboost cannot adaptively adjust the weight on the instance for subsequent learning, result from the simple comparison between the true output y (to an input x) and its corresponding predicted output y' cannot directly evaluate the learning performance on x. To address this issue, we develop the Adaboost with Auto-Evaluation (called AwE). In AwE, an auto-evaluator is proposed to evaluate the predicted results, which makes it applicable to conversational models. Furthermore, we present the theoretical analysis that the training error drops exponentially fast only if certain assumption over the proposed auto-evaluator holds. Finally, we empirically show that AwE visibly boosts the performance of existing single conversational models and also outperforms the other ensemble methods for conversational models.
Juncen Li, Ping Luo 0001, Ganbin Zhou, Cheng Niu
IJCAI5
2010 Exploiting query logs for cross-lingual query suggestions
abstract
Query suggestion aims to suggest relevant queries for a given query, which helps users better specify their information needs. Previous work on query suggestion has been limited to the same language. In this article, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to the scenarios of cross-language information retrieval (CLIR) and other related cross-lingual applications. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, and so on, are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly outperforms a baseline system that uses dictionary-based query translation. Besides, we evaluate CLQS with French-English and Chinese-English CLIR tasks on TREC-6 and NTCIR-4 collections, respectively. The CLIR experiments using typical retrieval models demonstrate that the CLQS-based approach has significantly higher effectiveness than several traditional query translation methods. We find that when combined with pseudo-relevance feedback, the effectiveness of CLIR using CLQS is enhanced for different pairs of languages.
Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon
ACM Trans. Inf. Syst.2
2009 Joint Ranking for Multilingual Web Search
Wei Gao 0001, Cheng Niu, Ming Zhou 0001, Kam-Fai Wong
ECIR2
2008 Combining Multiple Resources to Improve SMT-based Paraphrasing Model
Cheng Niu, Ming Zhou 0001, Ting Liu 0001, Sheng Li 0003
ACL2
2008 InfoXtract: A customizable intermediate level information extraction engine
abstract
Abstract Information Extraction (IE) systems assist analysts to assimilate information from electronic documents. This paper focuses on IE tasks designed to support information discovery applications. Since information discovery implies examining large volumes of heterogeneous documents for situations that cannot be anticipated a priori, they require IE systems to have breadth as well as depth. This implies the need for a domain-independent IE system that can easily be customized for specific domains: end users must be given tools to customize the system on their own. It also implies the need for defining new intermediate level IE tasks that are richer than the subject-verb-object (SVO) triples produced by shallow systems, yet not as complex as the domain-specific scenarios defined by the Message Understanding Conference (MUC). This paper describes InfoXtract, a robust, scalable, intermediate-level IE engine that can be ported to various domains. It describes new IE tasks such as synthesis of entity profiles, and extraction of concept-based general events which represent realistic near-term goals focused on deriving useful, actionable information. Entity profiles consolidate information about a person/organization/location etc. within a document and across documents into a single template; this takes into account aliases and anaphoric references as well as key relationships and events pertaining to that entity. Concept-based events attempt to normalize information such as time expressions (e.g., yesterday) as well as ambiguous location references (e.g., Buffalo). These new tasks facilitate the correlation of output from an IE engine with structured data to enable text mining. InfoXtract's hybrid architecture comprised of grammatical processing and machine learning is described in detail. Benchmarking results for the core engine and applications utilizing the engine are presented.
Rohini K. Srihari, Wei Li 0003, Thomas L. Cornell, Cheng Niu
Nat. Lang. Eng.4
2007 Named Entity Translation with Web Mining and Transliteration
Long Jiang, Ming Zhou 0001, Lee-Feng Chien, Cheng Niu
IJCAI4
2007 Cross-lingual query suggestion using query logs of different languages
abstract
Query suggestion aims to suggest relevant queries for a given query, which help users better specify their information needs. Previously, the suggested terms are mostly in the same language of the input query. In this paper, we extend it to cross-lingual query suggestion (CLQS): for a query in one language, we suggest similar or relevant queries in other languages. This is very important to scenarios of cross-language information retrieval (CLIR) and cross-lingual keyword bidding for search engine advertisement. Instead of relying on existing query translation technologies for CLQS, we present an effective means to map the input query of one language to queries of the other language in the query log. Important monolingual and cross-lingual information such as word translation relations and word co-occurrence statistics, etc. are used to estimate the cross-lingual query similarity with a discriminative model. Benchmarks show that the resulting CLQS system significantly out performs a baseline system based on dictionary-based query translation. Besides, the resulting CLQS is tested with French to English CLIR tasks on TREC collections. The results demonstrate higher effectiveness than the traditional query translation methods.
Wei Gao 0001, Cheng Niu, Jian-Yun Nie, Ming Zhou 0001, Kam-Fai Wong, Hsiao-Wuen Hon
SIGIR2
2007 Demographic prediction based on user's browsing behavior
abstract
Demographic information plays an important role in personalized web applications. However, it is usually not easy to obtain this kind of personal data such as age and gender. In this paper, we made a first approach to predict users' gender and age from their Web browsing behaviors, in which the Webpage view information is treated as a hidden variable to propagate demographic information between different users. There are three main steps in our approach: First, learning from the Webpage click-though data, Webpages are associated with users' (known) age and gender tendency through a discriminative model; Second, users' (unknown) age and gender are predicted from the demographic information of the associated Webpages through a Bayesian framework; Third, based on the fact that Webpages visited by similar users may be associated with similar demographic tendency, and users with similar demographic information would visit similar Webpages, a smoothing component is employed to overcome the data sparseness of web click-though log. Experiments are conducted on a real web click-through log to demonstrate the effectiveness of the proposed approach. The experimental results show that the proposed algorithm can achieve up to 30.4% improvements on gender prediction and 50.3% on age prediction in terms of macro F1, compared to baseline algorithms.
Jian Hu 0001, Hua-Jun Zeng, Hua Li 0001, Cheng Niu, Zheng Chen 0001
WWW4
2006 A DOM Tree Alignment Model for Mining Parallel Data from the Web
abstract
This paper presents a new web mining scheme for parallel data acquisition.Based on the Document Object Model (DOM), a web page is represented as a DOM tree.Then a DOM tree alignment model is proposed to identify the translationally equivalent texts and hyperlinks between two parallel DOM trees.By tracing the identified parallel hyperlinks, parallel web documents are recursively mined.Compared with previous mining schemes, the benchmarks show that this new mining scheme improves the mining coverage, reduces mining bandwidth, and enhances the quality of mined parallel sentences.
Cheng Niu, Ming Zhou 0001, Jianfeng Gao 0001
ACL2
2005 Word Independent Context Pair Classification Model for Word Sense Disambiguation
Cheng Niu, Wei Li 0003, Rohini K. Srihari
CoNLL1
2004 Weakly Supervised Learning for Cross-document Person Name Disambiguation Supported by Information Extraction
abstract
It is fairly common that different people are associated with the same name.In tracking person entities in a large document pool, it is important to determine whether multiple mentions of the same name across documents refer to the same entity or not.Previous approach to this problem involves measuring context similarity only based on co-occurring words.This paper presents a new algorithm using information extraction support in addition to co-occurring words.A learning scheme with minimal supervision is developed within the Bayesian framework.Maximum entropy modeling is then used to represent the probability distribution of context similarities based on heterogeneous features.Statistical annealing is applied to derive the final entity coreference chains by globally fitting the pairwise context similarities.Benchmarking shows that our new approach significantly outperforms the existing algorithm by 25 percentage points in overall F-measure.
Cheng Niu, Wei Li 0003, Rohini K. Srihari
ACL1
2003 An Expert Lexicon Approach to Identifying English Phrasal Verbs
abstract
Phrasal Verbs are an important feature of the English language. Properly identifying them provides the basis for an English parser to decode the related structures. Phrasal verbs have been a challenge to Natural Language Processing (NLP) because they sit at the borderline between lexicon and syntax. Traditional NLP frameworks that separate the lexicon module from the parser make it difficult to handle this problem properly. This paper presents a finite state approach that integrates a phrasal verb expert lexicon between shallow parsing and deep parsing to handle morpho-syntactic interaction. With precision/recall combined performance benchmarked consistently at 95.8%-97.5%, the Phrasal Verb identification problem has basically been solved with the presented method.
Wei Li 0003, Xiuhong Zhang, Cheng Niu, Yuankai Jiang, Rohini K. Srihari
ACL3
2003 A Bootstrapping Approach to Named Entity Classification Using Successive Learners
abstract
This paper presents a new bootstrapping approach to named entity (NE) classification. This approach only requires a few common noun/pronoun seeds that correspond to the concept for the target NE type, e.g. he/she/man/woman for PERSON NE. The entire bootstrapping procedure is implemented as training two successive learners: (i) a decision list is used to learn the parsing-based high precision NE rules; (ii) a Hidden Markov Model is then trained to learn string sequence-based NE patterns. The second learner uses the training corpus automatically tagged by the first learner. The resulting NE system approaches supervised NE performance for some NE types. The system also demonstrates intuitive support for tagging user-defined NE types. The differences of this approach from the co-training-based NE bootstrapping are also discussed.
Cheng Niu, Wei Li 0003, Jihong Ding, Rohini K. Srihari
ACL1
2003 A Case Restoration Approach to Named Entity Tagging in Degraded Documents
abstract
This paper describes a novel approach to named entity (NE) tagging on degraded documents. NE tagging is the process of identifying salient text strings in unstructured text, corresponding to names of people, places, organizations, times/dates, etc. Although NE tagging is typically part of a larger information extraction process, it has other applications, such as improving search in an information retrieval system, and post-processing the results of an OCR system. We focus on degraded documents, i.e. case insensitive documents that lack orthographic information. Examples include output of speech recognition systems, as well as e-mail. The traditional approach involves retraining an NE tagger on degraded text, a cumbersome operation. This paper describes an approach whereby text is first “restored ” to its implicit case sensitive form, and subsequently processed by the original NE tagger. Results show that this new approach leads to far less precision loss in NE tagging of degraded documents. 1.
Rohini K. Srihari, Cheng Niu, Wei Li 0003, Jihong Ding
ICDAR2
2003 Bootstrapping for Named Entity Tagging Using Concept-based Seeds
Cheng Niu, Wei Li 0003, Jihong Ding, Rohini K. Srihari
HLT-NAACL1
2002 Location Normalization for Information Extraction
Rohini K. Srihari, Cheng Niu, Wei Li 0003
COLING3