EDBT 2026 Demo / reviewers in the wild / expert
Chunping Li
dblp:69/1535
· DBLP profile ↗
50ranked-venue papers
1as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 7 since 2021Databases, data management, data science and information retrieval · 20 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CSE-SFP: Enabling Unsupervised Sentence Representation Learning via a Single Forward PassabstractAs a fundamental task in Information Retrieval and Computational Linguistics, sentence representation has profound implications for a wide range of practical applications such as text clustering, content analysis, question-answering systems, and web search. Recent advances in pre-trained language models (PLMs) have driven remarkable progress in this field, particularly through unsupervised embedding derivation methods centered on discriminative PLMs like BERT. However, due to time and computational constraints, few efforts have attempted to integrate unsupervised sentence representation with generative PLMs, which typically possess much larger parameter sizes. Given that state-of-the-art models in both academia and industry are predominantly based on generative architectures, there is a pressing need for an efficient unsupervised text representation framework tailored to decoder-only PLMs. To address this concern, we propose CSE-SFP, an innovative method that exploits the structural characteristics of generative models. Compared to existing strategies, CSE-SFP requires only a single forward pass to perform effective unsupervised contrastive learning. Rigorous experimentation demonstrates that CSE-SFP not only produces higher-quality embeddings but also significantly reduces both training time and memory consumption. Furthermore, we introduce two ratio metrics that jointly assess alignment and uniformity, thereby providing a more robust means for evaluating the semantic spatial properties of encoding models. Our code and checkpoints are available at https://github.com/ZBWpro/CSE-SFP. Bowen Zhang 0014, Zixin Song, Chunping Li |
SIGIR | 3 |
| 2025 | Prompting large language model for multi-location multi-step zero-shot wind power forecasting
Zhiyu Duan, Chong Bian, Shunkun Yang, Chunping Li |
Expert Syst. Appl. | 4 |
| 2025 | A Macroscopic-Fundamental-Function-Aided Neural Networks for Traffic Flows Prediction From Mobile Signaling DataabstractThe widespread deployment of the Internet of Things (IoT) allows for the utilization of mobile signaling data (MSD) in cellular networks to perceive the underlying traffic states on road networks. Nevertheless, the application of MSD for traffic flow prediction can be impeded by the limited positioning accuracy of MSD and stringent privacy policies. To surmount these obstacles, we propose a traffic flow prediction method that integrates an estimation process using a macroscopic fundamental function (M) with a positional factor (P), combined with the sequence-to-sequence (S2S) neural networks model (MPS2S). First, we estimate traffic flow using a macroscopic fundamental function (M). In this approach, we introduce a positional factor (P) that fuses road inflows and outflows to capture the relationship between cellular network and road network. Next, we employ an S2S neural network model, which considers the spatial and temporal aggregation information to infer future traffic flows. Specifically, an encoder estimates traffic parameters from MSD, a decoder infers multistep traffic flows, and a coefficients estimation method introduces the macroscopic fundamental function into the neural network. To validate the prediction performance of our proposed model, we collect video data from roadside cameras to measure ground-truth values. In addition, we analyze the impact of the positional factor on our estimation method and the impact of spatial and temporal aggregation on our prediction model in real-world freeway scenarios. The results indicate that MSD can link the potential traffic information, and our method contributes to better traffic flow prediction. Xiongwei Wu, Mengyun Xu, Chunping Li, Xuesong Wu 0004, Li Li 0082 |
IEEE Internet Things J. | 4 |
| 2025 | Turn Waste Into Wealth: On Efficient Clustering and Cleaning Over Dirty DataabstractDirty data commonly exist. Simply discarding a large number of inaccurate points (as noises) could greatly affect clustering results. We argue that dirty data can be repaired and utilized as strong supports in clustering. To this end, we study a novel problem of clustering and repairing over dirty data at the same time. Referring to the minimum change principle in data repairing, the objective is to find a minimum modification of inaccurate points such that the large amount of dirty data can enhance clustering. We show that the problem isNP-hard and can be formulated as an integer linear programming (ILP) problem. A constant factor approximation algorithmGDORCis devised based on grid, with high efficiency. In experiments,GDORChas great repairing and clustering results with low time consumption. Empirical results demonstrate thatboth the clustering and cleaning accuraciescan be improved by our approach of repairing and utilizing the dirty data in clustering. Kenny Ye Liang, Yunxiang Su, Shaoxu Song, Chunping Li |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2024 | Reinforcement Learning with Balanced Clinical Reward for Sepsis Treatment
Jingming Liu, Ruihong Luo, Chunping Li |
AIME (1) | 4 |
| 2024 | Advancing Semantic Textual Similarity Modeling: A Regression Framework with Translated ReLU and Smooth K2 LossabstractSince the introduction of BERT and RoBERTa, research on Semantic Textual Similarity (STS) has made groundbreaking progress.Particularly, the adoption of contrastive learning has substantially elevated state-of-the-art performance across various STS benchmarks.However, contrastive learning categorizes text pairs as either semantically similar or dissimilar, failing to leverage fine-grained annotated information and necessitating large batch sizes to prevent model collapse.These constraints pose challenges for researchers engaged in STS tasks that involve nuanced similarity levels or those with limited computational resources, compelling them to explore alternatives like Sentence-BERT.Despite its efficiency, Sentence-BERT tackles STS tasks from a classification perspective, overlooking the progressive nature of semantic relationships, which results in suboptimal performance.To bridge this gap, this paper presents an innovative regression framework and proposes two simple yet effective loss functions: Translated ReLU and Smooth K2 Loss.Experimental results demonstrate that our method achieves convincing performance across seven established STS benchmarks and offers the potential for further optimization of contrastive learning pretrained models. 1 Bowen Zhang 0014, Chunping Li |
EMNLP | 2 |
| 2024 | Pcc-tuning: Breaking the Contrastive Learning Ceiling in Semantic Textual SimilarityabstractSemantic Textual Similarity (STS) constitutes a critical research direction in computational linguistics and serves as a key indicator of the encoding capabilities of embedding models.Driven by advances in pre-trained language models and contrastive learning, leading sentence representation methods have reached an average Spearman's correlation score of approximately 86 across seven STS benchmarks in SentEval.However, further progress has become increasingly marginal, with no existing method attaining an average score higher than 86.5 on these tasks.This paper conducts an in-depth analysis of this phenomenon and concludes that the upper limit for Spearman's correlation scores under contrastive learning is 87.5.To transcend this ceiling, we propose an innovative approach termed Pcc-tuning, which employs Pearson's correlation coefficient as a loss function to refine model performance beyond contrastive learning.Experimental results demonstrate that Pcc-tuning can markedly surpass previous state-of-the-art strategies with only a minimal amount of fine-grained annotated samples. 1 Bowen Zhang 0014, Chunping Li |
EMNLP | 2 |
| 2024 | CoT-BERT: Enhancing Unsupervised Sentence Representation Through Chain-of-Thought
Bowen Zhang 0014, Kehua Chang, Chunping Li |
ICANN (7) | 3 |
| 2024 | COREX: Document-level Relation Extraction Framework with Consistent Two-Hop Reasoning and Evidence Sentence PredictionabstractDocument-level relation extraction (DocRE) focuses on identifying relationships between entities within a document. It differs from sentence-level extraction as it involves understanding relationships over longer distances, including across sentences or paragraphs. A major challenge in DocRE is to focus on crucial evidence and bridge entities, defined as entities that link others in relationships spanning multiple sentences. Therefore, we propose a novel, efficient representation to support two-hop reasoning. Supervised by evidence prediction and relation extraction, this explicit multi-hop reasoning representation notably improves long-distance reasoning capabilities in DocRE. To address the issue of class imbalance in DocRE, we revise the Adaptive Focal Loss (AFL) and incorporate it into our framework. Moreover, we develop a Weighted Fusion Layer, designed to optimize the utilization of evidence, thereby ensuring a comprehensive yet focused analytical perspective. Our framework, named COREX, has shown superior performance in extensive experiments on three benchmark datasets. Its performance exceeds the baseline by 1.8/1.58 in Ign F1/Inter F1 score on the DocRED leaderboard. Silan Zhao, Chunping Li |
IJCNN | 2 |
| 2023 | PT3: A Transformer-based Model for Sepsis Death Risk Prediction via Vital Signs Time SeriesabstractSepsis is a life-threatening systemic syndrome with a high mortality rate. It is critical for doctors to identify sepsis patients at high risk of death in real time, saving patients and reducing in-hospital mortality. However, current clinical methods use traditional machine learning models to predict the death risk of sepsis patients based on the vital signs and lab test results, which is difficult to achieve real-time prediction. In this work, we propose a novel Transformer-based model named PT3 to predict the death risk of sepsis patients within$k$hours (k = 6, 24, 48) in the future based only on the vital signs time series that can be collected in real time. In clinical settings, the collection intervals of vital signs time series are usually irregular. To address this challenge, we design a time-aware mechanism by using a time decay function to explore temporal correlations between records at different moments. We further introduce an auto-imputing mechanism to our model by using the masked prediction pre-training task. To enhance the representation learning ability, we propose a similarity prediction task, a self-supervised pre-training method, to pre-train our model with triplet-loss function. We validate the effectiveness of PT3 on two public clinical databases, MIMIC-IV and eICU. Experiments results show that our model has effective prediction performance, whose AUC on MIMIC-IV and eICU datasets achieve 0.9067 and 0.8733 respectively, especially in the next 6 hours, outperforming other state-of-the-art deep learning methods. Ruihong Luo, Minghui Gong, Chunping Li |
IJCNN | 3 |
| 2023 | Emotion-regulatory chatbots for enhancing consumer servicing: An interpersonal emotion management approach
Bei Luo, Raymond Y. K. Lau, Chunping Li |
Inf. Manag. | 3 |
| 2022 | The Transition Law of Sepsis Patients' Illness States Based on Complex Network
Ruolin Wang, Jingming Liu, Minghui Gong, Chunping Li |
AIME | 5 |
| 2022 | Cognitive Load Measurement in the Impact of VR Intervention in LearningabstractThe rapid development of VR technology in training and learning is based on the assumption that it is beneficial for skill training within an immersive environment. However, extra cognitive load may be induced due to the additional sensory information and hence learning ability might be affected. In this study, we examined and compared the impact of cognitive load and task performance in real-world and VR environments through an empirical quadrant model. Forty-six participants completed the tasks with and without the secondary task in realworld and VR environments. The detection response task (DRT), as the secondary task, was adopted to estimate cognitive load based on response time and omission rate. No statistically significant differences were found in cognitive load and task performance in the comparison of VR and non-VR environment settings. There was an encouraging trend observed that VR environments have some advantages over the real-world, such as a higher level of immersion, which suggests that VR can benefit trainees with improved concentration levels and task performance. As evidenced by the variation in performance between females and males in our study, it appears that females tend to perform less well in VR environments, with a slightly higher cognitive load. Chunping Li, Soonja Yeom, Julian R. Dermoudy, Kristy de Salas |
ICALT | 1 |
| 2021 | Coarse- and Fine-grained Attention Network with Background-aware Loss for Crowd Density Map EstimationabstractIn this paper, we present a novel method Coarse- and Fine-grained Attention Network (CFANet) for generating high-quality crowd density maps and people count estimation by incorporating attention maps to better focus on the crowd area. We devise a from-coarse-to-fine progressive attention mechanism by integrating Crowd Region Recognizer (CRR) and Density Level Estimator (DLE) branch, which can suppress the influence of irrelevant background and assign attention weights according to the crowd density levels, because generating accurate fine-grained attention maps directly is normally difficult. We also employ a multi-level supervision mechanism to assist the backpropagation of gradient and reduce overfitting. Besides, we propose a Background-aware Structural Loss (BSL) to reduce the false recognition ratio while improving the structural similarity to groundtruth. Extensive experiments on commonly used datasets show that our method can not only outperform previous state-of-the-art methods in terms of count accuracy but also improve the image quality of density maps as well as reduce the false recognition ratio. Liangzi Rong, Chunping Li |
WACV | 2 |
| 2021 | A semi-supervised deep learning image caption model based on Pseudo Label and N-gram
Chunping Li, Youfang Han, Yan Zhu 0007 |
Int. J. Approx. Reason. | 2 |
| 2020 | Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling DecoderabstractTopic models have been prevailing for many years on discovering latent semantics while modeling long documents.However, for short texts they generally suffer from data sparsity because of extremely limited word cooccurrences; thus tend to yield repetitive or trivial topics with low quality.In this paper, to address this issue, we propose a novel neural topic model in the framework of autoencoding with a new topic distribution quantization approach generating peakier distributions that are more appropriate for modeling short texts.Besides the encoding, to tackle this issue in terms of decoding, we further propose a novel negative sampling decoder learning from negative samples to avoid yielding repetitive topics.We observe that our model can highly improve short text topic modeling performance.Through extensive experiments on real-world datasets, we demonstrate our model can outperform both strong traditional and neural baselines under extreme data sparsity scenes, producing high-quality topics. Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao |
EMNLP (1) | 2 |
| 2020 | Learning Multilingual Topics with Neural Variational Inference
Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao |
NLPCC (1) | 2 |
| 2019 | Short Text Topic Modeling with Flexible Word PatternsabstractSince effective semantic representations are utilized in many practical applications, inferring discriminative and coherent latent topics from short texts is a critical and basic task. Traditional topic models like Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) behave not well on short texts due to data sparsity problem. One novel model called Biterm Topic Model (BTM) which models unordered word-pairs (i.e., biterms) from whole corpus was proposed to solve this problem. However, both the performance and efficiency of BTM are reduced because of many irrelevant and useless biterms. In this paper, we propose a Multiterm Topic Model (MTM) for short text topic modeling. MTM extracts variable-length and more correlative word patterns (i.e., multiterms) from the whole corpus. By directly modeling the generative process of multiterms, MTM can infer the word distributions of each topic and the topic distribution of each short text to alleviate the sparsity problem in short text modeling. With the the proper amount of flexible multiterms, learning process of MTM is enhanced. Through extensive experiments on two real-world short text collections, we show that MTM is more efficient and outperforms the baseline models in terms of topic coherence and text classification. Xiaobao Wu, Chunping Li |
IJCNN | 2 |
| 2018 | An Adaptive Dirichlet Multinomial Mixture Model for Short Text Streaming ClusteringabstractIn this paper, we propose an adaptive Dirichlet Multinomial Mixture model for short text clustering along the time slices. A hyperparameters adjusting algorithm is utilized to capture the temporal dynamics automatically, and a collapsed Gibbs sampling algorithm for the extended Dirichlet Multinomial Mixture (DMM) model (e-GSDMM algorithm), is proposed to infer the changes of topic and word distributions along the time slices. Our extensive experiments over three different datasets show that the proposed model is efficient and performs better than the existing GSDMM approach for short text clustering on the streaming data. Ruting Duan, Chunping Li |
WI | 2 |
| 2017 | A Neural Autoregressive Framework for Collaborative Filtering
Zhen Ouyang, Chunping Li |
ISNN (2) | 3 |
| 2016 | Image retrieval by geological proximity using deep neural networkabstractAssociating image content with their geo-location has long been pursued. Many solutions focus on descriptor extraction and correspondence matching to measure image similarity by visual features, and infer the location from similar images. Other solutions utilize the correspondence of descriptors to reconstruct 3D model for image location calculation. However, the descriptor based solutions are usu-ally computationally heavy and the similarity measurement is unreliable for buildings of similar textures, which is quite common if the image comes from a wide area. In this paper, we leverage a modified deep neural network to estimate images geo-location similarity rather than visual appearance similarity, to deal with cases where images of different location but with resembling visual features. Our model builds upon the ability of neural network to extract high level features and the back-propagation can maximize the tiny difference of visually-similar images. We collected tens of thousands photos on campus, with many visually resembling, but different location ones. Experiments show that using geo-location similarity measurement can correctly deal with such cases. Daoyuan Jia, Yongchi Su, Chunping Li |
INISTA | 3 |
| 2015 | Turn Waste into Wealth: On Simultaneous Clustering and Cleaning over Dirty DataabstractDirty data commonly exist. Simply discarding a large number of inaccurate points (as noises) could greatly affect clustering results. We argue that dirty data can be repaired and utilized as strong supports in clustering. To this end, we study a novel problem of clustering and repairing over dirty data at the same time. Referring to the minimum change principle in data repairing, the objective is to find a minimum modification of inaccurate points such that the large amount of dirty data can enhance the clustering. We show that the problem can be formulated as an integer linear programming (ILP) problem. Efficient approximation is then devised by a linear programming (LP) relaxation. In particular, we illustrate that an optimal solution of the LP problem can be directly obtained without calling a solver. A quadratic time approximation algorithm is developed based on the aforesaid LP solution. We further advance the algorithm to linear time cost, where a trade-off between effectiveness and efficiency is enabled. Empirical results demonstrate that both the clustering and cleaning accuracies can be improved by our approach of repairing and utilizing the dirty data in clustering. Shaoxu Song, Chunping Li, Xiaoquan Zhang |
KDD | 2 |
| 2014 | Object typicality for effective Web of Things recommendations
Yi Cai 0001, Raymond Y. K. Lau, Stephen Shaoyi Liao, Chunping Li, Ho-fung Leung, Louis C. K. Ma |
Decis. Support Syst. | 4 |
| 2014 | Social analytics: Learning fuzzy product ontologies for aspect-oriented sentiment analysis
Raymond Y. K. Lau, Chunping Li, Stephen Shaoyi Liao |
Decis. Support Syst. | 2 |
| 2013 | Social recommendation incorporating topic mining and social trust analysisabstractWe study the problem of social recommendation incorporating topic mining and social trust analysis. Different from other works related to social recommendation, we merge topic mining and social trust analysis techniques into recommender systems for finding topics from the tags of the items and estimating the topic-specific social trust. We propose a probabilistic matrix factorization (TTMF) algorithm and try to enhance the recommendation accuracy by utilizing the estimated topic-specific social trust relations. Moreover, TTMF is also convenient to solve the item cold start problem by inferring the feature (topic) of new items from their tags. Experiments are conducted on three different data sets. The results validate the effectiveness of our method for improving recommendation performance and its applicability to solve the cold start problem. Chunping Li, Li Li 0102 |
CIKM | 2 |
| 2013 | Combining Lexical and Semantic Features for Short Text ClassificationabstractIn this paper, we propose a novel approach to classify short texts by combining both their lexical and semantic features. We present an improved measurement method for lexical feature selection and furthermore obtain the semantic features with the background knowledge repository which covers target category domains. The combination of lexical and semantic features is achieved by mapping words to topics with different weights. In this way, the dimensionality of feature space is reduced to the number of topics. We here use Wikipedia as background knowledge and employ Support Vector Machine (SVM) as classifier. The experiment results show that our approach has better effectiveness compared with existing methods for classifying short texts. Chunping Li, Li Li 0102 |
KES | 2 |
| 2013 | Topic-Level Expert Modeling in Community Question AnsweringabstractCommunity Question Answering (CQA) services provide an open platform for people to share their knowledge and have attracted great attention for its rapidly increasing popularity. As more knowledge is shared in CQA, how to use the repository for solving new questions has become a crucial problem. In this paper, we tackle the problem by finding experts from the question answering history first and then recommending the appropriate experts to answer the new questions. We develop the Topic-level Expert Learning (TEL) model to find experts on topic level in CQA. Our proposed model incorporates link analysis into content analysis. The main difference between TEL and other generative models is that TEL can automatically adjust and update the sampling parameters during iterations in order to better model the experts on topic level. The experiments are conducted on datasets crawled from Yahoo! Answers and the results show that our method can effectively find experts to answer new questions and can better predict best responders for new questions. Our model achieves significant improvement over the baseline methods on multiple metrics. Beyond these metric performances, TEL converges fast within a few iterations. Naiwen Bian, Chunping Li |
SDM | 3 |
| 2013 | Discovering Correlated Entities from News Archives
Chunping Li, Li Li 0102 |
WISE (2) | 2 |
| 2013 | A random-walk based recommendation algorithm considering item categories
Liyan Zhang 0001, Chunping Li |
Neurocomputing | 3 |
| 2013 | Generating virtual ratings from chinese reviews to augment online recommendationsabstractCollaborative filtering (CF) recommenders based on User-Item rating matrix as explicitly obtained from end users have recently appeared promising in recommender systems. However, User-Item rating matrix is not always available or very sparse in some web applications, which has critical impact to the application of CF recommenders. In this article we aim to enhance the online recommender system by fusing virtual ratings as derived from user reviews. Specifically, taking into account of Chinese reviews' characteristics, we propose to fuse the self-supervised emotion-integrated sentiment classification results into CF recommenders, by which the User-Item Rating Matrix can be inferred by decomposing item reviews that users gave to the items. The main advantage of this approach is that it can extend CF recommenders to some web applications without user rating information. In the experiments, we have first identified the self-supervised sentiment classification's higher precision and recall by comparing it with traditional classification methods. Furthermore, the classification results, as behaving as virtual ratings, were incorporated into both user-based and item-based CF algorithms. We have also conducted an experiment to evaluate the proximity between the virtual and real ratings and clarified the effectiveness of the virtual ratings. The experimental results demonstrated the significant impact of virtual ratings on increasing system's recommendation accuracy in different data conditions (i.e., conditions with real ratings and without). Weishi Zhang, Guiguang Ding, Li Chen 0009, Chunping Li |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2012 | A Customized Dependency Tree Kernel for Effective Sentiment ClassificationabstractThis paper introduces a kernel over dependency trees for sentiment classification. In order to classify a text as positive or negative, syntactic and word dependency information should be exploited besides words. Dependency parse trees, generated by automatic sentence parser, contain much syntactic information which would be helpful for sentiment classification. On the other hand, dependency trees contain word dependency information of sentences, and give a deeper understanding of natural language than BOW (bag-of-word) and n-gram schemas. In this paper, we present an approach which exploits such syntactic and word dependency information for sentiment classification. Our approach achieves good performance. We also compared the dependency tree kernel we proposed with some other tree kernels. Zhou Sun, Chunping Li |
KES | 3 |
| 2012 | Latent Business Networks Mining: A Probabilistic Generative ModelabstractThough numerous research has been devoted to social network discovery and analysis, relatively little research has been conducted on business network discovery. The main contribution of our research is the development of a novel probabilistic generative model for latent business networks mining. Our experimental results confirm that the proposed method outperforms the well-known vector space based model by 24% in terms of AUC value. Wenping Zhang, Raymond Y. K. Lau, Yunqing Xia, Chunping Li, Wenjie Li 0002 |
Web Intelligence | 4 |
| 2012 | Predicting Best Responder in Community Question Answering Using Topic Model MethodabstractCommunity question answering (CQA) services provide an open platform for people to share their knowledge and have attracted great attention for its rapidly increasing popularity. As the more knowledge people provided are shared in CQA, how to use the historical knowledge for solving new questions has become a crucial problem. In this paper, we investigate the problem as predicting best responders for new questions and tackle the problem from two perspectives, one is from the asker of the new question, and the other is from the question itself. We propose two supervised topic models, Asker-Responder Topic Model (ARTM) and Question-Responder Topic Model (QRTM) for both two perspectives by tracking people's answering history as background knowledge. Our experiments show that the two supervised topic models can effectively predict best responders for new questions in CQA without any additional works and have significant improvement over the baseline method. Chunping Li, Li Li 0102 |
Web Intelligence | 2 |
| 2012 | Workflow simulation for operational decision support using event graph through process mining
Ying Liu 0004, Chunping Li, Roger Jianxin Jiao |
Decis. Support Syst. | 3 |
| 2010 | Identifying new categories in community question answering archives: a topic modeling approachabstractCommunity Question Answering (CQA) services have evolved into a popular way of information seeking and providing. User-posted questions in CQA are generally organized into hierarchical categories. In this paper, we define and study a novel problem which is referred to as New Category Identification (NCI) in CQA question archives. New Category Identification is primarily concerned with detecting and characterizing new or emerging categories which are not included in the existing category hierarchy. We define this problem formally, and propose both unsupervised and semi-supervised topic modeling methods to solve it. Experiments with a ground-truth set built from Yahoo! Answers show that our methods identify and interpret new categories effectively. Yajie Miao, Chunping Li, Jie Tang 0001 |
CIKM | 2 |
| 2010 | A Domain-related Authority Model for Web Pages based on Source and Related Information
Chunping Li, Ming Gu 0001 |
ICSOFT (1) | 2 |
| 2010 | Improving Question Answering Based on Query Expansion with WikipediaabstractAs an emerging area in information retrieval, question answering aims at retrieving answers to user-posted questions from a given sentence collection or text corpus. In question answering, the queries are usually submitted in the form of short sentences which are unable to represent user intentions sufficiently. In this study, we present a novel framework which improves question answering through query expansion. We enrich representation of queries with Wikipedia concepts generated by the proposed QRWiki retrieval model. Then the enriched queries are exploited to benefit the process of question answering. The experiments with benchmark datasets show that the proposed framework performs significantly better than the baseline system, and is effective in boosting the performance of question answering. Yajie Miao, Chunping Li |
ICTAI (2) | 3 |
| 2010 | A Novel Approach of Process Mining with Event Graph
Ying Liu 0004, Chunping Li, Roger Jianxin Jiao |
KES (1) | 3 |
| 2010 | MindDigger: Feature Identification and Opinion Association for Chinese Movie Reviews
Chunping Li |
KSEM | 2 |
| 2010 | Mining Wikipedia and Yahoo! Answers for Question Expansion in Opinion QA
Yajie Miao, Chunping Li |
PAKDD (1) | 2 |
| 2010 | Evaluating Importance of Websites on News Topics
Yajie Miao, Chunping Li, Ming Gu 0001 |
PRICAI | 2 |
| 2010 | Automatically Grouping Questions in Yahoo! AnswersabstractIn this paper, we define and study a novel problem which is referred to as Community Question Grouping (CQG). Online QA services such as Yahoo! Answers contain large archives of community questions which are posted by users. Community Question Grouping is primarily concerned with grouping a collection of community questions into predefined categories. We first investigate the effectiveness of two basic methods, i.e., K-means and PLSA, in solving this problem. Then, both methods are extended in different ways to include user information. The experimental results with real datasets show that incorporation of user information improves the basic methods significantly. In addition, performance comparison reveals that PLSA with regularization is the most effective solution to the CQG problem. Yajie Miao, Chunping Li, Jie Tang 0001 |
Web Intelligence | 3 |
| 2009 | Ontology Based Opinion Mining for Movie Reviews
Chunping Li |
KSEM | 2 |
| 2009 | Zero-Sum Reward and Punishment Collaborative Filtering Recommendation AlgorithmabstractIn this paper, we propose a novel memory-based collaborative filtering recommendation algorithm. Our algorithm use a new metric named influence weight, which is adjusted with zero-sum reward and punishment mechanism whenever the active user provides a new rating, to select neighbors and weight their opinions. Since the weight of personalized ratings, which contain more value for searching similar neighbors, is magnified appropriately in the formation of influence weight, our algorithm can find similar neighbors more effectively and filter the fake users introduced by shilling attacks automatically. When predicting for the active user, our algorithm select neighbors with the Top-N largest positive influence weights and predict their missing ratings. This rating smoothing method can alleviate data sparsity more efficiently. Then it computes the weighted average of all the selected neighbors' opinions and generates recommendations. Empirical results confirm that our algorithm achieves significant progress in all aspects of accuracy, scalability, robustness against data sparsity and shilling attacks simultaneously. Chunping Li |
Web Intelligence | 2 |
| 2008 | Combining Context Features by Canonical Belief Network for Chinese Part-Of-Speech Tagging
Hongzhi Xu, Chunping Li |
IJCNLP | 2 |
| 2008 | A topical PageRank based algorithm for recommender systemsabstractIn this paper, we propose a Topical PageRank based algorithm for recommender systems, which aim to rank products by analyzing previous user-item relationships, and recommend top-rank items to potentially interested users. We evaluate our algorithm on MovieLens dataset and empirical experiments demonstrate that it outperforms other state-of-the-art recommending algorithms. Chunping Li |
SIGIR | 3 |
| 2007 | A Novel Term Weighting Scheme for Automated Text CategorizationabstractTerm weighting is an important task for text classification. Inverse document frequency (IDF) is one of the most popular methods for this task; however, in some situations, such as supervised learning for text categorization, it doesn 't weight terms properly, because it neglects the category information and assumes that a term that occurs in smaller set of documents should get a higher weight. There have been several term weighting schemes that consider the category information. In this paper, we present a new term weighting scheme that considers more information provided by the term distribution among different categories. The experiments show that our method is more effective than three other popular schemes. Hongzhi Xu, Chunping Li |
ISDA | 2 |
| 2006 | Semi-supervised Classification Based on Smooth Graphs
Xueyuan Zhou, Chunping Li |
DASFAA | 2 |
| 2006 | Combining Smooth Graphs with Semi-supervised Classification
Xueyuan Zhou, Chunping Li |
PAKDD | 2 |
| 2006 | Improved ROCK for Text Clustering Using Asymmetric Proximity
Shaoxu Song, Chunping Li |
SOFSEM | 2 |