Qiudan Li

dblp:62/546 · DBLP profile ↗
← Back
54ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-8714-4562ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 1 since 2021Security and privacy · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Theory of computation · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Reinforcement Learning-Guided Adaptive Tuning for Out-of-Distribution Harmful Text Detection
abstract
As social media grows, harmful information spreads rapidly across platforms and evolves over time, showing cross-platform and crosstemporal variations.Existing methods rely on fixed model parameters during training, which fail to handle substantial semantic discrepancies, leading to Out-Of-Distribution (OOD) problems.While test-time tuning enables dynamic parameter adjustment, it may lead to excessive adaptation to individual samples.The key challenge is how to adapt to semantic variations during testing while preventing overfitting from continuous tuning.To tackle this issue, this paper proposes RLAT, a reinforcement learning (RL)-guided adaptive tuning method for harmful text detection.First, a tuning joint optimization module is designed to update parameters and adapt to semantic variations during testing.It tunes the model by optimizing consistency loss and applying word-level attention constraints to reduce over-reliance on local words and learn a more robust global representation.Then, to mitigate overfitting caused by continuous tuning, a RL-guided adaptive decision model is introduced to direct the tuning process.It reduces the influence of local samples by selecting data and controlling parameter updates, thereby improving overall test performance.Experimental results show that the RLAT outperforms state-of-the-art baselines in cross-platform and cross-temporal scenarios across multiple public datasets.
Mengyu Xiang, Tinghao Chen, Boxu Han, Qiudan Li, Daniel Dajun Zeng
ACL (1)4
2025 Knowledge-Enhanced Hierarchical Heterogeneous Graph for Personality Identification with Limited Training Data
abstract
Personality identification plays important roles in understanding user behavior and offering foresight ability for downstream applications. The key challenge is how to address the scarcity of labeled personality data. Recently, some studies have adopted data augmentation and prompt learning to perform personality identification. However, they still heavily require a large amount of labeled data to learn an appropriate distance strategy, which limits the generalization and flexibility of the model. This study proposes a knowledge-enhanced hierarchical heterogeneous graph model, which adopts a global multi-view graph node encoding to acquire comprehensive personality features and their inherent associations, where three types of knowledge including part-of-speech (POS) tag, entity, and Linguistic Inquiry and Word Count (LIWC) are introduced. Then, a hierarchical heterogeneous graph with a “post-word-diverse knowledge” structure is constructed for each post to obtain enhanced representation. Finally, a relation guided representation optimization that considers intra-user relationships and inter-label relationships is further developed to learn more discriminative semantic representation. Experimental results on three widely used datasets demonstrate that the model outperforms state-of-the-art methods when training with only 100 samples (approximately 1% of the total data set).
Qiudan Li, Yilin Wu 0005, David Jingjun Xu, Daniel Dajun Zeng
AAAI2
2025 A Fusion Pretrained Approach for Identifying the Cause of Sarcasm Remarks
abstract
Sarcastic remarks often appear in social media and e-commerce platforms to express almost exclusively negative emotions and opinions on certain instances, such as dissatisfaction with a purchased product or service. Thus, the detection of sarcasm allows merchants to timely resolve users’ complaints. However, detecting sarcastic remarks is difficult because of its common form of using counterfactual statements. The few studies that are dedicated to detecting sarcasm largely ignore what sparks these sarcastic remarks, which could be because of an empty promise of a merchant’s product description. This study formulates a novel problem of sarcasm cause detection that leverages domain information, dialogue context information, and sarcasm sentences by proposing a pretrained language model-based approach equipped with a novel hybrid multihead fusion-attention mechanism that combines self-attention, target-attention, and a feed-forward neural network. The domain information and the dialogue context information are then interactively fused to obtain the domain-specific dialogue context representation, and bidirectionally enhanced sarcasm-cause pair representations are generated for detecting sarcasm spark. Experimental results on real-world data sets demonstrate the efficacy of the proposed model. The findings of this study contribute to the literature on sarcasm cause detection and provide business value to relevant stakeholders and consumers. History: Accepted by Ram Ramesh, Area Editor for Data Science and Machine Learning. Funding: This work was partially supported by the National Natural Science Foundation of China [Grants 72293575, 62071467, and 62141608] and the Research Grant Council of the Hong Kong Special Administrative Region, China [Grants 11500322 and 11500421]. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2022.0285 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2022.0285 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
Qiudan Li, David Jingjun Xu, Haoda Qian, Linzi Wang, Minjie Yuan, Daniel Dajun Zeng
INFORMS J. Comput.1
2024 Graph Representation Learning Based on Cognitive Spreading Activations
abstract
Graph representation learning is an emerging area for graph analysis and inference. However, existing approaches for large-scale graphs either sample nodes in sequential walks or manipulate the adjacency matrices of graphs. The former approach can cause sampling bias against less-connected nodes, whereas the latter may suffer from sparsity that exists in many real-world graphs. To learn from structural information in a graph more efficiently and comprehensively, this paper proposes a new graph representation learning approach inspired by the cognitive model of spreading-activation mechanisms in human memory. This approach learns node embeddings by adopting a graph activation model that allows nodes to “activate” their neighbors and spread their own structural information to other nodes through the paths simultaneously. Comprehensive experiments demonstrate that the proposed model performs better than existing methods on several empirical datasets for multiple graph inference tasks. Meanwhile, the spreading-activation-based model is computationally more efficient than existing approaches–the training process converges after only a small number of iterations, and the training time is linear in the number of edges in a graph. The proposed method works for both homogeneous and heterogeneous graphs.
Kang Zhao 0001, Linjing Li, Daniel Dajun Zeng, Qiudan Li, Quannan Zu
IEEE Trans. Knowl. Data Eng.5
2023 A Mutually Enhanced Bidirectional Approach for Jointly Mining User Demand and Sentiment (Student Abstract)
abstract
User demand mining aims to identify the implicit demand from the e-commerce reviews, which are always irregular, vague and diverse. Existing sentiment analysis research mainly focuses on aspect-opinion-sentiment triplet extraction, while the deeper user demands remain unexplored. In this paper, we formulate a novel research question of jointly mining aspect-opinion-sentiment-demand, and propose a Mutually Enhanced Bidirectional Extraction (MEMB) framework for capturing the dynamic interaction among different types of information. Finally, experiments on Chinese e-commerce data demonstrate the efficacy of the proposed model.
Haoda Qian, Minjie Yuan, Qiudan Li
AAAI4
2023 Style-Driven Multi-Perspective Relevance Mining Model for Hotspot Reprint Paragraph Prediction
abstract
Accurately predicting hotspot reprint paragraphs can timely provide valuable clues for topic selection, thereby improving the influence of the disseminated content. Most existing works in media reprint analysis focus on mining reprint relationships and reprint patterns. Meanwhile, few works predict the hotspot reprint paragraph from a fine-grained level. The writing style reflects the structure and semantic logic of the article to some extent. Thus, the challenge is to determine how to effectively incorporate writing style features into the semantic analysis while also reasoning deeply about the semantic relevance between sections of the article. This paper proposes a multi-perspective relevance collaborative modeling method called MPRCM-TS. It integrates writing styles of titles into the semantic representations and deeply mines the multi-perspective semantic relevance between the title and paragraphs on the basis of the attention mechanism. Simultaneously, multiple loss functions collaborate to enhance the parameter optimization ability. We evaluate the performance of the proposed model on a real-world dataset, and the experimental results demonstrate the efficacy.
Linzi Wang, Haoda Qian, Qiudan Li, David Jingjun Xu, Daniel Dajun Zeng
ISI3
2023 A Two-Stage Prompt Learning Method for Jointly Predicting Topic and Personality
abstract
Accurate personality prediction can help management departments analyze users' behaviors and make informed decisions effectively. Existing text-based personality prediction studies mainly rely on deep neural networks or pre-trained language models to extract semantic information and personality traits. However, the text's topic and label description may provide additional personality clues. This paper proposes a topic and personality prediction method based on a large language model (LLM), which utilizes a two-stage prompt strategy to mine the interaction between the topic and personality information. Additionally, labels' descriptions are incorporated to construct cue-based prompts, and a fine-tuning approach is adopted to optimize the model's performance. Experiments on two datasets show the efficacy of the proposed model.
Yilin Wu 0005, Minjie Yuan, Chenyu Yuan, Qiudan Li
ISI6
2022 A Transformer-based Approach for Identifying Target-oriented Opinions from Travel Reviews
abstract
Performing target-oriented opinion word extraction (TOWE) from online travel reviews is a valuable reference for both tourists and attraction administration department. This paper formulates a novel research topic of identifying target-opinion pair from Chinese travel review corpus. Learning target-oriented representation accurately, locating the opinion word and extracting the complete opinion are three major challenges. Hence, we leverage aspect-based query, pos-tag and relative position and devise appropriate structure to fuse them in an encoder-decoder framework. Specifically, in the encoder, the target-fused (aspect, review) pair and the pos-tag label are encoded by transformers to model the global dependency, in the decoder, a BiLSTM is adopted to enhance contextual representation by incorporating relative position information. A real-world Chinese travel dataset for TOWE task is constructed, and the experimental results demonstrate the efficacy of the proposed model. Extensive ablation experiments are also conducted to study the effect of different components of the model.
Haoda Qian, Zaichuan Tang, Yajun Ren, Qiudan Li, Daniel Dajun Zeng
IJCNN4
2022 A BERT-based Heterogeneous Graph Convolution Approach for Mining Organization-Related Topics
abstract
Mining organization-related topics is helpful to analyze the information dissemination situation. Existing methods based on graph neural networks mainly consider the association between words and documents, they ignore the semantic interactions between documents, and do not consider the heterogeneity of edges which are difficult to solve the challenge of blurred topic boundaries in real scenarios, resulting in performance loss. This paper proposes a BERT-based Heterogeneous Graph Convolution Network (BERT-HGCN) approach for semi-supervised topic mining that comprehensively considers multi-semantic relations between words and documents. It deeply combines the advantages of transductive learning with pre-training models. We model documents as graph-structured data and capture multiple semantic dependencies among word-word, word-doc, and doc-doc via information propagation mechanism. During the model learning process, a two-stream encoding mechanism is used to learn the structural and semantic representations, which combines a hierarchical graph convolution network (HGCN) and a BERT-based auto-encoder. It considers both edges heterogeneity and semantics of original documents. Finally, a dual-supervision loss is used to train the classifier based on graph nodes and semantic representations for topic mining. We empirically evaluate the performance of the proposed model on a real-world organization-related dataset, and the experimental results demonstrate the efficacy of the model.
Haoda Qian, Minjie Yuan, Qiudan Li, Daniel Dajun Zeng
IJCNN3
2022 Detecting Product Adoption Intentions via Multiview Deep Learning
abstract
Detecting product adoption intentions on social media could yield significant value in a wide range of applications, such as personalized recommendations and targeted marketing. In the literature, no study has explored the detection of product adoption intentions on social media, and only a few relevant studies have focused on purchase intention detection for products in one or several categories. Focusing on a product category rather than a specific product is too coarse-grained for precise advertising. Additionally, existing studies primarily focus on using one type of text representation in target social media posts, ignoring the major yet unexplored potential of fusing different text representations. In this paper, we first formulate the problem of product adoption intention mining and demonstrate the necessity of studying this problem and its practical value. To detect a product adoption intention for an individual product, we propose a novel and general multiview deep learning model that simultaneously taps into the capability of multiview learning in leveraging different representations and deep learning in learning latent data representations using a flexible nonlinear transformation. Specifically, the proposed model leverages three different text representations from a multiview perspective and takes advantage of local and long-term word relations by integrating convolutional neural network (CNN) and long short-term memory (LSTM) modules. Extensive experiments on three Twitter datasets demonstrate the effectiveness of the proposed multiview deep learning model compared with the existing benchmark methods. This study also significantly contributes research insights to the literature about intention mining and provides business value to relevant stakeholders such as product providers.
Zhu (Drew) Zhang, Xuan Wei 0001, Xiaolong Zheng 0001, Qiudan Li, Daniel Dajun Zeng
INFORMS J. Comput.4
2021 An Attention Based Multi-view Model for Sarcasm Cause Detection (Student Abstract)
abstract
Sarcasm often relates to people’s implicit discontent with certain products and policies. Existing research mainly focus on sarcasm detection, while the deep causal relationships in the full conversation remained unexplored. This paper formulates a novel research question of sarcasm cause detection, and proposes an attention based model that simultaneously captures different semantic associations as well as the inner causal logics in multi-view manner. Experiments on public Reddit dataset prove the efficacy of the proposed model.
Hejing Liu, Qiudan Li, Zaichuan Tang
AAAI2
2021 A Multi-Task MRC Framework for Chinese Emotion Cause and Experiencer Extraction
Haoda Qian, Qiudan Li, Zaichuan Tang
ICANN (4)2
2021 A Multi-level Semantic Fusion Approach for News Reprint Pattern Detection
abstract
News reprint analysis is gradually becoming a hot research topic. Most existing studies in news reprint analysis mainly focus on mining news reprint relations, there is little study exploring the detection of the patterns of news reprint yet. To fill this gap, we aim to identify news reprint patterns such as word variant, sentence conversion, and topic restatement, which can provide deep insights into the news propagation mechanism. The challenge of this question lies in how to dig domain information and obtain the deep semantic representation of news to discover the writing styles of various reprint patterns. This paper proposes a media domain information-driven multi-level semantic fusion (MDID-MLSF) approach. It simultaneously considers news media information and extracts word-sentence-paragraph-level news semantic information by the interactive matching mechanism. Specifically, Word Mover's Distance (WMD) algorithm and Multilayer Perceptron (MLP) are employed to obtain semantic information at the word level and sentence level. Then, the model utilizes attention-based hierarchical Bi-LSTM to obtain the paragraph level semantic information. Finally, news media information and different levels' semantic information are jointly modeled to detect the patterns of news reprint. We empirically evaluate the performance of the proposed model on a real-world dataset, the experimental results demonstrate the efficacy of the model.
Linzi Wang, Riheng Yao, Qiudan Li, Xina Zhang, Hejing Liu
IJCNN3
2020 Session-Level User Satisfaction Prediction for Customer Service Chatbot in E-Commerce (Student Abstract)
abstract
This paper aims to predict user satisfaction for customer service chatbot in session level, which is of great practical significance yet rather untouched. It requires to explore the relationship between questions and answers across different rounds of interactions, and handle user bias. We propose an approach to model multi-round conversations within one session and take user information into account. Experimental results on a dataset from a real-world industrial customer service chatbot Alime demonstrate the good performance of our proposed model.
Riheng Yao, Shuangyong Song, Qiudan Li, Chao Wang 0057, Haiqing Chen, Daniel Dajun Zeng
AAAI3
2020 Understanding and Predicting Users' Rating Behavior: A Cognitive Perspective
abstract
Online reviews are playing an increasingly important role in understanding and predicting users’ rating behavior, which brings great opportunities for users and organizations to make better decisio...
Qiudan Li, Daniel Dajun Zeng, David Jingjun Xu, Ruoran Liu, Riheng Yao
INFORMS J. Comput.1
2019 Multimodal Data Enhanced Representation Learning for Knowledge Graphs
abstract
Knowledge graph, or knowledge base, plays an important role in a variety of applications in the field of artificial intelligence. In both research and application of knowledge graph, knowledge representation learning is one of the fundamental tasks. Existing representation learning approaches are mainly based on structural knowledge between entities and relations, while knowledge among entities per se is largely ignored. Though a few approaches integrated entity knowledge while learning representations, these methods lack the flexibility to apply to multimodalities. To tackle this problem, in this paper, we propose a new representation learning method, TransAE, by combining multimodal autoencoder with TransE model, where TransE is a simple and effective representation learning method for knowledge graphs. In TransAE, the hidden layer of autoencoder is used as the representation of entities in the TransE model, thus it encodes not only the structural knowledge, but also the multimodal knowledge, such as visual and textural knowledge, into the final representation. Compared with traditional methods based on only structural knowledge, TransAE can significantly improve the performance in the sense of link prediction and triplet classification. Also, TransAE has the ability to learn representations for entities out of knowledge base in zero-shot. Experiments on various tasks demonstrate the effectiveness of our proposed TransAE method.
Zikang Wang, Linjing Li, Qiudan Li, Daniel Dajun Zeng
IJCNN3
2019 A Novel Neural Approach for News Reprint Prediction
abstract
News media has become a prevalent information spreading platform, where news sites can reprint news from other sites. To better understand the mechanism of news propagation, it is necessary to model reprint behavior and predict whether a news site will reprint a piece of news. Most existing works in news reprint analysis focus on analyzing the semantic of news content, little work has been done on integrating reprint relationship among sites and news content for reprint prediction from the perspective of sites. The challenge of improving prediction performance lies in how to effectively incorporate these two kinds of information to learn a more comprehensive reprint behavior model. In this paper, we propose an Integrated Neural Reprint Prediction (INRP) model that considers both reprint relationship and news content. It models the reprint relationships as a directed weighted graph and maps it into a latent space to learn sites representations. During news content modeling process, sites representations are embedded as attention guidance to build up more site-specific content representations. Finally, sites and news representations are jointly modeled to predict whether a piece of news will be reprinted by a site. We empirically evaluate the performance of the proposed model on a real world dataset. Experimental results show that taking both the reprint relationship and news content information into consideration could allow us make more accurate analysis of reprint patterns. The mined patterns could serve as a feedback channel for both corporations and management departments.
Riheng Yao, Qiudan Li, Lei Wang 0062, Daniel Dajun Zeng
IJCNN2
2019 Analyzing Topics of JUUL Discussions on Social Media Using a Semantics-assisted NMF model
abstract
JUUL has become a widely used brand of e-cigarettes which takes more than 70% of the market. Social media provides a popular platform for users to discuss the preference and perceptions of JUUL. The discussions are valuable for real-time monitoring of JUUL use. Current research on topic analysis of JUUL discussions mainly relies on human work, which takes much time and effort. This paper adopts a Semantics-assisted NMF topic analysis model to automatically discover topics from JUUL-related short posts on Reddit. By successfully merging the semantic relationships into traditional NMF, this model outperforms in discovering topics with keywords that are important but have a lower word frequency among the posts. Experimental results show the potential of this model in JUUL surveillance and control practice.
Hejing Liu, Qiudan Li, Riheng Yao, Daniel Dajun Zeng
ISI2
2019 Inferring Users' Usage Patterns for Drug Abuse Surveillance
abstract
Inferring drug usage patterns includes age of drug abuse and intention of rehabilitation, which is of much importance for drug abuse surveillance. The challenges are how to mine patterns from posts and interaction relationships between users. In this paper, we propose a novel drug usage pattern inference method, which improves the inference accuracy by integrating the semantic features and interaction relationships effectively. Experimental results on a real-world dataset demonstrate the efficacy of the proposed method.
Ruoran Liu, Qiudan Li, Daniel Dajun Zeng
ISI2
2019 A Prior Knowledge Based Neural Attention Model for Opioid Topic Identification
abstract
The opioid epidemic has become a serious public health crisis in the United States. Social media sources such as Reddit containing user-generated content may be a valuable safety surveillance platform to evaluate discussions discerning opioid use. This paper proposes a prior knowledge based neural attention model for opioid topics identification, which considers prior knowledge with attention mechanism. Experimental results on a real-world dataset show that our model can extract coherent topics, the identified less discussed but important topics provide more comprehensive information for opioid safety surveillance.
Riheng Yao, Qiudan Li, Wei-Hsuan Lo-Ciganic, Daniel Dajun Zeng
ISI2
2018 A Novel Embedding Method for News Diffusion Prediction
abstract
News diffusion prediction aims to predict a sequence of news sites which will quote a particular piece of news. Most of previous propagation models make efforts to estimate propagation probabilities along observed links and ignore the characteristics of news diffusion processes, and they fail to capture the implicit relationships between news sites. In this paper, we propose an algorithm to model the news diffusion processes in a continuous space and take the attributes of news into account. Experiments performed on a real-world news dataset show that our model can take advantage of news’ attributes and predict news diffusion accurately.
Ruoran Liu, Qiudan Li, Lei Wang 0062, Daniel Dajun Zeng
AAAI2
2018 Catching Dynamic Heterogeneous User Data for Identity Linkage Learning
abstract
Benefitting from the development of social platforms, more and more users tend to register multiple accounts on different social networks. Linking user identities across multiple online social networks based on user behavior patterns is considerable for network supervision and information tracking. However, a user's online behavior in a social network is dynamic. The user profile may be changed due to some specific reasons such as user migration or job changes. Thus, catching the dynamics of evolutionary user data and collecting the latest user features are important and challenging issues in the area of user identity linkage. Inspired by deep learning models such as word2vec and Deep Walk, this paper proposes an integrated framework to catch the dynamic user data by supplementing vacant features and updating outdated features in data sources. The framework firstly represents all textual and structural user data into Iow- dimensional latent spaces by utilizing word2vec and DeepWalk, then, integrates different user features and predicts vacant data fields based on late fusion approach and cosine similarity computation. We then explore and evaluate the application of our proposed method in a user identity mapping task. The results proved that our framework can successfully catch the dynamic user data and enhance the performance of identity linkage models by supplementing and updating data sources advance with the times.
Qiudan Li, Lei Wang 0062, Daniel Dajun Zeng
IJCNN2
2017 Incorporating message embedding into co-factor matrix factorization for retweeting prediction
abstract
With the rapid growth of Web 2.0, social media has become a prevalent information sharing and spreading platform, where users can retweet interesting messages. To better understand the propagation mechanism for information diffusion, it is necessary to model the user retweeting behavior and predict future retweets. Some existing work in retweeting prediction based on matrix factorization focuses on using user-message interaction information, user information and social influence information, etc. The challenge of improving prediction performance is how to jointly perform deep representation of these information to solve the sparsity problem and then learn a more comprehensive retweeting behavior model. Inspired by word2vec and co-factor matrix factorization model, this paper proposes a hybrid model, called HCFMF, for learning users' retweeting behavior, it first computes the message content similarity by considering the message co-occurrence, the author information and word2vec based low-dimensional representation of content, then, jointly decomposes the user-message matrix and message-message similarity matrix based on a co-factorization model. We empirically evaluate the performance of the proposed model on real world weibo datasets. Experimental results show that taking the dense representation of author and content information into consideration could allow us make more accurate analysis of users' retweeting patterns. The mined patterns could serve as a feedback channel for both consumers and management departments.
Qiudan Li, Lei Wang 0062, Daniel Dajun Zeng
IJCNN2
2017 Mapping users across social media platforms by integrating text and structure information
abstract
With the development of social media technology, users often register accounts, post messages and create friend links on several different platforms. Performing user identity mapping on multi-platform based on the behavior patterns of users is considerable for network supervision and personalization service. The existing methods focus on utilizing either text information or structure information alone. However, text information and structure information reflect different aspects of a user. An organic combination of them is beneficial to mining user behavior patterns, thus help identify users across platforms accurately. The challenging problems are the effective representation and similarity computation of the text and structure information. We propose a mapping method which integrates text and structure information. At first, the model represents user name, description, location information based on word2vec or string matching, and friend information represented as relation network is regarded as structure information. Then these information are used for similarity computation using Jaccard index or cosine similarity. After similarity computation, a linear model is adopted to get the overall similarity of user pairs to perform user mapping. Based on the proposed method, we develop a prototype system, which allows users to set and adjust the weights of different information, or set expected index. The experimental results on a real-world dataset demonstrate the efficiency of the proposed model.
Qiudan Li, Daniel Dajun Zeng
ISI2
2017 Mining phase evolution for hot topics: A case study from multiple social media platforms
abstract
Monitoring the evolution phases of real-time event including occurrence, development, climax, decline and ending is crucial for management department to intuitively and comprehensively understand the event and then make better decisions. However, there have been very few studies on performing phase evolution analysis of event using the number of posts at the specific time unit. The challenge of this problem is how to identify temporal pattern and mine topic of different phases automatically. In this paper, we propose a unified phase evolution mining model, it firstly identifies the temporal patterns of phases based on k-means and empirical rules, then, burst detection algorithm is adopted to discover peak interval of all phases, finally, we use a summarization technique TextRank to extract keywords from contents to summarize the topics in each phase. In addition, we perform experiments on two real-world datasets collected from different social media platform to understand the event evolution in a more comprehensive way. Experimental results show the characteristics of event evolution on different social media platforms and verify the efficacy of the proposed model.
Ruoran Liu, Qiudan Li, Lei Wang 0062, Daniel Dajun Zeng, Hongyuan Ma
SMC2
2017 Associated Activation-Driven Enrichment: Understanding Implicit Information from a Cognitive Perspective
abstract
In this paper, we propose a novel text representation paradigm and a set of follow-up text representation models based on cognitive psychology theories. The intuition of our study is that the knowledge implied in a large collection of documents may improve the understanding of single documents. Based on cognitive psychology theories, we propose a general text enrichment framework, study the key factors to enable activation of implicit information, and develop new text representation methods to enrich text with the implicit information. Our study aims to mimic some aspects of human cognitive procedure in which given stimulant words serve to activate understanding implicit concepts. By incorporating human cognition into text representation, the proposed models advance existing studies by mining implicit information from given text and coordinating with most existing text representation approaches at the same time, which essentially bridges the gap between explicit and implicit information. Experiments on multiple tasks show that the implicit information activated by our proposed models matches human intuition and significantly improves the performance of the text mining tasks as well.
Linjing Li, Daniel Dajun Zeng, Qiudan Li
IEEE Trans. Knowl. Data Eng.4
2016 Predicting user's multi-interests with network embedding in health-related topics
abstract
With the rapid growth of Web 2.0, social media has become a prevalent information sharing and seeking channel for health surveillance, in which users form interactive networks by posting and replying messages, providing and rating reviews, attending multiple discussion boards on health-related topics. Users' behaviors in these interactive networks reflect users' multiple interests. To provide better information service for users, it is necessary to analyze the user interactions and predict users' multi-interests. Most existing work in predicting users' multi-interests based on multi label network classification focuses on using approximate inference methods to leverage the dependency information to improve classification results. Inspired by deep learning techniques, DEEPWALK learns label independent latent representations of vertices in a network using local information obtained from truncated random walks, which provides an efficient way for predicting users multi-interests from user interactions. In this paper, we develop a user's multi-interests prediction model based on DEEPWALK, weight information of user interactions is considered when modeling a stream of short constrained random walks and SkipGram is employed to generate more accurate representations of user vertices, which help identify users' interests. Experimental results on two real world health-related datasets show the efficacy of the proposed model.
Zhipeng Jin, Ruoran Liu, Qiudan Li, Daniel Dajun Zeng, Yongcheng Zhan, Lei Wang 0062
IJCNN3
2016 Jointly Modeling Review Content and Aspect Ratings for Review Rating Prediction
abstract
Review rating prediction is of much importance for sentiment analysis and business intelligence. Existing methods work well when aspect-opinion pairs can be accurately extracted from review texts and aspect ratings are complete. The challenges of improving prediction accuracy are how to capture the semantics of review content and how to fill in the missing values of aspect ratings. In this paper, we propose a novel review rating prediction method, which improves the prediction accuracy by capturing deep semantics of review content and alleviating data missing problem of aspect ratings. The method firstly learns the latent vector representation of review content using skip-thought vectors, a state-of-the-art deep learning method, then, the missing values of aspect ratings are filled in based on users? history reviewing behaviors, finally, a novel optimization framework is proposed to predict the review rating. Experimental results on two real-world datasets demonstrate the efficacy of the proposed method.
Zhipeng Jin, Qiudan Li, Daniel Dajun Zeng, Yongcheng Zhan, Ruoran Liu, Lei Wang 0062, Hongyuan Ma
SIGIR2
2016 Mining opinion summarizations using convolutional neural networks in Chinese microblogging systems
Qiudan Li, Zhipeng Jin, Daniel Dajun Zeng
Knowl. Based Syst.1
2015 Filtering spam in Weibo using ensemble imbalanced classification and knowledge expansion
abstract
Weibo has become an important information sharing platform in our daily life in China. Many applications utilize Weibo data to analyze hot topic and opinion evolution patterns to gain insights into user behavior. However, various spam messages degrade the performance of these applications and thus are essential to be filtered. In this paper, we propose a unified spam detection approach, which utilizes external knowledge sources to expand keywords features and applies an ensemble under-sampling based strategy to handle the class-imbalance problem. The experimental results show the effectiveness and robustness of our approach in Weibo data.
Zhipeng Jin, Qiudan Li, Daniel Dajun Zeng, Lei Wang 0062
ISI2
2014 Entity attribute discovery and clustering from online reviews
Qingliang Miao, Qiudan Li, Daniel Dajun Zeng, Shu Zhang 0004, Hao Yu 0005
Frontiers Comput. Sci.2
2014 Extracting evolutionary communities in community question answering
abstract
With the rapid growth of Web 2.0, community question answering (CQA) has become a prevalent information seeking channel, in which users form interactive communities by posting questions and providing answers. Communities may evolve over time, because of changes in users' interests, activities, and new users joining the network. To better understand user interactions inCQAcommunities, it is necessary to analyze the community structures and track community evolution over time. Existing work inCQAfocuses on question searching or content quality detection, and the important problems of community extraction and evolutionary pattern detection have not been studied. In this article, we propose a probabilistic community model (PCM) to extract overlapping community structures and capture their evolution patterns inCQA. The empirical results show that our algorithm appears to improve the community extraction quality. We show empirically, using the iPhone data set, that interesting community evolution patterns can be discovered, with each evolution pattern reflecting the variation of users' interests over time. Our analysis suggests that individual users could benefit to gain comprehensive information from tracking the transition of products. We also show that the communities provide a decision‐making basis for business.
Zhongfeng Zhang, Qiudan Li, Daniel Dajun Zeng
J. Assoc. Inf. Sci. Technol.2
2013 A new temporal and social PMF-based method to predict users' interests in micro-blogging
Hongyun Bao, Qiudan Li, Stephen Shaoyi Liao, Shuangyong Song
Decis. Support Syst.2
2013 User community discovery from multi-relational networks
Zhongfeng Zhang, Qiudan Li, Daniel Dajun Zeng
Decis. Support Syst.2
2012 Detecting popular topics in micro-blogging based on a user interest-based model
abstract
The rapid increasing popularity of micro-blogging has made it an important information seeking channel. By detecting recent popular topics from micro-blogging, we have opportunities to gain insights into internet hotspots. Generally, a topic's popularity is determined by two primary factors. One is how frequently a topic is discussed by users, and the other is how much influence those users have, since topics shown in the influential users' posts are more likely to attract others' attention. However, existing approaches interpret a topic's popularity with only the number of keywords related to it, which neglect the importance of the user influence to information diffusion in micro-blogging. In this paper, drawing upon the Cognitive Authority Theory and Social Network Theory, we propose a novel model that detects the most popular topics in micro-blogging with a user interest-based method. The proposed model first constructs a topic graph according to users' interests and their following relationship, and then calculates the topics' popularity with a link-based ranking algorithm. The popular topics detected by the method can reflect the relationship among users' interests, and the topics in the posts of influential users can be highlighted. Experimental results on the data of Twitter, a well-known and feature-rich micro-blogging service, show that the proposed method is effective in popular topic discovery.
Shuangyong Song, Qiudan Li, Xiaolong Zheng 0001
IJCNN2
2012 A graph-based action network framework to identify prestigious members through member's prestige evolution
Dongyuan Lu, Qiudan Li, Stephen Shaoyi Liao
Decis. Support Syst.2
2011 QuestionHolic: Hot topic discovery and trend analysis in community question answering systems
Zhongfeng Zhang, Qiudan Li
Expert Syst. Appl.2
2011 A recommender system based on tag and time information for social tagging systems
Qiudan Li
Expert Syst. Appl.2
2011 Mining Evolutionary Topic Patterns in Community Question Answering Systems
abstract
Community Question Answering (CQA) is becoming a popular Web 2.0 application. By analyzing evolutionary topic patterns from CQA applications, one can gain insights into user interests and user responses to external events. This paper proposes a novel evolutionary topic pattern mining approach. This approach consists of three components: 1) extraction of the topics being discussed through a temporal analysis; 2) discovery of topic evolutions and construction of evolutionary graphs of extracted topics; and 3) life cycle modeling of the extracted topics. We show empirically the effectiveness of our approach using two real-world data sets.
Zhongfeng Zhang, Qiudan Li, Daniel Dajun Zeng
IEEE Trans. Syst. Man Cybern. Part A2
2010 Flickr group recommendation based on tensor decomposition
abstract
Over the last few years, Flickr has gained massive popularity and groups in Flickr are one of the main ways for photo diffusion. However, the huge volume of groups brings troubles for users to decide which group to choose. In this paper, we propose a tensor decomposition-based group recommendation model to suggest groups to users which can help tackle this problem. The proposed model measures the latent associations between users and groups by considering both semantic tags and social relations. Experimental results show the usefulness of the proposed model.
Qiudan Li, Shengcai Liao, Leiming Zhang
SIGIR2
2010 Mining Fine Grained Opinions by Using Probabilistic Models and Domain Knowledge
abstract
The explosive growth of the user-generated content on the Web has offered a rich data source for mining opinions. However, the large number of diverse review sources challenges the individual users and organizations on how to use the opinion information effectively. Therefore, automated opinion mining and summarization techniques have become increasingly important. Different from previous approaches that have mostly treated product feature and opinion extraction as two independent tasks, we merge them together in a unified process by using probabilistic models. Specifically, we treat the problem of product feature and opinion extraction as a sequence labeling task and adopt Conditional Random Fields models to accomplish it. As part of our work, we develop a computational approach to construct domain specific sentiment lexicon by combining semi-structured reviews with general sentiment lexicon, which helps to identify the sentiment orientations of opinions. Experimental results on two real world datasets show that the proposed method is effective.
Qingliang Miao, Qiudan Li, Daniel Dajun Zeng
Web Intelligence2
2010 Fine-grained opinion mining by integrating multiple review sources
abstract
Abstract With the rapid development of Web 2.0, online reviews have become extremely valuable sources for mining customers' opinions. Fine‐grained opinion mining has attracted more and more attention of both applied and theoretical research. In this article, the authors study how to automatically mine product features and opinions from multiple review sources. Specifically, they propose an integration strategy to solve the issue. Within the integration strategy, the authors mine domain knowledge from semistructured reviews and then exploit the domain knowledge to assist product feature extraction and sentiment orientation identification from unstructured reviews. Finally, feature‐opinion tuples are generated. Experimental results on real‐world datasets show that the proposed approach is effective.
Qingliang Miao, Qiudan Li, Daniel Dajun Zeng
J. Assoc. Inf. Sci. Technol.2
2009 Link based small sample learning for web spam detection
abstract
Robust statistical learning based web spam detection system often requires large amounts of labeled training data. However, labeled samples are more difficult, expensive and time consuming to obtain than unlabeled ones. This paper proposed link based semi-supervised learning algorithms to boost the performance of a classifier, which integrates the traditional Self-training with the topological dependency based link learning. The experiments with a few labeled samples on standard WEBSPAM-UK2006 benchmark showed that the algorithms are effective.
Guanggang Geng, Qiudan Li, Xinchang Zhang 0001
WWW2
2009 AMAZING: A sentiment mining and retrieval system
Qingliang Miao, Qiudan Li, Ruwei Dai
Expert Syst. Appl.2
2008 An integration strategy for mining product features and opinions
abstract
With the development of Web 2.0, the web has become an extremely valuable source for mining opinions. In this paper, we study how to automatically mine product features and opinions by integrating multiple review sources. We propose an integration strategy to solve the problem. Experiments show that the proposed strategy is effective.
Qingliang Miao, Qiudan Li, Ruwei Dai
CIKM2
2008 An opinion search system for consumer products
abstract
With the rapid progress of e-commerce, many people like purchasing product on the e-commerce website, and giving their personal reviews to the product they purchased, so the number of customer reviews grows rapidly. Generally, a potential customer will browse product reviews before they purchase the product. However, retrieving opinions relevant to customer’s desire still remains challenging. To provide efficient opinion information for customers, we propose an opinion search system for consumer products, which utilizes data mining and information retrieval technology. A ranking mechanism taking temporal dimension into account and a method for results visualization are developed in the system. Experimental results on a real-world data set show the system is feasible and effective.
Qingliang Miao, Qiudan Li
IJCNN2
2008 A no reference image quality assessment method for JPEG2000
abstract
This paper presents a novel no reference method to assess image quality. Firstly, the image is divided into many blocks. Textured blocks are selected and their amplitude fall-off curves are employed for quality prediction based on natural scene statistics. Secondly, projections of wavelet coefficients between adjacent scales with the same orientation are utilized to measure the positional similarity. At last, general regression neural network is adopted to conduct quality prediction according to features from above two aspects. The performance of our method is evaluated on a public data set and experimental results confirm its effectiveness.
Jingchao Zhou, Baihua Xiao, Qiudan Li
IJCNN3
2008 A Unified Framework for Opinion Retrieval
abstract
The popularity of Web 2.0 has promoted the web to be a valuable source for accessing opinions. Unfortunately, due to the large number of user generated content, it is difficult to access and utilize the opinion resource efficiently. Developing an opinion retrieval system is a promising way to overcome the problem of overloaded opinion information. In this paper, we propose a unified framework for opinion retrieval, which is based on generative model and opinion mining technologies. Within this framework, relevance, quality and temporal dimension information of user generated content are incorporated in a unified language model, and on the top of which an opinion summary is proposed. We have developed an Opinion Retrieval System, (ORS), to retrieve opinions in customer review domain. Our evaluation on a real-world data set shows that ORS can effectively retrieve, summarize and visualize customer opinions.
Qingliang Miao, Qiudan Li, Ruwei Dai
Web Intelligence2
2008 Improving web spam detection with re-extracted features
abstract
Web spam detection has become one of the top challenges for the Internet search industry. Instead of using some heuristic rules, we propose a feature re-extraction strategy to optimize the detection result. Based on the predicted spamicity obtained by the preliminary detection, through the host level web graph, three types of features are extracted. Experiments on WEBSPAM-UK2006 benchmark show that with this strategy, the performance of web spam detection can be improved evidently. Categories and Subject Descriptors
Guanggang Geng, Chunheng Wang, Qiudan Li
WWW3
2008 Improving personalized services in mobile commerce by a novel multicriteria rating approach
abstract
With the rapid growth of wireless technologies and mobile devices, there is a great demand for personalized services in m-commerce. Collaborative filtering (CF) is one of successful techniques to produce personalized recommendations for users. This paper proposes a novel approach to improve CF algorithms, where the contextual information of a user and the multicriteria ratings of an item are considered besides the typical information on users and items. The multilinear singular value decomposition (MSVD) technique is utilized to explore both explicit relations and implicit relations among user, item and criterion. We implement the approach in an existing m-commerce platform, and encouraging experimental results demonstrate its effectiveness.
Qiudan Li, Chunheng Wang, Guanggang Geng
WWW1
2008 Combining empirical experimentation and modeling techniques: A design research approach for personalized mobile advertising applications
David Jingjun Xu, Stephen Shaoyi Liao, Qiudan Li
Decis. Support Syst.3
2007 Fighting Link Spam with a Two-Stage Ranking Strategy
Guanggang Geng, Chunheng Wang, Qiudan Li, Yuanping Zhu
ECIR3
2007 A Concept Lattice-Based Kernel Method for Mining Knowledge in an M-Commerce System
Qiudan Li, Chunheng Wang, Guanggang Geng, Ruwei Dai
ISNN (1)1
2007 A novel collaborative filtering-based framework for personalized services in m-commerce
abstract
With the rapid growth of wireless technologies and handheld devices, m-commerce is becoming a promising research area. Personalization is especially important to the success of m-commerce. This paper proposes a novel collaborative filtering-based framework for personalized services in m-commerce. The framework extends our previous work by using Online Analytical Processing (OLAP) to represent the relations among user, content and context information, and adopting a multi-dimensional collaborative filtering model to perform inference. It provides a powerful and well-founded mechanism to personalization for m-commerce. We implemented it in an existing m-commerce platform, and experimental results demonstrate its feasibility and correctness.
Qiudan Li, Chunheng Wang, Guanggang Geng, Ruwei Dai
WWW1