Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Huanhuan Cao

dblp:48/5943 · DBLP profile ↗
← Back
34ranked-venue papers
7as first author
3since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 26 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Data mining · 41% Recommender systems · 32% Information retrieval · 18%
Artificial intelligence
3 papers
Face, body and person analysis · 49% Transfer learning and domain adaptation · 32% Representation and self-supervised learning · 19%

Topics — the 29 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems › automated model design
embedding dimension search
0.812024
I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024
Data mining
feature engineering
0.812024
I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024
Data mining › dimensionality reduction
feature selection
0.812024
I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024
Computer vision › Face, body and person analysis
person re-identification
0.512021
Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training
0.512021
Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021
Computer vision › Face, body and person analysis › person re-identification › unsupervised person re-identification
unsupervised domain adaptive person re-identification
0.512021
Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.412019
Zero-shot Metric Learning · IJCAI 2019
Recommender systems
click-through rate prediction
0.212024
I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024
Web and social media mining
app categorization
0.212014
Mobile App Classification with Enriched Contextual Information · IEEE Trans. Mob. Comput. 2014
Data mining › text mining
text classification
0.212014
Mobile App Classification with Enriched Contextual Information · IEEE Trans. Mob. Comput. 2014
Information retrieval
query suggestion
0.222009
Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009
Context-aware query suggestion by mining click-through and session data · KDD 2008
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-network transfer
0.112012
Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012
Data mining › pattern mining
behavioral pattern mining
0.112012
A habit mining approach for discovering similar mobile users · WWW 2012
Recommender systems
context-aware recommendation
0.112012
Mining Personal Context-Aware Preferences for Mobile Users · ICDM 2012
Recommender systems › cross-domain recommendation
cross-network recommendation
0.112012
Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012
Knowledge graphs
link prediction
0.112012
Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012
Data mining
pattern mining
0.112012
A habit mining approach for discovering similar mobile users · WWW 2012
Recommender systems › user modeling
user preference learning
0.112012
Mining Personal Context-Aware Preferences for Mobile Users · ICDM 2012
Recommender systems
user similarity
0.112012
A habit mining approach for discovering similar mobile users · WWW 2012
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
multimedia annotation
0.112012
Towards Annotating Media Contents through Social Diffusion Analysis · ICDM 2012
Information retrieval
contextual search
0.112009
Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009
Information retrieval › reranking
document re-ranking
0.112009
Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009
Information retrieval › query understanding
query classification
0.112009
Context-aware query classification · SIGIR 2009
Information retrieval
query understanding
0.112009
Context-aware query classification · SIGIR 2009
Information retrieval › query log analysis › clickthrough data
click-through data mining
0.112008
Context-aware query suggestion by mining click-through and session data · KDD 2008
Information retrieval › query suggestion
context-aware query suggestion
0.112008
Context-aware query suggestion by mining click-through and session data · KDD 2008
Information retrieval
query log analysis
0.112008
Context-aware query suggestion by mining click-through and session data · KDD 2008
Information retrieval › query understanding › query modeling
query context modeling
0.012009
Context-aware query classification · SIGIR 2009
Information retrieval › user behavior
search session analysis
0.012009
Context-aware query classification · SIGIR 2009

Methods — techniques the papers use, named apart from their topics

pruning · 0.8differentiable neural architecture search · 0.8view-invariant representation learning · 0.5pseudo-labeling · 0.5multi-scale representation · 0.5gradient reverse layer · 0.5convolutional neural network · 0.4web search enrichment · 0.4maximum entropy model · 0.4contextual feature extraction · 0.4transfer learning · 0.3ranking factor graph · 0.3optimization · 0.3probabilistic modeling · 0.1context log mining · 0.1constraint-based factorization · 0.1bayesian matrix factorization · 0.1
YearPublicationVenuePosition
2024 I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems
abstract
Input features play a crucial role in DNN-based recommender systems with thousands of categorical and continuous fields from users, items, contexts, and interactions. Noisy features and inappropriate embedding dimension assignments can deteriorate the performance of recommender systems and introduce unnecessary complexity in model training and online serving. Optimizing the input configuration of DNN models, including feature selection and embedding dimension assignment, has become one of the essential topics in feature engineering. However, in existing industrial practices, feature selection and dimension search are optimized sequentially, i.e., feature selection is performed first, followed by dimension search to determine the optimal dimension size for each selected feature. Such a sequential optimization mechanism increases training costs and risks generating suboptimal input configurations. To address this problem, we propose a differentiable neuralinputrazor(i-Razor) that enables joint optimization of feature selection and dimension search. Concretely, we introduce an end-to-end differentiable model to learn the relative importance of different embedding regions of each feature. Furthermore, a flexible pruning algorithm is proposed to achieve feature filtering and dimension derivation simultaneously. Extensive experiments on two large-scale public datasets in the Click-Through-Rate (CTR) prediction task demonstrate the efficacy and superiority of i-Razor in balancing model complexity and performance.
Yao Yao 0006, Bin Liu 0072, Haoxun He, Dakui Sheng, Li Xiao 0006, Huanhuan Cao
IEEE Trans. Knowl. Data Eng.7
2022 How Online Reviews Interact with a Firm's Free Version Strategy
Huanhuan Cao, Jinhu Jiang, Xianjun Geng
Inf. Manag.1
2021 Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification
abstract
In recent years, person re-identification (re-ID) has achieved relatively good performance, benefiting from the revival of deep neural networks. However, due to the existence of domain bias which refers to the different data distributions between two domains, it remains challenging to directly deploy a model trained on a labeled source domain to a target domain only with unlabeled data available. In this paper, a Self-Training with Progressive Representation Enhancement (PREST) framework, which comprises a multi-scale self-training method and a view-invariant representation learning module, is proposed to promote re-ID performance on the target domain in an unsupervised manner. More specifically, multi-scale representations, including the global body and local parts of pedestrian images, are utilized to obtain pseudo-labels. Then, some images are selected according to the pseudo-labels to create a new dataset for supervising the fine-tuning process, which is operated iteratively to progressively promote the performance. Furthermore, to mitigate the influence of different styles among sub-domains, in cases where a single sub-domain is captured by one camera, a classifier with a gradient reverse layer is first employed to learn view-invariant representation for pedestrian images with the same identity taken by different cameras; this can further enhance the reliability of the predicted labels and improve the cross-domain re-ID performance. Extensive experiments on three large-scale re-ID datasets demonstrate that our framework achieves significantly better performance than existing approaches.
Huanhuan Cao, Xu Yang 0019, Cheng Deng 0002, Dacheng Tao
IEEE Trans. Image Process.2
2020 Online review manipulation by asymmetrical firms: Is a firm's manipulation of online reviews always detrimental to its competitor?
Huanhuan Cao
Inf. Manag.1
2020 Object and background disentanglement for unsupervised cross-domain person re-identification
Yuehua Zhu, Cheng Deng 0002, Huanhuan Cao, Hao Wang 0062
Neurocomputing3
2019 Zero-shot Metric Learning
abstract
In this work, we tackle the zero-shot metric learning problem and propose a novel method abbreviated as ZSML, with the purpose to learn a distance metric that measures the similarity of unseen categories (even unseen datasets). ZSML achieves strong transferability by capturing multi-nonlinear yet continuous relation among data. It is motivated by two facts: 1) relations can be essentially described from various perspectives; and 2) traditional binary supervision is insufficient to represent continuous visual similarity. Specifically, we first reformulate a collection of specific-shaped convolutional kernels to combine data pairs and generate multiple relation vectors. Furthermore, we design a new cross-update regression loss to discover continuous similarity. Extensive experiments including intra-dataset transfer and inter-dataset transfer on four benchmark datasets demonstrate that ZSML can achieve state-of-the-art performance.
Huanhuan Cao, Yanhua Yang, Erkun Yang, Cheng Deng 0002
IJCAI2
2018 Question Headline Generation for News Articles
abstract
In this paper, we introduce and tackle the Question Headline Generation (QHG) task. The motivation comes from the investigation of a real-world news portal where we find that news articles with question headlines often receive much higher click-through ratio than those with non-question headlines. The QHG task can be viewed as a specific form of the Question Generation (QG) task, with the emphasis on creating a natural question from a given news article by taking the entire article as the answer. A good QHG model thus should be able to generate a question by summarizing the essential topics of an article. Based on this idea, we propose a novel dual-attention sequence-to-sequence model (DASeq2Seq) for the QHG task. Unlike traditional sequence-to-sequence models which only employ the attention mechanism in the decoding phase for better generation, our DASeq2Seq further introduces a self-attention mechanism in the encoding phase to help generate a good summary of the article. We investigate two ways of the self-attention mechanism, namely global self-attention and distributed self-attention. Besides, we employ a vocabulary gate over both generic and question vocabularies to better capture the question patterns. Through the offline experiments, we show that our approach can significantly outperform the state-of-the-art question generation or headline generation models. Furthermore, we also conduct online evaluation to demonstrate the effectiveness of our approach using A/B test.
Ruqing Zhang 0001, Jiafeng Guo, Yixing Fan, Yanyan Lan, Jun Xu 0001, Huanhuan Cao, Xueqi Cheng 0001
CIKM6
2014 Learning to detect subway arrivals for passengers on a train
Kuifei Yu, Hengshu Zhu, Huanhuan Cao, Baoxian Zhang, Enhong Chen, Jilei Tian, Jinghai Rao
Frontiers Comput. Sci.3
2014 Mining Mobile User Preferences for Personalized Context-Aware Recommendation
abstract
Recent advances in mobile devices and their sensing capabilities have enabled the collection of rich contextual information and mobile device usage records through the device logs. These context-rich logs open a venue for mining the personal preferences of mobile users under varying contexts and thus enabling the development of personalized context-aware recommendation and other related services, such as mobile online advertising. In this article, we illustrate how to extract personal context-aware preferences from the context-rich device logs, or context logs for short, and exploit these identified preferences for building personalized context-aware recommender systems. A critical challenge along this line is that the context log of each individual user may not contain sufficient data for mining his or her context-aware preferences. Therefore, we propose to first learn common context-aware preferences from the context logs of many users. Then, the preference of each user can be represented as a distribution of these common context-aware preferences. Specifically, we develop two approaches for mining common context-aware preferences based on two different assumptions, namely, context-independent and context-dependent assumptions, which can fit into different application scenarios. Finally, extensive experiments on a real-world dataset show that both approaches are effective and outperform baselines with respect to mining personal context-aware preferences for mobile users.
Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Kuifei Yu, Huanhuan Cao, Jilei Tian
ACM Trans. Intell. Syst. Technol.5
2014 Mobile App Classification with Enriched Contextual Information
abstract
The study of the use of mobile Apps plays an important role in understanding the user preferences, and thus provides the opportunities for intelligent personalized context-based services. A key step for the mobile App usage analysis is to classify Apps into some predefined categories. However, it is a nontrivial task to effectively classify mobile Apps due to the limited contextual information available for the analysis. For instance, there is limited contextual information about mobile Apps in their names. However, this contextual information is usually incomplete and ambiguous. To this end, in this paper, we propose an approach for first enriching the contextual information of mobile Apps by exploiting the additional Web knowledge from the Web search engine. Then, inspired by the observation that different types of mobile Apps may be relevant to different real-world contexts, we also extract some contextual features for mobile Apps from the context-rich device logs of mobile users. Finally, we combine all the enriched contextual information into the Maximum Entropy model for training a mobile App classifier. To validate the proposed method, we conduct extensive experiments on 443 mobile users' device logs to show both the effectiveness and efficiency of the proposed approach. The experimental results clearly show that our approach outperforms two state-of-the-art benchmark methods with a significant margin.
Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Huanhuan Cao, Jilei Tian
IEEE Trans. Mob. Comput.4
2014 Ranking user authority with relevant knowledge categories for expert finding
Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Huanhuan Cao, Jilei Tian
World Wide Web4
2013 Learning to Detect the Subway Station Arrival for Mobile Users
Kuifei Yu, Hengshu Zhu, Huanhuan Cao, Baoxian Zhang, Enhong Chen, Jilei Tian, Jinghai Rao
IDEAL3
2013 Predicting Stay Time of Mobile Users With Contextual Information
abstract
Mobile service providers and manufacturers continue to provide services and devices that take advantage of the location information associated with devices to provide a more personalized experience for users. For many such services, the user experience can be dramatically improved if a mobile device can predict how long a mobile user will stay at the current location. In this paper, we propose to take advantage of contextual information for predicting the stay time of mobile users. Specially, we investigate two strategies for modeling the relevance between it and contextual information, i.e., Stay Status Prediction (SSP) and Stay Time Prediction (STP). SSP is to predict whether a mobile user will stay at the current location at time point ti+naccording to the contextual information at ti, while STP is to directly predict how long a mobile user will stay at the current location. Moreover, we study several typical machine learning models which can be extended for implementing SSP and STP and evaluate their performance with respect to prediction accuracy. We also conduct extensive experiments on real data sets to evaluate several implementations of the proposed strategies in terms of both effectiveness and efficiency for STP.
Huanhuan Cao, Lei Li 0009, MengChu Zhou
IEEE Trans Autom. Sci. Eng.2
2013 A vlHMM approach to context-aware search
abstract
Capturing the context of a user's query from the previous queries and clicks in the same session leads to a better understanding of the user's information need. A context-aware approach to document reranking, URL recommendation, and query suggestion may substantially improve users' search experience. In this article, we propose a general approach to context-aware search by learning avariable length hidden Markov model(vlHMM) from search sessions extracted from log data. While the mathematical model is powerful, the huge amounts of log data present great challenges. We develop several distributed learning techniques to learn a very large vlHMM under themap-reduceframework. Moreover, we construct feature vectors for each state of the vlHMM model to handle users' novel queries not covered by the training data. We test our approach on a raw dataset consisting of 1.9 billion queries, 2.9 billion clicks, and 1.2 billion search sessions before filtering, and evaluate the effectiveness of the vlHMM learned from the real data on three search applications: document reranking, query suggestion, and URL recommendation. The experiment results validate the effectiveness of vlHMM in the applications of document reranking, URL recommendation, and query suggestion.
Zhen Liao, Daxin Jiang, Jian Pei 0001, Yalou Huang, Enhong Chen, Huanhuan Cao, Hang Li 0001
ACM Trans. Web6
2012 Exploiting enriched contextual information for mobile app classification
abstract
A key step for the mobile app usage analysis is to classify apps into some predefined categories. However, it is a nontrivial task to effectively classify mobile apps due to the limited contextual information available for the analysis. To this end, in this paper, we propose an approach to first enrich the contextual information of mobile apps by exploiting the additional Web knowledge from the Web search engine. Then, inspired by the observation that different types of mobile apps may be relevant to different real-world contexts, we also extract some contextual features for mobile apps from the context-rich device logs of mobile users. Finally, we combine all the enriched contextual information into a Maximum Entropy model for training a mobile app classifier. The experimental results based on 443 mobile users' device logs clearly show that our approach outperforms two state-of-the-art benchmark methods with a significant margin.
Hengshu Zhu, Huanhuan Cao, Enhong Chen, Hui Xiong 0001, Jilei Tian
CIKM2
2012 Link Prediction and Recommendation across Heterogeneous Social Networks
abstract
Link prediction and recommendation is a fundamental problem in social network analysis. The key challenge of link prediction comes from the sparsity of networks due to the strong disproportion of links that they have potential to form to links that do form. Most previous work tries to solve the problem in single network, few research focus on capturing the general principles of link formation across heterogeneous networks. In this work, we give a formal definition of link recommendation across heterogeneous networks. Then we propose a ranking factor graph model (RFG) for predicting links in social networks, which effectively improves the predictive performance. Motivated by the intuition that people make friends in different networks with similar principles, we find several social patterns that are general across heterogeneous networks. With the general social patterns, we develop a transfer-based RFG model that combines them with network structure information. This model provides us insight into fundamental principles that drive the link formation and network evolution. Finally, we verify the predictive performance of the presented transfer model on 12 pairs of transfer cases. Our experimental results demonstrate that the transfer of general social patterns indeed help the prediction of links.
Yuxiao Dong, Jie Tang 0001, Sen Wu 0001, Jilei Tian, Nitesh V. Chawla, Jinghai Rao, Huanhuan Cao
ICDM7
2012 Towards Annotating Media Contents through Social Diffusion Analysis
abstract
Recently, the boom of media contents on the Internet raises challenges in managing them effectively and thus requires automatic media annotation techniques. Motivated by the observation that media contents are usually shared frequently in online communities and thus have a lot of social diffusion records, we propose a novel media annotating approach depending on these social diffusion records instead of metadata. The basic assumption is that the social diffusion records reflect the common interests (CI) between users, which can be analyzed for generating annotations. With this assumption, we present a novel CI-based social diffusion model and translate the automatic annotating task into the CI-based diffusion maximization (CIDM) problem. Moreover, we propose to solve the CIDM problem through two optimization tasks, corresponding to the training and test stages in supervised learning. Extensive experiments on real-world data sets show that our approach can effectively generate high quality annotations, and thus demonstrate the capability of social diffusion analysis in annotating media.
Tong Xu 0001, Dong Liu 0002, Enhong Chen, Huanhuan Cao, Jilei Tian
ICDM4
2012 Mining Personal Context-Aware Preferences for Mobile Users
abstract
In this paper, we illustrate how to extract personal context-aware preferences from the context-rich device logs (i.e., context logs) for building novel personalized context-aware recommender systems. A critical challenge along this line is that the context log of each individual user may not contain sufficient data for mining his/her context-aware preferences. Therefore, we propose to first learn common context-aware preferences from the context logs of many users. Then, the preference of each user can be represented as a distribution of these common context-aware preferences. Specifically, we develop two approaches for mining common context-aware preferences based on two different assumptions, namely, context independent and context dependent assumptions, which can fit into different application scenarios. Finally, extensive experiments on a real-world data set show that both approaches are effective and outperform baselines with respect to mining personal context-aware preferences for mobile users.
Hengshu Zhu, Enhong Chen, Kuifei Yu, Huanhuan Cao, Hui Xiong 0001, Jilei Tian
ICDM4
2012 Mining Significant Places from Cell ID Trajectories: A Geo-grid Based Approach
abstract
Mining the frequently visited places of single mobile users, i.e., significant places, is crucial for supporting personalized location-based services. Most of existing works for significance place mining have a need to take advantage the GPS trajectories of users. However, it is difficult to encourage mobile users to contribute GPS trajectories because of the high power consumption of GPS. In this paper, we propose a geo-grid based approach for mining significant places from cell ID trajectories. In our approach, the mined significant places are represented as sets of geo-grids which are much smaller than the coverage areas of cell-sites. To be specific, we firstly extract the stay areas where the mobile user used to stay and map them to many geo-grids. Then we mine significant places from the geo-grids by considering their significance. We evaluate the approach on real word data sets and the experimental results clearly show that the proposed approach outperforms two baselines.
Tengfei Bao, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian
MDM2
2012 A Demonstration of Mining Significant Places from Cell ID Trajectories through a Geo-grid Based Approach
abstract
Mining the frequently visited places of single mobile users, i.e., significant places, is crucial for supporting personalized location-based services. Most of existing works for significance place mining have a need to take advantage the GPS trajectories of users. However, it is difficult to encourage mobile users to contribute GPS trajectories because of the high power consumption of GPS. In this demonstration, we propose a geo-grid based approach for mining significant places from cell ID trajectories. In our approach, the mined significant places are represented as sets of geo-grids which are much smaller than the coverage areas of cell-sites. To be specific, we firstly extract the stay areas where the mobile user used to stay and map them to many geogrids. Then we mine significant places from the geo-grids by considering their significance.
Tengfei Bao, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian
MDM2
2012 BP-growth: Searching Strategies for Efficient Behavior Pattern Mining
abstract
User habit mining plays an important role in user understanding, which is critical for improving a wide range of personalized intelligence services. Recently, some researchers proposed to mine user behavior patterns which characterize the habits of mobile users and account for the associations between user interactions and context captured by mobile devices. However, the existing approaches for mining these behavior patterns are not practical in mobile environments due to limited computing resources on mobile devices. To fulfill this crucial void, we investigate optimizing strategies which can be used for improving the efficiency of behavior pattern mining in terms of computing and memory needs. Specifically, we examine typical optimizing strategies for association rule mining and study the feasibility of applying them to behavior pattern mining, since these two problems are similar in many aspects. Moreover, we develop an efficient algorithm, named BP-Growth, for behavior pattern mining by combining two promising strategies. Finally, experimental results show that BP-Growth outperforms benchmark methods with a significant margin in terms of both computing and memory cost.
Xueying Li 0004, Huanhuan Cao, Enhong Chen, Hui Xiong 0001, Jilei Tian
MDM2
2012 Towards Personalized Context-Aware Recommendation by Mining Context Logs through Topic Models
Kuifei Yu, Baoxian Zhang, Hengshu Zhu, Huanhuan Cao, Jilei Tian
PAKDD (1)4
2012 A habit mining approach for discovering similar mobile users
abstract
Discovering similar users with respect to their habits plays an important role in a wide range of applications, such as collaborative filtering for recommendation, user segmentation for market analysis, etc. Recently, the progressing ability to sense user contexts of smart mobile devices makes it possible to discover mobile users with similar habits by mining their habits from their mobile devices. However, though some researchers have proposed effective methods for mining user habits such as behavior pattern mining, how to leverage the mined results for discovering similar users remains less explored. To this end, we propose a novel approach for conquering the sparseness of behavior pattern space and thus make it possible to discover similar mobile users with respect to their habits by leveraging behavior pattern mining. To be specific, first, we normalize the raw context log of each user by transforming the location-based context data and user interaction records to more general representations. Second, we take advantage of a constraint-based Bayesian Matrix Factorization model for extracting the latent common habits among behavior patterns and then transforming behavior pattern vectors to the vectors of mined common habits which are in a much more dense space. The experiments conducted on real data sets show that our approach outperforms three baselines in terms of the effectiveness of discovering similar mobile users with respect to their habits.
Haiping Ma, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian
WWW2
2012 An unsupervised approach to modeling personalized contexts of mobile users
Tengfei Bao, Huanhuan Cao, Enhong Chen, Jilei Tian, Hui Xiong 0001
Knowl. Inf. Syst.2
2012 Learning to Infer the Status of Heavy-Duty Sensors for Energy-Efficient Context-Sensing
abstract
With the prevalence of smart mobile devices with multiple sensors, the commercial application of intelligent context-aware services becomes more and more attractive. However, limited by the battery capacity, the energy efficiency of context-sensing is the bottleneck for the success of context-aware applications. Though several previous studies for energy-efficient context-sensing have been reported, none of them can be applied to multiple types of high-energy-consuming sensors. Moreover, applying machine learning technologies to energy-efficient context-sensing is underexplored too. In this article, we propose to leverage machine learning technologies for improving the energy efficiency of multiple high-energy-consuming context sensors by trading off the sensing accuracy. To be specific, we try to infer the status of high-energy-consuming sensors according to the outputs of software-based sensors and the physical sensors that are necessary to work all the time for supporting the basic functions of mobile devices. If the inference indicates the high-energy-consuming sensor is in a stable status, we avoid the unnecessary invocation and instead use the latest invoked value as the estimation. The experimental results on real datasets show that the energy efficiency of GPS sensing and audio-level sensing are significantly improved by the proposed approach while the sensing accuracy is over 90%.
Xueying Li 0004, Huanhuan Cao, Enhong Chen, Jilei Tian
ACM Trans. Intell. Syst. Technol.2
2011 Towards expert finding by leveraging relevant categories in authority ranking
abstract
How to improve authority ranking is a crucial research problem for expert finding. In this paper, we propose a novel framework for expert finding based on the authority information in the target category as well as the relevant categories. First, we develop a scalable method for measuring the relevancy between categories through topic models. Then, we provide a link analysis approach for ranking user authority by considering the information in both the target category and the relevant categories. Finally, the extensive experiments on two large-scale real-world Q&A data sets clearly show that the proposed method outperforms the baseline methods with a significant margin.
Hengshu Zhu, Huanhuan Cao, Hui Xiong 0001, Enhong Chen, Jilei Tian
CIKM2
2011 Finding Experts in Tag Based Knowledge Sharing Communities
Hengshu Zhu, Enhong Chen, Huanhuan Cao
KSEM3
2011 Mining Concept Sequences from Large-Scale Search Logs for Context-Aware Query Suggestion
abstract
Query suggestion plays an important role in improving usability of search engines. Although some recently proposed methods provide query suggestions by mining query patterns from search logs, none of them models the immediately preceding queries as context systematically, and uses context information effectively in query suggestions. Context-aware query suggestion is challenging in both modeling context and scaling up query suggestion using context. In this article, we propose a novel context-aware query suggestion approach. To tackle the challenges, our approach consists of two stages. In the first, offline model-learning stage , to address data sparseness, queries are summarized into concepts by clustering a click-through bipartite. A concept sequence suffix tree is then constructed from session data as a context-aware query suggestion model. In the second, online query suggestion stage , a user’s search context is captured by mapping the query sequence submitted by the user to a sequence of concepts. By looking up the context in the concept sequence suffix tree, we suggest to the user context-aware queries. We test our approach on large-scale search logs of a commercial search engine containing 4.0 billion Web queries, 5.9 billion clicks, and 1.87 billion search sessions. The experimental results clearly show that our approach outperforms three baseline methods in both coverage and quality of suggestions.
Zhen Liao, Daxin Jiang, Enhong Chen, Jian Pei 0001, Huanhuan Cao, Hang Li 0001
ACM Trans. Intell. Syst. Technol.5
2010 An effective approach for mining mobile user habits
abstract
The user interaction with the mobile device plays an important role in user habit understanding. In this paper, we propose to mine the associations between user interactions and contexts captured by mobile devices, or behavior patterns for short, from context logs to characterize the habits of mobile users. The extensive experiments on the collected real life data clearly validate the ability of our approach for mining effective behavior patterns.
Huanhuan Cao, Tengfei Bao, Qiang Yang 0001, Enhong Chen, Jilei Tian
CIKM1
2009 Enhancing recommender systems under volatile userinterest drifts
abstract
This paper presents a systematic study of how to enhance recommender systems under volatile user interest drifts. A key development challenge along this line is how to track user interests dynamically. To this end, we first define four types of interest patterns to understand users' rating behaviors and analyze the properties of these patterns. We also propose a rating graph and rating chain based approach for detecting these interest patterns. For each users' rating series, a rating graph and a rating chain are constructed based on the similarities between rated items. The type of a given user's interest pattern is identified through the density of the corresponding rating graph and the continuity of the corresponding rating chain. In addition, we propose a general algorithm framework for improving recommender systems by exploiting these identified patterns. Finally, experimental results on a real-world data set show that the proposed rating graph based approach is effective for detecting user interest patterns, which in turn help to improve the performance of recommender systems.
Huanhuan Cao, Enhong Chen, Jie Yang 0004, Hui Xiong 0001
CIKM1
2009 Context-aware query classification
abstract
Understanding users'search intent expressed through their search queries is crucial to Web search and online advertisement. Web query classification (QC) has been widely studied for this purpose. Most previous QC algorithms classify individual queries without considering their context information. However, as exemplified by the well-known example on query "jaguar", many Web queries are short and ambiguous, whose real meanings are uncertain without the context information. In this paper, we incorporate context information into the problem of query classification by using conditional random field (CRF) models. In our approach, we use neighboring queries and their corresponding clicked URLs (Web pages) in search sessions as the context information. We perform extensive experiments on real world search logs and validate the effectiveness and effciency of our approach. We show that we can improve the F1 score by 52% as compared to other state-of-the-art baselines.
Huanhuan Cao, Derek Hao Hu, Dou Shen, Daxin Jiang, Jian-Tao Sun, Enhong Chen, Qiang Yang 0001
SIGIR1
2009 Towards context-aware search by learning a very large variable length hidden markov model from search logs
abstract
Capturing the context of a user's query from the previous queries and clicks in the same session may help understand the user's information need. A context-aware approach to document re-ranking, query suggestion, and URL recommendation may improve users' search experience substantially. In this paper, we propose a general approach to context-aware search. To capture contexts of queries, we learn a variable length Hidden Markov Model (vlHMM) from search sessions extracted from log data. Although the mathematical model is intuitive, how to learn a large vlHMM with millions of states from hundreds of millions of search sessions poses a grand challenge. We develop a strategy for parameter initialization in vlHMM learning which can greatly reduce the number of parameters to be estimated in practice. We also devise a method for distributed vlHMM learning under the map-reduce model. We test our approach on a real data set consisting of 1.8 billion queries, 2.6 billion clicks, and 840 million search sessions, and evaluate the effectiveness of the vlHMM learned from the real data on three search applications: document re-ranking, query suggestion, and URL recommendation. The experimental results show that our approach is both effective and efficient.
Huanhuan Cao, Daxin Jiang, Jian Pei 0001, Enhong Chen, Hang Li 0001
WWW1
2008 Context-aware query suggestion by mining click-through and session data
abstract
Query suggestion plays an important role in improving the usability of search engines. Although some recently proposed methods can make meaningful query suggestions by mining query patterns from search logs, none of them are context-aware - they do not take into account the immediately preceding queries as context in query suggestion. In this paper, we propose a novel context-aware query suggestion approach which is in two steps. In the offine model-learning step, to address data sparseness, queries are summarized into concepts by clustering a click-through bipartite. Then, from session data a concept sequence suffix tree is constructed as the query suggestion model. In the online query suggestion step, a user's search context is captured by mapping the query sequence submitted by the user to a sequence of concepts. By looking up the context in the concept sequence sufix tree, our approach suggests queries to the user in a context-aware manner. We test our approach on a large-scale search log of a commercial search engine containing 1:8 billion search queries, 2:6 billion clicks, and 840 million query sessions. The experimental results clearly show that our approach outperforms two baseline methods in both coverage and quality of suggestions.
Huanhuan Cao, Daxin Jiang, Jian Pei 0001, Qi He 0002, Zhen Liao, Enhong Chen, Hang Li 0001
KDD1
2008 Efficient strategies for tough aggregate constraint-based sequential pattern mining
Enhong Chen, Huanhuan Cao, Qing Li 0001, Tieyun Qian
Inf. Sci.2