EDBT 2026 Demo / reviewers in the wild / expert
Huanhuan Cao
dblp:48/5943
· DBLP profile ↗
34ranked-venue papers
7as first author
3since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 26 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 12 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Data mining · 41% Recommender systems · 32% Information retrieval · 18% | |
| Artificial intelligence
3 papers |
Face, body and person analysis · 49% Transfer learning and domain adaptation · 32% Representation and self-supervised learning · 19% |
Topics — the 29 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems › automated model design
embedding dimension search |
0.8 | 1 | 2024 | I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024 |
Data mining
feature engineering |
0.8 | 1 | 2024 | I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024 |
Data mining › dimensionality reduction
feature selection |
0.8 | 1 | 2024 | I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024 |
Computer vision › Face, body and person analysis
person re-identification |
0.5 | 1 | 2021 | Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.5 | 1 | 2021 | Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021 |
Computer vision › Face, body and person analysis › person re-identification › unsupervised person re-identification
unsupervised domain adaptive person re-identification |
0.5 | 1 | 2021 | Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-Identification · IEEE Trans. Image Process. 2021 |
Machine learning › Representation and self-supervised learning › representation learning
metric learning |
0.4 | 1 | 2019 | Zero-shot Metric Learning · IJCAI 2019 |
Recommender systems
click-through rate prediction |
0.2 | 1 | 2024 | I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender Systems · IEEE Trans. Knowl. Data Eng. 2024 |
Web and social media mining
app categorization |
0.2 | 1 | 2014 | Mobile App Classification with Enriched Contextual Information · IEEE Trans. Mob. Comput. 2014 |
Data mining › text mining
text classification |
0.2 | 1 | 2014 | Mobile App Classification with Enriched Contextual Information · IEEE Trans. Mob. Comput. 2014 |
Information retrieval
query suggestion |
0.2 | 2 | 2009 | Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009 Context-aware query suggestion by mining click-through and session data · KDD 2008 |
Machine learning › Transfer learning and domain adaptation › cross-domain transfer
cross-network transfer |
0.1 | 1 | 2012 | Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012 |
Data mining › pattern mining
behavioral pattern mining |
0.1 | 1 | 2012 | A habit mining approach for discovering similar mobile users · WWW 2012 |
Recommender systems
context-aware recommendation |
0.1 | 1 | 2012 | Mining Personal Context-Aware Preferences for Mobile Users · ICDM 2012 |
Recommender systems › cross-domain recommendation
cross-network recommendation |
0.1 | 1 | 2012 | Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012 |
Knowledge graphs
link prediction |
0.1 | 1 | 2012 | Link Prediction and Recommendation across Heterogeneous Social Networks · ICDM 2012 |
Data mining
pattern mining |
0.1 | 1 | 2012 | A habit mining approach for discovering similar mobile users · WWW 2012 |
Recommender systems › user modeling
user preference learning |
0.1 | 1 | 2012 | Mining Personal Context-Aware Preferences for Mobile Users · ICDM 2012 |
Recommender systems
user similarity |
0.1 | 1 | 2012 | A habit mining approach for discovering similar mobile users · WWW 2012 |
Multimedia analysis and retrieval › multimedia analysis › multimedia content description
multimedia annotation |
0.1 | 1 | 2012 | Towards Annotating Media Contents through Social Diffusion Analysis · ICDM 2012 |
Information retrieval
contextual search |
0.1 | 1 | 2009 | Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009 |
Information retrieval › reranking
document re-ranking |
0.1 | 1 | 2009 | Towards context-aware search by learning a very large variable length hidden markov model from search logs · WWW 2009 |
Information retrieval › query understanding
query classification |
0.1 | 1 | 2009 | Context-aware query classification · SIGIR 2009 |
Information retrieval
query understanding |
0.1 | 1 | 2009 | Context-aware query classification · SIGIR 2009 |
Information retrieval › query log analysis › clickthrough data
click-through data mining |
0.1 | 1 | 2008 | Context-aware query suggestion by mining click-through and session data · KDD 2008 |
Information retrieval › query suggestion
context-aware query suggestion |
0.1 | 1 | 2008 | Context-aware query suggestion by mining click-through and session data · KDD 2008 |
Information retrieval
query log analysis |
0.1 | 1 | 2008 | Context-aware query suggestion by mining click-through and session data · KDD 2008 |
Information retrieval › query understanding › query modeling
query context modeling |
0.0 | 1 | 2009 | Context-aware query classification · SIGIR 2009 |
Information retrieval › user behavior
search session analysis |
0.0 | 1 | 2009 | Context-aware query classification · SIGIR 2009 |
Methods — techniques the papers use, named apart from their topics
pruning · 0.8differentiable neural architecture search · 0.8view-invariant representation learning · 0.5pseudo-labeling · 0.5multi-scale representation · 0.5gradient reverse layer · 0.5convolutional neural network · 0.4web search enrichment · 0.4maximum entropy model · 0.4contextual feature extraction · 0.4transfer learning · 0.3ranking factor graph · 0.3optimization · 0.3probabilistic modeling · 0.1context log mining · 0.1constraint-based factorization · 0.1bayesian matrix factorization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | I-Razor: A Differentiable Neural Input Razor for Feature Selection and Dimension Search in DNN-Based Recommender SystemsabstractInput features play a crucial role in DNN-based recommender systems with thousands of categorical and continuous fields from users, items, contexts, and interactions. Noisy features and inappropriate embedding dimension assignments can deteriorate the performance of recommender systems and introduce unnecessary complexity in model training and online serving. Optimizing the input configuration of DNN models, including feature selection and embedding dimension assignment, has become one of the essential topics in feature engineering. However, in existing industrial practices, feature selection and dimension search are optimized sequentially, i.e., feature selection is performed first, followed by dimension search to determine the optimal dimension size for each selected feature. Such a sequential optimization mechanism increases training costs and risks generating suboptimal input configurations. To address this problem, we propose a differentiable neuralinputrazor(i-Razor) that enables joint optimization of feature selection and dimension search. Concretely, we introduce an end-to-end differentiable model to learn the relative importance of different embedding regions of each feature. Furthermore, a flexible pruning algorithm is proposed to achieve feature filtering and dimension derivation simultaneously. Extensive experiments on two large-scale public datasets in the Click-Through-Rate (CTR) prediction task demonstrate the efficacy and superiority of i-Razor in balancing model complexity and performance. Yao Yao 0006, Bin Liu 0072, Haoxun He, Dakui Sheng, Li Xiao 0006, Huanhuan Cao |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | How Online Reviews Interact with a Firm's Free Version Strategy
Huanhuan Cao, Jinhu Jiang, Xianjun Geng |
Inf. Manag. | 1 |
| 2021 | Self-Training With Progressive Representation Enhancement for Unsupervised Cross-Domain Person Re-IdentificationabstractIn recent years, person re-identification (re-ID) has achieved relatively good performance, benefiting from the revival of deep neural networks. However, due to the existence of domain bias which refers to the different data distributions between two domains, it remains challenging to directly deploy a model trained on a labeled source domain to a target domain only with unlabeled data available. In this paper, a Self-Training with Progressive Representation Enhancement (PREST) framework, which comprises a multi-scale self-training method and a view-invariant representation learning module, is proposed to promote re-ID performance on the target domain in an unsupervised manner. More specifically, multi-scale representations, including the global body and local parts of pedestrian images, are utilized to obtain pseudo-labels. Then, some images are selected according to the pseudo-labels to create a new dataset for supervising the fine-tuning process, which is operated iteratively to progressively promote the performance. Furthermore, to mitigate the influence of different styles among sub-domains, in cases where a single sub-domain is captured by one camera, a classifier with a gradient reverse layer is first employed to learn view-invariant representation for pedestrian images with the same identity taken by different cameras; this can further enhance the reliability of the predicted labels and improve the cross-domain re-ID performance. Extensive experiments on three large-scale re-ID datasets demonstrate that our framework achieves significantly better performance than existing approaches. Huanhuan Cao, Xu Yang 0019, Cheng Deng 0002, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2020 | Online review manipulation by asymmetrical firms: Is a firm's manipulation of online reviews always detrimental to its competitor?
Huanhuan Cao |
Inf. Manag. | 1 |
| 2020 | Object and background disentanglement for unsupervised cross-domain person re-identification
Yuehua Zhu, Cheng Deng 0002, Huanhuan Cao, Hao Wang 0062 |
Neurocomputing | 3 |
| 2019 | Zero-shot Metric LearningabstractIn this work, we tackle the zero-shot metric learning problem and propose a novel method abbreviated as ZSML, with the purpose to learn a distance metric that measures the similarity of unseen categories (even unseen datasets). ZSML achieves strong transferability by capturing multi-nonlinear yet continuous relation among data. It is motivated by two facts: 1) relations can be essentially described from various perspectives; and 2) traditional binary supervision is insufficient to represent continuous visual similarity. Specifically, we first reformulate a collection of specific-shaped convolutional kernels to combine data pairs and generate multiple relation vectors. Furthermore, we design a new cross-update regression loss to discover continuous similarity. Extensive experiments including intra-dataset transfer and inter-dataset transfer on four benchmark datasets demonstrate that ZSML can achieve state-of-the-art performance. Huanhuan Cao, Yanhua Yang, Erkun Yang, Cheng Deng 0002 |
IJCAI | 2 |
| 2018 | Question Headline Generation for News ArticlesabstractIn this paper, we introduce and tackle the Question Headline Generation (QHG) task. The motivation comes from the investigation of a real-world news portal where we find that news articles with question headlines often receive much higher click-through ratio than those with non-question headlines. The QHG task can be viewed as a specific form of the Question Generation (QG) task, with the emphasis on creating a natural question from a given news article by taking the entire article as the answer. A good QHG model thus should be able to generate a question by summarizing the essential topics of an article. Based on this idea, we propose a novel dual-attention sequence-to-sequence model (DASeq2Seq) for the QHG task. Unlike traditional sequence-to-sequence models which only employ the attention mechanism in the decoding phase for better generation, our DASeq2Seq further introduces a self-attention mechanism in the encoding phase to help generate a good summary of the article. We investigate two ways of the self-attention mechanism, namely global self-attention and distributed self-attention. Besides, we employ a vocabulary gate over both generic and question vocabularies to better capture the question patterns. Through the offline experiments, we show that our approach can significantly outperform the state-of-the-art question generation or headline generation models. Furthermore, we also conduct online evaluation to demonstrate the effectiveness of our approach using A/B test. Ruqing Zhang 0001, Jiafeng Guo, Yixing Fan, Yanyan Lan, Jun Xu 0001, Huanhuan Cao, Xueqi Cheng 0001 |
CIKM | 6 |
| 2014 | Learning to detect subway arrivals for passengers on a train
Kuifei Yu, Hengshu Zhu, Huanhuan Cao, Baoxian Zhang, Enhong Chen, Jilei Tian, Jinghai Rao |
Frontiers Comput. Sci. | 3 |
| 2014 | Mining Mobile User Preferences for Personalized Context-Aware RecommendationabstractRecent advances in mobile devices and their sensing capabilities have enabled the collection of rich contextual information and mobile device usage records through the device logs. These context-rich logs open a venue for mining the personal preferences of mobile users under varying contexts and thus enabling the development of personalized context-aware recommendation and other related services, such as mobile online advertising. In this article, we illustrate how to extract personal context-aware preferences from the context-rich device logs, or context logs for short, and exploit these identified preferences for building personalized context-aware recommender systems. A critical challenge along this line is that the context log of each individual user may not contain sufficient data for mining his or her context-aware preferences. Therefore, we propose to first learn common context-aware preferences from the context logs of many users. Then, the preference of each user can be represented as a distribution of these common context-aware preferences. Specifically, we develop two approaches for mining common context-aware preferences based on two different assumptions, namely, context-independent and context-dependent assumptions, which can fit into different application scenarios. Finally, extensive experiments on a real-world dataset show that both approaches are effective and outperform baselines with respect to mining personal context-aware preferences for mobile users. Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Kuifei Yu, Huanhuan Cao, Jilei Tian |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2014 | Mobile App Classification with Enriched Contextual InformationabstractThe study of the use of mobile Apps plays an important role in understanding the user preferences, and thus provides the opportunities for intelligent personalized context-based services. A key step for the mobile App usage analysis is to classify Apps into some predefined categories. However, it is a nontrivial task to effectively classify mobile Apps due to the limited contextual information available for the analysis. For instance, there is limited contextual information about mobile Apps in their names. However, this contextual information is usually incomplete and ambiguous. To this end, in this paper, we propose an approach for first enriching the contextual information of mobile Apps by exploiting the additional Web knowledge from the Web search engine. Then, inspired by the observation that different types of mobile Apps may be relevant to different real-world contexts, we also extract some contextual features for mobile Apps from the context-rich device logs of mobile users. Finally, we combine all the enriched contextual information into the Maximum Entropy model for training a mobile App classifier. To validate the proposed method, we conduct extensive experiments on 443 mobile users' device logs to show both the effectiveness and efficiency of the proposed approach. The experimental results clearly show that our approach outperforms two state-of-the-art benchmark methods with a significant margin. Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Huanhuan Cao, Jilei Tian |
IEEE Trans. Mob. Comput. | 4 |
| 2014 | Ranking user authority with relevant knowledge categories for expert finding
Hengshu Zhu, Enhong Chen, Hui Xiong 0001, Huanhuan Cao, Jilei Tian |
World Wide Web | 4 |
| 2013 | Learning to Detect the Subway Station Arrival for Mobile Users
Kuifei Yu, Hengshu Zhu, Huanhuan Cao, Baoxian Zhang, Enhong Chen, Jilei Tian, Jinghai Rao |
IDEAL | 3 |
| 2013 | Predicting Stay Time of Mobile Users With Contextual InformationabstractMobile service providers and manufacturers continue to provide services and devices that take advantage of the location information associated with devices to provide a more personalized experience for users. For many such services, the user experience can be dramatically improved if a mobile device can predict how long a mobile user will stay at the current location. In this paper, we propose to take advantage of contextual information for predicting the stay time of mobile users. Specially, we investigate two strategies for modeling the relevance between it and contextual information, i.e., Stay Status Prediction (SSP) and Stay Time Prediction (STP). SSP is to predict whether a mobile user will stay at the current location at time point ti+naccording to the contextual information at ti, while STP is to directly predict how long a mobile user will stay at the current location. Moreover, we study several typical machine learning models which can be extended for implementing SSP and STP and evaluate their performance with respect to prediction accuracy. We also conduct extensive experiments on real data sets to evaluate several implementations of the proposed strategies in terms of both effectiveness and efficiency for STP. Huanhuan Cao, Lei Li 0009, MengChu Zhou |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2013 | A vlHMM approach to context-aware searchabstractCapturing the context of a user's query from the previous queries and clicks in the same session leads to a better understanding of the user's information need. A context-aware approach to document reranking, URL recommendation, and query suggestion may substantially improve users' search experience. In this article, we propose a general approach to context-aware search by learning avariable length hidden Markov model(vlHMM) from search sessions extracted from log data. While the mathematical model is powerful, the huge amounts of log data present great challenges. We develop several distributed learning techniques to learn a very large vlHMM under themap-reduceframework. Moreover, we construct feature vectors for each state of the vlHMM model to handle users' novel queries not covered by the training data. We test our approach on a raw dataset consisting of 1.9 billion queries, 2.9 billion clicks, and 1.2 billion search sessions before filtering, and evaluate the effectiveness of the vlHMM learned from the real data on three search applications: document reranking, query suggestion, and URL recommendation. The experiment results validate the effectiveness of vlHMM in the applications of document reranking, URL recommendation, and query suggestion. Zhen Liao, Daxin Jiang, Jian Pei 0001, Yalou Huang, Enhong Chen, Huanhuan Cao, Hang Li 0001 |
ACM Trans. Web | 6 |
| 2012 | Exploiting enriched contextual information for mobile app classificationabstractA key step for the mobile app usage analysis is to classify apps into some predefined categories. However, it is a nontrivial task to effectively classify mobile apps due to the limited contextual information available for the analysis. To this end, in this paper, we propose an approach to first enrich the contextual information of mobile apps by exploiting the additional Web knowledge from the Web search engine. Then, inspired by the observation that different types of mobile apps may be relevant to different real-world contexts, we also extract some contextual features for mobile apps from the context-rich device logs of mobile users. Finally, we combine all the enriched contextual information into a Maximum Entropy model for training a mobile app classifier. The experimental results based on 443 mobile users' device logs clearly show that our approach outperforms two state-of-the-art benchmark methods with a significant margin. Hengshu Zhu, Huanhuan Cao, Enhong Chen, Hui Xiong 0001, Jilei Tian |
CIKM | 2 |
| 2012 | Link Prediction and Recommendation across Heterogeneous Social NetworksabstractLink prediction and recommendation is a fundamental problem in social network analysis. The key challenge of link prediction comes from the sparsity of networks due to the strong disproportion of links that they have potential to form to links that do form. Most previous work tries to solve the problem in single network, few research focus on capturing the general principles of link formation across heterogeneous networks. In this work, we give a formal definition of link recommendation across heterogeneous networks. Then we propose a ranking factor graph model (RFG) for predicting links in social networks, which effectively improves the predictive performance. Motivated by the intuition that people make friends in different networks with similar principles, we find several social patterns that are general across heterogeneous networks. With the general social patterns, we develop a transfer-based RFG model that combines them with network structure information. This model provides us insight into fundamental principles that drive the link formation and network evolution. Finally, we verify the predictive performance of the presented transfer model on 12 pairs of transfer cases. Our experimental results demonstrate that the transfer of general social patterns indeed help the prediction of links. Yuxiao Dong, Jie Tang 0001, Sen Wu 0001, Jilei Tian, Nitesh V. Chawla, Jinghai Rao, Huanhuan Cao |
ICDM | 7 |
| 2012 | Towards Annotating Media Contents through Social Diffusion AnalysisabstractRecently, the boom of media contents on the Internet raises challenges in managing them effectively and thus requires automatic media annotation techniques. Motivated by the observation that media contents are usually shared frequently in online communities and thus have a lot of social diffusion records, we propose a novel media annotating approach depending on these social diffusion records instead of metadata. The basic assumption is that the social diffusion records reflect the common interests (CI) between users, which can be analyzed for generating annotations. With this assumption, we present a novel CI-based social diffusion model and translate the automatic annotating task into the CI-based diffusion maximization (CIDM) problem. Moreover, we propose to solve the CIDM problem through two optimization tasks, corresponding to the training and test stages in supervised learning. Extensive experiments on real-world data sets show that our approach can effectively generate high quality annotations, and thus demonstrate the capability of social diffusion analysis in annotating media. Tong Xu 0001, Dong Liu 0002, Enhong Chen, Huanhuan Cao, Jilei Tian |
ICDM | 4 |
| 2012 | Mining Personal Context-Aware Preferences for Mobile UsersabstractIn this paper, we illustrate how to extract personal context-aware preferences from the context-rich device logs (i.e., context logs) for building novel personalized context-aware recommender systems. A critical challenge along this line is that the context log of each individual user may not contain sufficient data for mining his/her context-aware preferences. Therefore, we propose to first learn common context-aware preferences from the context logs of many users. Then, the preference of each user can be represented as a distribution of these common context-aware preferences. Specifically, we develop two approaches for mining common context-aware preferences based on two different assumptions, namely, context independent and context dependent assumptions, which can fit into different application scenarios. Finally, extensive experiments on a real-world data set show that both approaches are effective and outperform baselines with respect to mining personal context-aware preferences for mobile users. Hengshu Zhu, Enhong Chen, Kuifei Yu, Huanhuan Cao, Hui Xiong 0001, Jilei Tian |
ICDM | 4 |
| 2012 | Mining Significant Places from Cell ID Trajectories: A Geo-grid Based ApproachabstractMining the frequently visited places of single mobile users, i.e., significant places, is crucial for supporting personalized location-based services. Most of existing works for significance place mining have a need to take advantage the GPS trajectories of users. However, it is difficult to encourage mobile users to contribute GPS trajectories because of the high power consumption of GPS. In this paper, we propose a geo-grid based approach for mining significant places from cell ID trajectories. In our approach, the mined significant places are represented as sets of geo-grids which are much smaller than the coverage areas of cell-sites. To be specific, we firstly extract the stay areas where the mobile user used to stay and map them to many geo-grids. Then we mine significant places from the geo-grids by considering their significance. We evaluate the approach on real word data sets and the experimental results clearly show that the proposed approach outperforms two baselines. Tengfei Bao, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian |
MDM | 2 |
| 2012 | A Demonstration of Mining Significant Places from Cell ID Trajectories through a Geo-grid Based ApproachabstractMining the frequently visited places of single mobile users, i.e., significant places, is crucial for supporting personalized location-based services. Most of existing works for significance place mining have a need to take advantage the GPS trajectories of users. However, it is difficult to encourage mobile users to contribute GPS trajectories because of the high power consumption of GPS. In this demonstration, we propose a geo-grid based approach for mining significant places from cell ID trajectories. In our approach, the mined significant places are represented as sets of geo-grids which are much smaller than the coverage areas of cell-sites. To be specific, we firstly extract the stay areas where the mobile user used to stay and map them to many geogrids. Then we mine significant places from the geo-grids by considering their significance. Tengfei Bao, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian |
MDM | 2 |
| 2012 | BP-growth: Searching Strategies for Efficient Behavior Pattern MiningabstractUser habit mining plays an important role in user understanding, which is critical for improving a wide range of personalized intelligence services. Recently, some researchers proposed to mine user behavior patterns which characterize the habits of mobile users and account for the associations between user interactions and context captured by mobile devices. However, the existing approaches for mining these behavior patterns are not practical in mobile environments due to limited computing resources on mobile devices. To fulfill this crucial void, we investigate optimizing strategies which can be used for improving the efficiency of behavior pattern mining in terms of computing and memory needs. Specifically, we examine typical optimizing strategies for association rule mining and study the feasibility of applying them to behavior pattern mining, since these two problems are similar in many aspects. Moreover, we develop an efficient algorithm, named BP-Growth, for behavior pattern mining by combining two promising strategies. Finally, experimental results show that BP-Growth outperforms benchmark methods with a significant margin in terms of both computing and memory cost. Xueying Li 0004, Huanhuan Cao, Enhong Chen, Hui Xiong 0001, Jilei Tian |
MDM | 2 |
| 2012 | Towards Personalized Context-Aware Recommendation by Mining Context Logs through Topic Models
Kuifei Yu, Baoxian Zhang, Hengshu Zhu, Huanhuan Cao, Jilei Tian |
PAKDD (1) | 4 |
| 2012 | A habit mining approach for discovering similar mobile usersabstractDiscovering similar users with respect to their habits plays an important role in a wide range of applications, such as collaborative filtering for recommendation, user segmentation for market analysis, etc. Recently, the progressing ability to sense user contexts of smart mobile devices makes it possible to discover mobile users with similar habits by mining their habits from their mobile devices. However, though some researchers have proposed effective methods for mining user habits such as behavior pattern mining, how to leverage the mined results for discovering similar users remains less explored. To this end, we propose a novel approach for conquering the sparseness of behavior pattern space and thus make it possible to discover similar mobile users with respect to their habits by leveraging behavior pattern mining. To be specific, first, we normalize the raw context log of each user by transforming the location-based context data and user interaction records to more general representations. Second, we take advantage of a constraint-based Bayesian Matrix Factorization model for extracting the latent common habits among behavior patterns and then transforming behavior pattern vectors to the vectors of mined common habits which are in a much more dense space. The experiments conducted on real data sets show that our approach outperforms three baselines in terms of the effectiveness of discovering similar mobile users with respect to their habits. Haiping Ma, Huanhuan Cao, Qiang Yang 0001, Enhong Chen, Jilei Tian |
WWW | 2 |
| 2012 | An unsupervised approach to modeling personalized contexts of mobile users
Tengfei Bao, Huanhuan Cao, Enhong Chen, Jilei Tian, Hui Xiong 0001 |
Knowl. Inf. Syst. | 2 |
| 2012 | Learning to Infer the Status of Heavy-Duty Sensors for Energy-Efficient Context-SensingabstractWith the prevalence of smart mobile devices with multiple sensors, the commercial application of intelligent context-aware services becomes more and more attractive. However, limited by the battery capacity, the energy efficiency of context-sensing is the bottleneck for the success of context-aware applications. Though several previous studies for energy-efficient context-sensing have been reported, none of them can be applied to multiple types of high-energy-consuming sensors. Moreover, applying machine learning technologies to energy-efficient context-sensing is underexplored too. In this article, we propose to leverage machine learning technologies for improving the energy efficiency of multiple high-energy-consuming context sensors by trading off the sensing accuracy. To be specific, we try to infer the status of high-energy-consuming sensors according to the outputs of software-based sensors and the physical sensors that are necessary to work all the time for supporting the basic functions of mobile devices. If the inference indicates the high-energy-consuming sensor is in a stable status, we avoid the unnecessary invocation and instead use the latest invoked value as the estimation. The experimental results on real datasets show that the energy efficiency of GPS sensing and audio-level sensing are significantly improved by the proposed approach while the sensing accuracy is over 90%. Xueying Li 0004, Huanhuan Cao, Enhong Chen, Jilei Tian |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Towards expert finding by leveraging relevant categories in authority rankingabstractHow to improve authority ranking is a crucial research problem for expert finding. In this paper, we propose a novel framework for expert finding based on the authority information in the target category as well as the relevant categories. First, we develop a scalable method for measuring the relevancy between categories through topic models. Then, we provide a link analysis approach for ranking user authority by considering the information in both the target category and the relevant categories. Finally, the extensive experiments on two large-scale real-world Q&A data sets clearly show that the proposed method outperforms the baseline methods with a significant margin. Hengshu Zhu, Huanhuan Cao, Hui Xiong 0001, Enhong Chen, Jilei Tian |
CIKM | 2 |
| 2011 | Finding Experts in Tag Based Knowledge Sharing Communities
Hengshu Zhu, Enhong Chen, Huanhuan Cao |
KSEM | 3 |
| 2011 | Mining Concept Sequences from Large-Scale Search Logs for Context-Aware Query SuggestionabstractQuery suggestion plays an important role in improving usability of search engines. Although some recently proposed methods provide query suggestions by mining query patterns from search logs, none of them models the immediately preceding queries as context systematically, and uses context information effectively in query suggestions. Context-aware query suggestion is challenging in both modeling context and scaling up query suggestion using context. In this article, we propose a novel context-aware query suggestion approach. To tackle the challenges, our approach consists of two stages. In the first, offline model-learning stage , to address data sparseness, queries are summarized into concepts by clustering a click-through bipartite. A concept sequence suffix tree is then constructed from session data as a context-aware query suggestion model. In the second, online query suggestion stage , a user’s search context is captured by mapping the query sequence submitted by the user to a sequence of concepts. By looking up the context in the concept sequence suffix tree, we suggest to the user context-aware queries. We test our approach on large-scale search logs of a commercial search engine containing 4.0 billion Web queries, 5.9 billion clicks, and 1.87 billion search sessions. The experimental results clearly show that our approach outperforms three baseline methods in both coverage and quality of suggestions. Zhen Liao, Daxin Jiang, Enhong Chen, Jian Pei 0001, Huanhuan Cao, Hang Li 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2010 | An effective approach for mining mobile user habitsabstractThe user interaction with the mobile device plays an important role in user habit understanding. In this paper, we propose to mine the associations between user interactions and contexts captured by mobile devices, or behavior patterns for short, from context logs to characterize the habits of mobile users. The extensive experiments on the collected real life data clearly validate the ability of our approach for mining effective behavior patterns. Huanhuan Cao, Tengfei Bao, Qiang Yang 0001, Enhong Chen, Jilei Tian |
CIKM | 1 |
| 2009 | Enhancing recommender systems under volatile userinterest driftsabstractThis paper presents a systematic study of how to enhance recommender systems under volatile user interest drifts. A key development challenge along this line is how to track user interests dynamically. To this end, we first define four types of interest patterns to understand users' rating behaviors and analyze the properties of these patterns. We also propose a rating graph and rating chain based approach for detecting these interest patterns. For each users' rating series, a rating graph and a rating chain are constructed based on the similarities between rated items. The type of a given user's interest pattern is identified through the density of the corresponding rating graph and the continuity of the corresponding rating chain. In addition, we propose a general algorithm framework for improving recommender systems by exploiting these identified patterns. Finally, experimental results on a real-world data set show that the proposed rating graph based approach is effective for detecting user interest patterns, which in turn help to improve the performance of recommender systems. Huanhuan Cao, Enhong Chen, Jie Yang 0004, Hui Xiong 0001 |
CIKM | 1 |
| 2009 | Context-aware query classificationabstractUnderstanding users'search intent expressed through their search queries is crucial to Web search and online advertisement. Web query classification (QC) has been widely studied for this purpose. Most previous QC algorithms classify individual queries without considering their context information. However, as exemplified by the well-known example on query "jaguar", many Web queries are short and ambiguous, whose real meanings are uncertain without the context information. In this paper, we incorporate context information into the problem of query classification by using conditional random field (CRF) models. In our approach, we use neighboring queries and their corresponding clicked URLs (Web pages) in search sessions as the context information. We perform extensive experiments on real world search logs and validate the effectiveness and effciency of our approach. We show that we can improve the F1 score by 52% as compared to other state-of-the-art baselines. Huanhuan Cao, Derek Hao Hu, Dou Shen, Daxin Jiang, Jian-Tao Sun, Enhong Chen, Qiang Yang 0001 |
SIGIR | 1 |
| 2009 | Towards context-aware search by learning a very large variable length hidden markov model from search logsabstractCapturing the context of a user's query from the previous queries and clicks in the same session may help understand the user's information need. A context-aware approach to document re-ranking, query suggestion, and URL recommendation may improve users' search experience substantially. In this paper, we propose a general approach to context-aware search. To capture contexts of queries, we learn a variable length Hidden Markov Model (vlHMM) from search sessions extracted from log data. Although the mathematical model is intuitive, how to learn a large vlHMM with millions of states from hundreds of millions of search sessions poses a grand challenge. We develop a strategy for parameter initialization in vlHMM learning which can greatly reduce the number of parameters to be estimated in practice. We also devise a method for distributed vlHMM learning under the map-reduce model. We test our approach on a real data set consisting of 1.8 billion queries, 2.6 billion clicks, and 840 million search sessions, and evaluate the effectiveness of the vlHMM learned from the real data on three search applications: document re-ranking, query suggestion, and URL recommendation. The experimental results show that our approach is both effective and efficient. Huanhuan Cao, Daxin Jiang, Jian Pei 0001, Enhong Chen, Hang Li 0001 |
WWW | 1 |
| 2008 | Context-aware query suggestion by mining click-through and session dataabstractQuery suggestion plays an important role in improving the usability of search engines. Although some recently proposed methods can make meaningful query suggestions by mining query patterns from search logs, none of them are context-aware - they do not take into account the immediately preceding queries as context in query suggestion. In this paper, we propose a novel context-aware query suggestion approach which is in two steps. In the offine model-learning step, to address data sparseness, queries are summarized into concepts by clustering a click-through bipartite. Then, from session data a concept sequence suffix tree is constructed as the query suggestion model. In the online query suggestion step, a user's search context is captured by mapping the query sequence submitted by the user to a sequence of concepts. By looking up the context in the concept sequence sufix tree, our approach suggests queries to the user in a context-aware manner. We test our approach on a large-scale search log of a commercial search engine containing 1:8 billion search queries, 2:6 billion clicks, and 840 million query sessions. The experimental results clearly show that our approach outperforms two baseline methods in both coverage and quality of suggestions. Huanhuan Cao, Daxin Jiang, Jian Pei 0001, Qi He 0002, Zhen Liao, Enhong Chen, Hang Li 0001 |
KDD | 1 |
| 2008 | Efficient strategies for tough aggregate constraint-based sequential pattern mining
Enhong Chen, Huanhuan Cao, Qing Li 0001, Tieyun Qian |
Inf. Sci. | 2 |