VLDB 2026 Research / reviewers in the wild / expert
Hao Ma 0001
dblp:86/4227-1
· DBLP profile ↗
43ranked-venue papers
17as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 26 · 15 first-authorArtificial intelligence and machine learning · 22 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorSoftware engineering, systems software and programming languages · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Language models and text generation · 26% Question answering and dialogue systems · 16% Representation and self-supervised learning · 15% | |
| Databases, data mining, and information retrieval
22 papers |
Recommender systems · 52% Information retrieval · 27% Web and social media mining · 16% | |
| Software engineering, system software, and programming languages
3 papers |
Services computing and microservices · 82% Software testing · 18% |
Topics — the 30 heaviest of 73, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
masked language modeling |
1.0 | 2 | 2021 | Luna: Linear Unified Nested Attention · NeurIPS 2021 On the Influence of Masking Policies in Intermediate Pre-training · EMNLP (1) 2021 |
Machine learning › Representation and self-supervised learning
pre-training |
0.9 | 2 | 2021 | Luna: Linear Unified Nested Attention · NeurIPS 2021 To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich Tasks · ACL 2020 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization · ICML 2025 |
Recommender systems
social recommendation |
0.8 | 6 | 2014 | On measuring social friend interest similarities in recommender systems · SIGIR 2014 An experimental study on implicit social recommendation · SIGIR 2013 Improving Recommender Systems by Incorporating Social Contextual Information · ACM Trans. Inf. Syst. 2011 |
Machine learning › Graph learning
network embedding |
0.7 | 2 | 2019 | NetSMF: Large-Scale Network Embedding as Sparse Matrix Factorization · WWW 2019 Network Embedding as Matrix Factorization: Unifying DeepWalk, LINE, PTE, and node2vec · WSDM 2018 |
Recommender systems
collaborative filtering |
0.6 | 7 | 2014 | QoS-Aware Web Service Recommendation by Collaborative Filtering · IEEE Trans. Serv. Comput. 2011 Improving Recommender Systems by Incorporating Social Contextual Information · ACM Trans. Inf. Syst. 2011 Recommender systems with social regularization · WSDM 2011 |
Machine learning › Transfer learning and domain adaptation
dual learning |
0.6 | 1 | 2022 | Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word Generation · EMNLP 2022 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.6 | 1 | 2022 | Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word Generation · EMNLP 2022 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.5 | 1 | 2021 | Luna: Linear Unified Nested Attention · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.5 | 1 | 2021 | Luna: Linear Unified Nested Attention · NeurIPS 2021 |
Natural language and speech › Language models and text generation › pre-trained language model
intermediate pre-training |
0.5 | 1 | 2021 | On the Influence of Masking Policies in Intermediate Pre-training · EMNLP (1) 2021 |
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention |
0.5 | 1 | 2021 | Luna: Linear Unified Nested Attention · NeurIPS 2021 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.5 | 2 | 2016 | Question Answering with Knowledge Base, Web and Beyond · SIGIR 2016 Open Domain Question Answering via Semantic Enrichment · WWW 2015 |
Recommender systems › knowledge-aware recommendation
entity recommendation |
0.4 | 2 | 2015 | Learning to Recommend Related Entities to Search Users · WSDM 2015 On building entity recommender systems using user click log and freebase knowledge · WSDM 2014 |
Machine learning › Representation and self-supervised learning
matrix factorization |
0.4 | 1 | 2019 | NetSMF: Large-Scale Network Embedding as Sparse Matrix Factorization · WWW 2019 |
Machine learning › Graph learning › network embedding
scalable network embedding |
0.4 | 1 | 2019 | NetSMF: Large-Scale Network Embedding as Sparse Matrix Factorization · WWW 2019 |
Web and social media mining
social network analysis |
0.4 | 2 | 2018 | DeepInf: Social Influence Prediction with Deep Learning · KDD 2018 Introduction to social recommendation · WWW 2010 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2018 | DeepInf: Social Influence Prediction with Deep Learning · KDD 2018 |
Machine learning › Representation and self-supervised learning › word representation › word embedding
skip-gram negative sampling |
0.3 | 1 | 2018 | Network Embedding as Matrix Factorization: Unifying DeepWalk, LINE, PTE, and node2vec · WSDM 2018 |
Web and social media mining › social influence analysis
social influence prediction |
0.3 | 1 | 2018 | DeepInf: Social Influence Prediction with Deep Learning · KDD 2018 |
Information retrieval
ranking |
0.3 | 2 | 2016 | User Fatigue in Online News Recommendation · WWW 2016 Exploring and exploiting user search behavior on mobile and tablet devices to improve search relevance · WWW 2013 |
Natural language and speech › Question answering and dialogue systems
open-domain question answering |
0.3 | 2 | 2016 | Open Domain Question Answering via Semantic Enrichment · WWW 2015 Table Cell Search for Question Answering · WWW 2016 |
Computational social science and digital humanities
science of science |
0.3 | 1 | 2017 | A Century of Science: Globalization of Scientific Collaborations, Citations, and Innovations · KDD 2017 |
Recommender systems › collaborative filtering
matrix factorization |
0.3 | 3 | 2011 | Improving Recommender Systems by Incorporating Social Contextual Information · ACM Trans. Inf. Syst. 2011 Recommender systems with social regularization · WSDM 2011 Learning to recommend with social trust ensemble · SIGIR 2009 |
Natural language and speech › Question answering and dialogue systems
table question answering |
0.2 | 1 | 2016 | Table Cell Search for Question Answering · WWW 2016 |
Natural language and speech › Question answering and dialogue systems › open-domain question answering
web question answering |
0.2 | 1 | 2016 | Question Answering with Knowledge Base, Web and Beyond · SIGIR 2016 |
Recommender systems
news recommendation |
0.2 | 1 | 2016 | User Fatigue in Online News Recommendation · WWW 2016 |
Information retrieval › search engines › structured data search
table retrieval |
0.2 | 1 | 2016 | Table Cell Search for Question Answering · WWW 2016 |
Web and social media mining › social network analysis
social network |
0.2 | 1 | 2014 | On measuring social friend interest similarities in recommender systems · SIGIR 2014 |
Machine learning › Transfer learning and domain adaptation
parameter-efficient transfer learning |
0.2 | 1 | 2022 | UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning · ACL (1) 2022 |
Methods — techniques the papers use, named apart from their topics
utility maximization · 0.9self-consistency · 0.9primal-dual framework · 0.6prefix-tuning · 0.6knowledge distillation · 0.6adapter modules · 0.6LoRA · 0.6softmax attention approximation · 0.5nested attention · 0.5meta-learning · 0.5matrix factorization · 0.5spectral analysis · 0.3skip-gram · 0.3negative sampling · 0.3deep learning · 0.3collaborative filtering · 0.3search engine snippets · 0.2feature engineering · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Think Smarter not Harder: Adaptive Reasoning with Inference Aware OptimizationabstractSolving mathematics problems has been an intriguing capability of large language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through self-correction and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoning chain models exhibit an undesired single-modal behavior, where trivial questions require unnecessarily tedious long chains of thought. In this work, we propose a way to allow models to be aware of inference budgets by formulating it as utility maximization with respect to an inference budget constraint, hence naming our algorithm Inference Budget-Constrained Policy Optimization (IBPO). In a nutshell, models fine-tuned through IBPO learn to ``understand'' the difficulty of queries and allocate inference budgets to harder ones. With different inference budgets, our best models are able to have a $4.14$\% and $5.74$\% absolute improvement ($8.08$\% and $11.2$\% relative improvement) on MATH500 using $2.16$x and $4.32$x inference budgets respectively, relative to LLaMA3.1 8B Instruct. These improvements are approximately $2$x those of self-consistency under the same budgets. Zishun Yu, Tengyu Xu, Di Jin 0005, Karthik Abinav Sankararaman, Zhouhao Zeng, Eryk Helenowski, Sinong Wang, Hao Ma 0001 |
ICML | 11 |
| 2024 | Effective Long-Context Scaling of Foundation ModelsabstractWenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wenhan Xiong, Igor Molybog, Prajjwal Bhargava, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma 0001 |
NAACL-HLT | 21 |
| 2022 | UniPELT: A Unified Framework for Parameter-Efficient Language Model TuningabstractYuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Scott Yih, Madian Khabsa. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yuning Mao, Lambert Mathias, Amjad Almahairi, Hao Ma 0001, Jiawei Han 0001, Scott Yih, Madian Khabsa |
ACL (1) | 5 |
| 2022 | Learning to Generate Question by Asking Question: A Primal-Dual Approach with Uncommon Word GenerationabstractAutomatic question generation (AQG) is the task of generating a question from a given passage and an answer.Most existing AQG methods aim at encoding the passage and the answer to generate the question.However, limited work has focused on modeling the correlation between the target answer and the generated question.Moreover, unseen or rare word generation has not been studied in previous works.In this paper, we propose a novel approach which incorporates question generation with its dual problem, question answering, into a unified primal-dual framework.Specifically, the question generation component consists of an encoder that jointly encodes the answer with the passage, and a decoder that produces the question.The question answering component then re-asks the generated question on the passage to ensure that the target answer is obtained.We further introduce a knowledge distillation module to improve the model generalization ability.We conduct an extensive set of experiments on SQuAD and HotpotQA benchmarks.Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods. Qifan Wang 0001, Xiaojun Quan, Fuli Feng, Dongfang Liu, Zenglin Xu, Sinong Wang, Hao Ma 0001 |
EMNLP | 8 |
| 2022 | IDPG: An Instance-Dependent Prompt Generation MethodabstractZhuofeng Wu, Sinong Wang, Jiatao Gu, Rui Hou, Yuxiao Dong, V.G.Vinod Vydiswaran, Hao Ma. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Zhuofeng Wu 0001, Sinong Wang, Jiatao Gu, Yuxiao Dong, V. G. Vinod Vydiswaran, Hao Ma 0001 |
NAACL-HLT | 7 |
| 2021 | On the Influence of Masking Policies in Intermediate Pre-trainingabstractCurrent NLP models are predominantly trained through a two-stage "pre-train then fine-tune" pipeline.Prior work has shown that inserting an intermediate pre-training stage, using heuristic masking policies for masked language modeling (MLM), can significantly improve final performance.However, it is still unclear (1) in what cases such intermediate pre-training is helpful, (2) whether hand-crafted heuristic objectives are optimal for a given task, and (3) whether a masking policy designed for one task is generalizable beyond that task.In this paper, we perform a large-scale empirical study to investigate the effect of various masking policies in intermediate pre-training with nine selected tasks across three categories.Crucially, we introduce methods to automate the discovery of optimal masking policies via direct supervision or meta-learning.We conclude that the success of intermediate pre-training is dependent on appropriate pre-train corpus, selection of output format (i.e., masked spans or full sentence), and clear understanding of the role that MLM plays for the downstream task.In addition, we find our learned masking policies outperform the heuristic of masking named entities on TriviaQA, and policies learned from one task can positively transfer to other tasks in certain cases, inviting future research in this direction. Qinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte, Hao Ma 0001, Scott Yih, Xiang Ren 0001, Madian Khabsa |
EMNLP (1) | 5 |
| 2021 | On Unifying Misinformation DetectionabstractNayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa |
NAACL-HLT | 5 |
| 2021 | Luna: Linear Unified Nested AttentionabstractThe quadratic computational and memory complexities of the Transformer's attention mechanism have limited its scalability for modeling long sequences. In this paper, we propose Luna, a linear unified nested attention mechanism that approximates softmax attention with two nested linear attention functions, yielding only linear (as opposed to quadratic) time and space complexity. Specifically, with the first attention function, Luna packs the input sequence into a sequence of fixed length. Then, the packed sequence is unpacked using the second attention function. As compared to a more traditional attention mechanism, Luna introduces an additional sequence with a fixed length as input and an additional corresponding output, which allows Luna to perform attention operation linearly, while also storing adequate contextual information. We perform extensive evaluations on three benchmarks of sequence modeling tasks: long-context sequence modelling, neural machine translation and masked language modeling for large-scale pretraining. Competitive or even better experimental results demonstrate both the effectiveness and efficiency of Luna compared to a variety of strong baseline methods including the full-rank attention and other efficient sparse and dense attention methods. Xuezhe Ma, Xiang Kong, Sinong Wang, Chunting Zhou, Jonathan May, Hao Ma 0001, Luke Zettlemoyer |
NeurIPS | 6 |
| 2020 | To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich TasksabstractPretraining NLP models with variants of Masked Language Model (MLM) objectives has recently led to a significant improvements on many tasks.This paper examines the benefits of pretrained models as a function of the number of training samples used in the downstream task.On several text classification tasks, we show that as the number of training examples grow into the millions, the accuracy gap between finetuning BERT-based model and training vanilla LSTM from scratch narrows to within 1%.Our findings indicate that MLM-based models might reach a diminishing return point as the supervised data size increases significantly. Sinong Wang, Madian Khabsa, Hao Ma 0001 |
ACL | 3 |
| 2019 | NetSMF: Large-Scale Network Embedding as Sparse Matrix FactorizationabstractWe study the problem of large-scale network embedding, which aims to learn latent representations for network mining applications. Previous research shows that 1) popular network embedding benchmarks, such as DeepWalk, are in essence implicitly factorizing a matrix with a closed form, and 2) the explicit factorization of such matrix generates more powerful embeddings than existing methods. However, directly constructing and factorizing this matrix-which is dense-is prohibitively expensive in terms of both time and space, making it not scalable for large networks. Jiezhong Qiu, Yuxiao Dong, Hao Ma 0001, Jian Li 0015, Chi Wang 0001, Kuansan Wang, Jie Tang 0001 |
WWW | 3 |
| 2018 | DeepInf: Social Influence Prediction with Deep LearningabstractSocial and information networking activities such as on Facebook, Twitter, WeChat, and Weibo have become an indispensable part of our everyday life, where we can easily access friends' behaviors and are in turn influenced by them. Consequently, an effective social influence prediction for each user is critical for a variety of applications such as online recommendation and advertising. Jiezhong Qiu, Jian Tang 0005, Hao Ma 0001, Yuxiao Dong, Kuansan Wang, Jie Tang 0001 |
KDD | 3 |
| 2018 | GaAN: Gated Attention Networks for Learning on Large and Spatiotemporal Graphs
Jiani Zhang 0001, Xingjian Shi, Junyuan Xie, Hao Ma 0001, Irwin King, Dit-Yan Yeung |
UAI | 4 |
| 2018 | Network Embedding as Matrix Factorization: Unifying DeepWalk, LINE, PTE, and node2vecabstractSince the invention of word2vec, the skip-gram model has significantly advanced the research of network embedding, such as the recent emergence of the DeepWalk, LINE, PTE, and node2vec approaches. In this work, we show that all of the aforementioned models with negative sampling can be unified into the matrix factorization framework with closed forms. Our analysis and proofs reveal that: (1) DeepWalk empirically produces a low-rank transformation of a network's normalized Laplacian matrix; (2) LINE, in theory, is a special case of DeepWalk when the size of vertices' context is set to one; (3) As an extension of LINE, PTE can be viewed as the joint factorization of multiple networks» Laplacians; (4) node2vec is factorizing a matrix related to the stationary distribution and transition probability tensor of a 2nd-order random walk. We further provide the theoretical connections between skip-gram based network embedding algorithms and the theory of graph Laplacian. Finally, we present the NetMF method as well as its approximation algorithm for computing network embedding. Our method offers significant improvements over DeepWalk and LINE for conventional network mining tasks. This work lays the theoretical foundation for skip-gram based network embedding methods, leading to a better understanding of latent network representation learning. Jiezhong Qiu, Yuxiao Dong, Hao Ma 0001, Jian Li 0015, Kuansan Wang, Jie Tang 0001 |
WSDM | 3 |
| 2017 | A Century of Science: Globalization of Scientific Collaborations, Citations, and InnovationsabstractProgress in science has advanced the development of human society across history, with dramatic revolutions shaped by information theory, genetic cloning, and artificial intelligence, among the many scientific achievements produced in the 20th century. However, the way that science advances itself is much less well-understood. In this work, we study the evolution of scientific development over the past century by presenting an anatomy of 89 million digitalized papers published between 1900 and 2015. We find that science has benefited from the shift from individual work to collaborative effort, with over 90% of the world-leading innovations generated by collaborations in this century, nearly four times higher than they were in the 1900s. We discover that rather than the frequent myopic- and self-referencing that was common in the early 20th century, modern scientists instead tend to look for literature further back and farther around. Finally, we also observe the globalization of scientific development from 1900 to 2015, including 25-fold and 7-fold increases in international collaborations and citations, respectively, as well as a dramatic decline in the dominant accumulation of citations by the US, the UK, and Germany, from ~95% to ~50% over the same period. Our discoveries are meant to serve as a starter for exploring the visionary ways in which science has developed throughout the past century, generating insight into and an impact upon the current scientific innovations and funding policies. Yuxiao Dong, Hao Ma 0001, Zhihong Shen, Kuansan Wang |
KDD | 2 |
| 2016 | Question Answering with Knowledge Base, Web and BeyondabstractIn this tutorial, we give the audience a coherent overview of the research of question answering (QA). We first introduce a variety of QA problems proposed by pioneer researchers and briefly describe the early efforts. By contrasting with the current research trend in this domain, the audience can easily comprehend what technical problems remain challenging and what the main breakthroughs and opportunities are during the past half century. For the rest of the tutorial, we select three categories of the QA problems that have recently attracted a great deal of attention in the research community, and present the tasks with the latest technical survey. We conclude the tutorial by discussing the new opportunities and future directions of QA research. Scott Yih, Hao Ma 0001 |
SIGIR | 2 |
| 2016 | User Fatigue in Online News RecommendationabstractMany aspects and properties of Recommender Systems have been well studied in the past decade, however, the impact of User Fatigue has been mostly ignored in the literature. User fatigue represents the phenomenon that a user quickly loses the interest on the recommended item if the same item has been presented to this user multiple times before. The direct impact caused by the user fatigue is the dramatic decrease of the Click Through Rate (CTR, i.e., the ratio of clicks to impressions). In this paper, we present a comprehensive study on the research of the user fatigue in online recommender systems. By analyzing user behavioral logs from Bing Now news recommendation, we find that user fatigue is a severe problem that greatly affects the user experience. We also notice that different users engage differently with repeated recommendations. Depending on the previous users' interaction with repeated recommendations, we illustrate that under certain condition the previously seen items should be demoted, while some other times they should be promoted. We demonstrate how statistics about the analysis of the user fatigue can be incorporated into ranking algorithms for personalized recommendations. Our experimental results indicate that significant gains can be achieved by introducing features that reflect users' interaction with previously seen recommendations (up to 15% enhancement on all users and 34% improvement on heavy users). Hao Ma 0001, Xueqing Liu 0001, Zhihong Shen |
WWW | 1 |
| 2016 | Table Cell Search for Question AnsweringabstractTables are pervasive on the Web. Informative web tables range across a large variety of topics, which can naturally serve as a significant resource to satisfy user information needs. Driven by such observations, in this paper, we investigate an important yet largely under-addressed problem: Given millions of tables, how to precisely retrieve table cells to answer a user question. This work proposes a novel table cell search framework to attack this problem. We first formulate the concept of a relational chain which connects two cells in a table and represents the semantic relation between them. With the help of search engine snippets, our framework generates a set of relational chains pointing to potentially correct answer cells. We further employ deep neural networks to conduct more fine-grained inference on which relational chains best match the input question and finally extract the corresponding answer cells. Based on millions of tables crawled from the Web, we evaluate our framework in the open-domain question answering (QA) setting, using both the well-known WebQuestions dataset and user queries mined from Bing search engine logs. On WebQuestions, our framework is comparable to state-of-the-art QA systems based on knowledge bases (KBs), while on Bing queries, it outperforms other systems with a 56.7% relative gain. Moreover, when combined with results from our framework, KB-based QA performance can obtain a relative improvement of 28.1% to 66.7%, demonstrating that web tables supply rich knowledge that might not exist or is difficult to be identified in existing KBs. Huan Sun 0001, Hao Ma 0001, Xiaodong He 0001, Scott Yih, Yu Su 0001, Xifeng Yan |
WWW | 2 |
| 2015 | Learning to Recommend Related Entities to Search UsersabstractOver the past few years, major web search engines have introduced knowledge bases to offer popular facts about people, places, and things on the entity pane next to regular search results. In addition to information about the entity searched by the user, the entity pane often provides a ranked list of related entities. To keep users engaged, it is important to develop a recommendation model that tailors the related entities to individual user interests. We propose a probabilistic Three-way Entity Model (TEM) that provides personalized recommendation of related entities using three data sources: knowledge base, search click log, and entity pane log. Specifically, TEM is capable of extracting hidden structures and capturing underlying correlations among users, main entities, and related entities. Moreover, the TEM model can also exploit the click signals derived from the entity pane log. We further provide an inference technique to learn the parameters in TEM, and propose a principled preference learning method specifically designed for ranking related entities. Extensive experiments with two real-world datasets show that TEM with our probabilistic framework significantly outperforms a state of the art baseline, confirming the effectiveness of TEM and our probabilistic framework in related entity recommendation. Bin Bi, Hao Ma 0001, Bo-June Paul Hsu, Kuansan Wang, Junghoo Cho |
WSDM | 2 |
| 2015 | Open Domain Question Answering via Semantic EnrichmentabstractMost recent question answering (QA) systems query large-scale knowledge bases (KBs) to answer a question, after parsing and transforming natural language questions to KBs-executable forms (e.g., logical forms). As a well-known fact, KBs are far from complete, so that information required to answer questions may not always exist in KBs. In this paper, we develop a new QA system that mines answers directly from the Web, and meanwhile employs KBs as a significant auxiliary to further boost the QA performance. Specifically, to the best of our knowledge, we make the first attempt to link answer candidates to entities in Freebase, during answer candidate generation. Several remarkable advantages follow: (1) Redundancy among answer candidates is automatically reduced. (2) The types of an answer candidate can be effortlessly determined by those of its corresponding entity in Freebase. (3) Capitalizing on the rich information about entities in Freebase, we can develop semantic features for each answer candidate after linking them to Freebase. Particularly, we construct answer-type related features with two novel probabilistic models, which directly evaluate the appropriateness of an answer candidate's types under a given question. Overall, such semantic features turn out to play significant roles in determining the true answers from the large answer candidate pool. The experimental results show that across two testing datasets, our QA system achieves an 18%~54% improvement under F_1 metric, compared with various existing QA systems. Huan Sun 0001, Hao Ma 0001, Scott Yih, Chen-Tse Tsai, Jingjing Liu 0001, Ming-Wei Chang |
WWW | 2 |
| 2014 | On measuring social friend interest similarities in recommender systemsabstractSocial recommender system has become an emerging research topic due to the prevalence of online social networking services during the past few years. In this paper, aiming at providing fundamental support to the research of social recommendation problem, we conduct an in-depth analysis on the correlations between social friend relations and user interest similarities. When evaluating interest similarities without distinguishing different friends a user has, we surprisingly observe that social friend relations generally cannot represent user interest similarities. A user's average similarity on all his/her friends is even correlated with the average similarity on some other randomly selected users. However, when measuring interest similarities using a finer granularity, we find that the similarities between a user and his/her friends are actually controlled by the network structure in the friend network. Factors that affect the interest similarities include subgraph topology, connected components, number of co-friends, etc. We believe our analysis provides substantial impact for social recommendation research and will benefit ongoing research in both recommender systems and other social applications. Hao Ma 0001 |
SIGIR | 1 |
| 2014 | On building entity recommender systems using user click log and freebase knowledgeabstractDue to their commercial value, search engines and recommender systems have become two popular research topics in both industry and academia over the past decade. Although these two fields have been actively and extensively studied separately, researchers are beginning to realize the importance of the scenarios at their intersection: providing an integrated search and information discovery user experience. In this paper, we study a novel application, i.e., personalized entity recommendation for search engine users, by utilizing user click log and the knowledge extracted from Freebase. Xiao Yu 0007, Hao Ma 0001, Bo-June Paul Hsu, Jiawei Han 0001 |
WSDM | 2 |
| 2013 | An experimental study on implicit social recommendationabstractSocial recommendation problems have drawn a lot of attention recently due to the prevalence of social networking sites. The experiments in previous literature suggest that social information is very effective in improving traditional recommendation algorithms. However, explicit social information is not always available in most of the recommender systems, which limits the impact of social recommendation techniques. In this paper, we study the following two research problems: (1) In some systems without explicit social information, can we still improve recommender systems using implicit social information? (2) In the systems with explicit social information, can the performance of using implicit social information outperform that of using explicit social information? In order to answer these two questions, we conduct comprehensive experimental analysis on three recommendation datasets. The result indicates that: (1) Implicit user and item social information, including similar and dissimilar relationships, can be employed to improve traditional recommendation methods. (2) When comparing implicit social information with explicit social information, the performance of using implicit information is slightly worse. This study provides additional insights to social recommendation techniques, and also greatly widens the utility and spreads the impact of previous and upcoming social recommendation approaches. Hao Ma 0001 |
SIGIR | 1 |
| 2013 | Exploring and exploiting user search behavior on mobile and tablet devices to improve search relevanceabstractIn this paper, we present a log-based study on user search behavior comparisons on three different platforms: desktop, mobile and tablet. We use three-month search logs in 2012 from a commercial search engine for our study. Our objective is to better understand how and to what extent mobile and tablet searchers behave differently than desktop users. Our study spans a variety of aspects including query categorization, query length, search time distribution, search location distribution, user click patterns and so on. From our data set, we reveal that there are significant differences between user search patterns in these three platforms, and therefore use the same ranking system is not an optimal solution for all of them. Consequently, we propose a framework that leverages a set of domain-specific features, along with the training data from desktop search, to further improve the search relevance for mobile and tablet platforms. Experimental results demonstrate that by transferring knowledge from desktop search, search relevance on mobile and tablet can be greatly improved. Yang Song 0008, Hao Ma 0001, Hongning Wang, Kuansan Wang |
WWW | 2 |
| 2013 | Collaborative Web Service QoS Prediction via Neighborhood Integrated Matrix FactorizationabstractWith the increasing presence and adoption of web services on the World Wide Web, the demand of efficient web service quality evaluation approaches is becoming unprecedentedly strong. To avoid the expensive and time-consuming web service invocations, this paper proposes a collaborative quality-of-service (QoS) prediction approach for web services by taking advantages of the past web service usage experiences of service users. We first apply the concept of user-collaboration for the web service QoS information sharing. Then, based on the collected QoS data, a neighborhood-integrated approach is designed for personalized web service QoS value prediction. To validate our approach, large-scale real-world experiments are conducted, which include 1,974,675 web service invocations from 339 service users on 5,825 real-world web services. The comprehensive experimental studies show that our proposed approach achieves higher prediction accuracy than other approaches. The public release of our web service QoS data set provides valuable real-world data for future research. Zibin Zheng, Hao Ma 0001, Michael R. Lyu, Irwin King |
IEEE Trans. Serv. Comput. | 2 |
| 2012 | Mining Web Graphs for RecommendationsabstractAs the exponential explosion of various contents generated on the Web, Recommendation techniques have become increasingly indispensable. Innumerable different kinds of recommendations are made on the Web every day, including movies, music, images, books recommendations, query suggestions, tags recommendations, etc. No matter what types of data sources are used for the recommendations, essentially these data sources can be modeled in the form of various types of graphs. In this paper, aiming at providing a general framework on mining Web graphs for recommendations, (1) we first propose a novel diffusion method which propagates similarities between different nodes and generates recommendations; (2) then we illustrate how to generalize different recommendation problems into our graph diffusion framework. The proposed framework can be utilized in many recommendation tasks on the World Wide Web, including query suggestions, tag recommendations, expert finding, image recommendations, image annotations, etc. The experimental analysis on large data sets shows the promising future of our work. Hao Ma 0001, Irwin King, Michael R. Lyu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Mining test oracles of web search enginesabstractWeb search engines have major impact in people's everyday life. It is of great importance to test the retrieval effectiveness of search engines. However, it is labor-intensive to judge the relevance of search results for a large number of queries, and these relevance judgments may not be reusable since the Web data change all the time. In this work, we propose to mine test oracles of Web search engines from existing search results. The main idea is to mine implicit relationships between queries and search results, e.g., some queries may have fixed top 1 result while some may not, and some Web domains may appear together in top 10 results. We define a set of items of queries and search results, and mine frequent association rules between these items as test oracles. Experiments on major search engines show that our approach mines many high-confidence rules that help understand search engines and detect suspicious search results. Wujie Zheng, Hao Ma 0001, Michael R. Lyu, Tao Xie 0001, Irwin King |
ASE | 2 |
| 2011 | Probabilistic factor models for web site recommendationabstractDue to the prevalence of personalization and information filteringapplications, modelingusers ’ interests on theWeb has become increasingly important duringthe past few years. In this paper, aiming at providing accurate personalized Web site recommendations for Web users, we propose a novel probabilistic factor model based on dimensionality reduction techniques. We also extend the proposed method to collective probabilistic factor modeling, which further improves model performance by incorporating heterogeneous data sources. The proposed method is general, and can be applied to not only Web site recommendations, but also a wide range of Web applications, including behavioral targeting, sponsored search, etc. The experimental analysis on Web site recommendation shows that our method outperforms other traditional recommendation approaches. Moreover, the complexity analysis indicates that our approach can be applied to very large datasets since it scales linearly with the number of observations. Categories and Subject Descriptors Hao Ma 0001, Chao Liu 0001, Irwin King, Michael R. Lyu |
SIGIR | 1 |
| 2011 | Recommender systems with social regularizationabstractAlthough Recommender Systems have been comprehensively analyzed in the past decade, the study of social-based recommender systems just started. In this paper, aiming at providing a general method for improving recommender systems by incorporating social network information, we propose a matrix factorization framework with social regularization. The contributions of this paper are four-fold: (1) We elaborate how social network information can benefit recommender systems; (2) We interpret the differences between social-based recommender systems and trust-aware recommender systems; (3) We coin the term Social Regularization to represent the social constraints on recommender systems, and we systematically illustrate how to design a matrix factorization objective function with social regularization; and (4) The proposed method is quite general, which can be easily extended to incorporate other contextual information, like social tags, etc. The empirical analysis on two large datasets demonstrates that our approaches outperform other state-of-the-art methods. Hao Ma 0001, Dengyong Zhou, Chao Liu 0001, Michael R. Lyu, Irwin King |
WSDM | 1 |
| 2011 | Learning to recommend with explicit and implicit social relationsabstractRecommender systems have been well studied and developed, both in academia and in industry recently. However, traditional recommender systems assume that all the users are independent and identically distributed; this assumption ignores the connections among users, which is not consistent with the real-world observations where we always turn to our trusted friends for recommendations. Aiming at modeling recommender systems more accurately and realistically, we propose a novel probabilistic factor analysis framework which naturally fuses the users' tastes and their trusted friends' favors together. The proposed framework is quite general, and it can also be applied to pure user-item rating matrix even if we do not have explicit social trust information among users. In this framework, we coin the term social trust ensemble to represent the formulation of the social trust restrictions on the recommender systems. The complexity analysis indicates that our approach can be applied to very large datasets since it scales linearly with the number of observations, while the experimental results show that our method outperforms state-of-the-art approaches. Hao Ma 0001, Irwin King, Michael R. Lyu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Improving Recommender Systems by Incorporating Social Contextual InformationabstractDue to their potential commercial value and the associated great research challenges, recommender systems have been extensively studied by both academia and industry recently. However, the data sparsity problem of the involved user-item matrix seriously affects the recommendation quality. Many existing approaches to recommender systems cannot easily deal with users who have made very few ratings. In view of the exponential growth of information generated by online users, social contextual information analysis is becoming important for many Web applications. In this article, we propose a factor analysis approach based on probabilistic matrix factorization to alleviate the data sparsity and poor prediction accuracy problems by incorporating social contextual information, such as social networks and social tags. The complexity analysis indicates that our approach can be applied to very large datasets since it scales linearly with the number of observations. Moreover, the experimental results show that our method performs much better than the state-of-the-art approaches, especially in the circumstance that users have made few ratings. Hao Ma 0001, Tom Chao Zhou, Michael R. Lyu, Irwin King |
ACM Trans. Inf. Syst. | 1 |
| 2011 | QoS-Aware Web Service Recommendation by Collaborative FilteringabstractWith increasing presence and adoption of Web services on the World Wide Web, Quality-of-Service (QoS) is becoming important for describing nonfunctional characteristics of Web services. In this paper, we present a collaborative filtering approach for predicting QoS values of Web services and making Web service recommendation by taking advantages of past usage experiences of service users. We first propose a user-collaborative mechanism for past Web service QoS information collection from different service users. Then, based on the collected QoS data, a collaborative filtering approach is designed to predict Web service QoS values. Finally, a prototype called WSRec is implemented by Java language and deployed to the Internet for conducting real-world experiments. To study the QoS value prediction accuracy of our approach, 1.5 millions Web service invocation results are collected from 150 service users in 24 countries on 100 real-world Web services in 22 countries. The experimental results show that our algorithm achieves better prediction accuracy than other approaches. Our Web service QoS data set is publicly released for future research. Zibin Zheng, Hao Ma 0001, Michael R. Lyu, Irwin King |
IEEE Trans. Serv. Comput. | 2 |
| 2010 | Diversifying Query Suggestion ResultsabstractIn order to improve the user search experience, Query Suggestion, a technique for generating alternative queries to Web users, has become an indispensable feature for commercial search engines. However, previous work mainly focuses on suggesting relevant queries to the original query while ignoring the diversity in the suggestions, which will potentially dissatisfy Web users' information needs. In this paper, we present a novel unified method to suggest both semantically relevant and diverse queries to Web users. The proposed approach is based on Markov random walk and hitting time analysis on the query-URL bipartite graph. It can effectively prevent semantically redundant queries from receiving a high rank, hence encouraging diversities in the results. We evaluate our method on a large commercial clickthrough dataset in terms of relevance measurement and diversity measurement. The experimental results show that our method is very effective in generating both relevant and diverse query suggestions. Hao Ma 0001, Michael R. Lyu, Irwin King |
AAAI | 1 |
| 2010 | UserRec: A User Recommendation Framework in Social Tagging SystemsabstractSocial tagging systems have emerged as an effective way for users to annotate and share objects on the Web. However, with the growth of social tagging systems, users are easily overwhelmed by the large amount of data and it is very difficult for users to dig out information that he/she is interested in. Though the tagging system has provided interest-based social network features to enable the user to keep track of other users' tagging activities, there is still no automatic and effective way for the user to discover other users with common interests. In this paper, we propose a User Recommendation (UserRec) framework for user interest modeling and interest-based user recommendation, aiming to boost information sharing among users with similar interests. Our work brings three major contributions to the research community: (1) we propose a tag-graph based community detection method to model the users' personal interests, which are further represented by discrete topic distributions; (2) the similarity values between users' topic distributions are measured by Kullback-Leibler divergence (KL-divergence), and the similarity values are further used to perform interest-based user recommendation; and (3) by analyzing users' roles in a tagging system, we find users' roles in a tagging system are similar to Web pages in the Internet. Experiments on tagging dataset of Web pages (Yahoo!~Delicious) show that UserRec outperforms other state-of-the-art recommender system approaches. Tom Chao Zhou, Hao Ma 0001, Michael R. Lyu, Irwin King |
AAAI | 2 |
| 2010 | Introduction to social recommendationabstractAs the exponential growth of information generated on the World Wide Web, Social Recommendation has emerged as one of the hot research topics recently. Social Recommendation forms a specific type of information filtering technique that attempts to suggest information (blogs, news, music, travel plans, web pages, images, tags, etc.) that are likely to interest the users. Social Recommendation involves the investigation of collective intelligence by using computational techniques such as machine learning, data mining, natural language processing, etc. on social behavior data collected from blogs, wikis, recommender systems, question & answer communities, query logs, tags, etc. from areas such as social networks, social search, social media, social bookmarks, social news, social knowledge sharing, and social games. In this tutorial, we will introduce Social Recommendation and elaborate on how the various characteristics and aspects are involved in the social platforms for collective intelligence. Moreover, we will discuss the challenging issues involved in Social Recommendation in the context of theory and models of social networks, methods to improve recommender systems using social contextual information, ways to deal with partial and incomplete information in the social context, scalability and algorithmic issues with social computational techniques. Irwin King, Michael R. Lyu, Hao Ma 0001 |
WWW | 3 |
| 2010 | Bridging the Semantic Gap Between Image Contents and TagsabstractWith the exponential growth of Web 2.0 applications, tags have been used extensively to describe the image contents on the Web. Due to the noisy and sparse nature in the human generated tags, how to understand and utilize these tags for image retrieval tasks has become an emerging research direction. As the low-level visual features can provide fruitful information, they are employed to improve the image retrieval results. However, it is challenging to bridge the semantic gap between image contents and tags. To attack this critical problem, we propose a unified framework in this paper which stems from a two-level data fusions between the image contents and tags: 1) A unified graph is built to fuse the visual feature-based image similarity graph with the image-tag bipartite graph; 2) A novel random walk model is then proposed, which utilizes a fusion parameter to balance the influences between the image contents and tags. Furthermore, the presented framework not only can naturally incorporate the pseudo relevance feedback process, but also it can be directly applied to applications such as content-based image retrieval, text-based image retrieval, and image annotation. Experimental analysis on a large Flickr dataset shows the effectiveness and efficiency of our proposed framework. Hao Ma 0001, Jianke Zhu, Michael R. Lyu, Irwin King |
IEEE Trans. Multim. | 1 |
| 2009 | Semi-nonnegative matrix factorization with global statistical consistency for collaborative filteringabstractCollaborative Filtering, considered by many researchers as the most important technique for information filtering, has been extensively studied by both academic and industrial communities. One of the most popular approaches to collaborative filtering recommendation algorithms is based on low-dimensional factor models. The assumption behind such models is that a user's preferences can be modeled by linearly combining item factor vectors using user-specific coefficients. In this paper, aiming at several aspects ignored by previous work, we propose a semi-nonnegative matrix factorization method with global statistical consistency. The major contribution of our work is twofold: (1) We endow a new understanding on the generation or latent compositions of the user-item rating matrix. Under the new interpretation, our work can be formulated as the semi-nonnegative matrix factorization problem. (2) Moreover, we propose a novel method of imposing the consistency between the statistics given by the predicted values and the statistics given by the data. We further develop an optimization algorithm to determine the model complexity automatically. The complexity of our method is linear with the number of the observed ratings, hence it is scalable to very large datasets. Finally, comparing with other state-of-the-art methods, the experimental analysis on the EachMovie dataset illustrates the effectiveness of our approach. Hao Ma 0001, Haixuan Yang, Irwin King, Michael R. Lyu |
CIKM | 1 |
| 2009 | WSRec: A Collaborative Filtering Based Web Service Recommender SystemabstractAs the abundance of Web services on the World Wide Web increase, designing effective approaches for Web service selection and recommendation has become more and more important. In this paper, we present WSRec, a Web service recommender system, to attack this crucial problem. WSRec includes a user-contribution mechanism for Web service QoS information collection and an effective and novel hybrid collaborative filtering algorithm for Web service QoS value prediction. WSRec is implemented by Java language and deployed to the real-world environment. To study the prediction performance, a total of 21,197 public Web services are obtained from the Internet and a large-scale real-world experiment is conducted, where more than 1.5 millions test results are collected from 150 service users in different countries on 100 publicly available Web services located all over the world. The comprehensive experimental analysis shows that WSRec achieves better prediction accuracy than other approaches. Zibin Zheng, Hao Ma 0001, Michael R. Lyu, Irwin King |
ICWS | 2 |
| 2009 | Learning to recommend with trust and distrust relationshipsabstractWith the exponential growth of Web contents, Recommender System has become indispensable for discovering new information that might interest Web users. Despite their success in the industry, traditional recommender systems suffer from several problems. First, the sparseness of the user-item matrix seriously affects the recommendation quality. Second, traditional recommender systems ignore the connections among users, which loses the opportunity to provide more accurate and personalized recommendations. In this paper, aiming at providing more realistic and accurate recommendations, we propose a factor analysis-based optimization framework to incorporate the user trust and distrust relationships into the recommender systems. The contributions of this paper are three-fold: (1) We elaborate how user distrust information can benefit the recommender systems. (2) In terms of the trust relations, distinct from previous trust-aware recommender systems which are based on some heuristics, we systematically interpret how to constrain the objective function with trust regularization. (3) The experimental results show that the distrust relations among users are as important as the trust relations. The complexity analysis shows our method scales linearly with the number of observations, while the empirical analysis on a large Epinions dataset proves that our approaches perform better than the state-of-the-art approaches. Hao Ma 0001, Michael R. Lyu, Irwin King |
RecSys | 1 |
| 2009 | Learning to recommend with social trust ensembleabstractAs an indispensable technique in the field of Information Filtering, Recommender System has been well studied and developed both in academia and in industry recently. However, most of current recommender systems suffer the following problems: (1) The large-scale and sparse data of the user-item matrix seriously affect the recommendation quality. As a result, most of the recommender systems cannot easily deal with users who have made very few ratings. (2) The traditional recommender systems assume that all the users are independent and identically distributed; this assumption ignores the connections among users, which is not consistent with the real world recommendations. Aiming at modeling recommender systems more accurately and realistically, we propose a novel probabilistic factor analysis framework, which naturally fuses the users' tastes and their trusted friends' favors together. In this framework, we coin the term Social Trust Ensemble to represent the formulation of the social trust restrictions on the recommender systems. The complexity analysis indicates that our approach can be applied to very large datasets since it scales linearly with the number of observations, while the experimental results show that our method performs better than the state-of-the-art approaches. Hao Ma 0001, Irwin King, Michael R. Lyu |
SIGIR | 1 |
| 2008 | Learning latent semantic relations from clickthrough data for query suggestionabstractFor a given query raised by a specific user, the Query Suggestion technique aims to recommend relevant queries which potentially suit the information needs of that user. Due to the complexity of the Web structure and the ambiguity of users' inputs, most of the suggestion algorithms suffer from the problem of poor recommendation accuracy. In this paper, aiming at providing semantically relevant queries for users, we develop a novel, effective and efficient two-level query suggestion model by mining clickthrough data, in the form of two bipartite graphs (user-query and query-URL bipartite graphs) extracted from the clickthrough data. Based on this, we first propose a joint matrix factorization method which utilizes two bipartite graphs to learn the low-rank query latent feature space, and then build a query similarity graph based on the features. After that, we design an online ranking algorithm to propagate similarities on the query similarity graph, and finally recommend latent semantically relevant queries to users. Experimental analysis on the clickthrough data of a commercial search engine shows the effectiveness and the efficiency of our method. Hao Ma 0001, Haixuan Yang, Irwin King, Michael R. Lyu |
CIKM | 1 |
| 2008 | Mining social networks using heat diffusion processes for marketing candidates selectionabstractSocial Network Marketing techniques employ pre-existing social networks to increase brands or products awareness through word-of-mouth promotion. Full understanding of social network marketing and the potential candidates that can thus be marketed to certainly offer lucrative opportunities for prospective sellers. Due to the complexity of social networks, few models exist to interpret social network marketing realistically. We propose to model social network marketing using Heat Diffusion Processes. This paper presents three diffusion models, along with three algorithms for selecting the best individuals to receive marketing samples. These approaches have the following advantages to best illustrate the properties of real-world social networks: (1) We can plan a marketing strategy sequentially in time since we include a time factor in the simulation of product adoptions; (2) The algorithm of selecting marketing candidates best represents and utilizes the clustering property of real-world social networks; and (3) The model we construct can diffuse both positive and negative comments on products or brands in order to simulate the complicated communications within social networks. Our work represents a novel approach to the analysis of social network marketing, and is the first work to propose how to defend against negative comments within social networks. Complexity analysis shows our model is also scalable to very large social networks. Hao Ma 0001, Haixuan Yang, Michael R. Lyu, Irwin King |
CIKM | 1 |
| 2008 | SoRec: social recommendation using probabilistic matrix factorizationabstractData sparsity, scalability and prediction quality have been recognized as the three most crucial challenges that every collaborative filtering algorithm or recommender system confronts. Many existing approaches to recommender systems can neither handle very large datasets nor easily deal with users who have made very few ratings or even none at all. Moreover, traditional recommender systems assume that all the users are independent and identically distributed; this assumption ignores the social interactions or connections among users. In view of the exponential growth of information generated by online social networks, social network analysis is becoming important for many Web applications. Following the intuition that a person's social network will affect personal behaviors on the Web, this paper proposes a factor analysis approach based on probabilistic matrix factorization to solve the data sparsity and poor prediction accuracy problems by employing both users' social network information and rating records. The complexity analysis indicates that our approach can be applied to very large datasets since it scales linearly with the number of observations, while the experimental results shows that our method performs much better than the state-of-the-art approaches, especially in the circumstance that users have made few or no ratings. Hao Ma 0001, Haixuan Yang, Michael R. Lyu, Irwin King |
CIKM | 1 |
| 2007 | Effective missing data prediction for collaborative filteringabstractMemory-based collaborative filtering algorithms have been widely adopted in many popular recommender systems, although these approaches all suffer from data sparsity and poor prediction quality problems. Usually, the user-item matrix is quite sparse, which directly leads to inaccurate recommendations. This paper focuses the memory-based collaborative filtering problems on two crucial factors: (1) similarity computation between users or items and (2) missing data prediction algorithms. First, we use the enhanced Pearson Correlation Coefficient (PCC) algorithm by adding one parameter which overcomes the potential decrease of accuracy when computing the similarity of users or items. Second, we propose an effective missing data prediction algorithm, in which information of both users and items is taken into account. In this algorithm, we set the similarity threshold for users and items respectively, and the prediction algorithm will determine whether predicting the missing data or not. We also address how to predict the missing data by employing a combination of user and item information. Finally, empirical studies on dataset MovieLens have shown that our newly proposed method outperforms other state-of-the-art collaborative filtering algorithms and it is more robust against data sparsity. Hao Ma 0001, Irwin King, Michael R. Lyu |
SIGIR | 1 |