VLDB 2026 Research / reviewers in the wild / expert
Shenghua Bao
dblp:31/529
· DBLP profile ↗
38ranked-venue papers
9as first author
2since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 29 · 7 first-author · 2 since 2021Artificial intelligence and machine learning · 17 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
19 papers |
Information retrieval · 35% Data mining · 25% Web and social media mining · 17% | |
| Artificial intelligence
6 papers |
Information extraction and text analysis · 60% Language models and text generation · 26% Question answering and dialogue systems · 14% | |
| Computer graphics and multimedia
4 papers |
Multimedia analysis and retrieval · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 42, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining › text mining
topic modeling |
0.4 | 3 | 2014 | Probabilistic text modeling with orthogonalized topics · SIGIR 2014 Mining topics on participations for community discovery · SIGIR 2011 Joint Emotion-Topic Modeling for Social Affective Text Mining · ICDM 2009 |
Multimedia analysis and retrieval › image retrieval › instance-level image retrieval
landmark search |
0.3 | 2 | 2012 | Searching for diversified landmarks by photo · ACM Multimedia 2012 DLMSearch: diversified landmark search by photo · ACM Multimedia 2012 |
Recommender systems › e-commerce recommendation
product recommendation |
0.2 | 1 | 2016 | Boosting Recommendation in Unexplored Categories by User Price Preference · ACM Trans. Inf. Syst. 2016 |
Natural language and speech › Language models and text generation
large language model |
0.2 | 1 | 2024 | The 2nd International Workshop: From Innovation to Scale (I2S) - Successfully Build, Commercialize, and Scale AI Innovations · KDD 2024 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.2 | 1 | 2015 | Predicting Future Scientific Discoveries Based on a Networked Analysis of the Past Literature · KDD 2015 |
Knowledge graphs
knowledge graph reasoning |
0.2 | 1 | 2015 | Predicting Future Scientific Discoveries Based on a Networked Analysis of the Past Literature · KDD 2015 |
Recommender systems
cold-start recommendation |
0.2 | 1 | 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014 |
Recommender systems
cross-domain recommendation |
0.2 | 1 | 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014 |
Information retrieval
text analysis |
0.2 | 1 | 2014 | Probabilistic text modeling with orthogonalized topics · SIGIR 2014 |
Information retrieval
diversified retrieval |
0.2 | 2 | 2012 | DLMSearch: diversified landmark search by photo · ACM Multimedia 2012 Searching for diversified landmarks by photo · ACM Multimedia 2012 |
Multimedia analysis and retrieval › object recognition
landmark recognition |
0.2 | 1 | 2013 | Tell me what happened here in history · ACM Multimedia 2013 |
Web and social media mining › web mining
competitor mining |
0.1 | 2 | 2008 | Competitor Mining with the Web · IEEE Trans. Knowl. Data Eng. 2008 CoMiner: An Effective Algorithm for Mining Competitors from the Web · ICDM 2006 |
Information retrieval › search engines › semantic search › entity retrieval
entity ranking |
0.1 | 2 | 2008 | Competitor Mining with the Web · IEEE Trans. Knowl. Data Eng. 2008 CoMiner: An Effective Algorithm for Mining Competitors from the Web · ICDM 2006 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.1 | 1 | 2012 | Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2012 | Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012 |
Natural language and speech › Question answering and dialogue systems
community question answering |
0.1 | 1 | 2011 | Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 1 | 2011 | Analyzing and Predicting Not-Answered Questions in Community-based Question Answering Services · AAAI 2011 |
Web and social media mining
co-authorship networks |
0.1 | 1 | 2011 | Mining topics on participations for community discovery · SIGIR 2011 |
Data mining › structured data mining › graph mining
community detection |
0.1 | 1 | 2011 | Mining topics on participations for community discovery · SIGIR 2011 |
Data mining › structured data mining
graph mining |
0.1 | 1 | 2011 | Mining topics on participations for community discovery · SIGIR 2011 |
Data mining › text mining
sentiment analysis |
0.1 | 1 | 2011 | OOLAM: an opinion oriented link analysis model for influence persona discovery · WSDM 2011 |
Web and social media mining
social influence analysis |
0.1 | 1 | 2011 | OOLAM: an opinion oriented link analysis model for influence persona discovery · WSDM 2011 |
Data mining › text mining
text classification |
0.1 | 2 | 2014 | Probabilistic text modeling with orthogonalized topics · SIGIR 2014 Mining Social Emotions from Affective Text · IEEE Trans. Knowl. Data Eng. 2012 |
Web and social media mining
social tagging |
0.1 | 2 | 2008 | Towards effective browsing of large scale social annotations · WWW 2007 Exploring folksonomy for personalized search · SIGIR 2008 |
Information retrieval › search engines
expert finding |
0.1 | 1 | 2008 | A Probabilistic Model for Fine-Grained Expert Search · ACL 2008 |
Information retrieval
personalized search |
0.1 | 1 | 2008 | Exploring folksonomy for personalized search · SIGIR 2008 |
Information retrieval › retrieval models
query-document similarity |
0.1 | 1 | 2007 | Optimizing web search using social annotations · WWW 2007 |
Information retrieval
ranking |
0.1 | 1 | 2007 | Optimizing web search using social annotations · WWW 2007 |
Information retrieval
retrieval models |
0.1 | 1 | 2007 | Optimizing web search using social annotations · WWW 2007 |
Information retrieval
tag cloud |
0.1 | 1 | 2007 | Towards effective browsing of large scale social annotations · WWW 2007 |
Methods — techniques the papers use, named apart from their topics
network analysis · 0.7matrix factorization · 0.7graph diffusion · 0.7latent dirichlet allocation · 0.5tree search · 0.4objective function optimization · 0.4utility function · 0.2regularization · 0.2joint factorization · 0.2price factor modeling · 0.2multi-source information fusion · 0.2supervised learning · 0.1feature engineering · 0.1web-scale mining · 0.1sequential covering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | The 2nd International Workshop: From Innovation to Scale (I2S) - Successfully Build, Commercialize, and Scale AI InnovationsabstractIn recent years, there have been exciting and accelerated developments in AI with novel developments in foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, the pace of adoption of these innovations driven by both academic and industry research labs has sped up with both big tech companies and startups looking to deliver value-differentiated products and services. With Generative AI (GenAI) garnering significant attention, the second edition of the I2S workshop focuses on two aspects: First, bringing together AI thought leaders from academia, big tech, and startups to discuss the opportunities, use-case themes, challenges, and risks of GenAI in various business verticals; and Second, bringing together startup founders to share experiences and lessons learned in commercializing GenAI innovations into successful enterprises highlighting challenges through the entire commercial journey - from productization to acquiring customers, building a team, and securing funding. Ankur Teredesai, Michael Zeller, Mohak Shah, Shenghua Bao, Wee Hyong Tok, Linsey Pang |
KDD | 4 |
| 2023 | From Innovation to Scale (I2S) - Discuss and Learn How to Successfully Build, Commercialize, and Scale AI Innovations in Challenging Market ConditionsabstractIn recent years, the AI community has witnessed an exciting acceleration in innovation across foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, AI innovations driven by both academic and industry research labs have rapidly been adopted by big tech companies and startups to deliver value-differentiated products and services. Ankur Teredesai, Michael Zeller, Shenghua Bao, Wee Hyong Tok, Linsey Pang |
KDD | 3 |
| 2016 | Boosting Recommendation in Unexplored Categories by User Price PreferenceabstractState-of-the-art methods for product recommendation encounter a significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this article, we investigate the challenging problem of product recommendation in unexplored categories and discover that the price, a factor comparable across categories, can improve the recommendation performance significantly. We introduce the price utility concept to characterize users’ sense of price and propose three different utility functions. We show that user price preference in a category is a distribution and we mine typical user price preference patterns based on three different types of distance between distributions. We fuse user price preference through regularization and joint factorization to boost recommendation performance in both browsing and buying shopping orientations. Experimental results show that fusing user price preference improves performance in a series of recommendation tasks: unexplored category recommendation, product recommendation under a given unexplored category, and product recommendation under generic unexplored categories. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2015 | Predicting Future Scientific Discoveries Based on a Networked Analysis of the Past LiteratureabstractWe present KnIT, the Knowledge Integration Toolkit, a system for accelerating scientific discovery and predicting previously unknown protein-protein interactions. Such predictions enrich biological research and are pertinent to drug discovery and the understanding of disease. Unlike a prior study, KnIT is now fully automated and demonstrably scalable. It extracts information from the scientific literature, automatically identifying direct and indirect references to protein interactions, which is knowledge that can be represented in network form. It then reasons over this network with techniques such as matrix factorization and graph diffusion to predict new, previously unknown interactions. The accuracy and scope of KnIT's knowledge extractions are validated using comparisons to structured, manually curated data sources as well as by performing retrospective studies that predict subsequent literature discoveries using literature available prior to a given date. The KnIT methodology is a step towards automated hypothesis generation from text, with potential application to other scientific domains. Meena Nagarajan, Angela D. Wilkins, Benjamin J. Bachman, Ilya B. Novikov, Shenghua Bao, Peter J. Haas, María E. Terrón-Díaz, Sumit Bhatia, Anbu K. Adikesavan, Jacques J. Labrie, Sam Regenbogen, Christie M. Buchovecky, Curtis R. Pickering, Linda Kato, Andreas Martin Lisewski, Ana Lelescu, Houyin Zhang, Stephen Boyer, Griff Weber, Ying Chen 0001, Lawrence A. Donehower, W. Scott Spangler, Olivier Lichtarge |
KDD | 5 |
| 2015 | Lead curve detection in drawings with complex cross-points
Jia Chen 0001, Min Li 0022, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001 |
Neurocomputing | 4 |
| 2015 | Exploitation and Exploration Balanced Hierarchical Summary for Landmark ImagesabstractWhile we have made significant progress over image understanding and search, how to meet the ultimate goal of satisfying both exploration and exploitation in one single system is still an open challenge. In the context of landmark images, it means that a system should not only be able to help users to quickly locate the photo they are interested in (exploitation), but also to discover different parts of the landmark which have never been seen before (exploration), which is a common request as evidenced by many recent multimedia studies. To the best of our knowledge, existing systems mainly focus on either exploration (e.g., photo browsing) or exploitation (e.g., representative photo identification), while users' need of exploration and exploitation is dynamically mixed. In this paper, we tackle the challenge by organizing landmark images into a hierarchical summary which gives user the flexibility of conducting both exploration and exploitation. In the hierarchical summary construction, we introduce two principles: the coherence principle and the diversity principle. Behind these two principles, the intrinsic concept is “detail-level,” which measures how much detail that an image reflects for a certain landmark. A new objective function is derived from the definition of both exploration and exploitation experience on detail-level. The problem of finding an optimal hierarchical summary is formulated as searching over a space of trees for the one that achieves the best objective score. Extensive quantitative experimental results and comprehensive user studies show that the optimized hierarchical summary is able to satisfy both experiences simultaneously. Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Shimin Chen, Yong Yu 0001 |
IEEE Trans. Multim. | 3 |
| 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to helpabstractState-of-the-art methods for product recommendation encounter significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this paper, we investigate the challenge problem of product recommendation in unexplored categories and discover that the price, a factor transferrable across categories, can improve the recommendation performance significantly. Through our investigation, we address four research questions progressively: 1) what is the impact of unexplored category on recommendation performance? 2) How to represent the price factor from the recommendation point of view? 3) What does price factor across categories mean to recommendation? 4) How to utilize price factor across categories for recommendation in unexplored categories? Based on a series of experiments and analysis conducted on a dataset collected from a leading E-commerce website, we discover valuable findings for the above four questions: first, unexplored categories cause performance drop by 40% relatively for current recommendation systems; second, the price factor can be represented as either a quantity for a product or a distribution for a user to improve performance; third, consumer behavior with respect to price factor across categories is complicated and needs to be carefully modeled; finally and most importantly, we propose a new method which encodes the two perspectives of the price factor. The proposed method significantly improves the recommendation performance in unexplored categories over the state-of-the-art baseline systems and shortens the performance gap by 43% relatively. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
SIGIR | 4 |
| 2014 | Probabilistic text modeling with orthogonalized topicsabstractTopic models have been widely used for text analysis. Previous topic models have enjoyed great success in mining the latent topic structure of text documents. With many efforts made on endowing the resulting document-topic distributions with different motivations, however, none of these models have paid any attention on the resulting topic-word distributions.Since topic-word distribution also plays an important role in the modeling performance,topic models which emphasize only the resulting document-topic representations but pay less attention to the topic-term distributions are limited. In this paper, we propose the Orthogonalized Topic Model(OTM) which imposes an orthogonality constraint on the topic-term distributions. We also propose a novel model fitting algorithm based on the generalized Expectation-Maximization algorithm and the Newthon-Raphson method. Quantitative evaluation of text classification demonstrates that OTM outperforms other baseline models and indicates the important role played by topic orthogonalizing. Enpeng Yao, Guoqing Zheng, Ou Jin, Shenghua Bao, Kailong Chen, Zhong Su, Yong Yu 0001 |
SIGIR | 4 |
| 2013 | Head-shoulder based gender recognitionabstractThis paper proposes a novel gender recognition method based on the head-shoulder part of human body. The head-shoulder area contains much information that could be cues to infer the gender of a person, such as hair-style, face, neckline style and so on. A rich high-dimensional feature descriptor is designed to extract gradient, texture and orientation information from the head-shoulder area, then Partial Least Squares (PLS) is employed to learn a very low dimensional discriminative subspace. Features are projected into the low dimensional subspace and linear SVM is employed to learn an efficient classification model between the male and female categories. Experimental results on a large real-world dataset demonstrate the effectiveness of the proposed method. Min Li 0022, Shenghua Bao, Weishan Dong, Yu Wang 0021, Zhong Su |
ICIP | 2 |
| 2013 | Tell me what happened here in historyabstractThis demo shows our system that takes a landmark image as input, recognizes the landmark from the image and returns historical events of the landmark with related photos. Different from existing landmark related researches, we focus on the temporal dimension of a landmark. Our system automatically recognizes the landmark, shows historical events chronologically and provides detailed photos for the events. To build these functions, we fuse information from multiple online resources. Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 4 |
| 2012 | DLMSearch: diversified landmark search by photoabstractThis paper focuses on the problem of searching for diversified landmarks with photos. More particularly, we propose a system called DLMSearch which handles image query, searches for diversified landmarks and provides representative visual summaries. DLMSearch allows a user to upload a query photo and searches for landmarks with high relevance and diversity in real time. Then DLMSearch presents a delicate photo summary for each returned landmark, considering both visual representativeness and diversity. Quantative evaluations on a web-scale landmark photo collection demonstrate the effectiveness of the DLMSearch system. Experimental results verify the merits of the proposed system. Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 5 |
| 2012 | Searching for diversified landmarks by photoabstractThis demo focuses on the problem of searching for diversified landmarks with photos as input. More particularly, we propose a system called DLMSearch that allows a user to upload a photo as a query and searches for a diverse set of relevant landmarks in real time. It also presents a photo summary for each retrieved landmark, considering both visual representativeness and diversity. Our online demo is available at http://lm.apexlab.org/landmark/demo. Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 5 |
| 2012 | Mining Social Emotions from Affective TextabstractThis paper is concerned with the problem of mining social emotions from text. Recently, with the fast development of web 2.0, more and more documents are assigned by social users with emotion labels such as happiness, sadness, and surprise. Such emotions can provide a new aspect for document categorization, and therefore help online users to select related documents based on their emotional preferences. Useful as it is, the ratio with manual emotion labels is still very tiny comparing to the huge amount of web/enterprise documents. In this paper, we aim to discover the connections between social emotions and affective terms and based on which predict the social emotion from text content automatically. More specifically, we propose a joint emotion-topic model by augmenting Latent Dirichlet Allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model. Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2011 | Analyzing and Predicting Not-Answered Questions in Community-based Question Answering ServicesabstractThis paper focuses on analyzing and predicting not-answered questions in Community based Question Answering (CQA) services, such as Yahoo! Answers. In CQA services, users express their information needs by submitting natural language questions and await answers from other human users. Comparing to receiving results from web search engines using keyword queries, CQA users are likely to get more specific answers, because human answerers may catch the main point of the question. However, one of the key problems of this pattern is that sometimes no one helps to give answers, while web search engines hardly fail to response. In this paper, we analyze the not-answered questions and give a first try of predicting whether questions will receive answers. More specifically, we first analyze the questions of Yahoo Answers based on the features selected from different perspectives. Then, we formalize the prediction problem as supervised learning – binary classification problem and leverage the proposed features to make predictions. Extensive experiments are made on 76,251 questions collected from Yahoo! Answers. We analyze the specific characteristics of not-answered questions and try to suggest possible reasons for why a question is not likely to be answered. As for prediction, the experimental results show that classification based on the proposed features outperforms the simple word-based approach significantly. Lichun Yang, Shenghua Bao, Qingliang Lin, Xian Wu 0001, Dingyi Han, Zhong Su, Yong Yu 0001 |
AAAI | 2 |
| 2011 | Mining topics on participations for community discoveryabstractCommunity discovery on large-scale linked document corpora has been a hot research topic for decades. There are two types of links. The first one, which we call d2d-link, indicates connectiveness among different documents, such as blog references and research paper citations. The other one, which we call u2u-link, represents co-occurrences or simultaneous participations of different users in one document and typically each document from u2u-link corpus has more than one user/author. Examples of u2u-link data covers email archives and research paper co-authorship networks. Community discovery in d2d-link data has achieved much success, while methods for that in u2u-link data either make no use of the textual content of the documents or make oversimplified assumptions about the users and the textual content. In this paper we propose a general approach of community discovery for u2u-link data, i.e., multiple user data, by placing topical variables on multiple authors' participations in documents. Experiments on a research proceeding co-authorship corpus and a New York Times news corpus show the effectiveness of our model. Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001 |
SIGIR | 5 |
| 2011 | A Classification Framework for Disambiguating Web People Search Result Using Feedback
Ou Jin, Shenghua Bao, Zhong Su, Yong Yu 0001 |
WAIM | 2 |
| 2011 | Finding Appropriate Experts for Collaboration
Zhenjiang Zhan, Lichun Yang, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001 |
WAIM | 3 |
| 2011 | OOLAM: an opinion oriented link analysis model for influence persona discoveryabstractSocial influence is a complex and subtle force that governs the dynamics of social networks. In the past years, a lot of research work has been conducted to understand the spread patterns of social influence. However, most of approaches assume that influence exists between users with active social interactions, but ignore the question of what kind of influence happens between them. As such one interesting and also fundamental question is raised here: "in a social network, could the social connection reflect users'influence from both positive and negative aspects?". To this end, an Opinion Oriented Link Analysis Model (OOLAM) is proposed in this paper to characterize users' influence personae in order to exhibit their distinguishing influence ability in the social network. In particular, three types of influence personae are generalized and the problem of influence persona discovery is formally defined. Within the OOLAM model, two factors, i.e., opinion consistency and opinion creditability, are defined to capture the persona information from public opinion perspective. Extensive experimental studies have been performed to demonstrate the effectiveness of the proposed approach on influence persona analysis using real web data sets. Keke Cai, Shenghua Bao, Jie Tang 0001, Li Zhang 0007, Zhong Su |
WSDM | 2 |
| 2010 | A topical link model for community discovery in textual interaction graphabstractThis paper is concerned with community discovery in textual interaction graph, where the links between entities are indicated by textual documents. Specifically, we propose a Topical Link Model(TLM), which leverages Hierarchical Dirichlet Process(HDP) to introduce hidden topical variable of the links. Other than the use of links, TLM can look into the documents on the links in detail to recover sound communities. Moreover, TLM is a nonparametric model, which is able to learn the number of communities from the data. Extensive experiments on two real world corpora show TLM outperforms two state-of-the-art baseline models, which verify the effectiveness of TLM in determining the proper number of communities and generating sound communities. Guoqing Zheng, Jinwen Guo, Lichun Yang, Shengliang Xu, Shenghua Bao, Zhong Su, Dingyi Han, Yong Yu 0001 |
CIKM | 5 |
| 2009 | A study of information retrieval on accumulative social descriptions using the generation featuresabstractThis paper is concerned with the study of information retrieval (IR) on Accumulative Social Descriptions (ASDs). ASDs refer to Web texts that accumulated by many Web users describing certain Web resources, such as anchor texts, search logs and social annotations. There have been some studies working on leveraging ASDs for improving search performance in both internet and intranet. However, to the best of our knowledge, no prior study has concerned the specific generation features of ASDs, which are the focus point of this paper. Specifically, we consider the generation features from two perspectives, the generation processes and the generated distributions. Further, three probabilistic IR models are derived based on them. The three models are first demonstrated with one toy dataset and then empirically evaluated with two real datasets: an internet dataset consisting of 90,295 Web pages, with 25,845,818 social annotations crawled from Del.icio.us and 31,320,005 pieces of anchor texts crawled through Yahoo! API, and an intranet dataset consisting of 179,835 Web pages with 1,245,522 annotations dumped from the intranet tagging system in IBM, named as Dogear. Extensive experimental results show that the proposed methods, which fully leverage the generation features of ASDs, improve the performance of both internet and intranet search significantly. Lichun Yang, Shengliang Xu, Shenghua Bao, Dingyi Han, Zhong Su, Yong Yu 0001 |
CIKM | 3 |
| 2009 | sDoc: exploring social wisdom for document enhancement in web miningabstractWeb document could be seen to be composed of textual content as well as social metadata of various forms (e.g., anchor text, search query and social annotation), both of which are valuable to indicate the semantic content of the document. However, due to the free nature of the web, the two streams of web data suffer from the serious problems of noise and sparseness, which have actually become the major challenges to the success of many web mining applications. Previous work has shown that it could enhance the content of web document by integrating anchor text and search query. In this paper, we study the problem of exploring emergent social annotation for document enhancement and propose a novel reinforcement framework to generate "social representation" of document. Distinguishing from prior work, textual content and social annotation are enhanced simultaneously in our framework, which is achieved by exploiting a kind of mutual reinforcement relationship behind them. Two convergent models, social content model and social annotation model, are symmetrically derived from the framework to represent enhanced textual content and enhanced social annotation respectively. The enhanced document is referred to as Social Document or sDoc in that it could embed complementary viewpoints from many web authors and many web visitors. In this sense, the document semantics is enhanced exactly by exploring social wisdom. We build the framework on a large Del.icio.us data and evaluate it through three typical web mining applications: annotation, classification and retrieval. Experimental results demonstrate that social representation of web document could boost the performance of these applications significantly. Xiaoxun Zhang, Lichun Yang, Xian Wu 0001, Zhili Guo, Shenghua Bao, Yong Yu 0001, Zhong Su |
CIKM | 6 |
| 2009 | Joint Emotion-Topic Modeling for Social Affective Text MiningabstractThis paper is concerned with the problem of social affective text mining, which aims to discover the connections between social emotions and affective terms based on user-generated emotion labels. We propose a joint emotion-topic model by augmenting latent Dirichlet allocation with an additional layer for emotion modeling. It first generates a set of latent topics from emotions, followed by generating affective terms from each topic. Experimental results on an online news collection show that the proposed model can effectively identify meaningful latent topics for each emotion. Evaluation on emotion prediction further verifies the effectiveness of the proposed model. Shenghua Bao, Shengliang Xu, Li Zhang 0007, Zhong Su, Dingyi Han, Yong Yu 0001 |
ICDM | 1 |
| 2009 | Social Propagation: Boosting Social Annotations for Web Mining
Shenghua Bao, Bohai Yang, Ben Fei, Shengliang Xu, Zhong Su, Yong Yu 0001 |
World Wide Web | 1 |
| 2008 | A Probabilistic Model for Fine-Grained Expert Search
Shenghua Bao, Huizhong Duan, Qi Zhou 0001, Miao Xiong, Yunbo Cao, Yong Yu 0001 |
ACL | 1 |
| 2008 | Boosting social annotations using propagationabstractThis paper is concerned with the problem of boosting social annotations using propagation, which is also called social propagation. In particular, we focus on propagating social annotations of web pages (e.g., annotations in Del.icio.us). Although social annotations are developing fast, they cover only a small proportion of Web pages on the World Wide Web. To alleviate the low coverage problem, a general propagation model based on Random Surfer is proposed. Specifically, four steps are included: basic propagation, multiple-annotation propagation, multiple-link-type propagation, and constraint-guided propagation. Experimental results show that the proposed model is very effective in increasing coverage of annotations as well as preserving property of social annotations. Shenghua Bao, Bohai Yang, Ben Fei, Shengliang Xu, Zhong Su, Yong Yu 0001 |
CIKM | 1 |
| 2008 | Tapping on the potential of q&a community by recommending answer providersabstractThe rapidly increasing popularity of community-based Question Answering (cQA) services, e.g. Yahoo! Answers, Baidu Zhidao, etc. have attracted great attention from both academia and industry. Besides the basic problems, like question searching and answer finding, it should be noted that the low participation rate of users in cQA service is the crucial problem which limits its development potential. In this paper, we focus on addressing this problem by recommending answer providers, in which a question is given as a query and a ranked list of users is returned according to the likelihood of answering the question. Based on the intuitive idea for recommendation, we try to introduce topic-level model to improve heuristic term-level methods, which are treated as the baselines. The proposed approach consists of two steps: (1) discovering latent topics in the content of questions and answers as well as latent interests of users to build user profiles; (2) recommending question answerers for new arrival questions based on latent topics and term-level model. Specifically, we develop a general generative model for questions and answers in cQA, which is then altered to obtain a novel computationally tractable Bayesian network model. Experiments are carried out on a real-world data crawled from Yahoo! Answers during Jun 12 2007 to Aug 04 2007, which consists of 118510 questions, 772962 answers and 150324 users. The experimental results reveal significant improvements over the baseline methods and validate the positive influence of topic-level information. Jinwen Guo, Shengliang Xu, Shenghua Bao, Yong Yu 0001 |
CIKM | 3 |
| 2008 | SEM: Mining Spatial Events from the Web
Kaifeng Xu, Rui Li 0049, Shenghua Bao, Dingyi Han, Yong Yu 0001 |
PAKDD | 3 |
| 2008 | Exploring folksonomy for personalized searchabstractAs a social service in Web 2.0, folksonomy provides the users the ability to save and organize their bookmarks online with "social annotations" or "tags". Social annotations are high quality descriptors of the web pages' topics as well as good indicators of web users' interests. We propose a personalized search framework to utilize folksonomy for personalized search. Specifically, three properties of folksonomy, namely the categorization, keyword, and structure property, are explored. In the framework, the rank of a web page is decided not only by the term matching between the query and the web page's content but also by the topic matching between the user's interests and the web page's topics. In the evaluation, we propose an automatic evaluation framework based on folksonomy data, which is able to help lighten the common high cost in personalized search evaluations. A series of experiments are conducted using two heterogeneous data sets, one crawled from Del.icio.us and the other from Dogear. Extensive experimental results show that our personalized search approach can significantly improve the search quality. Shengliang Xu, Shenghua Bao, Ben Fei, Zhong Su, Yong Yu 0001 |
SIGIR | 2 |
| 2008 | Competitor Mining with the WebabstractThis paper is concerned with the problem of mining competitors from the Web automatically. Nowadays the fierce competition in the market necessitates every company not only to know which companies are its primary competitors, but also in which fields the company's rivals compete with itself and what its competitors' strength is in a specific competitive domain. The task of competitor mining that we address in the paper includes mining all the information such as competitors, competing fields and competitors' strength. A novel algorithm called CoMiner is proposed, which tries to conduct a Web-scale mining in a domain-independent manner. The CoMiner algorithm consists of three parts: 1) given an input entity, extracting a set of comparative candidates and then ranking them according to comparability; 2) extracting the fields in which the given entity and its competitors play against each other; 3) identifying and summarizing the competitive evidence that details the competitors' strength. As for evaluation, a prototype system implementing the CoMiner algorithm is presented. An evaluation data set consisting of 70 entities is constructed. 728 competitors and 3,640 competitive fields with 6,381 competitive evidences are discovered with the prototype. The experimental results show that the proposed algorithm is highly effective. Shenghua Bao, Rui Li 0049, Yong Yu 0001, Yunbo Cao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2007 | Using social annotations to improve language model for information retrievalabstractThis poster is concerned with the problem of exploring the use of social annotations for improving language models for information retrieval (denoted as LMIR). Two properties of social annotations, namely keyword property and structure property are studied for this aim. The keyword property improves LMIR by concatenating all the annotations of a document to generate a summary of the document. The structure property can boost LMIR further when similarity among annotations and similarity among documents are taken into consideration simultaneously. The two properties of social annotations are leveraged for the use of language modeling with a mixture model named as "Language Annotation Model" (denoted as LAM). Evaluations using del.icio.us data show that LAM outperforms the traditional LMIR approaches significantly. Shengliang Xu, Shenghua Bao, Yunbo Cao, Yong Yu 0001 |
CIKM | 2 |
| 2007 | CCRM: An Effective Algorithm for Mining Commodity Information from Threaded Chinese Customer Reviews
Huizhong Duan, Shenghua Bao, Yong Yu 0001 |
PAKDD | 2 |
| 2007 | Using Social Annotations to Smooth the Language Model for IR
Shengliang Xu, Shenghua Bao, Yong Yu 0001, Yunbo Cao |
PAKDD | 2 |
| 2007 | Optimizing web search using social annotationsabstractThis paper explores the use of social annotations to improve web search. Nowadays, many services, e.g. del.icio.us, have been developed for web users to organize and share their favorite web pages on line by using social annotations. We observe that the social annotations can benefit web search in two aspects: 1) the annotations are usually good summaries of corresponding web pages; 2) the count of annotations indicates the popularity of web pages. Two novel algorithms are proposed to incorporate the above information into page ranking: 1) SocialSimRank (SSR) calculates the similarity between social annotations and web queries; 2) SocialPageRank (SPR) captures the popularity of web pages. Preliminary experimental results show that SSR can find the latent semantic association between queries and annotations, while SPR successfully measures the quality (popularity) of a web page from the web users ’ perspective. We further evaluate the proposed methods empirically with 50 manually constructed queries and 3000 auto-generated queries on a dataset crawled from del.icio.us. Experiments show that both SSR and SPR benefit web search significantly. Shenghua Bao, Gui-Rong Xue, Xiaoyuan Wu, Yong Yu 0001, Ben Fei, Zhong Su |
WWW | 1 |
| 2007 | Towards effective browsing of large scale social annotationsabstractThis paper is concerned with the problem of browsing social annotations. Today, a lot of services (e.g., Del.icio.us, Filckr) have been provided for helping users to manage and share their favorite URLs and photos based on social annotations. Due to the exponential increasing of the social annotations, more and more users, however, are facing the problem how to effectively find desired resources from large annotation data. Existing methods such as tag cloud and annotation matching work well only on small annotation sets. Thus, an effective approach for browsing large scale annotation sets and the associated resources is in great demand by both ordinary users and service providers. In this paper, we propose a novel algorithm, namely Effective Large Scale Annotation Browser (ELSABer), to browse large-scale social annotation data. ELSABer helps the users browse huge number of annotations in a semantic, hierarchical and efficient way. More specifically, ELSABer has the following features: 1) the semantic relations between annotations are explored for browsing of similar resources; 2) the hierarchical relations between annotations are constructed for browsing in a top-down fashion; 3) the distribution of social annotations is studied for efficient browsing. By incorporating the personal and time information, ELSABer can be further extended for personalized and time-related browsing. A prototype system is implemented and shows promising results. Rui Li 0049, Shenghua Bao, Yong Yu 0001, Ben Fei, Zhong Su |
WWW | 2 |
| 2006 | Web Scale Competitor Discovery Using Mutual Information
Rui Li 0049, Shenghua Bao, Yuanjie Liu, Yong Yu 0001 |
ADMA | 2 |
| 2006 | Mining Latent Associations of Objects Using a Typed Mixture Model--A Case Study on Expert/Expertise MiningabstractThis paper studies the problem of discovering latent associations among objects in text documents. Specifically, given two sets of objects and various types of co-occurrence data concerning the objects existing in texts, we aim to discover the hidden or latent associative relationships between the two sets of objects. Existing methods are not directly applicable as they are unable to consider all this information. For example, the probabilistic mixture model called Separable Mixture Model (SMM) proposed by Hofmann can use only one type of co-occurrences to mine latent associations. This paper proposes a more general probabilistic mixture model called the Typed Separable Mixture Model (TSMM), which is able to use all types of co-occurrences within a single framework. Experimental results based on the expert/expertise mining task show that TSMM outperforms SMM significantly. Shenghua Bao, Yunbo Cao, Bing Liu 0001, Yong Yu 0001, Hang Li 0001 |
ICDM | 1 |
| 2006 | CoMiner: An Effective Algorithm for Mining Competitors from the WebabstractThis paper attempts to accomplish a novel task of mining competitive information with respect to an entity (such as a company, product, person) from the web. An algorithm called "CoMiner" is proposed, which first extracts a set of comparative candidates of the input entity and then ranks them according to the comparability, and finally extracts the competitive fields. The experimental results show that the proposed algorithm drafts a complete picture of competitive relation of a given entity effectively. Rui Li 0049, Shenghua Bao, Yong Yu 0001, Yunbo Cao |
ICDM | 2 |
| 2006 | LSM: Language Sense Model for Information Retrieval
Shenghua Bao, Lei Zhang 0007, Erdong Chen, Rui Li 0049, Yong Yu 0001 |
WAIM | 1 |