VLDB 2026 Research / reviewers in the wild / expert
Baoshi Yan
dblp:51/1876
· DBLP profile ↗
9ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 3 first-authorArtificial intelligence and machine learning · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 50% Information extraction and text analysis · 50% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Natural language and speech › Information extraction and text analysis › text classification
spam detection |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.2machine translation · 0.2feature generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A context-aware approach to detection of short irrelevant textsabstractThis paper presents a simple and effective framework that can detect irrelevant short text contents following blogs and news articles, etc. in a context-aware and timely fashion. Nowadays, websites such as Linkedin.com and CNN.com allow their visitors to leave comments after articles, and spammers are exploiting this feature to post irrelevant contents. Visited by millions of readers per day, these websites have extremely high visibility, and irrelevant comments have a detrimental effect on the visiting traffic and revenue of these websites. Therefore, it is critical to eliminate these irrelevant comments as accurately and early as possible. Different from traditional text mining tasks, comments following news and blog articles are characterized by briefness and context-dependent semantics, making it difficult to measure semantic relevance. What's worse, there could be only a handful of comments soon after an article is posted, leading to a severe lack of information for semantics and relevance measurement. We propose to infer “context-aware semantics” to address the above challenges in a unified framework. Specifically, we construct contexts for comments using either blocks of surrounding comments, or comments collected via a principled transfer learning approach. The constructed contexts mitigate the sparseness and sharply define context-dependent semantics of comments, even at the early stage of commenting activities, allowing traditional dimension reduction methods to better capture the semantics of short texts in a context-aware way. We confirm the effectiveness of the proposed method on two real world datasets consisting of news and blog articles and comments, with a maximal improvement of 20% in Area Under Precision-Recall Curve. Sihong Xie, Jing Wang 0102, Mohammad Shafkat Amin, Baoshi Yan, Anmol Bhasin, Clement T. Yu, Philip S. Yu |
DSAA | 4 |
| 2015 | Transfer Learning for Bilingual Content ClassificationabstractLinkedIn Groups provide a platform on which professionals with similar background, target and specialities can share content, take part in discussions and establish opinions on industry topics. As in most online social communities, spam content in LinkedIn Groups poses great challenges to the user experience and could eventually lead to substantial loss of active users. Building an intelligent and scalable spam detection system is highly desirable but faces difficulties such as lack of labeled training data, particularly for languages other than English. In this paper, we take the spam (Spanish) job posting detection as the target problem and build a generic machine learning pipeline for multi-lingual spam detection. The main components are feature generation and knowledge migration via transfer learning. Specifically, in the feature generation phase, a relatively large labeled data set is generated via machine translation. Together with a large set of unlabeled human written Spanish data, unigram features are generated based on the frequency. In the second phase, machine translated data are properly reweighted to capture the discrepancy from human written ones and classifiers can be built on top of them. To make effective use of a small portion of labeled data available in human written Spanish, an adaptive transfer learning algorithm is proposed to further improve the performance. We evaluate the proposed method on LinkedIn's production data and the promising results verify the efficacy of our proposed algorithm. The pipeline is ready for production. Qian Sun 0002, Mohammad Shafkat Amin, Baoshi Yan, Craig Martell, Vita Markman, Anmol Bhasin, Jieping Ye |
KDD | 3 |
| 2013 | Generating supplemental content information using virtual profilesabstractWe describe a hybrid recommendation system at LinkedIn that seeks to optimally extract relevant information pertaining to items to be recommended. By extending the notion of an item profile, we propose the concept of a "virtual profile" that augments the content of the item with rich set of features inherited from members who have already shown explicit interest in it. Unlike item-based collaborative filtering, we focus on discovering the characteristic descriptors that underlie the item-user association. Such information is used as supplemental features in a content-based filtering system. The main objective of virtual profiles is to provide a means to tap into rich-content information from one type of entity and propagate features extracted from which to other affiliated entities that may suffer from relative data scarcity. We empirically evaluate the proposed method on a real-world community recommendation problem at LinkedIn. The result shows that the virtual profiles outperform a collaborative filtering based approach (user who likes this also likes that). In particular, the improvement is more significant for new users with only limited connections, demonstrating the capability of the method to address the cold-start problem in pure collaborative filtering systems. Haishan Liu, Mohammad Shafkat Amin, Baoshi Yan, Anmol Bhasin |
RecSys | 3 |
| 2013 | Pairwise learning in recommendation: experiments with community recommendation on linkedinabstractMany online systems present a list of recommendations and infer user interests implicitly from clicks or other contextual actions. For modeling user feedback in such settings, a common approach is to consider items acted upon to be relevant to the user, and irrelevant otherwise. However, clicking some but not others conveys an implicit ordering of the presented items. Pairwise learning, which leverages such implicit ordering between a pair of items, has been successful in areas such as search ranking. In this work, we study whether pairwise learning can improve community recommendation. We first present two novel pairwise models adapted from logistic regression. Both offline and online experiments in a large real-world setting show that incorporating pairwise learning improves the recommendation performance. However, the improvement is only slight. We find that users' preferences regarding the kinds of communities they like can differ greatly, which adversely affect the effectiveness of features derived from pairwise comparisons. We therefore propose a probabilistic latent semantic indexing model for pairwise learning (Pairwise PLSI), which assumes a set of users' latent preferences between pairs of items. Our experiments show favorable results for the Pairwise PLSI model and point to the potential of using pairwise learning for community recommendation. Amit Sharma 0007, Baoshi Yan |
RecSys | 2 |
| 2012 | Social referral: leveraging network connections to deliver recommendationsabstractMuch work has been done to study the interplay between recommender systems and social networks. This creates a very powerful coupling in presenting highly relevant recommendations to the users. However, to our knowledge, little attention has been paid to leverage a user's social network to deliver these recommendations. We present a novel approach to aid delivery of recommendations using the recipient's friends or connections. Our contributions with this study are 1) A novel recommendation delivery paradigm called Social Referral, which utilizes a user's social network for the delivery of relevant content. 2) An implementation of the paradigm is described in a real industrial production setting of a large online professional network. 3) A study of the interaction between the trifecta of the recommender system, the trusted connections and the end consumer of the recommendation by comparing and contrasting the proposed approach's performance with the direct recommender system. Mohammad Shafkat Amin, Baoshi Yan, Sripad Sriram, Anmol Bhasin, Christian Posse |
RecSys | 2 |
| 2011 | Entity Resolution Using Social Graphs for Business ApplicationsabstractSocial network such as Linked In maintains profiles for its members in a semi-structured format. A lot of business applications like ad targeting and content recommendations rely on canonicalization of data elements like companies, titles and schools for enabling fine grained advertising or recommending candidates for job postings. In this paper we explore the issues around resolving company names for hundreds of millions of member positions to known company entities using the social graph. We proposed a machine learning approach leveraging three dimensional feature sets including the social graph, social behavior and various content and demographic features. The experiments showed that our approach achieved high precision at a reasonable coverage and is significantly superior to a baseline content based approach. Baoshi Yan, Lokesh Bajaj, Anmol Bhasin |
ASONAM | 1 |
| 2005 | Aligning Class Hierarchies with Grass-Roots Class AlignmentabstractThe performance of an ontology alignment technique largely depends on the amount of information that can be leveraged for the alignment task. On the semantic Web, end-users may explicitly or implicitly generate ontology alignments during their use of the semantic data. This kind of end-user-generated ontology alignment, which we call grass-roots ontology alignment, is an important source of information that is yet to be taken into account by current ontology alignment techniques. Grass-roots ontology alignment, often generated as a side effect of other data manipulations, could be user-specific, task-specific, approximate, or even contradictory. This paper reports our work on reusing grass-roots class alignment for aligning class hierarchies. A grass-roots class alignment, though approximate, still reveals some facts about relationships between different classes. We formalize facts about class relationships that can be inferred from an alignment under different cases. We then apply forward-chaining inference to the facts knowledge base to infer more facts. The facts KB is then leveraged for ontology alignment purposes. To deal with uncertainty and inconsistency, each fact is associated with an evidence that tells how the fact is obtained. The evidences are used to select better-supported facts in case of inconsistency. Baoshi Yan |
Web Intelligence | 1 |
| 2004 | A subscribable peer-to-peer RDF repository for distributed metadata management
Min Cai, Martin R. Frank, Baoshi Yan, Robert M. MacGregor |
J. Web Semant. | 3 |
| 2003 | WebScripter: Grass-Roots Ontology Alignment via End-User Report Creation
Baoshi Yan, Martin R. Frank, Pedro A. Szekely, Robert Neches, Juan Lopez |
ISWC | 1 |