VLDB 2026 Research / reviewers in the wild / expert
Mohammad Shafkat Amin
dblp:18/7248
· DBLP profile ↗
8ranked-venue papers
3as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 1 first-authorArtificial intelligence and machine learning · 3Software engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Transfer learning and domain adaptation · 50% Information extraction and text analysis · 50% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Natural language and speech › Information extraction and text analysis › text classification
spam detection |
0.2 | 1 | 2015 | Transfer Learning for Bilingual Content Classification · KDD 2015 |
Methods — techniques the papers use, named apart from their topics
transfer learning · 0.2machine translation · 0.2feature generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | A context-aware approach to detection of short irrelevant textsabstractThis paper presents a simple and effective framework that can detect irrelevant short text contents following blogs and news articles, etc. in a context-aware and timely fashion. Nowadays, websites such as Linkedin.com and CNN.com allow their visitors to leave comments after articles, and spammers are exploiting this feature to post irrelevant contents. Visited by millions of readers per day, these websites have extremely high visibility, and irrelevant comments have a detrimental effect on the visiting traffic and revenue of these websites. Therefore, it is critical to eliminate these irrelevant comments as accurately and early as possible. Different from traditional text mining tasks, comments following news and blog articles are characterized by briefness and context-dependent semantics, making it difficult to measure semantic relevance. What's worse, there could be only a handful of comments soon after an article is posted, leading to a severe lack of information for semantics and relevance measurement. We propose to infer “context-aware semantics” to address the above challenges in a unified framework. Specifically, we construct contexts for comments using either blocks of surrounding comments, or comments collected via a principled transfer learning approach. The constructed contexts mitigate the sparseness and sharply define context-dependent semantics of comments, even at the early stage of commenting activities, allowing traditional dimension reduction methods to better capture the semantics of short texts in a context-aware way. We confirm the effectiveness of the proposed method on two real world datasets consisting of news and blog articles and comments, with a maximal improvement of 20% in Area Under Precision-Recall Curve. Sihong Xie, Jing Wang 0102, Mohammad Shafkat Amin, Baoshi Yan, Anmol Bhasin, Clement T. Yu, Philip S. Yu |
DSAA | 3 |
| 2015 | Transfer Learning for Bilingual Content ClassificationabstractLinkedIn Groups provide a platform on which professionals with similar background, target and specialities can share content, take part in discussions and establish opinions on industry topics. As in most online social communities, spam content in LinkedIn Groups poses great challenges to the user experience and could eventually lead to substantial loss of active users. Building an intelligent and scalable spam detection system is highly desirable but faces difficulties such as lack of labeled training data, particularly for languages other than English. In this paper, we take the spam (Spanish) job posting detection as the target problem and build a generic machine learning pipeline for multi-lingual spam detection. The main components are feature generation and knowledge migration via transfer learning. Specifically, in the feature generation phase, a relatively large labeled data set is generated via machine translation. Together with a large set of unlabeled human written Spanish data, unigram features are generated based on the frequency. In the second phase, machine translated data are properly reweighted to capture the discrepancy from human written ones and classifiers can be built on top of them. To make effective use of a small portion of labeled data available in human written Spanish, an adaptive transfer learning algorithm is proposed to further improve the performance. We evaluate the proposed method on LinkedIn's production data and the promising results verify the efficacy of our proposed algorithm. The pipeline is ready for production. Qian Sun 0002, Mohammad Shafkat Amin, Baoshi Yan, Craig Martell, Vita Markman, Anmol Bhasin, Jieping Ye |
KDD | 2 |
| 2013 | Generating supplemental content information using virtual profilesabstractWe describe a hybrid recommendation system at LinkedIn that seeks to optimally extract relevant information pertaining to items to be recommended. By extending the notion of an item profile, we propose the concept of a "virtual profile" that augments the content of the item with rich set of features inherited from members who have already shown explicit interest in it. Unlike item-based collaborative filtering, we focus on discovering the characteristic descriptors that underlie the item-user association. Such information is used as supplemental features in a content-based filtering system. The main objective of virtual profiles is to provide a means to tap into rich-content information from one type of entity and propagate features extracted from which to other affiliated entities that may suffer from relative data scarcity. We empirically evaluate the proposed method on a real-world community recommendation problem at LinkedIn. The result shows that the virtual profiles outperform a collaborative filtering based approach (user who likes this also likes that). In particular, the improvement is more significant for new users with only limited connections, demonstrating the capability of the method to address the cold-start problem in pure collaborative filtering systems. Haishan Liu, Mohammad Shafkat Amin, Baoshi Yan, Anmol Bhasin |
RecSys | 2 |
| 2012 | Social referral: leveraging network connections to deliver recommendationsabstractMuch work has been done to study the interplay between recommender systems and social networks. This creates a very powerful coupling in presenting highly relevant recommendations to the users. However, to our knowledge, little attention has been paid to leverage a user's social network to deliver these recommendations. We present a novel approach to aid delivery of recommendations using the recipient's friends or connections. Our contributions with this study are 1) A novel recommendation delivery paradigm called Social Referral, which utilizes a user's social network for the delivery of relevant content. 2) An implementation of the paradigm is described in a real industrial production setting of a large online professional network. 3) A study of the interaction between the trifecta of the recommender system, the trusted connections and the end consumer of the recommendation by comparing and contrasting the proposed approach's performance with the direct recommender system. Mohammad Shafkat Amin, Baoshi Yan, Sripad Sriram, Anmol Bhasin, Christian Posse |
RecSys | 1 |
| 2012 | Top-k Similar Graph Matching Using TraM in Biological NetworksabstractMany emerging database applications entail sophisticated graph-based query manipulation, predominantly evident in large-scale scientific applications. To access the information embedded in graphs, efficient graph matching tools and algorithms have become of prime importance. Although the prohibitively expensive time complexity associated with exact subgraph isomorphism techniques has limited its efficacy in the application domain, approximate yet efficient graph matching techniques have received much attention due to their pragmatic applicability. Since public domain databases are noisy and incomplete in nature, inexact graph matching techniques have proven to be more promising in terms of inferring knowledge from numerous structural data repositories. In this paper, we propose a novel technique called TraM for approximate graph matching that off-loads a significant amount of its processing on to the database making the approach viable for large graphs. Moreover, the vector space embedding of the graphs and efficient filtration of the search space enables computation of approximate graph similarity at a throw-away cost. We annotate nodes of the query graphs by means of their global topological properties and compare them with neighborhood biased segments of the datagraph for proper matches. We have conducted experiments on several real data sets, and have demonstrated the effectiveness and efficiency of the proposed method Mohammad Shafkat Amin, Russell L. Finley Jr., Hasan M. Jamil |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2010 | An Efficient Web-Based Wrapper and Annotator for Tabular DataabstractIn the last few years, several works in the literature have addressed the problem of data extraction from web pages. The importance of this problem derives from the fact that, once extracted, data can be handled in a way similar to instances of a traditional database, which in turn can facilitate application of web data integration and various other domain specific problems. In this paper, we propose a novel table extraction technique that works on web pages generated dynamically from a back-end database. The proposed system can automatically discover table structure by relevant pattern mining from web pages in an efficient way, and can generate regular expression for the extraction process. Moreover, the proposed system can assign intuitive column names to the columns of the extracted table by leveraging Wikipedia knowledge base for the purpose of table annotation. To improve accuracy of the assignment, we exploit the structural homogeneity of the column values and their co-location information to weed out less likely candidates. This approach requires no human intervention and experimental results have shown its accuracy to be promising. Moreover, the wrapper generation algorithm works in linear time. Mohammad Shafkat Amin, Hasan M. Jamil |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2009 | On-the-Fly Integration and Ad Hoc Querying of Life Sciences Databases Using LifeDB
Anupam Bhattacharjee, Aminul Islam 0004, Mohammad Shafkat Amin, Shahriyar Hossain, Shazzad Hosain, Hasan M. Jamil, Leonard Lipovich |
DEXA | 3 |
| 2009 | A Model for Contextual Cooperative Query Answering in E-Commerce Applications
Kazi Zakia Sultana, Anupam Bhattacharjee, Mohammad Shafkat Amin, Hasan M. Jamil |
FQAS | 3 |