VLDB 2026 Research / reviewers in the wild / expert
Shike Mei
dblp:142/3319
· DBLP profile ↗
5ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 67% Recommender systems · 33% | |
| Artificial intelligence
2 papers |
Trustworthy machine learning · 42% Knowledge representation and reasoning · 36% Information extraction and text analysis · 11% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
adversarial attack |
0.2 | 1 | 2015 | Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015 |
Security and privacy of machine learning
poisoning attack |
0.2 | 1 | 2015 | Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015 |
Security and privacy of machine learning
training-time attack |
0.2 | 1 | 2015 | Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
first-order logic |
0.2 | 1 | 2014 | Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014 |
Information retrieval › user behavior › search behavior
click model |
0.2 | 1 | 2014 | Exploiting contextual factors for click modeling in sponsored search · WSDM 2014 |
Recommender systems
click-through rate prediction |
0.2 | 1 | 2014 | Exploiting contextual factors for click modeling in sponsored search · WSDM 2014 |
Information retrieval › online advertising
sponsored search |
0.2 | 1 | 2014 | Exploiting contextual factors for click modeling in sponsored search · WSDM 2014 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent generative model
latent topic model |
0.1 | 1 | 2014 | Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014 |
Natural language and speech › Information extraction and text analysis
topic model |
0.1 | 1 | 2014 | Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014 |
Methods — techniques the papers use, named apart from their topics
implicit function theorem · 0.4gradient methods · 0.4bi-level optimization · 0.4regularized bayesian framework · 0.2regression · 0.2first-order logic · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine LearnersabstractWe investigate a problem at the intersection of machine learning and security: training-set attacks on machine learners. In such attacks an attacker contaminates the training data so that a specific learning algorithm would produce a model profitable to the attacker. Understanding training-set attacks is important as more intelligent agents (e.g. spam filters and robots) are equipped with learning capability and can potentially be hacked via data they receive from the environment. This paper identifies the optimal training-set attack on a broad family of machine learners. First we show that optimal training-set attack can be formulated as a bilevel optimization problem. Then we show that for machine learners with certain Karush-Kuhn-Tucker conditions we can solve the bilevel problem efficiently using gradient methods on an implicit function. As examples, we demonstrate optimal training-set attacks on Support VectorMachines, logistic regression, and linear regression with extensive experiments. Finally, we discuss potential defenses against such attacks. Shike Mei, Xiaojin Zhu 0001 |
AAAI | 1 |
| 2015 | The Security of Latent Dirichlet AllocationabstractLatent Dirichlet allocation (LDA) is an increasingly popular tool for data analysis in many domains. If LDA output affects decision making (especially when money is involved), there is an incentive for attackers to compromise it. We ask the question: how can an attacker minimally poison the corpus so that LDA produces topics that the attacker wants the LDA user to see? Answering this question is important to characterize such attacks, and to develop defenses in the future. We give a novel bilevel optimization formulation to identify the optimal poisoning attack. We present an efficient solution (up to local optima) using descent method and implicit functions. We demonstrate poisoning attacks on LDA with extensive experiments, and discuss possible defenses. Shike Mei, Xiaojin Zhu 0001 |
AISTATS | 1 |
| 2014 | Inferring air pollution by sniffing social mediaabstractThe first step to deal with the significant issue of air pollution in China and elsewhere in the world is to monitor it. While more physical monitoring stations are built, current coverage is limited to large cities with most other places under-monitored. In this paper we propose a complementary approach to monitor Air Quality Index (AQI): using machine learning models to estimate AQI from social media posts. We propose a series of progressively more sophisticated machine learning models, culminating in a Markov Random Field model that utilizes the text content in social media as well as the spatiotemporal correlation among cities and days. Our extensive experiments on Sina Weibo data from 108 cities during a one-month period demonstrate the accurate AQI prediction performance of our approach. Shike Mei, Xiaojin Zhu 0001, Charles R. Dyer |
ASONAM | 1 |
| 2014 | Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian ModelsabstractMuch research in Bayesian modeling has been done to elicit a prior distribution that incorporates domain knowledge. We present a novel and more direct approach by imposing First-Order Logic (FOL) rules on the posterior distribution. Our approach unifies FOL and Bayesian modeling under the regularized Bayesian framework. In addition, our approach automatically estimates the uncertainty of FOL rules when they are produced by humans, so that reliable rules are incorporated while unreliable ones are ignored. We apply our approach to latent topic modeling tasks and demonstrate that by combining FOL knowledge and Bayesian modeling, we both improve the task performance and discover more structured latent representations in unsupervised and supervised learning. Shike Mei, Jun Zhu 0001, Jerry Zhu |
ICML | 1 |
| 2014 | Exploiting contextual factors for click modeling in sponsored searchabstractSponsored search is the primary business for today's commercial search engines. Accurate prediction of the Click-Through Rate (CTR) for ads is key to displaying relevant ads to users. In this paper, we systematically study the two kinds of contextual factors influencing the CTR: 1) In micro factors, we focus on the factors for mainline ads, including ad depth, query diversity, ad interaction. 2) In macro factors, we try to understand the correlations of clicks between organic search and sponsored search. Based on this data analysis, we propose novel click models which harvest these new explored factors. To the best of our knowledge, this is the first paper to examine and model the effects of the above contextual factors in sponsored search. Extensive experiments on large-scale real-world datasets show that by incorporating these contextual factors, our novel click models can outperform state-of-the-art methods. Dawei Yin 0001, Shike Mei, Bin Cao 0001, Jian-Tao Sun, Brian D. Davison 0001 |
WSDM | 2 |