Shike Mei

dblp:142/3319 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 67% Recommender systems · 33%
Artificial intelligence
2 papers
Trustworthy machine learning · 42% Knowledge representation and reasoning · 36% Information extraction and text analysis · 11%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.212015
Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015
Security and privacy of machine learning
poisoning attack
0.212015
Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015
Security and privacy of machine learning
training-time attack
0.212015
Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners · AAAI 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › logic-based reasoning
first-order logic
0.212014
Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014
Information retrieval › user behavior › search behavior
click model
0.212014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Recommender systems
click-through rate prediction
0.212014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Information retrieval › online advertising
sponsored search
0.212014
Exploiting contextual factors for click modeling in sponsored search · WSDM 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent generative model
latent topic model
0.112014
Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014
Natural language and speech › Information extraction and text analysis
topic model
0.112014
Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models · ICML 2014

Methods — techniques the papers use, named apart from their topics

implicit function theorem · 0.4gradient methods · 0.4bi-level optimization · 0.4regularized bayesian framework · 0.2regression · 0.2first-order logic · 0.2
YearPublicationVenuePosition
2015 Using Machine Teaching to Identify Optimal Training-Set Attacks on Machine Learners
abstract
We investigate a problem at the intersection of machine learning and security: training-set attacks on machine learners. In such attacks an attacker contaminates the training data so that a specific learning algorithm would produce a model profitable to the attacker. Understanding training-set attacks is important as more intelligent agents (e.g. spam filters and robots) are equipped with learning capability and can potentially be hacked via data they receive from the environment. This paper identifies the optimal training-set attack on a broad family of machine learners. First we show that optimal training-set attack can be formulated as a bilevel optimization problem. Then we show that for machine learners with certain Karush-Kuhn-Tucker conditions we can solve the bilevel problem efficiently using gradient methods on an implicit function. As examples, we demonstrate optimal training-set attacks on Support VectorMachines, logistic regression, and linear regression with extensive experiments. Finally, we discuss potential defenses against such attacks.
Shike Mei, Xiaojin Zhu 0001
AAAI1
2015 The Security of Latent Dirichlet Allocation
abstract
Latent Dirichlet allocation (LDA) is an increasingly popular tool for data analysis in many domains. If LDA output affects decision making (especially when money is involved), there is an incentive for attackers to compromise it. We ask the question: how can an attacker minimally poison the corpus so that LDA produces topics that the attacker wants the LDA user to see? Answering this question is important to characterize such attacks, and to develop defenses in the future. We give a novel bilevel optimization formulation to identify the optimal poisoning attack. We present an efficient solution (up to local optima) using descent method and implicit functions. We demonstrate poisoning attacks on LDA with extensive experiments, and discuss possible defenses.
Shike Mei, Xiaojin Zhu 0001
AISTATS1
2014 Inferring air pollution by sniffing social media
abstract
The first step to deal with the significant issue of air pollution in China and elsewhere in the world is to monitor it. While more physical monitoring stations are built, current coverage is limited to large cities with most other places under-monitored. In this paper we propose a complementary approach to monitor Air Quality Index (AQI): using machine learning models to estimate AQI from social media posts. We propose a series of progressively more sophisticated machine learning models, culminating in a Markov Random Field model that utilizes the text content in social media as well as the spatiotemporal correlation among cities and days. Our extensive experiments on Sina Weibo data from 108 cities during a one-month period demonstrate the accurate AQI prediction performance of our approach.
Shike Mei, Xiaojin Zhu 0001, Charles R. Dyer
ASONAM1
2014 Robust RegBayes: Selectively Incorporating First-Order Logic Domain Knowledge into Bayesian Models
abstract
Much research in Bayesian modeling has been done to elicit a prior distribution that incorporates domain knowledge. We present a novel and more direct approach by imposing First-Order Logic (FOL) rules on the posterior distribution. Our approach unifies FOL and Bayesian modeling under the regularized Bayesian framework. In addition, our approach automatically estimates the uncertainty of FOL rules when they are produced by humans, so that reliable rules are incorporated while unreliable ones are ignored. We apply our approach to latent topic modeling tasks and demonstrate that by combining FOL knowledge and Bayesian modeling, we both improve the task performance and discover more structured latent representations in unsupervised and supervised learning.
Shike Mei, Jun Zhu 0001, Jerry Zhu
ICML1
2014 Exploiting contextual factors for click modeling in sponsored search
abstract
Sponsored search is the primary business for today's commercial search engines. Accurate prediction of the Click-Through Rate (CTR) for ads is key to displaying relevant ads to users. In this paper, we systematically study the two kinds of contextual factors influencing the CTR: 1) In micro factors, we focus on the factors for mainline ads, including ad depth, query diversity, ad interaction. 2) In macro factors, we try to understand the correlations of clicks between organic search and sponsored search. Based on this data analysis, we propose novel click models which harvest these new explored factors. To the best of our knowledge, this is the first paper to examine and model the effects of the above contextual factors in sponsored search. Extensive experiments on large-scale real-world datasets show that by incorporating these contextual factors, our novel click models can outperform state-of-the-art methods.
Dawei Yin 0001, Shike Mei, Bin Cao 0001, Jian-Tao Sun, Brian D. Davison 0001
WSDM2