VLDB 2026 Research / reviewers in the wild / expert
Yan Zhou 0001
dblp:60/5157-1
· DBLP profile ↗
19ranked-venue papers
9as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 8 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Security and privacy · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Using AI Uncertainty Quantification to Improve Human Decision-MakingabstractAI Uncertainty Quantification (UQ) has the potential to improve human decision-making beyond AI predictions alone by providing additional probabilistic information to users. The majority of past research on AI and human decision-making has concentrated on model explainability and interpretability, with little focus on understanding the potential impact of UQ on human decision-making. We evaluated the impact on human decision-making for instance-level UQ, calibrated using a strict scoring rule, in two online behavioral experiments. In the first experiment, our results showed that UQ was beneficial for decision-making performance compared to only AI predictions. In the second experiment, we found UQ had generalizable benefits for decision-making across a variety of representations for probabilistic information. These results indicate that implementing high quality, instance-level UQ for AI may improve decision-making with real systems compared to AI predictions alone. Laura Marusich, Jonathan Z. Bakdash, Yan Zhou 0001, Murat Kantarcioglu |
ICML | 3 |
| 2023 | Attack Some while Protecting Others: Selective Attack Strategies for Attacking and Protecting Multiple ConceptsabstractMachine learning models are vulnerable to adversarial attacks. Existing research focuses on attack-only scenarios. In practice, one dataset may be used for learning different concepts, and the attacker may be incentivized to attack some concepts but protect the others. For example, the attacker might tamper a profile image for the "age'' model to predict "young'', while the "attractiveness'' model still predicts "pretty''. In this work, we empirically demonstrate that attacking the classifier for one learning task may negatively impact classifiers learning other tasks on the same data. This raises an interesting research question: is it possible to attack one set of classifiers while protecting the others trained on the same data? Vibha Belavadi, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
CCS | 2 |
| 2023 | On Improving Fairness of AI Models with Synthetic Minority Oversampling TechniquesabstractBiased AI models result in unfair decisions. In response, a number of algorithmic solutions have been engineered to mitigate bias, among which the Synthetic Minority Oversampling Technique (SMOTE) has been studied, to an extent. Although the SMOTE technique and its variants have great potentials to help improve fairness, there is little theoretical justification for its success. In addition, formal error and fairness bounds are not clearly given. This paper attempts to address both issues. We prove and demonstrate that synthetic data generated by oversampling underrepresented groups can mitigate algorithmic bias in AI models, while keeping the predictive errors bounded. We further compare this technique to the existing state-of-the-art fair AI techniques on five datasets using a variety of fairness metrics. We show that this approach can effectively improve fairness even when there is a significant amount of label and selection bias, regardless of the baseline AI algorithm. Yan Zhou 0001, Murat Kantarcioglu, Chris Clifton |
SDM | 1 |
| 2023 | Exploring the Effect of Randomness on Transferability of Adversarial Samples Against Deep Neural NetworksabstractWe investigate the transferability of adversarial attacks against deep neural networks (DNNs)—the contagion effect of adversarial attacks that, once deceiving one DNN model, can easily deceive other DNN models built on similar data. We demonstrate that introducing randomness to DNN models can break the curse of the transferability of adversarial attacks, given that the adversary does not have an unlimited attack budget. Two randomization schemes are explored: 1.) a random selection—single or ensemble—from a set of DNNs is surprisingly more robust against the strongest form of complete-knowledge attacks (a.k.a, white box attacks); 2.) after a small Gaussian random noise is added to its learned weights, a DNN model can potentially increase its resilience to adversarial attacks by as much as 74.2%. We compare the two randomization techniques to the Ensemble Adversarial Training technique and show that our randomization techniques are superior under different attack budget constraints. Furthermore, we explore the relationship between attack severity and decision boundary robustness in the version space. Finally, we connect the dots between the effectiveness of randomization to prevent attack transferability and the variability of DNN models through analyzing the differential entropy of sample hypotheses in the hypothesis space. Yan Zhou 0001, Murat Kantarcioglu, Bowei Xi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Unfair AI: It Isn't Just Biased DataabstractConventional wisdom holds that discrimination in machine learning is a result of historical discrimination: biased training data leads to biased models. We show that the reality is more nuanced; machine learning can be expected to induce types of bias not found in the training data. In particular, if different groups have different optimal models, and the optimal model for one group has higher accuracy, the optimal accuracy joint model will induce disparate impact even when the training data does not display disparate impact. We argue that due to systemic bias, this is a likely situation, and simply ensuring training data appears unbiased is insufficient to ensure fair machine learning. Chowdhury Mohammad Rakin Haider, Chris Clifton, Yan Zhou 0001 |
ICDM | 3 |
| 2021 | Does Explainable Artificial Intelligence Improve Human Decision-Making?abstractExplainable AI provides insights to users into the why for model predictions, offering potential for users to better understand and trust a model, and to recognize and correct AI predictions that are incorrect. Prior research on human and explainable AI interactions has focused on measures such as interpretability, trust, and usability of the explanation. There are mixed findings whether explainable AI can improve actual human decision-making and the ability to identify the problems with the underlying model. Using real datasets, we compare objective human decision accuracy without AI (control), with an AI prediction (no explanation), and AI prediction with explanation. We find providing any kind of AI prediction tends to improve user decision accuracy, but no conclusive evidence that explainable AI has a meaningful impact. Moreover, we observed the strongest predictor for human decision accuracy was AI accuracy and that users were somewhat able to detect when the AI was correct vs. incorrect, but this was not significantly affected by including an explanation. Our results indicate that, at least in some situations, the why information provided in explainable AI may not enhance user decision-making, and further research may be needed to understand how to integrate explainable AI into real systems. Yasmeen Alufaisan, Laura Marusich, Jonathan Z. Bakdash, Yan Zhou 0001, Murat Kantarcioglu |
AAAI | 4 |
| 2021 | Robust Transparency Against Model Inversion AttacksabstractTransparency has become a critical need in machine learning (ML) applications. Designing transparent ML models helps increase trust, ensure accountability, and scrutinize fairness. Some organizations may opt-out of transparency to protect individuals’ privacy. Therefore, there is a great demand for transparency models that consider both privacy and security risks. Such transparency models can motivate organizations to improve their credibility by making the ML-based decision-making process comprehensible to end-users. Differential privacy (DP) provides an important technique to disclose information while protecting individual privacy. However, it has been shown that DP alone cannot prevent certain types of privacy attacks against disclosed ML models. DP with low$\epsilon$values can provide high privacy guarantees, but may result in significantly weaker ML models in terms of accuracy. On the other hand, setting$\epsilon$value too high may lead to successful privacy attacks. This raises the question whether we can disclose accurate transparent ML models while preserving privacy. In this article we introduce a novel technique that complements DP to ensure model transparency and accuracy while being robust against model inversion attacks. We show that combining the proposed technique with DP provide highly transparent and accurate ML models while preserving privacy against model inversion attacks. Yasmeen Alufaisan, Murat Kantarcioglu, Yan Zhou 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | Data Mining with Algorithmic Transparency
Yan Zhou 0001, Yasmeen Alufaisan, Murat Kantarcioglu |
PAKDD (1) | 1 |
| 2018 | Privacy Preserving Synthetic Data Release Using Deep Learning
Nazmiye Ceren Abay, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham, Latanya Sweeney |
ECML/PKDD (1) | 2 |
| 2017 | Hacking social network data miningabstractOver the years social network data has been mined to predict individuals' traits such as intelligence and sexual orientation. While mining social network data can provide many beneficial services to the user such as personalized experiences, it can also harm the user when used in making critical decisions such as employment. In this work, we investigate the reliability of applying data mining techniques on social network data to predict various individual traits. In spite of the preliminary success of such data mining applications, in this paper, we demonstrate the vulnerabilities of existing state of the art social network data mining techniques when they are facing malicious attacks. Our results indicate that making critical decisions, such as employment or credit approval, based solely on social network data mining results is still premature at this stage. Specifically, we explore Facebook likes data for predicting the traits of a Facebook user, including their political views and sexual orientation. We perform several types of malicious attacks on the predictive models to measure and understand their potential vulnerabilities. We find that existing predictive models built on social network data can be easily manipulated and suggest some countermeasures to prevent some of the proposed attacks. Yasmeen Alufaisan, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ISI | 2 |
| 2016 | Modeling Adversarial Learning as Nested Stackelberg Games
Yan Zhou 0001, Murat Kantarcioglu |
PAKDD (2) | 1 |
| 2014 | Shingled Graph Disassembly: Finding the Undecideable Path
Richard Wartell, Yan Zhou 0001, Kevin W. Hamlen, Murat Kantarcioglu |
PAKDD (1) | 2 |
| 2014 | Adversarial Learning with Bayesian Hierarchical Mixtures of ExpertsabstractMany data mining applications operate in adversarial environment, for example, webpage ranking in the presence of web spam. A growing number of adversarial data mining techniques are recently developed, providing robust solutions under specific defense-attack models. Existing techniques are tied to distributional assumptions geared towards minimizing the undesirable impact of given attack models. However, the large variety of attack strategies renders the adversarial learning problem multimodal. Therefore, it calls for a more flexible modeling ideology for equivocal input. In this paper we present a Bayesian hierarchical mixtures of experts for adversarial learning. The technique groups data into soft partitions and fits simple function approximators, referred to as “experts”, within each. Experts are ranked using gating functions for each input. Ambiguous input is predicted competitively by multiple experts, while unambiguous input is effectively predicted by a single expert. Optimal attacks minimizing the likelihood of malicious data are modeled interactively at both expert and gating levels in the learning hierarchy. We demonstrate that our adversarial hierarchical-mixtures-of-experts learning model is robust against adversarial attacks on both artificial and real data. Yan Zhou 0001, Murat Kantarcioglu |
SDM | 1 |
| 2012 | Randomizing Smartphone Malware Profiles against Statistical Mining Techniques
Abhijith Shastry, Murat Kantarcioglu, Yan Zhou 0001, Bhavani Thuraisingham |
DBSec | 3 |
| 2012 | Self-Training with Selection-by-RejectionabstractPractical machine learning and data mining problems often face shortage of labeled training data. Self-training algorithms are among the earliest attempts of using unlabeled data to enhance learning. Traditional self-training algorithms label unlabeled data on which classifiers trained on limited training data have the highest confidence. In this paper, a self-training algorithm that decreases the disagreement region of hypotheses is presented. The algorithm supplements the training set with self-labeled instances. Only instances that greatly reduce the disagreement region of hypotheses are labeled and added to the training set. Empirical results demonstrate that the proposed self-training algorithm can effectively improve classification performance. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ICDM | 1 |
| 2012 | Sparse Bayesian Adversarial Learning Using Relevance Vector Machine EnsemblesabstractData mining tasks are made more complicated when adversaries attack by modifying malicious data to evade detection. The main challenge lies in finding a robust learning model that is insensitive to unpredictable malicious data distribution. In this paper, we present a sparse relevance vector machine ensemble for adversarial learning. The novelty of our work is the use of individualized kernel parameters to model potential adversarial attacks during model training. We allow the kernel parameters to drift in the direction that minimizes the likelihood of the positive data. This step is interleaved with learning the weights and the weight priors of a relevance vector machine. Our empirical results demonstrate that an ensemble of such relevance vector machine models is more robust to adversarial attacks. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ICDM | 1 |
| 2012 | Adversarial support vector machine learningabstractMany learning tasks such as spam filtering and credit card fraud detection face an active adversary that tries to avoid detection. For learning problems that deal with an active adversary, it is important to model the adversary's attack strategy and develop robust learning models to mitigate the attack. These are the two objectives of this paper. We consider two attack models: a free-range attack model that permits arbitrary data corruption and a restrained attack model that anticipates more realistic attacks that a reasonable adversary would devise under penalties. We then develop optimal SVM learning strategies against the two attack models. The learning algorithms minimize the hinge loss while assuming the adversary is modifying data to maximize the loss. Experiments are performed on both artificial and real data sets. We demonstrate that optimal solutions may be overly pessimistic when the actual attacks are much weaker than expected. More important, we demonstrate that it is possible to develop a much more resilient SVM learning model while making loose assumptions on the data corruption models. When derived under the restrained attack model, our optimal SVM learning strategy provides more robust overall performance under a wide range of attack parameters. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham, Bowei Xi |
KDD | 1 |
| 2011 | Compression for Anti-Adversarial Learning
Yan Zhou 0001, W. Meador Inge, Murat Kantarcioglu |
PAKDD (2) | 1 |
| 2011 | Differentiating Code from Data in x86 Binaries
Richard Wartell, Yan Zhou 0001, Kevin W. Hamlen, Murat Kantarcioglu, Bhavani Thuraisingham |
ECML/PKDD (3) | 2 |