Skyler Speakman

dblp:140/9492 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
9since 2021 · last 2024
0000-0003-0337-2312ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Racial and Neighborhood Disparities in Legal Financial Obligations in Jefferson County, Alabama
abstract
Legal financial obligations (LFOs) such as court fees and fines are commonly levied on individuals who are convicted of crimes. It is expected that LFO amounts should be similar across social, racial, and geographic subpopulations convicted of the same crime. This work analyzes the distribution of LFOs in Jefferson County, Alabama and highlights disparities across different individual and neighborhood demographic characteristics. Data-driven discovery methods are used to detect subpopulations that experience higher LFOs than the overall population of offenders. Critically, these discovery methods do not rely on pre-specified groups and can assist scientists and researchers investigate socially-sensitive hypotheses in a disciplined way. Some findings, such as individuals who are Black, live in Black-majority neighborhoods, or live in low-income neighborhoods tending to experience higher LFOs, are commensurate with prior expectation. However others, such as high LFO amounts in worthless instrument (bad check) cases experienced disproportionately by individuals living in affluent majority-white neighborhoods, are more surprising. More broadly than the specific findings, the methodology is shown to identify structural weaknesses that undermine the goal of equal justice under law that can be addressed through policy interventions.
Óscar Lara Yejas, Aakanksha Joshi, Andrew Martinez, Leah Nelson, Skyler Speakman, Krysten Thompson, Yuki Nishimura, Jordan Bond, Kush R. Varshney
AIES (1)5
2023 Spatially Constrained Adversarial Attack Detection and Localization in the Representation Space of Optical Flow Networks
abstract
Optical flow estimation have shown significant improvements with advances in deep neural networks. However, these flow networks have recently been shown to be vulnerable to patch-based adversarial attacks, which poses security risks in real-world applications, such as self-driving cars and robotics. We propose SADL, a Spatially constrained adversarial Attack Detection and Localization framework, to detect and localize these patch-based attack without requiring a dedicated training. The detection of an attacked input sequence is performed via iterative optimization on the features from the inner layers of flow networks, without any prior knowledge of the attacks. The novel spatially constrained optimization ensures that the detected anomalous subset of features comes from a local region. To this end, SADL provides a subset of nodes within a spatial neighborhood that contribute more to the detection, which will be utilized to localize the attack in the input sequence. The proposed SADL is validated across multiple datasets and flow networks. With patch attacks 4.8% of the size of the input image resolution on RAFT, our method successfully detects and localizes them with an average precision of 0.946 and 0.951 for KITTI-2015 and MPI-Sintel datasets, respectively. The results show that SADL consistently achieves higher detection rates than existing methods and provides new localization capabilities.
Hannah Kim 0002, Celia Cintas, Girmaw Abebe, Skyler Speakman
IJCAI4
2022 Principled Subpopulation Analysis of the BetterBirth Study and the Impact of WHO's Safe Childbirth Checklist Intervention
Girmaw Abebe, Megan Marx Delaney, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Katherine Semrau, Skyler Speakman
AMIA7
2022 Model-free feature selection to facilitate automatic discovery of divergent subgroups in tabular data
abstract
Data-centric AI encourages the need for cleaning, evaluating, and understanding data in order to achieve trustworthy AI. Existing technologies, such as AutoML, make it easier to design and train models automatically, but there is a lack of a similar level of capability to extract data-centric insights. Manual stratification of tabular data per a given feature of interest (e.g., gender) is limited to scaling up for higher feature dimension, which could be addressed using automatic discovery of divergent/anomalous subgroups. Nonetheless, these automatic discovery techniques often search across potentially exponential combinations of features which could be simplified using a preceding feature selection step. Existing feature selection techniques for tabular data often involve fitting a particular model (e.g., XGBoost) in order to select important features. However, such model-based selection is prone to model-bias and spurious correlations in addition to requiring extra resources to design, fine-tune and train a model. In this paper, we propose a model-free and sparsity-based automatic feature selection (SAFS) framework to facilitate automatic discovery of divergent subgroups. Different to filter-based selection techniques, we exploit the sparsity of objective measures among feature values to rank and select features. We validated SAFS across two publicly available datasets (MIMIC-III and Allstate Claims) and compared it with six existing feature selection methods. SAFS achieves a reduction of the feature selection time by a factor of 81× and 104×, averaged cross the existing methods in the MIMIC-III and Claims datasets, respectively. SAFS-selected features are also shown to achieve competitive detection performance, e.g., 18.3% of features selected by SAFS detected similar divergent group compared to using the whole features, in the Claims dataset, with a Jaccard similarity of 0.95 but with a 16× reduction in detection time.
Girmaw Abebe, William Ogallo, Celia Cintas, Skyler Speakman
IEEE Big Data4
2022 Systematic Discovery of Bias in Data
abstract
Detecting bias in data is an integral component of trustworthy and responsible ML. For researchers and data scientists, investigating, detecting, and becoming aware of biases present in data is an important step to correcting and making better ML decisions. Bias exists in the form of subsets that deviate from global expectations. Typically, researchers begin with a set of pre-defined protected/sensitive attributes and use them as the basis upon which deviation from expectation is examined. For instance, a researcher may examine under- or over-representation of a particular gender or race and adjust ML models accordingly. While this works for most settings, it is suboptimal, because it does not cover the true scale of all possible enumerations of subsets in the data. In this paper, we argue for a different approach to bias discovery. Instead of performing stratification across a pre-defined set of features, we ask the more open-ended question — which subset has the highest deviation between observed and expected outcomes? To answer this question, we leverage subset scanning, which efficiently maximizes measures of divergence over exponentially many combinations of feature values. We demonstrate the capabilities and advantages of subset scanning over pre-defined stratification by analyzing scanning results on the Stanford Open Policing dataset. In so doing, we uncover anomalous subsets within the data which, to the best of our knowledge, have not been discovered before and show that it is impossible to uncover such anomalies by stratifying across a set of pre-defined features.
John Wamburu, Girmaw Abebe, Celia Cintas, Adebayo Oshingbesan, Tanya Akumu, Skyler Speakman
IEEE Big Data6
2022 Towards Creativity Characterization of Generative Models via Group-Based Subset Scanning
abstract
Deep generative models, such as Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs), have been employed widely in computational creativity research. However, such models discourage out-of-distribution generation to avoid spurious sample generation, thereby limiting their creativity. Thus, incorporating research on human creativity into generative deep learning techniques presents an opportunity to make their outputs more compelling and human-like. As we see the emergence of generative models directed toward creativity research, a need for machine learning-based surrogate metrics to characterize creative output from these models is imperative. We propose group-based subset scanning to identify, quantify, and characterize creative processes by detecting a subset of anomalous node-activations in the hidden layers of the generative models. Our experiments on the standard image benchmarks and their ``creatively generated'' variants reveal that the proposed subset scores distribution is more useful for detecting novelty in creative processes in the activation space rather than the pixel space. Further, we found that creative samples generate larger subsets of anomalies than normal or non-creative samples across datasets. The node activations highlighted during the creative decoding process are different from those responsible for the normal sample generation. Lastly, we assess if the images from the subsets selected by our method were also found creative by human evaluators, presenting a link between creativity perception in humans and node activations within deep neural nets.
Celia Cintas, Brian Quanz, Girmaw Abebe, Skyler Speakman
IJCAI5
2022 Pattern detection in the activation space for identifying synthesized content
Celia Cintas, Skyler Speakman, Girmaw Abebe, Victor Akinwande, Edward McFowland, Komminist Weldemariam
Pattern Recognit. Lett.2
2021 Automatic Stratification of Tabular Health Data
Skyler Speakman, Girmaw Abebe, Victor Akinwande, William Ogallo, Claire-Helene Mershon, Nosa Orobaton, Daniel B. Neill
AMIA1
2021 Racial Representation Analysis in Dermatology Academic Materials
Girmaw Abebe, Celia Cintas, Roxana Daneshjou, Kush R. Varshney, Peter W. J. Staar, Skyler Speakman, Kenya S. Andrews, Chinyere Agunwa, Justin Jia, Elizabeth E. Bailey, Jules Lipoff, Ginikanwa Onyekaba, Veronica Rotemberg, Ademide Adelekun, James Zou 0001
AMIA6
2020 Identifying Factors Associated with Neonatal Mortality in Sub-Saharan Africa using Machine Learning
William Ogallo, Skyler Speakman, Victor Akinwande, Kush R. Varshney, Aisha Walcott-Bryant, Charity Wayua, Komminist Weldemariam, Claire-Helene Mershon, Nosa Orobaton
AMIA2
2020 Preservation of Anomalous Subgroups On Variational Autoencoder Transformed Data
abstract
We investigate the effect of variational autoencoder (VAE) based data anonymization and its ability to preserve anomalous subgroup properties. We present a Utility Guaranteed Deep Privacy (UGDP) system which casts existing anomalous pattern detection methods as a new utility measure for data synthesis. UGDP's approach shows that properties of an anomalous subset of records, identified in the original data set, are preserved through the anonymization of a VAE. This is despite the newly generated records being completely synthetic. More specifically, the Bias-Scan algorithm identifies a subgroup of records that are consistently over- (or under-) risked by a black-box classifier as an area of 'poor fit'. This scanning process is applied on both pre- and post- VAE synthesized data. The areas of poor fit (i.e. anomalous records) persist in both settings. We evaluate our approach using publicly available datasets from the financial industry. Our evaluation confirmed that the approach is able to produce synthetic datasets that preserved a high level of subgroup differentiation as identified initially in the original dataset. Such a distinction was maintained while having distinctly different records between the synthetic and original dataset.
Samuel C. Maina, Reginald E. Bryant, William Ogallo, Kush R. Varshney, Skyler Speakman, Celia Cintas, Aisha Walcott-Bryant, Robert-Florian Samoilescu, Komminist Weldemariam
ICASSP5
2020 Detecting Adversarial Attacks via Subset Scanning of Autoencoder Activations and Reconstruction Error
abstract
Reliably detecting attacks in a given set of inputs is of high practical relevance because of the vulnerability of neural networks to adversarial examples. These altered inputs create a security risk in applications with real-world consequences, such as self-driving cars, robotics and financial services. We propose an unsupervised method for detecting adversarial attacks in inner layers of autoencoder (AE) networks by maximizing a non-parametric measure of anomalous node activations. Previous work in this space has shown AE networks can detect anomalous images by thresholding the reconstruction error produced by the final layer. Furthermore, other detection methods rely on data augmentation or specialized training techniques which must be asserted before training time. In contrast, we use subset scanning methods from the anomalous pattern detection domain to enhance detection power without labeled examples of the noise, retraining or data augmentation methods. In addition to an anomalous “score” our proposed method also returns the subset of nodes within the AE network that contributed to that score. This will allow future work to pivot from detection to visualisation and explainability. Our scanning approach shows consistently higher detection power than existing detection methods across several adversarial noise models and a wide range of perturbation strengths.
Celia Cintas, Skyler Speakman, Victor Akinwande, William Ogallo, Komminist Weldemariam, Srihari Sridharan, Edward McFowland
IJCAI2
2020 Inspection of Blackbox Models for Evaluating Vulnerability in Maternal, Newborn, and Child Health
abstract
Improving maternal, newborn, and child health (MNCH) outcomes is a critical target for global sustainable development. Our research is centered on building predictive models, evaluating their interpretability, and generating actionable insights about the markers (features) and triggers (events) associated with vulnerability in MNCH. In this work, we demonstrate how a tool for inspecting "black box" machine learning models can be used to generate actionable insights from models trained on demographic health survey data to predict neonatal mortality.
William Ogallo, Skyler Speakman, Victor Akinwande, Kush R. Varshney, Aisha Walcott-Bryant, Charity Wayua, Komminist Weldemariam
IJCAI2
2019 Fair Transfer Learning with Missing Protected Attributes
abstract
Risk assessment is a growing use for machine learning models. When used in high-stakes applications, especially ones regulated by anti-discrimination laws or governed by societal norms for fairness, it is important to ensure that learned models do not propagate and scale any biases that may exist in training data. In this paper, we add on an additional challenge beyond fairness: unsupervised domain adaptation to covariate shift between a source and target distribution. Motivated by the real-world problem of risk assessment in new markets for health insurance in the United States and mobile money-based loans in East Africa, we provide a precise formulation of the machine learning with covariate shift and score parity problem. Our formulation focuses on situations in which protected attributes are not available in either the source or target domain. We propose two new weighting methods: prevalence-constrained covariate shift (PCCS) which does not require protected attributes in the target domain and target-fair covariate shift (TFCS) which does not require protected attributes in the source domain. We empirically demonstrate their efficacy in two applications.
Amanda Coston, Karthikeyan Natesan Ramamurthy, Dennis Wei, Kush R. Varshney, Skyler Speakman, Zairah Mustahsan, Supriyo Chakraborty
AIES5
2018 Three Population Covariate Shift for Mobile Phone-based Credit Scoring
abstract
Mobile money platforms are gaining traction across developing markets as a convenient way of sending and receiving money over mobile phones. Recent joint collaborations between banks and mobile-network operators leverage a customer's past mobile phone transactions in order to create a credit score for the individual. In this work, we address the problem of launching a mobile-phone based credit scoring system in a new market without the marginal distribution of features of borrowers in the new market. This challenge rules out traditional transfer learning approaches such as a direct covariate shift. We apply a market-based re-weighting scheme of Original Market Borrowers that accounts for the differences in the original and new markets. The goal of applying this generalized covariate shift to three populations is to understand the repayment behavior of a fourth: New Market Borrowers who will self-select into a loan product when it becomes available. To test the approach we use real-world data sets from two Sub-Saharan countries in Africa consisting of 200,000 customers' telephone records. Our results demonstrate that the market-based re-weighting scheme improves the credit scoring model in the new market compared to other more direct methods.
Skyler Speakman, Srihari Sridharan, Isaac M. Markus
COMPASS1
2013 Dynamic Pattern Detection with Temporal Consistency and Connectivity Constraints
abstract
We explore scalable and accurate dynamic pattern detection methods in graph-based data sets. We apply our proposed Dynamic Subset Scan method to the task of detecting, tracking, and source-tracing contaminant plumes spreading through a water distribution system equipped with noisy, binary sensors. While static patterns affect the same subset of data over a period of time, dynamic patterns may affect different subsets of the data at each time step. These dynamic patterns require a new approach to define and optimize penalized likelihood ratio statistics in the subset scan framework, as well as new computational techniques that scale to large, real-world networks. To address the first concern, we develop new subset scan methods that allow the detected subset of nodes to change over time, while incorporating temporal consistency constraints to reward patterns that do not dramatically change between adjacent time steps. Second, our Additive Graph Scan algorithm allows our novel scan statistic to process small graphs (500 nodes) in 4.1 seconds on average while maintaining an approximation ratio over 99% compared to an exact optimization method, and to scale to large graphs with over 12,000 nodes in 30 minutes on average. Evaluation results across multiple detection, tracking, and source-tracing tasks demonstrate substantial performance gains achieved by the Dynamic Subset Scan approach.
Skyler Speakman, Daniel B. Neill
ICDM1
2013 Fast generalized subset scan for anomalous pattern detection
Edward McFowland, Skyler Speakman, Daniel B. Neill
J. Mach. Learn. Res.2