George O. Mohler

dblp:77/9171 · DBLP profile ↗
← Back
15ranked-venue papers in the field
0as first author
5since 2021 · last 2024
0000-0003-4293-5106ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 8Data Mining & Knowledge Discovery · 6Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2024 Accurate Estimation of Cross-Excitation in Multivariate Hawkes Process Models of Infectious Diseases
abstract
Multivariate Hawkes processes are a popular model for estimating Granger causality from event sequences on networks. In this work we show that under certain parameter regimes, such as those that arise when modeling infectious disease transmission, false discovery of cross-excitation becomes a major problem. We first provide evidence through simulation that substantial spurious cross-excitation is present when the largest eigenvalue of the productivity matrix approaches the critical value of 1, which leads to multicollinearity. We then propose and compare several methods for mitigating false cross-excitation, through different types of regularization and staged estimation. Our experimental results include both synthetic data as well as transmission data from the Covid-19 pandemic.
Youness Diouane, Frederic Schoenberg, George O. Mohler
DSAA3
2023 Low-Cost Gunshot Detection System with Localization for Community Based Violence Interruption
abstract
There is growing interest in U.S. cities to shift resources towards community-led solutions to crime and disorder. However, there is a simultaneous need to provide community organizations with access to real-time data to facilitate decision making, to which only the police normally have access. In this work we present a low-cost gunshot detection system with localization that has been developed for community-based violence interruption. The distributed real-time gunshot detection sensor network is linked to a mobile phone-based alert and tasking system for exclusive use by civilian gang interventionists. Here we present details on the system architecture and gunshot detection model, which consists of an Audio Spectrogram Transformer (AST) neural network. We then combine gradient maps of the input to the AST for time of arrival identification with a Bayesian maximum a posteriori estimation procedure to identify the location of gunshots. We conduct several experiments using simulated data, open data from the commercial ShotSpotter detection system in Pittsburgh, and data collected using our devices during live-fire experiments at the Indianapolis Metropolitan Police Department (IMPD) gun firing range. We then discuss potential applications of the system and directions for future research.
Isaac Manring, James H. Hill, George O. Mohler, P. Jeffrey Brantingham, Thomas Williams, Bruce White
DSAA3
2023 Rewiring Police Officer Training Networks to Reduce Forecasted Use of Force
abstract
Research has shown that police officer involved shootings, misconduct and excessive use of force complaints exhibit network effects, where officers are at greater risk of being involved in these incidents when they socialize with officers who have a history of use of force and misconduct. In this work, we first construct a network survival model for the time-to-event of use of force incidents involving new police trainees. The model includes network effects of the diffusion of risk from field training officer (FTO) to trainee. We then introduce a network rewiring algorithm to maximize the expected time to use of force events upon completion of field training. We study several versions of the algorithm, including constraints that encourage demographic diversity of FTOs. Using data from Indianapolis, we show that rewiring the network can increase the expected time (in days) of a recruit's first use of force incident by 8%. We then discuss the potential benefits and challenges associated with implementing such an algorithm in practice.
Ritika Pandey, Jeremy G. Carter, James H. Hill, George O. Mohler
KDD4
2021 Source detection on networks using spatial temporal graph convolutional networks
abstract
Detecting the source of an outbreak cluster during a pandemic like COVID-19 can provide insights into the transmission process, associated risk factors, and help contain the spread. In this work we study the problem of source detection from multiple snapshots of spreading on an arbitrary network structure. We use a spatial temporal graph convolutional network based model (SD-STGCN) to produce a source probability distribution, by fusing information from temporal and topological spaces. We perform extensive experiments using popular compartmental simulation models over synthetic networks and empirical contact networks. We also demonstrate the applicability of our approach with real COVID-19 case data.
Hao Sha 0003, Mohammad Al Hasan, George O. Mohler
DSAA3
2021 Group Link Prediction Using Conditional Variational Autoencoder
Hao Sha 0003, Mohammad Al Hasan, George O. Mohler
ICWSM3
2020 Repurposing recidivism models for forecasting police officer use of force
abstract
We review several concepts and modeling techniques from statistical and machine learning that have been developed to forecast recidivism. We show how these methods might be repurposed for forecasting police officer use of force. Using open Chicago police department use-of-force complaint data for illustration, we discuss feature engineering, construction of black-box models, interpretable forecasts, and fairness.
Samira Khorshidi, Jeremy G. Carter, George O. Mohler
IEEE BigData3
2020 Interpretable Hawkes Process Spatial Crime Forecasting with TV-Regularization
abstract
Interpretable models for criminal justice forecasting are desirable due to the high-stakes nature of the application. While interpretable models have been developed for individual level forecasts of recidivism, interpretable models are lacking for the application of space-time crime hotspot forecasting. Here we introduce an interpretable Hawkes process model of crime that allows forecasts to capture near-repeat effects and spatial heterogeneity while being consumable in the form of easy-to-read score cards. For this purpose we employ penalized likelihood estimation of the point process with a total-variation regularization that enforces the triggering kernel to be piece-wise constant. We derive an efficient expectation-maximization algorithm coupled with forward backward splitting for the TV constraint to estimate the model. We apply our methodology to synthetic data and space-time crime data from Indianapolis. The TV-Hawkes process achieves similar accuracy to standard Hawkes process models of crime while increasing interpretability and transparency.
Hao Sha 0003, Mohammad Al Hasan, Jeremy G. Carter, George O. Mohler
IEEE BigData4
2020 Automated Corn Ear Height Prediction Using Video-Based Deep Learning
abstract
In corn breeding, hand-measurement of ear height is a labor-intensive process, thus limiting scalability. Here we show that it is feasible to automate estimation of the average ear height of a row of corn in experimental fields used for corn breeding. For this purpose we use point pattern analysis on predicted shank-node locations extracted from video captured on uncalibrated cameras moving through a plot at a fixed height from the ground (4 feet and 2 feet). First, a convolutional neural network-based object detection system (YOLOv3) was trained to detect the ear-stalk connection point and applied to the collected videos. Detected ear position and time information from each frame were super-imposed into a point pattern and point-features were then extracted. Using ridge regression to predict the average ear height per plot, we achieved 0.772 concordance, 2.989 inches root mean squared error, and 2.263 inches mean absolute error compared with hand-measured average ear height. Feature weight importance suggests that one camera may be sufficient for prediction without significant decrease in accuracy. This deep learning system can be utilized by mounting cameras onto the plot combine harvester to collect the necessary videos during harvest and could be expanded to quantify other phenotype measurements of interest that are labor-intensive to collect.
Johnson Wong, Hao Sha 0003, Mohammad Al Hasan, George O. Mohler, Steve Becker, Curtis Wiltse
IEEE BigData4
2019 Into the Reverie: Exploration of the Dream Market
abstract
Since the emergence of the Silk Road market in the early 2010s, dark web `cryptomarkets' have proliferated and offered people an online platform to buy and sell illicit drugs, relying on cryptocurrencies such as Bitcoin for anonymous transactions. However, recent studies have highlighted the potential for de-anonymization of bitcoin transactions, bringing into question the level of anonymity afforded by cryptomarkets. We examine a set of over 100,000 product reviews from several cryptomarkets collected in 2018 and 2019 and conduct a comprehensive analysis of the markets, including an examination of the distribution of drug sales and revenue among vendors, and a comparison of incidences of opioid sales to overdose deaths in a US city. We explore the potential for de-anonymization of vendors by implementing a Naïve-Bayes classifier to predict the vendor from a given product review, and attempt to link vendors' sales to specific Bitcoin transactions. On the buyer side, we evaluate the efficacy of hierarchical agglomerative clustering for grouping together transactions corresponding to the same buyer. We find that the high degree of specialization among the small subset of high-revenue vendors may render these vendors susceptible to de-anonymization. Further research is necessary to confirm these findings, which are restricted by the scarcity of ground-truth data for validation.
Theo Carr, Jun Zhuang 0004, Dwight Sablan, Emma LaRue, Yubao Wu, Mohammad Al Hasan, George O. Mohler
IEEE BigData7
2019 Low Cost Gunshot Detection using Deep Learning on the Raspberry Pi
abstract
Many cities using gunshot detection technology depend on expensive systems that ultimately rely on humans differentiating between gunshots and non-gunshots, such as ShotSpotter. Thus, a scalable gunshot detection system that is low in cost and high in accuracy would be advantageous for a variety of cities across the globe, in that it would favorably promote the delegation of tasks typically worked by humans to machines. A repository of audio data was created from sound clips collected from online audio databases as well as from clips recorded using a USB microphone in residential areas and at a gun range. One-dimensional as well as two-dimensional convolutional neural networks were then trained on this sound data, and spectrograms created from this sound data, to recognize gunshots. These models were deployed to a Raspberry Pi 3 Model B+ with a short message service modem and a USB microphone attached, using a software pipeline to continuously analyze discrete two-second chunks of audio and alert a set of phone numbers if a gunshot is detected in that chunk. Testing found that a majority-rules ensemble of our one-dimensional and two-dimensional models fared best, with an accuracy above 99% on validation data as well as when distinguishing gunshots from fireworks. Besides increasing the safety standards for a city's residents, the findings generated by this research project expand the current state of knowledge regarding sound-based applications of convolutional neural networks.
Alex Morehead, Lauren Ogden, Gabe Magee, Ryan Hosler, Bruce White, George O. Mohler
IEEE BigData6
2019 Group Link Prediction
abstract
Due to its universal applications in the domain of social network analysis, e-commerce, and recommendation systems, the task of link prediction has received enormous attention from the data mining and machine learning communities over the last decade. In its original setting, the task only predicts whether a pair of entities who are not connected at present time will form a connection in future. However, in real-life an entity sometimes join a group (or a community), thus making a connection with the group (or the community), instead of connecting with an individual. Existing solutions to link prediction are inadequate for solving this prediction task. To overcome this challenge, in this work we propose a novel problem named group link prediction which focuses on evaluating the likelihood for a candidate to become a member of a group at a given time. The problem has potential applications such as friendship or group suggestions on Facebook or other social networks, as well as co-authorship suggestion, or group email recommendations. To solve the problem, we propose a Long Short-term Memory based model that inputs the embedding vectors of the group and outputs the conditional probability distributions for the candidates. We also introduce a composite long short-term memory model that integrates keyword information. Experimental results on real-world data sets validate the superiority of our proposed model in comparison to various baseline methods.
Andrew Stanhope, Hao Sha 0003, Danielle Barman, Mohammad Al Hasan, George O. Mohler
IEEE BigData5
2019 Investigate Transitions into Drug Addiction through Text Mining of Reddit Data
abstract
Increasing rates of opioid drug abuse and heightened prevalence of online support communities underscore the necessity of employing data mining techniques to better understand drug addiction using these rapidly developing online resources. In this work, we obtained data from Reddit, an online collection of forums, to gather insight into drug use/misuse using text snippets from users narratives. Specifically, using users' posts, we trained a binary classifier which predicts a user's transitions from casual drug discussion forums to drug recovery forums. We also proposed a Cox regression model that outputs likelihoods of such transitions. In doing so, we found that utterances of select drugs and certain linguistic features contained in one's posts can help predict these transitions. Using unfiltered drug-related posts, our research delineates drugs that are associated with higher rates of transitions from recreational drug discussion to support/recovery discussion, offers insight into modern drug culture, and provides tools with potential applications in combating the opioid crisis.
John Lu, Sumati Sridhar, Ritika Pandey, Mohammad Al Hasan, George O. Mohler
KDD5
2018 Predicting Virality on Networks Using Local Graphlet Frequency Distribution
abstract
The task of predicting virality has far-reaching consequences, from the world of advertising to more recent attempts to reduce the spread of fake news. Previous work has shown that graphlet distribution is an effective feature for predicting virality. Here, we investigate the use of aggregated edge-centric local graphlets around source nodes as features for virality prediction. These prediction features are used to predict expected virality for both a time-independent Hawkes model and an independent cascade model of virality. In the Hawkes model, we use linear regression to predict the number of Hawkes events and node ranking, while in the independent cascade model we use logistic regression to predict whether a k-size cascade will multiply by a factor X in size. Our study indicates that local graphlet frequency distribution can effectively capture the variances of the viral processes simulated by Hawkes process and independent-cascade process. Furthermore, we identify a group of local graphlets which might be significant in the viral processes. We compare the effectiveness of our methods with eigenvector centrality-based node choice.
Andrew Baas, Frances Hung, Hao Sha 0003, Mohammad Al Hasan, George O. Mohler
IEEE BigData5
2018 Coupled IGMM-GANs for improved generative adversarial anomaly detection
abstract
Detecting anomalies and outliers in data has a number of applications including hazard sensing, fraud detection, and systems management. While generative adversarial networks seem like a natural fit for addressing these challenges, we find that existing GAN based anomaly detection algorithms perform poorly due to their inability to handle multimodal patterns. For this purpose we introduce an infinite Gaussian mixture model coupled with (bi-directional) generative adversarial networks, IGMM-GAN, that facilitates multimodal anomaly detection. We illustrate our methodology and its improvement over existing GAN anomaly detection on the MNIST dataset.
Kathryn Gray, Daniel Smolyak, Sarkhan Badirli, George O. Mohler
IEEE BigData4
2018 Forecasting Retweet Count during Elections Using Graph Convolution Neural Networks
abstract
A retweet refers to sharing a tweet posted by another user on Twitter and is primary way information spreads on the Twitter network. Political parties use Twitter extensively as a part of their campaign to promote their presence, announce their propaganda, and at times debating with opponents. In this work we consider the problem of early prediction of the final retweet count using information from the network during the first several minutes after a post is made. Such predictions are useful for ranking and promoting posts and also can be used in combination with fake news detection. From a machine learning perspective, the task can be viewed as a regression problem. We introduce a novel graph convolution neural network for forecasting retweet count that combines network level features through graph convolution layers as well as tweet level features at a higher dense layer in the network. We first will provide an overview of the graph convolution network architecture and then perform several experiments on Twitter data collected during presidential elections in South Africa (2014) and Kenya (2013). We show that the model outperforms baseline models including a feed forward neural network and the popular point process based model SEISMIC.
Raghavendran Vijayan, George O. Mohler
DSAA2