James G. Shanahan

dblp:88/5622 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
3since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 8 first-author · 2 since 2021Databases, data management, data science and information retrieval · 14 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Online Advertising Incrementality Testing: Practical Lessons, Paid Search and Emerging Challenges
Joel Barajas, Narayan L. Bhamidipati, James G. Shanahan
ECIR (2)3
2021 Online Advertising Incrementality Testing: Practical Lessons And Emerging Challenges
abstract
Online advertising has historically been approached as an ad-to-user matching problem within sophisticated optimization algorithms. As the research and ad-tech industries have progressed, advertisers have increasingly emphasized the causal effect estimation of their ads (incrementality) using controlled experiments (A/B testing). With low lift effects and sparse conversion, the development of incrementality testing platforms at scale suggests tremendous engineering challenges in measurement precision. Similarly, the correct interpretation of results addressing a business goal requires significant data science and experimentation research expertise. We propose a practical tutorial in the incrementality testing landscape, including: The business need; Literature solutions and industry practices; Designs in the development of testing platforms; The testing cycle, case studies, and recommendations. We provide first-hand lessons based on the development of such a platform in a major combined DSP and ad network, and after running several tests for up to two months each over recent years.
Joel Barajas, Narayan L. Bhamidipati, James G. Shanahan
CIKM3
2021 Online Advertising Incrementality Testing And Experimentation: Industry Practical Lessons
abstract
Online advertising has historically been approached as user targeting and ad-to-user matching problems within sophisticated optimization algorithms. As the research area and ad tech industry have progressed over the last couple of decades, advertisers have increasingly emphasized the causal effect estimation of their ads (aka incrementality) using controlled experiments (or A/B testing). Even though observational approaches have been derived in marketing science since the 80s including media mix models, the availability of online advertising personalization has enabled the deployment of more rigorous randomized controlled experiments with millions of individuals. These evolutions in marketing science, online advertising, and the ad tech industry have posed incredible challenges for engineers, data scientists, and marketers alike. With low effect percentage differences (or lift) and often sparse conversion rates, the development of incrementality testing platforms at scale suggests tremendous engineering challenges in the measurement precision and detailed implementation. Similarly, the correct interpretation of results addressing a business goal within the marketing science domain requires significant data science and experimentation research expertise. All these challenges on the ongoing evolution of the online advertising industry and the heterogeneity of its sources (social, paid search, native, programmatic, etc). In the current tutorial, we propose a practical, grounded view in the incrementality testing landscape, including: The business need Solutions in the literature Design and choices in the development of incrementality testing platform The testing cycle, case studies, and recommendations to effective results delivery Incrementality testing evolution in the industry We will provide first-hand lessons on developing and operationalizing such a platform in a major combined DSP and ad network; these are based on running tens of experiments for up to two months each over the last couple of years.
Joel Barajas, Narayan L. Bhamidipati, James G. Shanahan
KDD3
2020 Introduction to Computer Vision and Realtime Deep Learning-based Object Detection
abstract
The main focus of object detection, one of the most challenging problems in computer vision (CV), is to predict a set of bounding boxes and category labels for each object of interest in an image or in a point cloud. As such, object detection has a variety of exciting downstream applications such as self-driving cars, checkout-less shopping, smart cities, cancer detection, and more. This field has been revolutionized by deep learning over the past five years, where during this time, two-stage approaches to object detection have given way to simpler, more efficient, one-stage models. Mean average precision (mAP) on benchmark problems such as the COCO Object Detection dataset has improved almost 4X over the course of five years from 15% (Fast RCNN, a two-stage approach) to 55% (EfficientDet7x, a one-stage approach). This tutorial looks under the hood of state-of-the-art object detection systems, such as two-stage, one-stage, and also more recent approaches based upon transformers. It builds out some of their associated detection pipelines in a Jupyter Notebook using Python, OpenCV, PyTorch, Keras and Tensorflow. While the primary focus is on object detection in digital images from cameras and videos, this tutorial will also introduce object detection in 3D point clouds.
James G. Shanahan
CIKM1
2020 Introduction to Computer Vision and Real Time Deep Learning-based Object Detection
abstract
Computer vision (CV) is a field of artificial intelligence that trains computers to interpret and understand the visual world for a variety of exciting downstream tasks such as self-driving cars, checkout-less shopping, smart cities, cancer detection, and more. The field of CV has been revolutionized by deep learning over the last decade. This tutorial looks under the hood of modern day CV systems, and builds out some of these tech pipelines in a Jupyter Notebook using Python, OpenCV, Keras and Tensorflow. While the primary focus is on digital images from cameras and videos, this tutorial will also introduce 3D point clouds, and classification and segmentation algorithms for processing them.
James G. Shanahan, Liang Dai 0003
KDD1
2019 Realtime Object Detection via Deep Learning-based Pipelines
abstract
Ever wonder how the Tesla Autopilot system works (or why it fails)? In this tutorial we will look under the hood of self-driving cars and of other applications of computer vision and review state-of-the-art tech pipelines for object detection such as two-stage approaches (e.g., Faster R-CNN) or single-stage approaches (e.g., YOLO/SSD). This is accomplished via a series of Jupyter Notebooks that use Python, OpenCV, Keras, and Tensorflow. No prior knowledge of computer vision is assumed (although it will be help!). To this end we begin this tutorial with a review of computer vision and traditional approaches to object detection such as Histogram of oriented gradients (HOG).
James G. Shanahan, Liang Dai 0003
CIKM1
2015 Large Scale Distributed Data Science using Apache Spark
abstract
Apache Spark is an open-source cluster computing framework for big data processing. It has emerged as the next generation big data processing engine, overtaking Hadoop MapReduce which helped ignite the big data revolution. Spark maintains MapReduce's linear scalability and fault tolerance, but extends it in a few important ways: it is much faster (100 times faster for certain applications), much easier to program in due to its rich APIs in Python, Java, Scala (and shortly R), and its core data abstraction, the distributed data frame, and it goes far beyond batch applications to support a variety of compute-intensive tasks, including interactive queries, streaming, machine learning, and graph processing. This tutorial will provide an accessible introduction to Spark and its potential to revolutionize academic and commercial data science practices.
James G. Shanahan, Liang Dai 0003
KDD1
2012 Estimating the Expected Effectiveness of Text Classification Solutions under Subclass Distribution Shifts
abstract
Automated text classification is one of the most important learning technologies to fight information overload. However, the information society is not only confronted with an information flood but also with an increase in "information volatility", by which we understand the fact that kind and distribution of a data source's emissions can significantly vary. In this paper we show how to estimate the expected effectiveness of a classification solution when the underlying data source undergoes a shift in the distribution of its subclasses (modes). Subclass distribution shifts are observed among others in online media such as tweets, blogs, or news articles, where document emissions follow topic popularity. To estimate the expected effectiveness of a classification solution we partition a test sample by means of clustering. Then, using repetitive resampling with different margin distributions over the clustering, the effectiveness characteristics is studied. We show that the effectiveness is normally distributed and introduce a probabilistic lower bound that is used for model selection. We analyze the relation between our notion of expected effectiveness and the mean effectiveness over the clustering both theoretically and on standard text corpora. An important result is a heuristic for expected effectiveness estimation that is solely based on the initial test sample and that can be computed without resampling.
Nedim Lipka, Benno Stein 0001, James G. Shanahan
ICDM3
2010 Exploiting sequential relationships for familial classification
abstract
The pervasive nature of the internet has caused a significant transformation in the field of genealogical research. This has impacted not only how research is conducted, but has also dramatically increased the number of people discovering their family history. Recent market research (Maritz Marketing 2000, Harris Interactive 2009) indicates that general interest in the United States has increased from 45% in 1996, to 60% in 2000, and 87% in 2009. Increased popularity has caused a dramatic need for improvements in algorithms related to extracting, accessing, and processing genealogical data for use in building family trees. This paper presents one approach to algorithmic improvement in the family history domain, where we infer the familial relationships of households found in human transcribed United States census data. By applying advances made in natural language processing, exploiting the sequential nature of the census, and using state of the art machine learning algorithms, we were able to decrease the error by 35% over a hand coded baseline system. The resulting system is immediately applicable to hundreds of millions of other genealogical records where families are represented, but the familial relationships are missing.
Lee S. Jensen, James G. Shanahan
CIKM2
2010 Location disambiguation in local searches using gradient boosted decision trees
abstract
Local search is a specialization of the web search that allows users to submit geographically constrained queries. However, one of the challenges for local search engines is to uniquely understand and locate the geographical intent of the query. Geographical constraints (or location references) in a local search are often incomplete and thereby suffer from the referent ambiguity problem where the same location name can mean several different possibilities. For instance, just the term "Springfield" by itself can refer to 30 different cities in the USA. Previous approaches to location disambiguation have generally been hand compiled heuristic models. In this paper, we examine a data-driven, machine learning approach to location disambiguation. Essentially, we separately train a Gradient Boosted Decision Tree (GBDT) model on thousands of desktop and mobile-based local searches and compare the performance to one of our previous heuristic based location disambiguation system (HLDS). The GBDT based approach shows promising results with statistically significant improvements over the HLDS approach. The error rate reduction is about 9% and 22% for the desktop-based and the mobile-based local searches respectively. Additionally, we examine the relative influence of various geographic and non-geographic features that help with the location disambiguation task. It is interesting to note that while the distance between the user and the intended location has been considered as an important variable, the relative influence of distance is secondary to the popularity of the location in the GBDT learnt models.
Ritesh Agrawal, James G. Shanahan
GIS2
2005 Probabilistic workflow mining
abstract
In several organizations, it has become increasingly popular to document and log the steps that makeup a typical business process. In some situations, a normative workflow model of such processes is developed, and it becomes important to know if such a model is actually being followed by analyzing the available activity logs. In other scenarios, no model is available and, with the purpose of evaluating cases or creating new production policies, one is interested in learning a workflow representation of such activities. In either case, machine learning tools that can mine workflow models are of great interest and still relatively unexplored. We present here a probabilistic workflow model and a corresponding learning algorithm that runs in polynomial time. We illustrate the algorithm on example data derived from a real world workflow.
Ricardo Bezerra de Andrade e Silva, Jiji Zhang, James G. Shanahan
KDD3
2003 Boosting support vector machines for text classification through parameter-free threshold relaxation
abstract
Support vector machine (SVM) learning algorithms focus on finding the hyperplane that maximizes the margin (the distance from the separating hyperplane to the nearest examples) since this criterion provides a good upper bound of the generalization error. When applied to text classification, these learning algorithms lead to SVMs with excellent precision but poor recall. Various relaxation approaches have been proposed to counter this problem including: asymmetric SVM learning algorithms (soft SVMs with asymmetric misclassification costs); uneven margin based learning; and thresholding. A review of these approaches is presented here. In addition, in this paper, we describe a new threshold relaxation algorithm. This approach builds on previous thresholding work based upon the beta-gamma algorithm. The proposed thresholding strategy is parameter free, relying on a process of retrofitting and cross validation to set algorithm parameters empirically, whereas our previous approach required the specification of two parameters (beta and gamma). The proposed approach is more efficient, does not require the specification of any parameters, and similarly to the parameter-based approach, boosts the performance of baseline SVMs by at least 20% for standard information retrieval measures.
James G. Shanahan, Norbert Roma
CIKM1
2003 Improving SVM Text Classification Performance through Threshold Adjustment
James G. Shanahan, Norbert Roma
ECML1
2002 Topic structure modeling
abstract
In this paper, we present a method based on document probes to quantify and diagnose topic structure, distinguishing topics as monolithic, structured, or diffuse. The method also yields a structure analysis that can be used directly to optimize filter (classifier) creation. Preliminary results illustrate the predictive value of the approach on TREC/Reuters-96 topics.
David A. Evans 0001, James G. Shanahan, Victor Sheftel
SIGIR2
2001 Modeling with Words: an Approach to Text Categorization
abstract
Traditionally, fuzzy set-based approaches have performed excellently in modeling small to medium scale problem domains. This paper examines the scalability of fuzzy systems to a large-scale problem that is inherently vague and of text categorization. The paper presents two fuzzy probabilistic approaches to text classification and the corresponding machine learning algorithms to learn such systems from example data. The first approach follows the traditional fuzzy set paradigm, while the second approach fits within the modeling with words paradigm using granule features to represent the text problem domain.
James G. Shanahan
FUZZ-IEEE1
1999 Road Recognition Using Fuzzy Classifiers
abstract
Current learning approaches to computer vision have mainly focussed on low-level image processing and object recognition, while tending to ignore higher level processing for understanding. We propose an approach to scene analysis that facilitates the transition from recognition to understanding. It begins by segmenting the image into regions using standard approaches, which are then classified using a discovered fuzzy Cartesian granule feature classifier. Understanding is made possible through the transparent and succinct nature of the discovered models. The recognition of roads in images is taken as an illustrative problem. The discovered fuzzy models while providing high levels of accuracy (97%), also provide understanding of the problem domain through the transparency of the learnt models. The learning step in the proposed approach is compared with other techniques such as decision trees, naive Bayes and neural networks using a variety of performance criteria such as accuracy, understandability, and efficiency.
James G. Shanahan, Barry T. Thomas, Majid Mirmehdi, Trevor P. Martin, Jim F. Baldwin
BMVC1
1999 Constructive induction of fuzzy Cartesian granule feature models using genetic programming with applications
abstract
Cartesian granule features are derived features that are formed over the cross product of words that linguistically partition the universes of the constituent input features. Both classification and prediction problems can be modelled quite naturally in terms of Cartesian granule features incorporated into rule based models. The induction of Cartesian granule feature model involves discovering which input features should be combined to form Cartesian granule features in order to model a domain effectively; an exponential search problem. We present the G-DACG (Genetic Discovery of Additive Cartesian Granule feature models) constructive induction algorithm as a means of automatically identifying additive Cartesian granule feature models from example data. G-DACG combines the powerful optimisation capabilities of genetic programming with a rather novel and cheap fitness function which relies on the semantic separation of learnt concepts expressed in terms of Cartesian granule fuzzy sets. G-DACG is illustrated on a variety of artificial and real world classification problems.
James G. Shanahan, Jim F. Baldwin, Trevor P. Martin
CEC1
1999 Controlling with words using automatically identified fuzzy Cartesian granule feature models
Jim F. Baldwin, Trevor P. Martin, James G. Shanahan
Int. J. Approx. Reason.3
1999 Perceptual organization for inferring object boundaries in an image1
Anca L. Ralescu, James G. Shanahan
Pattern Recognit.2
1996 Editorial
Anca L. Ralescu, James G. Shanahan
Fuzzy Sets Syst.2