Robert Ball

dblp:27/640 · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0002-1609-7420ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 6 since 2021Human-computer interaction and ubiquitous computing · 12 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-authorComputer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Deduplicating the FDA adverse event reporting system with a novel application of network-based grouping
Kory Kreimeyer, Jonathan Spiker, Oanh Dang, Suranjan De, Robert Ball, Taxiarchis Botsis
J. Biomed. Informatics5
2024 A general framework for developing computable clinical phenotype algorithms
abstract
OBJECTIVE: To present a general framework providing high-level guidance to developers of computable algorithms for identifying patients with specific clinical conditions (phenotypes) through a variety of approaches, including but not limited to machine learning and natural language processing methods to incorporate rich electronic health record data. MATERIALS AND METHODS: Drawing on extensive prior phenotyping experiences and insights derived from 3 algorithm development projects conducted specifically for this purpose, our team with expertise in clinical medicine, statistics, informatics, pharmacoepidemiology, and healthcare data science methods conceptualized stages of development and corresponding sets of principles, strategies, and practical guidelines for improving the algorithm development process. RESULTS: We propose 5 stages of algorithm development and corresponding principles, strategies, and guidelines: (1) assessing fitness-for-purpose, (2) creating gold standard data, (3) feature engineering, (4) model development, and (5) model evaluation. DISCUSSION AND CONCLUSION: This framework is intended to provide practical guidance and serve as a basis for future elaboration and extension.
David Carrell, James S. Floyd, Susan Gruber, Brian L. Hazlehurst, Patrick J. Heagerty, Jennifer C. Nelson, Brian D. Williamson, Robert Ball
J. Am. Medical Informatics Assoc.8
2023 Trends and opportunities in computable clinical phenotyping: A scoping review
Anas Belouali, Jessica Patricoski, Harold P. Lehmann, Robert Ball, Valsamo Anagnostou, Kory Kreimeyer, Taxiarchis Botsis
J. Biomed. Informatics5
2022 The US Food and Drug Administration Sentinel System: a national resource for a learning health system
abstract
The US Food and Drug Administration (FDA) created the Sentinel System in response to a requirement in the FDA Amendments Act of 2007 that the agency establish a system for monitoring risks associated with drug and biologic products using data from disparate sources. The Sentinel System has completed hundreds of analyses, including many that have directly informed regulatory decisions. The Sentinel System also was designed to support a national infrastructure for a learning health system. Sentinel governance and guiding principles were designed to facilitate Sentinel's role as a national resource. The Sentinel System infrastructure now supports multiple non-FDA projects for stakeholders ranging from regulated industry to other federal agencies, international regulators, and academics. The Sentinel System is a working example of a learning health system that is expanding with the potential to create a global learning health system that can support medical product safety assessments and other research.
Jeffrey S. Brown, Aaron B. Mendelsohn, Young Hee-Nam, Judith C. Maro, Noelle M. Cocoros, Carla Rodriguez-Watson, Catherine M. Lockhart, Richard Platt, Robert Ball, Gerald J. Dal Pan, Sengwee Toh
J. Am. Medical Informatics Assoc.9
2021 Reliable insights regarding medication outcomes from real world data- key principles of an evidence generation framework
Rishi J. Desai, Sebastian Schneeweiss, Keith Marsolo, Shirley V. Wang, Robert Ball
AMIA5
2021 Specialized Neural Network Pruning for Boolean Abstractions
Jarren Briscoe, Brian W. Rague, Kyle D. Feuz, Robert Ball
KEOD4
2021 Electronic phenotyping of health outcomes of interest using a linked claims-electronic health record database: Findings from a machine learning pilot project
abstract
OBJECTIVE: Claims-based algorithms are used in the Food and Drug Administration Sentinel Active Risk Identification and Analysis System to identify occurrences of health outcomes of interest (HOIs) for medical product safety assessment. This project aimed to apply machine learning classification techniques to demonstrate the feasibility of developing a claims-based algorithm to predict an HOI in structured electronic health record (EHR) data. MATERIALS AND METHODS: We used the 2015-2019 IBM MarketScan Explorys Claims-EMR Data Set, linking administrative claims and EHR data at the patient level. We focused on a single HOI, rhabdomyolysis, defined by EHR laboratory test results. Using claims-based predictors, we applied machine learning techniques to predict the HOI: logistic regression, LASSO (least absolute shrinkage and selection operator), random forests, support vector machines, artificial neural nets, and an ensemble method (Super Learner). RESULTS: The study cohort included 32 956 patients and 39 499 encounters. Model performance (positive predictive value [PPV], sensitivity, specificity, area under the receiver-operating characteristic curve) varied considerably across techniques. The area under the receiver-operating characteristic curve exceeded 0.80 in most model variations. DISCUSSION: For the main Food and Drug Administration use case of assessing risk of rhabdomyolysis after drug use, a model with a high PPV is typically preferred. The Super Learner ensemble model without adjustment for class imbalance achieved a PPV of 75.6%, substantially better than a previously used human expert-developed model (PPV = 44.0%). CONCLUSIONS: It is feasible to use machine learning methods to predict an EHR-derived HOI with claims-based predictors. Modeling strategies can be adapted for intended uses, including surveillance, identification of cases for chart review, and outcomes research.
Teresa B. Gibson, Michael D. Nguyen, Timothy Burrell, Frank Yoon, Jenna Wong, Sai Dharmarajan, Rita Ouellet-Hellstrom, Elande Baro, Sarah Bloemers, Cory Pack, Adee Kennedy, Sengwee Toh, Robert Ball
J. Am. Medical Informatics Assoc.15
2020 Using and improving distributed data networks to generate actionable evidence: the case of real-world outcomes in the Food and Drug Administration's Sentinel system
abstract
The US Food and Drug Administration (FDA) Sentinel System uses a distributed data network, a common data model, curated real-world data, and distributed analytic tools to generate evidence for FDA decision-making. Sentinel system needs include analytic flexibility, transparency, and reproducibility while protecting patient privacy. Based on over a decade of experience, a critical system limitation is the inability to identify enough medical conditions of interest in observational data to a satisfactory level of accuracy. Improving the system's ability to use computable phenotypes will require an "all of the above" approach that improves use of electronic health data while incorporating the growing array of complementary electronic health record data sources. FDA recently funded a Sentinel System Innovation Center and a Community Building and Outreach Center that will provide a platform for collaboration across disciplines to promote better use of real-world data for decision-making.
Jeffrey S. Brown, Judith C. Maro, Michael D. Nguyen, Robert Ball
J. Am. Medical Informatics Assoc.4
2019 Predicting adverse event reports in FDAs Adverse Event Reporting System FAERS more likely to be useful for causality assessment
Abhivyakti Sawarkar, Robert Ball
AMIA2
2019 Applying Machine Learning to Improve Curriculum Design
abstract
Creating curriculum with an ever-changing student body is difficult. Faculty members in a given department will have different perspectives on the composition and academic needs of the student body based on their personal instructional experiences. We present an approach to curriculum development that is designed to be objective by performing a comprehensive analysis of the preparation of declared majors in Computer Science (CS) BS programs at two universities. Our strategy for improving curriculum is twofold. First, we analyze the characteristics and academic needs of the student body by using a statistical, machine learning approach, which involves examining institutional data and understanding what factors specifically affect graduation. Second, we use the results of the analysis as the basis for applying necessary changes to the curriculum in order to maximize graduation rates. To validate our approach, we analyzed two four-year open enrollment universities, which share many trends that help or hinder students' progress toward graduating. Finally, we describe proposed changes to both curriculum and faculty mindsets that are a result of our findings. Although the specifics of this study are applied only to CS majors, we believe that the methods outlined in this paper can be applied to any curriculum regardless of the major.
Robert Ball, Linda P. DuHadway, Kyle D. Feuz, Joshua Jensen, Brian W. Rague, Drew Weidman
SIGCSE1
2018 GUI-Based vs. Text-Based Assignments in CS1
abstract
Teaching CS1 can be daunting. The first courses in the CS curriculum help determine which students will ultimately matriculate into the program. There have been various studies on how to improve motivation and reduce attrition by using visual-based environments and assignments. We performed a year-long study in which we addressed two research questions: 1) How is student performance affected by drag-and-drop GUI assignments when compared to traditional text-based assignments? 2) If given the choice, would students select GUI-based or text-based assignments? For the first question, there was no statistical significance, indicating that student performance is not affected by this visual component. For the second question, we discovered more students selected the text-based assignments over the GUI-assignments. Separating the students into groups based on what they chose revealed that the students that selected the GUI-assignments scored on average one letter grade higher, enjoyed the assignments more and spent less time on the assignments. We recorded the reported motivations behind why students chose to do the GUI-based assignments versus the text-based assignments: Overall, the GUI Group's responses trended toward self-improvement (e.g. more like the real world, improve skills, more challenging) while the Text Group's responses trended toward ease (e.g. easier/simpler, save time). Lastly, at the end of each course we asked the students if, given the hypothetical case in which they were not pressed for time, they would create the Java application with or without a GUI? 93% of the students responded that they would create a GUI Java application.
Robert Ball, Linda P. DuHadway, Spencer Hilton, Brian W. Rague
SIGCSE1
2018 Evaluation of Natural Language Processing (NLP) systems to annotate drug product labeling with MedDRA terminology
Thomas Ly, Carol A. Pamer, Oanh Dang, Sonja Brajovic, Shahrukh Haider, Taxiarchis Botsis, David Milward, Andrew Winter, Susan Lu, Robert Ball
J. Biomed. Informatics10
2017 Development of an automated assessment tool for MedWatch reports in the FDA adverse event reporting system
abstract
OBJECTIVE: As the US Food and Drug Administration (FDA) receives over a million adverse event reports associated with medication use every year, a system is needed to aid FDA safety evaluators in identifying reports most likely to demonstrate causal relationships to the suspect medications. We combined text mining with machine learning to construct and evaluate such a system to identify medication-related adverse event reports. METHODS: FDA safety evaluators assessed 326 reports for medication-related causality. We engineered features from these reports and constructed random forest, L1 regularized logistic regression, and support vector machine models. We evaluated model accuracy and further assessed utility by generating report rankings that represented a prioritized report review process. RESULTS: Our random forest model showed the best performance in report ranking and accuracy, with an area under the receiver operating characteristic curve of 0.66. The generated report ordering assigns reports with a higher probability of medication-related causality a higher rank and is significantly correlated to a perfect report ordering, with a Kendall's tau of 0.24 ( P = .002). CONCLUSION: Our models produced prioritized report orderings that enable FDA safety evaluators to focus on reports that are more likely to contain valuable medication-related adverse event information. Applying our models to all FDA adverse event reports has the potential to streamline the manual review process and greatly reduce reviewer workload.
Lichy Han, Robert Ball, Carol A. Pamer, Russ B. Altman, Scott Proestel
J. Am. Medical Informatics Assoc.2
2016 Use of data mining at the Food and Drug Administration
abstract
OBJECTIVES: This article summarizes past and current data mining activities at the United States Food and Drug Administration (FDA). TARGET AUDIENCE: We address data miners in all sectors, anyone interested in the safety of products regulated by the FDA (predominantly medical products, food, veterinary products and nutrition, and tobacco products), and those interested in FDA activities. SCOPE: Topics include routine and developmental data mining activities, short descriptions of mined FDA data, advantages and challenges of data mining at the FDA, and future directions of data mining at the FDA.
Hesha J. Duggirala, Joseph M. Tonning, Ella Smith, Roselie A. Bright, John D. Baker, Robert Ball, Carlos Bell, Susan J. Bright-Ponte, Taxiarchis Botsis, Khaled Bouri, Marc Boyer, Keith Burkhart, G. Steven Condrey, James J. Chen, Stuart Chirtel, Ross W. Filice, Henry Francis, Hongying Jiang, Jonathan Levine, Taiye Oladipo, Rene O'Neill, Lee Anne M. Palmer, Antonio Paredes, George Rochester, Deborah Sholtes, Ana Szarfman, Hui-Lee Wong, Zhiheng Xu, Taha A. Kass-Hout
J. Am. Medical Informatics Assoc.6
2016 A new algorithmic approach for the extraction of temporal associations from clinical narratives with an application to medical product safety surveillance reports
Wei Wang 0249, Kory Kreimeyer, Emily Jane Woo, Robert Ball, Matthew Foster, John Scott, Taxiarchis Botsis
J. Biomed. Informatics4
2015 Identifying Similar Cases in Document Networks Using Cross-Reference Structures
abstract
Our objective was to explore the creation of document networks based on different thresholds of shared information and different clustering algorithms on those networks to identify document clusters describing similar clinical cases. We created networks from vaccine adverse event report sets using seven approaches for linking reports. We then applied three clustering algorithms [visualization of similarities (VOS), Louvain, k-means] to these networks and evaluated their ability to identify known clusters. The report sets included one simulated set and three sets from the Vaccine Adverse Event Reporting System; each was split into training and testing subsets. Training subsets were used to estimate parameter values for the clustering algorithms and testing subsets to evaluate clusters. We created the networks by linking reports based on shared information in the form either of individual Medical Dictionary for Regulatory Activities Preferred Terms (PTs) or of dyads, triplets, quadruplets, quintuplets, and sextuplets of PTs; we created another network by weighting the single PT network connections by Lin's information theoretic approach to similarity. We then repeated this entire process using networks based on text mining output rather than structured data. We evaluated report clustering using recall, precision, and f-measure. The VOS algorithm outperformed Louvain and k-means in general. The best weighting scheme appeared to be related to the complexity of the known cluster. For example, singleton weighting performed best for an intussusception cluster driven by a single PT. We observed marginal differences between the code- and textual-based clustering. In conclusion, our approach supported identification of similar nodes in a document network.
Taxiarchis Botsis, John Scott, Emily Jane Woo, Robert Ball
IEEE J. Biomed. Health Informatics4
2013 Don't Search, Just Show Me What I Did: Visualizing Provenance of Documents and Applications
abstract
Computer documents have evolved over time. As a result, this article presents the results from a survey that redefines what a “document” is. The study discovered that people think a document is something that holds information; is manipulated by people (not the system); and is anything that can be seen, heard, or touched. Second, based on the results of the survey, we present a novel, real-time visualization tool that shows the results of nonintrusive tracking of documents and applications for a 6-month period. The tool focuses on document provenance—the history or genealogy of a document. It shows every document and application used as well as what happened to those documents (e.g., if the documents were moved, renamed, and/or deleted). These evaluations of the visualization tool are promising in that it helped with refinding documents, finding behavior workflow patterns, finding insight into general document usage, and performing forensic-type activities.
Robert Ball
Int. J. Hum. Comput. Interact.1
2012 Vaccine adverse event text mining system for extracting features from vaccine safety reports
abstract
OBJECTIVE: To develop and evaluate a text mining system for extracting key clinical features from vaccine adverse event reporting system (VAERS) narratives to aid in the automated review of adverse event reports. DESIGN: Based upon clinical significance to VAERS reviewing physicians, we defined the primary (diagnosis and cause of death) and secondary features (eg, symptoms) for extraction. We built a novel vaccine adverse event text mining (VaeTM) system based on a semantic text mining strategy. The performance of VaeTM was evaluated using a total of 300 VAERS reports in three sequential evaluations of 100 reports each. Moreover, we evaluated the VaeTM contribution to case classification; an information retrieval-based approach was used for the identification of anaphylaxis cases in a set of reports and was compared with two other methods: a dedicated text classifier and an online tool. MEASUREMENTS: The performance metrics of VaeTM were text mining metrics: recall, precision and F-measure. We also conducted a qualitative difference analysis and calculated sensitivity and specificity for classification of anaphylaxis cases based on the above three approaches. RESULTS: VaeTM performed best in extracting diagnosis, second level diagnosis, drug, vaccine, and lot number features (lenient F-measure in the third evaluation: 0.897, 0.817, 0.858, 0.874, and 0.914, respectively). In terms of case classification, high sensitivity was achieved (83.1%); this was equal and better compared to the text classifier (83.1%) and the online tool (40.7%), respectively. CONCLUSION: Our VaeTM implementation of a semantic text mining strategy shows promise in providing accurate and efficient extraction of key features from VAERS narratives.
Taxiarchis Botsis, Thomas Buttolph, Michael D. Nguyen, Scott Winiecki, Emily Jane Woo, Robert Ball
J. Am. Medical Informatics Assoc.6
2011 Rethinking Reading for Age From Paper and Computers
abstract
The human–computer interaction research community has long been interested in the role of age in the use of computing devices. The current availability of large high-resolution displays and a growing community of older adults who use computers on a regular basis are examples of a changing landscape of users and devices that calls for a reevaluation of what we know about age and computers. This article presents two studies comparing the performance of young and older adults in reading tasks. In our studies, older adults outperformed young adults in terms of reading times and reading comprehension regardless of medium (paper or computer display). In addition, the overlap of confidence intervals for reading times and comprehension by medium suggests participants performed equally well regardless of medium. Both findings are in contrast to results from similar studies from the 1980s and 1990s.
Robert Ball, Juan Pablo Hourcade
Int. J. Hum. Comput. Interact.1
2011 Text mining for the Vaccine Adverse Event Reporting System: medical text classification using informative feature selection
abstract
OBJECTIVE: The US Vaccine Adverse Event Reporting System (VAERS) collects spontaneous reports of adverse events following vaccination. Medical officers review the reports and often apply standardized case definitions, such as those developed by the Brighton Collaboration. Our objective was to demonstrate a multi-level text mining approach for automated text classification of VAERS reports that could potentially reduce human workload. DESIGN: We selected 6034 VAERS reports for H1N1 vaccine that were classified by medical officers as potentially positive (N(pos)=237) or negative for anaphylaxis. We created a categorized corpus of text files that included the class label and the symptom text field of each report. A validation set of 1100 labeled text files was also used. Text mining techniques were applied to extract three feature sets for important keywords, low- and high-level patterns. A rule-based classifier processed the high-level feature representation, while several machine learning classifiers were trained for the remaining two feature representations. MEASUREMENTS: Classifiers' performance was evaluated by macro-averaging recall, precision, and F-measure, and Friedman's test; misclassification error rate analysis was also performed. RESULTS: Rule-based classifier, boosted trees, and weighted support vector machines performed well in terms of macro-recall, however at the expense of a higher mean misclassification error rate. The rule-based classifier performed very well in terms of average sensitivity and specificity (79.05% and 94.80%, respectively). CONCLUSION: Our validated results showed the possibility of developing effective medical text classifiers for VAERS reports by combining text mining with informative feature selection; this strategy has the potential to reduce reviewer workload considerably.
Taxiarchis Botsis, Michael D. Nguyen, Emily Jane Woo, Marianthi Markatou, Robert Ball
J. Am. Medical Informatics Assoc.5
2008 The effects of peripheral vision and physical navigation on large scale visualization
Robert Ball, Chris North 0001
Graphics Interface1
2007 Move to improve: promoting physical navigation to increase user performance with large displays
abstract
In navigating large information spaces, previous work indicates potential advantages of physical navigation (moving eyes, head, body) over virtual navigation (zooming, panning, flying). However, there is also indication of users preferring or settling into the less efficient virtual navigation. We present a study that examines these issues in the context of large, high resolution displays. The study identifies specific relationships between display size, amount of physical and virtual navigation, and user task performance. Increased physical navigation on larger displays correlates with reduced virtual navigation and improved user performance. Analyzing the differences between this study and previous results helps to identify design factors that afford and promote the use of physical navigation in the user interface.
Robert Ball, Chris North 0001, Doug A. Bowman
CHI1
2007 Realizing embodied interaction for visual analytics through large displays
Robert Ball, Chris North 0001
Comput. Graph.1
2007 High-resolution gaming: Interfaces, notifications, and the user experience
abstract
Advances in technology and display hardware have allowed the resolution of monitors – and video games – to incrementally improve over the past three decades. However, little research has been done in preparation for the resolutions that will be available in the future if this trend continues. We developed a number of display prototypes to explore the different aspects of gaming on large, high-resolution displays. By running a series of experiments, we were not only able to evaluate the benefits of these displays for gaming, but also identify potential user interface and hardware issues that can arise. Building on these results, various interface designs were developed to better notify the user of passive and critical game information as well as to overcome difficulties with mouse-based interaction on these displays. Different display form factors and user input devices are also explored in order to determine how they can further enhance the gaming experience. In many cases, the new techniques can be applied to single-monitor games and solve the same problems in real-world, high-resolution applications.
Andrew J. Sabri, Robert Ball, Alain Fabian, Saurabh Bhatia, Chris North 0001
Interact. Comput.2
2006 Evaluation of viewport size and curvature of large, high-resolution displays
Lauren Shupp, Robert Ball, Beth Yost, John Booker, Chris North 0001
Graphics Interface2
2006 A Survey of Large High-Resolution Display Technologies, Techniques, and Applications
abstract
Continued advances in display hardware, computing power, networking, and rendering algorithms have all converged to dramatically improve large high-resolution display capabilities. We present a survey on prior research with large high-resolution displays. In the hardware configurations section we examine systems including multi-monitor workstations, reconfigurable projector arrays, and others. Rendering and the data pipeline are addressed with an overview of current technologies. We discuss many applications for large high-resolution displays such as automotive design, scientific visualization, control centers, and others. Quantifying the effects of large high-resolution displays on human performance and other aspects is important as we look toward future advances in display technology and how it is applied in different situations. Interacting with these displays brings a different set of challenges for HCI professionals, so an overview of some of this work is provided. Finally, we present our view of the top ten greatest challenges in large highresolution displays.
Tao Ni 0002, Greg S. Schmidt, Oliver G. Staadt, Mark A. Livingston, Robert Ball, Richard May 0001
VR5
2006 PC Clusters for Virtual Reality
abstract
In the late 90’s the emergence of high performance 3D commodity graphics cards opened the way to use PC clusters for high performance Virtual Reality (VR) applications. Today PC clusters are broadly used to drive multi projector immersive environments. In this paper, we survey the different approaches that have been developed to use PC clusters for VR applications. We review the most common software tools that enable to take advantage of the power of clusters. We also discuss some new trends.
Bruno Raffin, Luciano P. Soares, Tao Ni 0002, Robert Ball, Greg S. Schmidt, Mark A. Livingston, Oliver G. Staadt, Richard May 0001
VR4
2005 Analysis of User Behavior on High-Resolution Tiled Displays
Robert Ball, Chris North 0001
INTERACT1
2004 Aggressive telecommunications overbooking ratios
abstract
The Internet is comprised of vast networks of wires and fiber. A common misconception is that there is an unlimited amount of bandwidth; in reality there exists only a finite amount. Each length of wire and fiber is owned by a company, and every company wants to maximize its profit. One means of improving profit is to overbook existing transmission lines in order to increase income without increasing expenses. If too much overbooking is performed, the quality of service (QoS) seen by customers declines. This paper explains a process to achieve an optimal overbooking ratio (OR) for admission control in network routers. By optimizing the overbooking ratio, profits can be increased while minimizing QoS problems for users.
Robert Ball, Mark J. Clement, Quinn Snell, Casey T. Deccio
IPCCC1
2004 Home-centric visualization of network traffic for security administration
abstract
Today's system administrators, burdened by rapidly increasing network activity, must quickly perceive the security state of their networks, but they often have only text-based tools to work with. These tools often provide no overview to help users grasp the big-picture. Our interviews with administrators have revealed that they need visualization tools; thus, we present VISUAL (Visual Information Security Utility for Administration Live), a network security visualization tool that allows users to see communication patterns between their home (or internal) networks and external hosts. VISUAL is part of our Network Eye security visualization architecture, also described in this paper.
Robert Ball, Glenn A. Fink, Chris North 0001
VizSEC1