Samantha Kleinberg

dblp:02/522 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0001-6964-3272ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 8 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Causal and Counterfactual Reasoning about Gradual and Abrupt Events
Vanessa Cheung, Cristina Leone, Samantha Kleinberg, David A. Lagnado
CogSci3
2025 Understanding the Impact of Metacognitive Ability on Decision-Making with Causal Diagrams
Elena Korshakova, Samantha Kleinberg
CogSci2
2025 Go Big or Go Hoax: Explanatory Scope and the Believability of Conspiracy Theories
Jessecae K. Marsh, Samantha Kleinberg
CogSci2
2025 Causal inference for time series datasets with partially overlapping variables
Louis Adedapo Gomez, Jan Claassen, Samantha Kleinberg
J. Biomed. Informatics3
2023 How Beliefs Influence Perceptions of Choices
Samantha Kleinberg, Elena Korshakova, Jessecae K. Marsh
CogSci1
2023 Quantifying the Utility of Complexity and Feedback Loops in Causal Models for Decision Making
Elena Korshakova, Jessecae K. Marsh, Samantha Kleinberg
CogSci3
2022 Absence Makes the Trust in Causal Models Grow Stronger
Samantha Kleinberg, Eren Alay, Jessecae K. Marsh
CogSci1
2022 The Compelling Complexity of Conspiracy Theories
Jessecae K. Marsh, Cayse Coachys, Samantha Kleinberg
CogSci3
2021 It's Complicated: Improving Decisions on Causally Complex Topics
Samantha Kleinberg, Jessecae K. Marsh
CogSci1
2021 Collaborative Graph Learning with Auxiliary Text for Temporal Event Prediction in Healthcare
abstract
Accurate and explainable health event predictions are becoming crucial for healthcare providers to develop care plans for patients. The availability of electronic health records (EHR) has enabled machine learning advances in providing these predictions. However, many deep-learning-based methods are not satisfactory in solving several key challenges: 1) effectively utilizing disease domain knowledge; 2) collaboratively learning representations of patients and diseases; and 3) incorporating unstructured features. To address these issues, we propose a collaborative graph learning model to explore patient-disease interactions and medical domain knowledge. Our solution is able to capture structural features of both patients and diseases. The proposed model also utilizes unstructured text data by employing an attention manipulating strategy and then integrates attentive text features into a sequential learning process. We conduct extensive experiments on two important healthcare problems to show the competitive prediction performance of the proposed method compared with various state-of-the-art models. We also confirm the effectiveness of learned representations and model interpretability by a set of ablation and case studies.
Chang Lu 0004, Chandan K. Reddy, Prithwish Chakraborty, Samantha Kleinberg, Yue Ning 0001
IJCAI4
2020 Tell me something I don't know: How perceived knowledge influences the use of information during decision making
Samantha Kleinberg, Jessecae K. Marsh
CogSci1
2019 Lagged Correlations among Physiological Variables as Indicators of Consciousness in Stroke Patients
Tahsin T. Yavuz, Jan Claassen, Samantha Kleinberg
AMIA3
2019 The Role of Causal Information and Perceived Knowledge in Decision-Making
Min Zheng 0001, Jessecae K. Marsh, Samantha Kleinberg
CogSci3
2019 Automated meal detection from continuous glucose monitor data through simulation and explanation
abstract
BACKGROUND: Artificial pancreas systems aim to reduce the burden of type 1 diabetes by automating insulin dosing. These systems link a continuous glucose monitor (CGM) and insulin pump with a control algorithm, but require users to announce meals, without which the system can only react to the rise in blood glucose. OBJECTIVE: We investigate whether CGM data can be used to automatically infer meals in daily life even in the presence of physical activity, which can raise or lower blood glucose. MATERIALS AND METHODS: We propose a novel meal detection algorithm that combines simulations with CGM, insulin pump, and heart rate monitor data. When observed and predicted glucose differ, our algorithm uses simulations to test whether a meal may explain this difference. We evaluated our method on simulated data and real-world data from individuals with type 1 diabetes. RESULTS: In simulated data, we detected meals earlier and with higher accuracy than was found in prior work (25.7 minutes, 1.2 g error; compared with 48.3 minutes, 17.2 g error). In real-world data, we discovered a larger number of plausible meals than was found in prior work (30 meals, 76.7% accepted; compared with 33 meals, 39.4% accepted). DISCUSSION: Prior research attempted meal detection from CGM, but had delays and lower accuracy in real data or did not allow for physical activity. Our approach can be used to improve insulin dosing in an artificial pancreas and trigger reminders for missed meal boluses. CONCLUSIONS: We demonstrate that meal information can be robustly inferred from CGM and body-worn sensor data, even in challenging environments of daily life.
Min Zheng 0001, Baohua Ni, Samantha Kleinberg
J. Am. Medical Informatics Assoc.3
2017 Replicability, Reproducibility, and Agent-based Simulation of Interventions
R. Stanley Hum, Samantha Kleinberg
AMIA2
2016 Causal Explanation Under Indeterminism: A Sampling Approach
abstract
One of the key uses of causes is to explain why things happen. Explanations of specific events, like an individual's heart attack on Monday afternoon or a particular car accident, help assign responsibility and inform our future decisions. Computational methods for causal inference make use of the vast amounts of data collected by individuals to better understand their behavior and improve their health. However, most methods for explanation of specific events have provided theoretical approaches with limited applicability. In contrast we make two main contributions: an algorithm for explanation that calculates the strength of token causes, and an evaluation based on simulated data that enables objective comparison against prior methods and ground truth. We show that the approach finds the correct relationships in classic test cases (causal chains, common cause, and backup causation) and in a realistic scenario (explaining hyperglycemic episodes in a simulation of type 1 diabetes).
Christopher A. Merck, Samantha Kleinberg
AAAI2
2016 Automated estimation of food type and amount consumed from body-worn audio and motion sensors
abstract
Determining when an individual is eating can be useful for tracking behavior and identifying patterns, but to create nutrition logs automatically or provide real-time feedback to people with chronic disease, we need to identify both what they are consuming and in what quantity. However, food type and amount have mainly been estimated using image data (requiring user involvement) or acoustic sensors (tested with a restricted set of foods rather than representative meals). As a result, there is not yet a highly accurate automated nutrition monitoring method that can be used with a variety of foods. We propose that multi-modal sensing (in-ear audio plus head and wrist motion) can be used to more accurately classify food type, as audio and motion features provide complementary information. Further, we propose that knowing food type is critical for estimating amount consumed in combination with sensor data. To test this we use data from people wearing audio and motion sensors, with ground truth annotated from video and continuous scale data. With data from 40 unique foods we achieve a classification accuracy of 82.7% with a combination of sensors (versus 67.8% for audio alone and 76.2% for head and wrist motion). Weight estimation error was reduced from a baseline of 127.3% to 35.4% absolute relative error. Ultimately, our estimates of food type and amount can be linked to food databases to provide automated calorie estimates from continuously-collected data.
Mark Mirtchouk, Christopher A. Merck, Samantha Kleinberg
UbiComp3
2016 Using uncertain data from body-worn sensors to gain insight into type 1 diabetes
abstract
The amount of observational data available for research is growing rapidly with the rise of electronic health records and patient-generated data. However, these data bring new challenges, as data collected outside controlled environments and generated for purposes other than research may be error-prone, biased, or systematically missing. Analysis of these data requires methods that are robust to such challenges, yet methods for causal inference currently only handle uncertainty at the level of causal relationships - rather than variables or specific observations. In contrast, we develop a new approach for causal inference from time series data that allows uncertainty at the level of individual data points, so that inferences depend more strongly on variables and individual observations that are more certain. In the limit, a completely uncertain variable will be treated as if it were not measured. Using simulated data we demonstrate that the approach is more accurate than the state of the art, making substantially fewer false discoveries. Finally, we apply the method to a unique set of data collected from 17 individuals with type 1 diabetes mellitus (T1DM) in free-living conditions over 72h where glucose levels, insulin dosing, physical activity and sleep are measured using body-worn sensors. These data often have high rates of error that vary across time, but we are able to uncover the relationships such as that between anaerobic activity and hyperglycemia. Ultimately, better modeling of uncertainty may enable better translation of methods to free-living conditions, as well as better use of noisy and uncertain EHR data.
Nathaniel Heintzman, Samantha Kleinberg
J. Biomed. Informatics2
2015 Combining Fourier and lagged k-nearest neighbor imputation for biomedical time series data
Shah Atiqur Rahman, Jan Claassen, Nathaniel Heintzman, Samantha Kleinberg
J. Biomed. Informatics5
2013 Lessons Learned in Replicating Data-Driven Experiments in Multiple Medical Systems and Patient Populations
Samantha Kleinberg, Noémie Elhadad
AMIA1
2013 Causal Inference with Rare Events in Large-Scale Time-Series Data
Samantha Kleinberg
IJCAI1
2011 A Logic for Causal Inference in Time Series with Discrete and Continuous Variables
abstract
Many applications of causal inference, such as finding the relationship between stock prices and news reports, involve both discrete and continuous variables observed over time. Inference with these complex sets of temporal data, though, has remained difficult and required a number of simplifications. We show that recent approaches for inferring temporal relationships (represented as logical formulas) can be adapted for inference with continuous valued effects. Building on advances in logic, PCTLc (an extension of PCTL with numerical constraints) is introduced here to allow representation and inference of relationships with a mixture of discrete and continuous components. Then, finding significant relationships in the continuous case can be done using the conditional expectation of an effect, rather than its conditional probability. We evaluate this approach on both synthetically generated and actual financial market data, demonstrating that it can allow us to answer different questions than the discrete approach can.
Samantha Kleinberg
IJCAI1
2011 A review of causal inference for biomedical informatics
Samantha Kleinberg, George Hripcsak
J. Biomed. Informatics1
2010 The Temporal Logic of Token Causes
Samantha Kleinberg, Bud Mishra
KR1
2010 Predicting malaria interactome classifications from time-course transcriptomic data along the intraerythrocytic developmental cycle
Antonina Mitrofanova, Samantha Kleinberg, Jane Carlton, Simon Kasif, Bud Mishra
Artif. Intell. Medicine2
2009 The Temporal Logic of Causal Structures
Samantha Kleinberg, Bud Mishra
UAI1
2008 Systems Biology via Redescription and Ontologies (III): Protein Classification Using Malaria Parasite's Temporal Transcriptomic Profiles
abstract
This paper addresses the protein classification problem, andexplores how its accuracy can be improved by using information fromtime-course gene expression data. The methods are tested on datafrom the most deadly species of the parasite responsible for malariainfections, Plasmodium falciparum. Even though avaccination for Malaria infections has been under intense study formany years, more than half of Plasmodiumproteins still remain uncharacterized and therefore are exemptedfrom clinical trials. The task is further complicated by arapid life cycle of the parasite, thus making precisetargeting of the appropriate proteins for vaccination a technicalchallenge. We propose to integrate protein-protein interactions (PPIs),sequence similarity, metabolic pathway, andgene expression, to produce a suitable set of predicted proteinfunctions for P.falciparum. Further,we treat gene expression data withrespect to various changes that occur during the five phases of theintraerythrocytic developmental cycle (IDC) (as determinedby our segmentation algorithm) ofP.falciparum and show that this analysis yields asignificantly improved protein function prediction, e.g., whencompared to analysis based on Pearson correlation coefficients seenin the data. The algorithm is able to assign ``meaningful''functions to 628 out of 1439 previously unannotated proteins, whichare first-choice candidates for experimental vaccine research.
Antonina Mitrofanova, Samantha Kleinberg, Jane Carlton, Simon Kasif, Bud Mishra
BIBM2
2008 Psst: a web-based system for tracking political statements
abstract
Determining candidates' views on important issues is critical in deciding whom to support and vote for; but finding their statements and votes on an issue can be laborious. In this paper we present PSST, (Political Statement and Support Tracker), a search engine to facilitate analysis of political statements and votes over time. We show that prior tools for text analysis can be combined with minimal manual processing to provide a first step in the full automation of this process.
Samantha Kleinberg, Bud Mishra
WWW1