Sarah Masud Preum

dblp:165/8174 · also Sarah M. Preum, Sarah Preum · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-7771-8323ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
abstract
Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, Sarah Masud Preum. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, Sarah Masud Preum
ACL (1)7
2026 Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance
abstract
LLMs are increasingly deployed in dynamic, real-world settings, where the distribution of user prompts can shift substantially over time as new tasks, prompts, and users are introduced to a deployed model.Such natural prompt distribution shift poses a major challenge to LLM reliability, particularly for specialized models designed for narrow domains or user populations.Despite attention to out-of-distribution robustness, there is very limited exploration of measuring natural prompt distribution shift in prior work, and its impact on deployed LLMs remains poorly understood.We introduce the LLM Evaluation under Natural prompt Shift (LENS) framework: a data-centric approach for quantifying natural prompt distribution shift and evaluating its effect on the performance of deployed LLMs.We perform a large-scale evaluation using 192 real-world post-deployment prompt shift settings over time, user group, and geographic axes, training a total of 81 models on 4.68M training prompts, and evaluating on 57.6k prompts.We find that even moderate shifts in user prompt behavior correspond with large performance drops (73% average loss) in deployed LLMs.This performance degradation is particularly prevalent when users from different latent groups and geographic regions interact with models and is correlated with natural prompt distribution shift over time.We systematically characterize how LLM instruction following ability degrades over time and between user groups.Our findings highlight the critical need for data-driven monitoring to ensure LLM performance remains stable across diverse and evolving user populations.
Parker Seegmiller, Sarah Masud Preum
ACL (1)2
2025 Follow-up Question Generation For Enhanced Patient-Provider Conversations
abstract
Joseph Gatto, Parker Seegmiller, Timothy E. Burdick, Inas S. Khayal, Sarah DeLozier, Sarah Masud Preum. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Joseph Gatto, Parker Seegmiller, Timothy E. Burdick, Inas Khayal, Sarah DeLozier, Sarah Masud Preum
ACL (1)6
2025 Document-Level Event-Argument Data Augmentation for Challenging Role Types
abstract
Event Argument Extraction (EAE) is a daunting information extraction problem -with significant limitations in few-shot cross-domain (FSCD) settings.A common solution to FSCD modeling is data augmentation.Unfortunately, existing augmentation methods are not wellsuited to a variety of real-world EAE contexts, including (i) modeling long documents (documents with over 10 sentences), and (ii) modeling challenging role types (i.e., event roles with little to no training data and semantically outlying roles).We introduce two novel LLMpowered data augmentation methods for generating extractive document-level EAE samples using zero in-domain training data.We validate the generalizability of our approach on four datasets -showing significant performance increases in low-resource settings.Our highest performing models provide a 13-pt increase in F1 score on zero-shot role extraction in FSCD evaluation.
Joseph Gatto, Omar Sharif, Parker Seegmiller, Sarah Masud Preum
ACL (1)4
2024 Characterizing Information Seeking Events in Health-Related Social Discourse
abstract
Social media sites have become a popular platform for individuals to seek and share health information. Despite the progress in natural language processing for social media mining, a gap remains in analyzing health-related texts on social discourse in the context of events. Event-driven analysis can offer insights into different facets of healthcare at an individual and collective level, including treatment options, misconceptions, knowledge gaps, etc. This paper presents a paradigm to characterize health-related information-seeking in social discourse through the lens of events. Events here are board categories defined with domain experts that capture the trajectory of the treatment/medication. To illustrate the value of this approach, we analyze Reddit posts regarding medications for Opioid Use Disorder (OUD), a critical global health concern. To the best of our knowledge, this is the first attempt to define event categories for characterizing information-seeking in OUD social discourse. Guided by domain experts, we develop TREAT-ISE, a novel multilabel treatment information-seeking event dataset to analyze online discourse on an event-based framework. This dataset contains Reddit posts on information-seeking events related to recovery from OUD, where each post is annotated based on the type of events. We also establish a strong performance benchmark (77.4% F1 score) for the task by employing several machine learning and deep learning classifiers. Finally, we thoroughly investigate the performance and errors of ChatGPT on this task, providing valuable insights into the LLM's capabilities and ongoing characterization efforts.
Omar Sharif, Madhusudan Basak, Tanzia Parvin, Ava Scharfstein, Alphonso Bradham, Jacob T. Borodovsky, Sarah E. Lord, Sarah Masud Preum
AAAI8
2024 Deciphering Hate: Identifying Hateful Memes and Their Targets
abstract
Internet memes have become a powerful means for individuals to express emotions, thoughts, and perspectives on social media.While often considered a source of humor and entertainment, memes can also disseminate hateful content targeting individuals or communities.Most existing research focuses on the negative aspects of memes in high-resource languages, overlooking the distinctive challenges associated with low-resource languages like Bengali (also known as Bangla).Furthermore, while previous work on Bengali memes has focused on detecting hateful memes, there has been no work on detecting their targeted entities.To bridge this gap and facilitate research in this arena, we introduce a novel multimodal dataset for Bengali, BHM (Bengali Hateful Memes).The dataset consists of 7,148 memes with Bengali as well as code-mixed captions, tailored for two tasks: (i) detecting hateful memes, and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).To solve these tasks, we propose DORA (Dual cO-attention fRAmework), a multimodal deep neural network that systematically extracts the significant modality features from the memes and jointly evaluates them with the modality-specific features to understand the context better.Our experiments show that DORA is generalizable on other low-resource hateful meme datasets and outperforms several state-of-the-art rivaling baselines.
Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque, Sarah Masud Preum
ACL (1)4
2024 Sketching AI Concepts with Capabilities and Examples: AI Innovation in the Intensive Care Unit
abstract
Advances in artificial intelligence (AI) have enabled unprecedented capabilities, yet innovation teams struggle when envisioning AI concepts. Data science teams think of innovations users do not want, while domain experts think of innovations that cannot be built. A lack of effective ideation seems to be a breakdown point. How might multidisciplinary teams identify buildable and desirable use cases? This paper presents a first hand account of ideating AI concepts to improve critical care medicine. As a team of data scientists, clinicians, and HCI researchers, we conducted a series of design workshops to explore more effective approaches to AI concept ideation and problem formulation. We detail our process, the challenges we encountered, and practices and artifacts that proved effective. We discuss the research implications for improved collaboration and stakeholder engagement, and discuss the role HCI might play in reducing the high failure rate experienced in AI innovation.
Nur Yildirim, Susanna Zlotnikov, Deniz Sayar, Jeremy M. Kahn, Leigh A. Bukowski, Sher Shah Amin, Kathryn A. Riman, Billie S. Davis, John S. Minturn, Andrew J. King 0002, Dan Ricketts, Lu Tang 0003, Venkatesh Sivaraman, Adam Perer, Sarah Masud Preum, James McCann, John Zimmerman
CHI15
2024 Explicit, Implicit, and Scattered: Revisiting Event Extraction to Capture Complex Arguments
abstract
Prior works formulate the extraction of eventspecific arguments as a span extraction problem, where event arguments are explicit -i.e.assumed to be contiguous spans of text in a document.In this study, we revisit this definition of Event Extraction (EE) by introducing two key argument types that cannot be modeled by existing EE frameworks.First, implicit arguments are event arguments which are not explicitly mentioned in the text, but can be inferred through context.Second, scattered arguments are event arguments that are composed of information scattered throughout the text.These two argument types are crucial to elicit the full breadth of information required for proper event modeling.To support the extraction of explicit, implicit, and scattered arguments, we develop a novel dataset, DiscourseEE, which includes 7,464 argument annotations from online health discourse.Notably, 51.2% of the arguments are implicit, and 17.4% are scattered, making Dis-courseEE a unique corpus for complex event extraction.Additionally, we formulate argument extraction as a text generation problem to facilitate the extraction of complex argument types.We provide a comprehensive evaluation of state-of-the-art models and highlight critical open challenges in generative event extraction.Our data and codebase are available at https://omar-sharif03.github.io/DiscourseEE.
Omar Sharif, Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
EMNLP4
2024 Theme-Driven Keyphrase Extraction to Analyze Social Media Discourse
abstract
Social media platforms are vital resources for sharing self-reported health experiences, offering rich data on various health topics. Despite advancements in Natural Language Processing (NLP) enabling large-scale social media data analysis, a gap remains in applying keyphrase extraction to health-related content. Keyphrase extraction is used to identify salient concepts in social media discourse without being constrained by predefined entity classes. This paper introduces a theme-driven keyphrase extraction framework tailored for social media, a pioneering approach designed to capture clinically relevant keyphrases from user-generated health texts. Themes are defined as broad categories determined by the objectives of the extraction task. We formulate this novel task of theme-driven keyphrase extraction and demonstrate its potential for efficiently mining social media text for the use case of treatment for opioid use disorder. This paper leverages qualitative and quantitative analysis to demonstrate the feasibility of extracting actionable insights from social media data and efficiently extracting keyphrases using minimally supervised NLP models. Our contributions include the development of a novel data collection and curation framework for theme-driven keyphrase extraction and the creation of SuboxoPhrase, the first dataset of its kind comprising human-annotated keyphrases from a Reddit community. We also identify the scope of minimally supervised NLP models to extract keyphrases from social media data efficiently. Lastly, we found that a large language model (ChatGPT) outperforms unsupervised keyphrase extraction models, showcasing its efficacy in this task.
William Romano, Omar Sharif, Madhusudan Basak, Joseph Gatto, Sarah Masud Preum
ICWSM5
2023 Statistical Depth for Ranking and Characterizing Transformer-Based Text Embeddings
abstract
The popularity of transformer-based text embeddings calls for better statistical tools for measuring distributions of such embeddings.One such tool would be a method for ranking texts within a corpus by centrality, i.e. assigning each text a number signifying how representative that text is of the corpus as a whole.However, an intrinsic center-outward ordering of high-dimensional text representations is not trivial.A statistical depth is a function for ranking k-dimensional objects by measuring centrality with respect to some observed kdimensional distribution.We adopt a statistical depth to measure distributions of transformerbased text embeddings, transformer-based text embedding (TTE) depth, and introduce the practical use of this depth for both modeling and distributional inference in NLP pipelines.We first define TTE depth and an associated rank sum test for determining whether two corpora differ significantly in embedding space.We then use TTE depth for the task of in-context learning prompt selection, showing that this approach reliably improves performance over statistical baseline approaches across six text classification tasks.Finally, we use TTE depth and the associated rank sum test to characterize the distributions of synthesized and human-generated corpora, showing that five recent synthetic data augmentation processes cause a measurable distributional shift away from associated humangenerated text.
Parker Seegmiller, Sarah Masud Preum
EMNLP2
2023 Scope of Pre-trained Language Models for Detecting Conflicting Health Information
abstract
An increasing number of people now rely on online platforms to meet their health information needs. Thus identifying inconsistent or conflicting textual health information has become a safety-critical task. Health advice data poses a unique challenge where information that is accurate in the context of one diagnosis can be conflicting in the context of another. For example, people suffering from diabetes and hypertension often receive conflicting health advice on diet. This motivates the need for technologies which can provide contextualized, user-specific health advice. A crucial step towards contextualized advice is the ability to compare health advice statements and detect if and how they are conflicting. This is the task of health conflict detection (HCD). Given two pieces of health advice, the goal of HCD is to detect and categorize the type of conflict. It is a challenging task, as (i) automatically identifying and categorizing conflicts requires a deeper understanding of the semantics of the text, and (ii) the amount of available data is quite limited. In this study, we are the first to explore HCD in the context of pre-trained language models. We find that DeBERTa-v3 performs best with a mean F1 score of 0.68 across all experiments. We additionally investigate the challenges posed by different conflict types and how synthetic data improves a model's understanding of conflict-specific semantics. Finally, we highlight the difficulty in collecting real health conflicts and propose a human-in-the-loop synthetic data augmentation approach to expand existing HCD datasets. Our HCD training dataset is over 2x bigger than the existing HCD dataset and is made publicly available on Github.
Joseph Gatto, Madhusudan Basak, Sarah Masud Preum
ICWSM3
2023 HealthE: Recognizing Health Advice & Entities in Online Health Communities
abstract
The task of extracting and classifying entities is at the core of important Health-NLP systems such as misinformation detection, medical dialogue modeling, and patient-centric information tools. Granular knowledge of textual entities allows these systems to utilize knowledge bases, retrieve relevant information, and build graphical representations of texts. Unfortunately, most existing works on health entity recognition are trained on clinical notes, which are both lexically and semantically different from public health information found in online health resources or social media. In other words, existing health entity recognizers vastly under-represent the entities relevant to public health data, such as those provided by sites like WebMD. It is crucial that future Health-NLP systems be able to model such information, as people rely on online health advice for personal health management and clinically relevant decision making. In this work, we release a new annotated dataset, HealthE, which facilitates the large-scale analysis of online textual health advice. HealthE consists of 3,400 health advice statements with token-level entity annotations. Additionally, we release 2,256 health statements which are not health advice to facilitate health advice mining. HealthE is the first dataset with an entity-recognition label space designed for the modeling of online health advice. We motivate the need for HealthE by demonstrating the limitations of five widely-used health entity recognizers on HealthE, such as those offered by Google and Amazon. We additionally benchmark three pre-trained language models on our dataset as reference for future research. All data is made publicly available.
Joseph Gatto, Parker Seegmiller, Garrett Johnston, Madhusudan Basak, Sarah Masud Preum
ICWSM5
2023 CitySpec with shield: A secure intelligent assistant for requirement formalization
Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma
Pervasive Mob. Comput.4
2022 CitySpec: An Intelligent Assistant System for Requirement Specification in Smart Cities
abstract
An increasing number of monitoring systems have been developed in smart cities to ensure that a city's real-time operations satisfy safety and performance requirements. However, many existing city requirements are written in English with missing, inaccurate, or ambiguous information. There is a high demand for assisting city policy makers in converting human-specified requirements to machine-understandable formal specifications for monitoring systems. To tackle this limitation, we build CitySpec, the first intelligent assistant system for requirement specification in smart cities. To create CitySpec, we first collect over 1,500 real-world city requirements across different domains from over 100 cities and extract city-specific knowledge to generate a dataset of city vocabulary with 3,061 words. We also build a translation model and enhance it through requirement synthesis and develop a novel online learning framework with validation under uncertainty. The evaluation results on real-world city requirements show that CitySpec increases the sentence-level accuracy of requirement specification from 59.02 % to 86.64 %, and has strong adaptability to a new city and a new domain (e.g., F1 score for requirements in Seattle increases from 77.6 % to 93.75% with online learning).
Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma
SMARTCOMP4
2022 An Intelligent Assistant for Converting City Requirements to Formal Specification
abstract
As more and more monitoring systems have been deployed to smart cities, there comes a higher demand for converting new human-specified requirements to machine-understandable formal specifications automatically. However, these human-specific requirements are often written in English and bring missing, inaccurate, or ambiguous information. In this paper, we present City Spec [1], an intelligent assistant system for requirement specification in smart cities. CitySpec not only helps overcome the language differences brought by English requirements and formal specifications, but also offers solutions to those missing, inaccurate, or ambiguous information. The goal of this paper is to demonstrate how CitySpec works. Specifically, we present three demos: (1) interactive completion of requirements in CitySpec; (2) human-in-the-loop correction while CitySepc encounters exceptions; (3) online learning in CitySpec.
Zirong Chen, Isaac Li, Haoxiang Zhang 0003, Sarah Masud Preum, John A. Stankovic, Meiyi Ma
SMARTCOMP4
2020 EMSContExt: EMS Protocol-Driven Concept Extraction for Cognitive Assistance in Emergency Response
abstract
This paper presents a technique for automated curation of a domain-specific knowledge base or lexicon for resource-constrained domains, such as Emergency Medical Services (EMS) and its application to real-time concept extraction and cognitive assistance in emergency response. The EMS responders often verbalize critical information describing the situations at an incident scene, including patients' physical condition and medical history. Automated extraction of EMS protocol-specific concepts from responders' speech data can facilitate cognitive support through the selection and execution of the proper EMS protocols for patient treatment. Although this task is similar to the traditional NLP task of concept extraction, the underlying application domain poses major challenges, including low training resources availability (e.g., no existing EMS ontology, lexicon, or annotated EMS corpus) and domain mismatch. Hence, we develop EMSContExt, a weakly-supervised concept extraction approach for EMS concepts. It utilizes different knowledge bases and a semantic concept model based on a corpus of over 9400 EMS narratives for lexicon expansion. The expanded EMS lexicon is then used to automatically extract critical EMS protocol-specific concepts from real-time EMS speech narratives. Our experimental results show that EMSContExt achieves 0.85 recall and 0.82 F1-score for EMS concept extraction and significantly outperforms MetaMap, a state-of-the-art medical concept extraction tool. We also demonstrate the application of EMSContExt to EMS protocol selection and execution and real-time recommendation of protocol-specific interventions to the EMS responders. Here, EMSContExt outperforms MetaMap with a 6% increase and six times speedup in weighted recall and execution time, respectively.
Sarah Masud Preum, Sile Shu, Homa Alemzadeh, John A. Stankovic
AAAI1
2020 GRACE: Generating Summary Reports Automatically for Cognitive Assistance in Emergency Response
abstract
EMS (emergency medical service) plays an important role in saving lives in emergency and accident situations. When first responders, including EMS providers and firefighters, arrive at an incident, they communicate with the patients (if conscious), family members and other witnesses, other first responders, and the command center. The first responders utilize a microphone and headset to support these communications. After the incident, the first responders are required to document the incident by filling out a form. Today, this is performed manually. Manual documentation of patient summary report is time-consuming, tedious, and error-prone. We have addressed these form filling problems by transcribing the audio from the scene, identifying the relevant information from all the conversations, and automatically filling out the form. Informal survey of first responders indicate that this application would be exceedingly helpful to them. Results show that we can fill out a model summary report form with an F1 score as high as 94%, 78%, 96%, and 83% when the data is noise-free audio, noisy audio, noise-free textual narratives, and noisy textual narratives, respectively.
M. Arif Imtiazur Rahman, Sarah Masud Preum, Ronald D. Williams, Homa Alemzadeh, John A. Stankovic
AAAI2
2020 IMACS - an interactive cognitive assistant module for cardiac arrest cases in emergency medical service: demo abstract
abstract
IMACS is an intelligent, interactive cognitive assistant dedicated to cardiac arrest cases in Emergency Medical Service (EMS). EMS providers deal with many cardiac cases. IMACS interacts with EMS providers in real-time and collects vital information from the providers' conversation, including names of interventions, timestamps of interventions, and dosage amount. Throughout the process, IMACS provides necessary reminders and creates a summary report afterward. Using the dynamic behavioral model of two different cardiac arrest recovery protocols, we have developed a critical risk-index based approach to provide time-sensitive feedback and suggest alternatives to the providers in real-time. Our experiments reveal an F1-score of 83% with 300 test cases. A qualitative study also reflects that seven out of ten of the EMS providers rate the system as very helpful in correctly executing cardiac arrest EMS protocols.
M. Arif Imtiazur Rahman, Sarah Masud Preum, John A. Stankovic, Leon Jia, Eimara Mirza, Ronald D. Williams, Homa Alemzadeh
SenSys2
2020 Data Sets, Modeling, and Decision Making in Smart Cities: A Survey
abstract
Cities are deploying tens of thousands of sensors and actuators and developing a large array of smart services. The smart services use sophisticated models and decision-making policies supported by Cyber Physical Systems and Internet of Things technologies. The increasing number of sensors collects a large amount of city data across multiple domains. The collected data have great potential value, but has not yet been fully exploited. This survey focuses on the domains of transportation, environment, emergency and public safety, energy, and social sensing. This article carefully reviews both the data sets being collected across 14 smart cities and the state-of-the-art work in modeling and decision making methodologies. The article also points out the characteristics, challenges faced today, and those challenges that will be exacerbated in the future. Key data issues addressed include heterogeneity, interdisciplinary, integrity, completeness, real-timeliness, and interdependencies. Key decision making issues include safety and service conflicts, security, uncertainty, humans in the loop, and privacy.
Meiyi Ma, Sarah Masud Preum, Mohsin Y. Ahmed, William Tärneberg, Abdeltawab M. Hendawi, John A. Stankovic
ACM Trans. Cyber Phys. Syst.2
2019 A Behavior Tree Cognitive Assistant System for Emergency Medical Services
abstract
This paper presents a cognitive assistant system for emergency medical services (EMS) that can serve as a rescue robot or virtual assistant, helping with improving situational awareness of the first responders through automated collection and analysis of data from the incident scene and providing suggestions to them. The proposed system relies on a Behavior Tree (BT) framework that combines the knowledge of EMS protocol guidelines with speech recognition, natural language processing, and machine learning methods to (i) extract critical information from responders' conversations and verbalized observations, (ii) infer the incident context, and (iii) decide on safe and effective response interventions to perform. We use a data-set of 8302 real EMS call records from an urban, high volume regional ambulance agency in the U.S. to evaluate the responsiveness and cognitive ability of the system and assess the safety of the suggestions provided to the responders. The experimental results show that the developed cognitive assistant achieves an average top-3 accuracy of 89% in selecting the correct EMS protocols and an average F1-score of 71% in suggesting the protocol specific interventions while providing transparency and evidence for the suggestions.
Sile Shu, Sarah Masud Preum, Haydon M. Pitchford, Ronald D. Williams, John A. Stankovic, Homa Alemzadeh
IROS2
2018 A Corpus of Drug Usage Guidelines Annotated with Type of Advice
Sarah Masud Preum, Md. Rizwan Parvez, Kai-Wei Chang 0001, John A. Stankovic
LREC1
2017 User authentication using wrist mounted inertial sensors: poster abstract
abstract
Smart wrist devices available today like smart watches and fitness trackers are usually enriched with inertial sensors such as accelerometers and gyroscopes that can be used to capture wrist motions. This opens an opportunity to use these devices for user authentication, exploiting an important biometric trait of a user: the wrist gestures during performing a signature in the air. In contrast to traditional authentication methods, the gestures can be captured while signing in the air freely with no need for any input media like keypads and signature capture devices. This paper presents result from a preliminary study that shows the potential of the proposed approach to be used for robust user authentication.
Md. Abu Sayeed Mondol, Ifat Afrin Emi, Sarah Masud Preum, John A. Stankovic
IPSN3
2017 Conflict detection in online textual health advice: demo abstract
abstract
Textual health advice generated from different online sources (e.g., health apps and websites) can be conflicting. Conflicts can occur due to lexical features, (such as, negation, antonyms, or numerical mismatch) or can be conditioned upon time and/or physiological status. Detecting conflicts from textual health advice poses several challenges, including, large structural variation between text and hypothesis pairs, finding conceptual overlap between pairs of advice, and inference of the semantics of an advice (i.e., what to do, why, and how). In this demonstration, we present a semantic rule-based system to detect different types of conflicts in online textual health advice statements in a context-aware and interpretable manner.
Sarah Masud Preum, Md. Abu Sayeed Mondol, Meiyi Ma, Hongning Wang, John A. Stankovic
IPSN1
2017 Preclude: Conflict detection in textual health advice
abstract
With the rapid digitalization of the health sector, people often turn to mobile apps and online health websites for health advice. Health advice generated from different sources can be conflicting as they address different aspects of health (e.g., weight loss, diet, disease) or as they are unaware of the context of a user (e.g., age, gender, physiological condition). Conflicts can occur due to lexical features, (such as, negation, antonyms, or numerical mismatch) or can be conditioned upon time and/or physiological status. We formulate the problem of finding conflicting health advice and develop a comprehensive taxonomy of conflicts. While a similar research area in the natural language processing domain explores the problem of textual contradiction identification, finding conflicts in health advice poses its own unique lexical and semantic challenges. These include large structural variation between text and hypothesis pairs, finding conceptual overlap between pairs of advice, and inference of the semantics of an advice (i.e., what to do, why and how). Hence, we develop Preclude, a novel semantic rule-based solution to detect conflicting health advice derived from heterogeneous sources utilizing linguistic rules and external knowledge bases. As our solution is interpretable and comprehensive, it can guide users towards conflict resolution too. We evaluate Preclude using 1156 real advice statements covering 8 important health topics that are collected from smart phone health apps and popular health websites. Preclude results in 90% accuracy and outperforms the accuracy and F1 score of the baseline approach by about 1.5 times and 3 times, respectively.
Sarah Masud Preum, Md. Abu Sayeed Mondol, Meiyi Ma, Hongning Wang, John A. Stankovic
PerCom1
2017 Preclude2 : Personalized conflict detection in heterogeneous health applications
Sarah Masud Preum, Md. Abu Sayeed Mondol, Meiyi Ma, Hongning Wang, John A. Stankovic
Pervasive Mob. Comput.1
2016 Detection of Runtime Conflicts among Services in Smart Cities
abstract
The populations of large cities around the world are growing rapidly. Cities are beginning to address this problem by implementing significant sensing and actuation infrastructure and building services on this infrastructure. However, as the density of sensing and actuation increases and as the complexities of services grow there is an increasing potential for conflicts across Smart City services. These conflicts can cause unsafe situations and disrupt the benefits that the services were originally intended to provide. Although some of the conflicts can be detected and avoided during designing the services, many can still occur unpredictably during runtime. This paper carefully defines and enumerates the main issues regarding the detection and resolution of runtime conflicts in smart cities. In particular, it focuses on conflicts that arise across services. This issue is becoming more and more important as Smart City designs attempt to integrate services from different domains (transportation, energy, public safety, emergency, medical, and many others). Research challenges are identified and then addressed that deal with uncertainty, dynamism, real-time, mobility and spatio-temporal availability, duration and scale of effect, efficiency, and ownership. A watchdog architecture is also described that oversees the services operating in a Smart City. This watchdog solution detects and resolves conflicts, it learns and adapts, and it provides additional inputs to decision making aspects of services. Using data from a Smart City dataset, an emulated set of services and activities using those services are created to perform a conflict analysis. A second analysis hypothesizes 41 future services across 5 domains. Both of these evaluations demonstrate the high probability of conflicts in smart cities of the future.
Meiyi Ma, Sarah Masud Preum, William Tärneberg, Mohsin Y. Ahmed, Matthew Ruiters, John A. Stankovic
SMARTCOMP2
2015 MAPer: A Multi-scale Adaptive Personalized Model for Temporal Human Behavior Prediction
abstract
The primary objective of this research is to develop a simple and interpretable predictive framework to perform temporal modeling of individual user's behavior traits based on each person's past observed traits/behavior. Individual-level human behavior patterns are possibly influenced by various temporal features (e.g., lag, cycle) and vary across temporal scales (e.g., hour of the day, day of the week). Most of the existing forecasting models do not capture such multi-scale adaptive regularity of human behavior or lack interpretability due to relying on hidden variables. Hence, we build a multi-scale adaptive personalized (MAPer) model that quantifies the effect of both lag and behavior cycle for predicting future behavior. MAper includes a novel basis vector to adaptively learn behavior patterns and capture the variation of lag and cycle across multi-scale temporal contexts. We also extend MAPer to capture the interaction among multiple behaviors to improve the prediction performance.
Sarah Masud Preum, John A. Stankovic, Yanjun Qi
CIKM1
2015 Holmes: A Comprehensive Anomaly Detection System for Daily In-home Activities
abstract
Advances in wireless sensor networks have enabled the monitoring of daily activities of elderly people. The goal of these monitoring applications is to learn normal behavior in terms of daily activities and look for any deviation, i.e., Anomalies, so that alerts can be sent to relatives or caregivers. However, human behavior is very complex, and many existing anomaly detection systems are too simplistic which cause many false alarms, resulting in unreliable systems. We present Holmes, a comprehensive anomaly detection system for daily in-home activities. Holmes accurately learns a resident's normal behavior by considering variability in daily activities based not only on a per day basis, but also considering specific days of the week, different time periods such as per week and per month, and collective, temporal, and correlation based features. This approach of learning complicated normal behaviors reduces false alarms. Also, based on resident and expert feedback, Holmes learns semantic rules that explain specific variations of activities in specific scenarios to further reduce false alarms. We evaluate Holmes using data collected from our own deployed system, public data sets, and data collected by a senior safety system provider company from an elderly resident's home. Our evaluation shows that compared to state of the art systems, Holmes reduces false positives and false negatives by at least 46% and 27%, respectively.
Enamul Hoque 0002, Robert F. Dickerson, Sarah Masud Preum, Mark A. Hanson, Adam T. Barth, John A. Stankovic
DCOSS3