VLDB 2026 Research / reviewers in the wild / expert
Anahita Davoudi
dblp:178/3124
· DBLP profile ↗
10ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-4345-3889ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorArtificial intelligence and machine learning · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Identifying stigmatizing and positive/preferred language in obstetric clinical notes using natural language processingabstractOBJECTIVE: To identify stigmatizing language in obstetric clinical notes using natural language processing (NLP). MATERIALS AND METHODS: We analyzed electronic health records from birth admissions in the Northeast United States in 2017. We annotated 1771 clinical notes to generate the initial gold standard dataset. Annotators labeled for exemplars of 5 stigmatizing and 1 positive/preferred language categories. We used a semantic similarity-based search approach to expand the initial dataset by adding additional exemplars, composing an enhanced dataset. We employed traditional classifiers (Support Vector Machine, Decision Trees, and Random Forest) and a transformer-based model, ClinicalBERT (Bidirectional Encoder Representations from Transformers) and BERT base. Models were trained and validated on initial and enhanced datasets and were tested on enhanced testing dataset. RESULTS: In the initial dataset, we annotated 963 exemplars as stigmatizing or positive/preferred. The most frequently identified category was marginalized language/identities (n = 397, 41%), and the least frequent was questioning patient credibility (n = 51, 5%). After employing a semantic similarity-based search approach, 502 additional exemplars were added, increasing the number of low-frequency categories. All NLP models also showed improved performance, with Decision Trees demonstrating the greatest improvement (21%). ClinicalBERT outperformed other models, with the highest average F1-score of 0.78. DISCUSSION: Clinical BERT seems to most effectively capture the nuanced and context-dependent stigmatizing language found in obstetric clinical notes, demonstrating its potential clinical applications for real-time monitoring and alerts to prevent usages of stigmatizing language use and reduce healthcare bias. Future research should explore stigmatizing language in diverse geographic locations and clinical settings to further contribute to high-quality and equitable perinatal care. CONCLUSION: ClinicalBERT effectively captures the nuanced stigmatizing language in obstetric clinical notes. Our semantic similarity-based search approach to rapidly extract additional exemplars enhanced the performances while reducing the need for labor-intensive annotation. Jihye Kim Scroggins, Ismael I Hulchafo, Sarah Harkins, Danielle Scharp, Hans Moen, Anahita Davoudi, Kenrick Cato, Michele Tadiello, Maxim Topaz, Veronica Barcelona |
J. Am. Medical Informatics Assoc. | 6 |
| 2024 | Exploring home healthcare clinicians' needs for using clinical decision support systems for early risk warningabstractOBJECTIVES: To explore home healthcare (HHC) clinicians' needs for Clinical Decision Support Systems (CDSS) information delivery for early risk warning within HHC workflows. METHODS: Guided by the CDS "Five-Rights" framework, we conducted semi-structured interviews with multidisciplinary HHC clinicians from April 2023 to August 2023. We used deductive and inductive content analysis to investigate informants' responses regarding CDSS information delivery. RESULTS: Interviews with thirteen HHC clinicians yielded 16 codes mapping to the CDS "Five-Rights" framework (right information, right person, right format, right channel, right time) and 11 codes for unintended consequences and training needs. Clinicians favored risk levels displayed in color-coded horizontal bars, concrete risk indicators in bullet points, and actionable instructions in the existing EHR system. They preferred non-intrusive risk alerts requiring mandatory confirmation. Clinicians anticipated risk information updates aligned with patient's condition severity and their visit pace. Additionally, they requested training to understand the CDSS's underlying logic, and raised concerns about information accuracy and data privacy. DISCUSSION: While recognizing CDSS's value in enhancing early risk warning, clinicians highlighted concerns about increased workload, alert fatigue, and CDSS misuse. The top risk factors identified by machine learning algorithms, especially text features, can be ambiguous due to a lack of context. Future research should ensure that CDSS outputs align with clinical evidence and are explainable. CONCLUSION: This study identified HHC clinicians' expectations, preferences, adaptations, and unintended uses of CDSS for early risk warning. Our findings endorse operationalizing the CDS "Five-Rights" framework to optimize CDSS information delivery and integration into HHC workflows. Zidu Xu, Lauren Evans, Jiyoun Song, Sena Chae, Anahita Davoudi, Kathryn H. Bowles, Margaret V. McDonald, Maxim Topaz |
J. Am. Medical Informatics Assoc. | 5 |
| 2023 | Predicting emergency department visits and hospitalizations for patients with heart failure in home healthcare using a time series risk modelabstractOBJECTIVES: Little is known about proactive risk assessment concerning emergency department (ED) visits and hospitalizations in patients with heart failure (HF) who receive home healthcare (HHC) services. This study developed a time series risk model for predicting ED visits and hospitalizations in patients with HF using longitudinal electronic health record data. We also explored which data sources yield the best-performing models over various time windows. MATERIALS AND METHODS: We used data collected from 9362 patients from a large HHC agency. We iteratively developed risk models using both structured (eg, standard assessment tools, vital signs, visit characteristics) and unstructured data (eg, clinical notes). Seven specific sets of variables included: (1) the Outcome and Assessment Information Set, (2) vital signs, (3) visit characteristics, (4) rule-based natural language processing-derived variables, (5) term frequency-inverse document frequency variables, (6) Bio-Clinical Bidirectional Encoder Representations from Transformers variables, and (7) topic modeling. Risk models were developed for 18 time windows (1-15, 30, 45, and 60 days) before an ED visit or hospitalization. Risk prediction performances were compared using recall, precision, accuracy, F1, and area under the receiver operating curve (AUC). RESULTS: The best-performing model was built using a combination of all 7 sets of variables and the time window of 4 days before an ED visit or hospitalization (AUC = 0.89 and F1 = 0.69). DISCUSSION AND CONCLUSION: This prediction model suggests that HHC clinicians can identify patients with HF at risk for visiting the ED or hospitalization within 4 days before the event, allowing for earlier targeted interventions. Sena Chae, Anahita Davoudi, Jiyoun Song, Lauren Evans, Mollie Hobensack, Kathryn H. Bowles, Margaret V. McDonald, Yolanda Barrón, Sarah Collins Rossetti, Kenrick Cato, Sridevi Sridharan, Maxim Topaz |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Uncovering hidden trends: identifying time trajectories in risk factors documented in clinical notes and predicting hospitalizations and emergency department visits during home health careabstractOBJECTIVE: This study aimed to identify temporal risk factor patterns documented in home health care (HHC) clinical notes and examine their association with hospitalizations or emergency department (ED) visits. MATERIALS AND METHODS: Data for 73 350 episodes of care from one large HHC organization were analyzed using dynamic time warping and hierarchical clustering analysis to identify the temporal patterns of risk factors documented in clinical notes. The Omaha System nursing terminology represented risk factors. First, clinical characteristics were compared between clusters. Next, multivariate logistic regression was used to examine the association between clusters and risk for hospitalizations or ED visits. Omaha System domains corresponding to risk factors were analyzed and described in each cluster. RESULTS: Six temporal clusters emerged, showing different patterns in how risk factors were documented over time. Patients with a steep increase in documented risk factors over time had a 3 times higher likelihood of hospitalization or ED visit than patients with no documented risk factors. Most risk factors belonged to the physiological domain, and only a few were in the environmental domain. DISCUSSION: An analysis of risk factor trajectories reflects a patient's evolving health status during a HHC episode. Using standardized nursing terminology, this study provided new insights into the complex temporal dynamics of HHC, which may lead to improved patient outcomes through better treatment and management plans. CONCLUSION: Incorporating temporal patterns in documented risk factors and their clusters into early warning systems may activate interventions to prevent hospitalizations or ED visits in HHC. Jiyoun Song, Se Hee Min, Sena Chae, Kathryn H. Bowles, Margaret V. McDonald, Mollie Hobensack, Yolanda Barrón, Sridevi Sridharan, Anahita Davoudi, Sungho Oh, Lauren Evans, Maxim Topaz |
J. Am. Medical Informatics Assoc. | 9 |
| 2022 | Identifying Barriers to Post-Acute Care Referral and Characterizing Negative Patient Preferences Among Hospitalized Older Adults Using Natural Language Processing
Erin E. Kennedy, Anahita Davoudi, Sy Hwang, Ryan J. Urbanowicz, Philip J. Freda, Kathryn H. Bowles, Danielle L. Mowery |
AMIA | 2 |
| 2020 | A Preliminary Characterization of Canonicalized and Non-Canonicalized Section Headers Across Variable Clinical Note Types
Shun Yu, Anahita Davoudi, Danielle L. Mowery |
AMIA | 3 |
| 2017 | Detection of profile injection attacks in social recommender systems using outlier analysisabstractAs systems based on social networks grow, they get affected by huge number of fake user profiles. Particularly, social recommender systems are vulnerable to profile injection attacks where malicious profiles are injected into the rating system to affect user's opinion. The objective of attackers is to inject a large set of biased profiles that provide favorable or unfavorable recommendations for a product. In this paper, we propose a classification technique for detection of attackers. First, we define the attributes that provide the likelihood of a user having a profile of that of an attacker. Using user-item rating matrix, user-connection matrix, and similarity between users, we find if the ratings are abnormal and if there are random connections in the network. Then, we use fc-means clustering to categorize users into authentic users and attackers. To evaluate our framework, we use Epinions dataset and inject intelligent push and nuke attacks. These attacks make arbitrary connections to existing users and provide biased ratings. To evaluate the performance, we use precision and recall to show that fc-means clustering can identify the attackers with high accuracy and low false positives. Anahita Davoudi, Mainak Chatterjee |
IEEE BigData | 1 |
| 2017 | Effects of User Interactions on Online Social Recommender SystemsabstractWe analyze online social data to model social interactions of users in recommender systems: i) Rating prediction, and ii) detecting spammers and abnormal user rating behaviors. We propose a social trust model using matrix factorization method to estimate users taste by incorporating user-item matrix. The effect of users friends tastes is modeled based on centrality metrics and similarity algorithms between users. The proposed method is validated using Epinions Dataset. To identify abnormal users in social recommender systems, we propose a classification approach. We define attributes to provide likelihood of a user having a profile of that of an attacker. Using user-item rating matrix and user-connection matrix, we find if the ratings are abnormal and if connections are random. We use k-means clustering to categorize users into authentic users and attackers. We use Epinions dataset to test the profile injection attacks. Anahita Davoudi |
ICDE | 1 |
| 2016 | Prediction of information diffusion in social networks using dynamic carrying capacityabstractOnline social networks have become an effective channel for influencing millions of users by facilitating exchange and spread of information. Despite recent works on modeling information diffusion in social networks, the complexity of social interactions makes quantification of any spreading phenomenon in social networks a challenging task. Most of the research in this area rely on empirical or statistical approaches without considering the temporal aspects and the carrying capacity of the networks. In this paper, we capture the temporal evolution of information spread in a social network using linear ordinary differential equations (ODEs). Our proposed model shows the influence of users and their temporal actions on the carrying capacity. We validate the diffusion process across the network using a dataset collected from Digg which is a popular social news sharing website. The results show that our dynamic carrying capacity PDE model is able to predict with high accuracy how the information diffuses in the network during the different phases of the lifetime of a news story. We also propose a model to represent the carrying capacity based on the portion of the influenced users. Anahita Davoudi, Mainak Chatterjee |
IEEE BigData | 1 |
| 2016 | Product rating prediction using trust relationships in social networksabstractTraditional recommender systems assume that all users are independent and identically distributed, and ignores the social interactions and connections between users. These issues hinder the recommender systems from providing more personalized recommendations to the users. In this paper, we propose a social trust model and use the probabilistic matrix factorization method to estimate users taste by incorporating user-item rating matrix. The effect of users friends tastes is modeled using a trust model which is defined based on importance (i.e., centrality) and similarity between users. Similarity is modeled using Vector Space Similarity (VSS) algorithm and centrality is quantified using two different centrality measures (degree and eigen-vector centrality). To validate the proposed method, rating estimation is performed on the Epinions dataset. Experiments show that our method provides better prediction when using trust relationship based on centrality and similarity values rather than using the binary values. The contributions of centrality and similarity in the trust values vary with different measures of centrality. Anahita Davoudi, Mainak Chatterjee |
CCNC | 1 |