EDBT 2026 Demo / reviewers in the wild / expert
Mary Regina Boland
dblp:119/7465
· DBLP profile ↗
26ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0001-8576-6408ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IRIS: Interpretable Risk Clustering Intelligence for Survival AnalysisabstractSurvival analysis models have evolved significantly with deep learning approaches, yet often lack interpretability and meaningful risk stratification capabilities. We present Interpretable Risk Clustering Intelligence for Survival Analysis (IRIS), a novel framework that addresses the critical task of risk clustering while enhancing both input-level and model-body interpretability. Unlike traditional survival models that perform post-hoc risk clustering, IRIS learns to cluster patients into meaningful risk groups directly from data while providing transparent feature importance estimation through feature contribution functions. We validate IRIS on several benchmark datasets, a real-world Alzheimer's disease dataset, and an electronic health record dataset, showing superior performance in risk clustering and predictive reliability with only a modest decrease in time-to-event prediction accuracy compared to state-of-the-art methods. Our results show that IRIS successfully balances the trade-off between interpretability and prediction performance in risk-based survival analysis, offering clinicians actionable insights for treatment planning and resource allocation. Kazi Noshin, Bojian Hou, Mary Regina Boland, Zixuan Wen, Boning Tong, Li Shen 0001, Aidong Zhang 0001 |
IEEE Big Data | 3 |
| 2025 | Predicting explainable dementia types with LLM-aided feature engineeringabstractMOTIVATION: The integration of Machine Learning and Artificial Intelligence (AI) into healthcare has immense potential due to the rapidly growing volume of clinical data. However, existing AI models, particularly Large Language Models (LLMs) like GPT-4, face significant challenges in terms of explainability and reliability, particularly in high-stakes domains like healthcare. RESULTS: This paper proposes a novel LLM-aided feature engineering approach that enhances interpretability by extracting clinically relevant features from the Oxford Textbook of Medicine. By converting clinical notes into concept vector representations and employing a linear classifier, our method achieved an accuracy of 0.72, outperforming a traditional n-gram Logistic Regression baseline (0.64) and the GPT-4 baseline (0.48), while focusing on high-level clinical features. We also explore using Text Embeddings to reduce the overall time and cost of our approach by 97%. AVAILABILITY AND IMPLEMENTATION: All code relevant to this paper is available at: https://github.com/AdityaKashyap423/Dementia_LLM_Feature_Engineering/tree/main. Aditya Kashyap, Delip Rao, Mary Regina Boland, Li Shen 0001, Chris Callison-Burch |
Bioinform. | 3 |
| 2024 | One-shot distributed algorithms for addressing heterogeneity in competing risks data across clinical sites
Dazheng Zhang, Jiayi Tong, Ronen Stein, Naimin Jing, Mary Regina Boland, Chongliang Luo, Robert N. Baldassano, Raymond J. Carroll, Christopher B. Forrest, Yong Chen 0016 |
J. Biomed. Informatics | 7 |
| 2023 | Scalable high-dimensional Bayesian varying coefficient models with unknown within-subject covarianceabstractNonparametric varying coefficient (NVC) models are useful for modeling time-varying effects on responses that are measured repeatedly for the same subjects. When the number of covariates is moderate or large, it is desirable to perform variable selection from the varying coefficient functions. However, existing methods for variable selection in NVC models either fail to account for within-subject correlations or require the practitioner to specify a parametric form for the correlation structure. In this paper, we introduce the nonparametric varying coefficient spike-and-slab lasso (NVC-SSL) for Bayesian high dimensional NVC models. Through the introduction of functional random effects, our method allows for flexible modeling of within-subject correlations without needing to specify a parametric covariance function. We further propose several scalable optimization and Markov chain Monte Carlo (MCMC) algorithms. For variable selection, we propose an Expectation Conditional Maximization (ECM) algorithm to rapidly obtain maximum a posteriori (MAP) estimates. Our ECM algorithm scales linearly in the total number of observations $N$ and the number of covariates $p$. For uncertainty quantification, we introduce an approximate MCMC algorithm that also scales linearly in both $N$ and $p$. We demonstrate the scalability, variable selection performance, and inferential capabilities of our method through simulations and a real data application. These algorithms are implemented in the publicly available R package NVCSSL on the Comprehensive R Archive Network. Ray Bai, Mary Regina Boland, Yong Chen 0016 |
J. Mach. Learn. Res. | 2 |
| 2022 | Informatics for sex- and gender-related health: understanding the problems, developing new methods, and designing new solutionsabstractInvestment in medical research continues to grow; however, the sex and gender gap of health outcomes persists1,2 with poor support for the health of cisgender and transgender women,3–5 intersex people,6 and all gender-diverse people (ie, people whose gender identity and sex assigned at birth do not fully align). Informatics approaches have the potential to identify, address, and mitigate these disparities. For example, methods that elucidate the impact of sex as a biological variable, clinical decision support (CDS) systems, and personal health informatics tools are needed to specifically fit the needs of women, intersex people, and all gender-diverse people. The goal of this issue is to highlight such informatics research. Included in this issue are 19 outstanding articles that focus on 4 key areas including gender disparities, gender diversity, maternal health, and sex differences (Table 1). The majority of these articles were focused in the clinical informatics domain with 13 clinical informatics papers. The remaining 6 papers were from the fields of consumer health informatics (N = 4) and translational informatics (N = 2). We also organized the included articles by sex- and gender-related health themes focusing on 4 areas: gender disparities, gender diversity, maternal health, and sex differences. Interestingly, the majority of the consumer health informatics papers (3 out of the 4 included in this issue) were studying gender disparities. No consumer health informatics papers focused on maternal health, and therefore, this could be an area that warrants further investigation. Mary Regina Boland, Noémie Elhadad, Wanda Pratt |
J. Am. Medical Informatics Assoc. | 1 |
| 2022 | Neighborhood deprivation increases the risk of Post-induction cesarean deliveryabstractOBJECTIVE: The purpose of this study was to measure the association between neighborhood deprivation and cesarean delivery following labor induction among people delivering at term (≥37 weeks of gestation). MATERIALS AND METHODS: We conducted a retrospective cohort study of people ≥37 weeks of gestation, with a live, singleton gestation, who underwent labor induction from 2010 to 2017 at Penn Medicine. We excluded people with a prior cesarean delivery and those with missing geocoding information. Our primary exposure was a nationally validated Area Deprivation Index with scores ranging from 1 to 100 (least to most deprived). We used a generalized linear mixed model to calculate the odds of postinduction cesarean delivery among people in 4 equally-spaced levels of neighborhood deprivation. We also conducted a sensitivity analysis with residential mobility. RESULTS: Our cohort contained 8672 people receiving an induction at Penn Medicine. After adjustment for confounders, we found that people living in the most deprived neighborhoods were at a 29% increased risk of post-induction cesarean delivery (adjusted odds ratio = 1.29, 95% confidence interval, 1.05-1.57) compared to the least deprived. In a sensitivity analysis, including residential mobility seemed to magnify the effect sizes of the association between neighborhood deprivation and postinduction cesarean delivery, but this information was only available for a subset of people. CONCLUSIONS: People living in neighborhoods with higher deprivation had higher odds of postinduction cesarean delivery compared to people living in less deprived neighborhoods. This work represents an important first step in understanding the impact of disadvantaged neighborhoods on adverse delivery outcomes. Jessica R. Meeker, Heather Burris, Ray Bai, Lisa Levine, Mary Regina Boland |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Towards deep phenotyping pregnancy: a systematic review on artificial intelligence and machine learning methods to improve pregnancy outcomesabstractOBJECTIVE: Development of novel informatics methods focused on improving pregnancy outcomes remains an active area of research. The purpose of this study is to systematically review the ways that artificial intelligence (AI) and machine learning (ML), including deep learning (DL), methodologies can inform patient care during pregnancy and improve outcomes. MATERIALS AND METHODS: We searched English articles on EMBASE, PubMed and SCOPUS. Search terms included ML, AI, pregnancy and informatics. We included research articles and book chapters, excluding conference papers, editorials and notes. RESULTS: We identified 127 distinct studies from our queries that were relevant to our topic and included in the review. We found that supervised learning methods were more popular (n = 69) than unsupervised methods (n = 9). Popular methods included support vector machines (n = 30), artificial neural networks (n = 22), regression analysis (n = 17) and random forests (n = 16). Methods such as DL are beginning to gain traction (n = 13). Common areas within the pregnancy domain where AI and ML methods were used the most include prenatal care (e.g. fetal anomalies, placental functioning) (n = 73); perinatal care, birth and delivery (n = 20); and preterm birth (n = 13). Efforts to translate AI into clinical care include clinical decision support systems (n = 24) and mobile health applications (n = 9). CONCLUSIONS: Overall, we found that ML and AI methods are being employed to optimize pregnancy outcomes, including modern DL methods (n = 13). Future research should focus on less-studied pregnancy domain areas, including postnatal and postpartum care (n = 2). Also, more work on clinical adoption of AI methods and the ethical implications of such adoption is needed. Lena M. Davidson, Mary Regina Boland |
Briefings Bioinform. | 2 |
| 2020 | Harnessing Electronic Health Records to Study Emerging Environmental Disasters: A Proof of Concept with PFAS
Mary Regina Boland, Lena M. Davidson, Silvia P. Canelón, Jessica R. Meeker, Trevor M. Penning, John H. Holmes, Jason H. Moore |
AMIA | 1 |
| 2020 | Development and Evaluation of an Algorithm to Automatically Extract Delivery Episodes from Electronic Health Records
Silvia P. Canelón, Heather Burris, Lisa Levine, Mary Regina Boland |
AMIA | 4 |
| 2020 | Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithmabstractOBJECTIVES: We propose a one-shot, privacy-preserving distributed algorithm to perform logistic regression (ODAL) across multiple clinical sites. MATERIALS AND METHODS: ODAL effectively utilizes the information from the local site (where the patient-level data are accessible) and incorporates the first-order (ODAL1) and second-order (ODAL2) gradients of the likelihood function from other sites to construct an estimator without requiring iterative communication across sites or transferring patient-level data. We evaluated ODAL via extensive simulation studies and an application to a dataset from the University of Pennsylvania Health System. The estimation accuracy was evaluated by comparing it with the estimator based on the combined individual participant data or pooled data (ie, gold standard). RESULTS: Our simulation studies revealed that the relative estimation bias of ODAL1 compared with the pooled estimates was <3%, and the ratio of standard errors was <1.25 for all scenarios. ODAL2 achieved higher accuracy (with relative bias <0.1% and ratio of standard errors <1.05). In real data analysis, we investigated the associations of 100 medications with fetal loss during pregnancy. We found that ODAL1 provided estimates with relative bias <10% for 85% of medications, and ODAL2 has relative bias <10% for 99% of medications. For communication cost, ODAL1 requires transferring p numbers from each site to the local site and ODAL2 requires transferring (p×p+p) numbers from each site to the local site, where p is the number of parameters in the regression model. CONCLUSIONS: This study demonstrates that ODAL is privacy-preserving and communication-efficient with small bias and high statistical efficiency. Rui Duan 0004, Mary Regina Boland, Howard H. Chang, Hua Xu 0001, Haitao Chu, Christopher H. Schmid, Christopher B. Forrest, John H. Holmes, Martijn J. Schuemie, Jesse A. Berlin, Jason H. Moore, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Learning from local to global: An efficient distributed algorithm for modeling time-to-event dataabstractOBJECTIVE: We developed and evaluated a privacy-preserving One-shot Distributed Algorithm to fit a multicenter Cox proportional hazards model (ODAC) without sharing patient-level information across sites. MATERIALS AND METHODS: Using patient-level data from a single site combined with only aggregated information from other sites, we constructed a surrogate likelihood function, approximating the Cox partial likelihood function obtained using patient-level data from all sites. By maximizing the surrogate likelihood function, each site obtained a local estimate of the model parameter, and the ODAC estimator was constructed as a weighted average of all the local estimates. We evaluated the performance of ODAC with (1) a simulation study and (2) a real-world use case study using 4 datasets from the Observational Health Data Sciences and Informatics network. RESULTS: On the one hand, our simulation study showed that ODAC provided estimates nearly the same as the estimator obtained by analyzing, in a single dataset, the combined patient-level data from all sites (ie, the pooled estimator). The relative bias was <0.1% across all scenarios. The accuracy of ODAC remained high across different sample sizes and event rates. On the other hand, the meta-analysis estimator, which was obtained by the inverse variance weighted average of the site-specific estimates, had substantial bias when the event rate is <5%, with the relative bias reaching 20% when the event rate is 1%. In the Observational Health Data Sciences and Informatics network application, the ODAC estimates have a relative bias <5% for 15 out of 16 log hazard ratios, whereas the meta-analysis estimates had substantially higher bias than ODAC. CONCLUSIONS: ODAC is a privacy-preserving and noniterative method for implementing time-to-event analyses across multiple sites. It provides estimates on par with the pooled estimator and substantially outperforms the meta-analysis estimator when the event is uncommon, making it extremely suitable for studying rare events and diseases in a distributed manner. Rui Duan 0004, Chongliang Luo, Martijn J. Schuemie, Jiayi Tong, C. Jason Liang, Howard H. Chang, Mary Regina Boland, Jiang Bian 0001, Hua Xu 0001, John H. Holmes, Christopher B. Forrest, Sally C. Morton, Jesse A. Berlin, Jason H. Moore, Kevin B. Mahoney, Yong Chen 0016 |
J. Am. Medical Informatics Assoc. | 7 |
| 2019 | Investigating Pregnancy-Related Health Outcomes Among Patients with Sickle Cell Disease and Linking with Health Disparities
Silvia P. Canelón, Mary Regina Boland |
AMIA | 2 |
| 2018 | Construction of PEPPER: Prenatal Exposure Pubmed ParsER
Mary Regina Boland, Aditya Kashyap, Jiadi Xiong, John H. Holmes, Scott Lorch |
AMIA | 1 |
| 2018 | Development and validation of the PEPPER framework (Prenatal Exposure PubMed ParsER) with applications to food additivesabstractBackground: Globally, 36% of deaths among children can be attributed to environmental factors. However, no comprehensive list of environmental exposures exists. We seek to address this gap by developing a literature-mining algorithm to catalog prenatal environmental exposures. Methods: We designed a framework called. PEPPER: Prenatal Exposure PubMed ParsER to a) catalog prenatal exposures studied in the literature and b) identify study type. Using PubMed Central, PEPPER classifies article type (methodology, systematic review) and catalogs prenatal exposures. We coupled PEPPER with the FDA's food additive database to form a master set of exposures. Results: We found that of 31 764 prenatal exposure studies only 53.0% were methodology studies. PEPPER consists of 219 prenatal exposures, including a common set of 43 exposures. PEPPER captured prenatal exposures from 56.4% of methodology studies (9492/16 832 studies). Two raters independently reviewed 50 randomly selected articles and annotated presence of exposures and study methodology type. Error rates for PEPPER's exposure assignment ranged from 0.56% to 1.30% depending on the rater. Evaluation of the study type assignment showed agreement ranging from 96% to 100% (kappa = 0.909, p < .001). Using a gold-standard set of relevant prenatal exposure studies, PEPPER achieved a recall of 94.4%. Conclusions: Using curated exposures and food additives; PEPPER provides the first comprehensive list of 219 prenatal exposures studied in methodology papers. On average, 1.45 exposures were investigated per study. PEPPER successfully distinguished article type for all prenatal studies allowing literature gaps to be easily identified. Mary Regina Boland, Aditya Kashyap, Jiadi Xiong, John H. Holmes, Scott Lorch |
J. Am. Medical Informatics Assoc. | 1 |
| 2018 | Uncovering exposures responsible for birth season - disease effects: a global studyabstractOBJECTIVE: Birth month and climate impact lifetime disease risk, while the underlying exposures remain largely elusive. We seek to uncover distal risk factors underlying these relationships by probing the relationship between global exposure variance and disease risk variance by birth season. MATERIAL AND METHODS: This study utilizes electronic health record data from 6 sites representing 10.5 million individuals in 3 countries (United States, South Korea, and Taiwan). We obtained birth month-disease risk curves from each site in a case-control manner. Next, we correlated each birth month-disease risk curve with each exposure. A meta-analysis was then performed of correlations across sites. This allowed us to identify the most significant birth month-exposure relationships supported by all 6 sites while adjusting for multiplicity. We also successfully distinguish relative age effects (a cultural effect) from environmental exposures. RESULTS: Attention deficit hyperactivity disorder was the only identified relative age association. Our methods identified several culprit exposures that correspond well with the literature in the field. These include a link between first-trimester exposure to carbon monoxide and increased risk of depressive disorder (R = 0.725, confidence interval [95% CI], 0.529-0.847), first-trimester exposure to fine air particulates and increased risk of atrial fibrillation (R = 0.564, 95% CI, 0.363-0.715), and decreased exposure to sunlight during the third trimester and increased risk of type 2 diabetes mellitus (R = -0.816, 95% CI, -0.5767, -0.929). CONCLUSION: A global study of birth month-disease relationships reveals distal risk factors involved in causal biological pathways that underlie them. Mary Regina Boland, Pradipta Parhi, Li Li 0062, Riccardo Miotto, Robert J. Carroll, Usman Iqbal, Phung Anh Nguyen, Martijn J. Schuemie, Seng Chan You, Donahue Smith, Sean D. Mooney, Patrick B. Ryan, Yu-Chuan Li, Rae Woong Park, Joshua C. Denny, Joel Dudley, George Hripcsak, Pierre Gentine, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 1 |
| 2017 | Biomedical informatics advancing the national health agenda: the AMIA 2015 year-in-review in clinical and consumer informaticsabstractThe field of biomedical informatics experienced a productive 2015 in terms of research. In order to highlight the accomplishments of that research, elicit trends, and identify shortcomings at a macro level, a 19-person team conducted an extensive review of the literature in clinical and consumer informatics. The result of this process included a year-in-review presentation at the American Medical Informatics Association Annual Symposium and a written report (see supplemental data). Key findings are detailed in the report and summarized here. This article organizes the clinical and consumer health informatics research from 2015 under 3 themes: the electronic health record (EHR), the learning health system (LHS), and consumer engagement. Key findings include the following: (1) There are significant advances in establishing policies for EHR feature implementation, but increased interoperability is necessary for these to gain traction. (2) Decision support systems improve practice behaviors, but evidence of their impact on clinical outcomes is still lacking. (3) Progress in natural language processing (NLP) suggests that we are approaching but have not yet achieved truly interactive NLP systems. (4) Prediction models are becoming more robust but remain hampered by the lack of interoperable clinical data records. (5) Consumers can and will use mobile applications for improved engagement, yet EHR integration remains elusive. Kirk Roberts, Mary Regina Boland, Lisiane Pruinelli, Jina J. Dcruz, Andrew B. L. Berry, Mattias Georgsson, Rebecca Hazen, Raymond Francis Sarmiento, Uba Backonja, Kun-Hsing Yu, Patricia Flatley Brennan |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Ten Simple Rules to Enable Multi-site Collaborations through Data SharingabstractOpen access, open data, and software are critical for advancing science and enabling collaboration across multiple institutions and throughout the world.Despite near universal recognition of its importance, major barriers still exist to sharing raw data, software, and research products throughout the scientific community.Many of these barriers vary by specialty [1], increasing the difficulties for interdisciplinary and/or translational researchers to engage in collaborative research.Multi-site collaborations are vital for increasing both the impact and the generalizability of research results.However, they often present unique data sharing challenges.We discuss enabling multi-site collaborations through enhanced data sharing in this set of Ten Simple Rules.Collaboration is an essential component of research [2] that takes many forms, including internal (across departments within a single institution) and external collaborations (across institutions).However, multi-site collaborations with more than two institutions encounter more complex challenges because of institutional-specific restrictions and guidelines [3].Vicens and Bourne focus on collaborators working together on a shared research grant [4].They do not discuss the specific complexities of multi-site collaborations and the vital need for enhanced data sharing in the multi-site and large-scale collaboration context, in which participants may or may not have the same funding source and/or research grant.While challenging, multi-site collaborations are equally rewarding and result in increased research productivity [5,6].One highly successful multi-site and translational collaboration is the Electronic Medical Records and Genomics (eMERGE) network (URL: https://emerge.mc. vanderbilt.edu/)initiated in 2007 [7].The eMERGE network links biorepository data with clinical information from Electronic Health Records (EHRs).They were able to find novel associations and replicate many known associations between genetic variants and clinical phenotypes that would have been more difficult without the collaboration [8].eMERGE members also collaborated with other consortiums and networks, including the Alzheimer's Disease Genetics Consortium [9] and the NINDS Stroke Genetics Network [10], to name a few.Other successful collaborations include OHDSI: Observational Health Data Sciences and Informatics (http://www.ohdsi.org/),which builds off of the methodology from the Observational Medical Outcomes Partnership (OMOP) [11], and CIRCLE: Clinical Informatics Research Collaborative (http://circleinformatics.org/).In genetics, there are many consortiums, including ExAC: The Exome Aggregation Consortium (http://exac.broadinstitute.org/), the 1000 Genomes Project Consortium (http://www.1000genomes.org/),the Australian BioGRID Mary Regina Boland, Konrad J. Karczewski, Nicholas P. Tatonetti |
PLoS Comput. Biol. | 1 |
| 2016 | The digital revolution in phenotypingabstractPhenotypes have gained increased notoriety in the clinical and biological domain owing to their application in numerous areas such as the discovery of disease genes and drug targets, phylogenetics and pharmacogenomics. Phenotypes, defined as observable characteristics of organisms, can be seen as one of the bridges that lead to a translation of experimental findings into clinical applications and thereby support 'bench to bedside' efforts. However, to build this translational bridge, a common and universal understanding of phenotypes is required that goes beyond domain-specific definitions. To achieve this ambitious goal, a digital revolution is ongoing that enables the encoding of data in computer-readable formats and the data storage in specialized repositories, ready for integration, enabling translational research. While phenome research is an ongoing endeavor, the true potential hidden in the currently available data still needs to be unlocked, offering exciting opportunities for the forthcoming years. Here, we provide insights into the state-of-the-art in digital phenotyping, by means of representing, acquiring and analyzing phenotype data. In addition, we provide visions of this field for future research work that could enable better applications of phenotype data. Anika Oellrich, Nigel Collier, Tudor Groza, Dietrich Rebholz-Schuhmann, Nigam H. Shah, Olivier Bodenreider, Mary Regina Boland, Ivo I. Georgiev, Kevin M. Livingston, Augustin Luna, Ann-Marie Mallon, Prashanti Manda, Peter N. Robinson, Gabriella Rustici, Michelle Simon, Rainer Winnenburg, Michel Dumontier |
Briefings Bioinform. | 7 |
| 2016 | Improving condition severity classification with an efficient active learning based framework
Nir Nissim, Mary Regina Boland, Nicholas P. Tatonetti, Yuval Elovici, George Hripcsak, Yuval Shahar, Robert Moskovitch |
J. Biomed. Informatics | 2 |
| 2015 | An Active Learning Framework for Efficient Condition Severity Classification
Nir Nissim, Mary Regina Boland, Robert Moskovitch, Nicholas P. Tatonetti, Yuval Elovici, Yuval Shahar, George Hripcsak |
AIME | 2 |
| 2015 | Birth month affects lifetime disease risk: a phenome-wide methodabstractOBJECTIVE: An individual's birth month has a significant impact on the diseases they develop during their lifetime. Previous studies reveal relationships between birth month and several diseases including atherothrombosis, asthma, attention deficit hyperactivity disorder, and myopia, leaving most diseases completely unexplored. This retrospective population study systematically explores the relationship between seasonal affects at birth and lifetime disease risk for 1688 conditions. METHODS: We developed a hypothesis-free method that minimizes publication and disease selection biases by systematically investigating disease-birth month patterns across all conditions. Our dataset includes 1 749 400 individuals with records at New York-Presbyterian/Columbia University Medical Center born between 1900 and 2000 inclusive. We modeled associations between birth month and 1688 diseases using logistic regression. Significance was tested using a chi-squared test with multiplicity correction. RESULTS: We found 55 diseases that were significantly dependent on birth month. Of these 19 were previously reported in the literature (P < .001), 20 were for conditions with close relationships to those reported, and 16 were previously unreported. We found distinct incidence patterns across disease categories. CONCLUSIONS: Lifetime disease risk is affected by birth month. Seasonally dependent early developmental mechanisms may play a role in increasing lifetime risk of disease. Mary Regina Boland, Zach Shahn, David Madigan, George Hripcsak, Nicholas P. Tatonetti |
J. Am. Medical Informatics Assoc. | 1 |
| 2014 | From expert-derived user needs to user-perceived ease of use and usefulness: A two-phase mixed-methods evaluation framework
Mary Regina Boland, Alex Rusanov, Yat So, Carlos Lopez-Jimenez, Linda Busacca, Richard C. Steinman, Suzanne Bakken, J. Thomas Bigger, Chunhua Weng |
J. Biomed. Informatics | 1 |
| 2014 | Clustering clinical trials with similar eligibility criteria features
Tianyong Hao, Alex Rusanov, Mary Regina Boland, Chunhua Weng |
J. Biomed. Informatics | 3 |
| 2013 | Design and Evaluation of a Bacterial Clinical Infectious Diseases Ontology
Claire L. Gordon, Stephanie Pouch, Lindsay G. Cowell, Mary Regina Boland, Heather L. Platt, Albert Goldfain, Chunhua Weng |
AMIA | 4 |
| 2013 | An Integrated Model for Patient Care and Clinical Trials (IMPACT) to support clinical research visit scheduling workflow for future learning health systemsabstractWe describe a clinical research visit scheduling system that can potentially coordinate clinical research visits with patient care visits and increase efficiency at clinical sites where clinical and research activities occur simultaneously. Participatory Design methods were applied to support requirements engineering and to create this software called Integrated Model for Patient Care and Clinical Trials (IMPACT). Using a multi-user constraint satisfaction and resource optimization algorithm, IMPACT automatically synthesizes temporal availability of various research resources and recommends the optimal dates and times for pending research visits. We conducted scenario-based evaluations with 10 clinical research coordinators (CRCs) from diverse clinical research settings to assess the usefulness, feasibility, and user acceptance of IMPACT. We obtained qualitative feedback using semi-structured interviews with the CRCs. Most CRCs acknowledged the usefulness of IMPACT features. Support for collaboration within research teams and interoperability with electronic health records and clinical trial management systems were highly requested features. Overall, IMPACT received satisfactory user acceptance and proves to be potentially useful for a variety of clinical research settings. Our future work includes comparing the effectiveness of IMPACT with that of existing scheduling solutions on the market and conducting field tests to formally assess user adoption. Chunhua Weng, Solomon Berhe, Mary Regina Boland, Junfeng Gao, Gregory William Hruby, Richard C. Steinman, Carlos Lopez-Jimenez, Linda Busacca, George Hripcsak, Suzanne Bakken, J. Thomas Bigger |
J. Biomed. Informatics | 4 |
| 2011 | EliXR: an approach to eligibility criteria extraction and representationabstractOBJECTIVE: To develop a semantic representation for clinical research eligibility criteria to automate semistructured information extraction from eligibility criteria text. MATERIALS AND METHODS: An analysis pipeline called eligibility criteria extraction and representation (EliXR) was developed that integrates syntactic parsing and tree pattern mining to discover common semantic patterns in 1000 eligibility criteria randomly selected from http://ClinicalTrials.gov. The semantic patterns were aggregated and enriched with unified medical language systems semantic knowledge to form a semantic representation for clinical research eligibility criteria. RESULTS: The authors arrived at 175 semantic patterns, which form 12 semantic role labels connected by their frequent semantic relations in a semantic network. EVALUATION: Three raters independently annotated all the sentence segments (N=396) for 79 test eligibility criteria using the 12 top-level semantic role labels. Eight-six per cent (339) of the sentence segments were unanimously labelled correctly and 13.8% (55) were correctly labelled by two raters. The Fleiss' κ was 0.88, indicating a nearly perfect interrater agreement. CONCLUSION: This study present a semi-automated data-driven approach to developing a semantic network that aligns well with the top-level information structure in clinical research eligibility criteria text and demonstrates the feasibility of using the resulting semantic role labels to generate semistructured eligibility criteria with nearly perfect interrater reliability. Chunhua Weng, Xiaoying Wu 0001, Jake Luo, Mary Regina Boland, Dimitri Theodoratos, Stephen B. Johnson |
J. Am. Medical Informatics Assoc. | 4 |