Karthik Natarajan

dblp:30/6897 · DBLP profile ↗
← Back
41ranked-venue papers
3as first author
19since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 35 · 3 first-author · 16 since 2021Theory of computation · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2024 Using patient portals for large-scale recruitment of individuals underrepresented in biomedical research: an evaluation of engagement patterns throughout the patient portal recruitment process at a single site within the All of Us Research Program
abstract
OBJECTIVE: To evaluate the use of patient portal messaging to recruit individuals historically underrepresented in biomedical research (UBR) to the All of Us Research Program (AoURP) at a single recruitment site. MATERIALS AND METHODS: Patient portal-based recruitment was implemented at Columbia University Irving Medical Center. Patient engagement was assessed using patient's electronic health record (EHR) at four recruitment stages: Consenting to be contacted, opening messages, responding to messages, and showing interest in participating. Demographic and socioeconomic data were also collected from patient's EHR and univariate logistic regression analyses were conducted to assess patient engagement. RESULTS: Between October 2022 and November 2023, a total of 59 592 patients received patient portal messages inviting them to join the AoURP. Among them, 24 445 (41.0%) opened the message, 8983 (15.1%) responded, and 3765 (6.3%) showed interest in joining the program. Though we were unable to link enrollment data with EHR data, we estimate about 2% of patients contacted ultimately enrolled in the AoURP. Patients from underrepresented race and ethnicity communities had lower odds of consenting to be contacted and opening messages, but higher odds of showing interest after responding. DISCUSSION: Patient portal messaging provided both patients and recruitment staff with a more efficient approach to outreach, but patterns of engagement varied across UBR groups. CONCLUSION: Patient portal-based recruitment enables researchers to contact a substantial number of participants from diverse communities. However, more effort is needed to improve engagement from underrepresented racial and ethnic groups at the early stages of the recruitment process.
Maura Beaton, Xinzhuo Jiang, Elise L. Minto, Chun Yee Lau, Lennon Turner, George Hripcsak, Kanchan Chaudhari, Karthik Natarajan
J. Am. Medical Informatics Assoc.8
2024 Identifying erroneous height and weight values from adult electronic health records in the All of Us research program
abstract
INTRODUCTION: Electronic Health Records (EHR) are a useful data source for research, but their usability is hindered by measurement errors. This study investigated an automatic error detection algorithm for adult height and weight measurements in EHR for the All of Us Research Program (All of Us). METHODS: We developed reference charts for adult heights and weights that were stratified on participant sex. Our analysis included 4,076,534 height and 5,207,328 wt measurements from ∼ 150,000 participants. Errors were identified using modified standard deviation scores, differences from their expected values, and significant changes between consecutive measurements. We evaluated our method with chart-reviewed heights (8,092) and weights (9,039) from 250 randomly selected participants and compared it with the current cleaning algorithm in All of Us. RESULTS: The proposed algorithm classified 1.4 % of height and 1.5 % of weight errors in the full cohort. Sensitivity was 90.4 % (95 % CI: 79.0-96.8 %) for heights and 65.9 % (95 % CI: 56.9-74.1 %) for weights. Precision was 73.4 % (95 % CI: 60.9-83.7 %) for heights and 62.9 (95 % CI: 54.0-71.1 %) for weights. In comparison, the current cleaning algorithm has inferior performance in sensitivity (55.8 %) and precision (16.5 %) for height errors while having higher precision (94.0 %) and lower sensitivity (61.9 %) for weight errors. DISCUSSION: Our proposed algorithm outperformed in detecting height errors compared to weights. It can serve as a valuable addition to the current All of Us cleaning algorithm for identifying erroneous height values.
Andrew Guide, Lina M. Sulieman, Shawn Garbett, Robert M. Cronin, Matthew E. Spotnitz, Karthik Natarajan, Robert J. Carroll, Paul A. Harris, Qingxia Chen
J. Biomed. Informatics6
2023 A Nonparametric Approach with Marginals for Modeling Consumer Choice
abstract
Given data on choices made by consumers for different assortments, a key challenge is to develop parsimonious models that describe and predict consumer choice behavior. One such choice model is the marginal distribution model (MDM), which requires only the specification of the marginal distributions of the random utilities of the alternatives to explain choice data.
Yanqiu Ruan, Xiaobo Li 0002, Karthyek Murthy, Karthik Natarajan
EC4
2023 Clinical and temporal characterization of COVID-19 subgroups using patient vector embeddings of electronic health records
abstract
OBJECTIVE: To identify and characterize clinical subgroups of hospitalized Coronavirus Disease 2019 (COVID-19) patients. MATERIALS AND METHODS: Electronic health records of hospitalized COVID-19 patients at NewYork-Presbyterian/Columbia University Irving Medical Center were temporally sequenced and transformed into patient vector representations using Paragraph Vector models. K-means clustering was performed to identify subgroups. RESULTS: A diverse cohort of 11 313 patients with COVID-19 and hospitalizations between March 2, 2020 and December 1, 2021 were identified; median [IQR] age: 61.2 [40.3-74.3]; 51.5% female. Twenty subgroups of hospitalized COVID-19 patients, labeled by increasing severity, were characterized by their demographics, conditions, outcomes, and severity (mild-moderate/severe/critical). Subgroup temporal patterns were characterized by the durations in each subgroup, transitions between subgroups, and the complete paths throughout the course of hospitalization. DISCUSSION: Several subgroups had mild-moderate severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infections but were hospitalized for underlying conditions (pregnancy, cardiovascular disease [CVD], etc.). Subgroup 7 included solid organ transplant recipients who mostly developed mild-moderate or severe disease. Subgroup 9 had a history of type-2 diabetes, kidney and CVD, and suffered the highest rates of heart failure (45.2%) and end-stage renal disease (80.6%). Subgroup 13 was the oldest (median: 82.7 years) and had mixed severity but high mortality (33.3%). Subgroup 17 had critical disease and the highest mortality (64.6%), with age (median: 68.1 years) being the only notable risk factor. Subgroups 18-20 had critical disease with high complication rates and long hospitalizations (median: 40+ days). All subgroups are detailed in the full text. A chord diagram depicts the most common transitions, and paths with the highest prevalence, longest hospitalizations, lowest and highest mortalities are presented. Understanding these subgroups and their pathways may aid clinicians in their decisions for better management and earlier intervention for patients.
Casey N. Ta, Jason Zucker 0001, Po-Hsiang Chiu, Yilu Fang, Karthik Natarajan, Chunhua Weng
J. Am. Medical Informatics Assoc.5
2023 Representing and utilizing clinical textual data for real world studies: An OHDSI approach
Vipina Kuttichi Keloth, Juan M. Banda, Michael J. Gurley, Paul M. Heider, Georgina Kennedy, Timothy A. Miller, Karthik Natarajan, Olga V. Patterson, Yifan Peng 0002, Kalpana Raja, Ruth M. Reeves, Masoud Rouhizadeh, Jianlin Shi, Yanshan Wang, Wei-Qi Wei, Andrew E. Williams, Rui Zhang 0028, Rimma Belenkaya, Christian G. Reich, Clair Blacketer, Patrick B. Ryan, George Hripcsak, Noémie Elhadad, Hua Xu 0001
J. Biomed. Informatics9
2023 Tight Probability Bounds with Pairwise Independence
abstract
Abstract. While useful probability bounds for [Formula: see text] pairwise independent Bernoulli random variables adding up to at least an integer [Formula: see text] have been proposed in the literature, none of these bounds are tight in general. In this paper, we provide several results in this direction. First, when [Formula: see text], the tightest upper bound on the probability of the union of [Formula: see text] pairwise independent events is provided in closed-form for any input marginal probability vector [Formula: see text]. To prove the result, we show the existence of a positively correlated Bernoulli random vector with transformed bivariate probabilities, which is of independent interest. Building on this, we show that the ratio of the Boole union bound to the tight pairwise independent bound is upper bounded by [Formula: see text] and that the ratio is attained. Applications of the result in correlation gap analysis and distributionally robust bottleneck optimization are discussed. The result is extended to find the tightest lower bound on the probability of the intersection of [Formula: see text] pairwise independent events. Second, for any [Formula: see text] and input marginal probability vector [Formula: see text], new upper bounds are derived by exploiting ordering of probabilities. Numerical examples are provided to illustrate when the bounds provide improvement over existing bounds. Lastly, we identify specific instances when the existing and the new bounds are tight, for example, with identical marginal probabilities.
Arjun Kodagehalli Ramachandra, Karthik Natarajan
SIAM J. Discret. Math.2
2022 Analyzing healthcare-seeking behavior among All of Us enrollees in the era of COVID-19
Nripendra D. Acharya, Harry Reyes Nieva, Karthik Natarajan
AMIA3
2022 Feasibility of Linking Area Deprivation Index Data to the OMOP Common Data Model
Xinzhuo Jiang, Maura Beaton, Jake Gillberg, Andrew E. Williams, Karthik Natarajan
AMIA5
2022 Self-paced Training Modality to Promote the Use of All of Us Researcher Workbench in Educational and Research Settings
Hiral Master, Lina M. Sulieman, Paul A. Harris, Karthik Natarajan, Robert J. Carroll, Kayla Marginean, Kelsey R. Mayo, Aymone Kouame
AMIA4
2022 An interactive fitness-for-use data completeness tool to assess activity tracker data
abstract
OBJECTIVE: To design and evaluate an interactive data quality (DQ) characterization tool focused on fitness-for-use completeness measures to support researchers' assessment of a dataset. MATERIALS AND METHODS: Design requirements were identified through a conceptual framework on DQ, literature review, and interviews. The prototype of the tool was developed based on the requirements gathered and was further refined by domain experts. The Fitness-for-Use Tool was evaluated through a within-subjects controlled experiment comparing it with a baseline tool that provides information on missing data based on intrinsic DQ measures. The tools were evaluated on task performance and perceived usability. RESULTS: The Fitness-for-Use Tool allows users to define data completeness by customizing the measures and its thresholds to fit their research task and provides a data summary based on the customized definition. Using the Fitness-for-Use Tool, study participants were able to accurately complete fitness-for-use assessment in less time than when using the Intrinsic DQ Tool. The study participants perceived that the Fitness-for-Use Tool was more useful in determining the fitness-for-use of a dataset than the Intrinsic DQ Tool. DISCUSSION: Incorporating fitness-for-use measures in a DQ characterization tool could provide data summary that meets researchers needs. The design features identified in this study has potential to be applied to other biomedical data types. CONCLUSION: A tool that summarizes a dataset in terms of fitness-for-use dimensions and measures specific to a research question supports dataset assessment better than a tool that only presents information on intrinsic DQ measures.
Sylvia Cho, Ipek Ensari, Noémie Elhadad, Chunhua Weng, Jennifer M. Radin, Brinnae Bent, Pooja M. Desai, Karthik Natarajan
J. Am. Medical Informatics Assoc.8
2022 Comparing medical history data derived from electronic health records and survey answers in the All of Us Research Program
abstract
OBJECTIVE: A participant's medical history is important in clinical research and can be captured from electronic health records (EHRs) and self-reported surveys. Both can be incomplete, EHR due to documentation gaps or lack of interoperability and surveys due to recall bias or limited health literacy. This analysis compares medical history collected in the All of Us Research Program through both surveys and EHRs. MATERIALS AND METHODS: The All of Us medical history survey includes self-report questionnaire that asks about diagnoses to over 150 medical conditions organized into 12 disease categories. In each category, we identified the 3 most and least frequent self-reported diagnoses and retrieved their analogues from EHRs. We calculated agreement scores and extracted participant demographic characteristics for each comparison set. RESULTS: The 4th All of Us dataset release includes data from 314 994 participants; 28.3% of whom completed medical history surveys, and 65.5% of whom had EHR data. Hearing and vision category within the survey had the highest number of responses, but the second lowest positive agreement with the EHR (0.21). The Infectious disease category had the lowest positive agreement (0.12). Cancer conditions had the highest positive agreement (0.45) between the 2 data sources. DISCUSSION AND CONCLUSION: Our study quantified the agreement of medical history between 2 sources-EHRs and self-reported surveys. Conditions that are usually undocumented in EHRs had low agreement scores, demonstrating that survey data can supplement EHR data. Disagreement between EHR and survey can help identify possible missing records and guide researchers to adjust for biases.
Lina M. Sulieman, Robert M. Cronin, Robert J. Carroll, Karthik Natarajan, Kayla Marginean, Brandy Mapes, Dan M. Roden, Paul A. Harris, Andrea H. Ramirez
J. Am. Medical Informatics Assoc.4
2021 Designing a Data Quality Characterization Tool for Fitness Tracker Data
Sylvia Cho, Karthik Natarajan
AMIA2
2021 Harmonization of Measurement Codes for Concept-Oriented Lab Data Retrieval
Matthew E. Spotnitz, Jason Patterson, Vojtech Huser, Chunhua Weng, Karthik Natarajan
AMIA5
2021 Measuring the correctness of All of Us physical measurement
Lina M. Sulieman, Karthik Natarajan, Qingxia Chen, Robert J. Carroll, Kayla Marginean, Paul A. Harris, Andrea H. Ramirez
AMIA2
2021 REDHot OMOP: Facilitating Semantic Interoperability in REDCap with FHIR and the OMOP CDM
Salvatore G. Volpe, Karthik Natarajan
AMIA2
2021 Worst-Case Expected Shortfall with Univariate and Bivariate Marginals
abstract
Computing and minimizing the worst-case bound on the expected shortfall risk of a portfolio given partial information on the distribution of the asset returns is an important problem in risk management. One such bound that been proposed is for the worst-case distribution that is “close” to a reference distribution where closeness in distance among distributions is measured using [Formula: see text]-divergence. In this paper, we advocate the use of such ambiguity sets with a tree structure on the univariate and bivariate marginal distributions. Such an approach has attractive modeling and computational properties. From a modeling perspective, this provides flexibility for risk management applications where there are many more choices for bivariate copulas in comparison with multivariate copulas. Bivariate copulas form the basis of the nested tree structure that is found in vine copulas. Because estimating a vine copula is fairly challenging, our approach provides robust bounds that are valid for the tree structure that is obtained by truncating the vine copula at the top level. The model also provides flexibility in tackling instances when the lower dimensional marginal information is inconsistent that might arise when multiple experts provide information. From a computational perspective, under the assumption of a tree structure on the bivariate marginals, we show that the worst-case expected shortfall is computable in polynomial time in the input size when the distributions are discrete. The corresponding distributionally robust portfolio optimization problem is also solvable in polynomial time. In contrast, under the assumption of independence, the expected shortfall is shown to be #P-hard to compute for discrete distributions. We provide numerical examples with simulated and real data to illustrate the quality of the worst-case bounds in risk management and portfolio optimization and compare it with alternate probabilistic models such as vine copulas and Markov tree distributions.
Anulekha Dhara, Bikramjit Das, Karthik Natarajan
INFORMS J. Comput.3
2021 The National COVID Cohort Collaborative (N3C): Rationale, design, infrastructure, and deployment
abstract
OBJECTIVE: Coronavirus disease 2019 (COVID-19) poses societal challenges that require expeditious data and knowledge sharing. Though organizational clinical data are abundant, these are largely inaccessible to outside researchers. Statistical, machine learning, and causal analyses are most successful with large-scale data beyond what is available in any given organization. Here, we introduce the National COVID Cohort Collaborative (N3C), an open science community focused on analyzing patient-level data from many centers. MATERIALS AND METHODS: The Clinical and Translational Science Award Program and scientific community created N3C to overcome technical, regulatory, policy, and governance barriers to sharing and harmonizing individual-level clinical data. We developed solutions to extract, aggregate, and harmonize data across organizations and data models, and created a secure data enclave to enable efficient, transparent, and reproducible collaborative analytics. RESULTS: Organized in inclusive workstreams, we created legal agreements and governance for organizations and researchers; data extraction scripts to identify and ingest positive, negative, and possible COVID-19 cases; a data quality assurance and harmonization pipeline to create a single harmonized dataset; population of the secure data enclave with data, machine learning, and statistical analytics tools; dissemination mechanisms; and a synthetic data pilot to democratize data access. CONCLUSIONS: The N3C has demonstrated that a multisite collaborative learning health network can overcome barriers to rapidly build a scalable infrastructure incorporating multiorganizational clinical data for COVID-19 analytics. We expect this effort to save lives by enabling rapid collaboration among clinicians, researchers, and data scientists to identify treatments and specialized care and thereby reduce the immediate and long-term impacts of COVID-19.
Melissa A. Haendel, Christopher G. Chute, Tellen D. Bennett, David Eichmann, Justin Guinney, Warren A. Kibbe, Philip R. O. Payne, Emily R. Pfaff, Peter N. Robinson, Joel H. Saltz, Heidi Spratt, Christine Suver, John Wilbanks, Adam B. Wilcox, Andrew E. Williams, Chunlei Wu, Clair Blacketer, Robert L. Bradford, James J. Cimino, Marshall Clark, Evan W. Colmenares, Patricia A. Francis, Davera Gabriel, Alexis Graves, Raju Hemadri, Stephanie S. Hong, George Hripcsak, Dazhi Jiao, Jeffrey G. Klann, Kristin Kostka, Adam M. Lee, Harold P. Lehmann, Lora Lingrey, Robert T. Miller, Michele Morris, Shawn N. Murphy, Karthik Natarajan, Matvey Palchuk, Usman Sheikh, Harold R. Solbrig, Shyam Visweswaran, Anita Walden, Kellie M. Walters, Griffin M. Weber, Xiaohan Tanner Zhang, Richard L. Zhu, Benjamin R. C. Amor, Andrew T. Girvin, Amin Manna, Nabeel Qureshi, Michael G. Kurilla, Samuel G. Michael, Lili M. Portilla, Joni L. Rutter, Christopher P. Austin, Kenneth R. Gersing
J. Am. Medical Informatics Assoc.37
2021 Are synthetic clinical notes useful for real natural language processing tasks: A case study on clinical entity recognition
abstract
OBJECTIVE: : Developing clinical natural language processing systems often requires access to many clinical documents, which are not widely available to the public due to privacy and security concerns. To address this challenge, we propose to develop methods to generate synthetic clinical notes and evaluate their utility in real clinical natural language processing tasks. MATERIALS AND METHODS: : We implemented 4 state-of-the-art text generation models, namely CharRNN, SegGAN, GPT-2, and CTRL, to generate clinical text for the History and Present Illness section. We then manually annotated clinical entities for randomly selected 500 History and Present Illness notes generated from the best-performing algorithm. To compare the utility of natural and synthetic corpora, we trained named entity recognition (NER) models from all 3 corpora and evaluated their performance on 2 independent natural corpora. RESULTS: : Our evaluation shows GPT-2 achieved the best BLEU (bilingual evaluation understudy) score (with a BLEU-2 of 0.92). NER models trained on synthetic corpus generated by GPT-2 showed slightly better performance on 2 independent corpora: strict F1 scores of 0.709 and 0.748, respectively, when compared with the NER models trained on natural corpus (F1 scores of 0.706 and 0.737, respectively), indicating the good utility of synthetic corpora in clinical NER model development. In addition, we also demonstrated that an augmented method that combines both natural and synthetic corpora achieved better performance than that uses the natural corpus only. CONCLUSIONS: : Recent advances in text generation have made it possible to generate synthetic clinical notes that could be useful for training NER models for information extraction from natural clinical notes, thus lowering the privacy concern and increasing data availability. Further investigation is needed to apply this technology to practice.
Jianfu Li, Yujia Zhou 0003, Xiaoqian Jiang, Karthik Natarajan, Serguei V. S. Pakhomov, Hua Xu 0001
J. Am. Medical Informatics Assoc.4
2021 Development and validation of prediction models for mechanical ventilation, renal replacement therapy, and readmission in COVID-19 patients
abstract
OBJECTIVE: Coronavirus disease 2019 (COVID-19) patients are at risk for resource-intensive outcomes including mechanical ventilation (MV), renal replacement therapy (RRT), and readmission. Accurate outcome prognostication could facilitate hospital resource allocation. We develop and validate predictive models for each outcome using retrospective electronic health record data for COVID-19 patients treated between March 2 and May 6, 2020. MATERIALS AND METHODS: For each outcome, we trained 3 classes of prediction models using clinical data for a cohort of SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2)-positive patients (n = 2256). Cross-validation was used to select the best-performing models per the areas under the receiver-operating characteristic and precision-recall curves. Models were validated using a held-out cohort (n = 855). We measured each model's calibration and evaluated feature importances to interpret model output. RESULTS: The predictive performance for our selected models on the held-out cohort was as follows: area under the receiver-operating characteristic curve-MV 0.743 (95% CI, 0.682-0.812), RRT 0.847 (95% CI, 0.772-0.936), readmission 0.871 (95% CI, 0.830-0.917); area under the precision-recall curve-MV 0.137 (95% CI, 0.047-0.175), RRT 0.325 (95% CI, 0.117-0.497), readmission 0.504 (95% CI, 0.388-0.604). Predictions were well calibrated, and the most important features within each model were consistent with clinical intuition. DISCUSSION: Our models produce performant, well-calibrated, and interpretable predictions for COVID-19 patients at risk for the target outcomes. They demonstrate the potential to accurately estimate outcome prognosis in resource-constrained care sites managing COVID-19 patients. CONCLUSIONS: We develop and validate prognostic models targeting MV, RRT, and readmission for hospitalized COVID-19 patients which produce accurate, interpretable predictions. Additional external validation studies are needed to further verify the generalizability of our results.
Victor Alfonso Rodriguez, Shreyas Bhave, George Hripcsak, Soumitra Sengupta, Noémie Elhadad, Robert A. Green, Jason S. Adelman, Katherine Schlosser Metitiri, Pierre A. Elias, Holden Groves, Sumit Mohan, Karthik Natarajan, Adler J. Perotte
J. Am. Medical Informatics Assoc.14
2020 Identifying Use Cases of Consumer-Grade Fitness Trackers in Research Studies
Sylvia Cho, Karthik Natarajan
AMIA2
2020 Fitbit "Bring Your Own Device" data in the All of Us Research Program
Michelle Holko, Francis Ratsimbazafy, Kayla Marginean, Karthik Natarajan, Sylvia Cho, Josh Schilling, Aymone Kouame, Dan Webster, Shaquille Peters, Mark Begale, Kelly Gebo, Andrea H. Ramirez, Paul A. Harris
AMIA4
2020 Data Quality Assessment of Laboratory Data
Vojtech Huser, Clair Blacketer, Karthik Natarajan, Robert T. Miller, Andrew E. Williams, Selva Muthu Kumaran Sathappan, José D. Posada, Nigam H. Shah
AMIA3
2020 Characterizing database granularity using SNOMED-CT hierarchy
Anna Ostropolets, Christian G. Reich, Patrick B. Ryan, Chunhua Weng, Anthony Molinaro, Frank J. DeFalco, Jitendra Jonnagaddala, Siaw-Teng Liaw, Hokyun Jeon, Rae Woong Park, Matthew E. Spotnitz, Karthik Natarajan, Kristin Kostka, George Argyriou, Robert T. Miller, Andrew E. Williams, Evan P. Minty, José D. Posada, George Hripcsak
AMIA12
2020 Characterization and Comparison of Embedding Algorithms for Phenotyping across a Network of Observational Databases
Harry Reyes Nieva, Krishna Kalluri, Tony Sun, Xinzhuo Jiang, Victor Alfonso Rodriguez, Patrick B. Ryan, Karthik Natarajan
AMIA9
2020 Phenotype Concept Set Construction from Concept Pair Likelihoods
Victor Alfonso Rodriguez, Tony Sun, Phyllis Thangaraj, Krishna Kalluri, Xinzhuo Jiang, Karthik Natarajan, Patrick B. Ryan, Anna Ostropolets
AMIA7
2020 Bias in the Reuse and Analysis of Electronic Health Record Data
Nicole Gray Weiskopf, Melody L. Greer, Karthik Natarajan, Caroline A. Thompson, Harold P. Lehmann
AMIA3
2020 Normalizing Clinical Document Titles to LOINC Document Ontology: an Initial Study
Xu Zuo, Jianfu Li, Bo Zhao 0001, Yujia Zhou 0003, Jon D. Duke, Karthik Natarajan, George Hripcsak, Nigam H. Shah, Juan M. Banda, Ruth M. Reeves, Hua Xu 0001
AMIA7
2020 Correlation Robust Influence Maximization
abstract
We propose a distributionally robust model for the influence maximization problem. Unlike the classical independent cascade model of Kempe et al (2003), this model's diffusion process is adversarially adapted to the choice of seed set. So instead of optimizing under the assumption that all influence relationships in the network are independent, we seek a seed set whose expected influence under the worst correlation, i.e., the ``worst-case, expected influence", is maximized. We show that this worst-case influence can be efficiently computed, and though the optimization is NP-hard, a (1 - 1/e) approximation guarantee holds. We also analyze the structure to the adversary's choice of diffusion process, and contrast with established models. Beyond the key computational advantages, we also study the degree to which the independence assumption may be considered costly, and provide insights from numerical experiments comparing the adversarial and independent cascade model.
Louis Chen, Divya Padmanabhan, Karthik Natarajan
NeurIPS4
2020 Composite pattern to handle variation points in software architectural design of evolving application systems
abstract
The variation points in software architecture arise as a result of the availability of large number of filters and component libraries. An integration of different architectural styles is crucial and necessary in the development of large‐scale software application systems to handle the variation points. This article proposes a composite software architectural style for building application systems involving data streams, user interactivity, and dynamic mode. It uses a pattern within a pattern approach for combining the architectural styles. This approach provides flexibility to add or delete any filter or component at run time. In addition, the changes in the order of processing of the different filters or components can also be incorporated. The software architectural specification for any combination of input components and their order of processing is generated automatically. This specification acts as a baseline for the subsequent design and implementation phases of the application system. This model is generic and has been successfully validated for a prototype application system involving all the three modes of operation.
Milu Mary Philip, Karthik Natarajan, Anithkumar Ramanathan, Vijayakumar Balakrishnan
IET Softw.2
2020 COVID-19 TestNorm: A tool to normalize COVID-19 testing names to LOINC codes
abstract
Large observational data networks that leverage routine clinical practice data in electronic health records (EHRs) are critical resources for research on coronavirus disease 2019 (COVID-19). Data normalization is a key challenge for the secondary use of EHRs for COVID-19 research across institutions. In this study, we addressed the challenge of automating the normalization of COVID-19 diagnostic tests, which are critical data elements, but for which controlled terminology terms were published after clinical implementation. We developed a simple but effective rule-based tool called COVID-19 TestNorm to automatically normalize local COVID-19 testing names to standard LOINC (Logical Observation Identifiers Names and Codes) codes. COVID-19 TestNorm was developed and evaluated using 568 test names collected from 8 healthcare systems. Our results show that it could achieve an accuracy of 97.4% on an independent test set. COVID-19 TestNorm is available as an open-source package for developers and as an online Web application for end users (https://clamp.uth.edu/covid/loinc.php). We believe that it will be a useful tool to support secondary use of EHRs for research on COVID-19.
Jianfu Li, Ekin Soysal, Jiang Bian 0001, Scott L. DuVall, Elizabeth Hanchrow, Kristine E. Lynch, Michael E. Matheny, Karthik Natarajan, Lucila Ohno-Machado, Serguei V. S. Pakhomov, Ruth M. Reeves, Amy M. Sitapati, Swapna Abhyankar, Theresa A. Cullen, Jami Deckard, Xiaoqian Jiang, Robert Murphy, Hua Xu 0001
J. Am. Medical Informatics Assoc.10
2019 Curating EHR data in the All of Us Research Program
Karthik Natarajan, Robert J. Carroll, Thomas R. Campion Jr., Joan Grand, Shyam Visweswaran
AMIA1
2019 Facilitating phenotype transfer using a common data model
George Hripcsak, Ning Shang 0004, Peggy L. Peissig, Luke V. Rasmussen, Cong Liu 0020, Barbara Benoit, Robert J. Carroll, David Carrell, Joshua C. Denny, Ozan Dikilitas, Vivian S. Gainer, Kayla Marie Howell, Jeffrey G. Klann, Iftikhar J. Kullo, Todd Lingren, Frank D. Mentch, Shawn N. Murphy, Karthik Natarajan, Chunhua Weng
J. Biomed. Informatics18
2019 Distributionally robust project crashing with partial or no correlation information
abstract
Abstract Crashing is shortening the project makespan by reducing activity times in a project network by allocating resources to them. Activity durations are often uncertain and an exact probability distribution itself might be ambiguous. We study a class of distributionally robust project crashing problems where the objective is to optimize the first two marginal moments (means and SDs) of the activity durations to minimize the worst‐case expected makespan. Under partial correlation information and no correlation information, the problem is solvable in polynomial time as a semidefinite program and a second‐order cone program, respectively. However, solving semidefinite programs is challenging for large project networks. We exploit the structure of the distributionally robust formulation to reformulate a convex‐concave saddle point problem over the first two marginal moment variables and the arc criticality index variables. We then use a projection and contraction algorithm for monotone variational inequalities in conjunction with a gradient method to solve the saddle point problem enabling us to tackle large instances. Numerical results indicate that a manager who is faced with ambiguity in the distribution of activity durations has a greater incentive to invest resources in decreasing the variations rather than the means of the activity durations.
Selin Damla Ahipasaoglu, Karthik Natarajan, Dongjian Shi
Networks2
2018 Treatment Pathways in Patients with Cancer Using a Large-scale Observational Data Network
Patrick B. Ryan, Karthik Natarajan, Thomas Falconer, Christian G. Reich, Rohit Vashisht, Nigam H. Shah, George Hripcsak
AMIA3
2017 The Data and Research Center of the All of Us Research Program: Framework for a National Cohort Program and Research Opportunities
Robert J. Carroll, Joshua C. Mandel, Karthik Natarajan, Scott Sutherland, Joshua C. Denny
AMIA3
2017 Evaluation of OMOP Vocabulary on Transplant Registry Data
Sylvia Cho, Margaret Sin, Karthik Natarajan
AMIA3
2014 Diagnosis code assignment: models and evaluation metrics
abstract
BACKGROUND AND OBJECTIVE: The volume of healthcare data is growing rapidly with the adoption of health information technology. We focus on automated ICD9 code assignment from discharge summary content and methods for evaluating such assignments. METHODS: We study ICD9 diagnosis codes and discharge summaries from the publicly available Multiparameter Intelligent Monitoring in Intensive Care II (MIMIC II) repository. We experiment with two coding approaches: one that treats each ICD9 code independently of each other (flat classifier), and one that leverages the hierarchical nature of ICD9 codes into its modeling (hierarchy-based classifier). We propose novel evaluation metrics, which reflect the distances among gold-standard and predicted codes and their locations in the ICD9 tree. Experimental setup, code for modeling, and evaluation scripts are made available to the research community. RESULTS: The hierarchy-based classifier outperforms the flat classifier with F-measures of 39.5% and 27.6%, respectively, when trained on 20,533 documents and tested on 2282 documents. While recall is improved at the expense of precision, our novel evaluation metrics show a more refined assessment: for instance, the hierarchy-based classifier identifies the correct sub-tree of gold-standard codes more often than the flat classifier. Error analysis reveals that gold-standard codes are not perfect, and as such the recall and precision are likely underestimated. CONCLUSIONS: Hierarchy-based classification yields better ICD9 coding than flat classification for MIMIC patients. Automated ICD9 coding is an example of a task for which data and tools can be shared and for which the research community can work together to build on shared models and advance the state of the art.
Adler J. Perotte, Rimma Perotte, Karthik Natarajan, Nicole Gray Weiskopf, Frank D. Wood, Noémie Elhadad
J. Am. Medical Informatics Assoc.3
2013 Analyzing Requests for Clinical Data for Self-Service Penetration
Karthik Natarajan, Adam B. Wilcox, Niloo Sobhani, Aurelia Boyer
AMIA1
2013 From Boxes to bees: Active learning in freshmen calculus
abstract
The vehicles of education have seen significant broadening with the proliferation of new technologies such as social media, microblogs, online references, multimedia, and interactive teaching tools. This paper summarizes research on the effect of using active learning methods to facilitate student learning and describes our experiences implementing group activities for a calculus course for first year university students at Singapore University of Technology and Design, a new design-centric university established in collaboration with Massachusetts Institute of Technology (MIT). We describe the educational impact of different pedagogical techniques, such as real-time response tools, hands on activities, mathematical modeling, visualization activities and motivational competitions on students with differing learning preferences in a unique cohort classroom setting. Based on faculty reflection and survey data, we provide guidelines on how to adopt the right set of active and group learning techniques to handle the changing learning preferences in the current and future generation of students.
Flora S. Tsai, Karthik Natarajan, Selin Damla Ahipasaoglu, Chau Yuen, Hyowon Lee 0001, Ngai-Man Cheung, Justin Ruths, Shisheng Huang, Thomas L. Magnanti
EDUCON2
2007 Redesigning electronic health record systems to support public health
Rita Kukafka, Jessica S. Ancker, Connie V. Chan, John Chelico, Sharib A. Khan, Selasie Mortoti, Karthik Natarajan, Kempton Presley, Kayann Stephens
J. Biomed. Informatics7
2003 Microprocessor pipeline energy analysis
abstract
The increase in high-performance microprocessor power consumption is due in part to the large power overhead of wide-issue, highly speculative cores. Microarchitectural speculation, such as branch prediction, increases instruction throughput but carries a power burden due to wasted power for mis-speculated instructions. Pipeline over-provisioning supplies excess resources which often go unused. In this paper, we use our detailed performance and power model for an Alpha 21264 to measure both the useful energy and the wasted effort due to mis-speculation and over-provisioning. Our experiments show that flushed instructions account for approximately 6% of total energy, while over-provisioning imposes a tax of 17% on average. These results suggest opportunities for power savings and energy efficiency throughout microprocessor pipelines.
Karthik Natarajan, Heather Hanson, Stephen W. Keckler, Charles R. Moore, Doug Burger
ISLPED1