EDBT 2026 Demo / reviewers in the wild / expert
Prithwish Chakraborty
dblp:25/8076
· DBLP profile ↗
12ranked-venue papers in the field
5as first author
3since 2021 · last 2023
0000-0003-1407-7677ORCID · corroborated
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 9 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (2 first)Information Retrieval & Web Search · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Workshop on Applied Data Science for Healthcare: Applications and New Frontiers of Generative Models for HealthcareabstractBuilt on the success of the past five years, KDD DSHealth 2023 will further catalyze the development of links between academic and industrial data science groups. The workshop aims to stimulate discussion on strategic areas for development and to facilitate future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community via timely topics, this year the workshop will focus on the applications and new development of generative models in healthcare, including the new development and application of LLMs. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature two invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, we will invite community members to submit their research works and bring them for discussion. The summary gives a brief description of the half-day workshop to be held on August 7th, 2023. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 3 |
| 2022 | Workshop on Applied Data Science for Healthcare (DSHealth): Transparent and Human-centered AIabstractKDD DSHealth 2022, aims to build on the success of the past four years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the transparency and human-centered AI in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature four invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Journal of Healthcare Informatics Research. The summary gives a brief description of the full-day workshop to be held on August 14th, 2022. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 3 |
| 2021 | KDD Health Day/DSHealth 2021: Joint KDD 2021 Health Day and 2021 KDD Workshop on Applied Data Science for Healthcare: State of XAI and Trustworthiness in HealthabstractKDD Health Day/DSHealth 2021, aims to build on the success of the past 3 years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the state of explainability and trustworthiness in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature 8 invited talks from eminent speakers across academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Artificial Intelligence in Medicine journal. The summary gives a brief description of the full-day workshop to be held on August, 2021 virtually. Fei Wang 0001, Prithwish Chakraborty, Tao Xu 0020, Pei-Yun Sabrina Hsueh, Xudong Sun 0014, Gregor Stiglic, Gracy Crane, Jiang Bian 0001, Laleh Haghverdi, Lixia Yao, Florian Buettner 0001 |
KDD | 2 |
| 2020 | Tutorial on Human-Centered Explainability for HealthcareabstractIn recent years, the rapid advances in Artificial Intelligence (AI) techniques along with an ever-increasing availability of healthcare data have made many novel analyses possible. Significant successes have been observed in a wide range of tasks such as next diagnosis prediction, AKI prediction, adverse event predictions including mortality and unexpected hospital re-admissions. However, there has been limited adoption and use in the clinical practice of these methods due to their black-box nature. A significant amount of research is currently focused on making such methods more interpretable or to make post-hoc explanations more accessible. However, most of this work is done at a very low level and as a result, may not have a direct impact at the point-of-care. This tutorial will provide an overview of the landscape of different approaches that have been developed for explainability in healthcare. Specifically, we will present the problem of explainability as it pertains to various personas involved in healthcare viz. data scientists, clinical researchers, and clinicians. We will chart out the requirements for such personas and present an overview of the different approaches that can address such needs. We will also walk-through several use-cases for such approaches. In this process, we will provide a brief introduction to explainability, charting its different dimensions as well as covering some relevant interpretability methods spanning such dimensions. We will touch upon some practical guides for explainability and provide a brief survey of open source tools such as the IBM AI Explainability 360 Open Source Toolkit. Prithwish Chakraborty, Bum Chul Kwon, Sanjoy Dey, Amit Dhurandhar, Dan Gruen, Kenney Ng, Daby M. Sow, Kush R. Varshney |
KDD | 1 |
| 2019 | A Robust Framework for Accelerated Outcome-driven Risk Factor Identification from EHRabstractElectronic Health Records (EHR) containing longitudinal information about millions of patient lives are increasingly being utilized by organizations across the healthcare spectrum. Studies on EHR data have enabled real world applications like understanding of disease progression, outcomes analysis, and comparative effectiveness research. However, often every study is independently commissioned, data is gathered by surveys or specifically purchased per study by a long and often painful process. This is followed by an arduous repetitive cycle of analysis, model building, and generation of insights. This process can take anywhere between 1 - 3 years. In this paper, we present a robust end-to-end machine learning based SaaS system to perform analysis on a very large EHR dataset. The framework consists of a proprietary EHR datamart spanning ~55 million patient lives in USA and over ~20 billion data points. To the best of our knowledge, this framework is the largest in the industry to analyze medical records at this scale, with such efficacy and ease. We developed an end-to-end ML framework with carefully chosen components to support EHR analysis at scale and suitable for further downstream clinical analysis. Specifically, it consists of a ridge regularized Survival Support Vector Machine (SSVM) with a clinical kernel, coupled with Chi-square distance-based feature selection, to uncover relevant risk factors by exploiting the weak correlations in EHR. Our results on multiple real use cases indicate that the framework identifies relevant factors effectively without expert supervision. The framework is stable, generalizable over outcomes, and also found to contribute to better out-of-bound prediction over known expert features. Importantly, the ML methodologies used are interpretable which is critical for acceptance of our system in the targeted user base. With the system being operational, all of these studies were completed within a time frame of 3-4 weeks compared to the industry standard 12-36 months. As such our system can accelerate analysis and discovery, result in better ROI due to reduced investments as well as quicker turn around of studies. Prithwish Chakraborty, Faisal Farooq |
KDD | 1 |
| 2018 | Generating Realistic Synthetic Population DatasetsabstractModern studies of societal phenomena rely on the availability of large datasets capturing attributes and activities of synthetic, city-level, populations. For instance, in epidemiology, synthetic population datasets are necessary to study disease propagation and intervention measures before implementation. In social science, synthetic population datasets are needed to understand how policy decisions might affect preferences and behaviors of individuals. In public health, synthetic population datasets are necessary to capture diagnostic and procedural characteristics of patient records without violating confidentialities of individuals. To generate such datasets over a large set of categorical variables, we propose the use of the maximum entropy principle to formalize a generative model such that in a statistically well-founded way we can optimally utilize given prior information about the data, and are unbiased otherwise. An efficient inference algorithm is designed to estimate the maximum entropy model, and we demonstrate how our approach is adept at estimating underlying data distributions. We evaluate this approach against both simulated data and US census datasets, and demonstrate its feasibility using an epidemic simulation application. Hao Wu 0041, Yue Ning 0001, Prithwish Chakraborty, Jilles Vreeken, Nikolaj Tatti, Naren Ramakrishnan |
ACM Trans. Knowl. Discov. Data | 3 |
| 2017 | GELL: Automatic Extraction of Epidemiological Line Lists from Open SourcesabstractReal-time monitoring and responses to emerging public health threats rely on the availability of timely surveillance data. During the early stages of an epidemic, the ready availability of line lists with detailed tabular information about laboratory-confirmed cases can assist epidemiologists in making reliable inferences and forecasts. Such inferences are crucial to understand the epidemiology of a specific disease early enough to stop or control the outbreak. However, construction of such line lists requires considerable human supervision and therefore, difficult to generate in real-time. In this paper, we motivate Guided Epidemiological Line List (GELL), the first tool for building automated line lists (in near real-time) from open source reports of emerging disease outbreaks. Specifically, we focus on deriving epidemiological characteristics of an emerging disease and the affected population from reports of illness. GELL uses distributed vector representations (ala word2vec) to discover a set of indicators for each line list feature. This discovery of indicators is followed by the use of dependency parsing based techniques for final extraction in tabular form. We evaluate the performance of GELL against a human annotated line list provided by HealthMap corresponding to MERS outbreaks in Saudi Arabia. We demonstrate that GELL extracts line list features with increased accuracy compared to a baseline method. We further show how these automatically extracted line list features can be used for making epidemiological inferences, such as inferring demographics and symptoms-to-hospitalization period of affected individuals. Saurav Ghosh, Prithwish Chakraborty, Bryan L. Lewis, Maimuna S. Majumder, Emily Cohn, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
KDD | 2 |
| 2016 | Characterizing Diseases from Unstructured Text: A Vocabulary Driven Word2vec ApproachabstractTraditional disease surveillance can be augmented with a wide variety of real-time sources such as, news and social media. However, these sources are in general unstructured and, construction of surveillance tools such as taxonomical correlations and trace mapping involves considerable human supervision. In this paper, we motivate a disease vocabulary driven word2vec model (Dis2Vec) to model diseases and constituent attributes as word embeddings from the HealthMap news corpus. We use these word embeddings to automatically create disease taxonomies and evaluate our model against corresponding human annotated taxonomies. We compare our model accuracies against several state-of-the art word2vec methods. Our results demonstrate that Dis2Vec outperforms traditional distributed vector representations in its ability to faithfully capture taxonomical attributes across different class of diseases such as endemic, emerging and rare. Saurav Ghosh, Prithwish Chakraborty, Emily Cohn, John S. Brownstein, Naren Ramakrishnan |
CIKM | 2 |
| 2015 | Dynamic Poisson Autoregression for Influenza-Like-Illness Case Count PredictionabstractInfluenza-like-illness (ILI) is among of the most common diseases worldwide, and reliable forecasting of the same can have significant public health benefits. Recently, new forms of disease surveillance based upon digital data sources have been proposed and are continuing to attract attention over traditional surveillance methods. In this paper, we focus on short-term ILI case count prediction and develop a dynamic Poisson autoregressive model with exogenous inputs variables (DPARX) for flu forecasting. In this model, we allow the autoregressive model to change over time. In order to control the variation in the model, we construct a model similarity graph to specify the relationship between pairs of models at two time points and embed prior knowledge in terms of the structure of the graph. We formulate ILI case count forecasting as a convex optimization problem, whose objective balances the autoregressive loss and the model similarity regularization induced by the structure of the similarity graph. We then propose an efficient algorithm to solve this problem by block coordinate descent. We apply our model and the corresponding learning method on historical ILI records for 15 countries around the world using a variety of syndromic surveillance data sources. Our approach provides consistently better forecasting results than state-of-the-art models available for short-term ILI case count forecasting. Zheng Wang 0011, Prithwish Chakraborty, Sumiko R. Mekaru, John S. Brownstein, Jieping Ye, Naren Ramakrishnan |
KDD | 2 |
| 2014 | Forecasting a Moving Target: Ensemble Models for ILI Case Count PredictionsabstractModern epidemiological forecasts of common illnesses, such as the flu, rely on both traditional surveillance sources as well as digital surveillance data. However, most published studies have been retrospective. Concurrently, the reports about flu activity generally lags by several weeks and even when published are revised for several weeks more. We posit that effectively handling this uncertainty is one of the key challenges for a real-time prediction system in this sphere. In this paper, we present a detailed prospective analysis on the generation of robust quantitative predictions about temporal trends of flu activity, using several surrogate data sources for 15 Latin American countries. We present our findings about the limitations and possible advantages of correcting the uncertainty associated with official flu estimates. We also compare the prediction accuracy between model-level fusion of different surrogate data sources against data-level fusion. Finally, we present a novel matrix factorization approach using neighborhood embedding to predict flu case counts. Comparing our proposed ensemble method against several baseline methods helps us demarcate the importance of different data sources for the countries under consideration. Prithwish Chakraborty, Pejman Khadivi, Bryan L. Lewis, Aravindan Mahendiran, Jiangzhuo Chen, Patrick Butler, Elaine O. Nsoesie, Sumiko R. Mekaru, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
SDM | 1 |
| 2011 | On convergence of the multi-objective particle swarm optimizers
Prithwish Chakraborty, Swagatam Das, Gourab Ghosh Roy, Ajith Abraham |
Inf. Sci. | 1 |
| 2011 | Erratum to "On convergence of the multi-objective particle swarm optimizers" [Inform. Sci 181 (2011) 1411-1425]
Prithwish Chakraborty, Swagatam Das, Gourab Ghosh Roy, Ajith Abraham |
Inf. Sci. | 1 |