VLDB 2026 Research / reviewers in the wild / expert
Prithwish Chakraborty
dblp:25/8076
· DBLP profile ↗
24ranked-venue papers
10as first author
8since 2021 · last 2023
0000-0003-1407-7677ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Workshop on Applied Data Science for Healthcare: Applications and New Frontiers of Generative Models for HealthcareabstractBuilt on the success of the past five years, KDD DSHealth 2023 will further catalyze the development of links between academic and industrial data science groups. The workshop aims to stimulate discussion on strategic areas for development and to facilitate future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community via timely topics, this year the workshop will focus on the applications and new development of generative models in healthcare, including the new development and application of LLMs. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature two invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, we will invite community members to submit their research works and bring them for discussion. The summary gives a brief description of the half-day workshop to be held on August 7th, 2023. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 3 |
| 2023 | Informing clinical assessment by contextualizing post-hoc explanations of risk prediction models in type-2 diabetes
Shruthi Chari, Prasant Acharya, Dan Gruen, Olivia Zhang, Elif Eyigöz, Mohamed F. Ghalwash, Oshani Seneviratne, Fernando J. Suarez Saiz, Pablo Meyer 0001, Prithwish Chakraborty, Deborah L. McGuinness |
Artif. Intell. Medicine | 10 |
| 2022 | Workshop on Applied Data Science for Healthcare (DSHealth): Transparent and Human-centered AIabstractKDD DSHealth 2022, aims to build on the success of the past four years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the transparency and human-centered AI in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature four invited talks from eminent speakers, spanning academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Journal of Healthcare Informatics Research. The summary gives a brief description of the full-day workshop to be held on August 14th, 2022. Tao Xu 0020, Fei Wang 0001, Prithwish Chakraborty, Pei-Yun Sabrina Hsueh, Gregor Stiglic, Jiang Bian 0001, Lixia Yao, Alexej Gossmann, Florian Buettner 0001 |
KDD | 3 |
| 2021 | Towards Clinically Relevant Explanations for Type-2 Diabetes Risk Prediction with the Explanation Ontology
Shruthi Chari, Prithwish Chakraborty, Oshani Seneviratne, Mohamed F. Ghalwash, Dan Gruen, Daby M. Sow, Deborah L. McGuinness |
AMIA | 2 |
| 2021 | Impact of Clinical and Genomic Factors on COVID-19 Disease Severity
Sanjoy Dey, Aritra Bose, Subrata Saha, Prithwish Chakraborty, Mohamed F. Ghalwash, Filippo Utro, Aldo Guzmán-Sáenz, Kenney Ng, Jianying Hu, Laxmi Parida, Daby M. Sow |
AMIA | 4 |
| 2021 | A Comparative Time-to-Event Analysis Across Health Systems
Mohamed F. Ghalwash, Prithwish Chakraborty, Akira Koseki, Hiroki Yanagisawa, Toshiya Iwamori, Ray Tokumasu, Masaki Makino, Ryosuke Yanagiya, Michiharu Kudo, Daby M. Sow |
AMIA | 2 |
| 2021 | Collaborative Graph Learning with Auxiliary Text for Temporal Event Prediction in HealthcareabstractAccurate and explainable health event predictions are becoming crucial for healthcare providers to develop care plans for patients. The availability of electronic health records (EHR) has enabled machine learning advances in providing these predictions. However, many deep-learning-based methods are not satisfactory in solving several key challenges: 1) effectively utilizing disease domain knowledge; 2) collaboratively learning representations of patients and diseases; and 3) incorporating unstructured features. To address these issues, we propose a collaborative graph learning model to explore patient-disease interactions and medical domain knowledge. Our solution is able to capture structural features of both patients and diseases. The proposed model also utilizes unstructured text data by employing an attention manipulating strategy and then integrates attentive text features into a sequential learning process. We conduct extensive experiments on two important healthcare problems to show the competitive prediction performance of the proposed method compared with various state-of-the-art models. We also confirm the effectiveness of learned representations and model interpretability by a set of ablation and case studies. Chang Lu 0004, Chandan K. Reddy, Prithwish Chakraborty, Samantha Kleinberg, Yue Ning 0001 |
IJCAI | 3 |
| 2021 | KDD Health Day/DSHealth 2021: Joint KDD 2021 Health Day and 2021 KDD Workshop on Applied Data Science for Healthcare: State of XAI and Trustworthiness in HealthabstractKDD Health Day/DSHealth 2021, aims to build on the success of the past 3 years to further catalyze the development of links between academic and commercial data science groups and the rapidly developing translational medicine informatics community. The workshop will stimulate discussion as to strategic areas for development and will lead to future cross-disciplinary collaborations. In accordance with the multi-year goal to continue fostering this community as a series of KDD workshops via timely topics, this year the workshop will focus on the state of explainability and trustworthiness in healthcare. The workshop invites full papers, as well as work-in-progress on the application of data science in healthcare. The workshop will feature 8 invited talks from eminent speakers across academia, industry, clinical researchers, and governmental regulatory bodies. In addition, selected papers will be invited to publish in a special issue of Artificial Intelligence in Medicine journal. The summary gives a brief description of the full-day workshop to be held on August, 2021 virtually. Fei Wang 0001, Prithwish Chakraborty, Tao Xu 0020, Pei-Yun Sabrina Hsueh, Xudong Sun 0014, Gregor Stiglic, Gracy Crane, Jiang Bian 0001, Laleh Haghverdi, Lixia Yao, Florian Buettner 0001 |
KDD | 2 |
| 2020 | Tutorial on Human-Centered Explainability for HealthcareabstractIn recent years, the rapid advances in Artificial Intelligence (AI) techniques along with an ever-increasing availability of healthcare data have made many novel analyses possible. Significant successes have been observed in a wide range of tasks such as next diagnosis prediction, AKI prediction, adverse event predictions including mortality and unexpected hospital re-admissions. However, there has been limited adoption and use in the clinical practice of these methods due to their black-box nature. A significant amount of research is currently focused on making such methods more interpretable or to make post-hoc explanations more accessible. However, most of this work is done at a very low level and as a result, may not have a direct impact at the point-of-care. This tutorial will provide an overview of the landscape of different approaches that have been developed for explainability in healthcare. Specifically, we will present the problem of explainability as it pertains to various personas involved in healthcare viz. data scientists, clinical researchers, and clinicians. We will chart out the requirements for such personas and present an overview of the different approaches that can address such needs. We will also walk-through several use-cases for such approaches. In this process, we will provide a brief introduction to explainability, charting its different dimensions as well as covering some relevant interpretability methods spanning such dimensions. We will touch upon some practical guides for explainability and provide a brief survey of open source tools such as the IBM AI Explainability 360 Open Source Toolkit. Prithwish Chakraborty, Bum Chul Kwon, Sanjoy Dey, Amit Dhurandhar, Dan Gruen, Kenney Ng, Daby M. Sow, Kush R. Varshney |
KDD | 1 |
| 2019 | A Robust Framework for Accelerated Outcome-driven Risk Factor Identification from EHRabstractElectronic Health Records (EHR) containing longitudinal information about millions of patient lives are increasingly being utilized by organizations across the healthcare spectrum. Studies on EHR data have enabled real world applications like understanding of disease progression, outcomes analysis, and comparative effectiveness research. However, often every study is independently commissioned, data is gathered by surveys or specifically purchased per study by a long and often painful process. This is followed by an arduous repetitive cycle of analysis, model building, and generation of insights. This process can take anywhere between 1 - 3 years. In this paper, we present a robust end-to-end machine learning based SaaS system to perform analysis on a very large EHR dataset. The framework consists of a proprietary EHR datamart spanning ~55 million patient lives in USA and over ~20 billion data points. To the best of our knowledge, this framework is the largest in the industry to analyze medical records at this scale, with such efficacy and ease. We developed an end-to-end ML framework with carefully chosen components to support EHR analysis at scale and suitable for further downstream clinical analysis. Specifically, it consists of a ridge regularized Survival Support Vector Machine (SSVM) with a clinical kernel, coupled with Chi-square distance-based feature selection, to uncover relevant risk factors by exploiting the weak correlations in EHR. Our results on multiple real use cases indicate that the framework identifies relevant factors effectively without expert supervision. The framework is stable, generalizable over outcomes, and also found to contribute to better out-of-bound prediction over known expert features. Importantly, the ML methodologies used are interpretable which is critical for acceptance of our system in the targeted user base. With the system being operational, all of these studies were completed within a time frame of 3-4 weeks compared to the industry standard 12-36 months. As such our system can accelerate analysis and discovery, result in better ROI due to reduced investments as well as quicker turn around of studies. Prithwish Chakraborty, Faisal Farooq |
KDD | 1 |
| 2018 | What to know before forecasting the fluabstractAccurate and timely influenza (flu) forecasting has gained significant traction in recent times. If done well, such forecasting can aid in deploying effective public health measures. Unlike other statistical or machine learning problems, however, flu forecasting brings unique challenges and considerations stemming from the nature of the surveillance apparatus and the end utility of forecasts. This article presents a set of considerations for flu forecasters to take into account prior to applying forecasting algorithms. Prithwish Chakraborty, Bryan L. Lewis, Stephen G. Eubank, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
PLoS Comput. Biol. | 1 |
| 2018 | Generating Realistic Synthetic Population DatasetsabstractModern studies of societal phenomena rely on the availability of large datasets capturing attributes and activities of synthetic, city-level, populations. For instance, in epidemiology, synthetic population datasets are necessary to study disease propagation and intervention measures before implementation. In social science, synthetic population datasets are needed to understand how policy decisions might affect preferences and behaviors of individuals. In public health, synthetic population datasets are necessary to capture diagnostic and procedural characteristics of patient records without violating confidentialities of individuals. To generate such datasets over a large set of categorical variables, we propose the use of the maximum entropy principle to formalize a generative model such that in a statistically well-founded way we can optimally utilize given prior information about the data, and are unbiased otherwise. An efficient inference algorithm is designed to estimate the maximum entropy model, and we demonstrate how our approach is adept at estimating underlying data distributions. We evaluate this approach against both simulated data and US census datasets, and demonstrate its feasibility using an epidemic simulation application. Hao Wu 0041, Yue Ning 0001, Prithwish Chakraborty, Jilles Vreeken, Nikolaj Tatti, Naren Ramakrishnan |
ACM Trans. Knowl. Discov. Data | 3 |
| 2017 | Data-driven Risk Characterization and Prediction of Renal Failure among Diabetic Type 2 Patients using Electronic Medical Records
Prithwish Chakraborty, Vishrawas Gopalakrishnan, Sharon M. H. Alford, Faisal Farooq |
AMIA | 1 |
| 2017 | GELL: Automatic Extraction of Epidemiological Line Lists from Open SourcesabstractReal-time monitoring and responses to emerging public health threats rely on the availability of timely surveillance data. During the early stages of an epidemic, the ready availability of line lists with detailed tabular information about laboratory-confirmed cases can assist epidemiologists in making reliable inferences and forecasts. Such inferences are crucial to understand the epidemiology of a specific disease early enough to stop or control the outbreak. However, construction of such line lists requires considerable human supervision and therefore, difficult to generate in real-time. In this paper, we motivate Guided Epidemiological Line List (GELL), the first tool for building automated line lists (in near real-time) from open source reports of emerging disease outbreaks. Specifically, we focus on deriving epidemiological characteristics of an emerging disease and the affected population from reports of illness. GELL uses distributed vector representations (ala word2vec) to discover a set of indicators for each line list feature. This discovery of indicators is followed by the use of dependency parsing based techniques for final extraction in tabular form. We evaluate the performance of GELL against a human annotated line list provided by HealthMap corresponding to MERS outbreaks in Saudi Arabia. We demonstrate that GELL extracts line list features with increased accuracy compared to a baseline method. We further show how these automatically extracted line list features can be used for making epidemiological inferences, such as inferring demographics and symptoms-to-hospitalization period of affected individuals. Saurav Ghosh, Prithwish Chakraborty, Bryan L. Lewis, Maimuna S. Majumder, Emily Cohn, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
KDD | 2 |
| 2016 | Characterizing Diseases from Unstructured Text: A Vocabulary Driven Word2vec ApproachabstractTraditional disease surveillance can be augmented with a wide variety of real-time sources such as, news and social media. However, these sources are in general unstructured and, construction of surveillance tools such as taxonomical correlations and trace mapping involves considerable human supervision. In this paper, we motivate a disease vocabulary driven word2vec model (Dis2Vec) to model diseases and constituent attributes as word embeddings from the HealthMap news corpus. We use these word embeddings to automatically create disease taxonomies and evaluate our model against corresponding human annotated taxonomies. We compare our model accuracies against several state-of-the art word2vec methods. Our results demonstrate that Dis2Vec outperforms traditional distributed vector representations in its ability to faithfully capture taxonomical attributes across different class of diseases such as endemic, emerging and rare. Saurav Ghosh, Prithwish Chakraborty, Emily Cohn, John S. Brownstein, Naren Ramakrishnan |
CIKM | 2 |
| 2015 | Dynamic Poisson Autoregression for Influenza-Like-Illness Case Count PredictionabstractInfluenza-like-illness (ILI) is among of the most common diseases worldwide, and reliable forecasting of the same can have significant public health benefits. Recently, new forms of disease surveillance based upon digital data sources have been proposed and are continuing to attract attention over traditional surveillance methods. In this paper, we focus on short-term ILI case count prediction and develop a dynamic Poisson autoregressive model with exogenous inputs variables (DPARX) for flu forecasting. In this model, we allow the autoregressive model to change over time. In order to control the variation in the model, we construct a model similarity graph to specify the relationship between pairs of models at two time points and embed prior knowledge in terms of the structure of the graph. We formulate ILI case count forecasting as a convex optimization problem, whose objective balances the autoregressive loss and the model similarity regularization induced by the structure of the similarity graph. We then propose an efficient algorithm to solve this problem by block coordinate descent. We apply our model and the corresponding learning method on historical ILI records for 15 countries around the world using a variety of syndromic surveillance data sources. Our approach provides consistently better forecasting results than state-of-the-art models available for short-term ILI case count forecasting. Zheng Wang 0011, Prithwish Chakraborty, Sumiko R. Mekaru, John S. Brownstein, Jieping Ye, Naren Ramakrishnan |
KDD | 2 |
| 2014 | Forecasting a Moving Target: Ensemble Models for ILI Case Count PredictionsabstractModern epidemiological forecasts of common illnesses, such as the flu, rely on both traditional surveillance sources as well as digital surveillance data. However, most published studies have been retrospective. Concurrently, the reports about flu activity generally lags by several weeks and even when published are revised for several weeks more. We posit that effectively handling this uncertainty is one of the key challenges for a real-time prediction system in this sphere. In this paper, we present a detailed prospective analysis on the generation of robust quantitative predictions about temporal trends of flu activity, using several surrogate data sources for 15 Latin American countries. We present our findings about the limitations and possible advantages of correcting the uncertainty associated with official flu estimates. We also compare the prediction accuracy between model-level fusion of different surrogate data sources against data-level fusion. Finally, we present a novel matrix factorization approach using neighborhood embedding to predict flu case counts. Comparing our proposed ensemble method against several baseline methods helps us demarcate the importance of different data sources for the countries under consideration. Prithwish Chakraborty, Pejman Khadivi, Bryan L. Lewis, Aravindan Mahendiran, Jiangzhuo Chen, Patrick Butler, Elaine O. Nsoesie, Sumiko R. Mekaru, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
SDM | 1 |
| 2012 | Fine-Grained Photovoltaic Output Prediction Using a Bayesian EnsembleabstractLocal and distributed power generation is increasingly relianton renewable power sources, e.g., solar (photovoltaic or PV) andwind energy. The integration of such sources into the power grid ischallenging, however, due to their variable and intermittent energyoutput. To effectively use them on alarge scale, it is essential to be able to predict power generation at afine-grained level. We describe a novel Bayesian ensemble methodologyinvolving three diverse predictors. Each predictor estimates mixingcoefficients for integrating PV generation output profiles but capturesfundamentally different characteristics. Two of them employ classicalparameterized (naive Bayes) and non-parametric (nearest neighbor) methods tomodel the relationship between weather forecasts and PV output. The thirdpredictor captures the sequentiality implicit in PV generation and uses motifsmined from historical data to estimate the most likely mixture weights usinga stream prediction methodology. We demonstrate the success and superiority of ourmethods on real PV data from two locations that exhibit diverse weatherconditions. Predictions from our model can be harnessed to optimize schedulingof delay tolerant workloads, e.g., in a data center. Prithwish Chakraborty, Manish Marwah, Martin F. Arlitt, Naren Ramakrishnan |
AAAI | 1 |
| 2012 | Discrete harmony search based expert model for epileptic seizure detection in electroencephalography
Tapan Kumar Gandhi, Prithwish Chakraborty, Gourab Ghosh Roy, Bijaya K. Panigrahi |
Expert Syst. Appl. | 2 |
| 2011 | On convergence of the multi-objective particle swarm optimizers
Prithwish Chakraborty, Swagatam Das, Gourab Ghosh Roy, Ajith Abraham |
Inf. Sci. | 1 |
| 2011 | Erratum to "On convergence of the multi-objective particle swarm optimizers" [Inform. Sci 181 (2011) 1411-1425]
Prithwish Chakraborty, Swagatam Das, Gourab Ghosh Roy, Ajith Abraham |
Inf. Sci. | 1 |
| 2010 | On convergence of multi-objective Particle Swarm OptimizersabstractSeveral variants of the Particle Swarm Optimization (PSO) algorithm have been proposed in recent past to tackle the multi-objective optimization problems based on the concept of Pareto optimality. Although a plethora of significant research articles have so far been published on analysis of the stability and convergence properties of PSO as a single-objective optimizer, till date, to the best of our knowledge, no such analysis exists for the multi-objective PSO (MOPSO) algorithms. This paper presents a first, simple analysis of the general Pareto-based MOPSO and finds conditions on its most important control parameters (the inertia factor and acceleration coefficients) that control the convergence behavior of the algorithm to the Pareto front in the objective function space. Limited simulation supports have also been provided to substantiate the theoretical derivations. Prithwish Chakraborty, Swagatam Das, Ajith Abraham, Václav Snásel, Gourab Ghosh Roy |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | Artificial foraging weeds for global numerical optimization over continuous spacesabstractInvasive Weed Optimization (IWO) is a recently developed derivative-free metaheuristic algorithm that mimics the robust process of weeds colonization and distribution in an ecosystem. On the other hand central to an ecosystem is the foraging behavior that pertains to the act of searching for food and forms an integral part of the daily life of most of the living creatures. For over past two decades, a few significant optimization algorithms were developed by emulating the foraging behavior of creatures like ants, bacteria, fish, bees etc. This article presents a hybrid real-parameter optimizer developed by incorporating the principles of Optimal Foraging Theory (OFT) in IWO, with a view to improving the search mechanism of the latter over discontinuous and multi-modal fitness landscapes, riddled with local optima. The hybridization does not impose any serious computational burden on IWO in terms of increasing number of Function Evaluations (FEs). The performance of the resulting hybrid algorithm has been compared with eleven other state-of-the-art metaheuristic algorithms over a test-suite of 16 numerical benchmarks taken from the CEC (Congress on Evolutionary Computation) 2005 competition and special session on real parameter optimization. Our simulation experiments indicate that the proposed algorithm is able to attain comparable results against the nine other optimizers. Owing to its promising performance on benchmarks and ease of implementation (without requiring much programming overhead), the proposed algorithm may serve as an attractive alternative for a plethora of practical optimization problems. Gourab Ghosh Roy, Prithwish Chakraborty, Shi-Zheng Zhao, Swagatam Das, Ponnuthurai N. Suganthan |
IEEE Congress on Evolutionary Computation | 2 |
| 2009 | An Improved Harmony Search Algorithm with Differential Mutation OperatorabstractHarmony Search (HS) is a recently developed stochastic algorithm which imitates the music improvisation process. In this process, the musicians improvise their instrument pitches searching for the perfect state of harmony. Practical experiences, however, suggest that the algorithm suffers from the problems of slow and/or premature convergence over multimodal and rough fitness landscapes. This paper presents an attempt to improve the search performance of HS by hybridizing it with Differential Evolution (DE) algorithm. The performance of the resulting hybrid algorithm has been compared with classical HS, the global best HS, and a very popular variant of DE over a test-suite of six well known benchmark functions and one interesting practical optimization problem. The comparison is based on the following performance indices - (i) accuracy of final result, (ii) computational speed, and (iii) frequency of hitting the optima. Prithwish Chakraborty, Gourab Ghosh Roy, Swagatam Das, Dhaval Jain, Ajith Abraham |
Fundam. Informaticae | 1 |