VLDB 2026 Research / reviewers in the wild / expert
Praveen Madiraju
dblp:m/PraveenMadiraju
· DBLP profile ↗
13ranked-venue papers in the field
1as first author
10since 2021 · last 2025
0009-0006-9737-9601ORCID · verified
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 12Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explaining Pre-Trained Language Models in the Context of Higher Education
Kevin Chovanec, John Fields, Praveen Madiraju |
IEEE Big Data | 3 |
| 2025 | MultiMentalRoBERTa: A Fine-Tuned Multiclass Classifier for Mental Health Disorder
K. M. Sajjadul Islam, John Fields, Praveen Madiraju |
IEEE Big Data | 3 |
| 2025 | Prediction of Long COVID and Mortality Among Patients With Substance use Disorder
K. M. Sajjadul Islam, Praveen Madiraju |
IEEE Big Data | 3 |
| 2024 | Integrating categorical and continuous data in a cluster-then-classify methodology for predicting undergraduate student successabstractStudent retention in higher education remains a significant challenge despite decades of research. This study introduces a novel cluster-label-classify methodology to predict at-risk students and identify common characteristics among those who drop out. Using data from a small private Midwestern university in the United States, we first applied the K-Prototypes algorithm to cluster non-retained students into five groups based on both numeric and categorical variables. We then labeled these clusters and the retained students, creating a multi-class classification problem. Finally, we used a Gradient Boosting Classifier and XGBoost for classification, achieving F1 scores of 0.82 to 0.89 for predicting non-retained students and 0.96 for retained students after addressing class imbalance with SMOTE. This approach allows for customized labeling specific to each institution and enables more targeted interventions for at-risk students. Our methodology combines demographic, academic, and socioeconomic factors to provide a comprehensive view of student retention, potentially offering new insights into this longstanding issue in higher education. The paper also discusses algorithmic bias, examining potential fairness issues in the predictive models and their implications for different student populations. Finally, a discussion on Privacy Preserving Machine Learning (PPML) provides future strategies for testing how these technologies generalize to other institutions while enhancing the privacy of student data. John Fields, Kevin Chovanec, Praveen Madiraju |
IEEE Big Data | 3 |
| 2023 | Combining Demographic Tabular Data with BERT Outputs for Multilabel Text Classification in Higher Education Survey DataabstractInstitutions of Higher Education (HEI) often possess rich text data in the form of student surveys. However, because these data are expensive to process, many universities have not yet capitalized on this resource. When working with student text data, researchers often desire to first label student responses with common categories of interest, a task in Natural Language Processing known as Multi-label Text Classification (MLTC). BERT and other Large Language Models have produced state of the art results on MLTC tasks; yet because MLTC generally presents challenges of data scarcity and data sparsity, accuracy often remains too low to fully automate the task. Unlike many common MLTC datasets, these student survey data can usually be paired with rich tabular data, both academic and demographic. In this paper, we show that a fusion approach combining tabular data with BERT outputs derived from student responses significantly improves model performance, increasing label ranking average precision from.75 to.84. The paper thus contributes to the open academic discussion of whether fusing tabular demographic data with BERT outputs improves performance, and also offers a practical approach for HEIs to automate survey labeling and thus incorporate more student text data into institutional research. Kevin Chovanec, John Fields, Praveen Madiraju |
IEEE Big Data | 3 |
| 2023 | Autocompletion of Chief Complaints in the Electronic Health Records using Large Language ModelsabstractThe Chief Complaint (CC) is a crucial component of a patient’s medical record as it describes the main reason or concern for seeking medical care. It provides critical information for healthcare providers to make informed decisions about patient care. However, documenting CCs can be time-consuming for healthcare providers, especially in busy emergency departments. To address this issue, an autocompletion tool that suggests accurate and well-formatted phrases or sentences for clinical notes can be a valuable resource for triage nurses. In this study, we utilized text generation techniques to develop machine learning models using CC data. In our proposed work, we train a Long Short-Term Memory (LSTM) model and fine-tune three different variants of Biomedical Generative Pretrained Transformers (BioGPT), namely microsoft/biogpt, microsoft/BioGPT-Large, and microsoft/BioGPT-Large-PubMedQA. Additionally, we tune a prompt by incorporating exemplar CC sentences, utilizing the OpenAI API of GPT-4. We evaluate the models’ performance based on the perplexity score, modified BERTScore, and cosine similarity score. The results show that BioGPT-Large exhibits superior performance compared to the other models. It consistently achieves a remarkably low perplexity score of 1.65 when generating CC, whereas the baseline LSTM model achieves the best perplexity score of 170. Further, we evaluate and assess the proposed models’ performance and the outcome of GPT-4.0. Our study demonstrates that utilizing LLMs such as BioGPT, leads to the development of an effective autocompletion tool for generating CC documentation in healthcare settings. K. M. Sajjadul Islam, Ayesha Siddika Nipu, Praveen Madiraju, Priya Deshpande |
IEEE Big Data | 3 |
| 2023 | Predicting Mental Health Disorders Post Long COVID Diagnosis Using Advanced Machine Learning TechniquesabstractAfter the global spread of COVID-19, the enduring effects of Long COVID and its health implications have emerged as a significant global issue, affecting people worldwide. The lingering symptoms post a COVID-19 infection can significantly affect individuals who had previously contracted the virus, exerting considerable influence over their mental well-being. Prolonged recuperation associated with Long COVID has been connected with the emergence of symptoms such as depression and anxiety, all of which can have adverse effects on emotional health. This paper delves into an in-depth analysis of healthcare data pertaining to Long COVID from the Froedtert Health Medical System in Wisconsin. Through the application of advanced Machine Learning (ML) techniques, we present predictive models aimed at assessing the risk of developing Mental Health Disorders (MHD) in patients diagnosed with Long COVID. Our study also encompasses the identification of pivotal features impacting MHD. To thoroughly investigate the factors that have a substantial impact on MHD, we employed the Recursive Feature Elimination (RFE) technique to carefully pick out essential attributes from our dataset. Given the dataset’s inherent imbalance, we have employed the Synthetic Minority Over-sampling Technique and Edited Nearest Neighbors (SMO-TEEN) technique to effectively address this issue. Multiple ML models have been meticulously constructed and validated using cross-validation methodologies. The results indicate that Random Forest (RF) Classifier shows better performance in comparison to other models with an area under the ROC curve (AUC) of 0.97, precision of 0.90, and recall of 0.89. Remarkably, the XGBoost Classifier also demonstrates strong predictive abilities for MHD, achieving an AUC of 0.90, precision of 0.79, and recall of 0.82. Ultimately, the crucial features identified through our predictive models hold the potential to identify individuals at risk of MHD, facilitating the delivery of targeted preventive care and essential resources. Manoj Purohit, Praveen Madiraju |
IEEE Big Data | 2 |
| 2022 | Oversampling techniques for predicting COVID-19 patient length of stayabstractCOVID-19 is a respiratory disease that caused a global pandemic in 2019. It is highly infectious and has the following symptoms: fever or chills, cough, shortness of breath, fatigue, muscle or body aches, headache, the new loss of taste or smell, sore throat, congestion or runny nose, nausea or vomiting, and diarrhea. These symptoms vary in severity; some people with many risk factors have been known to have lengthy hospital stays or die from the disease. In this paper, we analyze patients’ electronic health records (EHR) to predict the severity of their COVID-19 infection using the length of stay (LOS) as our measurement of severity. This is an imbalanced classification problem, as many people have a shorter LOS rather than a longer one. To combat this problem, we synthetically create alternate oversampled training data sets. Once we have this oversampled data, we run it through an Artificial Neural Network (ANN), which during training has its hyperparameters tuned by using bayesian optimization. We select the model with the best F1 score and then evaluate it and discuss it. Zach Farahany, K. M. Sajjadul Islam, Praveen Madiraju |
IEEE Big Data | 4 |
| 2021 | Identifying Precursors to Long-Term Crisis in Veterans Using Associative ClassifierabstractPost-Traumatic Stress Disorder (PTSD) is one of the most common mental health disorders prevalent in the US. Most alarming, PTSD occurs at double the rate for combat veterans compared to the general population. Severity of PTSD is associated with risk taking behaviors such as substance abuse, non-suicidal self-injury, sexual risk behaviors, among other negative behaviors. Psychological disorders are often preceded by crisis events, thus monitoring for crisis events can help prevent risky behavior in veterans. Ecological momentary assessment techniques are effective in capturing possible crisis events for veterans. Mobile apps are commonly used to gather such behavioral changes in participants. Crisis events collected from m-health can be analyzed for the identification of long- term PTSD risk. Early identification of risk can help in planning intervention to mitigate the risk. Many scholars have used traditional statistical and machine learning methods for the prediction of mental health issues in individuals. But these models lack transparency in how decisions are made. Providing justifications for the predictions can increase the reliability of the model. Our research focused on developing an explainable prediction model using class association rules to identify veterans at risk of persistent PTSD. The generated association rules serve as precursors to the long-term crisis in veterans. Results of the analysis showed that having no family support, little or no interest in hobbies, stress and lack of sleep are some of the influencing factors of persistent PTSD in veterans. Priyanka Annapureddy, Zeno Franco, Praveen Madiraju, Sheikh Iqbal Ahamed, Mark Flower, Md Fitrat Hossain, Md. Romael Haque, Nadiyah Johnson, Sabirat Rubya, Natalie Danielle Baker, Niharika Jain, Otis Winstead |
IEEE BigData | 3 |
| 2021 | A Machine Learning Approach to Predict Length of Stay for Opioid Overdose Admitted PatientsabstractPeople are prone to develop opioid dependence and other health problems due to regular non-medical use, prolonged use, and misuse of opioids. The number of hospital admissions for opioid dependence is growing across the US. The length of stay (LOS) is an essential indicator that assesses the severity of opioid overdose admissions. In this paper, opioid-related healthcare data from Froedtert Health Medical System in Wisconsin are analyzed and machine learning models are proposed to predict the LOS of opioid overdose admitted patients. We also determine important features that impact the LOS. To explore the factors that significantly influence the LOS, we implemented recursive feature elimination (RFE) to select important features from the data. Since the data set is imbalanced, we applied two imbalanced learning approaches to tackle that, namely, SMOTE and imbalanced learning models. Several machine learning models were constructed and validated with 10 iterations of 10-fold cross validation, and we fine-tuned the models with the highest f1 score and AUC score. Random Forest Classifier outperforms other models trained on the oversampled data set (AUC = 0.81, precision = 0.72, recall = 0.69). Easy Ensemble Classifier, which is trained on the imbalanced data set, has a better performance in predicting longer stays (AUC = 0.80, precision = 0.70, recall = 0.74). Finally, important features identified from the predictive models can be used to determine at-risk patients of longer LOS and provide appropriate preventive care and resources. Priyanka Annapureddy, Zach Farahany, Praveen Madiraju |
IEEE BigData | 4 |
| 2020 | Predicting Opioid Overdose Readmission and Opioid Use Disorder with Machine LearningabstractOpioid use disorder (OUD) is a medical condition associated with problematic patterns of opioid use that cause interpersonal and social impairment. This research demonstrates how supervised machine learning can be used to predict patients at risk of hospital readmission following opioid overdose, and to predict patients at risk of developing OUD. Two labeled datasets were built from deidentified hospital data provided by a Level I Trauma Center Hospital. Several machine learning models were constructed (logistic regression, random forest, support vector machine, AdaBoost, XGBoost) and validated with 10 iterations of 10-fold cross validation. The XGBoost classifier can sufficiently predict patients at risk for OUD (AUC = 0.78, precision = 0.71, recall = 0.53). This work can assist providers in determining appropriate preventive care and resources for at-risk patients. Sarah McDougall, Priyanka Annapureddy, Praveen Madiraju, Nicole Fumo, Stephen Hargarten |
IEEE BigData | 3 |
| 2020 | Analyzing Adverse Event Signal Detection with Publicly Available Web SourcesabstractData mining for drug-reaction associations is a major topic in the pharmaceutical industry. Historically the focus has been on using privately owned and maintained datasets consisting of information available in the FDA Adverse Event Reporting System (FAERS) and privatized reporting systems that house the data from clinical trials. Our focus will be on building a pipeline that demonstrates an open source solution for outlining a drug's safety profile from data collection through signal detection. This pipeline primarily uses data from the openFDA and the social media site Reddit, both of which provide a well-documented API. All data analysis for this pipeline is done in the R statistical programming language. The aim was to collect the information available in these public sources and apply popular data mining methodologies used to identify and predict the occurrence of adverse events. The results show the ability of the openFDA and social media sites to create real-time drug safety profiles by applying the same statistical methods applied in clinical trials. The social media data does not perform equally across all drug types and gives best results when applied to common over-the-counter drugs as opposed to last line of defense medications. Alex Salamun, Stephanie Duque, Praveen Madiraju |
IEEE BigData | 3 |
| 2006 | Semantic Integrity Constraint Checking for Multiple XML DatabasesabstractGlobal semantic integrity constraints ensure integrity and consistency of data spanning multiple databases. In this paper, we take initial steps towards representing global semantic integrity constraints for XML databases. We also provide a general framework for checking global semantic integrity constraints for XML databases. Furthermore, we set forth an efficient algorithm for checking global semantic integrity constraints across multiple XML databases. Our algorithm is efficient for three reasons: (1) the algorithm does not require the update statement to be executed before the constraint check is carried out; hence, we avoid any potential problems associated with rollbacks, (2) sub constraint checks are executed in parallel, and (3) most of the processing of algorithm could happen at compile time; hence, we save time spent at run-time. As a proof of concept, we present a prototype of the system implementing the ideas discussed in this paper. Praveen Madiraju, Rajshekhar Sunderraman, Shamkant B. Navathe |
J. Database Manag. | 1 |