EDBT 2026 Demo / reviewers in the wild / expert
Vibhuti Gupta
dblp:213/1586
· DBLP profile ↗
7ranked-venue papers in the field
2as first author
5since 2021 · last 2024
0000-0002-6221-4712ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 7 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Predicting Cardiac Complications of Myocardial Infarction Patients Using Machine LearningabstractIn the United States, heart disease is the leading cause of death, killing about 695,000 people each year. Myocardial infarction (MI) is a cardiac complication which occurs when blood flow to a portion of the heart decreases or halts, leading to damage in the heart muscle. Heart failure and Atrial fibrillation (AF) are closely associated with MI. Heart failure is a common complication of MI and a risk factor for AF. Machine learning (ML) and deep learning techniques have shown potential in predicting cardiovascular conditions. However, developing a simplified predictive model, along with a thorough feature analysis, is challenging due to various factors, including lifestyle, age, family history, medical conditions, and clinical variables for cardiac complications prediction. This paper aims to develop simplified models with comprehensive feature analysis and data preprocessing for predicting cardiac complications, such as heart failure and atrial fibrillation linked with MI, using a publicly available dataset of myocardial infarction patients. This will help the students and health care professionals understand various factors responsible for cardiac complications through a simplified workflow. By prioritizing interpretability, this paper illustrates how simpler models, like decision trees and logistic regression, can provide transparent decision-making processes while still maintaining a balance with accuracy. Additionally, this paper examines how age-specific factors affect heart failure and atrial fibrillation conditions. Overall this research focuses on making machine learning accessible and interpretable. Its goal is to equip students and non-experts with practical tools to understand how ML can be applied in healthcare, particularly for the cardiac complications prediction for patients having MI. Shriyansh Baidya, Vibhuti Gupta |
IEEE Big Data | 2 |
| 2024 | Harnessing the Power of Vocal Signals in COVID-19 Detection Utilizing Machine LearningabstractThe global COVID-19 pandemic has strained health-care systems and highlighted the need for accessible and efficient diagnostic methods. Traditional diagnostic tools, such as nasal swabs and biosensors, while accurate, pose significant logistical challenges and high costs, limiting their scalability. This paper explores an alternative, non-invasive approach to COVID-19 detection using machine learning algorithms to analyze vocal patterns, particularly cough and breathing sounds. Leveraging a publicly available dataset, we developed machine learning models capable of classifying audio samples as COVID-19 positive or negative. Our models achieve an AUC of up to 85% and an F1-score of 81%, demonstrating the potential of machine learning in enabling rapid, cost-effective COVID-19 diagnosis. These findings suggest that audio-based diagnostics could be a practical and scalable solution, particularly in resource-limited settings where traditional methods are less feasible. Aleesa Mann, Ajinkya P. Jadhav, Richard Matovu, Vibhuti Gupta |
IEEE Big Data | 4 |
| 2024 | A Novel Pipeline for Virus Integration Sites Detection in Tumor Genomes Using Deep LearningabstractCancer is one of the leading causes of death worldwide. Pathogenic viruses are estimated to be responsible for 15% of all human cancers globally and pose significant threats to public health. Viruses integrate their genetic material into the host genome, increasing the risk of cancer promoting changes in it. To understand the molecular mechanisms of virus-mediated cancers, it is crucial to identify viral insertion sites in cancer genomes. However, this effort is hindered by the rapidly increasing volume of tumor sequencing data, along with the challenges of accurate data analysis caused by high viral mutation rates and the difficulty of aligning short reads to the reference genome. Thus it is crucial to develop an efficient method for virus integration site detection in tumor genomes. This paper proposes a novel pipeline to identify viral integration sites leveraging deep Convolutional Neural Networks (CNN). Our contributions are twofold: (i) We propose and integrate three novel matrix generation methods into the pipeline, developed after aligning the host and viral genomes with their respective reference genomes.; (ii) We employ one-hot encoded images with reduced computational complexity to represent viral integration sites and harness the capabilities of Deep CNN networks for detection. The paper illustrates our proposed approach and presents experiments conducted using both synthetic and real sequencing data. Our preliminary experimental results are promising, showcasing the effectiveness of the proposed methods in detecting viral integration sites. Lorrayya Williams, Vibhuti Gupta |
IEEE Big Data | 2 |
| 2024 | Preparing Wearable Data for AI-Powered Mood and Compliance Prediction in HCT Patients and CaregiversabstractHematopoietic stem cell transplantation (HCT) is a potentially life-saving treatment that uses healthy blood-forming cells from donors to replace dysfunctional or damaged hematopoietic cells in patients with various blood disorders. This procedure is often employed to treat conditions such as hematological malignancies (e.g., leukemia, lymphoma, myeloma) and other severe blood or immune system diseases. Monitoring post-transplant complications is essential for tracking physiological effects and aiding in clinical decision-making. Biobehavioral aspects of care partners (i.e., unpaid caregivers) can also be influenced during the post-transplant stage of HCT. Wearable devices offer a non-invasive way to continuously track physiological parameters, making them a valuable resource for health monitoring. However, the physiological data collected from wearables is highly unstructured, often containing missing values, outliers, redundant features, and erroneous measurements leading to false conclusions/prediction. Therefore, enhancing data quality is essential for deriving meaningful insights. This paper introduces novel pre-processing methods to build a high quality, comprehensive, standardized, AI/ML ready, and clinically meaningful wearable dataset of HCT patients and caregivers. To test our data cleaning implementation, our cleaned, high-quality dataset is utilized to predict mood and compliance in HCT patients and their caregivers using machine learning algorithms. The paper illustrates our proposed approach and presents experimental results conducted on the data collected from Michigan Medicine for HCT patients and caregivers. Our preliminary experimental results are promising, demonstrating the effectiveness of the proposed methods and the high-quality dataset in predicting mood and compliance for the participants. Charles B. Ziegenbein, Bengie L. Ortiz, Vibhuti Gupta, Sung Won Choi |
IEEE Big Data | 3 |
| 2021 | Measles Rash Identification Using Transfer Learning and Deep Convolutional Neural NetworksabstractMeasles is a highly contagious disease, one of the largest vaccine-preventable illnesses and leading causes of death in developing countries, claiming more than 140,000 lives each year. Measles was declared eliminated in the United States in the year 2000 due to decades of successful vaccination but it resurged in 2019 with 1,282 confirmed cases. Due to rapid spread of this disease among people in contact, rapid and automated diagnostic systems are required for early prevention. In this work, we employed transfer learning to build deep convolutional neural networks (CNNs) to distinguish measles rash from other skin conditions. Experiments with ResNet-50 model, trained on our diverse and curated skin rash image dataset, produce classification accuracy of 95.2%, sensitivity of 81.7%, and specificity of 97.1%, respectively. This indicates that our technique is effective in facilitating an accurate detection of measles to help contain outbreaks. The performance of a small CNN model MobileNet-V2 on our image data set is also discussed. Our work will facilitate healthcare professionals to effectively diagnose measles and accelerate the development of automated diagnostic tools to prevent the measles spread at various public venues. Kimberly Glock, Charlie Napier, Todd Gary, Vibhuti Gupta, Joseph Gigante, William Schaffner, Qingguo Wang |
IEEE BigData | 4 |
| 2018 | Unleashing the Power of Hashtags in Tweet Analytics with Distributed Framework on Apache StormabstractTwitter is a popular social network platform where users can interact and post texts of up to 280 characters called tweets. Hashtags, hyperlinked words in tweets, have increasingly become crucial for tweet retrieval and search. Using hashtags for tweet topic classification is a challenging problem because of context dependent among words, slangs, abbreviation and emoticons in a short tweet along with evolving use of hashtags. Since Twitter generates millions of tweets daily, tweet analytics is a fundamental problem of Big data stream that often requires a real-time Distributed processing. This paper proposes a distributed online approach to tweet topic classification with hashtags. Being implemented on Apache Storm, a distributed real time framework, our approach incrementally identifies and updates a set of strong predictors in the Naïve Bayes model for classifying each incoming tweet instance. Preliminary experiments show promising results with up to 97% accuracy and 37% increase in throughput on eight processors. Vibhuti Gupta, Rattikorn Hewett |
IEEE BigData | 1 |
| 2017 | Harnessing the power of hashtags in tweet analyticsabstractTwitter is one of the most popular microblogging platforms where users can interact with each other by posting texts of up to 140 characters called tweets. Because of the large and fast growing number of tweets being generated daily, tweet analytics is viewed as one of the fundamental problems of Big data stream. Recently, hashtags, hyperlinked words, in tweets have been applied for tweet retrieval, trend/event detection and advertisement. However, using hashtags for tweet classification remains challenging because we have to cope with context dependent words, slangs, abbreviations, and emoticons with a limited small number of words and an evolving use of hashtags. Most existing approaches deal with classifying tweet sentiments by using the lexicon and meaning of hashtags. Our research aims to classify tweets by topics. Unlike sentiment analytics, hashtags for describing a topic need to be more diverse to cover various aspects of a topic. This paper presents a tweet analytics approach that uses domain-specific knowledge to create a set of strong hashtag predictors for tweet topic classification. The paper describes the approach and preliminary experiments that show promising results toward Big data tweet analytics. Vibhuti Gupta, Rattikorn Hewett |
IEEE BigData | 1 |