Subbareddy Batreddy

dblp:295/4230 · DBLP profile ↗
← Back
5ranked-venue papers in the field
2as first author
5since 2021 · last 2025
0000-0002-3265-4100ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5 (2 first)
YearPublicationVenuePosition
2025 EHR Can-Flow: Contrastively Aligned Normalizing Flows for Conditional EHR Generation
Subbareddy Batreddy, Giridhar Pamisetty, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data1
2025 UniFi-LLM: A Unified Large Language Model for Financial Data Generation and Fraud Prediction
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data2
2025 HiCARE-EHR: Hierarchical Conditional AutoRegressive Events for High-Fidelity EHRs
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data2
2024 Adaptive Neighborhood Sampling and KAN-Based Edge Features for Enhanced Fraud Detection
abstract
Fraud detection models have shown immense improvement in performance with the incorporation of graph neural networks (GNNs), as they are naturally capable of modeling relational data and capturing complex relations between users. However, real-world data has a significant class imbalance with very few fraudulent labels, resulting in a bias toward predicting the majority class as the decision boundaries are skewed towards it. To alleviate this class imbalance, we propose adaptive neighborhood sampling, an effective and optimal neighborhood sampling approach that uses sampling probability comprising of class frequency and global importance of the node through the page rank to reduce the dilution of the fraud nodes. We also propose a novel Kolmogorov-Arnold Network (KAN) based edge feature extraction approach, which treats incoming and outgoing edges separately to effectively capture the directional patterns that are crucial in distinguishing the patterns of both labels. In addition to this, we employed topology-aware margin loss to further improve the class imbalanced node classification. Incorporating these proposed approaches in the GNN-based fraud detection model helped outperform the state-of-the-art fraud models on open-source and complex real-world datasets.
Subbareddy Batreddy, Giridhar Pamisetty, Manish Kothuri, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data1
2024 EHRGPT: Leveraging language models for generating synthetic health records
abstract
Electronic health records (EHRs) are the comprehensive digital records containing patient health information, which help in various domains like public health monitoring, predictive modeling, data-driven research, etc. However, there are strict regulations about data sharing due to the sensitivity of the EHR data. So, there is an increasing demand for generating high-quality synthetic EHR data that mitigates privacy concerns. Most of the existing approaches to EHR generation are limited as they predominantly generate binary decisions about affecting a disease or count-based summaries of disease occurrences for a patient, failing to capture the full patient trajectory contained in the real data and hence reducing the downstream applications. With the success of language models in generative approaches, we propose a language modeling approach using transformers to generate the complete health record of a patient for each admission. We propose to group the patient’s EHRs as it enhances the model to capture temporal coherence and the medical history of the corresponding patient. The proposed model can generate high-fidelity privacy preserved EHR data with the statistical properties of the real data.
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data2