EDBT 2026 Demo / reviewers in the wild / expert
Sobhan Babu Chintapalli
dblp:233/3456 · also Ch. Sobhan Babu
· DBLP profile ↗
11ranked-venue papers in the field
0as first author
7since 2021 · last 2025
0000-0003-1645-8502ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EHR Can-Flow: Contrastively Aligned Normalizing Flows for Conditional EHR Generation
Subbareddy Batreddy, Giridhar Pamisetty, Priya Verma, Sobhan Babu Chintapalli |
IEEE Big Data | 4 |
| 2025 | UniFi-LLM: A Unified Large Language Model for Financial Data Generation and Fraud Prediction
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli |
IEEE Big Data | 4 |
| 2025 | HiCARE-EHR: Hierarchical Conditional AutoRegressive Events for High-Fidelity EHRs
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli |
IEEE Big Data | 4 |
| 2024 | Adaptive Neighborhood Sampling and KAN-Based Edge Features for Enhanced Fraud DetectionabstractFraud detection models have shown immense improvement in performance with the incorporation of graph neural networks (GNNs), as they are naturally capable of modeling relational data and capturing complex relations between users. However, real-world data has a significant class imbalance with very few fraudulent labels, resulting in a bias toward predicting the majority class as the decision boundaries are skewed towards it. To alleviate this class imbalance, we propose adaptive neighborhood sampling, an effective and optimal neighborhood sampling approach that uses sampling probability comprising of class frequency and global importance of the node through the page rank to reduce the dilution of the fraud nodes. We also propose a novel Kolmogorov-Arnold Network (KAN) based edge feature extraction approach, which treats incoming and outgoing edges separately to effectively capture the directional patterns that are crucial in distinguishing the patterns of both labels. In addition to this, we employed topology-aware margin loss to further improve the class imbalanced node classification. Incorporating these proposed approaches in the GNN-based fraud detection model helped outperform the state-of-the-art fraud models on open-source and complex real-world datasets. Subbareddy Batreddy, Giridhar Pamisetty, Manish Kothuri, Priya Verma, Sobhan Babu Chintapalli |
IEEE Big Data | 5 |
| 2024 | EHRGPT: Leveraging language models for generating synthetic health recordsabstractElectronic health records (EHRs) are the comprehensive digital records containing patient health information, which help in various domains like public health monitoring, predictive modeling, data-driven research, etc. However, there are strict regulations about data sharing due to the sensitivity of the EHR data. So, there is an increasing demand for generating high-quality synthetic EHR data that mitigates privacy concerns. Most of the existing approaches to EHR generation are limited as they predominantly generate binary decisions about affecting a disease or count-based summaries of disease occurrences for a patient, failing to capture the full patient trajectory contained in the real data and hence reducing the downstream applications. With the success of language models in generative approaches, we propose a language modeling approach using transformers to generate the complete health record of a patient for each admission. We propose to group the patient’s EHRs as it enhances the model to capture temporal coherence and the medical history of the corresponding patient. The proposed model can generate high-fidelity privacy preserved EHR data with the statistical properties of the real data. Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli |
IEEE Big Data | 4 |
| 2022 | Representation Learning on Graphs to Identifying Circular Trading in Goods and Services TaxabstractCircular trading is a form of tax evasion in Goods and Services Tax where a group of fraudulent taxpayers (traders) aims to mask illegal transactions by superimposing several fictitious transactions ( where no value is added to the goods or service) among themselves in a short period. Due to the vast database of taxpayers, it is infeasible for authorities to manually identify groups of circular traders and the illegitimate transactions they are involved in. This work uses big data analytics and graph representation learning techniques to propose a framework to identify communities of circular traders and isolate the illegitimate transactions in the respective communities. Our approach is tested on real-life data provided by the Department of Commercial Taxes, Government of Telangana, India, where we uncovered several communities of circular traders. Priya Mehta, Sanat Bhargava, K. Sandeep Kumar, M. Ravi Kumar 0002, Sobhan Babu Chintapalli |
IEEE Big Data | 5 |
| 2022 | Enhancement to Training of Bidirectional GAN : An Approach to Demystify Tax FraudabstractOutlier detection is a challenging activity. Several machine learning techniques are proposed in the literature for outlier detection. In this article, we propose a new training approach for bidirectional GAN (BiGAN) to detect outliers. To validate the proposed approach, we train a BiGAN with the proposed training approach to detect taxpayers, who are manipulating their tax returns. For each taxpayer, we derive six correlation parameters and three ratio parameters from tax returns submitted by him/her. We train a BiGAN with the proposed training approach on this nine-dimensional derived ground-truth data set. Next, we generate the latent representation of this data set using the encoder (encode this data set using the encoder) and regenerate this data set using the generator (decode back using the generator) by giving this latent representation as the input. For each taxpayer, compute the cosine similarity between his/her ground-truth data and regenerated data. Taxpayers with lower cosine similarity measures are potential return manipulators. We applied our method to analyze the iron and steel taxpayer’s data set provided by the Commercial Taxes Department, Government of Telangana, India. Priya Mehta, M. Ravi Kumar 0002, Sobhan Babu Chintapalli |
IEEE Big Data | 4 |
| 2020 | DeepCatch: Predicting Return Defaulters in Taxation System using Example-Dependent Cost-Sensitive Deep Neural NetworksabstractTax evasion is most common in several nations. Taxpayers evade tax by using thoughtful and well-considered techniques, which hinders the economic progress of the nation. Delaying the filing of returns by taxpayers is the most primitive form of tax evasion. Taxpayers who delay the filing of returns are called return defaulters. It is the most brazen form of tax evasion. To tackle this problem, we introduce an example-dependent cost-sensitive deep learning model to identify potential return defaulters. This model takes example-dependent costs into account and makes predictions that aim to minimize the overall cost instead of minimizing the total number of misclassifications. Applying our method, we show cost savings of about 55%. This work is designed and implemented for the Commercial Taxes Department Government of Telangana, India. Priya Mehta, Sobhan Babu Chintapalli, S. V. Kasi Visweswara Rao, K. Sandeep Kumar |
IEEE BigData | 2 |
| 2018 | Predictive Modeling for Identifying Return Defaulters in Goods and Services TaxabstractTax evasion is an illegal practice where a person or a business entity intentionally avoids paying his/her true tax liability. Any business entity is required by the law to file their tax return statements following a periodical schedule. Avoiding to file the tax return statement is one among the most rudimentary forms of tax evasion. The dealers committing tax evasion in such a way are called return defaulters. In this paper, we construct a logistic regression model that predicts with high accuracy whether a business entity is a potential return defaulter for the upcoming tax-filing period. For the same, we analyzed the effect of the amount of sales/purchases transactions among the business entities (dealers) and the mean absolute deviation (MAD) value of the first digit Benford's law on sales transactions by a business entity. We developed this model for the commercial taxes department, government of Telangana, India. Priya Mehta, Jithin Mathews, K. Suryamukhi, K. Sandeep Kumar, Sobhan Babu Chintapalli |
DSAA | 5 |
| 2018 | Regression Analysis towards Estimating Tax Evasion in Goods and Services TaxabstractTax evasion is as old as tax itself. In this paper, we devise a technique to predict the amount of tax-revenue lost by the state due to unscrupulous actions from a particular set of suspicious dealers. For the same, we build a regression model using the tax-return information of genuine business dealers and predict the amount of tax evaded by suspicious business dealers. Dealers are classified as genuine or suspicious by applying Benford's analysis on the different group of dealers formed after running k-medoids clustering algorithm over a set of dealers. In addition to getting an estimate on the loss of tax-revenue, results obtained from this work aid the tax enforcement officers on taking precautionary measures against tax evasion. The dataset used in the work is provided by the commercial tax department of Telangana state, India. Jithin Mathews, Priya Mehta, Suryamukhi Kuchibhotla, Dikshant Bisht, Sobhan Babu Chintapalli, S. V. Kasi Visweswara Rao |
WI | 5 |
| 2012 | Which sort orders are interesting?
Ravindra Guravannavar, S. Sudarshan 0001, Ajit A. Diwan, Sobhan Babu Chintapalli |
VLDB J. | 4 |