Sobhan Babu Chintapalli

dblp:233/3456 · also Ch. Sobhan Babu · DBLP profile ↗
← Back
11ranked-venue papers in the field
0as first author
7since 2021 · last 2025
0000-0003-1645-8502ORCID · reported

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 8Database Systems & Data Management · 1Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2025 EHR Can-Flow: Contrastively Aligned Normalizing Flows for Conditional EHR Generation
Subbareddy Batreddy, Giridhar Pamisetty, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data4
2025 UniFi-LLM: A Unified Large Language Model for Financial Data Generation and Fraud Prediction
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data4
2025 HiCARE-EHR: Hierarchical Conditional AutoRegressive Events for High-Fidelity EHRs
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data4
2024 Adaptive Neighborhood Sampling and KAN-Based Edge Features for Enhanced Fraud Detection
abstract
Fraud detection models have shown immense improvement in performance with the incorporation of graph neural networks (GNNs), as they are naturally capable of modeling relational data and capturing complex relations between users. However, real-world data has a significant class imbalance with very few fraudulent labels, resulting in a bias toward predicting the majority class as the decision boundaries are skewed towards it. To alleviate this class imbalance, we propose adaptive neighborhood sampling, an effective and optimal neighborhood sampling approach that uses sampling probability comprising of class frequency and global importance of the node through the page rank to reduce the dilution of the fraud nodes. We also propose a novel Kolmogorov-Arnold Network (KAN) based edge feature extraction approach, which treats incoming and outgoing edges separately to effectively capture the directional patterns that are crucial in distinguishing the patterns of both labels. In addition to this, we employed topology-aware margin loss to further improve the class imbalanced node classification. Incorporating these proposed approaches in the GNN-based fraud detection model helped outperform the state-of-the-art fraud models on open-source and complex real-world datasets.
Subbareddy Batreddy, Giridhar Pamisetty, Manish Kothuri, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data5
2024 EHRGPT: Leveraging language models for generating synthetic health records
abstract
Electronic health records (EHRs) are the comprehensive digital records containing patient health information, which help in various domains like public health monitoring, predictive modeling, data-driven research, etc. However, there are strict regulations about data sharing due to the sensitivity of the EHR data. So, there is an increasing demand for generating high-quality synthetic EHR data that mitigates privacy concerns. Most of the existing approaches to EHR generation are limited as they predominantly generate binary decisions about affecting a disease or count-based summaries of disease occurrences for a patient, failing to capture the full patient trajectory contained in the real data and hence reducing the downstream applications. With the success of language models in generative approaches, we propose a language modeling approach using transformers to generate the complete health record of a patient for each admission. We propose to group the patient’s EHRs as it enhances the model to capture temporal coherence and the medical history of the corresponding patient. The proposed model can generate high-fidelity privacy preserved EHR data with the statistical properties of the real data.
Giridhar Pamisetty, Subbareddy Batreddy, Priya Verma, Sobhan Babu Chintapalli
IEEE Big Data4
2022 Representation Learning on Graphs to Identifying Circular Trading in Goods and Services Tax
abstract
Circular trading is a form of tax evasion in Goods and Services Tax where a group of fraudulent taxpayers (traders) aims to mask illegal transactions by superimposing several fictitious transactions ( where no value is added to the goods or service) among themselves in a short period. Due to the vast database of taxpayers, it is infeasible for authorities to manually identify groups of circular traders and the illegitimate transactions they are involved in. This work uses big data analytics and graph representation learning techniques to propose a framework to identify communities of circular traders and isolate the illegitimate transactions in the respective communities. Our approach is tested on real-life data provided by the Department of Commercial Taxes, Government of Telangana, India, where we uncovered several communities of circular traders.
Priya Mehta, Sanat Bhargava, K. Sandeep Kumar, M. Ravi Kumar 0002, Sobhan Babu Chintapalli
IEEE Big Data5
2022 Enhancement to Training of Bidirectional GAN : An Approach to Demystify Tax Fraud
abstract
Outlier detection is a challenging activity. Several machine learning techniques are proposed in the literature for outlier detection. In this article, we propose a new training approach for bidirectional GAN (BiGAN) to detect outliers. To validate the proposed approach, we train a BiGAN with the proposed training approach to detect taxpayers, who are manipulating their tax returns. For each taxpayer, we derive six correlation parameters and three ratio parameters from tax returns submitted by him/her. We train a BiGAN with the proposed training approach on this nine-dimensional derived ground-truth data set. Next, we generate the latent representation of this data set using the encoder (encode this data set using the encoder) and regenerate this data set using the generator (decode back using the generator) by giving this latent representation as the input. For each taxpayer, compute the cosine similarity between his/her ground-truth data and regenerated data. Taxpayers with lower cosine similarity measures are potential return manipulators. We applied our method to analyze the iron and steel taxpayer’s data set provided by the Commercial Taxes Department, Government of Telangana, India.
Priya Mehta, M. Ravi Kumar 0002, Sobhan Babu Chintapalli
IEEE Big Data4
2020 DeepCatch: Predicting Return Defaulters in Taxation System using Example-Dependent Cost-Sensitive Deep Neural Networks
abstract
Tax evasion is most common in several nations. Taxpayers evade tax by using thoughtful and well-considered techniques, which hinders the economic progress of the nation. Delaying the filing of returns by taxpayers is the most primitive form of tax evasion. Taxpayers who delay the filing of returns are called return defaulters. It is the most brazen form of tax evasion. To tackle this problem, we introduce an example-dependent cost-sensitive deep learning model to identify potential return defaulters. This model takes example-dependent costs into account and makes predictions that aim to minimize the overall cost instead of minimizing the total number of misclassifications. Applying our method, we show cost savings of about 55%. This work is designed and implemented for the Commercial Taxes Department Government of Telangana, India.
Priya Mehta, Sobhan Babu Chintapalli, S. V. Kasi Visweswara Rao, K. Sandeep Kumar
IEEE BigData2
2018 Predictive Modeling for Identifying Return Defaulters in Goods and Services Tax
abstract
Tax evasion is an illegal practice where a person or a business entity intentionally avoids paying his/her true tax liability. Any business entity is required by the law to file their tax return statements following a periodical schedule. Avoiding to file the tax return statement is one among the most rudimentary forms of tax evasion. The dealers committing tax evasion in such a way are called return defaulters. In this paper, we construct a logistic regression model that predicts with high accuracy whether a business entity is a potential return defaulter for the upcoming tax-filing period. For the same, we analyzed the effect of the amount of sales/purchases transactions among the business entities (dealers) and the mean absolute deviation (MAD) value of the first digit Benford's law on sales transactions by a business entity. We developed this model for the commercial taxes department, government of Telangana, India.
Priya Mehta, Jithin Mathews, K. Suryamukhi, K. Sandeep Kumar, Sobhan Babu Chintapalli
DSAA5
2018 Regression Analysis towards Estimating Tax Evasion in Goods and Services Tax
abstract
Tax evasion is as old as tax itself. In this paper, we devise a technique to predict the amount of tax-revenue lost by the state due to unscrupulous actions from a particular set of suspicious dealers. For the same, we build a regression model using the tax-return information of genuine business dealers and predict the amount of tax evaded by suspicious business dealers. Dealers are classified as genuine or suspicious by applying Benford's analysis on the different group of dealers formed after running k-medoids clustering algorithm over a set of dealers. In addition to getting an estimate on the loss of tax-revenue, results obtained from this work aid the tax enforcement officers on taking precautionary measures against tax evasion. The dataset used in the work is provided by the commercial tax department of Telangana state, India.
Jithin Mathews, Priya Mehta, Suryamukhi Kuchibhotla, Dikshant Bisht, Sobhan Babu Chintapalli, S. V. Kasi Visweswara Rao
WI5
2012 Which sort orders are interesting?
Ravindra Guravannavar, S. Sudarshan 0001, Ajit A. Diwan, Sobhan Babu Chintapalli
VLDB J.4