VLDB 2026 Research / reviewers in the wild / expert
Edward Bowen
dblp:312/6857
· DBLP profile ↗
14ranked-venue papers
0as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Query Efficiency and Accuracy of Transfer Learning-based Model Extraction Attack in Federated LearningabstractFederated Learning (FL) is a collaborative learning framework designed to protect client data, yet it remains highly vulnerable to Intellectual Property (IP) threats. Model extraction (ME) attack poses a significant risk to Machine-Learning-as-a-Service (MLaaS) platforms, enabling attackers to replicate confidential models by querying Black-Box (without internal insight) APIs. Despite FL’s privacy-preserving goals, its distributed nature makes it particularly susceptible to such attacks. This paper examines the vulnerability of the FL-based victim model to two types of model extraction attacks. For various federated clients built under NVFlare platform, we implemented ME attack across two deep-learning architectures and three image datasets. We evaluate the proposed ME attack performance using various metrics, including accuracy, fidelity, and KL divergence. The experiments show that for various FL clients, the accuracy and fidelity of the extraction model are closely related to the size of the attack query set. Additionally, we explore a transfer learning-based approach where pre-trained models serve as the starting point for the extraction process. The results indicate that the accuracy and fidelity of the fine-tuned pre-trained extraction models are notably higher, particularly with smaller query sets, highlighting potential advantages for attackers. Sayyed Farid Ahamed, Sandip Roy 0001, Soumya Banerjee 0001, Marc Vucovich, Kevin Choi, Abdul Rahman, Alison Hu, Edward Bowen, Sachin Shetty |
IWCMC | 8 |
| 2025 | RADEP: A Resilient Adaptive Defense Framework Against Model Extraction AttacksabstractMachine Learning as a Service (MLaaS) enables users to leverage powerful machine learning models through cloud-based APIs, offering scalability and ease of deployment. However, these services are vulnerable to model extraction attacks, where adversaries repeatedly query the application programming interface (API) to reconstruct a functionally similar model, compromising intellectual property and security. Despite various defense strategies being proposed, many suffer from high computational costs, limited adaptability to evolving attack techniques, and a reduction in performance for legitimate users. In this paper, we introduce a Resilient Adaptive Defense Framework for Model Extraction Attack Protection (RADEP), a multifaceted defense framework designed to counteract model extraction attacks through a multi-layered security approach. RADEP employs progressive adversarial training to enhance model resilience against extraction attempts. Malicious query detection is achieved through a combination of uncertainty quantification and behavioral pattern analysis, effectively identifying adversarial queries. Furthermore, we develop an adaptive response mechanism that dynamically modifies query outputs based on their suspicion scores, reducing the utility of stolen models. Finally, ownership verification is enforced through embedded watermarking and backdoor triggers, enabling reliable identification of unauthorized model use. Experimental evaluations demonstrate that RADEP significantly reduces extraction success rates while maintaining high detection accuracy with minimal impact on legitimate queries. Extensive experiments show that RADEP effectively defends against model extraction attacks and remains resilient even against adaptive adversaries, making it a reliable security framework for MLaaS models. Amit Chakraborty, Sayyed Farid Ahamed, Sandip Roy 0001, Soumya Banerjee 0001, Kevin Choi, Abdul Rahman, Alison Hu, Edward Bowen, Sachin Shetty |
IWCMC | 8 |
| 2025 | MentorPDM: Learning Data-Driven Curriculum for Multi-Modal Predictive MaintenanceabstractPredictive Maintenance (PDM) systems are essential for preemptive monitoring of sensor signals to detect potential machine component failures in industrial assets such as bearings in rotating machinery. Existing PDM systems face two primary challenges: 1) Irregular Signal Acquisition, where data collection from the sensors is intermittent, and 2) Signal Heterogeneity, where the full spectrum of sensor modalities is not effectively integrated. To address these challenges, we propose a Curriculum Learning Framework for Multi-Modal Predictive Maintenance - MentorPDM. MentorPDM consists of 1) a graph-augmented pretraining module that captures intrinsic and structured temporal correlations across time segments via a temporal contrastive learning objective and 2) a bi-level curriculum learning module that captures task complexities for weighing the importance of signal modalities and samples via modality and sample curricula. Empirical results from MentorPDM show promising performance with better generalizability in PDM tasks compared to existing benchmarks. The efficacy of the MentorPDM model will be further demonstrated in real industry testbeds and platforms. Shuaicheng Zhang, Sanmitra Bhattacharya, Sunil Reddy Tiyyagura, Edward Bowen, Balaji Veeramani, Dawei Zhou 0003 |
KDD (1) | 6 |
| 2025 | Fraud detection in healthcare claims using machine learning: A systematic reviewabstractOBJECTIVE: Identifying fraud in healthcare programs is crucial, as an estimated 3%-10% of the total healthcare expenditures are lost to fraudulent activities. This study presents a systematic literature review of machine learning techniques applied to fraud detection in health insurance claims. We aim to analyze the data and methodologies documented in the literature over the past two decades, providing insights into research challenges and opportunities. METHODS: We identified research studies on health insurance fraud detection using machine learning approaches from databases such as Google Scholar, Springer-Link journals, Elsevier, PubMed, Excerpta Medica Database (EMBASE), Scopus, the Association for Computing Machinery (ACM) Digital Library, and the Institute of Electrical and Electronics Engineers (IEEE) Xplore Digital Library. We included only articles that presented experimental results of machine learning-based approaches applied to healthcare claims. From the reviewed articles, 137 were selected for the final qualitative and quantitative analyses. RESULTS: In recent years, there has been a surge in publications centered on the use of machine learning to detect health insurance fraud. Among these studies, those focused on the detection of fraud committed by healthcare providers was the most prevalent, followed by fraud committed by patients. A wide variety of machine learning algorithms are highlighted in these studies, ranging from unsupervised (41 studies) and supervised methods (94 studies), to hybrid approaches (12 studies). While traditional machine learning approaches remain dominant in this research area, the adoption of advanced deep learning techniques is on the rise. Considering the type of healthcare claims data used, 30 studies utilized private data sources, while the rest used publicly available datasets. Data from 16 countries were utilized, with a majority coming from the United States (96 studies), followed by China (11 studies) and Australia (5 studies). DISCUSSION AND CONCLUSION: Detecting fraud in healthcare claims using machine learning presents several challenges. These include inconsistent data, absence of data standardization and integration, privacy concerns, and a limited number of labeled fraudulent cases to train models on. Future work should focus on enhancing transparency in data preparation, promoting the sharing of fraud investigation outcomes by authorities, and developing benchmark datasets to enhance accessibility and comparability. Furthermore, innovative techniques in data sampling, feature encoding methods for training machine learning models, and exploring the latest advancements in deep learning can significantly advance research in health insurance fraud detection. Anli du Preez, Sanmitra Bhattacharya, Peter A. Beling, Edward Bowen |
Artif. Intell. Medicine | 4 |
| 2024 | Data Composition for Continual Learning in Application of Cyberattack Detection
Jiayi Lian, Kevin Choi, Balaji Veeramani, Sathvik Murli, Alison Hu, Laura J. Freeman, Edward Bowen, Xinwei Deng |
ASONAM (4) | 8 |
| 2024 | EvoluNet: Advancing Dynamic Non-IID Transfer Learning on GraphsabstractNon-IID transfer learning on graphs is crucial in many high-stakes domains. The majority of existing works assume stationary distribution for both source and target domains. However, real-world graphs are intrinsically dynamic, presenting challenges in terms of domain evolution and dynamic discrepancy between source and target domains. To bridge the gap, we shift the problem to the dynamic setting and pose the question: given the *label-rich* source graphs and the *label-scarce* target graphs both observed in previous $T$ timestamps, how can we effectively characterize the evolving domain discrepancy and optimize the generalization performance of the target domain at the incoming $T+1$ timestamp? To answer it, we propose a generalization bound for *dynamic non-IID transfer learning on graphs*, which implies the generalization performance is dominated by domain evolution and domain discrepancy between source and target graphs. Inspired by the theoretical results, we introduce a novel generic framework named EvoluNet. It leverages a transformer-based temporal encoding module to model temporal information of the evolving domains and then uses a dynamic domain unification module to efficiently learn domain-invariant representations across the source and target domains. Finally, EvoluNet outperforms the state-of-the-art models by up to 12.1%, demonstrating its effectiveness in transferring knowledge from dynamic source graphs to dynamic target graphs. Haohui Wang, Yuzhen Mao, Yujun Yan, Yaoqing Yang 0002, Jianhui Sun, Kevin Choi, Balaji Veeramani, Alison Hu, Edward Bowen, Tyler Cody, Dawei Zhou 0003 |
ICML | 9 |
| 2024 | Toward Robust Generative AI Text Detection: Generalizable Neural ModelabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in generating text that closely resembles human writing across wide range of styles and genres. However, such capabilities are prone to potential misuse, such as fake news generation, spam email creation, and misuse in academic assignments. Hence, it is essential to build automated approaches capable of distinguishing between Artificial Intelligence-generated text and human-authored text. In this paper, we proposed a fine-tuning based neural model which is a combination of transformer models, linguistic features and state-of-the-art embedding models. We have also curated a training dataset encompassing diverse samples from different LLMs and domains to fine-tune pretrained language models. We evaluated our model's performance against state-of-the-art methods, and the comparative analysis demonstrates that our approach consistently outperforms other methods across various datasets using established evaluation metrics. Harika Abburi, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya |
ICMLA | 4 |
| 2024 | Pre-train. Mixup and Fine-tune: A Simple Strategy to Handle Domain ShiftabstractTransfer learning leverages models trained on large source datasets to target domains with limited datasets by fine- tuning pre-trained models. These approaches work well with minimal distribution shifts. However, the nature of distribution shift is unknown in real-world applications which may lead to worse performance of models in the field. Domain adaptation approaches handles this issue explicitly by adapting source trained models using discrepancy, adversarial or reconstruction based approaches, but doesn't tackle the limited target data issue. A data augmentation approach such as mixup helps train resilient models with limited datasets, and is recently being considered for its ability to handle distribution shifts. In this work, we investigate how mixup can be used along with transfer learning to improve model performance on target domains with distribution shifts. Our experimental results shows mixup is complementary to transfer learning which we demonstrate by varying the percentage of available target training data with publicly available source and target datasets. Our proposed approach of using mixup along with fine tuning shows improved performance than just fine tuning or mixup across varying percentage of target dataset sizes. Haider Ilyas, Harika Abburi, Edward Bowen, Balaji Veeramani |
ICMLA | 3 |
| 2024 | Domain Contextual and Relational Graph Model for Predictive MaintenanceabstractThe rotating machine is one of the common components which is examined under predictive maintenance across different industry sectors. Failure of bearings and gear boxes in such rotating machines are typical problems and their inspection is done using multiple sensors such as vibration, ultrasound, torque and temperature. The existing machine learning (ML) methods for bearing fault detection include traditional ML approaches and advanced deep learning algorithms. Nevertheless, these approaches often fail to account for the domain context and relationships between features, which are essential for building generalized, knowledge oriented, and trustworthy models. This work attempts to incorporate these aspects using a domain contextual and relational graph model (DCRG). It involves a graph convolutional network which is constructed with domain inspired features and their relations based on domain knowledge. The proposed method has been inspected using open-source bearing datasets such as Franche-Comté Électronique Mécanique Thermique et Optique (FEMTO), and Xi'an Jiaotong University (XJTU) bearing dataset. DCRG achieves stronger F1 scores (improvement by 5–10%) than the baseline machine learning approaches. Robert Schiller, Trupti Chavan, Akshay Kakkar, Viraj E, Don Williams, Derek Snaidauf, Edward Bowen, Deepak Mittal, Sunil Reddy Tiyyagura |
ICMLA | 7 |
| 2023 | Embedding Representations of Diagnosis Codes for Outlier Payment DetectionabstractModels for detecting payment outliers from healthcare claims often rely on sparse, high-dimensional feature vector encodings of diagnosis codes. These encodings tend to lose inherent relationships between the diagnosis codes, and lead to more complex and less efficient models. In this paper, we propose a novel approach that leverages word and graph embeddings to represent diagnosis codes, particularly when predicting healthcare claim payment amounts within an outlier detection model. Word embeddings are generated using BioSentVec, utilizing medical descriptions of diagnosis codes extracted from the Unified Medical Language System (UMLS) database. Graph embeddings are created using node2vec applied to a claims graph network that connects claims information with the diagnosis code hierarchy. On a dataset of 36 million claims, the graph embeddings outperformed other feature representations, improving$\mathrm{R}^{2}$by over 99% compared to sparse encodings. Embedding representations produce significantly smaller dense vectors that encapsulate more information than large, sparse multi-hot encoded diagnosis code vectors. These dense embedded vectors provide meaningful representations of diagnosis codes, significantly improving payment prediction and outlier detection capabilities. Anvesh Matta, Michael Suesserman, David McNamee, Daniel Lasaga, Dan Olson, Edward Bowen, Sanmitra Bhattacharya |
ICMLA | 6 |
| 2023 | How far is too far? Identifying suspicious travel patterns in healthcare claims using machine learningabstractFraud in healthcare services and claims poses a significant threat to healthcare expenditure, accessibility to health services, and quality of care of members. One important type of member-provider collusion is where members travel unreason-able distances seeking healthcare services. Such activities could be indicators of “pill mills”, doctor shopping, or referral kickback schemes. Previous research on the identification of suspicious travel distances have focused mostly on the billed amount and considered select diagnosis conditions and travel distances at zip code or county levels. Compared to these studies, our proposed framework focuses on claims across various diagnoses and takes into account population densities of members' zip codes, and provider densities for various specialties, among other features, which are critical to the prediction of travel distances. We exper-iment with two approaches - i) a regression model paired with a statistical anomalous distance detector, and ii) a neural network-based model paired with a likelihood estimator for anomalous distance detection. The evaluation of these models on a manually annotated dataset shows that the second approach outperforms the first one in identifying anomalous travel distances. Daniel Lasaga, John Helms, Edward Bowen, Sanmitra Bhattacharya |
ICMLA | 4 |
| 2023 | An Ensemble-Based Approach for Generative Language Model Attribution
Harika Abburi, Michael Suesserman, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya |
WISE | 5 |
| 2022 | Exposing Surveillance Detection Routes via Reinforcement Learning, Attack Graphs, and Cyber TerrainabstractReinforcement learning (RL) operating on attack graphs leveraging cyber terrain principles are used to develop reward and state associated with determination of surveillance detection routes (SDR). This work extends previous efforts on developing RL methods for path analysis within enterprise networks. This work focuses on building SDR where the routes focus on exploring the network services while trying to evade risk. RL is utilized to support the development of these routes by building a reward mechanism that would help in realization of these paths. The RL algorithm is modified to have a novel warm-up phase which decides in the initial exploration which areas of the network are safe to explore based on the rewards and penalty scale factor. Lanxiao Huang, Tyler Cody, Christopher Redino, Abdul Rahman, Akshay Kakkar, Deepak Kushwaha, Cheng Wang 0040, Ryan Clark, Daniel Radke, Peter A. Beling, Edward Bowen |
ICMLA | 11 |
| 2022 | Zero Day Threat Detection Using Metric Learning AutoencodersabstractThe proliferation of zero-day threats (ZDTs) to companies’ networks has been immensely costly and requires novel methods to scan traffic for malicious behavior at massive scale. The diverse nature of normal behavior along with the huge landscape of attack types makes deep learning methods an attractive option for their ability to capture highly-nonlinear behavior patterns. In this paper, the authors demonstrate an improvement upon a previously introduced methodology, which used a dual-autoencoder approach to identify ZDTs in network flow telemetry. In addition to the previously-introduced asset-level graph features, which help abstractly represent the role of a host in its network, this new model uses metric learning to train the second autoencoder on labeled attack data. This not only produces stronger performance, but it has the added advantage of improving the interpretability of the model by allowing for multiclass classification in the latent space. This can potentially save human threat hunters time when they investigate predicted ZDTs by showing them which known attack classes were nearby in the latent space. The models presented here are also trained and evaluated with two more datasets, and continue to show promising results even when generalizing to new network topologies. Dhruv Nandakumar, Robert Schiller, Christopher Redino, Kevin Choi, Abdul Rahman, Edward Bowen, Marc Vucovich, Joe Nehila, Matthew Weeks, Aaron Shaha |
ICMLA | 6 |