Karuna P. Joshi

dblp:161/7948 · also Karuna Pande Joshi · DBLP profile ↗
← Back
15ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0002-6354-1686ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 13 (1 first)Database Systems & Data Management · 1 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2025 LLM Based Knowledge Graph Approach to Automating Medical Device Regulatory Compliance
abstract
Advanced medical devices increasingly rely on AI-driven frameworks to automate compliance processes, ensuring safety and efficacy while reducing regulatory burdens. In the United States, software-based medical devices, including those utilizing AI/ML models, are regulated by the FDA's Center for Devices and Radiological Health (CDRH) under the Code of Federal Regulations (CFR) Title 21. These regulations are extensive, cross-referenced documents that require significant human effort to parse, leading to high compliance costs for manufacturers. We propose a novel, semantically rich framework that extracts regulatory knowledge from FDA documents and translates it into a machine-processable format. Our system encodes regulatory knowledge into an OWL/RDF-based knowledge graph and uses the Mistral 7B Instruct model to dynamically generate SPARQL queries, perform compliance reasoning, and produce structured reports. This enables automated device classification (Class I, II, or III) and real-time regulatory evaluation. Validated through real-world use cases, our framework significantly reduces manual review effort, enhances interpretability, and accelerates time-to-market. The proposed approach integrates AI reasoning and semantic technologies to achieve scalable, transparent, and automated regulatory compliance.
Subhankar Chattoraj, Karuna P. Joshi
IEEE Big Data2
2025 Security Compliance for Smart Manufacturing Using Knowledgegraph Based Digital Twin
Javed Tamboli, Karuna P. Joshi, Ommo Clark
IEEE Big Data2
2024 MedReg-KG: KnowledgeGraph for Streamlining Medical Device Regulatory Compliance
abstract
Healthcare providers are deploying a large number of AI-driven Medical devices to help monitor and medicate patients. For patients with chronic ailments, like diabetes or gastric diseases, usage of these devices becomes part of their daily lifestyle. These medical devices often capture personally identifiable information (PII) and hence are strictly regulated by the Food and Drug Administration (FDA) to ensure the safety and efficacy of the medical device. Medical device regulations are currently available as large textual documents, called Code of Federal Regulations (CFR) Title 21, that cross-reference other documents and so require substantial human effort and cost to parse and comprehend. We have developed a semantically rich framework MedReg-KG to extract the knowledge from the rules and policies for Medical devices and translate it into a machine-processable format that can be reasoned over. By applying Deontic Logic over the policies, we are able to identify the permissions and prohibitions in the regulation policies. This framework was developed using AI/Knowledge extraction techniques and Semantic Web technologies like OWL/RDF and SPARQL. This paper presents our Ontology/Knowledge graph and the Deontic rules integrated into the design. We include the results of our validation against the dataset of Gastroenterology Urology devices and demonstrate the efficiency gained by using our system.
Subhankar Chattoraj, Karuna P. Joshi
IEEE Big Data2
2023 IoT-Reg: A Comprehensive Knowledge Graph for Real-Time IoT Data Privacy Compliance
abstract
The proliferation of the Internet of Things (IoT) has led to an exponential increase in data generation, especially from wearable IoT devices. While this data influx offers unparalleled insights and connectivity, it also brings significant privacy and security challenges. Existing regulatory frameworks like the United States (US) National Institute of Standards and Technology Interagency or Internal Report (NISTIR) 8228, the US Health Insurance Portability and Accountability Act (HIPAA), and the European Union (EU) General Data Protection Regulation (GDPR) aim to address these challenges but often operate in isolation, making their compliance in the vast IoT ecosystem inconsistent. This paper presents the IoT-Reg ontology, a holistic semantic framework that amalgamates these regulations, offering a stratified approach based on the IoT data lifecycle stages and providing a comprehensive yet granular approach to IoT data handling practices. The IoT-Reg ontology aims to transform the IoT domain into a realm where regulatory controls are seamlessly integrated system components by emphasizing risk management, compliance, and the pivotal role of manufacturers’ privacy policies, ensuring consistent adherence, enhancing user trust, and promoting a privacy-centric IoT environment. We include the results of validating this framework against risk mitigation for Wearable IoT devices.
Kelvin Uzoma Echenim, Karuna P. Joshi
IEEE Big Data2
2021 Trusted Compliance Enforcement Framework for Sharing Health Big Data
abstract
COVID pandemic management via contact tracing and vaccine distribution has resulted in a large volume and high velocity of Health-related data being collected and exchanged among various healthcare providers, regulatory and government agencies, and people. This unprecedented sharing of sensitive health-related Big Data has raised technical challenges of ensuring robust data exchange while adhering to security and privacy regulations. We have developed a semantically rich and trusted Compliance Enforcement Framework for sharing large velocity Health datasets. This framework, built using Semantic Web technologies, defines a Trust Score for each participant in the data exchange process and includes ontologies combined with policy reasoners that ensure data access complies with health regulations, like Health Insurance Portability and Accountability Act (HIPAA). We have validated our framework by applying it to the Centers for Disease Control and Prevention (CDC) Contact Tracing Use case by exchanging over 1 million synthetic contact tracing records. This paper presents our framework in detail, along with the validation results against Contact Tracing data exchange. This framework can be used by all entities who need to exchange high velocity-sensitive data while ensuring real-time compliance with data regulations.
Dae-young Kim, Lavanya Elluri, Karuna P. Joshi
IEEE BigData3
2020 Measuring Semantic Similarity across EU GDPR Regulation and Cloud Privacy Policies
abstract
Data protection authorities formulate policies and rules which the service providers have to comply with to ensure security and privacy when they perform Big Data analytics using users Personally Identifiable Information (PII). The knowledge contained in the data regulations and organizational privacy policies are typically maintained as short unstructured text in HTML or PDF formats. Hence it is an open challenge to determine the specific regulation rules that are being addressed by a provider's privacy policies. We have developed a semantically rich framework, using techniques from Semantic Web and Natural Language Processing, to extract and compare the context of a short text in real-time. This framework allows automated incremental text comparison and identifying context from short text policy documents by determining the semantic similarity score and extracting semantically similar key terms. Additionally, we also created a knowledge graph to store the semantically similar comparison results while evaluating our framework across EU GDPR and privacy policies of 20 organizations complying with this regulation associated with various categories apply to Big Data stored in the cloud. Our approach can be utilized by Big Data practitioners to update their referential documents regularly based on the authority documents.
Lavanya Elluri, Karuna P. Joshi, Anantaa Kotal
IEEE BigData2
2020 Cloud-based Encrypted EHR System with Semantically Rich Access Control and Searchable Encryption
abstract
Cloud-based electronic health records (EHR) systems provide important security controls by encrypting patient data. However, these records cannot be queried without decrypting the entire record. This incurs a huge amount of burden in network bandwidth and the client-side computation. As the volume of cloud-based EHRs reaches Big Data levels, it is essential to search over these encrypted patient records without decrypting them to ensure that the medical caregivers can efficiently access the EHRs. This is especially critical if the caregivers have access to only certain sections of the patient EHR and should not decrypt the whole record. In this paper, we present our novel approach that facilitates searchable encryption of large EHR systems using Attribute-based Encryption (ABE) and multi-keyword search techniques. Our framework outsources key search features to the cloud side. This way, our system can perform keyword searches on encrypted data with significantly reduced costs of network bandwidth and client-side computation.
Redwan Walid, Karuna P. Joshi, Seung Geol Choi, Dae-young Kim
IEEE BigData2
2019 Scalability Analysis of Blockchain on a Serverless Cloud
abstract
While adopting Blockchain technologies to automate their enterprise functionality, organizations are recognizing the challenges of scalability and manual configuration that the state of art present. Scalability of Hyperledger Fabric is an open challenge recognized by the research community. We have automated many of the configuration steps of installing Hyperledger Fabric Blockchain on AWS infrastructure and have benchmarked the scalability of that system. We have used the UCR (University of California Riverside) Time Series Archive with 128 timeseries datasets containing over 191,177 rows of data totaling 76,453,742 numbers. Using an automated Serverless approach, we have loaded this dataset, by chunks, into different AWS instances, triggering the load by SQS messaging. In this paper, we present the results of this benchmarking study and describe the approach we took to automate the Hyperledger Fabric processes using serverless Lambda functions and SQS triggering. We will also discuss what is needed to make the Blockchain technology more robust and scalable.
Alex Kaplunovich, Karuna P. Joshi, Yelena Yesha
IEEE BigData2
2018 An Integrated Knowledge Graph to Automate GDPR and PCI DSS Compliance
abstract
Big data analytics related to consumer behavior, market analysis, opinions, and recommendation often deal with end user's derived and inferred data, along with the observed data. To ensure consumer data protection, rules defined by the European Union's General Data Protection Regulation (EU GDPR) must be adhered to by every organization using Personally Identifiable Information (PII) data for Big Data analysis. Similarly, Payment Card Industry Data Security Standard (PCI DSS) has policy guidelines specifically for organizations handling consumer's payment card data. Both data regulation policies are currently available only in textual format and require significant manual effort to ensure their compliance. We have developed an integrated, semantically rich Knowledge Graph (or Ontology) to represent the rules mandated by both PCI DSS and EU GDPR. In the Ontology, we have also identified the obligations defined in these regulations and related them with corresponding Cloud Security Alliance (CSA) controls. We have validated this Knowledge Graph against the data policies of major vendors that deal with Big Data. This Knowledge Graph that is available in the public domain can be used by Big Data practitioners to automate data protection compliance in their organization.
Lavanya Elluri, Ankur Nagar, Karuna P. Joshi
IEEE BigData3
2017 Link before you share: Managing privacy policies through blockchain
abstract
With the advent of numerous online content providers, utilities and applications, each with their own specific version of privacy policies and its associated overhead, it is becoming increasingly difficult for concerned users to manage and track the confidential information that they share with the providers. We have developed a novel framework to automatically track details about how a user's PII is stored, used and shared by the provider. We have integrated our data privacy ontology with the properties of blockchain, to develop an automated access-control and audit mechanism that enforces users' data privacy policies when sharing their data across third parties. We have also validated this framework by implementing a working system LinkShare. In this paper, we describe our framework on detail along with the LinkShare system. Our approach can be adopted by big data users to automatically apply their privacy policy on data operations and track the flow of that data across various stakeholders.
Agniva Banerjee, Karuna P. Joshi
IEEE BigData2
2017 Automated knowledge extraction from the federal acquisition regulations system (FARS)
abstract
With increasing regulation of Big Data, it is becoming essential for organizations to ensure compliance with various data protection standards. The Federal Acquisition Regulations System (FARS) within the Code of Federal Regulations (CFR) includes facts and rules for individuals and organizations seeking to do business with the US Federal government. Parsing and gathering knowledge from such lengthy regulation documents is currently done manually and is time and human intensive. Hence, developing a cognitive assistant for automated analysis of such legal documents has become a necessity. We have developed semantically rich approach to automate the analysis of legal documents and have implemented a system to capture various facts and rules contributing towards building an efficient legal knowledge base that contains details of the relationships between various legal elements, semantically similar terminologies, deontic expressions and cross-referenced legal facts and rules. In this paper, we describe our framework along with the results of automating knowledge extraction from the FARS document (Title 48, CFR). Our approach can be used by Big Data Users to automate knowledge extraction from Large Legal documents.
Srishty Saha, Karuna P. Joshi, Renee Frank, Michael Aebig, Jiayong Lin
IEEE BigData2
2016 Semantic approach to automating management of big data privacy policies
abstract
Ensuring privacy of Big Data managed on the cloud is critical to ensure consumer confidence. Cloud providers publish privacy policy documents outlining the steps they take to ensure data and consumer privacy. These documents are available as large text documents that require manual effort and time to track and manage. We have developed a semantically rich ontology to describe the privacy policy documents and built a database of several policy documents as instances of this ontology. We next extracted rules from these policy documents based on deontic logic which can be used to automate management of data privacy. In this paper we describe our ontology in detail along with the results of our analysis of privacy policies of prominent cloud services.
Karuna P. Joshi, Aditi Gupta 0003, Sudip Mittal, Claudia Pearce, Anupam Joshi, Tim Finin
IEEE BigData1
2015 Parallelizing natural language techniques for knowledge extraction from cloud service level agreements
abstract
To efficiently utilize their cloud based services, consumers have to continuously monitor and manage the Service Level Agreements (SLA) that define the service performance measures. Currently this is still a time and labor intensive process since the SLAs are primarily stored as text documents. We have significantly automated the process of extracting, managing and monitoring cloud SLAs using natural language processing techniques and Semantic Web technologies. In this paper we describe our prototype system that uses a Hadoop cluster to extract knowledge from unstructured legal text documents. For this prototype we have considered publicly available SLA/terms of service documents of various cloud providers. We use established natural language processing techniques in parallel to speed up cloud legal knowledge base creation. Our system considerably speeds up knowledge base creation and can also be used in other domains that have unstructured data.
Sudip Mittal, Karuna P. Joshi, Claudia Pearce, Anupam Joshi
IEEE BigData2
2011 DC Proposal: Automation of Service Lifecycle on the Cloud by Using Semantic Technologies
Karuna P. Joshi
ISWC (2)1
2003 On Using a Warehouse to Analyze Web Logs
Karuna P. Joshi, Anupam Joshi, Yelena Yesha
Distributed Parallel Databases1