Lalana Kagal

dblp:75/6949 · DBLP profile ↗
← Back
14ranked-venue papers in the field
1as first author
5since 2021 · last 2024
0000-0001-8469-1993ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 6 (1 first)Big Data, Cloud & Distributed Data Systems · 5Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2024 MediRAG: Secure Question Answering for Healthcare Data
abstract
Retrieval augmented generation (RAG) allows large language models to answer domain-specific questions by using external knowledge bases without training on this domain data or fine-tuning on its updates. This is especially promising for clinical tasks, as medical data tends to be dynamic, private, and distributed. By exposing the source documents that inform the response, RAG enables greater interpretability as well as reduced hallucination, both of which are crucial for safe deployment in healthcare. However, applying standard RAG to answer questions across patient data is complicated by strict access control to sensitive records enforced by privacy regulations such as HIPAA as well as the distributed nature of EHRs among institutions. We propose MediRAG, a clinical QA system that (i) supports a hierarchical design for federated document retrieval, and (ii) enables policy-based access control (PBAC) at retrieval time. For our experiments, we use the MIMIC-IV dataset, a publicly available EHR database that has been used in a wide array of research studies. We carry out a simulation of multiple federated hospitals and show that our scheme results in no loss of quality against a centralized baseline. We also evaluate performance with respect to key RAG metrics such as ROUGE, BLUE, context relevance, faithfulness and answer relevance and show that MediRAG is effective at clinical question answering across decentralized EHR documents while enforcing policies on sensitive data.
Emily Jiang, Alice Chen, Irene Tenison, Lalana Kagal
IEEE Big Data4
2024 Private Synthetic Data Generation for Mixed Type Datasets
abstract
In the face of escalating threats from privacy attacks on machine learning models, we propose a system that can artificially generate data that imitates real data but doesn’t contain any sensitive or personally identifiable information. The generated data, called synthetic data, will have the same semantic and statistical distribution as the original dataset but provide privacy guarantees. Compared to previous works that dealt with either structured or unstructured data separately, our work develops a complete hybrid pipeline for generating private synthetic datasets from complex datasets that consist of both structured (numerical or categorical) and unstructured data. The private synthetic data generated can be analyzed by collaborators and third parties without increasing the risks of leakage of sensitive data. We evaluate our system on Yelp reviews and drug side-effects datasets and calculate metrics for both quality and privacy. We introduce a context-aware exposure metric to quantify context-dependent memorization and use it along with exposure to evaluate privacy. Our evaluations demonstrate that our system generates meaningful private synthetic datasets that achieve good performance in characteristic similarity, utility, as well as privacy. Given these results, the generated synthetic data can be used by data scientists, researchers, and developers to address challenges related to data privacy, scarcity, diversity, and model testing in a wide range of applications including healthcare, insurance, and financial systems that rely on sensitive data.
Irene Tenison, Ashley Chen, Navpreet Singh, Omar Dahleh, Eliott Zemour, Lalana Kagal
IEEE Big Data6
2022 Overcoming Challenges of Synthetic Data Generation
abstract
There are several shortcomings in current methods of generating synthetic data using Generative Adversarial Networks (GANs). First, they tend to only emulate certain attributes of the original dataset. Second, they do not effectively model unbalanced discrete columns, long tails, or bimodal distributions of continuous columns. Lastly, these approaches often do not consider the potential for information leakage from the generated data. We propose UniformGAN, a GAN with a novel uniform loss function, which addresses these challenges and provides strong privacy guarantees using differential privacy. UniformGAN pre-processes datasets to transpose each column into a uniform distribution. We use a modified Deep Convolutional Generative Adversarial Network (DCGAN) architecture in which we replace ReLU activation functions with the more robust SeLU, which has significantly better performance and better convergence properties, and apply Dense-Sparse-Dense training to our network. We also use differential privacy to add noise to the discriminator during training. Along with UniformGAN, we provide a configurable command-line tool to generate and evaluate synthetic datasets on numerous metrics. It allows users to generate synthetic datasets from CTGAN, TableGAN, UniformGAN, or a custom framework and analyze the resultant datasets. This tool will help data scientists and industry users compare different synthetic dataset generation models and enable them to improve existing methods. We evaluated UniformGAN using multiple datasets, including the Adult, Covertype, and Credit Kaggle datasets, as well as two insurance-related Kaggle datasets. The results show that, when used on datatsets containing a large number of continuous columns, UniformGAN out performs other methods by producing synthetic data with similar correlations and distributions as the original dataset while ensuring privacy.
Kevin Fang, Vaikkunth Mugunthan, Vayd Ramkumar, Lalana Kagal
IEEE Big Data4
2021 Explaining Multimodal Errors in Autonomous Vehicles
abstract
Complex machines, such as autonomous vehicles, are unable to reconcile conflicting behaviors between their underlying subsystems, which leads to accidents and other negative consequences. Existing approaches to error and anomaly detection are not equipped to detect and mitigate inconsistencies among parts. In this paper, we present “Anomaly Detection through Explanations” or ADE, a multimodal monitoring architecture to reconcile critical discrepancies under uncertainty. ADE uses symbolic explanations as a debugging language, by examining underlying reasons for those decisions. Further, when decisions conflict, our method uses a synthesizer, along with a priority hierarchy, to process subsystem outputs along with their underlying reasons and transparently judges the conflicts. We show the accuracy and performance of ADE on autonomous vehicle scenarios and data, and discuss other error evaluations for future work.
Leilani H. Gilpin, Vishnu Penubarthi, Lalana Kagal
DSAA3
2021 The Punya Platform: Building Mobile Research Apps with Linked Data and Semantic Features
Evan W. Patton, William Van Woensel, Oshani Seneviratne, Giuseppe Loseto, Floriano Scioscia, Lalana Kagal
ISWC6
2020 PrivacyFL: A Simulator for Privacy-Preserving and Secure Federated Learning
abstract
Federated learning is a technique that enables distributed clients to collaboratively learn a shared machine learning model without sharing their training data. This reduces data privacy risks, however, privacy concerns still exist since it is possible to leak information about the training dataset from the trained model's weights or parameters. Therefore, it is important to develop federated learning algorithms that train highly accurate models in a privacy-preserving manner. Setting up a federated learning environment, especially with security and privacy guarantees, is a time-consuming process with numerous configurations and parameters that can be manipulated. In order to help clients ensure that collaboration is feasible and to check that it improves their model accuracy, a real-world simulator for privacy-preserving and secure federated learning is required.
Vaikkunth Mugunthan, Anton Peraire-Bueno, Lalana Kagal
CIKM3
2018 Explaining Explanations: An Overview of Interpretability of Machine Learning
abstract
There has recently been a surge of work in explanatory artificial intelligence (XAI). This research area tackles the important problem that complex machines and algorithms often cannot provide insights into their behavior and thought processes. XAI allows users and parts of the internal system to be more transparent, providing explanations of their decisions in some level of detail. These explanations are important to ensure algorithmic fairness, identify potential bias/problems in the training data, and to ensure that the algorithms perform as expected. However, explanations produced by these systems is neither standardized nor systematically assessed. In an effort to create best practices and identify open challenges, we describe foundational concepts of explainability and show how they can be used to classify existing literature. We discuss why current approaches to explanatory methods especially for deep neural networks are insufficient. Finally, based on our survey, we conclude with suggested future research directions for explanatory artificial intelligence.
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael A. Specter, Lalana Kagal
DSAA6
2017 Shade: A differentially-private wrapper for enterprise big data
abstract
Enterprises usually provide strong controls to prevent cyberattacks and inadvertent leakage of data to external entities. However, in the case where employees and data scientists have legitimate access to analyze and derive insights from the data, there are insufficient controls and employees are usually permitted access to all information about the customers of the enterprise including sensitive and private information. Though it is important to be able to identify useful patterns of one's customers for better customization and service, customers' privacy must not be sacrificed to do so. We propose an alternative - a framework that will allow privacy preserving data analytics over big data. In this paper, we present an efficient and scalable framework for Apache Spark, a cluster computing framework, that provides strong privacy guarantees for users even in the presence of an informed adversary, while still providing high utility for analysts. The framework, titled Shade, includes two mechanisms - SparkLAP, which provides Laplacian perturbation based on a user's query and SparkSAM, which uses the contents of the database itself in order to calculate the perturbation. We show that the performance of Shade is substantially better than earlier differential privacy systems without loss of accuracy, particularly when run on datasets small enough to fit in memory, and find that SparkSAM can even exceed performance of an identical nonprivate Spark query.
Alexander Heifetz, Vaikkunth Mugunthan, Lalana Kagal
IEEE BigData3
2015 AccountableMR: Toward accountable MapReduce systems
abstract
Traditional security techniques (e.g., authorization and encryption) have been extensively used in data management systems to provide security and privacy for many years. However, recent security breaches (e.g., WikiLeaks) showed that even if perfect access control is achieved, malicious insiders can still infer sensitive information and can misuse this sensitive information. To address this issue, accountability is introduced to deter inappropriate use of data through provision of usage control, privacy-aware interfaces, and careful monitoring and auditing. In this paper, we propose an accountable MapReduce architecture, where specific data usage is allowed after fine-grained transparent authorizations (i.e., individual record level), and such data usage are subject to effective accountability assessments by those who seek to assure privacy and security policy compliance. Our architecture enhances the MapReduce systems with the purpose concept (i.e., usage restrictions), authorize the users in fine-grained manner, and verifies the output of previously run jobs at post authorization time for detecting authorization and purpose breaches. Our empirical results show that in combination with traditional security features, AccountableMR can efficiently enhance the security and accountability of MapReduce model.
Huseyin Ulusoy, Murat Kantarcioglu, Erman Pattuk, Lalana Kagal
IEEE BigData4
2009 Policy-Aware Content Reuse on the Web
abstract
The Web allows users to share their work very effectively leading to the rapid re-use and remixing of content on the Web including text, images, and videos. Scientific research data, social networks, blogs, photo sharing sites and other such applications known collectively as the Social Web have lots of increasingly complex information. Such information from several Web pages can be very easily aggregated, mashed up and presented in other Web pages. Content generation of this nature inevitably leads to many copyright and license violations, motivating research into effective methods to detect and prevent such violations. This is supported by an experiment on Creative Commons (CC) attribution license violations from samples of Web pages that had at least one embedded Flickr image, which revealed that the attribution license violation rate of Flickr images on the Web is around 70-90%. Our primary objective is to enable users to do the right thing and comply with CC licenses associated with Web media, instead of preventing them from doing the wrong thing or detecting violations of these licenses. As a solution, we have implemented two applications: (1) Attribution License Violations Validator, which can be used to validate users’ derived work against attribution licenses of reused media and, (2) Semantic Clipboard, which provides license awareness of Web media and enables users to copy them along with the appropriate license metadata.
Oshani Seneviratne, Lalana Kagal, Tim Berners-Lee
ISWC2
2006 A Semantic Context-Aware Access Control Framework for Secure Collaborations in Pervasive Computing Environments
Alessandra Toninelli, Rebecca Montanari, Lalana Kagal, Ora Lassila
ISWC3
2003 Security for DAML Web Services: Annotation and Matchmaking
Grit Denker, Lalana Kagal, Tim Finin, Massimo Paolucci 0001, Katia P. Sycara
ISWC2
2003 A Policy Based Approach to Security for the Semantic Web
Lalana Kagal, Tim Finin, Anupam Joshi
ISWC1
2002 Intelligent Agents for Mobile and Embedded Devices
abstract
The pervasive computing environments of the near future will involve the interactions, coordination and cooperation of numerous, casually accessible, and often invisible computing devices. These devices, whether carried on our person or embedded in our homes, businesses and classrooms, will connect via wireless and wired links to one another and to the global networking infrastructure. The result will be a networking milieu with a new level of openness. The localized and dynamic nature of their interactions raises many new issues that draw on and challenge the disciplines of agents, distributed systems, and security. This paper describes recent work by the UMBC Ebiquity research group which addresses some of these issues.
Tim Finin, Anupam Joshi, Lalana Kagal, Olga Ratsimor, Sasikanth Avancha, Vlad Korolev, Harry Chen 0001, Filip Perich, R. Scott Cost
Int. J. Cooperative Inf. Syst.3