VLDB 2026 Research / reviewers in the wild / expert
Bhavani Thuraisingham
dblp:t/BMThuraisingham · also Bhavani M. Thuraisingham
· DBLP profile ↗
265ranked-venue papers
71as first author
34since 2021 · last 2026
0000-0003-4653-2080ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 110 · 29 first-author · 20 since 2021Databases, data management, data science and information retrieval · 66 · 16 first-author · 5 since 2021Artificial intelligence and machine learning · 41 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 33 · 13 first-author · 2 since 2021Software engineering, systems software and programming languages · 32 · 15 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-authorSystems, architecture and hardware · 4Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Theory of computation · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoE-T: Dependency Graph-Gated Mixture of Experts for Tabular Generation with Functional Dependency
Mary Grace Dhooghe, Murat Kantarcioglu, Bhavani Thuraisingham |
CODASPY | 3 |
| 2026 | An Automated Vulnerability Detection Framework for Smart ContractsabstractWith the increase of the adoption of blockchain technology in providing decentralized solutions to various problems, smart contracts have become more popular to the point that billions of US Dollars are currently exchanged every day through such technology. Meanwhile, various vulnerabilities in smart contracts have been exploited by attackers to steal cryptocurrencies worth millions of dollars. The automatic detection of smart contract vulnerabilities therefore is an essential research problem. Existing solutions to this problem particularly rely on human experts to define features or different rules to detect vulnerabilities. However, this often causes many vulnerabilities to be ignored, and they are inefficient in detecting new vulnerabilities. In this study, to overcome such challenges, we propose a framework to automatically detect vulnerabilities in smart contracts on the blockchain. More specifically, first, we utilize novel feature vector generation techniques from bytecode of smart contract as source code is rarely publicly available. These feature vectors are then analyzed using our innovative metric learning-based Deep Neural Networks (DNNs) to produce detection results. The framework’s predictions are further refined through a voting mechanism to achieve consensus. We conduct comprehensive experiments on large-scale benchmarks, and the quantitative results demonstrate the effectiveness and efficiency of our approach. Feng Mi, Chen Zhao 0010, Zhuoyi Wang, Sadaf Md. Halim, Xiaodi Li 0002, Zhouxiang Wu, Latifur Khan, Bhavani Thuraisingham |
Distributed Ledger Technol. Res. Pract. | 8 |
| 2025 | Mentoring for Career Success in Computer Science for All
Bhavani Thuraisingham |
MDM | 1 |
| 2025 | MUBox: A Critical Evaluation Framework of Deep Machine Unlearning [Systematization of Knowledge Paper]abstractRecent legal frameworks have mandated the right to be forgotten, obligating the removal of specific data upon user requests. Machine Unlearning has emerged as a promising solution by selectively removing learned information from machine learning models. This paper presents MUBox, a comprehensive platform designed to evaluate unlearning methods in deep learning. MUBox integrates 23 advanced unlearning techniques, tested across six practical scenarios with 11 diverse evaluation metrics. It allows researchers and practitioners to (1) assess and compare the effectiveness of different machine unlearning methods across various scenarios; (2) examine the impact of current evaluation metrics on unlearning performance; and (3) conduct detailed comparative studies on machine unlearning in a unified framework. Leveraging MUBox, we systematically evaluate these unlearning methods in deep learning and uncover a set of key insights: (a) Even state-of-the-art unlearning methods, including those published in top-tier venues and winners of unlearning competitions, demonstrate inconsistent effectiveness across diverse scenarios. Prior research has predominantly focused on simplified settings, such as random forgetting and class-wise unlearning, highlighting the need for broader evaluations across more complex and realistic unlearning tasks. (b) Assessing unlearning performance remains a non-trivial problem, as no single evaluation metric can comprehensively capture the effectiveness, efficiency, and preservation of model utility. Our findings emphasize the necessity of employing multiple metrics to achieve a balanced and holistic assessment of unlearning methods. (c) In the context of depoisoning-removing the adverse effects of poisoned data-our evaluation reveals significant variability in the effectiveness of existing approaches, which is highly dependent on the specific type of poisoning attack. We believe MUBox will serve as a valuable benchmark, advancing research in machine unlearning and highlighting areas for future improvement. Codes are available at https://github.com/Jessegator/MUBox. Xiang Li 0176, Wenqi Wei 0001, Bhavani Thuraisingham |
SACMAT | 3 |
| 2025 | Category-Based Administrative Access Control PoliciesabstractAs systems evolve, security administrators need to review and update access control policies. Such updates must be carefully controlled due to the risks associated with erroneous or malicious policy changes. We propose a category-based access control (CBAC) model, called Admin-CBAC , to control administrative actions. Since most of the access control models in use nowadays (including the popular RBAC and ABAC models) are instances of CBAC, from Admin-CBAC , we derive administrative models for RBAC and ABAC, too. We present a graph-based representation of Admin-CBAC policies and a formal operational semantics for administrative actions via graph rewriting. We also discuss implementations of Admin-CBAC exploiting the graph-based representation. Using the formal semantics, we show how properties (such as safety, liveness, and effectiveness of policies) and constraints (such as separation of duties) can be checked, and discuss the impact of policy changes. Although the most interesting properties of policies are generally undecidable in dynamic access control models, we identify particular cases where reachability properties are decidable and can be checked using our operational semantics, generalising previous results for RBAC and ABAC α . Clara Bertolissi, Maribel Fernández, Bhavani Thuraisingham |
ACM Trans. Priv. Secur. | 3 |
| 2024 | An Axiomatic Category-Based Access Control Model for Smart Homes
Clara Bertolissi, Maribel Fernández, Bhavani Thuraisingham |
LOPSTR | 3 |
| 2024 | Utilizing Threat Partitioning for More Practical Network Anomaly DetectionabstractAnomaly-based network intrusion detection would appear on the surface to be ideal for detection of zero-day network threats. Yet in practice, their often unacceptably high false positive rates keep them on the sideline in favor of signature-based methods, which typically detect known threats. We argue that an anomaly-based network intrusion detection system should not only be specialized to a specific class of related threats, but characteristics of the threat class itself should be utilized when designing both the detection system and structuring the network data to use with the system. To this end, we take two common network threat classes, DDoS-as-a-Smokescreen (DaaSS) and SYN flood, and analyze their characteristics for structure that we can use to specialize anomaly detection. We partition these threat classes into known behavior and unknown behavior, leaving the latter open-ended. Through experimentation on multiple datasets, we show that our proposed detection system based on this threat partitioning approach is capable of detecting DaaSS attacks and zero-day SYN flood variants with very low false positive rates, even in the face of concept drift, and can do so without having to collect large amounts of benign network traffic for training. Brian Ricks, Patrick Tague, Bhavani Thuraisingham, Sriraam Natarajan |
SACMAT | 3 |
| 2024 | Trustworthy Artificial Intelligence for Securing Transportation SystemsabstractArtificial Intelligence (AI) techniques are being applied to numerous applications from Healthcare to Cyber Security to Finance. For example, Machine Learning (ML) algorithms are being applied to solve security problems such as malware analysis and insider threat detection. However, there are many challenges in applying ML algorithms for various applications. For example, (i) the ML algorithms may violate the privacy of individuals. This is because we can gather massive amounts of data and apply ML algorithms to the data to extract highly sensitive information. (ii) ML algorithms may show bias and be unfair to various segments of the population. (iii) ML algorithms themselves may be attacked possibly resulting in catastrophic errors including in cyber-physical systems such as transportation systems. Finally, (iv) the ML algorithms must be safe and not harm society. Therefore, when ML algorithms are applied to transportation systems for handling congestion, preventing accidents, and giving advice to drivers, we must ensure that they are secure, ensure privacy and fairness, as well as provide for the safe operation of the transportation systems. Other AY techniques such as Generative AI (GenAI) are also being applied not only to secure systems design but also to determine the attacks and potential solutions. This presentation is divided into two parts. First, we describe our research over the past decade on Trustworthy ML systems. These are systems that are secure as well as ensure privacy, fairness, and safety. We discuss our ensemble-based ML models for detecting attacks as well as our research on developing Adversarial Machine Learning techniques. We also discuss securing the Internet of Transportation systems that are based on traditional methods such as Extended Kalman Filters to detect cyberattacks. Second, Second, we discuss our work on Finally, we discuss the research we recently started as part of the USDOT National University Technology Center TraCR (Transportation Cybersecurity and Resiliency) led by Clemson University. In particular, we describe (i) the application of federated machine learning techniques for detecting attacks in transportation systems; (ii) publishing synthetic transportation data sets that preserve privacy, (iii) fairness algorithms for transportation systems, and (iv) examining how GenAI systems are being integrated with transportation systems to provide security. Our focus includes the following: · Data Privacy: We are designing a Privacy-aware Policy-based Data Management Framework for Transportation Systems. Our work involves collecting the requisite data and developing analysis tools to identify and quantify privacy risks. Existing privacy-preserving, differentially private synthetic data generation techniques, which tailor data utility for generic ML accuracy, are not well suited for specific applications. We are developing synthetic data generation tools for transportation systems applications. We will develop new ML algorithms that can leverage these datasets. · Fairness: We have developed a novel adaptive fairness-aware online meta-learning algorithm, FairSAOML, which adapts to changing environments in both bias control and model precision. Our current work is focusing on adapting our framework to fairness in transportation systems. and control bias over time, especially ensuring group fairness across different protected sub-populations; identifying interesting attributes using explainable AI techniques that might help to mitigate bias and develop equitable algorithms. We have also developed a second system, FairDolce, that recognizes objects involving fairness constraints in a changing environment. We are adapting it to transportation applications. For example, pedestrian detection (whether or not the object being seen is a pedestrian) must be fair with respect to the race or gender of the individuals being detected under changing environments (e.g., rainy, cloudy sunny). Adversarial ML: Our prior work on adversarial ML models worked on traditional datasets such as network traffic data. Our current focus is on adapting our approach to AV-based sensor data. Our ML models are being applied to sensor data for object recognition and traffic management. These ML models may be attacked by the adversary. We will study various attack models and investigate ways of how interactions may occur between the model and the adversary and subsequently develop appropriate adversarial ML models that operate on the AV sensor data. · Attack Detection - Smart vehicles are often exposed to various attacks making it difficult for manufacturers to collaboratively train anomaly/attack detection models. Yet it would be ideal if all the data available across manufacturers could be used in building robust attack detection systems. To achieve this, we developed FAST-SV, which incorporates federated learning in conjunction with augmentation techniques to build a highly performant attack detection system for smart cars. Safety: Safety has been studied for cyber-physical systems and formal methods have been applied to specify safety properties and subsequently verify that the system satisfies the specifications. However, our goal is to ensure that the ML algorithms utilized by the transportation systems are safe. This would involve developing an AI Governance framework that would require transparency and explainability (among others) of the ML algorithms utilized by the transportation system. Bhavani Thuraisingham |
SACMAT | 1 |
| 2024 | Heterogeneous Domain Adaptation for Multistream Classification on Cyber Threat DataabstractUnder a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true, since data streams from different sources may contain distinct features. Indeed, they may even have different numbers of features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore if classes of labeled instances from other related streams are helpful to predict the classes of unlabeled instances in a different stream. Note that domains of source and target streams may have distinct feature spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream. We propose a framework of multistream classification by using projected data from a common latent feature space, which is embedded from both source and target domains. This framework is also crucial for enterprise system defenders to detect cross-platform attacks, such as Advanced Persistent Threats (APTs). Empirical valuation and analysis on both real-world and synthetic datasets are performed to validate the effectiveness of our proposed algorithm, comparing to state-of-the-art techniques. Experimental results show that our approach significantly outperforms other existing approaches. Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Latifur Khan, Anoop Singhal, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2023 | Attack Some while Protecting Others: Selective Attack Strategies for Attacking and Protecting Multiple ConceptsabstractMachine learning models are vulnerable to adversarial attacks. Existing research focuses on attack-only scenarios. In practice, one dataset may be used for learning different concepts, and the attacker may be incentivized to attack some concepts but protect the others. For example, the attacker might tamper a profile image for the "age'' model to predict "young'', while the "attractiveness'' model still predicts "pretty''. In this work, we empirically demonstrate that attacking the classifier for one learning task may negatively impact classifiers learning other tasks on the same data. This raises an interesting research question: is it possible to attack one set of classifiers while protecting the others trained on the same data? Vibha Belavadi, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
CCS | 4 |
| 2023 | The Design of an Ontology for ATT&CK and its Application to CybersecurityabstractThe spread of attacks in computer networks and within systems can have severe consequences for both individuals and organizations. One approach to preventing the spread of attacks is to use ontological aid, which is the use of ontologies to provide a structured representation of knowledge about the attack and its components, especially the ones who often disguise themselves to remain undetected for a long time within the system. As soon as one particular stage of such an attack is detected, it is imperative to reduce the amount of spread so that no permanent damage can be done. For this, the security analyst must boil down to technical details from a behavioral perspective so that proper defensive initiatives can be taken. We propose an ontology that will aid security analysts to find out the list of vulnerabilities to be patched so that an ongoing attack campaign can be prevented from spreading even more. Khandakar Ashrafi Akbar, Sadaf Md. Halim, Anoop Singhal, Basel Abdeen, Latifur Khan, Bhavani Thuraisingham |
CODASPY | 6 |
| 2023 | WokeGPT: Improving Counterspeech Generation Against Online Hate Speech by Intelligently Augmenting Datasets Using a Novel MetricabstractWith hate speech spreading rapidly online, it is increasingly important to respond automatically. However, there are some critical limitations in developing systems which produce these responses, which are known as counterspeeches. First, datasets containing paired instances of a hate speech and its appropriate response are very small. There is an abundance of hate speech on the web and in structured datasets, but quality counterspeeches are rare. Thus, since data is scarce, there is a need for automated methods to intelligently increase the size of existing paired datasets. Another critical challenge is that existing Natural Language Generation (NLG) metrics are not suitable for evaluating such systems, because these metrics do not accurately reflect how a human interprets the relationship between a hate speech and its counterspeech. Lastly, language models trained on internet text often exhibit a large amount of bias, which is unsuitable for sensitive tasks such as counterspeech generation. To address these challenges, we first introduce a technique to intelligently augment a small paired dataset of hate speech and counterspeech to make it substantially larger and varied, through a pairing technique that appropriately matches unpaired instances of hate speech with synthetic and existing counterspeeches. Next, we identified a need for a metric that evaluates counterspeech in the same way humans do, and propose a novel metric called PD-Score that leverages an advanced debating system. We empirically show through a large survey, that existing NLG metrics correlate poorly to human assessment and that our alternative is much more tightly bound to human assessment. Lastly, we curated a large domain-specific text corpus called WokeCorpus which we use to pretrain the language model before finetuning it for producing counterspeeches. We show that this both debiases the language model and aids performance. Sadaf Md. Halim, Saquib Irtiza, Yibo Hu 0002, Latifur Khan, Bhavani Thuraisingham |
IJCNN | 5 |
| 2023 | SAFE-PASS: Stewardship, Advocacy, Fairness and Empowerment in Privacy, Accountability, Security, and Safety for Vulnerable GroupsabstractOur vision is to achieve societally responsible secure and trustworthy cyberspace that puts algorithmic and technological checks and balances on the indiscriminate sharing and analysis of data. We achieve this vision in a holistic manner by framing research directions with four major considerations: (i) Expanding knowledge and understanding of security and privacy perceptions and expectations in vulnerable groups, which significantly contribute to their unwillingness to share data, and use that knowledge to drive research in (a) mitigating missing/imbalanced data problems, (b) understanding and modeling security and privacy risks of data sharing, and (c) modeling utility of data sharing. (ii) Developing a risk-adaptive, policy model capable of capturing and articulating security and privacy expectations of users that are relevant in a particular context and develops associated technology to ensure provenance and accountability. (iii) Developing robust AI/ML algorithms that are transparent and explainable with respect to fairness and bias to reduce/eliminate discrimination, misuse, privacy violations, or other cyber-crimes. (iv) Developing models and techniques for a nuanced, contextually adaptive, and graded privacy paradigm that allows trade-offs between privacy and utility. Towards this, in this paper we present the SAFE-PASS framework to provide Stewardship, Advocacy, Fairness and Empowerment in Privacy, Accountability, Security, and Safety for Vulnerable Groups. Indrajit Ray, Bhavani Thuraisingham, Jaideep Vaidya, Sharad Mehrotra, Vijayalakshmi Atluri, Indrakshi Ray, Murat Kantarcioglu, Ramesh Raskar, Babak Salimi, Steven J. Simske, Nalini Venkatasubramanian, Vivek K. Singh 0001 |
SACMAT | 2 |
| 2023 | Con2Mix: A semi-supervised method for imbalanced tabular security dataabstractCon2Mix (Contrastive Double Mixup) is a new semi-supervised learning methodology that innovates a triplet mixup data augmentation approach for finding code vulnerabilities in imbalanced, tabular security data sets. Tabular data sets in cybersecurity domains are widely known to pose challenges for machine learning because of their heavily imbalanced data (e.g., a small number of labeled attack samples buried in a sea of mostly benign, unlabeled data). Semi-supervised learning leverages a small subset of labeled data and a large subset of unlabeled data to train a learning model. While semi-supervised methods have been well studied in image and language domains, in security domains they remain underutilized, especially on tabular security data sets which pose especially difficult contextual information loss and balance challenges for machine learning. Experiments applying Con2Mix to collected security data sets show promise for addressing these challenges, achieving state-of-the-art performance on two evaluated data sets compared with other methods. Xiaodi Li 0002, Latifur Khan, Mahmoud Zamani, Shamila Wickramasuriya, Kevin W. Hamlen, Bhavani Thuraisingham |
J. Comput. Secur. | 6 |
| 2023 | Advanced Persistent Threat Detection Using Data Provenance and Metric LearningabstractAdvanced persistent threats (APT) have increased in recent times as a result of the rise in interest by nation-states and sophisticated corporations to obtain high-profile information. Typically, APT attacks are more challenging to detect since they leverage zero-day attacks and common benign tools. Furthermore, these attack campaigns are often prolonged to evade detection. We leverage an approach that uses a provenance graph to obtain execution traces of host nodes in order to detect anomalous behavior. By using the provenance graph, we extract features that are then used to train an online adaptive metric learning. Online metric learning is a deep learning method that learns a function to minimize the separation between similar classes and maximizes the separation between dis- similar instances. We compare our approach with baseline models and we show our method outperforms the baseline models by increasing detection accuracy on average by 11.3% and increases True positive rate (TPR) on average by 18.3%. We also show that our method outperforms several state-of-the-art models performances in comprehensive attack datasets in both binary and multi-class settings. Khandakar Ashrafi Akbar, Yigong Wang, Gbadebo Ayoade, Yang Gao 0027, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham, Kangkook Jee |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | A Privacy-Preserving Architecture and Data-Sharing Model for Cloud-IoT ApplicationsabstractMany service providers offer their services in exchange for users’ private data. Despite new regulations created to protect users privacy, users are often given little choice over the way their data is collected and used. To address privacy concerns in cloud-IoT applications, we propose to use an architecture, called Data Bank, which gives users fine-grained control over their data. Data Bank uses a category-based data access (CBDA) model which covers the whole data life-cycle, from data collection from IoT devices to data sharing with services. We show how dynamic policies can be specified using a new attribute-based instance of CBDA, and describe the use of policy graphs to visualise and analyse policies. Maribel Fernández, Jenjira Jaimunk, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2023 | Blockchain-Enabled Service Optimizations in Supply Chain Digital TwinabstractDigital twin is considered an alternative for optimizing real-world performance within virtual context, which also applies to the optimization ofSupply Chain Management(SCM). Blockchain, which facilitates data secure storage and trusted tracking, is deemed to be a proper assistant technology for achieving digital twin implementation. In this work, we propose a blockchain-based digital twin solution to reengineer SCM system, which promotes the digitization and intelligence of SCM to fit in massive service volumes in complex-intercrossed industry system. A strong-weak consensus mode is developed to achieve energy and time savings. We also design intelligent switch-based algorithms to generate time-saving consensus plans under energy constraints. Finally, we set up multiple experiments to compare our algorithm with three baseline algorithms, includingEffective Iterative Greedy(EIG),Two Dimensional Genetic(TDG), andHigh-level Task Scheduling Dynamic Programming(HTSDP). Findings from evaluation demonstrate the potential of our proposed model. Specifically, our algorithm reduces time and energy consumption of EIG algorithm in consensus by 46.84% and 16.25%, respectively. Compared with TDG algorithm, consensus time and energy consumption of our algorithm are reduced by 50.05% and 48.46%. Our algorithm cuts down time spent of HTSDP algorithm in generating consensus plan by a factor of 9.88. Keke Gai, Yue Zhang 0011, Meikang Qiu, Bhavani Thuraisingham |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Security and Privacy for Emerging IoT and CPS DomainsabstractThe proliferation of IoT and CPS technologies demand novel conceptual, foundational and applied cybersecurity solutions. The dynamic behaviour of these distributed systems augmented with physical and computational constraints of smart devices, require cybersecurity approaches for timely prevention and detection of attacks. This panel aims to discuss open challenges and highlight future research directions for cybersecurity in IoT and CPS. Elisa Bertino, Ravi S. Sandhu, Bhavani Thuraisingham, Indrakshi Ray, Wenjia Li, Maanak Gupta, Sudip Mittal |
CODASPY | 3 |
| 2022 | Knowledge Mining in Cybersecurity: From Attack to Defense
Khandakar Ashrafi Akbar, Sadaf Md. Halim, Yibo Hu 0002, Anoop Singhal, Latifur Khan, Bhavani Thuraisingham |
DBSec | 6 |
| 2022 | MCoM: A Semi-Supervised Method for Imbalanced Tabular Security Data
Xiaodi Li 0002, Latifur Khan, Mahmoud Zamani, Shamila Wickramasuriya, Kevin W. Hamlen, Bhavani Thuraisingham |
DBSec | 6 |
| 2022 | Imbalanced Adversarial Training with ReweightingabstractAdversarial training has been empirically proven to be one of the most effective and reliable defense methods against adversarial attacks. However, the majority of existing studies are focused on balanced datasets, where each class has a similar amount of training examples. Research on adversarial training with imbalanced training datasets is rather limited. As the initial effort to investigate this problem, we reveal the facts that adversarially trained models present two distinguished behaviors from naturally trained models in imbalanced datasets: (1) Compared to natural training, adversarially trained models can suffer much worse performance on under-represented classes, when the training dataset is extremely imbalanced. (2) Traditional reweighting strategies which assign large weights to underrepresented classes will drastically hurt the model’s performance on well-represented classes. In this paper, to further understand our observations, we theoretically show that the poor data separability is one key reason causing this strong tension between under-represented and well-represented classes. Motivated by this finding, we propose the Separable Reweighted Adversarial Training (SRAT) framework to facilitate adversarial training under imbalanced scenarios, by learning more separable features for different classes. Extensive experiments on various datasets verify the effectiveness of the proposed framework. Wentao Wang 0006, Han Xu 0002, Yaxin Li 0001, Bhavani Thuraisingham, Jiliang Tang |
ICDM | 5 |
| 2022 | Personalized Hashtag Recommendation with User-level Meta-learningabstractPersonalized hashtag recommendation aims to au-tomatically recommend user-specific hashtags to annotate the posts (e.g., tweets). It is actually an unwieldy procedure because only a small number of posts contain hashtags, which can hinder the quality of recommendation results. Moreover, users always have their own preferences while choosing the hashtag to describe the post contents. Thus, solving label scarcity and user bias both have raised researchers' attention toward hashtag recommendation. The existing methods can be generally divided into two groups: The first group models the relationship between hashtags and posts taking only post contents into account, without considering the user preferences. The second group models the relationship between users and posts assuming that each user has adequate posts tapped with hashtags. In this paper, we propose a Meta-learning based Personalized Hashtag Recommendation (MetaTag) framework to address above mentioned challenges at once. Our method can leverage the prior experience learned from other users and quickly adapt to a new user. In this framework, model is trained in an episodic manner with user-specific historical posts to learn semantic embeddings that can distinguish different hashtags even they have not been seen before. The results of experiments on the data collected from Twitter demonstrated that our method outperforms state-of-the-art approaches by a large margin. Hemeng Tao, Latifur Khan, Bhavani Thuraisingham |
IJCNN | 3 |
| 2022 | GCI: A GPU-Based Transfer Learning Approach for Detecting Cheats of Computer GameabstractCheating in massive multiple online games (MMOGs) adversely affect the game’s popularity and reputation among its users. Therefore, game developers invest large amount of efforts to detect and prevent cheats that provide an unfair advantage to cheaters over other naive users during game play. Particularly, MMOG clients share data with the server during game play. Game developers leverage this data to detect cheating. However, detecting cheats is challenging mainly due to the limited client-side information, along with unknown and complex cheating techniques. In this article, we aim to leverage machine learning-based models to predict cheats over encrypted game traffic during game play. Concretely, network game traffic during game play from each player can be used to determine whether a cheat is employed. A major challenge in developing such a prediction model is the availability of sufficient training data, which is sparingly available in practice. Game traffic obtained from a few known players can be easily labeled. However, if such players are not a good representation of the population (i.e., other players), then a supervised model trained on labeled game traffic from these set of players may not generalize well for the population. Here, we propose a Graphics Processing Unit (GPU) based scalable transfer learning approach to overcome the constraints of limited labeled data. Our empirical evaluation on a popular MMOG demonstrates significant improvement in cheat prediction compared to other competing methods. Md Shihabul Islam, Swarup Chandra, Latifur Khan, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2022 | Explainable Artificial Intelligence for Cyber Threat Intelligence (XAI-CTI)abstractThe papers in this special section focus on explainable artificial intelligence for cyber threat intelligence. Despite concerted efforts from industry, academia, and government on improving cybersecurity capabilities, cyber-threats such as ransomware, fake news, advanced malware, and others, continue to exact a substantial toll on modern infrastructure and day-to-day societal operations. To help combat the ever-growing quantity and severity of cyber-threats, many organizations are adopting Cyber Threat Intelligence (CTI). At its core, CTI is a data-driven process that aims to identify emerging threats and key threat actors to help enable effective cybersecurity decision-making. Sagar Samtani, Hsinchun Chen, Murat Kantarcioglu, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | SACCOS: A Semi-Supervised Framework for Emerging Class Detection and Concept Drift Adaption Over Data StreamsabstractIn this paper, we address challenges of detecting instances from emerging classes over a non-stationary data stream during data classification. In particular, data instances from an entirely unknown class may appear in a data stream over time. Existing classification techniques utilize unsupervised clustering to identify emergence of such data instances. Unfortunately, they make strong assumptions which are typically invalid in practice; (i) Most instances associated with a class are closer to each other in feature space than instances associated with different classes, (ii) Covariates of data are normalized through an oracle to overcome the effect of a few data instances having large feature values, and (iii) Labels of instances from emerging classes are readily available soon after detection. To address the challenges that occur in practice when the above assumptions are weak, i.e., instances of each class are scattered and the true labels of novel class instances are sparsely available, we propose a practical semi-supervised emerging class detection framework. Particularly, we aim to identify similar data instances within local regions in feature space by incorporating a mutual graph clustering mechanism. We also perform online normalization along the data stream instead of assuming an oracle, and propose a classification technique that uses only a small amount of true labels for training and emerging class detection. Our empirical evaluation of this framework on real-world datasets demonstrates its superiority of classification performance compared to existing methods, while using significantly fewer labeled instances. Yang Gao 0027, Swarup Chandra, Yifan Li 0003, Latifur Khan, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2021 | Progressive One-shot Human ParsingabstractPrior human parsing models are limited to parsing humans into classes pre-defined in the training data, which is not flexible to generalize to unseen classes, e.g., new clothing in fashion analysis. In this paper, we propose a new problem named one-shot human parsing (OSHP) that requires to parse human into an open set of reference classes defined by any single reference example. During training, only base classes defined in the training set are exposed, which can overlap with part of reference classes. In this paper, we devise a novel Progressive One-shot Parsing network (POPNet) to address two critical challenges , i.e., testing bias and small sizes. POPNet consists of two collaborative metric learning modules named Attention Guidance Module and Nearest Centroid Module, which can learn representative prototypes for base classes and quickly transfer the ability to unseen classes during testing, thereby reducing testing bias. Moreover, POPNet adopts a progressive human parsing framework that can incorporate the learned knowledge of parent classes at the coarse granularity to help recognize the descendant classes at the fine granularity, thereby handling the small sizes issue. Experiments on the ATR-OS benchmark tailored for OSHP demonstrate POPNet outperforms other representative one-shot segmentation models by large margins and establishes a strong baseline. Source code can be found at https://github.com/Charleshhy/One-shot-Human-Parsing. Haoyu He 0001, Jing Zhang 0037, Bhavani Thuraisingham, Dacheng Tao |
AAAI | 3 |
| 2021 | DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data AugmentationabstractPublic resources and services (e.g., datasets, training platforms, pre-trained models) have been widely adopted to ease the development of Deep Learning-based applications. However, if the third-party providers are untrusted, they can inject poisoned samples into the datasets or embed backdoors in those models. Such an integrity breach can cause severe consequences, especially in safety- and security-critical applications. Various backdoor attack techniques have been proposed for higher effectiveness and stealthiness. Unfortunately, existing defense solutions are not practical to thwart those attacks in a comprehensive way. Han Qiu 0001, Yi Zeng 0005, Shangwei Guo, Tianwei Zhang 0004, Meikang Qiu, Bhavani Thuraisingham |
AsiaCCS | 6 |
| 2021 | Graph-Based Specification of Admin-CBAC PoliciesabstractWe present a graph-based language for the specification of administrative access control policies in Admin-CBAC, an administrative model for Category-Based Access Control. More precisely, we propose a multi-level graph representation of policies and a graph-rewriting semantics for administrative actions, from which properties (such as safety, liveness and effectiveness of policies) and constraints (such as separation of duties) can be checked using graph traversal algorithms and rewriting properties. Since Admin-CBAC is a generic model, the techniques are directly applicable to a variety of access control models. In particular, we illustrate our techniques for the RBAC and ABAC instances of Admin-CBAC. Clara Bertolissi, Maribel Fernández, Bhavani Thuraisingham |
CODASPY | 3 |
| 2021 | An Episodic Learning based Geolocation Detection Framework for Imbalanced DataabstractA social media user's geographical location is vital to many applications like local search and event detection. The scarcity of publicly available location information motivates researchers to predict user geolocation based on information such as tweet text and social interaction data. In this paper, we investigate and improve on the task of predicting a Twitter user's city-level location based on the content of the user's historical tweets. In order to train a reliable location classifier, previous studies on this topic have typically assumed that there are sufficient amount of users living in each cities. However, they simply ignore the fact that different demographic groups may participate in social media platforms, which results in a highly imbalanced data distribution. Being aware of this population imbalance issue, we propose an episodic learning based framework to extract a single representative for each class (location), so that classifiers can later be trained on a balanced class distribution. To examine the effectiveness of our method, we design experiments which involve two kinds of baselines, the state-of-the-art geolocation detection methods and the well-known approaches handling imbalanced data in classification. The results of experiments on the data collected from Twitter demonstrated the superiority of our method when compared with baselines. Hemeng Tao, Yang Gao 0027, Zhuoyi Wang, Latifur Khan, Bhavani Thuraisingham |
IJCNN | 5 |
| 2021 | Fairness-Aware Online Meta-learningabstractIn contrast to offline working fashions, two research paradigms are devised for online learning: (1) Online Meta-Learning (OML)[6, 20, 26] learns good priors over model parameters (or learning to learn) in a sequential setting where tasks are revealed one after another. Although it provides a sub-linear regret bound, such techniques completely ignore the importance of learning with fairness which is a significant hallmark of human intelligence. (2) Online Fairness-Aware Learning [1, 8, 21]. This setting captures many classification problems for which fairness is a concern. But it aims to attain zero-shot generalization without any task-specific adaptation. This therefore limits the capability of a model to adapt onto newly arrived data. To overcome such issues and bridge the gap, in this paper for the first time we proposed a novel online meta-learning algorithm, namely FFML, which is under the setting of unfairness prevention. The key part of FFML is to learn good priors of an online fair classification model's primal and dual parameters that are associated with the model's accuracy and fairness, respectively. The problem is formulated in the form of a bi-level convex-concave optimization. The theoretic analysis provides sub-linear upper bounds O(log T)for loss regret and O(√log T)violation of cumulative fairness constraints. Our experiments demonstrate the versatility of FFML by applying it to classification on three real-world datasets and show substantial improvements over the best prior work on the tradeoff between fairness and classification accuracy. Chen Zhao 0010, Feng Chen 0001, Bhavani Thuraisingham |
KDD | 3 |
| 2021 | Special Issue on Robustness and Efficiency in the Convergence of Artificial Intelligence and IoTabstractToday, the Internet of Things (IoT) is increasingly flourishing with establishing ubiquitous connections between smart devices and objects, and by 2020, there will be a total of 30 billion connected things reported by IDC. The unprecedented data explosion provides immense opportunities for valuable information mining. At the same time, it also floods the infrastructure with tremendous values it necessarily handles and proposes high challenges to traditional data storing or processing techniques. On the other hand, artificial intelligence (AI) has become a key component for many applications that profoundly change our lives. Machine learning, especially deep learning (DL) technologies, vastly improves traditional computer science and networking technologies. The convergence of AI and IoT enables data to be quickly explored and turned into significant decisions. For companies and enterprises, AI enhances the speed and accuracy of data processing for instant market strategies. Meikang Qiu, Bhavani Thuraisingham, Mahmoud Daneshmand, Huansheng Ning, Payam M. Barnaghi |
IEEE Internet Things J. | 2 |
| 2021 | BiMorphing: A Bi-Directional Bursting Defense against Website Fingerprinting AttacksabstractNetwork traffic analysis has been increasingly used in various applications to either protect or threaten people, information, and systems. Website fingerprinting is a passive traffic analysis attack which threatens web navigation privacy. It is a set of techniques used to discover patterns from a sequence of network packets generated while a user accesses different websites. Internet users (such as online activists or journalists) may wish to hide their identity and online activity to protect their privacy. Typically, an anonymity network is utilized for this purpose. These anonymity networks such as Tor (The Onion Router) provide layers of data encryption which poses a challenge to the traffic analysis techniques. Although various defenses have been proposed to counteract this passive attack, they have been penetrated by new attacks that proved the ineffectiveness and/or impracticality of such defenses. In this work, we introduce a novel defense algorithm to counteract the website fingerprinting attacks. The proposed defense obfuscates original website traffic patterns through the use of double sampling and mathematical optimization techniques to deform packet sequences and destroy traffic flow dependency characteristics used by attackers to identify websites. We evaluate our defense against state-of-the-art studies and show its effectiveness with minimal overhead and zero-delay transmission to the real traffic. Khaled Al-Naami, Amir El-Ghamry, Md Shihabul Islam, Latifur Khan, Bhavani Thuraisingham, Kevin W. Hamlen, Mohammed F. Alrahmawy, Magdi Zakria Rashad |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2021 | Deep Residual Learning-Based Enhanced JPEG Compression in the Internet of ThingsabstractWith the development of big data and network technology, there are more use cases, such as edge computing, that require more secure and efficient multimedia big data transmission. Data compression methods can help achieving many tasks like providing data integrity, protection, as well as efficient transmission. Classical multimedia big data compression relies on methods like the spatial-frequency transformation for compressing with loss. Recent approaches use deep learning to further explore the limit of the data compression methods in communication constrained use cases like the Internet of Things (IoT). In this article, we propose a novel method to significantly enhance the transformation-based compression standards like JPEG by transmitting much fewer data of one image at the sender's end. At the receiver's end, we propose a two-step method by combining the state-of-the-art signal processing based recovery method with a deep residual learning model to recover the original data. Therefore, in the IoT use cases, the sender like edge device can transmit only 60% data of the original JPEG image without any additional calculation steps but the image quality can still be recovered at the receiver's end like cloud servers with peak signal-to-noise ratio over 31 dB. Han Qiu 0001, Qinkai Zheng, Gérard Memmi, Meikang Qiu, Bhavani Thuraisingham |
IEEE Trans. Ind. Informatics | 6 |
| 2021 | GraphBoot: Quantifying Uncertainty in Node Feature Learning on Large NetworksabstractIn recent years, as online social networks continue to grow in size, estimating node features, such as sociodemographics, preferences and health status, in a scalable and reliable way has become a primary research direction in social network mining. Although many techniques have been developed for estimating various node features, quantifying uncertainty in such estimations has received little attention. Furthermore, most existing methods study networks parametrically, which limits insights about necessary quantity of queried data, reliable feature estimation, and estimator uncertainty. Uncertainty quantification is critical for answering key questions, such as, given a limited availability of social network data, how much data should be queried from the network?, and which node features can be learned reliably? More importantly, how can we evaluate uncertainty of our estimators? Uncertainty quantification is not equivalent to network sampling but constitutes a key complementary concept to sampling and the associated reliability analysis. To our knowledge, this paper is the first work that sheds light on uncertainty quantification and uncertainty propagation in social network feature mining. We propose a novel non-parametric bootstrap method for uncertainty analysis of node features in social network mining, derive its asymptotic properties, and demonstrate its effectiveness with extensive experiments. Furthermore, we develop a new metric based on dispersion of estimations, enabling analysts to assess how much more information is needed for increasing prediction reliability based on the estimated uncertainty. We demonstrate the effectiveness of our new uncertainty quantification methodology with extensive experiments on real life social networks, and a case study of mental health on Twitter. Cuneyt Gurcan Akcora, Yulia R. Gel, Murat Kantarcioglu, Vyacheslav Lyubchich, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Cloud GovernanceabstractCorporate governance has received a lot of prominence since the accounting scandals of the early 2000s. Corporate governance not only includes maintaining correct financial records, but also the information technology aspects of the corporation. These include cyber security governance, artificial intelligence governance and cloud governance. Most corporations are migrating their data, software and processes to the cloud and therefore it is critical that the cloud be governed by appropriate policies and procedures. We also need cloud governance frameworks. The roles and responsibilities of the people involved have to be clearly identified. In this paper we discuss some of the essential points of cloud governance. Bhavani Thuraisingham |
CLOUD | 1 |
| 2020 | Data Science, COVID-19 Pandemic, Privacy and Civil LibertiesabstractThe world has seen pandemics, terrorism, hurricanes and other natural and man-made disasters. Each time such an event occurs we discuss technologies that can solve the problem and their impact on our privacy and civil liberties. Such discussions occurred after the 9/11 terrorist attacks and is happening now during the COVID-19 pandemic, the worst human crisis we have faced in a century. This paper discusses the applications of data science to detect and possibly prevent such pandemics and its impact on our privacy and civil liberties. Bhavani Thuraisingham |
IEEE BigData | 1 |
| 2020 | Admin-CBAC: An Administration Model for Category-Based Access ControlabstractWe present Admin-CBAC, an administrative model for Category- Based Access Control (CBAC). Since most of the access control models in use nowadays are instances of CBAC, in particular the popular RBAC and ABAC models, from Admin-CBAC we derive administrative models for RBAC and ABAC too. We define Admin- CBAC using Barker's metamodel, and use its axiomatic semantics to derive properties of administrative policies. Using an abstract operational semantics for administrative actions, we show how properties (such as safety, liveness and effectiveness of policies) and constraints (such as separation of duties) can be checked, and discuss the impact of policy changes. Although the most interesting properties of policies are generally undecidable in dynamic access control models, we identify particular cases where reachability based properties are decidable and can be checked using our operational semantics, generalising previous results for RBAC and ABACalpha. Clara Bertolissi, Maribel Fernández, Bhavani Thuraisingham |
CODASPY | 3 |
| 2020 | Can AI be for Good in the Midst of Cyber Attacks and Privacy Violations?: A Position PaperabstractArtificial Intelligence (AI) is affecting every aspect of our lives from healthcare to finance to driving to managing the home. Sophisticated machine learning techniques with a focus on deep learning are being applied successfully to detect cancer, to make the best choices for investments, to determine the most suitable routes for driving as well as to efficiently manage the electricity in our homes. We expect AI to have even more influence as advances are made with technology as well as in learning, planning, reasoning and explainable systems. While these advances will greatly advance humanity, organizations such as the United Nations have embarked on initiatives such as "AI for Good" and we can expect to see more emphasis on applying AI for the good of humanity especially in developing countries. However, the question that needs to be answered is Can AI be for Good when when the AI techniques can be attacked and the AI techniques themselves can cause privacy violations? This position paper will provide an overview of this topic with protecting children and children's rights as an example. Bhavani Thuraisingham |
CODASPY | 1 |
| 2020 | A Data Access Model for Privacy-Preserving Cloud-IoT ArchitecturesabstractWe propose a novel data collection and data sharing model for cloud-IoT architectures with an emphasis on data privacy. This model has been implemented in Privasee, an open source platform for privacy-aware web-application development, which provides a plug-in module to support IoT application development. Privasee uses a cloud-IoT architecture called DataBank. We provide examples and discuss future extensions. Maribel Fernández, Alex Franch Tapia, Jenjira Jaimunk, Manuel Martinez Chamorro, Bhavani Thuraisingham |
SACMAT | 5 |
| 2019 | Specification and Analysis of ABAC Policies via the Category-based MetamodelabstractThe Attribute-Based Access Control (ABAC) model is one of the most powerful access control models in use. It subsumes popular models, such as the Role-Based Access Control (RBAC) model, and can also enforce dynamic policies where authorisations depend on values of user, resource or environment attributes. However, in its general form, ABAC does not lend itself well to some operations, such as review queries, and ABAC policies are in general more difficult to specify and analyse than simpler RBAC policies. In this paper we propose a formal specification of ABAC in the category-based metamodel of access control, which adds structure to ABAC policies, making them easier to design and understand. We provide an axiomatic and an operational semantics for ABAC policies, and show how to use them to analyse policies and evaluate review queries. Maribel Fernández, Ian Mackie, Bhavani Thuraisingham |
CODASPY | 3 |
| 2019 | ChainNet: Learning on Blockchain Graphs with Topological FeaturesabstractWith emergence of blockchain technologies and the associated cryptocurrencies, such as Bitcoin, understanding network dynamics behind Blockchain graphs has become a rapidly evolving research direction. Unlike other financial networks, such as stock and currency trading, blockchain based cryptocurrencies have the entire transaction graph accessible to the public (i.e., all transactions can be downloaded and analyzed). A natural question is then to ask whether dynamics of the transaction graph impacts price of the underlying cryptocurrency. We show that standard graph features such as degree distribution of the transaction graph may not be sufficient to capture network dynamics and its potential impact on fluctuations of Bitcoin price. In contrast, topological features computed from the blockchain graph using the tools of persistent homology, are found to exhibit higher utility for predicting Bitcoin price dynamics. Nazmiye Ceren Abay, Cuneyt Gurcan Akcora, Yulia R. Gel, Murat Kantarcioglu, Umar Islambekov, Yahui Tian, Bhavani Thuraisingham |
ICDM | 7 |
| 2019 | Privacy-Preserving Architecture for Cloud-IoT PlatformsabstractWe propose a cloud-IoT architecture, called Data Bank, aiming at protecting users' sensitive data by allowing them to control which kind of data is transmitted by their devices and providing supportive tools for agreement visualisation and privacy-utility trade-off. The architecture consists of several layers, from IoT objects in the lower layer to web and mobile applications in the top layer, with regulated communication mechanisms to transfer data from the lower level to data processing services in the top level. We illustrate our proposal with an example in a smart vehicle environment. Maribel Fernández, Jenjira Jaimunk, Bhavani Thuraisingham |
ICWS | 3 |
| 2019 | Cyber Security and Data Governance Roles and Responsibilities at the C-Level and the BoardabstractCorporate governance and the roles and responsibilities of the corporate officers and the board of directors have received an increasing interest since the Enron scandal of the early 2000s. This scandal resulted in enacting policies, laws and regulations such as the Sarbanes-Oxley and others. More recently, the near daily cyber security attacks on the infrastructures and data assets of corporations have resulted in cyber security professionals taking a serious look at the roles and responsibilities of the corporate officers and the board with respect to cyber security and data governance. This paper discusses the issues and challenges for cyber security governance with an emphasis on data governance and the potential roles and responsibilities of the corporate officers and the board of directors. Bhavani Thuraisingham |
ISI | 1 |
| 2019 | An OpenRBAC Semantic Model for Access Control in Vehicular NetworksabstractInter-vehicle communication has the potential to significantly improve driving safety, but also raises security concerns. The fundamental mechanism to govern information sharing behaviors is access control. Since vehicular networks have a highly dynamic and open nature, access control becomes very challenging. Existing works are not applicable to the vehicular world. In this paper, we develop a new access control model, openRBAC, and the corresponding mechanisms for access control in vehicular systems. Our approach lets the accessee define a relative role hierarchy, specifying all potential accessor roles in terms of their relative perception to the accessees. Access control policies are defined for the relative roles in the hierarchy. Since the accessee has a clear understanding of the relative roles defined by itself, the policy definitions can be precise and less flawed. Sultan Alsarra, I-Ling Yen, Yongtao Huang, Farokh B. Bastani, Bhavani Thuraisingham |
SACMAT | 5 |
| 2019 | Towards Self-Adaptive Metric Learning On the FlyabstractGood quality similarity metrics can significantly facilitate the performance of many large-scale, real-world applications. Existing studies have proposed various solutions to learn a Mahalanobis or bilinear metric in an online fashion by either restricting distances between similar (dissimilar) pairs to be smaller (larger) than a given lower (upper) bound or requiring similar instances to be separated from dissimilar instances with a given margin. However, these linear metrics learned by leveraging fixed bounds or margins may not perform well in real-world applications, especially when data distributions are complex. We aim to address the open challenge of “Online Adaptive Metric Learning” (OAML) for learning adaptive metric functions on-the-fly. Unlike traditional online metric learning methods, OAML is significantly more challenging since the learned metric could be non-linear and the model has to be self-adaptive as more instances are observed. In this paper, we present a new online metric learning framework that attempts to tackle the challenge by learning a ANN-based metric with adaptive model complexity from a stream of constraints. In particular, we propose a novel Adaptive-Bound Triplet Loss (ABTL) to effectively utilize the input constraints, and present a novel Adaptive Hedge Update (AHU) method for online updating the model parameters. We empirically validates the effectiveness and efficacy of our framework on various applications such as real-world image classification, facial verification, and image retrieval. Yang Gao 0027, Yifan Li 0003, Swarup Chandra, Latifur Khan, Bhavani Thuraisingham |
WWW | 5 |
| 2019 | Multistream Classification for Cyber Threat Data with Heterogeneous Feature SpaceabstractUnder a newly introduced setting of multistream classification, two data streams are involved, which are referred to as source and target streams. The source stream continuously generates data instances from a certain domain with labels, while the target stream does the same task without labels from another domain. Existing approaches assume that domains for both data streams are identical, which is not quite true in real world scenario, since data streams from different sources may contain distinct features. Furthermore, obtaining labels for every instance in a data stream is often expensive and time-consuming. Therefore, it has become an important topic to explore whether labeled instances from other related streams can be helpful to predict those unlabeled instances in a given stream. Note that domains of source and target streams may have distinct features spaces and data distributions. Our objective is to predict class labels of data instances in the target stream by using the classifiers trained by the source stream. Yifan Li 0003, Yang Gao 0027, Gbadebo Ayoade, Hemeng Tao, Latifur Khan, Bhavani Thuraisingham |
WWW | 6 |
| 2018 | GCI: A Transfer Learning Approach for Detecting Cheats of Computer GameabstractCheating in massive multiple online games (MMOGs) adversely affect the game's popularity and reputation among its users. Therefore, game developers invest large amount of efforts to detect and prevent cheats that provide an unfair advantage to cheaters over other naive users during game play. Particularly, MMOG clients share data with the server during game play. Game developers leverage this data to detect cheating. However, detecting cheats is challenging mainly due to the limited client-side information, along with unknown and complex cheating techniques. In this paper, we aim to leverage machine learning based models to predict cheats over encrypted game traffic during game play. Concretely, network game traffic during game play from each player can be used to determine whether a cheat is employed. A major challenge in developing such a prediction model is the availability of sufficient training data, which is sparingly available in practice. Game traffic obtained from a few known players can be easily labeled. However, if such players are not a good representation of the population (i.e., other players), then a supervised model trained on labeled game traffic from these set of players may not generalize well for the population. Here, we propose a scalable transfer learning approach to overcome the constraints of limited labeled data. Our empirical evaluation on a popular MMOG demonstrates significant improvement in cheat prediction compared to other competing methods. Md Shihabul Islam, Swarup Chandra, Latifur Khan, Bhavani Thuraisingham |
IEEE BigData | 5 |
| 2018 | Attacklets: Modeling High Dimensionality in Real World CyberattacksabstractWe introduce attacklets, a novel approach to model the high dimensional interactions in cyberattacks. Attacklets are implemented using a real-world dataset of cyberattacks from the Verizon Data Breach Investigation Report. Whereas the commonly used attack graphs model the action sequences of attackers for specific exploits, attacklets model general attributes and states of each attack separately. Attacklets may inform the number and types of attributes across a wide range of cyberattacks. These structural properties can then be used in machine learning models to classify and predict future cyberattacks. Cuneyt Gurcan Akcora, Jonathan Z. Bakdash, Yulia R. Gel, Murat Kantarcioglu, Laura Marusich, Bhavani Thuraisingham |
ISI | 6 |
| 2018 | Lifting the Smokescreen: Detecting Underlying Anomalies During a DDoS AttackabstractWhile DDoS attacks have become an ever-growing threat in the last decade, a new variation is taking root in which the DDoS is used as a distraction or smokescreen to hide other malicious activity. This variation, which we call DDoS as a Smokescreen (DaaSS), often result in data theft and financial loss, and often are only detected because the theft is discovered independently, long after the attack has ceased. In this work, we set out to describe these attacks and present a novel approach to detect them using real-world network trace data. We present experimental results showing promise that DaaSS attacks can be detected in a manner conducive to practical deployment. Brian Ricks, Bhavani Thuraisingham, Patrick Tague |
ISI | 2 |
| 2018 | Privacy Preserving Synthetic Data Release Using Deep Learning
Nazmiye Ceren Abay, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham, Latanya Sweeney |
ECML/PKDD (1) | 4 |
| 2018 | Towards a Privacy-Aware Qunatified Self Data Management FrameworkabstractMassive amounts of data are being collected, stored, and analyzed for various business and marketing purposes. While such data analysis is critical for many applications, it could also violate the privacy of individuals. This paper describes the issues involved in designing a privacy aware data management framework for collecting, storing, and analyzing the data. We also discuss behavioral aspects of data sharing as well as aspects of a formal framework based on rewriting rules that encompasses the privacy aware data management framework. Bhavani Thuraisingham, Murat Kantarcioglu, Elisa Bertino, Jonathan Z. Bakdash, Maribel Fernández |
SACMAT | 1 |
| 2018 | Proactive user-centric secure data scheme using attribute-based semantic access controls for mobile clouds in financial industry
Meikang Qiu, Keke Gai, Bhavani Thuraisingham, Lixin Tao, Hui Zhao 0002 |
Future Gener. Comput. Syst. | 3 |
| 2017 | Focus location extraction from political news reports with bias correctionabstractAutomatic identification of geolocation mentioned in online news articles provide vital information for understanding associated events. While numerous open-source and commercial tools exist for geolocation extraction, they lack in reliable identification of fine-grained location, i.e., they identify location at country-level rather than a fine-grained city or locality level. The problem of location identification has been widely studied. Yet, most techniques depend on external knowledge-base or view the problem only in terms of Named Entity Recognition (NER), only to identify country-level location information. In this paper, we focus on news articles describing an event. A set of locations directly associated with the event are called focus locations. However, an event can occur only at a single location. Therefore, we aim to extract this location among focus locations, and call this as primary focus location. We propose a mechanism that utilizes the named entities to identify potential sentences containing focus locations, and then employ a supervised classification mechanism over sentence embedding to predict the primary focused geolocation. However, the main issue with such an approach is the unavailability of ground truth (i.e., whether words in a sentence is focus or non-focus) for training a classifier. In practice, labels from only a small number of news articles may be available for training due to high cost of manual labeling. If these articles are not a good representation of news articles in the wild, the classifier may not perform well. Therefore, we utilize an adaptation mechanism to overcome sampling bias in training data. Particularly, we train a classifier by using bias-corrected training data obtained from news articles published by an agency, while testing it on news articles published by a different agency. Our empirical results show superior performance compared to baseline approaches on real-world datasets consisting of news articles. Maryam Bahojb Imani, Swarup Chandra, Samuel Ma, Latifur Khan, Bhavani Thuraisingham |
IEEE BigData | 5 |
| 2017 | Unsupervised deep embedding for novel class detection over data streamabstractData streams are continuous flows of data points. Novel class detection is an important part of data stream mining. A novel class is a newly emerged class that has not previously been modeled by the classifier over the input stream. This paper proposes deep embedding for novel class detection - a novel approach that combines feature learning using denoising autoencoding with novel class detection. A denoising autoencoder is a neural network with hidden layers aiming to reconstruct the input vector from a corrupted version. A nonparametric multidimensional change point detection approach is also proposed, to detect concept-drift (the change of data feature values over time). Experiments on several real datasets show that the approach significantly improves the performance of novel class detection. Ahmad Mustafa 0001, Gbadebo Ayoade, Khaled Al-Naami, Latifur Khan, Kevin W. Hamlen, Bhavani Thuraisingham, Frederico Araujo |
IEEE BigData | 6 |
| 2017 | Securing Data Analytics on SGX with Randomization
Swarup Chandra, Vishal Karande, Zhiqiang Lin 0001, Latifur Khan, Murat Kantarcioglu, Bhavani Thuraisingham |
ESORICS (1) | 6 |
| 2017 | Hacking social network data miningabstractOver the years social network data has been mined to predict individuals' traits such as intelligence and sexual orientation. While mining social network data can provide many beneficial services to the user such as personalized experiences, it can also harm the user when used in making critical decisions such as employment. In this work, we investigate the reliability of applying data mining techniques on social network data to predict various individual traits. In spite of the preliminary success of such data mining applications, in this paper, we demonstrate the vulnerabilities of existing state of the art social network data mining techniques when they are facing malicious attacks. Our results indicate that making critical decisions, such as employment or credit approval, based solely on social network data mining results is still premature at this stage. Specifically, we explore Facebook likes data for predicting the traits of a Facebook user, including their political views and sexual orientation. We perform several types of malicious attacks on the predictive models to measure and understand their potential vulnerabilities. We find that existing predictive models built on social network data can be easily manipulated and suggest some countermeasures to prevent some of the proposed attacks. Yasmeen Alufaisan, Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ISI | 4 |
| 2016 | Adaptive encrypted traffic fingerprinting with bi-directional dependence
Khaled Al-Naami, Swarup Chandra, Ahmad Mustafa 0001, Latifur Khan, Zhiqiang Lin 0001, Kevin W. Hamlen, Bhavani Thuraisingham |
ACSAC | 7 |
| 2016 | Message from the SEPT Organizing CommitteeabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Bhavani Thuraisingham, Dianxiang Xu, Hiroki Takakura, Mohammad Zulkernine, Elisa Bertino |
COMPSAC | 1 |
| 2016 | Efficient handling of concept drift and concept evolution over Stream DataabstractTo decide if an update to a data stream classifier is necessary, existing sliding window based techniques monitor classifier performance on recent instances. If there is a significant change in classifier performance, these approaches determine a chunk boundary, and update the classifier. However, monitoring classifier performance is costly due to scarcity of labeled data. In our previous work, we presented a semi-supervised framework SAND, which uses change detection on classifier confidence to detect a concept drift. Unlike most approaches, it requires only a limited amount of labeled data to detect chunk boundaries and to update the classifier. However, SAND is expensive in terms of execution time due to exhaustive invocation of the change detection module. In this paper, we present an efficient framework, which is based on the same principle as SAND, but exploits dynamic programming and executes the change detection module selectively. Moreover, we provide theoretical justification of the confidence calculation, and show effect of a concept drift on subsequent confidence scores. Experiment results show efficiency of the proposed framework in terms of both accuracy and execution time. Ahsanul Haque, Latifur Khan, Michael Baron, Bhavani Thuraisingham, Charu C. Aggarwal |
ICDE | 4 |
| 2016 | Near real-time atrocity event codingabstractIn recent years, mass atrocities, terrorism, and political unrest have caused much human suffering. Thousands of innocent lives have been lost to these events. With the help of advanced technologies, we can now dream of a tool that uses machine learning and natural language processing (NLP) techniques to warn of such events. Detecting atrocities demands structured event data that contain metadata, with multiple fields and values (e.g. event date, victim, perpetrator). Traditionally, humans apply common sense and encode events from news stories but this process is slow, expensive, and ambiguous. To accelerate it, we use machine coding to generate an encoded event. In this paper, we develop a near-real-time supervised machine coding technique with an external knowledge base, WordNet, to generate a structured event. We design a Spark-based distributed framework with a web scraper to gather news reports periodically, process, and generate events. We use Spark to reduce the performance bottleneck while processing raw text news using CoreNLP. Mohiuddin Solaimani, Sayeed Salam, Ahmad Mustafa 0001, Latifur Khan, Patrick T. Brandt, Bhavani Thuraisingham |
ISI | 6 |
| 2016 | Online anomaly detection for multi-source VMware using a distributed streaming frameworkabstractSummary Anomaly detection refers to the identification of patterns in a dataset that do not conform to expected patterns. Such non‐conformant patterns typically correspond to samples of interest and are assigned to different labels in different domains, such as outliers, anomalies, exceptions, and malware. A daunting challenge is to detect anomalies in rapid voluminous streams of data. This paper presents a novel, generic real‐time distributed anomaly detection framework for multi‐source stream data. As a case study, we investigate anomaly detection for a multi‐source VMware‐based cloud data center, which maintains a large number of virtual machines (VMs). This framework continuously monitors VMware performance stream data related to CPU statistics (e.g., load and usage). It collects data simultaneously from all of the VMs connected to the network and notifies the resource manager to reschedule its CPU resources dynamically when it identifies any abnormal behavior from its collected data. A semi‐supervised clustering technique is used to build a model from benign training data only. During testing, if a data instance deviates significantly from the model, then it is flagged as an anomaly. Effective anomaly detection in this case demands a distributed framework with high throughput and low latency. Distributed streaming frameworks like Apache Storm, Apache Spark, S4, and others are designed for a lower data processing time and a higher throughput than standard centralized frameworks. We have experimentally compared the average processing latency of a tuple during clustering and prediction in both Spark and Storm and demonstrated that Spark processes a tuple much quicker than storm on average. Copyright © 2016 John Wiley & Sons, Ltd. Mohiuddin Solaimani, Mohammed Iftekhar, Latifur Khan, Bhavani Thuraisingham, Joey Burton Ingram, Sadi Evren Seker |
Softw. Pract. Exp. | 4 |
| 2016 | Role-Based Integrated Access Control and Data Provenance for SOA Based Net-Centric SystemsabstractIn multi-domain service-based systems, services from different domains are composed together to accomplish critical tasks. In these systems, data flow from one domain to another through the composed services. Thus, security and trustworthiness are the major concerns. Many access control models have been developed for service-based systems. Also, many data provenance schemes have been proposed in recent years to support data quality assessment and enhancement, data reproduction, etc. However, none of the existing mechanisms consider both access control and data provenance in an integrated model. In this paper, we propose an integrated role-based access control and data provenance model to secure the cross-domain interactions. We develop a role-based data provenance scheme which tracks the roles of originators/contributors of a data object and uses this information to help evaluate data trustworthiness. We also make use of the data provenance information and the derived data quality attributes to assist with cross domain access and information flow control. This integrated model mutually enhances data provenance and access control, providing better security and trustworthiness for many multi-domain service-based applications. Wei She, Wei Zhu 0002, I-Ling Yen, Farokh B. Bastani, Bhavani Thuraisingham |
IEEE Trans. Serv. Comput. | 5 |
| 2015 | Big Data Security and PrivacyabstractThis paper describes the issues surrounding big data security and privacy and provides a summary of the National Science Foundation sponsored workshop on this topic held in Dallas, Texas on September 16-17, 2014. Our goal is to build a community in big data security and privacy to explore the challenging research problems. Bhavani Thuraisingham |
CODASPY | 1 |
| 2015 | Message from SEPT Symposium Organizing CommitteeabstractPresents a listing of the Symposium organizing committee. Bhavani Thuraisingham, Dianxiang Xu, Hiroki Takakura, Mohammad Zulkernine, Elisa Bertino |
COMPSAC | 1 |
| 2015 | Honeypot based unauthorized data access detection in MapReduce systemsabstractThe data processing capabilities of MapReduce systems pioneered with the on-demand scalability of cloud computing have enabled the Big Data revolution. However, the data controllers/owners worried about the privacy and accountability impact of storing their data in the cloud infrastructures as the existing cloud computing solutions provide very limited control on the underlying systems. The intuitive approach - encrypting data before uploading to the cloud - is not applicable to MapReduce computation as the data analytics tasks are ad-hoc defined in the MapReduce environment using general programming languages (e.g, Java) and homomorphic encryption methods that can scale to big data do not exist. In this paper, we address the challenges of determining and detecting unauthorized access to data stored in MapReduce based cloud environments. To this end, we introduce alarm raising honeypots distributed over the data that are not accessed by the authorized MapReduce jobs, but only by the attackers and/or unauthorized users. Our analysis shows that unauthorized data accesses can be detected with reasonable performance in MapReduce based cloud environments. Huseyin Ulusoy, Murat Kantarcioglu, Bhavani Thuraisingham, Latifur Khan |
ISI | 3 |
| 2015 | Services in the CloudabstractProvides an in-depth review and analysis of cloud computing, applications for its use, cloud technologies, different kinds of cloud services, types of clouds, and major areas of research and development in the cloud computing field. Also provides an overview of the articles persented in this issue. Louise E. Moser, Bhavani Thuraisingham |
IEEE Trans. Serv. Comput. | 2 |
| 2015 | Emerging Web ServicesabstractThe articles in this special section focus on emerging Web Services. The papers address the latest advances in Web Services selection, discovery and recommendation. These advances consider context awareness, quality of service, quality of experience, service usage history and evolution, and cost effectiveness. Bhavani Thuraisingham, Liang-Jie Zhang, Louise E. Moser |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | Evolving Big Data Stream Classification with MapReduceabstractBig Data Stream mining has some inherent challenges which are not present in traditional data mining. Not only Big Data Stream receives large volume of data continuously, but also it may have different types of features. Moreover, the concepts and features tend to evolve throughout the stream. Traditional data mining techniques are not sufficient to address these challenges. In our current work, we have designed a multi-tiered ensemble based method HSMiner to address aforementioned challenges to label instances in an evolving Big Data Stream. However, this method requires building large number of AdaBoost ensembles for each of the numeric features after receiving each new data chunk which is very costly. Thus, HSMiner may face scalability issue in case of classifying Big Data Stream. To address this problem, we propose three approaches to build these large number of AdaBoost ensembles using MapReduce based parallelism. We compare each of these approaches from different aspects of design. We also empirically show that, these approaches are very useful for our base method to achieve significant scalability and speedup. Ahsanul Haque, Brandon Parker, Latifur Khan, Bhavani Thuraisingham |
IEEE CLOUD | 4 |
| 2014 | Statistical technique for online anomaly detection using Spark over heterogeneous data from multi-source VMware performance dataabstractAnomaly detection refers to the identification of patterns in a dataset that do not conform to expected patterns. Depending on the domain, the non-conformant patterns are assigned various tags, e.g. anomalies, outliers, exceptions, malwares and so forth. Online anomaly detection aims to detect anomalies in data flowing in a streaming fashion. Such stream data is commonplace in today's cloud data centers that house a large array of virtual machines(VM) producing vast amounts of performance data in real-time. Sophisticated detection mechanism will likely entail collation of data from heterogeneous sources with diversified data format and semantics. Therefore, detection of performance anomaly in this context requires a distributed framework with high throughput and low latency. Apache Spark is one such framework that represents the bleeding-edge amongst its contemporaries. In this paper, we have taken up the challenge of anomaly detection in VMware based cloud data centers. We have employed a Chi-square based statistical anomaly detection technique in Spark. We have demonstrated how to take advantage of the high processing power of Spark to perform anomaly detection on heterogeneous data using statistical techniques. Our approach is optimally designed to cope with the heterogeneity of input data streams and the experiments we conducted testify to its efficacy in online anomaly detection. Mohiuddin Solaimani, Mohammed Iftekhar, Latifur Khan, Bhavani Thuraisingham |
IEEE BigData | 4 |
| 2014 | Spark-based anomaly detection over multi-source VMware performance data in real-timeabstractAnomaly detection refers to identifying the patterns in data that deviate from expected behavior. These non-conforming patterns are often termed as outliers, malwares, anomalies or exceptions in different application domains. This paper presents a novel, generic real-time distributed anomaly detection framework for multi-source stream data. As a case study, we have decided to detect anomaly for multi-source VMware-based cloud data center. The framework monitors VMware performance stream data (e.g., CPU load, memory usage, etc.) continuously. It collects these data simultaneously from all the VMwares connected to the network. It notifies the resource manager to reschedule its resources dynamically when it identifies any abnormal behavior of its collected data. We have used Apache Spark, a distributed framework for processing performance stream data and making prediction without any delay. Spark is chosen over a traditional distributed framework (e.g., Hadoop and MapReduce, Mahout, etc.) that is not ideal for stream data processing. We have implemented a flat incremental clustering algorithm to model the benign characteristics in our distributed Spark based framework. We have compared the average processing latency of a tuple during clustering and prediction in Spark with Storm, another distributed framework for stream data processing. We experimentally find that Spark processes a tuple much quicker than Storm on average. Mohiuddin Solaimani, Mohammed Iftekhar, Latifur Khan, Bhavani Thuraisingham, Joey Burton Ingram |
CICS | 4 |
| 2014 | Evolving stream classification using change detectionabstractClassifying instances in evolving data stream is a challenging task because of its properties, e.g., infinite length, concept drift, and concept evolution. Most of the currently available approaches to classify stream data instances divide the stream data into fixed size chunks to fit the data in m Ahmad Mustafa 0001, Ahsanul Haque, Latifur Khan, Michael Baron, Bhavani Thuraisingham |
CollaborateCom | 5 |
| 2014 | Stream Mining Using Statistical Relational LearningabstractStream mining has gained popularity in recent years due to the availability of numerous data streams from sources such as social media and sensor networks. Data mining on such continuous streams possess a variety of challenges including concept drift and unbounded stream length. Traditional data mining approaches to these problems have difficulty incorporating relational domain knowledge and feature relationships, which can be used to improve the accuracy of a classifier. In this work, we model large data streams using statistical relational learning techniques for classification, in particular, we use a Markov Logic Network to capture relational features in structured data and show that this approach performs better for supervised learning than current state-of-the-art approaches. Additionally, we evaluate our approach with semi-supervised learning scenarios, where class labels are only partially available during training. Swarup Chandra, Justin Sahs, Latifur Khan, Bhavani Thuraisingham, Charu C. Aggarwal |
ICDM | 4 |
| 2014 | Towards fine grained RDF access controlabstractThe Semantic Web is envisioned as the future of the current web, where the information is enriched with machine understandable semantics. According to the World Wide Web Consortium (W3C), "The Semantic Web provides a common framework that allows data to be shared and reused across application, enterprise, and community boundaries". Among the various technologies that empower Semantic Web, the most significant ones are Resource Description Framework (RDF) and SPARQL, which facilitate data integration and a means to query respectively. Although Semantic Web is elegantly and effectively equipped for data sharing and integration via RDF, lack of efficient means to securely share data pose limitations in practice. In order to make data sharing and integration pragmatic for Semantic Web, we present a query language based secure data sharing mechanism. We extend SPARQL with a new query form called SANITIZE which comprises a set of sanitization operations that are used to sanitize or mask sensitive data within an RDF graph. The sanitization operations can be further leveraged towards RDF access control and anonymization, thus enabling secure sharing of RDF data. Jyothsna Rachapalli, Vaibhav Khadilkar, Murat Kantarcioglu, Bhavani Thuraisingham |
SACMAT | 4 |
| 2014 | Redaction based RDF access control languageabstractWe propose an access control language for securing RDF graphs which essentially leverages an underlying query language based redaction mechanism to provide fine grained RDF access control. The access control language presented is equipped with critical features such as policy resolution and cascading policies that are essential for fine grained RDF access control. We present the architecture of our system which primarily features a flexible, scalable and general purpose RDF access control mechanism. Jyothsna Rachapalli, Vaibhav Khadilkar, Murat Kantarcioglu, Bhavani Thuraisingham |
SACMAT | 4 |
| 2014 | A roadmap for privacy-enhanced secure data provenance
Elisa Bertino, Gabriel Ghinita, Murat Kantarcioglu, Dang Nguyen 0001, Jae Park, Ravi S. Sandhu, Salmin Sultana, Bhavani Thuraisingham, Shouhuai Xu |
J. Intell. Inf. Syst. | 8 |
| 2013 | Intelligent MapReduce Based Framework for Labeling Instances in Evolving Data StreamabstractIn our current work, we have proposed a multi-tiered ensemble based robust method to address all of the challenges of labeling instances in evolving data stream. Bottleneck of our current work is, it needs to build ADABOOST ensembles for each of the numeric features. This can face scalability issue as number of features can be very large at times in data stream. In this paper, we propose an intelligent approach to build these large number of ADABOOST ensembles with MapReduce based parallelism. We show that, this approach can help our base method to achieve significant scalability without compromising classification accuracy. We analyze different aspects of our design to depict advantages and disadvantages of the approach. We also compare and analyze performance of the proposed approach in terms of execution time, speedup and scale up. Ahsanul Haque, Brandon Parker, Latifur Khan, Bhavani Thuraisingham |
CloudCom (2) | 4 |
| 2013 | MapReduce-guided scalable compressed dictionary construction for evolving repetitive sequence streamsabstractUsers' repetitive daily or weekly activities may constitute user profiles. For example, a user's frequent command sequences may represent normative pattern of that user. To find normative patterns over dynamic data streams of unbounded length is challenging. For this, an unsupervised learning approa Pallabi Parveen, Pratik Desai, Bhavani Thuraisingham, Latifur Khan |
CollaborateCom | 3 |
| 2013 | Analysis of heuristic based access pattern obfuscationabstractAs cloud computing becomes popular, the security and privacy issues emerge as important hindrances to more widespread adoption of cloud computing. In particular, outsourcing sensitive data to untrusted cloud service providers creates important security and regulatory compliance challenges. Encryptio Huseyin Ulusoy, Murat Kantarcioglu, Bhavani Thuraisingham, Ebru Celikel Cankaya, Erman Pattuk |
CollaborateCom | 3 |
| 2013 | Least Cost Rumor Blocking in Social NetworksabstractIn many real-world scenarios, social network serves as a platform for information diffusion, alongside with positive information (truth) dissemination, negative information (rumor) also spread among the public. To make the social network as a reliable medium, it is necessary to have strategies to control rumor diffusion. In this article, we address the Least Cost Rumor Blocking (LCRB) problem where rumors originate from a community Cr in the network and a notion of protectors are used to limit the bad influence of rumors. The problem can be summarized as identifying a minimal subset of individuals as initial protectors to minimize the number of people infected in neighbor communities of Cr at the end of both diffusion processes. Observing the community structure property, we pay attention to a kind of vertex set, called bridge end set, in which each node has at least one direct in-neighbor in Cr and is reachable from rumors. Under the OOAO model, we study LCRB-P problem, in which α (0 < 1) fraction of bridge ends are required to be protected. We prove that the objective function of this problem is submodular and a greedy algorithm is adopted to derive a (1 -- 1/e)-approximation. Furthermore, we study LCRB-D problem over the DOAA model, in which all the bridge ends are required to be protected, we prove that there is no polynomial time o(ln n)-approximation for the LCRB-D problem unless P = NP, and propose a Set Cover Based Greedy (SCBG) algorithm which achieves a O(ln n)-approximation ratio. Finally, to evaluate the efficiency and effectiveness of our algorithm, we conduct extensive comparison simulations in three real-world datasets, and the results show that our algorithm outperforms other heuristics. Lidan Fan, Zaixin Lu, Weili Wu 0001, Bhavani Thuraisingham, Yuanjun Bi |
ICDCS | 4 |
| 2013 | Measuring expertise and bias in cyber security using cognitive and neuroscience approachesabstractToward the ultimate goal of enhancing human performance in cyber security, we attempt to understand the cognitive components of cyber security expertise. Our initial focus is on cyber security attackers - often called “hackers”. Our first aim is to develop behavioral measures of accuracy and response time to examine the cognitive processes of pattern-recognition, reasoning and decision-making that underlie the detection and exploitation of security vulnerabilities. Understanding these processes at a cognitive level will lead to theory development addressing questions about how cyber security expertise can be identified, quantified, and trained. In addition to behavioral measures our plan is to conduct a functional magnetic resonance imaging (fMRI) study of neural processing patterns that can differentiate persons with different levels of cyber security expertise. Our second aim is to quantitatively assess the impact of attackers' thinking strategies - conceptualized by psychologists as heuristics and biases - on their susceptibility to defensive techniques (e.g., “decoys,” “honeypots”). Honeypots are an established method to lure attackers into exploiting a dummy system containing misleading or false content, distracting their attention from genuinely sensitive information, and consuming their limited time and resources. We use the extensive research and experimentation that we have carried out to study the minds of successful chess players in order to study the minds of hackers with the ultimate goal of enhancing the security of current systems. This paper outlines our approach. Daniel C. Krawczyk, James Bartlett, Murat Kantarcioglu, Kevin W. Hamlen, Bhavani Thuraisingham |
ISI | 5 |
| 2013 | Adaptive Information Coding for Secure and Reliable Wireless Telesurgery Communications
M. Engin Tozal, Yongge Wang 0001, Ehab Al-Shaer, Kamil Saraç, Bhavani Thuraisingham, Bei-tseng Chu |
Mob. Networks Appl. | 5 |
| 2013 | Preventing Private Information Inference Attacks on Social NetworksabstractOnline social networks, such as Facebook, are increasingly utilized by many people. These networks allow users to publish details about themselves and to connect to their friends. Some of the information revealed inside these networks is meant to be private. Yet it is possible to use learning algorithms on released data to predict private information. In this paper, we explore how to launch inference attacks using released social networking data to predict private information. We then devise three possible sanitization techniques that could be used in various situations. Then, we explore the effectiveness of these techniques and attempt to use methods of collective inference to discover sensitive attributes of the data set. We show that we can decrease the effectiveness of both local and relational classification algorithms by using the sanitization methods we described. Raymond Heatherly, Murat Kantarcioglu, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2013 | Security-Aware Service Composition with Fine-Grained Information Flow ControlabstractEnforcing access control in composite services is essential in distributed multidomain environment. Many advanced access control models have been developed to secure web services at execution time. However, they do not consider access control validation at composition time, resulting in high execution-time failure rate of composite services due to access control violations. Performing composition-time access control validation is not straightforward. First, many candidate compositions need to be considered and validating them can be costly. Second, some service composers may not be trusted to access protected policies and validation has to be done remotely. Another major issue with existing models is that they do not consider information flow control in composite services, which may result in undesirable information leakage. To resolve all these problems, we develop a novel three-phase composition protocol integrating information flow control. To reduce the policy evaluation cost, we use historical information to efficiently evaluate and prune candidate compositions and perform local/remote policy evaluation only on top candidates. To achieve effective and efficient information flow control, we introduce the novel concept of transformation factor to model the computation effect of intermediate services. Experimental studies show significant performance benefit of the proposed mechanism. Wei She, I-Ling Yen, Bhavani Thuraisingham, Elisa Bertino |
IEEE Trans. Serv. Comput. | 3 |
| 2012 | Cloud Guided Stream Classification Using Class-Based EnsembleabstractWe propose a novel class-based micro-classifier ensemble classification technique (MCE) for classifying data streams. Traditional ensemble-based data stream classification techniques build a classification model from each data chunk and keep an ensemble of such models. Due to the fixed length of the ensemble, when a new model is trained, one existing model is discarded. This creates several problems. First, if a class disappears from the stream and reappears after a long time, it would be misclassified if a majority of the classifiers in the ensemble does not contain any model of that class. Second, discarding a model means discarding the corresponding data chunk completely. However, knowledge obtained from some classes might be still useful and if they are discarded, the overall error rate would increase. To address these problems, we propose an ensemble model where each class information is stored separately. From each data chunk, we train a model for each class of data. We call each such model a micro-classifier. This approach is more robust than existing chunk-based ensembles in handling dynamic changes in the data stream. To the best of our knowledge, this is the first attempt to classify data streams using the class-based ensembles approach. When the number of classes grow in the stream, class-based ensembles may degrade in performance (speed). Hence, we sketch a cloud-based solution of our class-based ensembles to handle a large number of classes effectively. We compare our technique with several state-of-the-art data stream classification techniques on both synthetic and benchmark data streams, and obtain much higher accuracy. Tahseen Al-Khateeb, Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham |
IEEE CLOUD | 4 |
| 2012 | Risk-Aware Workload Distribution in Hybrid CloudsabstractThis paper explores an efficient and secure mechanism to partition computations across public and private machines in a hybrid cloud setting. We propose a principled framework for distributing data and processing in a hybrid cloud that meets the conflicting goals of performance, sensitive data disclosure risk and resource allocation costs. The proposed solution is implemented as an add-on tool for a Hadoop and Hive based cloud computing infrastructure. Our experiments demonstrate that the developed mechanism can lead to a major performance gain by exploiting both the hybrid cloud components without violating any pre-determined public cloud usage constraints. Kerim Yasin Oktay, Vaibhav Khadilkar, Bijit Hore, Murat Kantarcioglu, Sharad Mehrotra, Bhavani Thuraisingham |
IEEE CLOUD | 6 |
| 2012 | Tweecalization: Efficient and intelligent location mining in twitter using semi-supervised learningabstractGeosocial Networking is the new hotness, with social networks providing services and capabilities to the users to associate location to their profiles. But, because of privacy and security reasons, most of the people on social networking sites like Twitter are unwilling to provide locations in their Satyen Abrol, Latifur Khan, Bhavani Thuraisingham |
CollaborateCom | 3 |
| 2012 | Randomizing Smartphone Malware Profiles against Statistical Mining Techniques
Abhijith Shastry, Murat Kantarcioglu, Yan Zhou 0001, Bhavani Thuraisingham |
DBSec | 4 |
| 2012 | Stream Classification with Recurring and Novel Class Detection Using Class-Based EnsembleabstractConcept-evolution has recently received a lot of attention in the context of mining data streams. Concept-evolution occurs when a new class evolves in the stream. Although many recent studies address this issue, most of them do not consider the scenario of recurring classes in the stream. A class is called recurring if it appears in the stream, disappears for a while, and then reappears again. Existing data stream classification techniques either misclassify the recurring class instances as another class, or falsely identify the recurring classes as novel. This increases the prediction error of the classifiers, and in some cases causes unnecessary waste in memory and computational resources. In this paper we address the recurring class issue by proposing a novel "class-based" ensemble technique, which substitutes the traditional "chunk-based" ensemble approaches and correctly distinguishes between a recurring class and a novel one. We analytically and experimentally confirm the superiority of our method over state-of-the-art techniques. Tahseen Al-Khateeb, Mohammad M. Masud 0001, Latifur Khan, Charu C. Aggarwal, Jiawei Han 0001, Bhavani Thuraisingham |
ICDM | 6 |
| 2012 | Self-Training with Selection-by-RejectionabstractPractical machine learning and data mining problems often face shortage of labeled training data. Self-training algorithms are among the earliest attempts of using unlabeled data to enhance learning. Traditional self-training algorithms label unlabeled data on which classifiers trained on limited training data have the highest confidence. In this paper, a self-training algorithm that decreases the disagreement region of hypotheses is presented. The algorithm supplements the training set with self-labeled instances. Only instances that greatly reduce the disagreement region of hypotheses are labeled and added to the training set. Empirical results demonstrate that the proposed self-training algorithm can effectively improve classification performance. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ICDM | 3 |
| 2012 | Sparse Bayesian Adversarial Learning Using Relevance Vector Machine EnsemblesabstractData mining tasks are made more complicated when adversaries attack by modifying malicious data to evade detection. The main challenge lies in finding a robust learning model that is insensitive to unpredictable malicious data distribution. In this paper, we present a sparse relevance vector machine ensemble for adversarial learning. The novelty of our work is the use of individualized kernel parameters to model potential adversarial attacks during model training. We allow the kernel parameters to drift in the direction that minimizes the likelihood of the positive data. This step is interleaved with learning the weights and the weight priors of a relevance vector machine. Our empirical results demonstrate that an ensemble of such relevance vector machine models is more robust to adversarial attacks. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham |
ICDM | 3 |
| 2012 | Design and implementation of SNODSOC: Novel class detection for social network analysisabstractThis paper describes a framework, SNODSOC (Stream based novel class detection for social network analysis), that detects evolving patterns and trends in social microblogs. SNODSOC extends our powerful data mining system, SNOD (Stream-based Novel Class Detection) for now detecting novel patterns and trends within microblogs. Satyen Abrol, Latifur Khan, Vaibhav Khadilkar, Bhavani Thuraisingham, Tyrone Cadenhead |
ISI | 4 |
| 2012 | Extracting semantic information structures from free text law enforcement dataabstractA detective distributes information on a current case to his law enforcement peers. He quickly receives a computer generated response with leads identified within hundreds of thousands of previously distributed free text documents from thousands of other detectives. The challenges lie in the nature of free text - unstructured formats, confusing word usage, cut-andpaste additions, abbreviations, inserted html/xml tags, multimedia content, and domain-specific terminology. This research proposes a new data structure, the semantic information structure, which encapsulates the extracted content information on classes of information such as people, vehicles, events, organizations, objects, and locations as well as the contextual information about the connections and measures to enable prioritization of files containing related pieces of content. The structure is organized to be a result of automated natural language processing methods that extract entities, expanded entity phrases and their links which are driven by ontologies, DLSafe rules, abductive hypotheses and semantic composition. Importance and significance measures aid in prioritization. James R. (Bob) Johnson, Anita Miller, Latifur Khan, Bhavani Thuraisingham |
ISI | 4 |
| 2012 | Towards cyber operations - The new role of academic cyber security research and educationabstractThe shift towards cyber operations represents a shift not only for the defense establishments worldwide but also cyber security research and education. Traditionally cyber security research and education has been founded on information assurance, expressed in underlying subfields such as forensics, network security, and penetration testing. Cyber security research and education is connected to the homeland security agencies and defense through funding, mutual interest in the outcome of the research, and the potential job market for graduates. The future of cyber security is both defensive information assurance measures and active defense driven information operations that jointly and coordinately are launched, in the pursuit of a cohesive and decisive execution of the national cyber defense strategy. The cohesive cyber defense requires universities to optimize their campus wide resources to fuse knowledge, intellectual capacity, and practical skills in an unprecedented way in cyber security. The future will require cyber defense research teams to address not only computer science, electrical engineering, software and hardware security, but also political theory, institutional theory, behavioral science, deterrence theory, ethics, international law, international relations, and additional social sciences. This paper is the result of an ocular survey of the U.S. 48 academic CAE-R research centers, evaluating the collective group of research centers' ability to adapt to the shift towards cyber operations, and the challenges therein. Jan Kallberg, Bhavani Thuraisingham |
ISI | 2 |
| 2012 | Unsupervised incremental sequence learning for insider threat detectionabstractInsider threat detection requires the identification of rare anomalies in contexts where evolving behaviors tend to mask such anomalies. This paper proposes and tests an incremental learning algorithm based on unsupervised learning that addresses this challenge by maintaining repetitive sequences in a compressed dictionary to identify anomaly over dynamic data streams of unbounded length. For unsupervised learning, compression-based techniques are used to model normal behavior sequences. The result is a classifier that exhibits substantially increased classification accuracy for insider threat streams relative to traditional static learning approaches and effectiveness over supervised learning approaches. Pallabi Parveen, Bhavani Thuraisingham |
ISI | 2 |
| 2012 | Adversarial support vector machine learningabstractMany learning tasks such as spam filtering and credit card fraud detection face an active adversary that tries to avoid detection. For learning problems that deal with an active adversary, it is important to model the adversary's attack strategy and develop robust learning models to mitigate the attack. These are the two objectives of this paper. We consider two attack models: a free-range attack model that permits arbitrary data corruption and a restrained attack model that anticipates more realistic attacks that a reasonable adversary would devise under penalties. We then develop optimal SVM learning strategies against the two attack models. The learning algorithms minimize the hinge loss while assuming the adversary is modifying data to maximize the loss. Experiments are performed on both artificial and real data sets. We demonstrate that optimal solutions may be overly pessimistic when the actual attacks are much weaker than expected. More important, we demonstrate that it is possible to develop a much more resilient SVM learning model while making loose assumptions on the data corruption models. When derived under the restrained attack model, our optimal SVM learning strategy provides more robust overall performance under a wide range of attack parameters. Yan Zhou 0001, Murat Kantarcioglu, Bhavani Thuraisingham, Bowei Xi |
KDD | 3 |
| 2012 | A cloud-based RDF policy engine for assured information sharingabstractIn this paper, we describe a general-purpose, scalable RDF policy engine. The innovations in our work include seamless support for a diverse set of security policies enforced by a highly available and scalable policy engine designed using a cloud-based platform. Our main goal is to demonstrate how coalition agencies can share information stored in multiple formats, through the enforcement of appropriate policies. Tyrone Cadenhead, Vaibhav Khadilkar, Murat Kantarcioglu, Bhavani Thuraisingham |
SACMAT | 4 |
| 2012 | PTAS for the minimum weighted dominating set in growth bounded graphs
Wei Wang 0032, Joonmo Kim, Bhavani Thuraisingham, Weili Wu 0001 |
J. Glob. Optim. | 4 |
| 2012 | Guest Editors' Introduction: Special Section on Data and Applications Security and PrivacyabstractThe four papers in this special section focus on the latest advancements in data and application systems in the information security and privacy industry. Elena Ferrari 0001, Bhavani Thuraisingham |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2012 | Effective Software Fault Localization Using an RBF Neural NetworkabstractWe propose the application of a modified radial basis function neural network in the context of software fault localization, to assist programmers in locating bugs effectively. This neural network is trained to learn the relationship between the statement coverage information of a test case and its corresponding execution result, success or failure. The trained network is then given as input a set of virtual test cases, each covering a single statement. The output of the network, for each virtual test case, is considered to be the suspiciousness of the corresponding covered statement. A statement with a higher suspiciousness has a higher likelihood of containing a bug, and thus statements can be ranked in descending order of their suspiciousness. The ranking can then be examined one by one, starting from the top, until a bug is located. Case studies on 15 different programs were conducted, and the results clearly show that our proposed technique is more effective than several other popular, state of the art fault localization techniques. Further studies investigate the robustness of the proposed technique, and illustrate how it can easily be applied to programs with multiple bugs as well. W. Eric Wong, Vidroha Debroy, Richard M. Golden, Bhavani Thuraisingham |
IEEE Trans. Reliab. | 5 |
| 2011 | Scalable Complex Query Processing over Large Semantic Web Data Using CloudabstractCloud computing solutions continue to grow increasingly popular both in research and in the commercial IT industry. With this popularity comes ever increasing challenges for the cloud computing service providers. Semantic web is another domain of rapid growth in both research and industry. RDF datasets are becoming increasingly large and complex and existing solutions do not scale adequately. In this paper, we will detail a scalable semantic web framework built using cloud computing technologies. We define solutions for generating and executing optimal query plans. We handle not only queries with Basic Graph Patterns (BGP) but also complex queries with optional blocks. We have devised a novel algorithm to handle these complex queries. Our algorithm minimizes binding triple patterns and joins between them by identifying common blocks by algorithms to find sub graph isomorphism and building a query plan utilizing that information. We utilize Hadoop's MapReduce framework to process our query plan. We will show that our framework is extremely scalable and efficiently answers complex queries. Mohammad Farhan Husain, James P. McGlothlin, Latifur Khan, Bhavani Thuraisingham |
IEEE CLOUD | 4 |
| 2011 | A language for provenance access controlabstractProvenance is a directed acyclic graph that explains how a resource came to be in its current form. Traditional access control does not support provenance graphs. We cannot achieve all the benefits of access control if the relationships between the data and their sources are not protected. In this paper, we propose a language that complements and extends existing access control languages to support provenance. This language also provides access to data based on integrity criteria. We have also built a prototype to show that this language can be implemented effectively using Semantic Web technologies. Tyrone Cadenhead, Vaibhav Khadilkar, Murat Kantarcioglu, Bhavani Thuraisingham |
CODASPY | 4 |
| 2011 | Towards privacy preserving access control in the cloudabstractIt is very costly and cumbersome to manage database systems in-house especially for small or medium organizations. Data-as-a-Service (DaaS) hosted in the cloud provides an attractive solution, which is flexible, reliable, easy and economical to operate, for such organizations. However security a Mohamed Nabeel, Elisa Bertino, Murat Kantarcioglu, Bhavani Thuraisingham |
CollaborateCom | 4 |
| 2011 | Detecting Recurring and Novel Classes in Concept-Drifting Data StreamsabstractConcept-evolution is one of the major challenges in data stream classification, which occurs when a new class evolves in the stream. This problem remains unaddressed by most state-of-the-art techniques. A recurring class is a special case of concept-evolution. This special case takes place when a class appears in the stream, then disappears for a long time, and again appears. Existing data stream classification techniques that address the concept-evolution problem, wrongly detect the recurring classes as novel class. This creates two main problems. First, much resource is wasted in detecting a recurring class as novel class, because novel class detection is much more computationally- and memory-intensive, as compared to simply recognizing an existing class. Second, when a novel class is identified, human experts are involved in collecting and labeling the instances of that class for future modeling. If a recurrent class is reported as novel class, it will be only a waste of human effort to find out whether it is really a novel class. In this paper, we address the recurring issue, and propose a more realistic novel class detection technique, which remembers a class and identifies it as "not novel" when it reappears after a long disappearance. Our approach has shown significant reduction in classification error over state-of-the-art stream classification techniques on several benchmark data streams. Mohammad M. Masud 0001, Tahseen Al-Khateeb, Latifur Khan, Charu C. Aggarwal, Jing Gao 0004, Jiawei Han 0001, Bhavani Thuraisingham |
ICDM | 7 |
| 2011 | Supervised Learning for Insider Threat Detection Using Stream MiningabstractInsider threat detection requires the identification of rare anomalies in contexts where evolving behaviors tend to mask such anomalies. This paper proposes and tests an ensemble-based stream mining algorithm based on supervised learning that addresses this challenge by maintaining an evolving collection of multiple models to classify dynamic data streams of unbounded length. The result is a classifier that exhibits substantially increased classification accuracy for real insider threat streams relative to traditional supervised learning (traditional SVM and one-class SVM) and other single-model approaches. Pallabi Parveen, Zackary R. Weger, Bhavani Thuraisingham, Kevin W. Hamlen, Latifur Khan |
ICTAI | 3 |
| 2011 | Ontology-Driven Query Expansion Using Map/Reduce Framework to Facilitate Federated QueriesabstractIn view of the need for a highly distributed and federated architecture, a robust query expansion has great impact on the performance of information retrieval. We aim to determine ontology-driven query expansion terms using different weighting techniques. For this, we consider each individual ontology and user query keywords to determine the Basic Expansion Terms (BET) using a number of semantic measures including Betweenness Measure (BM) and Semantic Similarity Measure (SSM). We propose a Map/Reduce distributed algorithm for calculating all the shortest paths in ontology graph. Map/Reduce algorithm will improve considerably the efficiency of BET calculation for large ontologies. Neda Alipanah, Pallabi Parveen, Latifur Khan, Bhavani Thuraisingham |
ICWS | 4 |
| 2011 | Rule-Based Run-Time Information Flow Control in Service CloudabstractService cloud provides added value to customers by allowing them to compose services from multiple providers. Most existing web service security models focus on the protection of individual web services. When multiple services from different domains are composed together, it is critical to ensure the proper information flow on the chain of services. In a service chain, each service needs to determine whether the sensitive information can be directly or indirectly disseminated to the subsequent services. Also, each service in the chain needs to decide whether to accept the data passed to it directly or indirectly from prior services. Moreover, the input data that service si receives from si-1, si. InF, may cause certain side effects inside si, such as updating si's backend database using data computed from si. InF. Service si may wish to allow such side effects in one situation while reject some side effects in another situation. All these decisions should be made based on the service's information flow control policies. To achieve fine-grained information flow control, it is also necessary to analyze the flow and processing of the data and derive the dependencies between the data dynamically generated or used in a service chain. In this paper, we develop a run-time information flow control model for service cloud. First, we develop a run-time dependency analysis mechanism which enables each service in the service chain to determine the correlation between the locally accessed data and the data dynamically generated by the services in the service chain. Then, we develop a model to enable each service in a service chain to specify policies on how its sensitive information can be released to its subsequent services and what types of input data from prior services can be accepted and how they can flow within the services. Finally, we design a run-time protocol to enforce these policies in a service chain. Wei She, I-Ling Yen, Bhavani Thuraisingham, San-Yih Huang |
ICWS | 3 |
| 2011 | RDFKB: A Semantic Web Knowledge BaseabstractThere are many significant research projects focused on providing semantic web repositories that are scalable and efficient. However, the true value of the semantic web architecture is its ability to represent meaningful knowledge and not just data. Therefore, a semantic web knowledge base should do more than retrieve collections of triples. We propose RDFKB (Resource Description Knowledge Base), a complete semantic web knowledge case. RDFKB is a solution for managing, persisting and querying semantic web knowledge. Our experiments with real world and synthetic datasets demonstrate that RDFKB achieves superior query performance to other state-of-the-art solutions. The key features of RDFKB that differentiate it from other solutions are: 1) a simple and efficient process for data additions, deletions and updates that does not involve reprocessing the dataset; 2) materialization of inferred triples at addition time without performance degradation; 3) materialization of uncertain information and support for queries involving probabilities; 4) distributed inference across datasets; 5) ability to apply alignments to the dataset and perform queries against multiple sources using alignment. RDFKB allows more knowledge to be stored and retrieved; it is a repository not just for RDF datasets, but also for inferred triples, probability information, and lineage information. RDFKB provides a complete and efficient RDF data repository and knowledge base. James P. McGlothlin, Latifur Khan, Bhavani Thuraisingham |
IJCAI | 3 |
| 2011 | Identification of related information of interest across free text documentsabstractAn approach is presented for finding information of interest in a free text document and then identifying and presenting related information of interest from other free text documents. The goal is to find specific related items of interest within documents whether the documents are of the same category or not. Information of interest is defined with respect to expanded entity phrases and their ontology mappings. Powerful techniques requiring minimal training are described for expanding an entity phrase to include attributes from components of a complex sentence; for measuring relatedness of same-name expanded entity phrases; and for detecting related expanded entity phrases through ontology inferences. A representative dataset is described and preliminary measurements of performance against ground truth are provided. James R. (Bob) Johnson, Anita Miller, Latifur Khan, Bhavani Thuraisingham, Murat Kantarcioglu |
ISI | 4 |
| 2011 | Extraction of expanded entity phrasesabstractThis research is part of a larger integrated approach for extraction of information of interest from free text and the visualization of semantic relatedness between phrases of interest. This paper defines a new structure which is a key component, the expanded entity phrase (EPx). This paper also presents an approach for extracting EPx's from free text. The structure of the EPx's facilitates quantitative comparison with other EPx's. A combination of part of speech-based template matching and ontology-driven NLP provides an effective technique for extracting complex entity structures that cross clause boundaries. This approach also uses ontology-based inferences to lay the ground work for linking EPx's for semantic relatedness assessments involving different named entities not explicitly stated in the text. The real world data used in this research were derived from a collection of law enforcement email messages submitted by hundreds of investigators seeking information or posting information about crimes, incidents, requests, and announcements. Performance data on the approaches used for extracting EPx's and links from this data are presented. James R. (Bob) Johnson, Anita Miller, Latifur Khan, Bhavani Thuraisingham, Murat Kantarcioglu |
ISI | 4 |
| 2011 | Differentiating Code from Data in x86 Binaries
Richard Wartell, Yan Zhou 0001, Kevin W. Hamlen, Murat Kantarcioglu, Bhavani Thuraisingham |
ECML/PKDD (3) | 5 |
| 2011 | Transforming provenance using redactionabstractOngoing mutual relationships among entities rely on sharing quality information while preventing release of sensitive content. Provenance records the history of a document for ensuring both, the quality and trustworthiness; while redaction identifies and removes sensitive information from a document. Traditional redaction techniques do not extend to the directed graph representation of provenance. In this paper, we propose a graph grammar approach for rewriting redaction policies over provenance. Our rewriting procedure converts a high level specification of a redaction policy into a graph grammar rule that transforms a provenance graph into a redacted provenance graph. Our prototype shows that this approach can be effectively implemented using Semantic Web technologies. Tyrone Cadenhead, Vaibhav Khadilkar, Murat Kantarcioglu, Bhavani Thuraisingham |
SACMAT | 4 |
| 2011 | Semantic web-based social network access control
Barbara Carminati, Elena Ferrari 0001, Raymond Heatherly, Murat Kantarcioglu, Bhavani Thuraisingham |
Comput. Secur. | 5 |
| 2011 | Heuristics-Based Query Processing for Large RDF Graphs Using Cloud ComputingabstractSemantic web is an emerging area to augment human reasoning. Various technologies are being developed in this arena which have been standardized by the World Wide Web Consortium (W3C). One such standard is the Resource Description Framework (RDF). Semantic web technologies can be utilized to build efficient and scalable systems for Cloud Computing. With the explosion of semantic web technologies, large RDF graphs are common place. This poses significant challenges for the storage and retrieval of RDF graphs. Current frameworks do not scale for large RDF graphs and as a result do not address these challenges. In this paper, we describe a framework that we built using Hadoop to store and retrieve large numbers of RDF triples by exploiting the cloud computing paradigm. We describe a scheme to store RDF data in Hadoop Distributed File System. More than one Hadoop job (the smallest unit of execution in Hadoop) may be needed to answer a query because a single triple pattern in a query cannot simultaneously take part in more than one join in a single Hadoop job. To determine the jobs, we present an algorithm to generate query plan, whose worst case cost is bounded, based on a greedy approach to answer a SPARQL Protocol and RDF Query Language (SPARQL) query. We use Hadoop's MapReduce framework to answer the queries. Our results show that we can store large RDF graphs in Hadoop clusters built with cheap commodity class hardware. Furthermore, we show that our framework is scalable and efficient and can handle large amounts of RDF data, unlike traditional approaches. Mohammad Farhan Husain, James P. McGlothlin, Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2011 | Classification and Novel Class Detection in Concept-Drifting Data Streams under Time ConstraintsabstractMost existing data stream classification techniques ignore one important aspect of stream data: arrival of a novel class. We address this issue and propose a data stream classification technique that integrates a novel class detection mechanism into traditional classifiers, enabling automatic detection of novel classes before the true labels of the novel class instances arrive. Novel class detection problem becomes more challenging in the presence of concept-drift, when the underlying data distributions evolve in streams. In order to determine whether an instance belongs to a novel class, the classification model sometimes needs to wait for more test instances to discover similarities among those instances. A maximum allowable wait time Tcis imposed as a time constraint to classify a test instance. Furthermore, most existing stream classification approaches assume that the true label of a data point can be accessed immediately after the data point is classified. In reality, a time delay Tlis involved in obtaining the true label of a data point since manual labeling is time consuming. We show how to make fast and correct classification decisions under these constraints and apply them to real benchmark data. Comparison with state-of-the-art stream classification techniques prove the superiority of our approach. Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2011 | Enhanced geographically typed semantic schema matching
Jeffrey Partyka, Pallabi Parveen, Latifur Khan, Bhavani Thuraisingham, Shashi Shekhar 0001 |
J. Web Semant. | 4 |
| 2010 | Data Intensive Query Processing for Large RDF Graphs Using Cloud Computing ToolsabstractCloud computing is the newest paradigm in the IT world and hence the focus of new research. Companies hosting cloud computing services face the challenge of handling data intensive applications. Semantic web technologies can be an ideal candidate to be used together with cloud computing tools to provide a solution. These technologies have been standardized by the World Wide Web Consortium (W3C). One such standard is the Resource Description Framework (RDF). With the explosion of semantic web technologies, large RDF graphs are common place. Current frameworks do not scale for large RDF graphs. In this paper, we describe a framework that we built using Hadoop, a popular open source framework for Cloud Computing, to store and retrieve large numbers of RDF triples. We describe a scheme to store RDF data in Hadoop Distributed File System. We present an algorithm to generate the best possible query plan to answer a SPARQL Protocol and RDF Query Language (SPARQL) query based on a cost model. We use Hadoop's MapReduce framework to answer the queries. Our results show that we can store large RDF graphs in Hadoop clusters built with cheap commodity class hardware. Furthermore, we show that our framework is scalable and efficient and can easily handle billions of RDF triples, unlike traditional approaches. Mohammad Farhan Husain, Latifur Khan, Murat Kantarcioglu, Bhavani Thuraisingham |
IEEE CLOUD | 4 |
| 2010 | A Token-Based Access Control System for RDF Data in the CloudsabstractThe Semantic Web is gaining immense popularity-and with it, the Resource Description Framework (RDF)broadly used to model Semantic Web content. However, access control on RDF stores used for single machines has been seldom discussed in the literature. One significant obstacle to using RDF stores defined for single machines is their scalability. Cloud computers, on the other hand, have proven useful for storing large RDF stores, but these system slack access control on RDF data to our knowledge. This work proposes a token-based access control system that is being implemented in Hadoop (an open source cloud computing framework). It defines six types of access levels and an enforcement strategy for the resulting access control policies. The enforcement strategy is implemented at three levels: Query Rewriting, Embedded Enforcement, and Post processing Enforcement. In Embedded Enforcement, policies are enforced during data selection using MapReduce, whereas in Post-processing Enforcement they are enforced during the presentation of data to users. Experiments show that Embedded Enforcement consistently outperforms Post processing Enforcement due to the reduced number of jobs required. Arindam Khaled, Mohammad Farhan Husain, Latifur Khan, Kevin W. Hamlen, Bhavani Thuraisingham |
CloudCom | 5 |
| 2010 | Secure data storage and retrieval in the cloudabstractWith the advent of the World Wide Web and the emergence of e-commerce applications and social networks, organizations across the world generate a large amount of data daily. This data would be more useful to cooperating organizations if they were able to share their data. Two major obstacles to this Bhavani Thuraisingham, Vaibhav Khadilkar, Anuj Gupta 0002, Murat Kantarcioglu, Latifur Khan |
CollaborateCom | 1 |
| 2010 | Challenges and Future Directions of Software Technology: Secure Software DevelopmentabstractDeveloping large scale software systems has major security challenges. This paper describes the issues involved and then addresses two topics: formal methods for emerging secure systems and secure services modeling. Bhavani Thuraisingham, Kevin W. Hamlen |
COMPSAC | 1 |
| 2010 | Scalable and Efficient Reasoning for Enforcing Role-Based Access Control
Tyrone Cadenhead, Murat Kantarcioglu, Bhavani Thuraisingham |
DBSec | 3 |
| 2010 | Addressing Concept-Evolution in Concept-Drifting Data StreamsabstractThe problem of data stream classification is challenging because of many practical aspects associated with efficient processing and temporal behavior of the stream. Two such well studied aspects are infinite length and concept-drift. Since a data stream may be considered a continuous process, which is theoretically infinite in length, it is impractical to store and use all the historical data for training. Data streams also frequently experience concept-drift as a result of changes in the underlying concepts. However, another important characteristic of data streams, namely, concept-evolution is rarely addressed in the literature. Concept-evolution occurs as a result of new classes evolving in the stream. This paper addresses concept-evolution in addition to the existing challenges of infinite-length and concept-drift. In this paper, the concept-evolution phenomenon is studied, and the insights are used to construct superior novel class detection techniques. First, we propose an adaptive threshold for outlier detection, which is a vital part of novel class detection. Second, we propose a probabilistic approach for novel class detection using discrete Gini Coefficient, and prove its effectiveness both theoretically and empirically. Finally, we address the issue of simultaneous multiple novel class occurrence, and provide an elegant solution to detect more than one novel classes at the same time. We also consider feature-evolution in text data streams, which occurs because new features (i.e., words) evolve in the stream. Comparison with state-of-the-art data stream classification techniques establishes the effectiveness of the proposed approach. Mohammad M. Masud 0001, Latifur Khan, Charu C. Aggarwal, Jing Gao 0004, Jiawei Han 0001, Bhavani Thuraisingham |
ICDM | 7 |
| 2010 | Policy-Driven Service Composition with Information Flow ControlabstractEnsuring secure information flow is a critical task for service composition in multi-domain systems. Research in security-aware service composition provides some preliminary solutions to this problem, but there are still issues to be addressed. In this paper, we develop a service composition mechanism specifically focusing on the secure information flow control issues. We first introduce a general model for information flow control in service chains, considering the transformation factors of services and security classes of data resources in a service chain. Then, we develop general rules to guide service composition satisfying secure information flow requirements. Finally, to achieve efficient service composition, we develop a three-phase protocol to allow rapid filtering of candidate compositions that are unlikely to satisfy the information flow constraints and thorough evaluation of highly promising candidates. Our approach can achieve effective and efficient service composition considering secure information flow. Wei She, I-Ling Yen, Bhavani Thuraisingham, Elisa Bertino |
ICWS | 3 |
| 2010 | Classification and Novel Class Detection in Data Streams with Active Mining
Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
PAKDD (2) | 5 |
| 2010 | Classification and Novel Class Detection of Data Streams in a Dynamic Feature Space
Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
ECML/PKDD (2) | 6 |
| 2010 | Ranking Ontologies Using Verified Entities to Facilitate Federated QueriesabstractIn view of the need for highly distributed and federated architecture, ranking ontologies from different data sources in a specific domain have great impact on the performance of web applications. Since ontologies for a same domain usually overlap, we aim to rank ontologies based on the commonality of overlapping entities and distance between each pair of ontologies. Overlapping entities are determined by considering entities and relationships between them in ontology's graph. First we find out the Common Subset of Entities (CSE) between two ontologies using the lexical and structural similarity of entities. Second, we propose a novel strategy to find the similarity between ontologies by an Entropy Based Distribution (EBD) measurement. Third, we rank some synthetic and non-synthetic ontologies based on EBD values by naive and clustering approaches. Finally, we compare our approach with an existing approach and show effectiveness of our EBD approach. Ranking ontologies has great impact in different web scenarios including the federated query expansion and knowledge searching. Neda Alipanah, Piyush Srivastava 0002, Pallabi Parveen, Bhavani Thuraisingham |
Web Intelligence | 4 |
| 2010 | Policy Enforcement System for Inter-Organizational Data SharingabstractSharing data among organizations plays an important role in security and data mining. In this paper, the authors describe a Data Sharing Miner and Analyzer (DASMA) system that simulates data sharing among N organizations. Each organization has its own enforced policy. The N organizations share their data based on trusted third party. The system collects the released data from each organization, processes it, mines it, and analyzes the results. Sharing in DASMA is based on trusted third parties. However, organizations may encode some attributes, for example. Each organization has its own policy represented in XML format. This policy states what attributes can be released, encoded, and randomized. DASMA processes the data set and collects the data, combines it, and prepares it for mining. After mining, a statistical report is produced stating the similarities between mining with data sharing and mining without sharing. The authors test, apply data sharing, enforce policy, and analyze the results of two separate datasets in different domains. The results indicate a fluctuation on the amount of information loss using different releasing factors. Mamoun A. Awad, Latifur Khan, Bhavani Thuraisingham |
Int. J. Inf. Secur. Priv. | 3 |
| 2010 | Security Issues for Cloud ComputingabstractIn this paper, the authors discuss security issues for cloud computing and present a layered framework for secure clouds and then focus on two of the layers, i.e., the storage layer and the data layer. In particular, the authors discuss a scheme for secure third party publications of documents in a cloud. Next, the paper will converse secure federated query processing with map Reduce and Hadoop, and discuss the use of secure co-processors for cloud computing. Finally, the authors discuss XACML implementation for Hadoop and discuss their beliefs that building trusted applications from untrusted components will be a major aspect of secure cloud computing. Kevin W. Hamlen, Murat Kantarcioglu, Latifur Khan, Bhavani Thuraisingham |
Int. J. Inf. Secur. Priv. | 4 |
| 2010 | Secure Data Objects Replication in Data GridabstractSecret sharing and erasure coding-based approaches have been used in distributed storage systems to ensure the confidentiality, integrity, and availability of critical information. To achieve performance goals in data accesses, these data fragmentation approaches can be combined with dynamic replication. In this paper, we consider data partitioning (both secret sharing and erasure coding) and dynamic replication in data grids, in which security and data access performance are critical issues. More specifically, we investigate the problem of optimal allocation of sensitive data objects that are partitioned by using secret sharing scheme or erasure coding scheme and/or replicated. The grid topology we consider consists of two layers. In the upper layer, multiple clusters form a network topology that can be represented by a general graph. The topology within each cluster is represented by a tree graph. We decompose the share replica allocation problem into two subproblems: the Optimal Intercluster Resident Set Problem (OIRSP) that determines which clusters need share replicas and the Optimal Intracluster Share Allocation Problem (OISAP) that determines the number of share replicas needed in a cluster and their placements. We develop two heuristic algorithms for the two subproblems. Experimental studies show that the heuristic algorithms achieve good performance in reducing communication cost and are close to optimal solutions. Manghui Tu, Peng Li 0033, I-Ling Yen, Bhavani Thuraisingham, Latifur Khan |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2010 | Editorial SACMAT 2007abstractNo abstract available. Bhavani Thuraisingham |
ACM Trans. Inf. Syst. Secur. | 1 |
| 2009 | Assured information sharing between trustworthy, semi-trustworthy and untrustworthy coalition partnersabstractThere is a critical need for organizations to share data within and across infospheres and form coalitions so that analysts could examine the data, mine the data, and make effective decisions. Each organization could share information within its infosphere. An infosphere may consist of the data, applications and services that are needed for its operation. Organizations may share data with one another across what is called a global infosphere that spans multiple infospheres. It is critical that the war fighters get timely information. Furthermore, secure data and information sharing is an important requirement. The challenge is for data processing techniques to meet timing constraints and at the same time ensure that security is maintained. Bhavani Thuraisingham |
AsiaCCS | 1 |
| 2009 | Storage and Retrieval of Large RDF Graph Using Hadoop and MapReduce
Mohammad Farhan Husain, Pankil Doshi, Latifur Khan, Bhavani Thuraisingham |
CloudCom | 4 |
| 2009 | Dynamic Service and Data Migration in the CloudsabstractCloud computing is an emerging computation paradigm. To support successful cloud computing, service oriented architecture (SOA) should play a major role. Due to the nature of widely distributed service providers in clouds, the service performance could be impacted when the network traffic is congested. This can be a major barrier for tasks with real-time requirements. In clouds, this problem can be solved by migrating services to different platforms such that the communication cost can be minimized. In this paper, we consider the problem of service selection and migration in clouds. We develop a framework to facilitate service migration and design a cost model and the decision algorithm to determine the tradeoffs on service selection and migration. Wei Hao 0001, I-Ling Yen, Bhavani Thuraisingham |
COMPSAC (2) | 3 |
| 2009 | Geographically-typed semantic schema matchingabstractResolving semantic heterogeneity across distinct data sources remains a highly relevant problem in the GIS domain requiring innovative solutions. Our approach, called GSim, semantically aligns tables from respective GIS databases by first choosing attributes for comparison. We then examine their instances and calculate a similarity value between them called entropy-based distribution (EBD) by combining two separate methods. Our primary method discerns the geographic types from instances of compared attributes. If geographic type matching is not possible, we then apply a generic schema matching method which employs normalized Google distance. We show the effectiveness of our approach over the traditional N-gram approach across multi-jurisdictional datasets by generating impressive results. Jeffrey Partyka, Latifur Khan, Bhavani Thuraisingham |
GIS | 3 |
| 2009 | The SCIFC Model for Information Flow Control in Web Service CompositionabstractExisting Web service access control models focus on individual Web services, and do not consider service composition. In composite services, a major issue is information flow control. Critical information may flow from one service to another in a service chain through requests and responses and there is no mechanism for verifying that the flow complies with the access control policies. In this paper, we propose an innovative access control model to empower the services in a service chain to control the flow of their sensitive information. Our model supports information flow control through a back-check procedure and pass-on certificates. We also introduce additional factors such as the carry-along policy, security class, and transformation factor, to improve the protocol efficiency. A formal analysis is also presented to show the power and complexity of our protocol. Wei She, I-Ling Yen, Bhavani Thuraisingham, Elisa Bertino |
ICWS | 3 |
| 2009 | Assured Information Sharing Life CycleabstractThis paper describes our approach to assured information sharing. The research is being carried out under a MURI 9Multiuniversiyt Research Initiative) project funded by the Air Force Office of Scientific Research (AFOSR). The main objective of our project is: define, design and develop an Assured Information Sharing Lifecycle (AISL) that realizes the DoD's information sharing value chain. In this paper we describe the problem faced by the Department of Defense and our solution to developing an AISL System. Tim Finin, Anupam Joshi, Hillol Kargupta, Yelena Yesha, Joel Sachs, Elisa Bertino, Ninghui Li 0001, Chris Clifton, Eugene H. Spafford, Bhavani Thuraisingham, Murat Kantarcioglu, Alain Bensoussan 0001, Nathan Berg, Latifur Khan, Jiawei Han 0001, ChengXiang Zhai, Ravi S. Sandhu, Shouhuai Xu, Jim Massaro, Lada A. Adamic |
ISI | 10 |
| 2009 | Social network classification incorporating link type valuesabstractClassification of nodes in a social network and its applications to security informatics have been extensively studied in the past. However, previous work generally does not consider the types of links (e.g., whether a person is friend or a close friend) that connect social networks members for classification purposes. Here, we propose modified Naive Bayes Classification schemes to make use of the link type information in classification tasks. Basically, we suggest two new Bayesian classification methods that extend a traditional relational Naive Bayes Classifier, namely, the Link Type relational Bayes Classifier and the Weighted Link Type Bayes Classifier. We then show the efficacy of our proposed techniques by conducting experiments on data obtained from the Internet Movie Database. Raymond Heatherly, Murat Kantarcioglu, Bhavani Thuraisingham |
ISI | 3 |
| 2009 | On the mitigation of bioterrorism through game theoryabstractBioterrorism represents a serious threat to the security of civilian populations. The nature of an epidemic requires careful consideration of all possible vectors over which an infection can spread. Our work takes the SIR model and creates a detailed hybridization of existing simulations to allow a large search space to be explored. We then create a Stackelberg game to evaluate all possibilities with respect to the investment of available resources and consider the resulting scenarios. Our analysis of our experimental results yields the opportunity to place an upper bound on the worst case scenario for a population center in the event of an attack, with consideration of defensive and offensive measures. Ryan Layfield, Murat Kantarcioglu, Bhavani Thuraisingham |
ISI | 3 |
| 2009 | Design and implementation of a secure social network systemabstractContext-based anomaly tracking represents a new approach to security enhancement of communication streams. By creating a system that develops an understanding of normal and abnormal based on communication history, it is possible to detect fluctuations in an evolving social network. Although more research is necessary to overcome current obstacles, the combination of social network analysis and anomaly detection techniques yields a promising set of applications for enhancing communication security. In this paper we will describe a system for context-based anomaly detection and then describe experiments for message surveillance application. Ryan Layfield, Bhavani Thuraisingham, Latifur Khan, Murat Kantarcioglu, Jyothsna Rachapalli |
ISI | 2 |
| 2009 | A Multi-partition Multi-chunk Ensemble Technique to Classify Concept-Drifting Data Streams
Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
PAKDD | 5 |
| 2009 | Integrating Novel Class Detection with Classification for Concept-Drifting Data Streams
Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
ECML/PKDD (2) | 5 |
| 2009 | A semantic web based framework for social network access controlabstractThe existence of on-line social networks that include person specific information creates interesting opportunities for various applications ranging from marketing to community organization. On the other hand, security and privacy concerns need to be addressed for creating such applications. Improving social network access control systems appears as the first step toward addressing the existing security and privacy concerns related to on-line social networks. To address some of the current limitations, we propose an extensible fine grained access control model based on semantic web tools. In addition, we propose authorization, admin and filtering policies that depend on trust relationships among various users, and are modeled using OWL and SWRL. Besides describing the model, we present the architecture of the framework in its support. Barbara Carminati, Elena Ferrari 0001, Raymond Heatherly, Murat Kantarcioglu, Bhavani Thuraisingham |
SACMAT | 5 |
| 2009 | R2D: Extracting Relational Structure from RDF StoresabstractThe enthusiastic acceptance of Resource Description Framework (RDF) as a data model has given birth to a new data storage paradigm, namely, the RDF Graph model. The pool of modeling and visualization tools available for RDF stores is limited due to the technology being in its fledgling stage. The work presented in this paper, called R2D (RDF-to-Database) is an effort to make available, to RDF data stores, the abundance of relational tools that are currently in the market. This is done in the form of a JDBC wrapper around RDF Stores that presents a relational view of the stores and their data to the modeling and visualization tools. This paper presents key R2D functionalities and mapping constructs, procedures for every stage of R2D deployment, and sample results in the form of screenshots and performance graphs. Sunitha Ramanujam, Anubha Gupta, Latifur Khan, Steven Seida, Bhavani Thuraisingham |
Web Intelligence | 5 |
| 2009 | Data Mining for Malicious Code Detection and Security ApplicationsabstractData mining is the process of posing queries and extracting patterns, often previously unknown from large quantities of data using pattern matching or other reasoning techniques. Data mining has many applications in security including for national security as well as for cyber security. The threats to national security include attacking buildings, destroying critical infrastructures such as power grids and telecommunication systems. Data mining techniques are being investigated to find out who the suspicious people are and who is capable of carrying out terrorist activities. Cyber security is involved with protecting the computer and network systems against corruption due to Trojan horses, worms and viruses. Data mining is also being applied to provide solutions such as intrusion detection and auditing. The first part of the presentation will discuss my joint research with Prof. Latifur Khan and our students at the University of Texas at Dallas on data mining for cyber security applications For example; anomaly detection techniques could be used to detect unusual patterns and behaviors. Link analysis may be used to trace the viruses to the perpetrators. Classification may be used to group various cyber attacks and then use the profiles to detect an attack when it occurs. Prediction may be used to determine potential future attacks depending in a way on information learnt about terrorists through email and phone conversations. Data mining is also being applied for intrusion detection and auditing. Other applications include data mining for malicious code detection such as worm detection and managing firewall policies. This second part of the presentation will discuss the various types of threats to national security and describe data mining techniques for handling such threats. Threats include non real-time threats and real-time threats. We need to understand the types of threats and also gather good data to carry out mining and obtain useful results. The challenge is to reduce false positives and false negatives. The third part of the presentation will discuss some of the research challenges. We need some form of real-time data mining, that is, the results have to be generated in real-time, we also need to build models in real-time for realtime intrusion detection. Data mining is also being applied for credit card fraud detection and biometrics related applications. While some progress has been made on topics such as stream data mining, there is still a lot of work to be done here. Another challenge is to mine multimedia data including surveillance video. Finally, we need to maintain the privacy of individuals. Much research has been carried out on privacy preserving data mining. In summary, the presentation will provide an overview of data mining, the various types of threats and then discuss the applications of data mining for malicious code detection, cyber security and national security. Then we will discuss the consequences to privacy. Bhavani Thuraisingham |
Web Intelligence | 1 |
| 2009 | Inferring private information using social network dataabstractOn-line social networks, such as Facebook, are increasingly utilized by many users. These networks allow people to publish details about themselves and connect to their friends. Some of the information revealed inside these networks is private and it is possible that corporations could use learning algorithms on the released data to predict undisclosed private information. In this paper, we explore how to launch inference attacks using released social networking data to predict undisclosed private information about individuals. We then explore the effectiveness of possible sanitization techniques that can be used to combat such inference attacks under different scenarios. Jack Lindamood, Raymond Heatherly, Murat Kantarcioglu, Bhavani Thuraisingham |
WWW | 4 |
| 2009 | Relationalizing RDF stores for tools reusabilityabstractThe emergence of Semantic Web technologies and standards such as Resource Description Framework (RDF) has introduced novel data storage models such as the RDF Graph Model. In this paper, we present a research effort called R2D, which attempts to bridge the gap between RDF and RDBMS concepts by presenting a relational view of RDF data stores. Thus, R2D is essentially a relational wrapper around RDF stores that aims to make the variety of stable relational tools that are currently in the market available to RDF stores without data duplication and synchronization issues. Sunitha Ramanujam, Anubha Gupta, Latifur Khan, Steven Seida, Bhavani Thuraisingham |
WWW | 5 |
| 2009 | Privacy preservation in wireless sensor networks: A state-of-the-art survey
Na Li 0008, Nan Zhang 0004, Sajal K. Das 0001, Bhavani Thuraisingham |
Ad Hoc Networks | 4 |
| 2009 | Necessary and sufficient conditions for transaction-consistent global checkpoints in a distributed database system
D. Manivannan 0001, Bhavani Thuraisingham |
Inf. Sci. | 3 |
| 2008 | Privacy/Analysis Tradeoffs in Sharing Anonymized Packet Traces: Single-Field CaseabstractNetwork data needs to be shared for distributed security analysis. Anonymization of network data for sharing sets up a fundamental tradeoff between privacy protection versus security analysis capability. This privacy/analysis tradeoff has been acknowledged by many researchers but this is the first paper to provide empirical measurements to characterize the privacy/analysis tradeoff for an enterprise dataset. Specifically we perform anonymization options on single-fields within network packet traces and then make measurements using intrusion detection system alarms as a proxy for security analysis capability. Our results show: (1) two fields have a zero sum tradeoff (more privacy lessens security analysis and vice versa) and (2) eight fields have a more complex tradeoff (that is not zero sum) in which both privacy and analysis can both be simultaneously accomplished. William Yurcik, Clay Woolam, Greg Hellings, Latifur Khan, Bhavani Thuraisingham |
ARES | 5 |
| 2008 | Incentive and Trust Issues in Assured Information Sharing
Ryan Layfield, Murat Kantarcioglu, Bhavani Thuraisingham |
CollaborateCom | 3 |
| 2008 | Panel Session: What Are the Key Challenges in Distributed Security?
Steve Barker, David W. Chadwick, Jason Crampton, Emil C. Lupu, Bhavani Thuraisingham |
DBSec | 5 |
| 2008 | Content-based ontology matching for GIS datasetsabstractThe alignment of separate ontologies by matching related concepts continues to attract great attention within the database and artificial intelligence communities, especially since semantic heterogeneity across data sources remains a widespread and relevant problem. In particular, the Geographic Information System (GIS) domain presents unique forms of semantic heterogeneity that require a variety of matching approaches. Jeffrey Partyka, Neda Alipanah, Latifur Khan, Bhavani Thuraisingham, Shashi Shekhar 0001 |
GIS | 4 |
| 2008 | A Practical Approach to Classify Evolving Data Streams: Training with Limited Amount of Labeled DataabstractRecent approaches in classifying evolving data streams are based on supervised learning algorithms, which can be trained with labeled data only. Manual labeling of data is both costly and time consuming. Therefore, in a real streaming environment, where huge volumes of data appear at a high speed, labeled data may be very scarce. Thus, only a limited amount of training data may be available for building the classification models, leading to poorly trained classifiers. We apply a novel technique to overcome this problem by building a classification model from a training set having both unlabeled and a small amount of labeled instances. This model is built as micro-clusters using semi-supervised clustering technique and classification is performed with kappa-nearest neighbor algorithm. An ensemble of these models is used to classify the unlabeled data. Empirical evaluation on both synthetic data and real botnet traffic reveals that our approach, using only a small amount of labeled data for training, outperforms state-of-the-art stream classification algorithms that use twenty times more labeled data than our approach. Mohammad M. Masud 0001, Jing Gao 0004, Latifur Khan, Jiawei Han 0001, Bhavani Thuraisingham |
ICDM | 5 |
| 2008 | Enhancing Security Modeling for Web Services Using Delegation and Pass-OnabstractIn recent years, the issues in web service security have been widely investigated and various security standards have been proposed. But most of these studies and standards focus on the access control policies for individual web services and do not consider the access issues in composed services. Consider a simple service chain where service s1accesses s2, and s2, in turn, accesses service s3. The information returned from s3to s2may be used to compute some results that are further returned to s1. The current web service security framework does not provide any mechanisms to control such an information flow, and hence, sensitive information may be leaked to s1without the consensus of s3. In this paper, we propose an enhanced security model to facilitate the control of information flow through service chains. It extends the basic security models by introducing the concepts of delegation and pass-on. Based on these concepts, new certificates, certificate chain, delegation and pass-on policies, and how they are used to control the information flow are discussed. Wei She, I-Ling Yen, Bhavani Thuraisingham |
ICWS | 3 |
| 2008 | Detecting Remote Exploits Using Data Mining
Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham, Peng Liu 0005, Sencun Zhu |
IFIP Int. Conf. Digital Forensics | 3 |
| 2008 | Simulating bioterrorism through epidemiology approximationabstractBioterrorism represents a significant threat to society. The lack of successful attacks that have resulted in true epidemics have created a need for data that can be generated from existing known factors. We have taken the popular susceptible-infected-recovery model and created a hybridized model that balances the simplicity of the original with an approximation of what more complex agent-based models already offer, with an emphasis on the exploration of large search spaces. Our experiments focus on more unconventional methods of intervention in the event of an epidemic, the results of which suggest a much more basic approach can be taken beyond the thoroughly explored realm of inoculation strategies. Ryan Layfield, Murat Kantarcioglu, Bhavani Thuraisingham |
ISI | 3 |
| 2008 | Data mining for security applications: Mining concept-drifting data streams to detect peer to peer botnet trafficabstractThe presentation first provide an overview for data mining for security applications and then discuss our research to the botnet problem which follows from an important observation that network traffic (as well as botnet traffic) is a continuous flow of data stream. Conventional data mining techniques are not directly applicable to stream data because of two vital problems associated with them: potentially infinite in length, and concept drift. We propose a technique that can efficiently handle both problems. Our main focus is to adapt three major data mining techniques: classification, clustering, and outlier detection to handle stream data. Our preliminary study on the development of new stream classification techniques for P2P bothnet detection has generated encouraging results. In addition to botnet detection, we also discuss our research on data mining for malicious code detection and intrusion detection. Bhavani Thuraisingham |
ISI | 1 |
| 2008 | QoS Aware Dependable Distributed Stream ProcessingabstractIn this paper we describe our approach for developing a QoS-aware, dependable execution environment for large-scale distributed stream processing applications. Distributed stream processing applications have strong timeliness and security demands. In particular, we address the following challenges: (1) propose a real-time dependable execution model by extending the component-based execution model with real-time and dependability properties, and (2) develop QoS-aware application composition and adaptation techniques that employ resource management strategies and security policies when discovering and selecting application components. Our approach enables us to develop a distributed stream processing environment that is predictable, secure, flexible and adaptable. Vana Kalogeraki, Dimitrios Gunopulos, Ravi S. Sandhu, Bhavani Thuraisingham |
ISORC | 4 |
| 2008 | The SCRUB security data sharing infrastructureabstractSharing data between organization is important aspect of network protection that is not currently occurring since it is unsafe. This talk is about a suite of tools that can be used to "scrub" data (using anonymization) so it can be safely shared. SCRUB* is an infrastructure because all the tools use the same anonymization algorithms for seamless sharing independent of the data source. William Yurcik, Clay Woolam, Greg Hellings, Latifur Khan, Bhavani Thuraisingham |
NOMS | 5 |
| 2008 | Measuring anonymization privacy/analysis tradeoffs inherent to sharing network dataabstractSharing of network data between organizations is desperately needed as attackers bounce between targets in different security domains and launch attacks across security domains. Anonymization to protect private/sensitive information has emerged as a promising approach to sharing network data between security domains. However, a fundamental tradeoff exists between the anonymization of data for privacy protection and the utility of anonymized for security analysis. While many researchers have referred to this tradeoff, no one has characterized it with testing. In this paper we present a testing framework we have developed to characterize privacy/analysis anonymization tradeoffs along with some preliminary results. William Yurcik, Clay Woolam, Greg Hellings, Latifur Khan, Bhavani Thuraisingham |
NOMS | 5 |
| 2008 | ROWLBAC: representing role based access control in OWLabstractThere have been two parallel themes in access control research in recent years. On the one hand there are efforts to develop new access control models to meet the policy needs of real world application domains. In parallel, and almost separately, researchers have developed policy languages for access control. This paper is motivated by the consideration that these two parallel efforts need to develop synergy. A policy language in the abstract without ties to a model gives the designer little guidance. Conversely a model may not have the machinery to express all the policy details of a given system or may deliberately leave important aspects unspecified. Our vision for the future is a world where advanced access control concepts are embodied in models that are supported by policy languages in a natural intuitive manner, while allowing for details beyond the models to be further specified in the policy language. Tim Finin, Anupam Joshi, Lalana Kagal, Jianwei Niu 0001, Ravi S. Sandhu, William H. Winsborough, Bhavani Thuraisingham |
SACMAT | 7 |
| 2008 | An Effective Evidence Theory Based K-Nearest Neighbor (KNN) ClassificationabstractIn this paper, we study various K nearest neighbor (KNN) algorithms and present a new KNN algorithm based on evidence theory. We introduce global frequency estimation of prior probability (GE) and local frequency estimation of prior probability (LE). A GE for a class is the prior probability of the class across the whole training data space based on frequency estimation; on the other hand, a LE for a class in a particular neighborhood is the prior probability of the class in this neighborhood space based on frequency estimation. By considering the difference between the GE and the LE of each class, we present a solution to the imbalanced data problem in some degree without doing re-sampling. We compare our algorithm with other KNN algorithms using two benchmark datasets. Results show that our KNN algorithm outperforms other KNN algorithms, including basic evidence based KNN. Lei Wang 0021, Latifur Khan, Bhavani Thuraisingham |
Web Intelligence | 3 |
| 2008 | The applicability of the perturbation based privacy preserving data mining for real-world data
Murat Kantarcioglu, Bhavani Thuraisingham |
Data Knowl. Eng. | 3 |
| 2008 | Design and Implementation of a Framework for Assured Information Sharing Across Organizational BoundariesabstractIn this article we have designed and developed a framework for sharing data in an assured manner in case of emergencies. We focus especially on a need to share environment. It is often required to divulge information when an emergency is flagged and then take necessary steps to handle the consequences of divulging information. This procedure involves the application of a wide range of policies to determine how much information can be divulges in case of an emergency depending on how trustworthy the requester of the information is. Bhavani Thuraisingham, Yashaswini Harsha Kumar, Latifur Khan |
Int. J. Inf. Secur. Priv. | 1 |
| 2008 | Predicting WWW surfing using multiple evidence combination
Mamoun A. Awad, Latifur Khan, Bhavani Thuraisingham |
VLDB J. | 3 |
| 2007 | Fingerprint Matching Algorithm Based on Tree Comparison using Ratios of Relational DistancesabstractWe present a fingerprint matching algorithm that initially identifies the candidate common unique (minutiae) points in both the base and the input images using ratios of relative distances as the comparing function. A tree like structure is then drawn connecting the common minutiae points from bottom up in both the base and the input images. Matching score is obtained by comparing the similarity of the two tree structures based on a threshold value. We define a new term called the 'M(i)-tuple' for each minutiae point which uniquely encodes details about the local surrounding region, where i = 1 to N, and N is the number of minutiae. The proposed algorithm requires no explicit alignment of the two to-be compared fingerprint images and also tolerates distortions caused by spurious minutiae points. The algorithm is also capable of comparing and producing matching scores between two images obtained from two different kinds of sensors, hence is sensor interoperable and also reduces the FNMR in cases where there is very little overlap region between the base and the input image. We conducted evaluations on the FVC-2000 datasets and have summarized the results in the concluding section Abinandhan Chandrasekaran, Bhavani Thuraisingham |
ARES | 2 |
| 2007 | Extended RBAC - Based Design and Implementation for a Secure Data WarehouseabstractThis paper first discusses security issues for data warehousing. In particular, issues on building a secure data warehouse, secure data warehousing technologies as well as design issues are discussed. Our design of a secure data warehouse that enforces an extended RBAC policy is described next. Finally directions for secure data warehouses are discussed Bhavani Thuraisingham, Srinivasan Iyer 0003 |
ARES | 1 |
| 2007 | Centralized Security Labels in Decentralized P2P NetworksabstractThis paper describes the design of a peer-to-peer network that supports integrity and confidentiality labeling of shared data. A notion of data ownership privacy is also enforced, whereby peers can share data without revealing which data they own. Security labels are global but the implementation does not require a centralized label server. The network employs a reputation-based trust management system to assess and update data labels, and to store and retrieve labels safely in the presence of malicious peers. The security labeling scheme preserves the efficiency of network operations; lookup cost including label retrieval is O(log N), where N is the number of agents in the network. Nathalie Tsybulnik, Kevin W. Hamlen, Bhavani Thuraisingham |
ACSAC | 3 |
| 2007 | Secure peer-to-peer networks for trusted collaborationabstractAn overview of recent advances in secure peer-to-peer networking is presented, toward enforcing data integrity, confidentiality, availability, and access control policies in these decentralized, distributed systems. These technologies are combined with reputation-based trust management systems to enforce integrity-based discretionary access control policies. Particular attention is devoted to the problem of developing secure routing protocols that constitute a suitable foundation for implementing this security system. The research is examined as a basis for developing a secure data management system for trusted collaboration applications such as e-commerce, situation awareness, and intelligence analysis. Kevin W. Hamlen, Bhavani Thuraisingham |
CollaborateCom | 2 |
| 2007 | Enforcing Honesty in Assured Information Sharing Within a Distributed System
Ryan Layfield, Murat Kantarcioglu, Bhavani Thuraisingham |
DBSec | 3 |
| 2007 | Geospatial data qualities as web services performance metricsabstractService discovery is the crucial phase in the emerging Geospatial Semantic Web to select functionally similar services for the user query. Quality of Service (QoS) based service discovery, popularly studied in traditional Web Services, applies also to Geospatial Web Services. QoS allows service clients to fine-tune their search according to their specific needs and criteria. In high-performance service-based geospatial applications, it becomes an interesting research challenge to identify geospatial parameters to further improve the search process. In this paper we have proposed a set of geospatial criteria that can be used alongside the regular QoS parameters in service discovery and invocation. We show that using this novel approach of incorporating domain-specific drill-down information in addition to the commonly used QoS parameters yield more accurate and trustable Web services platform. We use the proposed geospatial parameters as performance metrics in the experimental evaluation of our application. The parameters reflect geospatial data quality attributes already standardized and well-studied in geospatial literature. Ganesh Subbiah, Latifur Khan, Bhavani Thuraisingham |
GIS | 4 |
| 2007 | A Hybrid Model to Detect Malicious ExecutablesabstractWe present a hybrid data mining approach to detect malicious executables. In this approach we identify important features of the malicious and benign executables. These features are used by a classifier to learn a classification model that can distinguish between malicious and benign executables. We construct a novel combination of three different kinds of features: binary n-grams, assembly n-grams, and library function calls. Binary features are extracted from the binary executables, whereas assembly features are extracted from the disassembled executables. The function call features are extracted from the program headers. We also propose an efficient and scalable feature extraction technique. We apply our model on a large corpus of real benign and malicious executables. We extract the above mentioned features from the data and train a classifier using support vector machine. This classifier achieves a very high accuracy and low false positive rate in detecting malicious executables. Our model is compared with other feature-based approaches, and found to be more efficient in terms of detection accuracy and false alarm rate. Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham |
ICC | 3 |
| 2007 | Geospatial Data Mining for National Security: Land Cover Classification and Semantic GroupingabstractLand cover classification for the evaluation of land cover changes over certain areas or time periods is crucial for geospatial modeling, environmental crisis evaluation and urban open space planning. Remotely sensed images of various spatial and spectral resolutions make it possible to classify land covers on the level of pixels. Semantic meanings of large regions consisting of hundreds of thousands of pixels cannot be revealed by discrete and individual pixel classes, but can be derived by integrating various groups of pixels using ontologies. This paper combines data of different resolutions for pixel classification by support vector classifiers, and proposes an efficient algorithm to group pixels based on classes of neighboring pixels. The algorithm is linear in the number of pixels of the target area, and is scalable to very large regions. It also re-evaluates imprecise classifications according to neighboring classes for region level semantic interpretations. Experiments on advanced spaceborne thermal emission and reflection radiometer (ASTER) data of more than six million pixels show that the proposed approach achieves up to 99.8% cross validation accuracy and 89.25% test accuracy for pixel classification, and can effectively and efficiently group pixels to generate high level semantic concepts. Chuanjun Li, Latifur Khan, Bhavani Thuraisingham, M. Husain, Shaofei Chen, Fang Qiu |
ISI | 3 |
| 2007 | Design of Secure CAMIN Application System Based on Dependable and Secure TMO and RT-UCONabstractIncreasingly the need for protecting information from unauthorized access has lead to more attention in the field of information security. Access control mechanisms have been in place for the last four decades and are a powerful tool utilized to ensure security. In a real-time distributed computing environment, systems have to meet timing constraints for time sensitive applications, such as obtaining financial quotes and operating command and control systems. However, real-time distributed computing environment security is yet to be investigated. Earlier work examined a Time Triggered Message Triggered Object (TMO) scheme that provides real-time services in a distributed computing environment secured by applying a Role-Based Access Control (RBAC). In this paper we describe the application of a sophisticated Usage Control (UCON) model for a TMO and subsequently design an application that utilizes a Coordinated Anti-Missile Interceptor Network (CAMIN) system. This application is a secure CAMIN based upon a secure TMO that applies a Real Time UCON (RT-UCON) in the CAMIN environment Jungin Kim, Bhavani Thuraisingham |
ISORC | 2 |
| 2007 | Feature Based Techniques for Auto-Detection of Novel Email Worms
Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham |
PAKDD | 3 |
| 2007 | E-Mail Worm Detection Using Data MiningabstractThis work applies data mining techniques to detect e-mail worms. E-mail messages contain a number of different features such as the total number of words in message body/subject, presence/absence of binary attachments, type of attachments, and so on. The goal is to obtain an efficient classification model based on these features. The solution consists of several steps. First, the number of features is reduced using two different approaches: feature-selection and dimension-reduction. This step is necessary to reduce noise and redundancy from the data. The feature-selection technique is called Two-phase Selection (TPS), which is a novel combination of decision tree and greedy selection algorithm. The dimension-reduction is performed by Principal Component Analysis. Second, the reduced data is used to train a classifier. Different classification techniques have been used, such as Support Vector Machine (SVM), Naïve Bayes, and their combination. Finally, the trained classifiers are tested on a dataset containing both known and unknown types of worms. These results have been compared with published results. It is found that the proposed TPS selection along with SVM classification achieves the best accuracy in detecting both known and unknown types of worms. Mohammad M. Masud 0001, Latifur Khan, Bhavani Thuraisingham |
Int. J. Inf. Secur. Priv. | 3 |
| 2007 | Administering the Semantic Web: Confidentiality, Privacy, and Trust ManagementabstractThe Semantic Web is essentially a collection of technologies to support machine-understandable Web pages as well as Information Interoperability. There has been much progress made on the Semantic Web, including standards for eXtensible Markup Language, Resource Description Framework, and Ontologies. However, administration policies and techniques for enforcing them have received little attention. These policies include policies for security, privacy, data quality, integrity, trust, and timely information processing. This article discusses administration policies for the Semantic Web as well as techniques for enforcing them. In particular, we will discuss an approach for ensuring confidentiality, privacy, and trust for the Semantic Web. We will also discuss the inference and privacy problems within the context of administration policies. Bhavani Thuraisingham, Natasha Tsybulnik |
Int. J. Inf. Secur. Priv. | 1 |
| 2007 | A framework for a video analysis tool for suspicious event detection
Gal Lavee, Latifur Khan, Bhavani Thuraisingham |
Multim. Tools Appl. | 3 |
| 2007 | Security and privacy for multimedia database management systems
Bhavani Thuraisingham |
Multim. Tools Appl. | 1 |
| 2007 | PP-trust-X: A system for privacy preserving trust negotiationsabstractTrust negotiation is a promising approach for establishing trust in open systems, in which sensitive interactions may often occur between entities with no prior knowledge of each other. Although, to date several trust negotiation systems have been proposed, none of them fully address the problem of privacy preservation. Today, privacy is one of the major concerns of users when exchanging information through the Web and thus we believe that trust negotiation systems must effectively address privacy issues in order to be widely applicable. For these reasons, in this paper, we investigate privacy in the context of trust negotiations. We propose a set of privacy-preserving features for inclusion in any trust negotiation system, such as the support for the P3P standard, as well as a number of innovative features, such as a novel format for encoding digital credentials specifically designed for preserving privacy. Further, we present a variety of interoperable strategies to carry on the negotiation with the aim of improving both privacy and efficiency. Anna Cinzia Squicciarini, Elisa Bertino, Elena Ferrari 0001, Federica Paci, Bhavani Thuraisingham |
ACM Trans. Inf. Syst. Secur. | 5 |
| 2007 | A new intrusion detection system using support vector machines and hierarchical clustering
Latifur Khan, Mamoun A. Awad, Bhavani Thuraisingham |
VLDB J. | 3 |
| 2006 | Detection and Resolution of Anomalies in Firewall Policy Rules
Muhammad Arshad Ul Abedin, Syeda Nessa, Latifur Khan, Bhavani Thuraisingham |
DBSec | 4 |
| 2006 | Face Recognition Using Multiple ClassifiersabstractIn this paper, we propose a near real-time effective face recognition system for consumer applications. Since the nature of application domain requires real time result and better accuracy, it poses a serious challenge. To address this challenge, we study various classification techniques, namely, support vector machine (SVM), linear discriminant analysis (LDA) and K nearest neighbor (KNN). We observe that although KNN is as effective as SVM but KNN prohibits its usage due to high response time when data is high dimensional. To speed up KNN retrieval, we propose a feature reduction technique using principle component analysis (PCA) to facilitate near real time face recognition along with better accuracy. We apply KNN after we reduce the number of features by PCA. Hence, we test various classification approaches, namely, SVM, KNN, KNN with PCA, LDA, and LDA with PCA on a benchmark dataset and demonstrate the effectiveness of KNN with PCA over SVM and LDA Pallabi Parveen, Bhavani Thuraisingham |
ICTAI | 2 |
| 2006 | Dependable and Secure TMO SchemeabstractIn a real-time distributed computing environment, security is critical to protect the system from unauthorized access especially since such systems are being used in time critical applications. Access control mechanisms have been introduced during the last several decades and have offered a basic and powerful means for enforcing security. In this paper, we examine the concepts of the TMO (time triggered message triggered object) scheme that provides guaranteed real-time services in a distributed object computing environment. We also examine access control mechanisms; such as the traditional model, the RBAC (role-based access control) model and the UCON (usage control) model. The main contribution of this paper is applying the traditional, RBAC and UCON models to the TMO scheme in order to provide a secure real-time distributed environment. Jungin Kim, Bhavani Thuraisingham |
ISORC | 2 |
| 2006 | Data Mining for Surveillance Applications
Bhavani Thuraisingham |
PAKDD | 1 |
| 2006 | Access control, confidentiality and privacy for video surveillance databasesabstractIn this paper we have addressed confidentiality and privacy for video surveillance databases. First we discussed our overall approach for suspicious event detection. Next we discussed an access control model and accedes control algorithms for confidentiality. Finally we discuss privacy preserving video surveillance. Our goal is build a comprehensive system that can detect suspicious events, ensure confidentiality as well as privacy. Bhavani Thuraisingham, Gal Lavee, Elisa Bertino, Jianping Fan 0001, Latifur Khan |
SACMAT | 1 |
| 2006 | Secure knowledge management: confidentiality, trust, and privacyabstractKnowledge management enhances the value of a corporation by identifying the assets and expertise as well as efficiently managing the resources. Security for knowledge management is critical as organizations have to protect their intellectual assets. Therefore, only authorized individuals must be permitted to execute various operations and functions in an organization. In this paper, secure knowledge management will be discussed, focusing on confidentiality, trust, and privacy. In particular, certain access-control techniques will be investigated, and trust management as well as privacy control for knowledge management will be explored Elisa Bertino, Latifur Khan, Ravi S. Sandhu, Bhavani Thuraisingham |
IEEE Trans. Syst. Man Cybern. Part A | 4 |
| 2006 | Guest editorial: special issue on privacy preserving data management
Elena Ferrari 0001, Bhavani Thuraisingham |
VLDB J. | 2 |
| 2005 | Trust Management in a Distributed EnvironmentabstractCybercrime as well as threats to national security are costing U.S. organizations billions of dollars each year. These organizations could be government organizations, financial corporations, medical hospitals and academic institutions. There is a critical need for organizations to share data within and across the organizations so that analysts could analyze the data, mine the data, and make effective decisions. Each organization could share information within the infosphere of that organization. An infosphere may consist of the data, applications and services that are needed for the operation of the organization. Organizations may share data with one another across what is called a global infosphere that spans multiple infospheres. While access control is an important security concern for organizational data sharing, managing trust is also an important consideration. For example, A may have the authorization to share the data with B, but A may not trust B. Trust management and negotiation has been studied extensively by Winslett et al. and Bertino et al. in the systems TrustBuilder and TrustX. In this paper we will discuss the issues on managing trust in a distributed environment. Much of the discussion is based on the work on secure knowledge management (Bertino et al. 2005). Bhavani Thuraisingham |
COMPSAC (1) | 1 |
| 2005 | Secure Model Management Operations for the Web
Guang-Lei Song, Kang Zhang 0001, Bhavani Thuraisingham |
DBSec | 3 |
| 2005 | Multilevel Secure Teleconferencing over Public Switched Telephone Network
Inja Youn, Csilla Farkas, Bhavani Thuraisingham |
DBSec | 3 |
| 2005 | Dependable Real-Time Data MiningabstractIn this paper we discuss the need for real-time data mining for many applications in government and industry and describe resulting research issues. We also discuss dependability issues including incorporating security, integrity, timeliness and fault tolerance into data mining. Several different data mining outcomes are described with regard to their implementation in a real-time environment. These outcomes include clustering, association-rule mining, link analysis and anomaly detection. The paper describes how they would be used together in various parallel-processing architectures. Stream mining is discussed with respect to the challenges of performing data mining on stream data from sensors. The paper concludes with a summary and discussion of directions in this emerging area. Bhavani Thuraisingham, Latifur Khan, Chris Clifton, John A. Maurer, Marion G. Ceruti |
ISORC | 1 |
| 2005 | Privacy constraint processing in a privacy-enhanced database management system
Bhavani Thuraisingham |
Data Knowl. Eng. | 1 |
| 2005 | Privacy-Preserving Data Mining: Development and DirectionsabstractThis article first describes the privacy concerns that arise due to data mining, especially for national security applications. Then we discuss privacy-preserving data mining. In particular, we view the privacy problem as a form of inference problem and introduce the notion of privacy constraints. We also describe an approach for privacy constraint processing and discuss its relationship to privacy-preserving data mining. Then we give an overview of the developments on privacy-preserving data mining that attempt to maintain privacy and at the same time extract useful information from data mining. Finally, some directions for future research on privacy as related to data mining are given. Bhavani Thuraisingham |
J. Database Manag. | 1 |
| 2004 | Security and Privacy for Web Databases and Services
Elena Ferrari 0001, Bhavani Thuraisingham |
EDBT | 2 |
| 2004 | Data mining for security applicationsabstractData mining is the process of posing queries and extracting patterns, often previously unknown from large quantities of data using pattern matching or other reasoning techniques. Cyber security is the area that deals with cyber terrorism. We are hearing that cyber attacks will cause corporations billions of dollars. For example, one could masquerade as a legitimate user and swindle say a bank of billions of dollars. Data mining and web mining may be used to detect and possibly prevent security attacks including cyber attacks. For example, anomaly detection techniques could be used to detect unusual patterns and behaviors. Link analysis may be used to trace the viruses to the perpetrators. Classification may be used to group various cyber attacks and then use the profiles to detect an attack when it occurs. Prediction may be used to determine potential future attacks depending in a way on information learnt about terrorists through email and phone conversations. Also, for some threats non real-time data mining may suffice while for certain other threats such as for network intrusions we may need real-time data mining. Many researchers are investigating the use of data mining for intrusion detection. While we need some form of real-time data mining, that is, the results have to be generated in real-time, we also need to build models in real-time. For example, credit card fraud detection is a form of real-time processing. However, here models are built ahead of time. Building models in real-time remains a challenge. Data mining can also be used for analyzing web logs as well as analyzing the audit trails. Based on the results of the data mining tool, one can then determine whether any unauthorized intrusions have occurred and/or whether any unauthorized queries have been posed. There has been much research on data mining for intrusion detection. Data mining may also be applied for Biometrics related applications. Finally data mining has applications in national security including detecting and preventing terrorist activities. The presentation will provide an overview of data mining and security threats and then discuss the applications of data mining for cyber security and national security including in intrusion detection and biometrics. Privacy considerations including a discussion of privacy preserving data mining will also be given. Bhavani Thuraisingham |
ICMLA | 1 |
| 2004 | Editorial
Bhavani Thuraisingham |
J. Intell. Inf. Syst. | 1 |
| 2004 | Selective and Authentic Third-Party Distribution of XML DocumentsabstractThird-party architectures for data publishing over the Internet today are receiving growing attention, due to their scalability properties and to the ability of efficiently managing large number of subjects and great amount of data. In a third-party architecture, there is a distinction between the Owner and the Publisher of information. The Owner is the producer of information, whereas Publishers are responsible for managing (a portion of) the Owner information and for answering subject queries. A relevant issue in this architecture is how the Owner can ensure a secure and selective publishing of its data, even if the data are managed by a third-party, which can prune some of the nodes of the original document on the basis of subject queries and access control policies. An approach can be that of requiring the Publisher to be trusted with regard to the considered security properties. However, the serious drawback of this solution is that large Web-based systems cannot be easily verified to be secure and can be easily penetrated. For these reasons, we propose an alternative approach, based on the use of digital signature techniques, which does not require the Publisher to be trusted. The security properties we consider are authenticity and completeness of a query response, where completeness is intended with regard to the access control policies stated by the information Owner. In particular, we show that, by embedding in the query response one digital signature generated by the Owner and some hash values, a subject is able to locally verify the authenticity of a query response. Moreover, we present an approach that, for a wide range of queries, allows a subject to verify the completeness of query results. Elisa Bertino, Barbara Carminati, Elena Ferrari 0001, Bhavani Thuraisingham, Amar Gupta |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2003 | Security Issues for the Semantic WebabstractThis paper first describes the developments in semantic Web and then provides an overview of secure semantic Web. In particular XML security, RDF security, and secure information integration and trust on the semantic Seb are discussed. Finally directions for research on secure semantic Web are provided. Bhavani Thuraisingham |
COMPSAC | 1 |
| 2003 | Data and Applications Security: Past, Present and the Future
Bhavani Thuraisingham |
DBSec | 1 |
| 2003 | Dependable Computing for National SecurityabstractThis paper examines various security threats and then discusses the applications of dependable computing for national security. By a dependable computing system we mean a flexible system that enforces quality of service policies between security, real-time processing, integrity and fault tolerance. We also discuss infrastructures and business objects for counter-terrorism. Bhavani Thuraisingham |
ISADS | 1 |
| 2002 | Data and Applications Security: Developments and DirectionsabstractThis paper first describes the developments in data and applications security with a special emphasis on database security. Then it discusses direction for data and applications security which includes secure semantic web, XML security, and security for emerging applications such as bioinformatics, peer-to-peer computing, and stream information management. Bhavani Thuraisingham |
COMPSAC | 1 |
| 2002 | Web and Information Security: Workshop Summary
Bhavani Thuraisingham, Elena Ferrari 0001 |
COMPSAC | 1 |
| 2002 | Privacy and Civil Liberties
David W. Chadwick, Martin S. Olivier, Pierangela Samarati, Eleanor Sharpston, Bhavani Thuraisingham |
DBSec | 5 |
| 2002 | Building Secure Survivable Semantic WebsabstractThis paper describes some ideas for a secure survivable semantic web that follows some of our previous ideas on dependable semantic web. Semantic web is a technology for understanding Web pages. It is important that the semantic web is secure. In addition, data exchanged by the Web has to be of high quality and survive failures and errors. The processes that the Web supports have to meet certain timing constraints. This paper discusses these aspects, and describes how they provide a dependable semantic web. Bhavani Thuraisingham |
ICTAI | 1 |
| 2002 | Data and applications security (Guest editorial)
Bhavani Thuraisingham, Reind P. van de Riet |
Data Knowl. Eng. | 1 |
| 2001 | Panel on XML and Security
Sylvia L. Osborn, Bhavani Thuraisingham, Pierangela Samarati |
DBSec | 2 |
| 2001 | Real-Time Data Mining of Multimedia ObjectsabstractWhereas much of the previous work on data mining has focused on mining data in relational databases, we discuss mining objects. Object models are very popular for representing multimedia data, and therefore we need to mine object databases to extract useful information from the large quantities of multimedia data. We first describe the motivation for multimedia data mining with examples and then discuss object mining with focus on text, image, video and audio mining. We also address the need for real time data mining for multimedia applications. Bhavani Thuraisingham, Chris Clifton, John A. Maurer, Marion G. Ceruti |
ISORC | 1 |
| 2001 | Security for Distributed Databases
Bhavani Thuraisingham |
Inf. Secur. Tech. Rep. | 1 |
| 2001 | Scheduling and Priority Mapping for Static Real-Time Middleware
Lisa Cingiser DiPippo, Victor Fay Wolfe, Levon Esibov, Gregory Cooper, Ramachandra Bethmangalkar, Russell Johnston, Bhavani Thuraisingham, John A. Maurer |
Real Time Syst. | 7 |
| 2000 | Network and Web Security and E-Commerce and Other Applications
Bhavani Thuraisingham |
COMPSAC | 1 |
| 2000 | Understanding Data Mining and applying it to Command, Control, Communications and Intelligence EnvironmentsabstractThe paper describes data mining and the various ways in which it can be applied to command, control, communications and intelligence. It begins with an overview of data mining and the various steps in the data mining process. Examples of data mining outcomes are cited and an example of a text mining scenario for intelligence applications is described. Problems associated with data mining are described, including the conflict between security and privacy and the challenges of distributed mining. The paper concludes with a discussion of future trends in data mining. Bhavani Thuraisingham, Marion G. Ceruti |
COMPSAC | 1 |
| 2000 | Web Security and Privacy (Panel)
Bhavani Thuraisingham |
DBSec | 1 |
| 2000 | Benchmarking Real-Time Distributed Object Management Systems for Evolvable and Adaptable Command and Control ApplicationsabstractThe paper describes benchmarking for evolvable and adaptable real time command and control systems. MITRE's Evolvable Real-Time C3 initiative developed an approach that would enable current real time systems to evolve into the systems of the future. We designed and implemented an infrastructure and data manager so that various applications could be hosted on the infrastructure. Then we completed a follow-on effort to design flexible adaptable distributed object management systems for command and control (C2) systems. Such an adaptable system would switch scheduling algorithms, policies, and protocols depending on the need and the environment. Both initiatives were carried out for the United States Air Force. One of the key contributions of the work is the investigation of real time features for distributed object management systems. Partly as a result of our work we are now seeing various real time distributed object management products being developed. In selecting a real time distributed object management systems, we need to analyze various criteria. Therefore, we need benchmarking studies for real time distributed object management systems. Although benchmarking systems such as Hartstone and Distributed Hartstone have been developed for middleware systems, these systems are not developed specifically for distributed object based middleware. Since much of our work is heavily based on distributed objects, we developed benchmarking systems by adapting the Hartstone system. The paper describes our efforts in developing benchmarks. Richard Freedman, John A. Maurer, Steven Wohlever, Bhavani Thuraisingham, Victor Fay Wolfe, Michael Milligan |
ISORC | 4 |
| 2000 | Fundamental R&D Issues in Real-Time Distributed Computing
Insup Lee 0001, Hermann Kopetz, K. H. (Kane) Kim, Thomas F. Lawrence, Bhavani Thuraisingham |
ISORC | 6 |
| 2000 | Real-Time CORBAabstractThis paper presents a survey of results in developing Real-Time CORBA, a standard for real-time management of distributed objects. This paper includes background on two areas that have been combined to realize Real-Time CORBA: the CORBA standards that have been produced by the international Object Management Group; and techniques for distributed real-time computing that have been produced in the research community. The survey describes major RT CORBA research efforts, commercial development efforts, and standardization efforts by the Object Management Group. Victor Fay Wolfe, Lisa Cingiser DiPippo, Gregory Cooper, Russell Johnston, Peter Kortmann, Bhavani Thuraisingham |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 1999 | Adaptable Real-Time Distributed Object Management for Command and Control SystemsabstractBetween now and the early part of the next century, significant portions of today's real-time command, control, and communication (C3) systems will become either functionally inadequate or logistically insupportable. Furthermore, due to the continuing budget reductions, new developments of next generation real-time C3 systems may not be possible. Therefore, current real-time C3 systems need to become easier, faster and less costly to upgrade in capability, and easier to support. What is needed is an approach to evolve current real-time C3 systems into the extensible systems required for the future. This paper describes our Adaptable Real-time Distributed Object Management (ARTDOM) approach for command and control systems. We have constructed a testbed environment and experimented with a Common Object Request Broker Architecture (CORBA) Object Request Broker (ORB). This project is a three year initiative and is expected to be completed at the end of 1999. John A. Maurer, Roman Ginis, Richard Freedman, Michael Squadrito, Steven Wohlever, Bhavani Thuraisingham |
ISADS | 6 |
| 1999 | Object Technology for Building Adaptable and Evolvable Autonomous Decentralized Systems
Bhavani Thuraisingham |
ISADS | 1 |
| 1999 | CORBA-based Real-time Trader Service for Adaptable Command and Control SystemsabstractThe paper describes an approach to building adaptable real time command and control (C2) systems. In particular it presents an overview of the Adaptable Real-Time Distributed Object Management (ARTDOM) project in progress at the MITRE Corporation. This project is currently developing real time extensions for the Common Object Request Broker Architecture (CORBA) Trading Object Service. The goal of the project is to demonstrate how current C2 systems can be more easily upgraded and made more adaptable by using emerging distributed computing technology. This goal is being accomplished by investigating and developing real time middleware that is reflexive (i.e., capable of examining its current state and processing demands) and self adapting (i.e., capable of reconfiguring itself based on its reflective findings). Steven Wohlever, Victor Fay Wolfe, Bhavani Thuraisingham, Richard Freedman, John A. Maurer |
ISORC | 3 |
| 1999 | Information Survivability for Evolvable and Adaptable Real-Time Command and Control SystemsabstractMITRE's Evolvable Real-Time C3 (command, control and communications) project has developed an approach that would enable current real-time systems to evolve into the systems of the future. This paper first summarizes the design and implementation of an infrastructure for an evolvable real-time C3 system. Then, a detailed discussion of the infrastructure requirements for a survivable real-time C3 system is presented. Finally, security issues for survivability, as well as open implementation of the infrastructure are described. In particular, adaptable middleware for survivable systems is discussed. Bhavani Thuraisingham, John A. Maurer |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1998 | Panel Overview
Bhavani Thuraisingham |
COMPSAC | 1 |
| 1998 | Security and Privacy Issues for the World Wide Web: Panel Discussion
Bhavani Thuraisingham, Sushil Jajodia, Pierangela Samarati, John E. Dobson, Martin S. Olivier |
DBSec | 1 |
| 1998 | Migrating Legacy Databases and Applications (Panel)
Bhavani Thuraisingham, Sandra Heiler, Arnon Rosenthal, Susan Malaika |
ICDE | 1 |
| 1998 | Concurrency Control in Real-Time Object-Oriented Systems: The Affected Set Priority Ceiling ProtocolsabstractThis paper presents the affected set priority ceiling protocols (ASPCP) for concurrency control in real-time object-oriented systems. These protocols are based on a combination of semantic locking and priority ceiling techniques. This paper shows that the ASPCP protocols provide higher potential concurrency for object-oriented systems than existing priority ceiling protocols, while still bounding priority inversion and preventing deadlock. Michael Squadrito, Levon Esibov, Lisa Cingiser DiPippo, Victor Fay Wolfe, Gregory Cooper, Bhavani Thuraisingham, Peter C. Krupp, Michael Milligan, Russell Johnston |
ISORC | 6 |
| 1997 | Security Issues in Data Warehousing and Data Mining: Panel Discussion
Bhavani Thuraisingham, Linda Schlipper, Pierangela Samarati, Tsau Young Lin, Sushil Jajodia, Chris Clifton |
DBSec | 1 |
| 1997 | Issues on real-time object request brokersabstractFor many applications such as command and control, telecommunications, and process control, it is critical that they meet timing constraints and ensure predictable computation. Furthermore, some of these applications are also distributed in nature and utilize distributed object management technology such as object request brokers (ORB) for interoperability. Therefore, it is necessary to integrate real time systems technology with distributed object management technology for many distributed real time applications. The article describes some of the issues that need to be investigated in order to develop real time data processing extensions to the common object request broker architecture (CORBA). In particular, issues on extensions to ORBs are given. A discussion of some real time services and applications hosted on real time ORBs is also provided. Finally, an overview of the work of the Object Management Group's Real time Special Interest Group is provided. Bhavani Thuraisingham |
ISADS | 1 |
| 1997 | Data Manager for Evolvable Real-time Command and Control Systems
Eric Hughes, Roman Ginis, Bhavani Thuraisingham, Peter C. Krupp, John A. Maurer |
VLDB | 3 |
| 1996 | Security Issues for Data Warehousing and Data Mining
Bhavani Thuraisingham |
DBSec | 1 |
| 1996 | Evolvable Real-Time C3 Systems-II: Real-Time Infrastructure RequirementsabstractMITRE's Evolvable Real-Time Command Control, and Communications (C3) project, funded under the Air Force Mission Oriented Investigation and Experimentation (MOIE) program attempts to develop an approach that would enable current real-time systems to evolve into the systems of the future. The project has chosen the Airborne Warning and Control System (AWACS) as an example to test the concepts and architectures to be developed. We discuss the requirements for the infrastructure for next generation complex real-time command and control systems. This discussion also includes an overview of the infrastructure requirements for each of the three architectures that we have considered. Bhavani Thuraisingham, Arkady Kanevsky, Peter C. Krupp, Alice Schafer, Mike Gates, Thomas Wheeler, Edward H. Bensley, Ruth Ann Sigel, Michael Squadrito |
ICECCS | 1 |
| 1996 | Guest Editors' Introduction to the Special Issue on Secure Database Systems TechnologyabstractINCE the U.S. Air Force Summer Study in 1982, several S research and development efforts in secure database management systems have been initiated. These include efforts in Secure Relational DBMS, Secure Object-Oriented DBMS, Secure Distributed IDBMS, and other topics such as inference and aggregation, policies and models, polyinstantiation, concurrency control, auditing, and role-based security. In addition to military applications, security for commercial applications such as medical information systems and banking systems have received increased attention in recent years. Since security is becoming increasingly important to many government as well as commercial organizations, and database technology is a necessity for these organizations, it is important for the various communities to be aware of the developments made in securing database systems. Due to the considerable pressing interest and concern in this area, this special issue of IEEE Transactions on Knowledge and Data Engineering is devoted to this topic. This issue consists of seven papers addressing a variety of topics in secure database systems technology. The first paper by Qian and Lunt describes a MAC policy framework for multilevel relational databases. Much of the work in multilevel secure database management systems has focussed on the relational model. Various prototype systems as well as commercial products have been developed. This paper presents a formal framework to specify mandatory access control policies for relational database systems. More recently quite a few efforts have been reported on multilevel secure object-oriented database management systems. One such effort is reported in the second paper by Thomas and Sandhu. ‘They propose a trusted subject architecture for designing multilevel secure objectoriented databases. Transaction processing in multilevel secure database management systems is a major issue. Concurrency control algorithms such as locking are known to cause covert channels. The goal in secure transaction processing is to ensure consistency as well as security. The third paper by Smith, Blaustein, . Jajodia, and Notargiacomo describes a Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1995 | Security and Data Mining
Tsau Young Lin, Thomas H. Hinke, Donald G. Marks, Bhavani Thuraisingham |
DBSec | 4 |
| 1995 | An Adaptive Policy for Improved Timeliness in Secure Database Systems
Sang Hyuk Son, Rasikan David, Bhavani Thuraisingham |
DBSec | 3 |
| 1995 | Evolvable real-time C3 systemsabstractThis paper describes MITRE's Evolvable Real-Time Command, Control, and Communications (C3) systems initiative that attempts to develop an approach that would enable current real-time systems to evolve into the systems of the future. In particular, this article describes the infrastructure requirements that we have developed. We first provide an overview of the current real-time C3 systems and describe the systems of the future. Next, we describe some candidate architectures that we have examined for future systems. Then a detailed discussion of the requirements for the infrastructure are given. The main focus is on operating systems, data management systems, and communication systems requirements. The discussion is based on the candidate architectures that we have examined. The project has chosen Airborne Warning and Control System (AWACS) as an example to test out the concepts and architectures to be developed. Edward H. Bensley, Lawrence Fisher, Mike Gates, James Houchens, Arkady Kanevsky, Soohee Kim, Peter C. Krupp, Alice Schafer, Bhavani Thuraisingham |
ICECCS | 9 |
| 1995 | Multilevel security for information retrieval systems - II
Bhavani Thuraisingham |
Inf. Manag. | 1 |
| 1995 | Security Constraints in a Multilevel Secure Distributed Database Management SystemabstractIn a multilevel secure distributed database management system, users cleared at different security levels access and share a distributed database consisting of data at different sensitivity levels. An approach to assigning sensitivity levels, also called security levels, to data is one which utilizes constraints or classification rules. Security constraints provide an effective classification policy. They can be used to assign security levels to the data based on content, context, and time. We extend our previous work on security constraint processing in a centralized multilevel secure database management system by describing techniques for processing security constraints in a distributed environment during query, update, and database design operations.> Bhavani Thuraisingham, William R. Ford |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1994 | Hypersemantic Data Modeling for Inference Analysis
Donald G. Marks, Leonard J. Binns, Bhavani Thuraisingham |
DBSec | 3 |
| 1994 | A Fine-grained Access Control Model for Object-Oriented DBMSs
Arnon Rosenthal, James G. Williams 0002, William R. Herndon, Bhavani Thuraisingham |
DBSec | 4 |
| 1994 | Security issues for federated database systems
Bhavani Thuraisingham |
Comput. Secur. | 1 |
| 1993 | Design and implementation of a distributed databaseabstractWe describe an approach for controlling certain unauthorized inferences in a multilevel secure distributed database management system. In such a system, two or more multilevel secure database management systems are connected via a trusted network. Furthermore, the environment that we have considered is a limited heterogeneous one where not all of the nodes handle the same accreditation ranges. In our approach, security constraints, which are rules that assign security levels to the data, are processed during the distributed query, update, and database design operations in such a way that users do not acquire information to which they are not authorized via logical deductions. We describe the design and implementation of the distributed inference controller which functions during the query operation.> Bhavani Thuraisingham, Harvey H. Rubinovitz, David Foti, Andres Abreu |
COMPSAC | 1 |
| 1993 | Applying OMT for Designing Multilevel Database Applications
Peter J. Sell, Bhavani Thuraisingham |
DBSec | 2 |
| 1993 | Workshop Summary
Bhavani Thuraisingham |
DBSec | 1 |
| 1993 | Secure computing with the actor paradigmabstractThis paper describes the a.ctor model of concurrent computation and discusses some of t,he issues in securing such a model. Bhavani Thuraisingham |
NSPW | 1 |
| 1993 | Integrating Object-Oriented Technolgy and Security Technology: A Panel DiscussionabstractNo abstract available. Bhavani Thuraisingham |
OOPSLA | 1 |
| 1993 | Design and Implementation of a Database Inference Controller
Bhavani Thuraisingham, William R. Ford, Marie Collins, J. O'Keeffe |
Data Knowl. Eng. | 1 |
| 1993 | Multilevel security for information retrieval systems
Bhavani Thuraisingham |
Inf. Manag. | 1 |
| 1993 | Simulation of join query processing algorithms for a trusted distributed database management system
Harvey H. Rubinovitz, Bhavani Thuraisingham |
Inf. Softw. Technol. | 2 |
| 1993 | Design and implementation of a query processor for a trusted distributed data base management system
Harvey H. Rubinovitz, Bhavani Thuraisingham |
J. Syst. Softw. | 2 |
| 1992 | A Non-monotonic Typed Multilevel Logic for Multilevel Secure Data/Knowledge - IIabstractFor pt.I. see Proc. 4th Computer Security Foundations, Franconia, USA (1991). In pt.I the author described a logic called nonmonotonic typed multilevel logic (NTML) for multilevel database applications. They also described various approaches to viewing multilevel databases through NTML. In this paper he continues with his discussion of the applications of NTML. In particular, the use of NTML as a programming language, issues on handling negative information in multilevel databases, and approaches for integrity checking in multilevel database systems are described. His work on NTML will be of significance to multilevel data/knowledge base applications in the same way logic programming has been to the development of data/knowledge base applications.> Bhavani Thuraisingham |
CSFW | 1 |
| 1992 | Multilevel security issues in distributed database management systems - III
Bhavani Thuraisingham, Harvey H. Rubinovitz |
Comput. Secur. | 1 |
| 1991 | Security constraint processing during the update operation in a multilevel secure database management systemabstractIn a multilevel secure database management system (MLS/DBMS), users cleared at different security levels access and share a database consisting of data at different sensitivity levels (also called security levels) to data is one which utilizes security constraints or classification rules. Security constraints provide an effective and versatile classification policy. They can be used to assign security levels to the data depending on the content, context, and time. Security constraints are a special form of integrity constraints enforced in a MLS/DBMS. As such, they can be handled during query processing, during database updates or during database design. The authors describe in detail the design and implementation of a secure update processor which handles security constraints in a multilevel secure database management system.> Marie Collins, William R. Ford, Bhavani Thuraisingham |
ACSAC | 3 |
| 1991 | A Nonmonotonic Typed Multilevel Logic for Multilevel Secure Database/Knowledge-Based Management SystemsabstractThe paper describes nonmonotonic typed multilevel logic (NTML) for multilevel database applications. It also describes various approaches to viewing multilevel databases through NTML and discusses techniques for query evaluation and integrity checking.> Bhavani Thuraisingham |
CSFW | 1 |
| 1991 | Multilevel security issues in distributed database management systems II
Bhavani Thuraisingham |
Comput. Secur. | 1 |
| 1990 | Secure query processing in distributed database management systems-design and performance studiesabstractDistributed systems are vital for the efficient processing required in military applications. For these applications it is especially important that the distributed database management systems (DDBMS) should operate in a secure manner. For example, the DDBMS should allow users, who are cleared to different levels, access to the database consisting of data at a variety of sensitivity levels without compromising security. The authors focus on secure query processing in a DDBMS. Implementation of secure query processing algorithms in a DDBMS as well as an analysis of the performance of the algorithms is described.> Bhavani Thuraisingham, Ammiel Kamon |
ACSAC | 1 |
| 1990 | Towards the Design of a Secure Data/Knowledge Base Management System
Bhavani Thuraisingham |
Data Knowl. Eng. | 1 |
| 1990 | Design of LDV: A Multilevel Secure Relational Database Management SystemabstractThe authors describe the design of a secure database system,LDV (Lock Data Views), that builds upon the classical security policies for operating systems. LDV is hosted on the LOgical Coprocessing Kernel (LOCK) Trusted Computing Base (TCB). LDVs security policy builds on the security policy of LOCK. Its design is based on three assured pipelines for the query, update, and metadata management operations. The authors describe the security policy of LDV, its system architecture, the designs of the query processor, the update processor, the metadata manager, and the operating system issues. LDVs solutions to the inference and aggregation problems are also described.> Paul D. Stachour, Bhavani Thuraisingham |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1989 | Secure query processing in intelligent database management systemsabstractIn a multilevel secure database management system, users cleared at different security levels access and share a database with data at different sensitivity levels. A serious threat to database security, not adequately addressed at present, is the inference problem. That is users acquire unauthorized information from the responses that they legitimately receive. It is the inferencing capability and deductive power of intelligent database systems that could provide viable approaches for handling the inference problem. The author describes secure query processing in intelligent database systems. The types of database systems that are considered are augmented relational systems, fuzzy relational systems and object-oriented systems.> Bhavani Thuraisingham |
ACSAC | 1 |
| 1989 | Knowledge-Based Support for the Development of Database-Centered ApplicationsabstractUsing the Application Development Toolkit (ADT) as an example, it is shown that by borrowing some techniques from the artificial-intelligence field, database-centered application-development productivity tools can be made more acceptable to end users and more useful to expert developers. Experience with ADT has indicated that a more end-user-oriented approach, and, in particular, more accommodating and application-oriented interface, is needed. A characteristic set of problems that are found in the class of productivity tools similar to ADT is presented. A series of possible improvements that shed some light on deficiencies in current state-of-the-art application-generation systems and productivity tools is suggested.> Hany M. Atchan, Rob Bell, Bhavani Thuraisingham |
ICDE | 3 |
| 1989 | Mandatory Security in Object-Oriented Database SystemsabstractA multilevel secure object-oriented data model (using the ORION data model) is proposed for which mandatory security issues in the context of a database system is discussed. In particular the following issues are dealt with: (1) the security policy for the system, (2) handling polyinstantiation, and (3) handling the inference problem. Bhavani Thuraisingham |
OOPSLA | 1 |
| 1989 | SODA: A secure object-oriented database system
Thomas F. Keefe, Wei-Tek Tsai, Bhavani Thuraisingham |
Comput. Secur. | 3 |
| 1989 | Prototyping to explore MLS/DBMS design
Dan Thomsen, Wei-Tek Tsai, Bhavani Thuraisingham |
Comput. Secur. | 3 |
| 1989 | A functional view of multilevel databases
Bhavani Thuraisingham |
Comput. Secur. | 1 |
| 1989 | Recovery Point Selection on a Reverse Binary Tree Task ModelabstractAn analysis is conducted of the complexity of placing recovery points where the computation is modeled as a reverse binary tree task model. The objective is to minimize the expected computation time of a program in the presence of faults. The method can be extended to an arbitrary reverse tree model. For uniprocessor systems, an optimal placement algorithm is proposed. For multiprocessor systems, a procedure for computing their performance is described. Since no closed form solution is available, an alternative measurement is proposed that has a closed form formula. On the basis of this formula, algorithms are devised for solving the recovery point placement problem. The estimated formula can be extended to include communication delays where the algorithm devised still applies.> Shyh-Kwei Chen, Wei-Tek Tsai, Bhavani Thuraisingham |
IEEE Trans. Software Eng. | 3 |
| 1988 | Multilevel security issues in distributed database management systems
John McHugh, Bhavani Thuraisingham |
Comput. Secur. | 2 |
| 1987 | Multilevel Security in Database Management Systems
Patricia A. Dwyer, George D. Jelatis, Bhavani Thuraisingham |
Comput. Secur. | 3 |
| 1987 | Security checking in relational database management systems augmented with inference engines
Bhavani Thuraisingham |
Comput. Secur. | 1 |
| 1982 | Representation of One-One Degrees by Decision Problems for System Functions
Bhavani Thuraisingham |
J. Comput. Syst. Sci. | 1 |