EDBT 2026 Demo / reviewers in the wild / expert
Noman Mohammed
dblp:94/1458
· DBLP profile ↗
44ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0001-8547-9951ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 14 · 6 first-author · 4 since 2021Security and privacy · 7 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 4 first-author · 2 since 2021Computer networks · 3 · 1 first-authorTheory of computation · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Assessing the Effectiveness of Singling-Out Attacks on Synthetic Data: Is the Current GDPR Guidance Adequate?abstractPrivacy remains one of the central challenges to data sharing under contemporary data protection regulations. To mitigate privacy risks, real datasets are either sanitized using anonymization techniques such as generalization or suppression, or replaced with artificially generated synthetic data. The European General Data Protection Regulation (GDPR) defines a singling-out attack as the ability to isolate specific records and identify an individual based on the content of a released dataset. Data controllers must ensure that published data does not allow singling out. However, this concept was originally designed for anonymized datasets, where a singled-out record corresponds to a real individual. In contrast, synthetic datasets contain artificially generated records, which may or may not represent or coincide with real individuals. This distinction raises questions about the applicability and effectiveness of singling-out attacks on synthetic data. To assess the effectiveness of singling-out attacks, we evaluated the attack on synthetic data generated by three well-known generative models (CTGAN, PATEGAN, and TabDDPM) across four real-world datasets. Our study reveals inherent limitations of conventional singling-out attacks when applied to synthetic data. As the GDPR’s privacy risk framework was originally developed with anonymized data in mind rather than synthetic data, there is a critical need for further research to establish robust evaluation frameworks and metrics specifically designed to assess the privacy risks associated with synthetic data. Fatima Jahan Sarmin, Atiquer Rahman Sarkar, Noman Mohammed |
TrustCom | 3 |
| 2025 | Robust privacy amidst innovation with large language models through a critical assessment of the risksabstractOBJECTIVE: This study evaluates the integration of electronic health records (EHRs) and natural language processing (NLP) with large language models (LLMs) to enhance healthcare data management and patient care, focusing on using advanced language models to create secure, Health Insurance Portability and Accountability Act-compliant synthetic patient notes for global biomedical research. MATERIALS AND METHODS: The study used de-identified and re-identified versions of the MIMIC III dataset with GPT-3.5, GPT-4, and Mistral 7B to generate synthetic clinical notes. Text generation employed templates and keyword extraction for contextually relevant notes, with One-shot generation for comparison. Privacy was assessed by analyzing protected health information (PHI) occurrence and co-occurrence, while utility was evaluated by training an ICD-9 coder using synthetic notes. Text quality was measured using ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and cosine similarity metrics to compare synthetic notes with source notes for semantic similarity. RESULTS: The analysis of PHI occurrence and text utility via the ICD-9 coding task showed that the keyword-based method had low risk and good performance. One-shot generation exhibited the highest PHI exposure and PHI co-occurrence, particularly in geographic location and date categories. The Normalized One-shot method achieved the highest classification accuracy. Re-identified data consistently outperformed de-identified data. DISCUSSION: Privacy analysis revealed a critical balance between data utility and privacy protection, influencing future data use and sharing. CONCLUSION: This study shows that keyword-based methods can create synthetic clinical notes that protect privacy while retaining data usability, potentially improving clinical data sharing. The use of dummy PHIs to counter privacy attacks may offer better utility and privacy than traditional de-identification. Yao-Shun Chuang, Atiquer Rahman Sarkar, Yu-Chun Hsu, Noman Mohammed, Xiaoqian Jiang |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Privacy-Preserving Learning via Data and Knowledge DistillationabstractIn the current era of data science, deep learning, computer vision and image analysis have become ubiquitous across various sectors, ranging from government agencies and large corporations to small end devices, due to their ability to simplify people’s lives. However, the widespread use of sensitive image data and the high memorization capacity of deep learning present significant privacy risks. Now, a simple Google search can yield numerous images of a person, and the knowledge that a specific patient’s record was utilized for training a specific model associated with a disease may reveal the patient’s ailment, potentially leading to membership privacy leakage and other advanced attacks in the future. Furthermore, these unprotected models may also suffer from poor generalization due to this overfitting to train data. Previous state-of-the-art methods like differential privacy (DP) and regularizer-based defenses compromised functionality, i.e., task accuracy, to preserve privacy. Such an imbalanced trade-off raises concerns about the practicability of such defenses. Other existing knowledge-transfer-based methods either reuse private data or require more public data, which could compromise privacy and may not be viable in certain domains. To address these challenges, where membership privacy is of utmost importance and utility cannot be compromised, we propose a novel collaborative distillation approach that transfers the private model’s knowledge based on a minimal amount of distilled synthetic data, leading to a compact private model in an end-to-end fashion. Empirically, our proposed method guarantees superior performance compared to most advanced models currently in use, increasing utility by almost 8%, 34%, and 6% for CIFAR-10, CIFAR-100, and MNIST, respectively. The utility resembles non-private counterparts almost closely while maintaining a respectable level of membership privacy leakage of 50-53.5%, despite employing a smaller model with 50% fewer parameters. Fahim Faisal, Carson K. Leung, Noman Mohammed, Yang Wang 0003 |
DSAA | 3 |
| 2023 | Comparative Analysis of Membership Inference Attacks in Federated LearningabstractGiven a federated learning model and a record, a membership inference attack can determine whether this record is part of the model’s training dataset. Federated learning is a machine learning technique that enables different parties to train a model without the need to centralize or share that data. Membership inference attack risks the private datasets if those datasets are used to train the federated learning model and access to the generated model is available. There is a need for further study in a federated learning environment to develop effective countermeasures against the membership inference attack without compromising the utility of the target model. In this study, we empirically investigated and compared various membership inference attack approaches in a federated learning environment. We also evaluated these attacks on several optimizers and analyzed them with and without countermeasures. Saroj Dayal, Dima Alhadidi, Ali Abbasi Tadi, Noman Mohammed |
IDEAS | 4 |
| 2023 | FedShare: Secure Aggregation based on Additive Secret Sharing in Federated LearningabstractFederated learning is a machine learning technique where multiple clients with local data collaborate in training a machine learning model. In FedAvg, the main federated learning algorithm, clients train machine learning models locally and share the trained model with the server. While the sensitive data will never be sent to the server, a malicious server can construct the original training data by having access to the clients’ models in each training round. Secure aggregation techniques such as cryptography, trusted execution environment, or differential privacy are used to solve this problem. However, these techniques incur computation and communication overhead or affect the model’s accuracy. In this paper, we consider a secure multi-party computation setup where clients use additive secret sharing to send their models to multiple servers. Our solution provides secure aggregation as long as there are at least two non-colluding servers. Moreover, we provide mathematical proof to show that the securely aggregated model at the end of each training round is exactly equal to the one provided by FedAvg without affecting accuracy and with efficient communication and computation. In comparison with SCOTCH, the state-of-the-art secure aggregation solution, experimental results show that our approach is 557% faster compared to SCOTCH and at the same time it reduces the communication cost of clients by 25%. Additionally, the accuracy of the trained model is exactly as FedAvg under balanced, unbalanced, IID, and Non-IID data distributions while it is only 8% slower. Hamid Fazli Khojir, Dima Alhadidi, Sara Rouhani, Noman Mohammed |
IDEAS | 4 |
| 2022 | Private Federated Framework for Health DataabstractFederated Learning (FL) is an efficient way to train Machine Learning (ML) algorithms on distributed datasets where data owners are restricted by policies to share their raw data. Through local training and model aggregation to a central server, this method reduces the need to communicate raw data with parties outside of the premises. However, FL raises serious privacy concerns. Therefore, additional privacy measures are required. The differential privacy (DP) approach is a cutting-edge privacy method used to perturb the local models prior to transmission and add an additional layer of privacy. However, this technique can affect the utility of the framework. To balance the privacy-utility trade-off, we implement a hybrid private technique to sanitize raw data using a combination of a top-down taxonomy tree and DP noise. The generalized data using DP noise is used to train local models to be shared in the FL architecture. The proposed framework achieves improved utility with a moderate privacy budget. Tanzir Ul Islam, Noman Mohammed, Dima Alhadidi |
BIBM | 2 |
| 2022 | Generating Privacy Preserving Synthetic Medical DataabstractDue to the recent development in the deep learning community and the availability of state-of-the-art models, medical practitioners are getting more interested in computer vision and deep learning for diagnosis tasks. Moreover, those medical diagnostic models can also increase the reliability of conventional findings. As radiology images can convey a lot of information for a patient’s diagnosis task, the problem is that such medical data may contain sensitive private information in their content header. De-anonymization (i.e., removal of sensitive header information) does not work well due to the re-identification risk, which may link those images to essential details (e.g., birth date, SSN, institution name, etc.), and such an approach can also reduce utility. In the medical domain, utility is significant because a less accurate diagnosis may lead to the wrong course of treatment and/or loss of life. In this paper, we developed a differentially private approach that can generate high-quality and high dimensional synthetic medical image data with guaranteed differential privacy. It can be used to create sufficient quality data to train a deep model. Moreover, we used W-GAN for bounded gradient guarantee, which eliminates the need for an extensive clipping hyperparameter search. We also added noise selectively to the generator to maintain the privacy-utility trade-off. Due to a noise-free discriminator and such selective noise addition to the generator, high-quality and reliable generated radiology images can be utilized for diagnosis tasks. Moreover, our approach can work in a distributed system where different hospitals can contain their private images in the local server and use a central server to generate synthetic radiology images without storing patient data. Fahim Faisal, Noman Mohammed, Carson K. Leung, Yang Wang 0003 |
DSAA | 2 |
| 2022 | Differentially Private Medical Texts Generation Using Generative Neural NetworksabstractTechnological advancements in data science have offered us affordable storage and efficient algorithms to query a large volume of data. Our health records are a significant part of this data, which is pivotal for healthcare providers and can be utilized in our well-being. The clinical note in electronic health records is one such category that collects a patient’s complete medical information during different timesteps of patient care available in the form of free-texts. Thus, these unstructured textual notes contain events from a patient’s admission to discharge, which can prove to be significant for future medical decisions. However, since these texts also contain sensitive information about the patient and the attending medical professionals, such notes cannot be shared publicly. This privacy issue has thwarted timely discoveries on this plethora of untapped information. Therefore, in this work, we intend to generate synthetic medical texts from a private or sanitized (de-identified) clinical text corpus and analyze their utility rigorously in different metrics and levels. Experimental results promote the applicability of our generated data as it achieves more than 80\% accuracy in different pragmatic classification problems and matches (or outperforms) the original text data. Md Momin Al Aziz, Tanbir Ahmed, Tasnia Faequa, Xiaoqian Jiang, Yiyu Yao, Noman Mohammed |
ACM Trans. Comput. Heal. | 6 |
| 2022 | Privacy preserving collaborative learning of generalized linear mixed model
Md. Monowar Anjum, Noman Mohammed, Xiaoqian Jiang |
J. Biomed. Informatics | 2 |
| 2022 | Generalized genomic data sharing for differentially private federated learning
Md Momin Al Aziz, Md. Monowar Anjum, Noman Mohammed, Xiaoqian Jiang |
J. Biomed. Informatics | 3 |
| 2021 | De-identification of Unstructured Clinical Texts from Sequence to Sequence PerspectiveabstractIn this work, we propose a novel problem formulation for de-identification of unstructured clinical text. We formulate the de-identification problem as a sequence to sequence learning problem instead of a token classification problem. Our approach is inspired by the recent state-of -the-art performance of sequence to sequence learning models for named entity recognition. Early experimentation of our proposed approach achieved 98.91% recall rate on i2b2 dataset. This performance is comparable to current state-of-the-art models for unstructured clinical text de-identification. Md. Monowar Anjum, Noman Mohammed, Xiaoqian Jiang |
CCS | 2 |
| 2021 | Online Algorithm for Differentially Private Genome-wide Association StudiesabstractDigitization of healthcare records contributed to a large volume of functional scientific data that can help researchers to understand the behaviour of many diseases. However, the privacy implications of this data, particularly genomics data, have surfaced recently as the collection, dissemination, and analysis of human genomics data is highly sensitive. There have been multiple privacy attacks relying on the uniqueness of the human genome that reveals a participant or a certain group’s presence in a dataset. Therefore, the current data sharing policies have ruled out any public dissemination and adopted precautionary measures prior to genomics data release, which hinders timely scientific innovation. In this article, we investigate an approach that only releases the statistics from genomic data rather than the whole dataset and propose a generalized Differentially Private mechanism for Genome-wide Association Studies (GWAS). Our method provides a quantifiable privacy guarantee that adds noise to the intermediate outputs but ensures satisfactory accuracy of the private results. Furthermore, the proposed method offers multiple adjustable parameters that the data owners can set based on the optimal privacy requirements. These variables are presented as equalizers that balance between the privacy and utility of the GWAS. The method also incorporates Online Bin Packing technique [1], which further bounds the privacy loss linearly, growing according to the number of open bins and scales with the incoming queries. Finally, we implemented and benchmarked our approach using seven different GWAS studies to test the performance of the proposed methods. The experimental results demonstrate that for 1,000 arbitrary online queries, our algorithms are more than 80% accurate with reasonable privacy loss and exceed the state-of-the-art approaches on multiple studies (i.e., EigenStrat, LMM, TDT). Md Momin Al Aziz, Shahin Kamali, Noman Mohammed, Xiaoqian Jiang |
ACM Trans. Comput. Heal. | 3 |
| 2020 | Nearest neighbour search over encrypted data using intel SGX
Kazi Wasif Ahmed, Md Momin Al Aziz, Md. Nazmus Sadat, Dima Alhadidi, Noman Mohammed |
J. Inf. Secur. Appl. | 5 |
| 2020 | sf SecDM: privacy-preserving data outsourcing framework with differential privacy
Gaby G. Dagher, Benjamin C. M. Fung, Noman Mohammed, Jeremy Clark |
Knowl. Inf. Syst. | 3 |
| 2019 | Privacy-preserving techniques of genomic data - a surveyabstractGenomic data hold salient information about the characteristics of a living organism. Throughout the past decade, pinnacle developments have given us more accurate and inexpensive methods to retrieve genome sequences of humans. However, with the advancement of genomic research, there is a growing privacy concern regarding the collection, storage and analysis of such sensitive human data. Recent results show that given some background information, it is possible for an adversary to reidentify an individual from a specific genomic data set. This can reveal the current association or future susceptibility of some diseases for that individual (and sometimes the kinship between individuals) resulting in a privacy violation. Regardless of these risks, our genomic data hold much importance in analyzing the well-being of us and the future generation. Thus, in this article, we discuss the different privacy and security-related problems revolving around human genomic data. In addition, we will explore some of the cardinal cryptographic concepts, which can bring efficacy in secure and private genomic data computation. This article will relate the gaps between these two research areas-Cryptography and Genomics. Md Momin Al Aziz, Md. Nazmus Sadat, Dima Alhadidi, Shuang Wang 0002, Xiaoqian Jiang, Cheryl L. Brown, Noman Mohammed |
Briefings Bioinform. | 7 |
| 2019 | SecureLR: Secure Logistic Regression Model via a Hybrid Cryptographic ProtocolabstractMachine learning applications are intensively utilized in various science fields, and increasingly the biomedical and healthcare sector. Applying predictive modeling to biomedical data introduces privacy and security concerns requiring additional protection to prevent accidental disclosure or leakage of sensitive patient information. Significant advancements in secure computing methods have emerged in recent years, however, many of which require substantial computational and/or communication overheads, which might hinder their adoption in biomedical applications. In this work, we propose SecureLR, a novel framework allowing researchers to leverage both the computational and storage capacity of Public Cloud Servers to conduct learning and predictions on biomedical data without compromising data security or efficiency. Our model builds upon homomorphic encryption methodologies with hardware-based security reinforcement through Software Guard Extensions (SGX), and our implementation demonstrates a practical hybrid cryptographic solution to address important concerns in conducting machine learning with public clouds. Jenny Hamer, Chenghong Wang, Xiaoqian Jiang, Miran Kim, Yongsoo Song, Yuhou Xia, Noman Mohammed, Md. Nazmus Sadat, Shuang Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2019 | SAFETY: Secure gwAs in Federated Environment through a hYbrid SolutionabstractRecent studies demonstrate that effective healthcare can benefit from using the human genomic information. Consequently, many institutions are using statistical analysis of genomic data, which are mostly based on genome-wide association studies (GWAS). GWAS analyze genome sequence variations in order to identify genetic risk factors for diseases. These studies often require pooling data from different sources together in order to unravel statistical patterns, and relationships between genetic variants and diseases. Here, the primary challenge is to fulfill one major objective: accessing multiple genomic data repositories for collaborative research in a privacy-preserving manner. Due to the privacy concerns regarding the genomic data, multi-jurisdictional laws and policies of cross-border genomic data sharing are enforced among different countries. In this article, we present SAFETY, a hybrid framework, which can securely perform GWAS on federated genomic datasets using homomorphic encryption and recently introduced secure hardware component of Intel Software Guard Extensions to ensure high efficiency and privacy at the same time. Different experimental settings show the efficacy and applicability of such hybrid framework in secure conduction of GWAS. To the best of our knowledge, this hybrid use of homomorphic encryption along with Intel SGX is not proposed to this date. SAFETY is up to 4.82 times faster than the best existing secure computation technique. Md. Nazmus Sadat, Md Momin Al Aziz, Noman Mohammed, Feng Chen 0016, Xiaoqian Jiang, Shuang Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Secure Similar Patients Query on Encrypted Genomic DataabstractBoth individuals and enterprises produce genomic data rapidly and continuously. There is a need to outsource such data to the cloud for better flexibility. Outsourcing also helps data owners by eliminating the local storage management problem. To protect data privacy and security, data owners must encrypt the sensitive data before outsourcing. Since genomic data are enormous in volume, executing researchers queries securely, and efficiently is a challenging task. In this paper, we introduce an indexing algorithm based on the prefix-tree to support similar patient queries. The proposed method guarantees the following: data privacy, query privacy, and output privacy. The privacy is guaranteed through encryption and garbled circuits considering the semi-honest adversary model. The overall computation is scalable and fast enough for real-life biomedical applications. Moreover, experimental results show that our method performs better than existing state-of-art techniques in this domain. Md Safiur Rahman Mahdi, Md Momin Al Aziz, Dima Alhadidi, Noman Mohammed |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Privacy-Preserving Age Estimation for Content RatingabstractContent rating (aka. maturity rating) rates the suitability of kinds of media (e.g., movies and video games) to its audience. It is essential to prevent a specific age group of people such as children from inappropriate information. However, in practice the administration of content rating system is usually suggestion-based declaration by media sources or key-based password which can easily fail if someone ignores the suggestions or somehow knows the keys. In this paper, we propose to estimate user's age in a privacy-preserving manner for automatic content rating. Several privacy-preserving approaches on facial images with different degree of privacy are proposed and evaluated on a deep neural network architecture for age estimation accuracy. We also introduce an attention mechanism which can adaptively learn discriminative features from the processed facial images. Experiments show that the proposed attention-based model performs better than the baseline model and achieves a reasonable performance to that with raw images in testing. Linwei Ye, Noman Mohammed, Yang Wang 0003, Jie Liang 0001 |
MMSP | 3 |
| 2018 | Parallel Linear Regression on Encrypted DataabstractIn recent years, the advent of machine learning models on private data has been remarkable. However, in-corporating machine learning techniques to healthcare data is pretty challenging due to the privacy issues of sensitive data which restricts data sharing in plaintext. Ensuring the privacy of individuals in healthcare datasets while constructing a machine learning model is a challenging research problem today. This paper proposes an approximate mathematical model utilizing linear regression on homomorphically encrypted data to predict the disease association of an individual. Furthermore, as these encryption schemes are not efficient considering computation time, we incorporate the multi-core parallelism to make the framework realistic. We experimentally evaluate the performance of the proposed methods and report on the experimental results. Toufique Morshed, Dima Alhadidi, Noman Mohammed |
PST | 3 |
| 2018 | Secure count query on encrypted genomic data
Mohammad Zahidul Hasan, Md Safiur Rahman Mahdi, Md. Nazmus Sadat, Noman Mohammed |
J. Biomed. Informatics | 4 |
| 2017 | SCOTCH: Secure Counting Of encrypTed genomiC data using a Hybrid approach
Chenghong Wang, Feng Chen 0016, Noman Mohammed, Xiaoqian Jiang, Md Momin Al Aziz, Md. Nazmus Sadat, Shuang Wang 0002 |
AMIA | 4 |
| 2017 | Image-Centric Social Discovery Using Neural Network under Anonymity ConstraintabstractImage sharing is one of the most attractive features facilitated by different social media sites such as Facebook, Flickr, Pinterest, and Instagram. People frequently use these social media sites to express various aspects of their life with peers they are connected through these sites. The service providers of these sites sometimes use the image features for social discovery such as friend recommendation, group or community recommendation, etc. As images are rich in content and more expressive, it also reveals much sensitive information about a user and impedes their privacy. Due to storage constraints, many popular social media sites prefer to outsource their data to the cloud server. However, if the cloud server gets compromised, then an adversary can use these sensitive images for malicious purposes. In this paper, we propose a privacy-preserving image-centric social discovery framework using the neural network and efficient anonymization scheme based on optimum feature selection. Experimental results show that our proposed approach provides better accuracy than existing method as well as is scalable for big datasets. Kazi Wasif Ahmed, Mohammad Zahidul Hasan, Noman Mohammed |
IC2E | 3 |
| 2017 | Private and Efficient Query Processing on Outsourced Genomic DatabasesabstractApplications of genomic studies are spreading rapidly in many domains of science and technology such as healthcare, biomedical research, direct-to-consumer services, and legal and forensic. However, there are a number of obstacles that make it hard to access and process a big genomic database for these applications. First, sequencing genomic sequence is a time consuming and expensive process. Second, it requires large-scale computation and storage systems to process genomic sequences. Third, genomic databases are often owned by different organizations, and thus, not available for public usage. Cloud computing paradigm can be leveraged to facilitate the creation and sharing of big genomic databases for these applications. Genomic data owners can outsource their databases in a centralized cloud server to ease the access of their databases. However, data owners are reluctant to adopt this model, as it requires outsourcing the data to an untrusted cloud service provider that may cause data breaches. In this paper, we propose a privacy-preserving model for outsourcing genomic data to a cloud. The proposed model enables query processing while providing privacy protection of genomic databases. Privacy of the individuals is guaranteed by permuting and adding fake genomic records in the database. These techniques allow cloud to evaluate count and top-k queries securely and efficiently. Experimental results demonstrate that a count and a top-k query over 40 Single Nucleotide Polymorphisms (SNPs) in a database of 20 000 records takes around 100 and 150 s, respectively. Reza Ghasemi, Md Momin Al Aziz, Noman Mohammed, Massoud Hadian Dehkordi, Xiaoqian Jiang |
IEEE J. Biomed. Health Informatics | 3 |
| 2016 | Secure and Efficient Multiparty Computation on Genomic DataabstractLarge scale biomedical research projects involve analysis of huge amount of genomic data which is owned by different data owners. The collection and storing of genomic data is sometimes beyond the capability of a sole organization. Genomic data sharing is a feasible solution to overcome this problem. These scenarios can be generalized into the problem of aggregating data distributed among multiple databases and owned by different data owners. However, we should guarantee that an adversary cannot learn anything about the data or the individual contribution of each party towards the final output of the computation. In this paper, we propose a practical solution for secure sharing and computation of genomic data. We adopt the Paillier cryptosystem and the order preserving encryption to securely execute the count query and the ranked query. Experimental results demonstrate that the computation time is realistic enough to make our system adoptable in the real world. Md Momin Al Aziz, Mohammad Zahidul Hasan, Noman Mohammed, Dima Alhadidi |
IDEAS | 3 |
| 2015 | Secure and Private Management of Healthcare Databases for Data MiningabstractThere has been a tremendous growth in health data collection since the development of Electronic Medical Record (EMR) systems. Such collected data is further shared and analyzed for diverse purposes. Despite many benefits, data collection and sharing have become a big concern as it threatens individual privacy. In this paper, we propose a secure and private data management framework that addresses both the security and privacy issues in the management of medical data in outsourced databases. The proposed framework ensures the security of data by using semantically-secure encryption schemes to keep data encrypted in outsourced databases. The framework also provides a differentially-private query interface that can support a number of SQL queries and complex data mining tasks. We experimentally evaluate the performance of the proposed framework, and the results show that the proposed framework is practical and has low overhead. Noman Mohammed, Samira Barouti, Dima Alhadidi, Rui Chen 0012 |
CBMS | 1 |
| 2014 | Secure Two-Party Differentially Private Data Release for Vertically Partitioned DataabstractPrivacy-preserving data publishing addresses the problem of disclosing sensitive data when mining for useful information. Among the existing privacy models, ϵ-differential privacy provides one of the strongest privacy guarantees. In this paper, we address the problem of private data publishing, where different attributes for the same set of individuals are held by two parties. In particular, we present an algorithm for differentially private data release for vertically partitioned data between two parties in the semihonest adversary model. To achieve this, we first present a two-party protocol for the exponential mechanism. This protocol can be used as a subprotocol by any other algorithm that requires the exponential mechanism in a distributed setting. Furthermore, we propose a two-party algorithm that releases differentially private data in a secure way according to the definition of secure multiparty computation. Experimental results on real-life data suggest that the proposed algorithm can effectively preserve information for a data mining task. Noman Mohammed, Dima Alhadidi, Benjamin C. M. Fung, Mourad Debbabi |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2013 | Privacy-preserving trajectory data publishing by local suppression
Rui Chen 0012, Benjamin C. M. Fung, Noman Mohammed, Bipin C. Desai, Ke Wang 0001 |
Inf. Sci. | 3 |
| 2013 | Privacy-preserving heterogeneous health data sharingabstractOBJECTIVE: Privacy-preserving data publishing addresses the problem of disclosing sensitive data when mining for useful information. Among existing privacy models, ε-differential privacy provides one of the strongest privacy guarantees and makes no assumptions about an adversary's background knowledge. All existing solutions that ensure ε-differential privacy handle the problem of disclosing relational and set-valued data in a privacy-preserving manner separately. In this paper, we propose an algorithm that considers both relational and set-valued data in differentially private disclosure of healthcare data. METHODS: The proposed approach makes a simple yet fundamental switch in differentially private algorithm design: instead of listing all possible records (ie, a contingency table) for noise addition, records are generalized before noise addition. The algorithm first generalizes the raw data in a probabilistic way, and then adds noise to guarantee ε-differential privacy. RESULTS: We showed that the disclosed data could be used effectively to build a decision tree induction classifier. Experimental results demonstrated that the proposed algorithm is scalable and performs better than existing solutions for classification analysis. LIMITATION: The resulting utility may degrade when the output domain size is very large, making it potentially inappropriate to generate synthetic data for large health databases. CONCLUSIONS: Unlike existing techniques, the proposed algorithm allows the disclosure of health data containing both relational and set-valued data in a differentially private manner, and can retain essential information for discriminative analysis. Noman Mohammed, Xiaoqian Jiang, Rui Chen 0012, Benjamin C. M. Fung, Lucila Ohno-Machado |
J. Am. Medical Informatics Assoc. | 1 |
| 2012 | Secure Distributed Framework for Achieving ε-Differential Privacy
Dima Alhadidi, Noman Mohammed, Benjamin C. M. Fung, Mourad Debbabi |
Privacy Enhancing Technologies | 2 |
| 2011 | Differentially private data release for data miningabstractPrivacy-preserving data publishing addresses the problem of disclosing sensitive data when mining for useful information. Among the existing privacy models, ∈-differential privacy provides one of the strongest privacy guarantees and has no assumptions about an adversary's background knowledge. Most of the existing solutions that ensure ∈-differential privacy are based on an interactive model, where the data miner is only allowed to pose aggregate queries to the database. In this paper, we propose the first anonymization algorithm for the non-interactive setting based on the generalization technique. The proposed solution first probabilistically generalizes the raw data and then adds noise to guarantee ∈-differential privacy. As a sample application, we show that the anonymized data can be used effectively to build a decision tree induction classifier. Experimental results demonstrate that the proposed non-interactive anonymization algorithm is scalable and performs better than the existing solutions for classification analysis. Noman Mohammed, Rui Chen 0012, Benjamin C. M. Fung, Philip S. Yu |
KDD | 1 |
| 2011 | Publishing Set-Valued Data via Differential Privacy
Rui Chen 0012, Noman Mohammed, Benjamin C. M. Fung, Bipin C. Desai, Li Xiong 0001 |
Proc. VLDB Endow. | 2 |
| 2011 | Mechanism Design-Based Secure Leader Election Model for Intrusion Detection in MANETabstractIn this paper, we study leader election in the presence of selfish nodes for intrusion detection in mobile ad hoc networks (MANETs). To balance the resource consumption among all nodes and prolong the lifetime of an MANET, nodes with the most remaining resources should be elected as the leaders. However, there are two main obstacles in achieving this goal. First, without incentives for serving others, a node might behave selfishly by lying about its remaining resources and avoiding being elected. Second, electing an optimal collection of leaders to minimize the overall resource consumption may incur a prohibitive performance overhead, if such an election requires flooding the network. To address the issue of selfish nodes, we present a solution based on mechanism design theory. More specifically, the solution provides nodes with incentives in the form of reputations to encourage nodes in honestly participating in the election process. The amount of incentives is based on the Vickrey, Clarke, and Groves (VCG) model to ensure truth-telling to be the dominant strategy for any node. To address the optimal election issue, we propose a series of local election algorithms that can lead to globally optimal election results with a low cost. We address these issues in two possible application settings, namely, Cluster-Dependent Leader Election (CDLE) and Cluster-Independent Leader Election (CILE). The former assumes given clusters of nodes, whereas the latter does not require any preclustering. Finally, we justify the effectiveness of the proposed schemes through extensive experiments. Noman Mohammed, Hadi Otrok, Lingyu Wang 0001, Mourad Debbabi, Prabir Bhattacharya |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2011 | Anonymity meets game theory: secure data integration with malicious participants
Noman Mohammed, Benjamin C. M. Fung, Mourad Debbabi |
VLDB J. | 1 |
| 2010 | A Secure Mechanism Design-Based and Game Theoretical Model for MANETs
Abderrezak Rachedi, Abderrahim Benslimane, Hadi Otrok, Noman Mohammed, Mourad Debbabi |
Mob. Networks Appl. | 4 |
| 2010 | Centralized and Distributed Anonymization for High-Dimensional Healthcare DataabstractSharing healthcare data has become a vital requirement in healthcare system management; however, inappropriate sharing and usage of healthcare data could threaten patients’ privacy. In this article, we study the privacy concerns of sharing patient information between the Hong Kong Red Cross Blood Transfusion Service (BTS) and the public hospitals. We generalize their information and privacy requirements to the problems of centralized anonymization and distributed anonymization , and identify the major challenges that make traditional data anonymization methods not applicable. Furthermore, we propose a new privacy model called LKC-privacy to overcome the challenges and present two anonymization algorithms to achieve LKC-privacy in both the centralized and the distributed scenarios. Experiments on real-life data demonstrate that our anonymization algorithms can effectively retain the essential information in anonymous data for data analysis and is scalable for anonymizing large datasets. Noman Mohammed, Benjamin C. M. Fung, Patrick C. K. Hung, Cheuk-kwong Lee |
ACM Trans. Knowl. Discov. Data | 1 |
| 2009 | Walking in the crowd: anonymizing trajectory data for pattern analysisabstractRecently, trajectory data mining has received a lot of attention in both the industry and the academic research. In this paper, we study the privacy threats in trajectory data publishing and show that traditional anonymization methods are not applicable for trajectory data due to its challenging properties: high-dimensional, sparse, and sequential. Our primary contributions are (1) to propose a new privacy model called LKC-privacy that overcomes these challenges, and (2) to develop an efficient anonymization algorithm to achieve LKC-privacy while preserving the information utility for trajectory pattern mining. Noman Mohammed, Benjamin C. M. Fung, Mourad Debbabi |
CIKM | 1 |
| 2009 | Privacy-preserving data mashupabstractMashup is a web technology that combines information from more than one source into a single web application. This technique provides a new platform for different data providers to flexibly integrate their expertise and deliver highly customizable services to their customers. Nonetheless, combining data from different sources could potentially reveal person-specific sensitive information. In this paper, we study and resolve a real-life privacy problem in a data mashup application for the financial industry in Sweden, and propose a privacy-preserving data mashup (PPMashup) algorithm to securely integrate private data from different data providers, whereas the integrated data still retains the essential information for supporting general data exploration or a specific data mining task, such as classification analysis. Experiments on real-life data suggest that our proposed method is effective for simultaneously preserving both privacy and information usefulness, and is scalable for handling large volume of data. Noman Mohammed, Benjamin C. M. Fung, Ke Wang 0001, Patrick C. K. Hung |
EDBT | 1 |
| 2009 | Anonymizing healthcare data: a case study on the blood transfusion serviceabstractSharing healthcare data has become a vital requirement in healthcare system management; however, inappropriate sharing and usage of healthcare data could threaten patients' privacy. In this paper, we study the privacy concerns of the blood transfusion information-sharing system between the Hong Kong Red Cross Blood Transfusion Service (BTS) and public hospitals, and identify the major challenges that make traditional data anonymization methods not applicable. Furthermore, we propose a new privacy model called LKC-privacy, together with an anonymization algorithm, to meet the privacy and information requirements in this BTS case. Experiments on the real-life data demonstrate that our anonymization algorithm can effectively retain the essential information in anonymous data for data analysis and is scalable for anonymizing large datasets. Noman Mohammed, Benjamin C. M. Fung, Patrick C. K. Hung, Cheuk-kwong Lee |
KDD | 1 |
| 2008 | A Mechanism Design-Based Multi-Leader Election Scheme for Intrusion Detection in MANETabstractIn this paper, we study the election of multiple leaders for intrusion detection in the presence of selfish nodes in mobile ad hoc networks (MANETs). To balance the resource consumption and prolong the lifetime of all nodes, each cluster should elect a node with the most remaining resources as its leader. However, without incentives for serving others, a node may behave selfishly by lying about its remaining resource and avoiding being elected. We present a solution based on mechanism design theory. More specifically, we design a scheme for electing cluster leaders that have the following two advantages: First, the collection of elected leaders is the optimal in the sense that the overall resource consumption will be balanced among all nodes in the network overtime. Second, the scheme provides the leaders with incentives in the form of reputation so that nodes are encouraged to honestly participate in the election process. The design of such incentives is based on the Vickrey, Clarke, and Groves (VCG) model by which truth-telling is the dominant strategy for each node. Simulation results show that our scheme can effectively prolong the overall lifetime of IDS in MANET and balance the resource consumptions among all the nodes. Noman Mohammed, Hadi Otrok, Lingyu Wang 0001, Mourad Debbabi, Prabir Bhattacharya |
WCNC | 1 |
| 2008 | A Moderate to Robust Game Theoretical Model for Intrusion Detection in MANETsabstractOne popular solution for reducing the resource consumption of intrusion detection system (IDS) in MANET is to elect a head-cluster (leader) to provide intrusion detection service to other nodes in the same cluster. However, such a moderate mode is only suitable when the probability of attack is low. Once the probability of attack is high, victim nodes should launch their own IDSs to detect and thwart intrusions. Such a robust mode is, however, costly with respect to energy and leads nodes to die faster. Clearly, to reduce the resource consumption of IDSs and yet keep its effectiveness, a critical issue is: when should we shift from moderate to robust mode? In this paper, we formalize this issue as a nonzero-sum noncooperative game theoretical model that takes into consideration the tradeoff between security and IDS resource consumption. The game solution will guide the leader-IDS to find the right moment for notifying the victim node to launch its IDS once the security risk is high enough. To achieve this goal, the Bayesian game theory is used to analyze the interaction between the leader-IDS and intruder with incomplete information about the intruder. By solving such a game, we are able to find the threshold value for notifying the victim node to launch its IDS once the probability of attack exceeds that value. Simulation results show that our scheme can effectively reduce the IDS resource consumption without sacrificing security. Hadi Otrok, Noman Mohammed, Lingyu Wang 0001, Mourad Debbabi, Prabir Bhattacharya |
WiMob | 2 |
| 2008 | A Mechanism Design-Based Secure Architecture for Mobile Ad Hoc NetworksabstractTo avoid the single point of failure for the certificate authority (CA) in MANET, a decentralized solution is proposed where nodes are grouped into different clusters. Each cluster should contain at least two confident nodes. One is known as CA and the another as register authority RA. The Dynamic Demilitarized Zone (DDMZ) is proposed as a solution for protecting the CA node against potential attacks. It is formed from one or more RA node. The problems of such a model are: (1) Clusters with one confident node, CA, cannot be created and thus clusters' sizes are increased which negatively affect clusters' services and stability. (2) Clusters with high density of RA can cause channel collision at the CA. (3) Clusters' lifetime are reduced since RA monitors are always launched (i.e., resource consumption). In this paper, we propose a model based on mechanism design that will allow clusters with single trusted node (CA) to be created. Our mechanism will motivate nodes that does not belong to the confident community to participate by giving them incentives in the form of trust, which can be used for cluster's services. To achieve this goal, a RA selection algorithm is proposed that selects nodes based on a predefined selection criteria function. Finally, empirical results are provided to support our solutions. Abderrezak Rachedi, Abderrahim Benslimane, Hadi Otrok, Noman Mohammed, Mourad Debbabi |
WiMob | 4 |
| 2008 | A game-theoretic intrusion detection model for mobile ad hoc networks
Hadi Otrok, Noman Mohammed, Lingyu Wang 0001, Mourad Debbabi, Prabir Bhattacharya |
Comput. Commun. | 2 |
| 2007 | An Efficient and Truthful Leader IDS Election Mechanism for MANET
Hadi Otrok, Noman Mohammed, Lingyu Wang 0001, Mourad Debbabi, Prabir Bhattacharya |
WiMob | 2 |