Justin Zhijun Zhan

dblp:127/0551 · also Justin Z. Zhan, Justin Zhan · DBLP profile ↗
← Back
58ranked-venue papers
20as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 15 · 4 first-author · 2 since 2021Security and privacy · 14 · 9 first-authorApplied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 1 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Distribution-enhanced knowledge distillation
Calvin Kinateder, Usman Anjum, Justin Zhijun Zhan
Neurocomputing3
2025 GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
abstract
We introduce GuessingGame, a protocol for evaluating large language models (LLMs) as strategic question-askers in open-ended, opendomain settings.A Guesser LLM identifies a hidden object by posing free-form questions to an Oracle without predefined choices or candidate lists.To measure question quality, we propose two information gain (IG) metrics: a Bayesian method that tracks belief updates over semantic concepts using LLM-scored relevance, and an entropy-based method that filters candidates via ConceptNet.Both metrics are model-agnostic and support post hoc analysis.Across 858 games with multiple models and prompting strategies, higher IG strongly predicts efficiency: a one-standard-deviation IG increase reduces expected game length by 43%.Prompting constraints guided by IG, such as enforcing question diversity, enable weaker models to significantly improve performance.These results show that question-asking in LLMs is both measurable and improvable, and crucial for interactive reasoning.
Dylan Hutson, Daniel Vennemeyer, Aneesh Deshmukh, Justin Zhijun Zhan, Tianyu Jiang 0001
EMNLP4
2025 Multimodal Entity Alignment Via Siamese Network and Structural Attention
abstract
We consider the problem of entity alignment in multi-modal knowledge graphs (MMKGs). The explosion in interest in MMKGs has led to a wide array of techniques used to align entities within multiple MMKG’s. The nature of the different modalities available in the MMKG makes the act of embedding them in a shared space challenging. This paper proposes using optimal transport (OT) as a means of embedding and aligning multiple modalities into one unified representation [1]. After acquiring the unified representation, we propose using contrastive learning to train a model for better performance in accurately predicting links between entities. Experiments looking at the effect of the additional modality and OT data fusion show poor performance failed to meet, let alone exceed, current state of the art.
Andrew Wagner, Usman Anjum, Justin Zhijun Zhan
SMC3
2024 Identifying and training deep learning neural networks on biomedical-related datasets
abstract
This manuscript describes the development of a resources module that is part of a learning platform named 'NIGMS Sandbox for Cloud-based Learning' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox at the beginning of this Supplement. This module delivers learning materials on implementing deep learning algorithms for biomedical image data in an interactive format that uses appropriate cloud resources for data access and analyses. Biomedical-related datasets are widely used in both research and clinical settings, but the ability for professionally trained clinicians and researchers to interpret datasets becomes difficult as the size and breadth of these datasets increases. Artificial intelligence, and specifically deep learning neural networks, have recently become an important tool in novel biomedical research. However, use is limited due to their computational requirements and confusion regarding different neural network architectures. The goal of this learning module is to introduce types of deep learning neural networks and cover practices that are commonly used in biomedical research. This module is subdivided into four submodules that cover classification, augmentation, segmentation and regression. Each complementary submodule was written on the Google Cloud Platform and contains detailed code and explanations, as well as quizzes and challenges to facilitate user training. Overall, the goal of this learning module is to enable users to identify and integrate the correct type of neural network with their data while highlighting the ease-of-use of cloud computing for implementing neural networks. This manuscript describes the development of a resource module that is part of a learning platform named ``NIGMS Sandbox for Cloud-based Learning'' https://github.com/NIGMS/NIGMS-Sandbox. The overall genesis of the Sandbox is described in the editorial NIGMS Sandbox [1] at the beginning of this Supplement. This module delivers learning materials on the analysis of bulk and single-cell ATAC-seq data in an interactive format that uses appropriate cloud resources for data access and analyses.
Alan E. Woessner, Usman Anjum, Hadi Salman, Jacob Lear, Jeffrey T. Turner, Ross Campbell, Laura Beaudry, Justin Zhijun Zhan, Lawrence E. Cornett, Susan Gauch, Kyle P. Quinn
Briefings Bioinform.8
2022 Cyberbully Detection Using BERT with Augmented Texts
abstract
Detecting cyberbullying in texts is an essential task as it curtails and identifies s ocial p roblems. I n t his p aper, we propose an architecture called Augmented BERT which combines both data augmentation techniques and BERT for detecting cyberbullying content in texts. Many techniques have been used in prior works to augment existing data for classification tasks and BERT had been applied in many text classification problems. However, there is a lack of annotated cyberbullying texts and obtaining annotated texts is hard and expensive. We propose to use various GAN-based and autoencoder-based data augmentation techniques to generate annotated data. The augmented texts can be used to fine-tune BERT. We choose to use HateBERT which is already pre-trained on abusive language to detect cyberbullying texts. Experimental results show an increased improvement over other cyberbullying detection models.
Usman Anjum, Justin Zhijun Zhan
IEEE Big Data3
2022 WaveNets: Wavelet Channel Attention Networks
abstract
Channel Attention reigns supreme as an effective technique in the field of computer vision. However, the proposed channel attention by SENet suffers from information loss in feature learning caused by the use of Global Average Pooling (GAP) to represent channels as scalars. Thus, designing effective channel attention mechanisms requires finding a s olution t o enhance features preservation in modeling channel inter-dependencies. In this work, we utilize Wavelet transform compression as a solution to the channel representation problem. We first t est wavelet transform as a standalone channel compression method. We prove that global average pooling is equivalent to the recursive approximate Haar wavelet transform. With this proof, we generalize channel attention using Wavelet compression and name it WaveNet. Implementation of our method can be embedded within existing channel attention methods with a couple of lines of code. We test our proposed method using ImageNet dataset for image classification t ask. O ur m ethod o utperforms t he baseline SENet-34, and SOTA FcaNet-34. Our code implementation is publicly available at https://github.com/hady1011/WaveNet-C.
Hadi Salman, Caleb Parks, Shi Yin Hong, Justin Zhijun Zhan
IEEE Big Data4
2021 Neighborhood Random Walk Graph Sampling for Regularized Bayesian Graph Convolutional Neural Networks
abstract
In the modern age of social media and networks, graph representations of real-world phenomena have become an incredibly useful source to mine insights. Often, we are interested in understanding how entities in a graph are interconnected. The Graph Neural Network (GNN) has proven to be a very useful tool in a variety of graph learning tasks including node classification, link prediction, and edge classification. However, in most of these tasks, the graph data we are working with may be noisy and may contain spurious edges. That is, there is a lot of uncertainty associated with the underlying graph structure. Recent approaches to modeling uncertainty have been to use a Bayesian framework and view the graph as a random variable with probabilities associated with model parameters. Introducing the Bayesian paradigm to graph-based models, specifically for semi-supervised node classification, has been shown to yield higher classification accuracies. However, the method of graph inference proposed in recent work does not take into account the structure of the graph. In this paper, we propose a novel algorithm called Bayesian Graph Convolutional Network using Neighborhood Random Walk Sampling (BGCN-NRWS), which uses a Markov Chain Monte Carlo (MCMC) based graph sampling algorithm utilizing graph structure, reduces overfitting by using a variational inference layer, and yields consistently competitive classification results compared to the state-of-the-art in semi-supervised node classification.
Aneesh Komanduri, Justin Zhijun Zhan
ICMLA2
2021 Modeling Cell Communication with Time-Dependent Signaling Hypergraphs
abstract
Signaling pathways describe a group of molecules in a cell that collaborate to control one or more cell functions, such as cell division or cell death. The pathways communicate by sending signals between molecules, and this process is repeated until the terminal molecule is activated and the cell function is executed. Signaling pathways are often represented as directed graphs, which does not provide enough information when modeling cell functions and reactions. Recently, directed hypergraphs have been proposed to more accurately represent reactions such as protein activation and interaction. To further improve the representation of signaling pathways, time dependency must be considered to improve the representation of cell signaling at any given time. In this paper, the importance of time dependency in modeling signaling pathways is presented. An algorithm that finds the shortest a priori path using time-dependent hypergraphs to more robustly model signaling pathways is adopted. The shortest time-dependent hyperpaths representing signaling pathways are an improvement to the recent adoption of hypergraphs representing these pathways. The results display the improved representation of signaling pathways and motivate the adoption of time-dependent signaling hypergraphs.
Michael R. Schwob, Justin Zhijun Zhan, Aeren Dempsey
IEEE ACM Trans. Comput. Biol. Bioinform.2
2020 A Cancelable Multi-Modal Biometric Based Encryption Scheme for Medical Images
abstract
In this paper, a novel multi-modal biometric based encryption scheme for medical images is proposed. It is based upon the ever-secure Advanced Encryption Standard Cipher Block Chaining (AES-CBC) mode of encryption and Indexing-First One (IFO) hashing for the generation of the secret keys. The proposed system utilizes two biometrics of the user, the iris and the fingerprint, but instead of using the features of the biometrics directly, they are hashed through the IFO process to create a secret key that can be revoked and regenerated in case of a compromise ensuring the security of the users biometrics. The IFO hashes obtained from the biometric feature vectors are then used as two different secret keys in a two round AES-CBC system to encrypt the medical image. The medical image can be decrypted by performing the AES-CBC rounds in reverse with the correct keys, meaning that with high probability, only the correct user will be able to decrypt the image. This encryption technique improves upon many existing medical encryption schemes based on biometrics, which do not take the protection of the biometric template into account. We were able to perform one full key generation, encryption and decryption in.027 seconds. In addition, this processes is lossless, meaning that there is no change in the pixels of the medical image during the encryption or decryption process, which is necessary of a medical image encryption system.
Alycia N. Carey, Justin Zhijun Zhan
IEEE BigData2
2020 Semi-Supervised Learning and Feature Fusion for Multi-view Data Clustering
abstract
Generative Adversarial Networks GANs have become widely used in Single-view classification tasks. Nowadays, most of the data have multiple views and each view emphasizes a unique feature set of the data. In this paper, we investigate the application of GANs on Multi-view data for the task of clustering and few-shot learning. We propose mvSGAN, a deep learning approach to GAN multi-view clustering, where generator and classifier networks are in a competitive min-max game. A multi-view learning algorithm is implemented with a mini-batch which can handle large data sets. We test the accuracy of our method in clustering real-world data sets. The experimental results show that our method outperforms state-of-the-art research.
Hadi Salman, Justin Zhijun Zhan
IEEE BigData2
2018 Prediction of online social networks users' behaviors with a game theoretic approach
abstract
As the number of users on social media platforms increases, a problem of great significance becomes clear: the increasing amount of negativity within the communities. We create a game that has two social media users, or players, trying to extend their outreach while keeping a positive community. Using five variables collected from 50 Instagram profiles, we create cost and payoff functions for the decision making part of the game. We implement those functions and make several discoveries: contrary to popular belief, the comments and negative comment percentages are not correlated; the number of followers directly impacted the number of likes on a user's posts; the cost function output increases with the payoff function output in a strong positive correlation; and that an increase in the number of followers number increase on an account results in an increase in number of comments.
Felix Zhan 0001, Gabriella Laines, Sarah Deniz, Sahan Paliskara, Irvin Ochoa, Idania Guerra, Shahab Tayeb, Carter Chiu, Matin Pirouz, Elliott Ploutz, Justin Zhijun Zhan, Laxmi P. Gewali, Paul Y. Oh
CCNC11
2018 An efficient alternative to personalized page rank for friend recommendations
abstract
In this paper, a new algorithm is proposed to calculate the value of friend score in a graph. The main goal is decreasing the running time and accuracy of calculation. Even though similar estimations have been done, the proposed method is faster because it increases the speed of computation with ignoring calculation of not involved nodes. This feature makes this method an acceptable candidate for large networks analyses like search engines and social media friend recommendation techniques. Based on the experiments done in this study, the FOF algorithm increases the speed of calculation by a factor of 28 on the two test datasets, namely, Amazon and Facebook.
Felix Zhan 0001, Brandon Waters, Maria Mijangos, LeAnn Chung, Raghav Bhagat, Tanvi Bhagat, Matin Pirouz, Carter Chiu, Shahab Tayeb, Elliott Ploutz, Justin Zhijun Zhan, Laxmi P. Gewali
CCNC11
2018 Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001
Knowl. Inf. Syst.5
2017 Toward data quality analytics in signature verification using a convolutional neural network
abstract
Many studies have been conducted on Handwritten Signature Verification. Researchers have taken many different approaches to accurately identify valid signatures from skilled forgeries, which closely resemble the real signature. The purpose of this paper is to suggest a method for validating written signatures on bank checks. This model uses a convolutional neural network (CNN) to analyze pixels from a signature image to recognize abnormalities. We believe the feature extraction capabilities of a CNN can optimize processing time and feature analysis of signature verification. Unique characteristics from signatures can be accurately and rapidly analyzed with multiple layers of receptive fields and hidden layers. Our method was able to correctly detect the validity of the inputted signature approximately 83 percent of the time. We tested our method using the SIGCOMP 2011 dataset. The main contribution of this method is to detect and decrease fraud committed, especially in the banking industry. Future uses of signature verification could include legal documents and the justice system.
Shahab Tayeb, Matin Pirouz, Brittany Cozzens, Richard Huang, Maxwell Jay, Kyle Khembunjong, Sahan Paliskara, Felix Zhan 0001, Mark Zhang, Justin Zhijun Zhan, Shahram Latifi
IEEE BigData10
2017 Securing the positioning signals of autonomous vehicles
abstract
One of the fastest growing industries in America is autonomous vehicle technology. The main motivation is to decrease the number of accidents each year. This leads to the challenge presented in this paper, preventing the spoofing of signals coming into autonomous vehicles. The increased complexity of these vehicles creates more vulnerabilities for attackers to take advantage of. Authentication of vehicular ad hoc networks is one method to help stop the potential hacking of autonomous vehicles. We propose a two-factor authentication for GPS signals by synchronizing with Stratum-1 clocks and digital signatures to prevent man-in-the-middle attacks. The experiment uses a computer to simulate a GPS signal that is sent to a Raspberry Pi 3 along with a timestamp and hashed key using RSA-1024. The Raspberry Pi 3 represents the vehicle. The method presented will prevent GPS spoofing attacks that are using modified or corrupted messages, an impersonation attack, or a roadside unit replication attack.
Shahab Tayeb, Matin Pirouz, Gabriel Esguerra, Kimiya Ghobadi, Jimson Huang, Robin Hill, Derwin Lawson, Stone Li, Tiffany Zhan, Justin Zhijun Zhan, Shahram Latifi
IEEE BigData10
2017 Toward predicting medical conditions using k-nearest neighbors
abstract
As the healthcare industry becomes more reliant upon electronic records, the amount of medical data available for analysis increases exponentially. While this information contains valuable statistics, the sheer volume makes it difficult to analyze without efficient algorithms. By using machine learning to classify medical data, diagnoses can become more efficient, accurate, and accessible for the public. After choosing k-Nearest Neighbors for its simplicity, we applied it to datasets compiled by the University of California, Irvine Machine Learning Repository to diagnose two conditions - chronic kidney failure and heart disease - with an accuracy of approximately 90%. In the future, similar methods can be used on a larger scale to bring ease of use to the field of medical diagnostics.
Shahab Tayeb, Matin Pirouz, Johann Sun, Kaylee Hall, Jessica Li, Connor Song, Apoorva Chauhan, Michael Ferra, Theresa Sager, Justin Zhijun Zhan, Shahram Latifi
IEEE BigData11
2017 Extracting recent weighted-based patterns from uncertain temporal databases
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Jimmy Ming-Tai Wu, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.6
2017 Mining of frequent patterns with multiple minimum supports
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.5
2017 Efficient hiding of confidential high-utility itemsets with minimal side effects
abstract
Privacy preserving data mining (PPDM) is an emerging research problem that has become critical in the last decades. PPDM consists of hiding sensitive information to ensure that it cannot be discovered by data mining algorithms. Several PPDM algorithms have been developed. Most of them are designed for hiding sensitive frequent itemsets or association rules. Hiding sensitive information in a database can have several side effects such as hiding other non-sensitive information and introducing redundant information. Finding the set of itemsets or transactions to be sanitised that minimises side effects is an NP-hard problem. In this paper, a genetic algorithm (GA) using transaction deletion is designed to hide sensitive high-utility itemsets for PPUM. A flexible fitness function with three adjustable weights is used to evaluate the goodness of each chromosome for hiding sensitive high-utility itemsets. To speed up the evolution process, the pre-large concept is adopted in the designed algorithm. It reduces the number of database scans required for verifying the goodness of an evaluated chromosome. Substantial experiments are conducted to compare the performance of the designed GA approach (with/without the pre-large concept), with a GA-based approach relying on transaction insertion and a non-evolutionary algorithm, in terms of execution time, side effects, database integrity and utility integrity. Results demonstrate that the proposed algorithm hides sensitive high-utility itemsets with fewer side effects than previous studies, while preserving high database and utility integrity.
Jerry Chun-Wei Lin, Tzung-Pei Hong, Philippe Fournier-Viger, Qiankun Liu 0002, Jia-Wei Wong, Justin Zhijun Zhan
J. Exp. Theor. Artif. Intell.6
2017 An ACO-based approach to mine high-utility itemsets
Jimmy Ming-Tai Wu, Justin Zhijun Zhan, Jerry Chun-Wei Lin
Knowl. Based Syst.2
2017 A Framework for Community Detection in Large Networks Using Game-Theoretic Modeling
abstract
Community detection is a fundamental component of large network analysis. In both academia and industry, progressive research has been made on problems related to community network analysis. Community detection is gaining significant attention and importance in the area of network science. Regular and synthetic complex networks have motivated intense interest in studying the fundamental unifying principles of various complex networks. This paper presents a new game-theoretic approach towards community detection in large-scale complex networks based on modified modularity; this method was developed based on modified adjacency, modified Laplacian matrices and neighborhood similarity. This approach was used to partition a given network into dense communities. It is based on determining a Nash stable partition, which is a pure strategy Nash equilibrium of an appropriately defined strategic game in which the nodes of the network were the players and the strategy of a node was to decide to which community it ought to belong. Players chose to belong to a community according to a maximized fitness/payoff. Quality of the community networks was assessed using modified modularity along with a new fitness function. Community partitioning was performed using Normalized Mutual Information and a `modularity measure', which involved comparing the new game-theoretic community detection algorithm (NGTCDA) with well-studied and well-known algorithms, such as Fast Newman, Fast Modularity Detection, and Louvain Community. The quality of a network partition in communities was evaluated by looking at the contribution of each node and its neighbors against the strength of its community.
Pravin Chopade, Justin Zhijun Zhan
IEEE Trans. Big Data2
2016 An efficient algorithm to mine high average-utility itemsets
Jerry Chun-Wei Lin, Ting Li 0011, Philippe Fournier-Viger, Tzung-Pei Hong, Justin Zhijun Zhan, Miroslav Voznak
Adv. Eng. Informatics5
2016 A sanitization approach for hiding sensitive itemsets based on particle swarm optimization
Jerry Chun-Wei Lin, Qiankun Liu 0002, Philippe Fournier-Viger, Tzung-Pei Hong, Miroslav Voznak, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.6
2016 Fast algorithms for hiding sensitive high-utility itemsets in privacy-preserving utility mining
Jerry Chun-Wei Lin, Tsu-Yang Wu, Philippe Fournier-Viger, Justin Zhijun Zhan, Miroslav Voznak
Eng. Appl. Artif. Intell.5
2016 Mining high-utility itemsets based on particle swarm optimization
Jerry Chun-Wei Lin, Philippe Fournier-Viger, Jimmy Ming-Tai Wu, Tzung-Pei Hong, Shyue-Liang Wang, Justin Zhijun Zhan
Eng. Appl. Artif. Intell.7
2016 Efficient mining of high-utility itemsets using multiple minimum utility thresholds
Jerry Chun-Wei Lin, Wensheng Gan, Philippe Fournier-Viger, Tzung-Pei Hong, Justin Zhijun Zhan
Knowl. Based Syst.5
2014 An Information-Theoretic Approach for Secure Protocol Composition
Yi-Ting Chiang, Tsan-sheng Hsu, Churn-Jung Liau, Yun-Ching Liu, Chih-Hao Shen, Dawei Wang 0004, Justin Zhijun Zhan
SecureComm (1)7
2012 An artificial immune system for phishing detection
abstract
The amazing nature of biological immune systems on protecting humans from pathogens inspired people to develop artificial immune systems. Designed to simulate the functionalities of biological immune systems, artificial immune systems are suggested to be mainly applied in the domain of computer security. In this paper, we propose an artificial immune system for phishing detection. The system is to detect phishing emails through memory detectors and mature detectors. The memory detectors are generated from the training data set, which, in turn, contains the phishing emails previously seen by the system. The immature detectors are reproduced through the system's mutation process. To the best of our knowledge, this is the first time such a system is ever proposed. We believe that the system is more adaptive than any other existing phishing detection techniques.
Nicholas Koceja, Justin Zhijun Zhan, Gerry V. Dozier, Dipankar Dasgupta
IEEE Congress on Evolutionary Computation3
2011 Location privacy protection on social networks
abstract
Location information is considered as private in many scenarios. Location information protection on social networks has not been paid much attention. In this paper, we extend our previous proposed location privacy protection approach on the basis of user messages in social networks. Our approach grants flexibility to users by offering them multiple protecting options. The extension includes performance evaluation towards our approach.
Justin Zhijun Zhan, Naveen Bandaru
CICS1
2011 Trust optimization in task-oriented social networks
abstract
Trust is a human-related phenomenon in social networks. Trust research on social networks has gained much attention on its usefulness, and on modeling propagations. There is little focus on finding maximum trust in social networks which is particularly important when a social network is oriented by certain tasks. In this paper, we first propose a trust maximization algorithm based on the task-oriented social networks. We then take communication cost into account and introduce four different trust optimization algorithms. We also conduct extensive experiments to evaluate the proposed algorithms and test their performance. To our best knowledge, this is pioneering work on trust optimization in task-oriented social networks.
Justin Zhijun Zhan, Peter Killion
CICS1
2011 Phishing detection using stochastic learning-based weak estimators
abstract
Phishing attack has been a serious concern to online banking and e-commerce Websites. This paper proposes a method to detect and filter phishing emails in dynamic environment by applying a family of weak estimators. Anomaly detection identifies observations that deviate from the normal behavior of a system and is achieved by identifying the phenomena that characterize the “normal” observation. The new observations are classified either a normal or abnormal based on the characteristics of data learnt. Most of the anomaly detection works with the assumption that the underlying distributions of observations are stationary, where this assumption is relevant to many applications. However some detection problem occurs within environments that are non-stationary. One good example to demonstrate the information is by identifying anomalous temperature pattern in meteorology that takes into account the seasonal changes of normal observations. It is necessary that anomalous observations are identified even with the changes or acquire the ability to adapt to the variations in non-stationary environments. Our experimental results show the feasibility and effectiveness of our approach.
Justin Zhijun Zhan, Lijo Thomas
CICS1
2011 Using gaming strategies for attacker and defender in recommender systems
abstract
Ratings are the prominent factors to decide the fate of any product in the present Internet Market and many people follow the ratings in a genuine sense. Unfortunately, the Sibyl attacks can affect the credibility of the genuine product. Influence limiter algorithms in recommender systems have been used extensively to overcome the Sibyl attacks but the effort could not reach the safe mark. This paper highlights an approach to generating gaming strategies for the attacker and defender in a recommender system. In a given recommender system environment, attackers and defenders play the most crucial part in a gaming strategy. A sequence of decision rules that an attacker or defender may use to achieve their desired goal is represented in these strategies involved in the game theory. The valid approaches to avoid the Sibyl attacks from the attackers are efficiently defended by the defenders. In our approach, we define attack graphs, use cases, and misuses cases in our gaming framework to analyze the vulnerabilities and security measures incorporated in a recommender system.
Justin Zhijun Zhan, Lijo Thomas, Venkata Pasumarthi
CIDM1
2011 Integrating online social networks for enhancing reputation systems of e-commerce
abstract
Online shopping has been introduced to Internet users for decades. However, many people are still susceptible to the risks of online shopping. Multiple risk-mitigating solutions have been adopted by most online shopping sites. One such solution, the reputation system, is capable of providing an up-to-date reputation score for every online seller. Nevertheless, it is also vulnerable to fake reviews, which can, in turn, mislead prospective buyers. In order to quench the fire, previous research work mounted friendship annotations to certain reviews, so that the prospective buyers can choose to trust such particular reviews based on their close relationships with the reviewers. In this paper, we extend this friend-annotation scheme by integrating online social networks to online shopping sites. Our protocol is designed to provide three online review submission methods with different privacy-preserving requirements.
Justin Zhijun Zhan, Nicholas Koceja, Kenneth Williams 0002, Justin Brewton
ISI2
2011 Anomaly detection using weak estimators
abstract
Anomaly detection involves identifying observations that deviate from the normal behavior of a system. One of the ways to achieve this is by identifying the phenomena that characterize “normal” observations. Subsequently, based on the characteristics of data learned from the “normal” observations, new observations are classified as being either “normal” or not. Most state-of-the-art approaches, especially those which belong to the family parameterized statistical schemes, work under the assumption that the underlying distributions of the observations are stationary. That is, they assume that the distributions that are learned during the training (or learning) phase, though unknown, are not time-varying. They further assume that the same distributions are relevant even as new observations are encountered. Although such a “stationarity” assumption is relevant for many applications, there are some anomaly detection problems where stationarity cannot be assumed. For example, in network monitoring, the patterns which are learned to represent normal behavior may change over time due to several factors such as network infrastructure expansion, new services, growth of user population, etc. Similarly, in meteorology, identifying anomalous temperature patterns involves taking into account seasonal changes of normal observations. Detecting anomalies or outliers under these circumstances introduces several challenges. Indeed, the ability to adapt to changes in non-stationary environments is necessary so that anomalous observations can be identified even with changes in what would otherwise be classified as “normal” behavior. In this paper, we proposed to apply weak estimation theory for anomaly detection in dynamic environments. In particular, we apply this theory to detect anomaly activities in system calls. Our experimental results demonstrate that our proposal is both feasible and effective for the detection of such anomalous activities.
Justin Zhijun Zhan, B. John Oommen, Johanna Crisostomo
ISI1
2011 Anomaly Detection in Dynamic Systems Using Weak Estimators
abstract
Anomaly detection involves identifying observations that deviate from the normal behavior of a system. One of the ways to achieve this is by identifying the phenomena that characterize “normal” observations. Subsequently, based on the characteristics of data learned from the “normal” observations, new observations are classified as being either “normal” or not. Most state-of-the-art approaches, especially those which belong to the family of parameterized statistical schemes, work under the assumption that the underlying distributions of the observations are stationary. That is, they assume that the distributions that are learned during the training (or learning) phase, though unknown, are not time-varying. They further assume that the same distributions are relevant even as new observations are encountered. Although such a “stationarity” assumption is relevant for many applications, there are some anomaly detection problems where stationarity cannot be assumed. For example, in network monitoring, the patterns which are learned to represent normal behavior may change over time due to several factors such as network infrastructure expansion, new services, growth of user population, and so on. Similarly, in meteorology, identifying anomalous temperature patterns involves taking into account seasonal changes of normal observations. Detecting anomalies or outliers under these circumstances introduces several challenges. Indeed, the ability to adapt to changes in nonstationary environments is necessary so that anomalous observations can be identified even with changes in what would otherwise be classified as “normal” behavior. In this article we propose to apply a family of weak estimators for anomaly detection in dynamic environments. In particular, we apply this theory to spam email detection. Our experimental results demonstrate that our proposal is both feasible and effective for the detection of such anomalous emails.
Justin Zhijun Zhan, B. John Oommen, Johanna Crisostomo
ACM Trans. Internet Techn.1
2010 Multi-party k-Means Clustering with Privacy Consideration
abstract
The k-means clustering algorithm is a widely used scheme to solve the clustering problem which classifies a given set of n data points in m-dimensional space into k clusters, whose centers are obtained by the centroids of the points in the same cluster. The problem with privacy consideration has been studied, when the data is distributed among different parties and the privacy of the distributed data is to be preserved. In this paper, we apply the concept of parallel computing to solve the privacy-preserving multi-party k-means clustering problem, when the data is vertically partitioned and horizontally partitioned respectively among different parties. We present algorithms for solving the problems for these two data partition models that run in O(nk) time and in O(m(k + log(n=k))) time respectively. The time complexities of the algorithms are much better than others without parallel computing.
Teng-Kai Yu, D. T. Lee, Shih-Ming Chang, Justin Zhijun Zhan
ISPA4
2010 P4P: Practical Large-Scale Privacy-Preserving Distributed Computation Robust against Malicious Users
Yitao Duan, NetEase Youdao, John F. Canny, Justin Zhijun Zhan
USENIX Security Symposium4
2010 Combined Authentication-Based Multilevel Access Control in Mobile Application for DailyLifeService
abstract
In current computing environments, collaborative computing has been a central concern in Ubiquitous, Convergent, and Social Computing. "MobiLife¿ and "MyLifeBits¿ are the leading projects for representing dailylifeservices and their systems require complicate and collaborative network systems. The collaborative computing environments remain in high potential risks for users' security and privacy because of diverse attack routes. In order to solve the problems, we design combined authentication and multilevel access control, which deals with cryptographic methods in a personal database of "MyLifeBits¿ system. We propose a scheme which is flexible in dynamic access authorization changes, secure against all the attacks from various routes, a minimum round of protocol, privacy preserving access control, and multifunctional.
Hyun-A Park, Jong-Wook Hong, Jaehyun Park 0007, Justin Zhijun Zhan, Dong Hoon Lee 0001
IEEE Trans. Mob. Comput.4
2010 Secure Collaborative Social Networks
abstract
A social network is the mapping and measuring of relationships and flows between individuals, groups, organizations, computers, websites, and other information/knowledge processing entities. The nodes in the network are the people and groups, while the links show relationships or flows between the nodes. Social networks provide both a visual and a mathematical analysis of relationships. Social network has been researched for a while. However, to our best knowledge, privacy-preserving social networks have not been well explored. In this paper, we would like to address how to build up a social network involving multiple parties. Data collection is a necessary step in the social-network-construction process. Due to privacy reasons, collecting data from different parties becomes difficult and these concerns may prevent the parties from directly sharing the data. How multiple parties collaboratively construct a social network without breaching data privacy presents a challenge. The objective of this paper is to provide a solution for privacy-preserving collaborative social-network problem.
Justin Zhijun Zhan
IEEE Trans. Syst. Man Cybern. Part C1
2010 Privacy-Preserving Collaborative Recommender Systems
abstract
Collaborative recommender systems use various types of information to help customers find products of personalized interest. To increase the usefulness of collaborative recommender systems in certain circumstances, it could be desirable to merge recommender system databases between companies, thus expanding the data pool. This can lead to privacy disclosure hazards during the merging process. This paper addresses how to avoid privacy disclosure in collaborative recommender systems by comparing with major cryptology approaches and constructing a more efficient privacy-preserving collaborative recommender system based on the scalar product protocol.
Justin Zhijun Zhan, Chia-Lung Hsieh, I-Cheng Wang, Tsan-sheng Hsu, Churn-Jung Liau, Dawei Wang 0004
IEEE Trans. Syst. Man Cybern. Part C1
2009 Authentication Using Multi-level Social Networks
Justin Zhijun Zhan
IC3K1
2009 Toward Empirical Aspects of Secure Scalar Product
abstract
There is a fair amount of research about privacy, but few empirical studies about its cost have been conducted. In the area of secure multiparty computation, the scalar product has long been reckoned as one of the most promising alternatives to classic logic gates. The reason for this is that the scalar product is not only complete, which is as good as logic gates, but also much more efficient than logic gates. As a result, we set out to study the computation and communication resources needed for some of the most well-known and frequently referenced secure scalar product protocols, including the composite residuosity, the invertible matrix, the polynomial sharing, and the commodity-based approaches. In addition to the implementation details of these approaches, we analyze and compare their execution time, computation time, and memory and random number consumption. Moreover, Fairplay, the benchmark approach that implements Yao's circuit evaluation protocol, is also included in our experiments in order to demonstrate the potential for the scalar products to replace logic gates.
I-Cheng Wang, Chih-Hao Shen, Justin Zhijun Zhan, Tsan-sheng Hsu, Churn-Jung Liau, Dawei Wang 0004
IEEE Trans. Syst. Man Cybern. Part C3
2008 Computer-Aided Privacy Requirements Elicitation Technique
abstract
The legislative penalties and economic penalties for privacy violations are more serious for a service provider these days. In spite of demonstrating that it is willing and able to protect the privacy of information, a service provider developing a privacy-compliant system faces two challenges; technical complexities and legal complexities. In this paper, we propose a computer-aided privacy requirements elicitation technique (PRET) that helps software developers elicit privacy requirements more efficiently in the early stages of software development. The goal of the PRET tool is to accelerate the elicitation process and prevent privacy requirements leaks by using a general privacy requirements database derived from privacy laws and empirical privacy requirements. We also show the results of integrating the PRET tool with the security quality requirements engineering (SQUARE) methodology and provide evidence of the efficacy of the resultant tool.
Seiya Miyazaki, Nancy R. Mead, Justin Zhijun Zhan
APSCC3
2008 Efficient keyword index search over encrypted documents of groups
abstract
In a real world, it is often in a group setting that sensitive information has to be stored in databases of a server. Although personal information does not need to be stored in a server, the secret information shared by group members is likely to be stored there. The shared sensitive information requires more security and privacy protection. To our best knowledge, there is no paper which deals with applicable group search schemes to a real world over encrypted data. In this paper, we propose two schemes which can search the encrypted documents without re-encrypting all documents in a server even if group keys have to be updated. The schemes can support general database normalization for encrypted database. Our experiments show that our schemes are much more efficient than the existing schemes.
Hyun-A Park, Dong Hoon Lee 0001, Justin Zhijun Zhan, Gary Blosser
ISI3
2007 Efficient Privacy-Preserving Association Rule Mining: P4P Style
abstract
In this paper we introduce a new practical framework, called P4P (peers for privacy), for privacy-preserving data mining. P4P features a hybrid architecture combining P2P and client-server paradigms and provides practical private protocols for user data validation and general computation. The architecture is guided by the natural incentives of the participants and allows the computation to be based on verifiable secret sharing (VSS) where arithmetic operations are done over small fields (e.g. 32 or 64 bits), so that private arithmetic operations have the same cost as normal arithmetic. Verification of user data, which uses large-field public-key arithmetic (1024 bits or more) and homomorphic computation, only requires a small number (constant or logarithmic in the size of user data) of large integer operations. The solution is extremely efficient: In experiments with our implementation, verification of a million-element vector takes a few seconds of server or client time on commodity PCs (in contrast, using standard techniques takes hours). This verification can be used in many privacy-preserving data mining tasks to detect cheating users who attempt to bias the computation by submitting exaggerated values as their inputs. As an example, we demonstrate how association rule mining can be done in the P4P model with near-optimal efficiency and provable privacy
Yitao Duan, John F. Canny, Justin Zhijun Zhan
CIDM3
2007 Quantifying Privacy for Privacy Preserving Data Mining
abstract
Data privacy is an important issue in data mining. How to protect respondents' data privacy during the data collection and mining process is a challenge to the security and privacy community. In this paper, we describe two schemes for privacy preserving naive Bayesian classification which is one of data mining tasks. More importantly, for each scheme, we present a method to measure data privacy. We finally compare these two methods
Justin Zhijun Zhan
CIDM1
2007 Using Homomorphic Encryption For Privacy-Preserving Collaborative Decision Tree Classification
abstract
To conduct data mining, we often need to collect data from various parties. Privacy concerns may prevent the parties from directly sharing the data. A challenging problem is how multiple parties collaboratively conduct data mining without breaching data privacy. The goal of this paper is to provide solutions for privacy-preserving decision tree classification which is one of data mining tasks. Our goal is to obtain accurate data mining results without disclosing private data
Justin Zhijun Zhan
CIDM1
2007 Privacy Preserving Collaborative Data Mining
abstract
Data mining and knowledge discovery in databases are important research areas that investigate the automatic extraction of previously unknown patterns from large amounts of data. The field connects the three worlds of databases, artificial intelligence and statistics. However, the usefulness of this data is negligible if meaningful information or knowledge cannot be extracted from it. Data mining and knowledge discovery attempts to answer this need. Aiming at developing practical solutions to privacy-preserving data mining problems, we have applied the random perturbation technique and the randomized response technique. The idea is to add random noise to the original data so that it is hidden. In another field, the success of homeland security aiming to counter terrorism depends on a combination of strength across different mission areas, effective international collaboration and information sharing to support a coalition in which different organizations and nations must share some, but not all, information. Information privacy thus becomes extremely important and our technique can be applied. In the Internet era, collaborative data mining is becoming a popular way to extract useful knowledge from large databases.
Justin Zhijun Zhan
ISI1
2007 Using Homomorphic Encryption and Digital Envelope Techniques for Privacy Preserving Collaborative Sequential Pattern Mining
abstract
Nowadays, data mining is widely used in various applications. Privacy is an important issue in data mining systems. By privacy, we mean how to conduct data mining without compromising much data privacy. In particular, we consider the scenario where data sharing for data mining purpose is the main goal. However, we would like minimize the data disclosure during data mining process.
Justin Zhijun Zhan
ISI1
2007 Collaboratively mining sequential patterns over private data
abstract
To conduct data mining, we often need to collect data from various parties. Privacy concerns may prevent the parties from directly sharing the data. A challenging problem is how multiple parties collaboratively conduct data mining without breaching data privacy. The goal of this paper is to provide solutions for privacy-preserving sequential pattern mining for horizontal collaboration. Our goal is to obtain accurate mining results without disclosing private data.
Justin Zhijun Zhan
SMC1
2007 Privacy preserving K-Medoids clustering
abstract
Privacy is an important issue in the collaborative data mining since privacy concerns may prevent the parties from directly sharing the data and some types of information about the data. How multiple parties collaboratively conduct data mining without breaching data privacy presents a challenge. This paper seeks to investigate solutions for privacy- preserving K-Medoids clustering which is one of data mining tasks.
Justin Zhijun Zhan
SMC1
2007 Privacy-preserving collaborative association rule mining
Justin Zhijun Zhan, Stan Matwin, LiWu Chang
J. Netw. Comput. Appl.1
2006 Privacy-Oriented Collaborative Learning Systems
abstract
This paper addresses the problem of data sharing among multiple parties in the following scenario: without disclosing their private data to each other, multiple parties, each having a private data set, want to collaboratively construct support vector machines using a linear, polynomial or sigmoid kernel function. To tackle this problem, we develop a secure protocol for multiple parties to conduct the desired computation. In our solution, multiple parties use homomorphic encryption and digital envelope techniques to exchange the data while keeping it private. All the parties are treated symmetrically: they all participate in the encryption and in the computation involved in learning support vector machines.
Justin Zhijun Zhan, Stan Matwin
SMC1
2005 Privacy-Preserving Collaborative Association Rule Mining
Justin Zhijun Zhan, Stan Matwin, LiWu Chang
DBSec1
2005 Private Mining of Association Rules
Justin Zhijun Zhan, Stan Matwin, LiWu Chang
ISI1
2004 Privacy-Preserving Multi-Party Decision Tree Induction
abstract
Data mining is a process to extract useful knowledge from large amounts of data. To conduct data mining, we often need to collect data. However, sometimes the data are distributed among various parties. Privacy concerns may prevent the parties from directly sharing the data and some types of information about the data. How multiple parties can collaboratively conduct data mining without breaching data privacy presents a grand challenge. In this paper, we propose a randomisation-based scheme for multi-parties to conduct data mining computations without disclosing their actual data sets to each other.
Justin Zhijun Zhan, LiWu Chang, Stan Matwin
DBSec1
2003 Using randomized response techniques for privacy-preserving data mining
abstract
Privacy is an important issue in data mining and knowledge discovery. In this paper, we propose to use the randomized response techniques to conduct the data mining computation. Specially, we present a method to build decision tree classifiers from the disguised data. We conduct experiments to compare the accuracy of our decision tree with the one built from the original undisguised data. Our results show that although the data are disguised, our method can still achieve fairly high accuracy. We also show how the parameter used in the randomized response techniques affects the accuracy of the results.
Wenliang Du 0001, Justin Zhijun Zhan
KDD2
2002 A practical approach to solve Secure Multi-party Computation problems
abstract
Secure Multi-party Computation (SMC) problems deal with the following situation: Two (or many) parties want to jointly perform a computation. Each party needs to contribute its private input to this computation, but no party should disclose its private inputs to the other parties, or to any third party. With the proliferation of the Internet, SMC problems becomes more and more important. So far no practical solution has emerged, largely because SMC studies have been focusing on zero information disclosure, an ideal security model that is expensive to achieve.Aiming at developing practical solutions to SMC problems, we propose a new paradigm, in which we use an acceptable security model that allows partial information disclosure. Our conjecture is that by lowering the restriction on the security, we can achieve a much better performance. The paradigm is motivated by the observation that in practice people do accept a less secure but much more efficient solution because sometimes disclosing information about their private data to certain degree is a risk that many people would rather take if the performance gain is so significant. Moreover, in our paradigm, the security is adjustable, such that users can adjust the level of security based on their definition of the acceptable security. We have developed a number of techniques under this new paradigm, and are currently conducting extensive studies based on this new paradigm.
Wenliang Du 0001, Justin Zhijun Zhan
NSPW2