EDBT 2026 Demo / reviewers in the wild / expert
Justin Zhijun Zhan
dblp:127/0551 · also Justin Z. Zhan, Justin Zhan
· DBLP profile ↗
15ranked-venue papers in the field
4as first author
2since 2021 · last 2022
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 7Data Mining & Knowledge Discovery · 6 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Cyberbully Detection Using BERT with Augmented TextsabstractDetecting cyberbullying in texts is an essential task as it curtails and identifies s ocial p roblems. I n t his p aper, we propose an architecture called Augmented BERT which combines both data augmentation techniques and BERT for detecting cyberbullying content in texts. Many techniques have been used in prior works to augment existing data for classification tasks and BERT had been applied in many text classification problems. However, there is a lack of annotated cyberbullying texts and obtaining annotated texts is hard and expensive. We propose to use various GAN-based and autoencoder-based data augmentation techniques to generate annotated data. The augmented texts can be used to fine-tune BERT. We choose to use HateBERT which is already pre-trained on abusive language to detect cyberbullying texts. Experimental results show an increased improvement over other cyberbullying detection models. Usman Anjum, Justin Zhijun Zhan |
IEEE Big Data | 3 |
| 2022 | WaveNets: Wavelet Channel Attention NetworksabstractChannel Attention reigns supreme as an effective technique in the field of computer vision. However, the proposed channel attention by SENet suffers from information loss in feature learning caused by the use of Global Average Pooling (GAP) to represent channels as scalars. Thus, designing effective channel attention mechanisms requires finding a s olution t o enhance features preservation in modeling channel inter-dependencies. In this work, we utilize Wavelet transform compression as a solution to the channel representation problem. We first t est wavelet transform as a standalone channel compression method. We prove that global average pooling is equivalent to the recursive approximate Haar wavelet transform. With this proof, we generalize channel attention using Wavelet compression and name it WaveNet. Implementation of our method can be embedded within existing channel attention methods with a couple of lines of code. We test our proposed method using ImageNet dataset for image classification t ask. O ur m ethod o utperforms t he baseline SENet-34, and SOTA FcaNet-34. Our code implementation is publicly available at https://github.com/hady1011/WaveNet-C. Hadi Salman, Caleb Parks, Shi Yin Hong, Justin Zhijun Zhan |
IEEE Big Data | 4 |
| 2020 | A Cancelable Multi-Modal Biometric Based Encryption Scheme for Medical ImagesabstractIn this paper, a novel multi-modal biometric based encryption scheme for medical images is proposed. It is based upon the ever-secure Advanced Encryption Standard Cipher Block Chaining (AES-CBC) mode of encryption and Indexing-First One (IFO) hashing for the generation of the secret keys. The proposed system utilizes two biometrics of the user, the iris and the fingerprint, but instead of using the features of the biometrics directly, they are hashed through the IFO process to create a secret key that can be revoked and regenerated in case of a compromise ensuring the security of the users biometrics. The IFO hashes obtained from the biometric feature vectors are then used as two different secret keys in a two round AES-CBC system to encrypt the medical image. The medical image can be decrypted by performing the AES-CBC rounds in reverse with the correct keys, meaning that with high probability, only the correct user will be able to decrypt the image. This encryption technique improves upon many existing medical encryption schemes based on biometrics, which do not take the protection of the biometric template into account. We were able to perform one full key generation, encryption and decryption in.027 seconds. In addition, this processes is lossless, meaning that there is no change in the pixels of the medical image during the encryption or decryption process, which is necessary of a medical image encryption system. Alycia N. Carey, Justin Zhijun Zhan |
IEEE BigData | 2 |
| 2020 | Semi-Supervised Learning and Feature Fusion for Multi-view Data ClusteringabstractGenerative Adversarial Networks GANs have become widely used in Single-view classification tasks. Nowadays, most of the data have multiple views and each view emphasizes a unique feature set of the data. In this paper, we investigate the application of GANs on Multi-view data for the task of clustering and few-shot learning. We propose mvSGAN, a deep learning approach to GAN multi-view clustering, where generator and classifier networks are in a competitive min-max game. A multi-view learning algorithm is implemented with a mini-batch which can handle large data sets. We test the accuracy of our method in clustering real-world data sets. The experimental results show that our method outperforms state-of-the-art research. Hadi Salman, Justin Zhijun Zhan |
IEEE BigData | 2 |
| 2018 | Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001 |
Knowl. Inf. Syst. | 5 |
| 2017 | Toward data quality analytics in signature verification using a convolutional neural networkabstractMany studies have been conducted on Handwritten Signature Verification. Researchers have taken many different approaches to accurately identify valid signatures from skilled forgeries, which closely resemble the real signature. The purpose of this paper is to suggest a method for validating written signatures on bank checks. This model uses a convolutional neural network (CNN) to analyze pixels from a signature image to recognize abnormalities. We believe the feature extraction capabilities of a CNN can optimize processing time and feature analysis of signature verification. Unique characteristics from signatures can be accurately and rapidly analyzed with multiple layers of receptive fields and hidden layers. Our method was able to correctly detect the validity of the inputted signature approximately 83 percent of the time. We tested our method using the SIGCOMP 2011 dataset. The main contribution of this method is to detect and decrease fraud committed, especially in the banking industry. Future uses of signature verification could include legal documents and the justice system. Shahab Tayeb, Matin Pirouz, Brittany Cozzens, Richard Huang, Maxwell Jay, Kyle Khembunjong, Sahan Paliskara, Felix Zhan 0001, Mark Zhang, Justin Zhijun Zhan, Shahram Latifi |
IEEE BigData | 10 |
| 2017 | Securing the positioning signals of autonomous vehiclesabstractOne of the fastest growing industries in America is autonomous vehicle technology. The main motivation is to decrease the number of accidents each year. This leads to the challenge presented in this paper, preventing the spoofing of signals coming into autonomous vehicles. The increased complexity of these vehicles creates more vulnerabilities for attackers to take advantage of. Authentication of vehicular ad hoc networks is one method to help stop the potential hacking of autonomous vehicles. We propose a two-factor authentication for GPS signals by synchronizing with Stratum-1 clocks and digital signatures to prevent man-in-the-middle attacks. The experiment uses a computer to simulate a GPS signal that is sent to a Raspberry Pi 3 along with a timestamp and hashed key using RSA-1024. The Raspberry Pi 3 represents the vehicle. The method presented will prevent GPS spoofing attacks that are using modified or corrupted messages, an impersonation attack, or a roadside unit replication attack. Shahab Tayeb, Matin Pirouz, Gabriel Esguerra, Kimiya Ghobadi, Jimson Huang, Robin Hill, Derwin Lawson, Stone Li, Tiffany Zhan, Justin Zhijun Zhan, Shahram Latifi |
IEEE BigData | 10 |
| 2017 | Toward predicting medical conditions using k-nearest neighborsabstractAs the healthcare industry becomes more reliant upon electronic records, the amount of medical data available for analysis increases exponentially. While this information contains valuable statistics, the sheer volume makes it difficult to analyze without efficient algorithms. By using machine learning to classify medical data, diagnoses can become more efficient, accurate, and accessible for the public. After choosing k-Nearest Neighbors for its simplicity, we applied it to datasets compiled by the University of California, Irvine Machine Learning Repository to diagnose two conditions - chronic kidney failure and heart disease - with an accuracy of approximately 90%. In the future, similar methods can be used on a larger scale to bring ease of use to the field of medical diagnostics. Shahab Tayeb, Matin Pirouz, Johann Sun, Kaylee Hall, Jessica Li, Connor Song, Apoorva Chauhan, Michael Ferra, Theresa Sager, Justin Zhijun Zhan, Shahram Latifi |
IEEE BigData | 11 |
| 2016 | An efficient algorithm to mine high average-utility itemsets
Jerry Chun-Wei Lin, Ting Li 0011, Philippe Fournier-Viger, Tzung-Pei Hong, Justin Zhijun Zhan, Miroslav Voznak |
Adv. Eng. Informatics | 5 |
| 2011 | Using gaming strategies for attacker and defender in recommender systemsabstractRatings are the prominent factors to decide the fate of any product in the present Internet Market and many people follow the ratings in a genuine sense. Unfortunately, the Sibyl attacks can affect the credibility of the genuine product. Influence limiter algorithms in recommender systems have been used extensively to overcome the Sibyl attacks but the effort could not reach the safe mark. This paper highlights an approach to generating gaming strategies for the attacker and defender in a recommender system. In a given recommender system environment, attackers and defenders play the most crucial part in a gaming strategy. A sequence of decision rules that an attacker or defender may use to achieve their desired goal is represented in these strategies involved in the game theory. The valid approaches to avoid the Sibyl attacks from the attackers are efficiently defended by the defenders. In our approach, we define attack graphs, use cases, and misuses cases in our gaming framework to analyze the vulnerabilities and security measures incorporated in a recommender system. Justin Zhijun Zhan, Lijo Thomas, Venkata Pasumarthi |
CIDM | 1 |
| 2009 | Authentication Using Multi-level Social Networks
Justin Zhijun Zhan |
IC3K | 1 |
| 2007 | Efficient Privacy-Preserving Association Rule Mining: P4P StyleabstractIn this paper we introduce a new practical framework, called P4P (peers for privacy), for privacy-preserving data mining. P4P features a hybrid architecture combining P2P and client-server paradigms and provides practical private protocols for user data validation and general computation. The architecture is guided by the natural incentives of the participants and allows the computation to be based on verifiable secret sharing (VSS) where arithmetic operations are done over small fields (e.g. 32 or 64 bits), so that private arithmetic operations have the same cost as normal arithmetic. Verification of user data, which uses large-field public-key arithmetic (1024 bits or more) and homomorphic computation, only requires a small number (constant or logarithmic in the size of user data) of large integer operations. The solution is extremely efficient: In experiments with our implementation, verification of a million-element vector takes a few seconds of server or client time on commodity PCs (in contrast, using standard techniques takes hours). This verification can be used in many privacy-preserving data mining tasks to detect cheating users who attempt to bias the computation by submitting exaggerated values as their inputs. As an example, we demonstrate how association rule mining can be done in the P4P model with near-optimal efficiency and provable privacy Yitao Duan, John F. Canny, Justin Zhijun Zhan |
CIDM | 3 |
| 2007 | Quantifying Privacy for Privacy Preserving Data MiningabstractData privacy is an important issue in data mining. How to protect respondents' data privacy during the data collection and mining process is a challenge to the security and privacy community. In this paper, we describe two schemes for privacy preserving naive Bayesian classification which is one of data mining tasks. More importantly, for each scheme, we present a method to measure data privacy. We finally compare these two methods Justin Zhijun Zhan |
CIDM | 1 |
| 2007 | Using Homomorphic Encryption For Privacy-Preserving Collaborative Decision Tree ClassificationabstractTo conduct data mining, we often need to collect data from various parties. Privacy concerns may prevent the parties from directly sharing the data. A challenging problem is how multiple parties collaboratively conduct data mining without breaching data privacy. The goal of this paper is to provide solutions for privacy-preserving decision tree classification which is one of data mining tasks. Our goal is to obtain accurate data mining results without disclosing private data Justin Zhijun Zhan |
CIDM | 1 |
| 2003 | Using randomized response techniques for privacy-preserving data miningabstractPrivacy is an important issue in data mining and knowledge discovery. In this paper, we propose to use the randomized response techniques to conduct the data mining computation. Specially, we present a method to build decision tree classifiers from the disguised data. We conduct experiments to compare the accuracy of our decision tree with the one built from the original undisguised data. Our results show that although the data are disguised, our method can still achieve fairly high accuracy. We also show how the parameter used in the randomized response techniques affects the accuracy of the results. Wenliang Du 0001, Justin Zhijun Zhan |
KDD | 2 |