Junwei Zhou 0002

dblp:04/3398-2 · DBLP profile ↗
← Back
47ranked-venue papers
15as first author
28since 2021 · last 2026
0000-0002-6094-1203ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 15 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Generation of Hard SAT Instances and Its Application in Negative Databases for Privacy Enhancement
abstract
In recent years, machine learning and deep learning have made remarkable progress and are now widely applied in various fields, including image classification, autonomous driving, natural language processing, and medical diagnosis. However, training these models requires large datasets, which often contain substantial amounts of sensitive personal information, such as medical records and financial details. Without effective privacy protection measures during model training, the risk of sensitive data leakage increases, potentially resulting in severe privacy violations and a loss of trust. As an innovative data representation technique, the Negative Database has proven to be an effective solution in privacy-sensitive domains. Negative databases can be derived from SAT (Boolean Satisfiability Problem) instances, and the hardness of these instances is directly correlated with the level of data protection provided by the Negative Databases. Developing efficient SAT instance generation algorithms to create harder SAT instances can significantly enhance the privacy protection capabilities of the Negative Database. This paper analyzes the hardness conditions of SAT solvers using the Conflict-Driven Clause Learning strategy and proposes a two-stage SAT instance generation algorithm to generate harder SAT instances. These hard instances not only aid in constructing more secure negative databases for enhanced privacy protection but also provide valuable test cases for evaluating and improving SAT solvers.
Dongdong Zhao 0001, Pang Chen, Changtian Song, Jianwen Xiang, Junwei Zhou 0002, Zebo Tang, Baogang Song
IEEE Trans. Big Data5
2026 SHRD: A Scalable Scheme for Hierarchical File Sharing With Rank-Aware Dissemination
Shulan Wang, Jinghong Gan, Chenbin Zhao, Fuyi Wang, Junwei Zhou 0002, Kaitai Liang
IEEE Trans. Inf. Forensics Secur.5
2025 Practical Lossless Recompression of JPEG Images Using Transform Domain Prediction
abstract
The inefficiency of the decades-old JPEG standard imposes a significant maintenance and cost burden on largescale software ecosystems that handle trillions of legacy files. Current solutions are ineffective, as modern codecs are incompatible with JPEG’s unique artifacts, while existing recompression tools rely on ad-hoc, handcrafted algorithms with fundamental performance limitations. To address this challenge systematically, we introduce PLLR, a reusable software framework built upon a novel design paradigm: learned information decomposition. Instead of manual rule-making, our framework automates the process of identifying and separating redundancies. It employs a Variational Autoencoder to learn a global, probabilistic model of the image content, effectively decoupling the predictable signal components from a highly sparse, low-entropy residual. By isolating this essential information, the subsequent entropy coding stage becomes significantly more efficient. Our implementation of this framework establishes a new state-of-the-art, achieving a fully lossless file size reduction of 31.54% on the Kodak dataset. This result validates the superiority of our learned decomposition framework for legacy data compression and offers a practical pathway to reduce the immense economic and environmental costs associated with large-scale data systems.
Junwei Zhou 0002, Jianwen Xiang
APSEC4
2025 Poster: LogCADA: Cross-System Log Anomaly Detection based on Two-Stage Multi-Source Domain Adaptation
abstract
Deep learning-based log anomaly detection demands extensive labeled data, posing significant challenges for emerging systems with limited logs. Transfer learning mitigates this issue by leveraging knowledge from data-rich source domains, enabling effective adaptation to data-scarce target domains. The challenge in recent cross-domain research lies in the joint optimization of knowledge transfer efficacy and model generalization capability. To address these limitations, we propose LogCADA, a novel logarithmic anomaly detection framework based on transfer learning, which can obtain effective common features through double-layer adversarial training, and distinguish common features and unique features between different domains through multi-source domain contrast alignment to achieve better knowledge transfer.The results demonstrate that our method can be adapted from dual source to single target domain and effectively overcome the inherent limitations of traditional cross-domain anomaly detection methods, yielding significant practical value for real-world log analysis scenarios.
Junwei Zhou 0002, Linhao Wang, Jianwen Xiang, Yanchao Yang 0002
CCS1
2025 Poster: GLog: Self-Evolving Log Anomaly Type Prediction via Instruction-Tuned LLM and Clustering
abstract
Log anomaly detection is critical for maintaining system reliability and observability in complex cloud and microservice environments. However, existing methods often remain limited to binary classification, struggle to adapt to dynamic log patterns, and suffer from semantic loss due to log parsing. To address these challenges, we propose GLog, an end-to-end framework that enables dynamic anomaly type prediction without requiring manual type labels. GLog first fine-tunes instruction-tuned large language models using normal/abnormal labels to achieve high-accuracy anomaly detection on raw, unparsed log sequences. It then clusters the detected anomalies to automatically generate pseudo anomaly type labels and descriptions, which are further used for second-stage fine-tuning, enabling the model to predict specific anomaly types with interpretable outputs. By leveraging full log semantics and dynamically updating its anomaly type repository, GLog reduces manual annotation costs and adapts to evolving system behaviors in large-scale environments.
Junwei Zhou 0002, Yanchao Yang 0002, Jianwen Xiang
CCS1
2025 DVSTdetector: Dual-View Spatio-Temporal Representation Learning for Intrusion Detection in Industrial Control System
abstract
Industrial Control Systems (ICS) serve as the backbone of critical infrastructures, controlling and automating the stable operation of industrial processes. With the progressive integration of industrial processes and Information Technology (IT), ICS have evolved from closed, isolated systems to open, interconnected networks. This evolution has significantly expanded attack surfaces and increased security vulnerabilities, making ICS more susceptible to cyber attacks. Intrusion Detection Systems (IDS), particularly those based on deep learning that can learn spatio-temporal features from raw network traffic, are the most effective methods for protecting ICS. However, most existing DL-based IDS adopt a Traditional Spatio-Temporal Feature (TSTF) view for representation learning, which often ignores the unique characteristics of industrial protocols and ICS communication patterns, and loses important fine-grained discriminative information. As a result, the detection models perform poorer with higher false negatives and false positives when applied in ICS. To overcome the above issue, we propose a novel Segment-Based Spatio-Temporal Feature (SSTF) view, which leverages temporal dependencies among the same segments in different packets within a flow and spatial correlations between different segments. Additionally, we introduce a dual-view intrusion detection framework-DVSTdetector, that integrates both the TSTF and SSTF views and employs two workflows to promote better representation learning from both global and local perspectives in parallel, obtaining more robust spatiotemporal features. A publicly available dataset (WDT) and a private dataset (XLP) are used to evaluate our approach. The experimental results demonstrate its effectiveness and superiority, outperforming six state-of-the-art approaches and achieving high performance across six metrics: Accuracy ($99.35 \%$, 99.87%), Precision (99.28%, 99.93%), Recall (98.52%, 99.90%), F1-score ($\mathbf{9 8. 9 0 \%, ~} \mathbf{9 9. 9 1 \%}$), AUC-ROC ($\mathbf{9 9. 1 1 \%, ~ 9 9. 8 4 \%), ~ a n d ~ a ~ l o w ~ F a l s e ~}$ Positive Rate ($\mathbf{0. 2 9 \%, ~} \mathbf{0. 2 1 \%}$).
Qianrong Zheng, Zhe Xia, Junwei Zhou 0002, Jianwen Xiang
ISSRE5
2025 AMDIC: Adaptive Multi-Granularity Joint Context Transfer for Distributed Image Coding
Benyi Zhang, Junwei Zhou 0002, Yanchao Yang 0002, Jianwen Xiang
PRCV (9)2
2025 Semi-supervised method for anomaly detection in HTTP traffic
abstract
Anomaly detection in HTTP traffic is critical for securing web applications against evolving cyber threats. We propose a semi-supervised method that combines domain-specific language modeling with sequence reconstruction to identify anomalies in HTTP requests. Our approach leverages only benign traffic for training and uses reconstruction errors for detecting malicious activity. It achieves a strong balance between precision and recall while maintaining low computational requirements, making it suitable for real-time and edge deployments. Extensive evaluations on three public HTTP datasets show that our method outperforms traditional baselines and fine-tuned BERT models, with an F1-score of 0.92 and AUC of 0.96. We also introduce a simple interpretability mechanism by attributing anomalies to token-level reconstruction errors, providing insights into detected threats. The proposed solution is scalable, lightweight, and effective across diverse attack scenarios without requiring large labeled datasets.
Malki Ishara Wasundara, Junwei Zhou 0002, Yanchao Yang 0002, Dongdong Zhao 0001, Jianwen Xiang
EURASIP J. Inf. Secur.2
2025 Jpeg stereo image lossy recompression with mutual information enhancement
Junwei Zhou 0002, Benyi Zhang, Shengping Wu, Lei Zhou 0008, Yanchao Yang 0002, Jianwen Xiang
Multim. Syst.1
2025 DDVC: Deep Distributed Video Coding Using Quality Enhancement Network
abstract
Distributed video coding (DVC) transfers the complex process of the encoder to the decoder, which is suitable for video applications with limited encoding resources. Deep learning has shown impressive performance in video coding tasks in learning nonlinear compact representations of input frames and reconstructing video frame details. It is worth exploring whether deep learning implementation of the DVC paradigm is feasible and whether performance gains can be obtained. This paper proposes a deep DVC scheme (DDVC) using a quality enhancement network (QEN), which maps pixels to a more compressible latent space via an autoencoder resulting in a compact representation of Wyner-Ziv (WZ) frames. Moreover, considering the spatio-temporal correlation between the WZ frame and the Key frame, the QEN on the decoder side, using CNN and LSTM iteratively extracts common information between the WZ frame and the Key frame, which could further finetune the WZ frame reconstruction. We evaluated DDVC in limited encoding resources application scenarios with 19 related video sequences. Results on the video sequences with different motion intensity levels show that DDVC significantly outperforms existing schemes in reconstruction quality with the same compression ratio. We open-sourced the implementation at GitHub1.
Junwei Zhou 0002, Zhuang Ye, Xiangbo Yi, Qiuzhen Lin, Jianwen Xiang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Cross-Project Aging-Related Bug Prediction Based on Transfer Learning and Class Imbalance Learning
abstract
Software aging results from aging-related bugs (ARBs) in long-running systems, that usually causes performance decline and system crashes. Since collecting ARB data is challenging due to its scarcity, it hinders the development of effective prediction models. Moreover, existing cross-project ARB prediction methods often ignore project-specific distribution differences and neglect class imbalance and overlap issues between ARB and non-ARB classes. In this paper, a hybrid approach that combines the balanced distribution adaptation (BDA), the improved subclass discriminant analysis (ISDA), and the self-paced ensemble under-sampling (SPE) techniques, called BISP in short, is proposed to address the aforementioned problems. The main idea behind BISP is first to use BDA to adaptively reduce the difference of projects' marginal distribution and conditional distribution, and then employ ISDA and SPE to alleviate the severe class imbalance together with class overlap. Experimental results obtained for six classifiers and six cross-project datasets show that compared with the state-of-the-art approaches TLAP and JDA-ISDA based on transfer learning, BISP improves the average balance by 34.8% and 2.3% and improves the average AUC by 26.5% and 8.4%, respectively. Compared with the deep learning approach SRLA, BISP can improve the average balance value by 5.1%.
Bin Xu 0020, Dongdong Zhao 0001, Junwei Zhou 0002, Wenzhi Xie, Jianwen Xiang
IEEE Trans. Dependable Secur. Comput.3
2025 LogDLR: Unsupervised Cross-System Log Anomaly Detection Through Domain-Invariant Latent Representation
abstract
Log anomaly detection aims to discover abnormal events from massive log data to ensure the security and reliability of software systems. However, due to the heterogeneity of log formats and syntaxes across different systems, existing log anomaly detection methods often need to be designed and trained for specific systems, lacking generalization ability. To address this challenge, we propose LogDLR, a novel unsupervised cross-system log anomaly detection method. The core idea of LogDLR is to use universal sentence embeddings and a Transformer-based autoencoder to extract domain-invariant latent representations from log entries, which can effectively adapt to log format changes and capture semantic information and dependencies in log sequences. To obtain domain-invariant latent representations, we adopt a domain-adversarial training strategy, introducing a domain discriminator that competes with the Transformer-based encoder through a gradient reversal layer, forcing the encoder to learn shared knowledge between different system logs. Finally, the Transformer-based decoder detects anomalies based on the domain-invariant representations obtained by the encoder. We evaluate LogDLR in simulated cross-system scenarios using three publicly available log datasets. The experimental results show that LogDLR can handle heterogeneous logs effectively in cross-system scenarios and achieve efficient and accurate anomaly detection on both source and target systems.
Junwei Zhou 0002, Shaowen Ying, Shulan Wang, Dongdong Zhao 0001, Jianwen Xiang, Kaitai Liang, Peng Liu 0005
IEEE Trans. Dependable Secur. Comput.1
2025 Learning-Based Directional Improvement Prediction for Dynamic Multiobjective Optimization
abstract
In recent years, dynamic multiobjective evolutionary algorithms (DMOEAs) using the prediction strategy have shown promising performance for solving dynamic multiobjective optimization problems (DMOPs), as they can predict environmental changing trends in advance. However, most of them follow a regular change pattern and thus their performance is compromised when solving DMOPs with irregular change patterns (e.g., nonlinear correlations). To alleviate this challenge, this article proposes a DMOEA with a learnable prediction for tackling DMOPs. Specifically, a neural network is designed to effectively capture diverse change patterns of the environment. Based on the change patterns learned, a directional improvement prediction (DIP) is developed to guide the evolutionary search toward promising directions in the decision space. In this way, a superior initial population with good convergence and diversity is predicted by DIP, which can be more effective for solving various DMOPs. Comprehensive empirical studies show that the proposed DIP is effective and the proposed algorithm has some advantages over five competitive DMOEAs when solving three commonly used benchmarks and one real-world problem.
Yulong Ye, Songbai Liu, Junwei Zhou 0002, Qiuzhen Lin, Min Jiang 0005, Kay Chen Tan
IEEE Trans. Evol. Comput.3
2025 DRLLog: Deep Reinforcement Learning for Online Log Anomaly Detection
abstract
System logs record the system’s status and application behavior, providing support for various system management and diagnostic tasks. However, existing methods for log anomaly detection face several challenges, including limitations in recognizing current types of anomalous logs and difficulties in performing online incremental updates to the anomaly detection models. To address these challenges, this paper introduces DRLLog, which applies Deep Reinforcement Learning (DRL) networks to detect anomalous events. DRLLog uses Deep Q Network (DQN) as the agent, with log entries serving as reward signals. By interacting with the environment generated from log data and adopting various action behaviors, it aims to maximize the reward value obtained as feedback. Through this approach, DRLLog achieves learning from historical log data and perception of the current environment, enabling continuous learning and adaptation to different log sequence patterns. Additionally, DRLLog introduces low-rank adaptation by using two low-rank parameter matrices in the fully connected layer of the DQN to represent changes in its weight matrix. During online model learning, only low-rank parameter matrices of the model are updated, effectively reducing the model’s overhead. Furthermore, DRLLog introduces focal loss to focus more on learning the features of anomalous logs, effectively addressing the issue of imbalanced quantities between normal and anomalous logs. We evaluated the performance on widely used log datasets, including HDFS, BGL and ThunderBird, showing an average improvement of 3% in F1-Score compared to baseline methods. During online model learning, DRLLog achieves an average reduction of 90% in parameter count and a significant decrease in training and testing time as well.
Junwei Zhou 0002, Xiangtian Yu, Yanchao Yang 0002, Jianwen Xiang
IEEE Trans. Netw. Serv. Manag.1
2024 AEDD: Anomaly Edge Detection Defense for Visual Recognition in Autonomous Vehicle Systems
abstract
Visual recognition algorithms based on deep neural network (DNN) have been widely used in the design of automatic driving to recognize traffic sign images. However, there exists adversarial patches which are essentially the anormal image block that can be locally observed but not noticed by humans. And these visual recognition algorithms often suffer from the effect of adversarial patches, due to these patches can change the algorithms recognition result of the images. To solve the above issues, this work proposes the anomaly edge detection and image inpainting defense (AEDD) for visual recognition. This framework uses anomaly location to obtain the anomaly area, uses edge detection to get an accurate edge of the anomaly area, finally uses image inpainting to repair this area. We also combine two attack algorithms with three patch sizes, and generate six types of adversarial patches on the GTSRB dataset. We have demonstrated the effectiveness of our approach, resulting in an average 6.6% increase in defense accuracy compared to the state-of-the-art methods. Our code is available at https://github.com/drtt438/AEDD for the purpose of reproducibility.
Junwei Zhou 0002, Dongdong Zhao 0001, Dongqing Liao, Jianwen Xiang
CSCWD2
2024 Lightweight Autoencoder with Hierarchical Priors for Learned Image Compression
abstract
Image compression has become an important task for reducing storage and transmission costs. However, recent models for learned image compression have been developed to increase the network’s number of layers and channels to achieve better visual effects. This resulted in higher computing and memory resources, making deploying the model on compute-constrained platforms such as wireless devices impractical. In this paper, we propose a lightweight autoencoder with hierarchical priors. The lightweight autoencoder reduces the model’s parameter size and calculation amount based on ensuring high fidelity and low bit rates of the image. Simulation results indicate the proposed model yields a smaller size: the parameters are reduced by 81.66%, and the calculation amount is reduced by 94.7% over the benchmark. Besides, the proposed model results in a speed improvement of 200 times. At the same time, our model achieves nearly the same performance as the baseline on MS-SSIM and LPIPS distortion metrics.
Junwei Zhou 0002, Lei Zhou 0008, Yanchao Yang 0002, Jianwen Xiang
HPCC1
2024 CIDF: Combined Intrusion Detection Framework in Industrial Control Systems based on Packet Signature and Enhanced FSFDP
abstract
Industrial Control System (ICS) is vital to critical infrastructures, yet it faces increasing security threats. Current Intrusion Detection System (IDS) designed for ICS often overlooks the unbalanced resource distribution among devices at different layers and primarily focus on known attacks, rendering it difficult to be deployed on all key nodes and vulnerable to unknown threats. To address above issues, we propose a Combined Intrusion Detection Framework (CIDF). This innovative approach is based on strategy of “multi-level layered deployment, combined detection”, deploying the Packet Signature model and the Enhanced Fast Search and Find of Density Peaks (EFSFDP) model on devices at different layers. To achieve optimal use of resource and full protection for ICS and combining the advantages of multiple detection methods to effective detect both known and unknown attacks. The Evaluation using a public gas pipeline dataset and a private dataset shows our approach outperforms existing methods, achieving an average Accuracy, Precision, and Recall of 94%, 95.5%, and 86.5% respectively, and along with superior detection speed.
Jianwen Xiang, Qianrong Zheng, Longmin Deng, Dongdong Zhao 0001, Junwei Zhou 0002
Internetware6
2024 Block-Feature Fusion for Privacy-Protected Iris Recognition
abstract
Ensuring privacy often results in sacrificing the accuracy of iris recognition systems. A primary challenge in contemporary iris biometric privacy methods lies in striking a balance between recognition accuracy and privacy of protected iris templates. Hence, any proposition for privacy-protected iris recognition must prioritize irreversibility, revocability, and unlinkability to uphold robust privacy standards while achieving higher recognition accuracy. This research proposes an approach that stands as a robust solution with acceptable advancement in three challenges: recognition accuracy, privacy protection and computational efficiency. We experimented with an innovative technique that manipulates the columns of an iris template by fusing the bit pattern of the template using the XOR operation. The transformation process is non-linear. This fusion introduces randomness and variability to the fused templates. It also poses enhanced privacy protection. Two datasets were used to validate the proposed approach. Based on the results of dataset 1, the proposed approach accepts genuine users at a rate of 99.11% while it accepts 0.01% of imposters. For dataset 2, the Genuine Acceptance Rate (GAR) is depicted as 81.12% while FAR is at 0.01%. The proposed approach can be applied in practice due to its higher computational efficiency. As further improvements, the research can be extended to more widespread databases and higher-quality iris samples.
Wiraj Udara Wickramaarachchi, Junwei Zhou 0002, Dongdong Zhao 0001, Jianwen Xiang
TrustCom2
2024 Improving effort-aware defect prediction by directly learning to rank software modules
Xiao Yu 0008, Jiqing Rao, Lei Liu 0062, Guancheng Lin, Jacky W. Keung, Junwei Zhou 0002, Jianwen Xiang
Inf. Softw. Technol.7
2024 An effective iris biometric privacy protection scheme with renewability
Wiraj Udara Wickramaarachchi, Dongdong Zhao 0001, Junwei Zhou 0002, Jianwen Xiang
J. Inf. Secur. Appl.3
2024 Neural Net-Enhanced Competitive Swarm Optimizer for Large-Scale Multiobjective Optimization
abstract
The competitive swarm optimizer (CSO) classifies swarm particles into loser and winner particles and then uses the winner particles to efficiently guide the search of the loser particles. This approach has very promising performance in solving large-scale multiobjective optimization problems (LMOPs). However, most studies of CSOs ignore the evolution of the winner particles, although their quality is very important for the final optimization performance. Aiming to fill this research gap, this article proposes a new neural net-enhanced CSO for solving LMOPs, called NN-CSO, which not only guides the loser particles via the original CSO strategy, but also applies our trained neural network (NN) model to evolve winner particles. First, the swarm particles are classified into winner and loser particles by the pairwise competition. Then, the loser particles and winner particles are, respectively, treated as the input and desired output to train the NN model, which tries to learn promising evolutionary dynamics by driving the loser particles toward the winners. Finally, when model training is complete, the winner particles are evolved by the well-trained NN model, while the loser particles are still guided by the winner particles to maintain the search pattern of CSOs. To evaluate the performance of our designed NN-CSO, several LMOPs with up to ten objectives and 1000 decision variables are adopted, and the experimental results show that our designed NN model can significantly improve the performance of CSOs and shows some advantages over several state-of-the-art large-scale multiobjective evolutionary algorithms as well as over model-based evolutionary algorithms.
Qiuzhen Lin, Songbai Liu, Junwei Zhou 0002, Zhong Ming 0001, Carlos A. Coello Coello
IEEE Trans. Cybern.5
2024 Evolutionary Optimization with a Simplified Helper Task for High-Dimensional Expensive Multiobjective Problems
abstract
In recent years, surrogate-assisted evolutionary algorithms (SAEAs) have been sufficiently studied for tackling computationally expensive multiobjective optimization problems (EMOPs), as they can quickly estimate the qualities of solutions by using surrogate models to substitute for expensive evaluations. However, most existing SAEAs only show promising performance for solving EMOPs with no more than 10 dimensions, and become less efficient for tackling EMOPs with higher dimensionality. Thus, this article proposes a new SAEA with a simplified helper task for tackling high-dimensional EMOPs. In each generation, one simplified task will be generated artificially by using random dimension reduction on the target task (i.e., the target EMOPs). Then, two surrogate models are trained for the helper task and the target task, respectively. Based on the trained surrogate models, evolutionary multitasking optimization is run to solve these two tasks so that the experiences of solving the helper task can be transferred to speed up the convergence of tackling the target task. Moreover, an effective model management strategy is designed to select new promising samples for training the surrogate models. When compared to five competitive SAEAs on four well-known benchmark suites, the experiments validate the advantages of the proposed algorithm on most test cases.
Xunfeng Wu, Qiuzhen Lin, Junwei Zhou 0002, Songbai Liu, Carlos A. Coello Coello, Victor C. M. Leung
ACM Trans. Evol. Learn. Optim.3
2023 Secure genotype imputation using homomorphic encryption
abstract
Genotype imputation estimates missing genotypes from the haplotype or genotype reference panel in individual genetic sequences, which boosts the potential of genome-wide association and is essential in genetic data analysis. However, the genetic sequences involve people’s privacy, confirming an individual’s identification and even disease information. This work proposes a secure genotype imputation model, which uses a linear regression model and the homomorphic encryption scheme over ciphertext to impute missing genotypes. The inference model is trained with float plaintext parameters, which are round into integers to avoid high complexity homomorphic evaluation on float number operations without bootstrapping operations. Even though the rounding parameters in the inference model are not the same as those in the trained model, We find that it will no effect on the outcome of the homomorphic prediction. Thus, a high-efficiency genotype imputation inference model over the ciphertext is obtained while keeping the high-security level. The simulation results indicate that the accuracy of the secure inference model is almost the same as the original model trained on float parameters. The secure inference model’s accuracy is 98.6% for a single genotype.
Junwei Zhou 0002, Botian Lei, Huile Lang, Emmanouil A. Panaousis, Kaitai Liang, Jianwen Xiang
J. Inf. Secur. Appl.1
2022 End-to-end Distributed Video Coding
abstract
Existing distributed video coding (DVC) frameworks use manually designed and optimized modules when encoding and decoding video. Each module can set the appropriate parameters as much as possible to achieve its independent optimization. But there is no connection between the modules, and overall end-to-end optimization is not realized. Inspired by the application of neural networks to video coding, we try to implement DVC using neural networks. This article proposes an end-to-end DVC framework, which combines DVC and neural networks to perform end-to-end encoding and decoding of Wyner-Ziv (WZ) frames, while achieving variable compression rates. We employ four different kinds of side information to assist in decoding the WZ frames. Experimental results demonstrate that the proposed framework achieves excellent decoded performance on videos with varying degrees of motion intensity.
Junwei Zhou 0002, Ting Lv, XiangBo Yi
DCC1
2022 CBSDI: Cross-Architecture Binary Code Similarity Detection based on Index Table
abstract
Binary code similarity detection for cross-platform is widely used in plagiarism detection, malware detection and vulnerability search, aiming to detect whether two binary functions over different platforms are similar. Existing cross-architecture approaches mainly rely on the approximate matching calculation of complex high-dimensional features, such as graph, which are inevitably slow and unsuitable for large-scale applications. To solve this problem, we propose a novel approach based on index table called CBSDI, improving efficiency by screening a batch of mismatched functions before similarity detection. We select three features and compare them across architectures to select the most appropriate one to construct the index table, and this table can be embedded in other tools. The evaluation shows that the index table can roughly cut the computational costs in half when there are few errors. Moreover, compared with the related works in the literature, our proposed approach can improve not only the efficiency but also the accuracy.
Longmin Deng, Dongdong Zhao 0001, Junwei Zhou 0002, Zhe Xia, Jianwen Xiang
QRS3
2022 DeepSyslog: Deep Anomaly Detection on Syslog Using Sentence Embedding and Metadata
abstract
Anomaly events indicating the unhealthy status of the computer system are recorded in the system log (Syslog). Therefore, Syslog-based anomaly event detection is crucial for diagnosing system issues and problems. However, existing log-based anomaly detection approaches use raw and unstructured log entriesindependentlyandincompletely, i.e., without considering the context of each event and event metadata in the logs. They employ incomplete representation of unstructured log data, limiting the deep learning model’s capacity in the early stage, which tends to omit anomaly events and cause false alarms. In this work, we propose DeepSyslog, which represents Syslog with the context of log events and event metadata in the logs. Inspired by the sequence nature of the log stream, we employ unsupervised sentence embedding to extract the semantic and context information hidden in the log stream, rather than word embedding or one-hot embedding, which only capture the similarities between log words. The sentence embedding is further integrated with event metadata to form complete representations of Syslog, which can distinguish the anomaly caused by the correlated log entries and exceptional event metadata in the log. The simulation results on widely used log datasets show that DeepSyslog achieves high performance compared with the existing log-based anomaly event detection approaches.
Junwei Zhou 0002, Yijia Qian, Qingtian Zou, Peng Liu 0005, Jianwen Xiang
IEEE Trans. Inf. Forensics Secur.1
2021 Accurate and Robust Stereo Direct Visual Odometry for Agricultural Environment
abstract
Vision-based localization and mapping in the agricultural environment is challenging due to the unstructured scene with unstable features, illumination variations, bumpy roads, and dynamic environmental objects. To address these challenges, we propose an accurate and robust stereo direct visual odometry system with modifications on Stereo-DSO. We firstly select some well-matched static stereo points in the latest keyframe to improve the accuracy of inverse depth calculation for tracking. The inverse depth can further distinguish close objects from background, which will avoid large and far-away scene objects in keyframe determination. To boost efficiency and accuracy at the tracking stage, we propose a point selection method to sample map points and remove outliers. Furthermore, altitude smoothness verification with a local flat ground assumption and recovery method for tracking failure on bumpy roads are proposed to improve the system’s robustness. Finally, a far-away keyframe is reserved in the sliding window to alleviate the orientation drift since the agricultural robots usually move straightly following the crop row. Our system achieved new state-of-the-art results on Flourish dataset and the recently released Rosario dataset.
Junwei Zhou 0002, Liangliang Wang 0007, Shengwu Xiong 0001
ICRA2
2021 Face alignment using two-stage cascaded pose regression and mirror error correction
Ziye Tong, Junwei Zhou 0002
Pattern Recognit.2
2020 Modified Decoding Metric Of Distributed Arithmetic Coding
abstract
As an alternative implementation of Slepian-Wolf coding, distributed arithmetic coding is very competitive in short and medium block lengths. The existing distributed arithmetic coding decoder uses the maximum a posteriori metric and the M-algorithm to select the decoding sequence by considering the prior information. Since the prior probability distribution has been explored by the encoding process via the model stage and retained in the codeword, we propose a modified decoding metric ignoring the prior information. The reliability of the decoding paths is only determined by the correlation between the side information and the input source. Simulation results show that the modified metric can significantly reduce the decoding error.
Yanchao Yang 0002, Mingwei Qi, Junwei Zhou 0002, Lee-Ming Cheng
ICIP3
2020 Face Anti-Spoofing Based on Dynamic Color Texture Analysis Using Local Directional Number Pattern
abstract
Face anti-spoofing is becoming increasingly indispensable for face recognition systems, which are vulnerable to various spoofing attacks performed using fake photos and videos. In this paper, a novel “LDN-TOP representation followed by ProCRC classification” pipeline for face anti-spoofing is proposed. We use local directional number pattern (LDN) with the derivative-Gaussian mask to capture detailed appearance information resisting illumination variations and noises, which can influence the texture pattern distribution. To further capture motion information, we extend LDN to a spatial-temporal variant named local directional number pattern from three orthogonal planes (LDN- TOP). The multi-scale LDN- TOP capturing complete information is extracted from color images to generate the feature vector with powerful representation capacity. Finally, the feature vector is fed into the probabilistic collaborative representation based classifier (ProCRC) for face anti-spoofing. Our method is evaluated on three challenging public datasets, namely CASIA FASD, Replay-Attack database, and UVAD database using sequence-based evaluation protocol. The experimental results show that our method can achieve promising performance with 0.37% EER on CASIA and 5.73% HTER on UVAD. The performance on Replay-Attack database is also competitive.
Junwei Zhou 0002, Ke Shu, Peng Liu 0005, Jianwen Xiang, Shengwu Xiong 0001
ICPR1
2020 Cross-Project Aging-Related Bug Prediction Based on Joint Distribution Adaptation and Improved Subclass Discriminant Analysis
abstract
Software aging, which is caused by Aging-Related Bugs (ARBs), refers to the phenomenon of performance degradation and eventual crash in long running systems. In order to discover and remove ARBs, ARB prediction is proposed. However, due to the low presence and reproducing difficulty of ARBs, it is usually difficult to collect sufficient ARB data within a project. Therefore, cross-project ARB prediction is proposed as a solution to build the target project's ARB predictor by using the labeled data from the source project. A key point for cross-project ARB prediction is to reduce distribution difference between source and target project. However, existing approaches mainly focus on the marginal distribution difference while somehow overlook the conditional distribution difference, and they mainly use random oversampling to alleviate the class imbalance which may lead to overfitting. To address these problems, we propose a new crossproject ARB prediction approach based on Joint Distribution Adaptation (JDA) and Improved Subclass Discriminant Analysis (ISDA), called JDA-ISDA. The key idea of JDA-ISDA is first to use JDA to reduce the marginal distribution and conditional distribution difference jointly and then apply ISDA to alleviate the severe class imbalance problem. A set of experiments are carried out on two large open-source projects with six different machine learning (ML) classifiers. The experimental results demonstrate that compared with the state-of-the-art Transfer Learning based Aging-related bug Prediction (TLAP) and Supervised Representation Learning Approach (SRLA), JDA-ISDA is much more robust to different ML classifiers than TLAP, and the average improvement in terms of the balance value can be achieved up to 31.8%, and JDA-ISDA also outperforms TLAP and SRLA on average when logistic regression is chosen as the classifier for best performance prediction.
Bin Xu 0020, Dongdong Zhao 0001, Junwei Zhou 0002, Jianwen Xiang
ISSRE4
2020 Using deep learning to solve computer security challenges: a survey
abstract
Abstract Although using machine learning techniques to solve computer security challenges is not a new idea, the rapidly emerging Deep Learning technology has recently triggered a substantial amount of interests in the computer security community. This paper seeks to provide a dedicated review of the very recent research works on using Deep Learning techniques to solve computer security challenges. In particular, the review covers eight computer security problems being solved by applications of Deep Learning: security-oriented program analysis, defending return-oriented programming (ROP) attacks, achieving control-flow integrity (CFI), defending network attacks, malware classification, system-event-based anomaly detection, memory forensics, and fuzzing for software security.
Yoon-Ho Choi, Peng Liu 0005, Zitong Shang, Lan Zhang 0008, Junwei Zhou 0002, Qingtian Zou
Cybersecur.7
2019 Robust Facial Landmark Localization Based on Two-Stage Cascaded Pose Regression
Ziye Tong, Junwei Zhou 0002, Yanchao Yang 0002, Lee-Ming Cheng
AAAI2
2019 Novel Meaningful Image Encryption Based on Block Compressive Sensing
abstract
This paper proposes a new image compression-encryption algorithm based on a meaningful image encryption framework. In block compressed sensing, the plain image is divided into blocks, and subsequently, each block is rendered sparse. The zigzag scrambling method is used to scramble pixel positions in all the blocks, and subsequently, dimension reduction is undertaken via compressive sensing. To ensure the robustness and security of our algorithm and the convenience of subsequent embedding operations, each block is merged, quantized, and disturbed again to obtain the secret image. In particular, landscape paintings have a characteristic hazy beauty, and secret images can be camouflaged in them to some extent. For this reason, in this paper, a landscape painting is selected as the carrier image. After a 2-level discrete wavelet transform (DWT) of the carrier image, the low-frequency and high-frequency coefficients obtained are further subjected to a discrete cosine transform (DCT). The DCT is simultaneously applied to the secret image as well to split it. Next, it is embedded into the DCT coefficients of the low-frequency and high-frequency components, respectively. Finally, the encrypted image is obtained. The experimental results show that, under the same compression ratio, the proposed image compression-encryption algorithm has better reconstruction effect, stronger security and imperceptibility, lower computational complexity, shorter time consumption, and lesser storage space requirements than the existing ones.
Guodong Ye, Junwei Zhou 0002
Secur. Commun. Networks4
2019 Distributed video coding using interval overlapped arithmetic coding
Junwei Zhou 0002, Yincheng Fu, Yanchao Yang 0002, Anthony Tung Shuen Ho
Signal Process. Image Commun.1
2019 Multi-Channel Embedding Convolutional Neural Network Model for Arabic Sentiment Classification
abstract
With the advent of social network services, Arabs’ opinions on the web have attracted many researchers in recent years toward detecting and classifying sentiments in Arabic tweets and reviews. However, the impact of word embeddings vectors (WEVs) initialization and dataset balance on Arabic sentiment classification using deep learning has not been thoroughly studied. In this article, a multi-channel embedding convolutional neural network (MCE-CNN) is proposed to improve Arabic sentiment classification by learning sentiment features from different text domains, word, and character n-grams levels. MCE-CNN encodes a combination of different pre-trained word embeddings into the embedding block at each embedding channel and trains these channels in parallel. Besides, a separate feature extraction module implemented in a CNN block is used to extract more relevant sentiment features. These channels and blocks help to start training on high-quality WEVs and fine-tuning them. The performance of MCE-CNN is evaluated on several standard balanced and imbalanced datasets to reflect real-world use cases. Experimental results show that MCE-CNN provides a high classification accuracy and benefits from the second embedding channel on both standard Arabic and dialectal Arabic text, which outperforms state-of-the-art methods.
Abdelghani Dahou, Shengwu Xiong 0001, Junwei Zhou 0002, Mohamed E. Abd Elaziz
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2018 Image authentication using distributed arithmetic coding
Junwei Zhou 0002, Fang Liu 0024, Lee-Ming Cheng
Multim. Tools Appl.1
2018 Load Balancing Opportunistic Routing for Cognitive Radio Ad Hoc Networks
abstract
Recent research activities have shown that opportunistic routing can achieve considerable performance gains in Cognitive Radio Ad hoc Networks (CRAHNs). Most of these studies focused on designing appropriate metrics to select and prioritize the forwarding candidates. However, in multiple‐flow networks, a small number of nodes may always be with the higher priority order for different flows. Thus, some nodes may easily become overloaded with too much traffic and be severely congested. To overcome this problem, we propose a load balancing opportunistic routing (LBOR) scheme to maximize the total throughput of the whole network. We first formulate the problem of maximizing the total throughput of the network as a linear programming problem. Then, we develop heuristic load balancing candidate forwarder sorting and selection algorithms. Simulation results and comparisons demonstrate that our proposed LBOR scheme outperforms existing opportunistic routing protocols with nonload balancing methods in CRAHNs.
Wenxuan Duan, Xing Tang 0001, Junwei Zhou 0002, Jing Wang 0063, Guosheng Zhou
Wirel. Commun. Mob. Comput.3
2017 Robust Facial Landmark Localization Using LBP Histogram Correlation Based Initialization
abstract
Facial landmark localization on images with occlusions is an important and challenging task in many visual applications. Recently, the cascaded pose regression has attracted increasing attention, since it achieved superior performance in terms of facial landmark localization under occlusions. However, such approach is sensitive to initialization, where an improper initialization will decrease the performance sharply. In this paper, we propose a novel initialization method to get a robust initial shape by analysing correlation of Local Binary Patterns (LBP) histograms between the estimated face and training faces. The shape of the training face that is most correlated with the estimated face, will be selected as the initialization for the regression. The selected shape is closer to the real shape of the estimated face, which makes the landmark localization more accurate. Besides, in order to make the initial shape more robust to occlusions, we propose a boosted smart restarts technique by checking location and occlusion jointly instead of checking location only. We show that the proposed method significantly improves performance over existing landmark localization methods on the challenging dataset of COFW. The experimental results demonstrate that the proposed method reduces error by 11.9% and failure cases by 20.8% on COFW dataset. Moreover, it detects face occlusions with 85/40% precision/recall.
Yiyun Pan, Junwei Zhou 0002, Yongsheng Gao 0001, Jianwen Xiang, Shengwu Xiong 0001, Yanchao Yang 0002
FG2
2017 A Modified Segmentation Approach for Overlapping Elliptical Objects with Various Sizes
Guanghui Zhao 0004, Xingyan Zi, Kaitai Liang, Panyi Yun, Junwei Zhou 0002
GPC5
2017 Securing Outsourced Data in the Multi-Authority Cloud with Fine-Grained Access Control and Efficient Attribute Revocation
abstract
Data outsourcing is a promising service for data owners, where their data are stored on a cloud storage provider. Since the cloud is not fully trusted, data access control has become a challenging issue in the Cloud Storage System (CSS). Ciphertext-Policy Attribute-Based Encryption (CP-ABE) is a feasible technique for ensuring access control in the CSS, where an attribute authority is responsible to manage attributes and distribute keys. In this paper, we propose a novel revocable Multi-Authority CP-ABE scheme, in which the access policy can be constructed as an arbitrary tree rather than a matrix used by existing schemes. The tree-like policy makes our scheme more flexible. Consequently, the encryption, decryption and attribute revocation operations are also more efficient. Our scheme is also proved to be secure under the standard assumption. It can resist user collusion attack, while the attribute revocation operation also achieves both forward security and backward security. Simulation results show that our scheme is highly efficient.
Junwei Zhou 0002, Hui Duan, Kaitai Liang, Qiao Yan, Fei Chen 0003, F. Richard Yu, Jieming Wu, Jianyong Chen
Comput. J.1
2017 How to Share Secret Efficiently over Networks
abstract
In a secret-sharing scheme, the secret is shared among a set of shareholders, and it can be reconstructed if a quorum of these shareholders work together by releasing their secret shares. However, in many applications, it is undesirable for nonshareholders to learn the secret. In these cases, pairwise secure channels are needed among shareholders to exchange the shares. In other words, a shared key needs to be established between every pair of shareholders. But employing an additional key establishment protocol may make the secret-sharing schemes significantly more complicated. To solve this problem, we introduce a new type of secret-sharing, calledprotected secret-sharing(PSS), in which the shares possessed by shareholders not only can be used to reconstruct the original secret but also can be used to establish the shared keys between every pair of shareholders. Therefore, in the secret reconstruction phase, the recovered secret is only available to shareholders but not to nonshareholders. In this paper, an information theoretically secure PSS scheme is proposed, its security properties are analyzed, and its computational complexity is evaluated. Moreover, our proposed PSS scheme also can be applied to threshold cryptosystems to prevent nonshareholders from learning the output of the protocols.
Lein Harn, Ching-Fang Hsu 0001, Zhe Xia, Junwei Zhou 0002
Secur. Commun. Networks4
2016 Word Embeddings and Convolutional Neural Network for Arabic Sentiment Classification
abstract
With the development and the advancement of social networks, forums, blogs and online sales, a growing number of Arabs are expressing their opinions on the web. In this paper, a scheme of Arabic sentiment classification, which evaluates and detects the sentiment polarity from Arabic reviews and Arabic social media, is studied. We investigated in several architectures to build a quality neural word embeddings using a 3.4 billion words corpus from a collected 10 billion words web-crawled corpus. Moreover, a convolutional neural network trained on top of pre-trained Arabic word embeddings is used for sentiment classification to evaluate the quality of these word embeddings. The simulation results show that the proposed scheme outperforms the existed methods on 4 out of 5 balanced and unbalanced datasets.
Abdelghani Dahou, Shengwu Xiong 0001, Junwei Zhou 0002, Mohamed Houcine Haddoud, Pengfei Duan 0005
COLING3
2016 Wave atom transform based image hashing using distributed source coding
Yanchao Yang 0002, Junwei Zhou 0002, Feipeng Duan, Fang Liu 0024, Lee-Ming Cheng
J. Inf. Secur. Appl.2
2016 An Efficient File Hierarchy Attribute-Based Encryption Scheme in Cloud Computing
abstract
Ciphertext-policy attribute-based encryption (CP-ABE) has been a preferred encryption technology to solve the challenging problem of secure data sharing in cloud computing. The shared data files generally have the characteristic of multilevel hierarchy, particularly in the area of healthcare and the military. However, the hierarchy structure of shared files has not been explored in CP-ABE. In this paper, an efficient file hierarchy attribute-based encryption scheme is proposed in cloud computing. The layered access structures are integrated into a single access structure, and then, the hierarchical files are encrypted with the integrated access structure. The ciphertext components related to attributes could be shared by the files. Therefore, both ciphertext storage and time cost of encryption are saved. Moreover, the proposed scheme is proved to be secure under the standard assumption. Experimental simulation shows that the proposed scheme is highly efficient in terms of encryption and decryption. With the number of the files increasing, the advantages of our scheme become more and more conspicuous.
Shulan Wang, Junwei Zhou 0002, Joseph K. Liu, Jianyong Chen, Weixin Xie
IEEE Trans. Inf. Forensics Secur.2
2015 Distributed arithmetic coding with interval swapping
Junwei Zhou 0002, Kwok-Wo Wong, Yanchao Yang 0002
Signal Process.1
2011 Improvement of Security and Feasibility for Chaos-Based Multimedia Cryptosystem
Jianyong Chen, Junwei Zhou 0002
ICCSA (4)2