Zhen Qin 0002

dblp:06/864-0002 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
34since 2021 · last 2026
0000-0001-7857-9719ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Computer networks · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Revocable Policy Hiding Bilateral Access Control Scheme With Equality Test for IIoT Environment
abstract
With the development of the Industrial Internet of Things (IIoT), the amount of data generated by industrial manufacturing equipment is growing rapidly, creating a critical need for secure and efficient cloud-based data sharing. While bilateral access control enables both data senders and receivers to define their own fine-grained access policies, existing schemes transmit these policies in plaintext, exposing sensitive operational metadata and creating security vulnerabilities. Furthermore, they lack essential functionalities for practical IIoT deployment, including efficient duplicate data detection, a robust user revocation mechanism, and computationally lightweight operations suitable for resource-constrained devices. To address these challenges, this paper proposes a novel policy-hiding, lightweight, and revocable bilateral access control scheme with equality test (RBAC-ET) for cloud-enabled IIoT systems. RBAC-ET is the first scheme to integrally combine five critical capabilities: first, fine-grained bilateral access control using Linear Secret Sharing Schemes (LSSS) supporting AND/OR logical operations. Second, policy confidentiality is achieved by encrypting access policies within the ciphertext to prevent the leakage of sensitive relationships. Third, the equality test functionality enables cloud servers to identify duplicate ciphertexts for efficient storage and processing without requiring decryption. Fourth, a proposed lightweight design is based on elliptic curve scalar multiplication operations instead of the more resource-intensive bilinear pairing operations. Fifth, a robust user revocation mechanism that ensures both forward and backward secrecy in dynamic industrial environments. Theoretical analysis and experimental results demonstrate that RBAC-ET offers superior security and functionality while maintaining computational efficiency comparable to that of state-of-the-art schemes, making it a scalable and practical solution for modern IIoT data-sharing applications.
Junaid Hassan, Zhen Qin 0002, Muhammad Arslan Rauf, Muhammad Umar Aftab, Negalign Wake Hundera, Xinzhong Zhu
IEEE Internet Things J.2
2026 Nonprofiled Incremental Learning-Based Side-Channel Attack on Lattice-Based KEMs: The Case Study of Kyber
abstract
Post-quantum key-encapsulation mechanisms (KEMs) usually use the well-known Fujisaki-Okamoto (FO) transformation during key decapsulation to achieve chosen ciphertext attack (CCA) security. In the FO transformation, the re-encryption procedure depends on the message. The side-channel leakage of re-encryption can be exploited to recover the coefficients of the secret key under chosen ciphertexts, even if the KEM scheme is CCA secure. However, it still requires a large amount of trace during the profiling and attacking phase. In this work, we introduce the first non-profiled deep learning-based attack on lattice-based KEMs. We propose a chosen ciphertexts method that is suitable for non-profiled deep learning-based attacks. The horizontal chunking strategy is employed to partition the coefficients into chunks, enabling the independent recovery of multiple secret key coefficients within each chunk. Then, we adopt an incremental learning strategy to allow the deep learning model to gradually learn the knowledge of each chunk. Moreover, we push the limits of traditional non-profiled deep learning-based attacks by combining the unsupervised domain adaptation with the correlation distinguisher, eliminating the necessity of neural network training for each possible key guess. The feasibility of the attack is verified by practical experiments for the unprotected and masked implementations of Kyber on the ARM Cortex-M4.
Yanbin Li 0001, Yikang Guo, Xinru Cong, Chunpeng Ge 0001, Zhen Qin 0002, Willy Susilo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 Leakage-Resilient Multi-Party Signatures for Industrial IoT via Cryptographic Reverse Firewalls
abstract
The rapid growth of Industrial Internet of Things (IIoT) systems has heightened the need for secure cryptographic operations, particularly multi-party digital signatures for decentralized trust. However, existing multi-party signature schemes are vulnerable to insider attacks, and no current solutions address insider-induced data exfiltration effectively. Cryptographic Reverse Firewalls (CRFs) provide a promising solution but face challenges in integration with digital signatures, especially with hash-dependent components. We propose MCRF, a CRF-enhanced multi-party signature scheme designed for IIoT environments. MCRF uses a commitment-based mechanism to optimize the signing process, enabling output-side CRF to re-randomize signatures without compromising correctness. The two-stage CRF architecture-input-side CRF for sanitizing messages and output-side CRF for re-randomizing signatures-ensures efficient protection against both input-triggered and output-stealth exfiltration attacks. MCRF offers strong leakage resistance with minimal computational and communication overhead, making it a scalable solution for secure multi-party signing in IIoT systems.
Zengxiang Wang, Yunfan Hu, Zhen Qin 0002, Hu Xiong
IEEE Trans. Dependable Secur. Comput.3
2026 Blockchain-Oriented Certificateless Threshold Signature With Identifiable Abort for Federated Learning in Digital Twin-Assisted IoV
abstract
As a promising subdomain of intelligent transportation systems (ITS), Internet of Vehicles (IoV) can be empowered by digital twin (DT) technology for real-time traffic simulation and artificial intelligence (AI)-driven predictive analytics in evolutionary trend projection, demonstrating significant potential in dynamic transportation optimization. Among various machine learning paradigms, federated learning (FL) not only aligns well with IoV, but also provides it with privacy protection. Traditional FL faces single point of failure due to the existence of an aggregation center, so blockchain-based FL with multiple aggregators is utilized to mitigate this issue. Nevertheless, in such distributed environments, both aggregators and model parameters exposed to network are vulnerable to attacks, impeding the normal operation of FL. In this paper, for blockchain-enabled FL with multiple aggregators in IoV, we propose CLTSwNI&IA, the first non-interactive certificateless threshold signature with identifiable abort. This scheme eliminates certificate management and key escrow, adopts a blockchain-oriented approach by utilizing a fully distributed signing paradigm. Additionally, the proposed signing scheme is capable of identifying malicious FL aggregators during the entire process through distributed fine-grained verification and ensuring the integrity of aggregation results. Finally, theoretical and experimental comparisons with related literature demonstrate the advanced functionality and the acceptable efficiency of our approach.
Yunfan Hu, Zengxiang Wang, Hu Xiong, Liming Fang 0001, Changgen Peng, Abubaker Wahaballa, Zhen Qin 0002, Zhiguang Qin
IEEE Trans. Intell. Transp. Syst.8
2025 CPSNet: Comprehensive Enhancement Representation for Polyp Segmentation Task
abstract
Accurately segmenting polyp regions in colonoscopy images is crucial for the diagnosis and intervention of colorectal cancer. However, the task of polyp segmentation remains challenging due to the diverse size and shape variations among polyps, their extreme similarity to the background, and frequent rotation of the lens, which further increases the diversity in polyp presentation. To address these challenges effectively, we propose a comprehensive polyp segmentation network (CPSNet). Specifically, we introduce a Comprehensive Spatial Feature Extraction Module (CFEM) that progressively and densely integrates features while forming receptive windows with various shapes. This enables enhanced perception of polyps at manifold sizes and shapes. Additionally, we propose a Fine-grained Region Strengthen Module (FGSM) to supplement uncertain areas around polyps by mitigating background noise interference. In terms of training strategy, we further introduce a Rotation-augmented Constrained Loss (RC Loss), which reinforces consistency constraints on polyp images under multiple rotation angles. Qualitative and quantitative experiments conducted on five public datasets demonstrate both the plug-and-play capability of CFEM as well as the effectiveness and excellence achieved by CPSNet.
Jiati Cai, Hongjie Yang, Yi Ding 0003, Ting Zhong, Zhen Qin 0002
ICASSP6
2025 Multi-scale Graph Convolution with Corrective Contrastive Learning for Skeleton-based Action Recognition
abstract
For pursuing accurate skeleton-based action recognition, many existing graph-based approaches deploy the higher-order polynomials of the skeletal adjacency matrix to model the node correlations of distant neighbours. To further capture robust graphical patterns, a novel multi-scale graph convolution operator is proposed, which enables to aggregate multi-scale dependencies and capture long-range joint relationships on human skeleton graph. Additionally, a novel corrective contrastive learning strategy is proposed, which aims to distinguish the representative clues and calibrate the confused action clips in the feature space. Comprehensive experiments validate the effectiveness and superiority of our proposed method over state-of-the-art approaches on three large-scale datasets: NTU RGB+D 60, NTU RGB+D 120, and Kinetics Skeleton 400.
Tianming Zhuang, Erqiang Zhou, Hanwen Zhang 0009, Yi Ding 0003, Ji Geng 0001, Zhen Qin 0002
ICASSP6
2025 SymND: Detecting Backdoor Attacks in Self-Supervised Facial Representation Tasks
abstract
Facial image tasks present distinct challenges in self-supervised learning (SSL) that are not encountered in general image classification, with existing backdoor attacks and defenses often fail to handle these specific issues. This paper introduces SymND, the first defense framework specifically designed to counter backdoor attacks in facial image SSL scenarios. SymND innovatively assesses noise stability across images and dynamically adjusts noise placement, capitalizing on the symmetrical properties of facial triggers—a departure from traditional SSL defenses that presume static trigger locations. Our method’s efficacy is underscored by experiments conducted on RAF-DB and UTKFace datasets, which show a significant reduction in attack success rates, plummeting from 99.58% to 0.15%, across a variety of downstream tasks employing different encoders.
Liyue Zhu, Changchun Yin, Liming Fang 0001, Zhen Qin 0002
ICME4
2025 Open-world multi-modal machine learning decision model based on uncertain data analysis for fetal heart diagnosis
Guosong Zhu, Zhen Qin 0002, Hu Xiong, Saru Kumari, Mohammed J. F. Alenazi, Yingkun Guo, Chien-Ming Chen 0001
Inf. Sci.2
2025 RLL-SWE: A Robust Linked List Steganography Without Embedding for intelligence networks in smart environments
abstract
With the rapid development of technology, smart environments utilizing the Internet of Things, artificial intelligence, and big data are improving the quality of life and work efficiency through connected devices. However, these advances present significant security challenges. The data generated by these smart devices contains many private and sensitive information. In data transmission, crime and terrorism may intercept this sensitive information and use it for secret communications and illegal activities. Steganography hides information in media files and prevents information leakage and interception by criminal and terrorist networks in an intelligent environment. It is an important technology to protect data integrity and security. Traditional steganography techniques often cause detectable distortions, whereas Steganography Without Embedding (SWE) avoids direct modification of cover media, thereby minimizing detection risks. This paper introduces an innovative and robust technique called Robust Linked List (RLL)-SWE, which improves resistance to attacks compared to traditional methods. Using multiple median downsampling and gradient calculations, this method extracts stable features. It restructures them into a multi-head unidirectional linked list, ensuring accurate message retrieval and high resistance to adversarial attacks. Comprehensive analysis and simulation experiments confirm the technique’s exceptional effectiveness and steganographic capacity.
Pengbiao Zhao, Yuanjian Zhou, Salman Ijaz 0002, Fazlullah Khan, Jingxue Chen, Bandar Alshawi, Zhen Qin 0002, Md. Arafatur Rahman
J. Netw. Comput. Appl.7
2025 Driving mutual advancement of 3D reconstruction and inpainting for masked faces
Guosong Zhu, Zhen Qin 0002, Erqiang Zhou, Yi Ding 0003, Zhiguang Qin
Pattern Recognit.2
2025 Edge-Adaptive Dynamic Scalable Convolution for Efficient Remote Mobile Pathology Analysis
abstract
With the emergence of edge computing, there is a growing need for advanced technologies capable of real-time, efficient processing of complex data on edge devices, particularly in mobile health systems handling pathological images. On edge computing devices, the lightweighting of models and reduction of computational requirements not only save resources but also increase inference speed. Although many lightweight models and methods have been proposed in recent years, they still face many common challenges. This article introduces a novel convolution operation, Dynamic Scalable Convolution (DSC), which optimizes computational resources and accelerates inference on edge computing devices. DSC is shown to outperform traditional convolution methods in terms of parameter efficiency, computational speed, and overall performance, through comparative analyses in computer vision tasks like image classification and semantic segmentation. Experimental results demonstrate the significant potential of DSC in enhancing deep neural networks, particularly for edge computing applications in smart devices and remote healthcare, where it addresses the challenge of limited resources by reducing computational demands and improving inference speed. By integrating advanced convolution technology and edge computing applications, DSC offers a promising approach to support the rapidly developing mobile health field, especially in enhancing remote healthcare delivery through mobile multimedia communication.
Dajiang Chen, Zhen Qin 0002, Mingsheng Cao 0001, Rui-dong Chen
ACM Trans. Auton. Adapt. Syst.3
2025 DSDC-GCN: Decoupled Static-Dynamic Co-Occurrence Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
The existing approaches for skeleton-based action recognition based on graph convolutional networks (GCNs) primarily emphasize the construction of human skeletal structure by leveraging inherent connections. However, the static skeletal topology used across all action categories fails to capture discriminative relationships between joint pairs, while current graph structures struggle to model dynamic motion information, limiting their ability to represent both temporal and motion-specific dependencies. To address this limitation, we propose the decoupled static-dynamic co-occurrence graph convolution (DSDC-GConv), which specifically aims to learn and adapt the graph topology by refining the inter-frame and intra-frame joint dependencies through decomposed manner. Additionally, a multi-level context-aware module is proposed to comprehensively model the latent saliencies of multiple domains in skeletal sequences. This module refines the spatial nodes, temporal dynamics, channel-wise characteristics, and motional dependencies within the graph convolution block. Furthermore, a hierarchical densely connected temporal convolution is proposed to enhance the representation of local features through partial dense connections and enrich the temporal information during the convolution process. Findings from our evaluations on five large-scale benchmark datasets (i.e., NTU RGB+D 60, NTU RGB+D 120, Kinetics Skeleton 400, Northwestern-UCLA, PKU-MMD) demonstrate the effectiveness and superiority of our proposed method over competing approaches, with an recognition accuracy of 93.0% and 97.1% on NTU RGB+D 60, 89.9% and 90.6% on NTU RGB+D 120, 38.6% and 63.4% on Kinetics Skeleton 400, 97.4% on Northwestern-UCLA, 97.6% and 63.6% on PKU-MMD.
Tianming Zhuang, Zhen Qin 0002, Yi Ding 0003, Zhiguang Qin, Ji Geng 0001, Kim-Kwang Raymond Choo
IEEE Trans. Circuits Syst. Video Technol.2
2025 E-harnet: an efficient hybrid transformer network for human activity recognition
Muhammad Arslan Rauf, Salim, M. D. Shakib Mahamud, Mian Muhammad Yasir Khalil, Zhen Qin 0002
J. Supercomput.6
2025 TransMatch: Employing Bridging Strategy to Overcome Large Deformation for Feature Matching in Gastroscopy Scenario
abstract
Feature matching is widely applied in the image processing field. However, both traditional feature matching methods and previous deep learning-based methods struggle to accurately match the features with severe deformations and large displacements, particularly in gastroscopy scenario. To fill this gap, an effective feature matching framework named TransMatch is proposed, which addresses the largely displacements issue by matching features with global information leveraged via Transformer structure. To address the severe deformation of features, an effective bridging strategy with a novel bidirectional quadratic interpolation network is employed. This bridging strategy decomposes and simplifies the matching of features undergoing severe deformations. A deblurring module for gastroscopy scenario is specifically designed to address the potential blurriness. Experiments have illustrated that proposed method achieves state-of-the-art performance of feature matching and frame interpolation in gastroscopy scenario. Moreover, a large-scale gastroscopy dataset is also constructed for multiple tasks.
Guosong Zhu, Zhen Qin 0002, Linfang Yu, Yi Ding 0003, Zhiguang Qin
IEEE Trans. Medical Imaging2
2024 Gradient Saliency-aware CutMix for Semi-Supervised Medical Image Segmentation
abstract
In semi-supervised medical image segmentation, the use of CutMix in the Mean Teacher architecture is considered an effective strong data augmentation strategy. However, we believe that randomly selecting patches from the source image might mislead the model into learning unexpected feature representations. Therefore, we propose Gradient Saliency-aware CutMix for semi-supervised medical image segmentation (GSC-Seg). Utilizing the gradient from pre-trained models to detect salient regions and then copies and pastes the large gradient areas from labeled data into corresponding areas of unlabeled data based on the gradient, and vice versa, guiding the model to learn more appropriate feature representations. Furthermore, we propose a gradient augmentation strategy, which generates disruptions in the gradient through the network itself and enhances the gradient representation abilities of the network. Experiment results show that our approach achieves the state-of-the-art performance on three medical image segmentation datasets. Code is available at https://github.com/UESTC-Med424-JYX/GSC-Seg.
Guobin Zhu, Yi Ding 0003, Zhen Qin 0002, Minghui Pang
ICME4
2024 A privacy protection scheme for green communication combining digital steganography
Pengbiao Zhao, Bintao Wang, Zhen Qin 0002, Yi Ding 0003, Kim-Kwang Raymond Choo
Peer Peer Netw. Appl.3
2024 C2FResMorph: A high-performance framework for unsupervised 2D medical image registration
Yi Ding 0003, Junjian Bu, Zhen Qin 0002, Mingsheng Cao 0001, Zhiguang Qin, Minghui Pang
Pattern Recognit.3
2024 A cascaded framework with cross-modality transfer learning for whole heart segmentation
Yi Ding 0003, Dan Mu, Zhen Qin 0002, Zhiguang Qin, Yingkun Guo
Pattern Recognit.4
2024 MFSSE: Multi-Keyword Fuzzy Ranked Symmetric Searchable Encryption With Pattern Hidden in Mobile Cloud Computing
abstract
In this paper, we propose a novel Multi-keyword Fuzzy Symmetric Searchable Encryption (SSE) with patterns hidden, namely MFSSE. In MFSSE, the search trapdoor can be modified differently each time even if the keywords are the same when performing multi-keyword search to prevent the leakage of search patterns. Moreover, MFSSE modifies the search trapdoor by introducing random false negative and false positive errors to resist access pattern leakage. Furthermore, MFSSE utilizes efficient cryptographic algorithms (e.g., Locality-Sensitive Hashing) and lightweight operations (such as, integer addition, matrix multiplication, etc.) to minimize computational and communication, and storage overheads on mobile devices while meeting security and functional requirements. Specifically, its query process requires only a single round of communication, in which, the communication cost is linearly related to the number of the documents in the database, and is independent of the total number of keywords and the number of queried keywords; its computational complexity for matching a document is$O(1)$; and it requires only a small amount of fixed local storage (i.e., secret key) to be suitable for mobile scenarios. The experimental results demonstrate that MFSSE can prevent the leakage of access patterns and search patterns, while keeping a low communication and computation overheads.
Dajiang Chen, Zeyu Liao, Zhidong Xie, Rui-dong Chen, Zhen Qin 0002, Mingsheng Cao 0001, Hongning Dai, Kuan Zhang 0001
IEEE Trans. Cloud Comput.5
2024 Enhanced Pseudo-Label Generation With Self-Supervised Training for Weakly- Supervised Semantic Segmentation
abstract
Due to the high cost of pixel-level labels required for fully-supervised semantic segmentation, weakly-supervised segmentation has emerged as a more viable option recently. Existing weakly-supervised methods tried to generate pseudo-labels without pixel-level labels for semantic segmentation, but a common problem is that the generated pseudo-labels contain insufficient semantic information, resulting in poor accuracy. To address this challenge, a novel method is proposed, which generates class activation/attention maps (CAMs) containing sufficient semantic information as pseudo-labels for the semantic segmentation training without pixel-level labels. In this method, the attention-transfer module is designed to preserve salient regions on CAMs while avoiding the suppression of inconspicuous regions of the targets, which results in the generation of pseudo-labels with sufficient semantic information. A pixel relevance focused-unfocused module has also been developed for better integrating contextual information, with both attention mechanisms employed to extract focused relevant pixels and multi-scale atrous convolution employed to expand receptive field for establishing distant pixel connections. The proposed method has been experimentally demonstrated to achieve competitive performance in weakly-supervised segmentation, and even outperforms many saliency-joined methods.
Zhen Qin 0002, Guosong Zhu, Erqiang Zhou, Yingjie Zhou 0001, Yicong Zhou, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.1
2024 Backdoor Attack on Deep Learning-Based Medical Image Encryption and Decryption Network
abstract
Medical images often contain sensitive information, and one typical security measure is to encrypt medical images prior to storage and analysis. A number of solutions, such as those utilizing deep learning, have been proposed for medical image encryption and decryption. However, our research shows that deep learning-based encryption models can potentially be vulnerable to backdoor attacks. In this paper, a backdoor attack paradigm for encryption and decryption network is proposed and corresponding attacks are respectively designed for encryption and decryption scenarios. For attacking the encryption model, a backdoor discriminator is adopted, which is randomly trained with the normal discriminator to confuse the encryption process. In the decryption scenario, a number of subnetwork parameters are replaced and the subnetwork can be activated when detecting the trigger embedded into the input (encrypted image) to degrade the decryption performance. Considering the model performance degradation due to parameter replacement, the model pruning is also adopted to further strengthen the attacking performance. Furthermore, the image steganography is adopted to generate invisible triggers for each image; subsequently, improving the stealthiness of backdoor attacks. Our research on designing backdoor attacks for encryption and decryption network can serve as an attacking mode for such networks, and provides another research direction for improving the security of such models. This research is also one of the earliest works to realize the backdoor attack on the deep learning based medical encryption and decryption network to evaluate the security performance of these networks. Extensive experimental results show that the proposed method can effectively threaten the security performance both for the encryption and decryption network.
Yi Ding 0003, Zhen Qin 0002, Erqiang Zhou, Guobin Zhu, Zhiguang Qin, Kim-Kwang Raymond Choo
IEEE Trans. Inf. Forensics Secur.3
2024 LightNet: A Novel Lightweight Convolutional Network for Brain Tumor Segmentation in Healthcare
abstract
Diagnosis, treatment planning, surveillance, and the monitoring of clinical trials for brain diseases all benefit greatly from neuroimaging-based tumor segmentation. Recently, Convolutional Neural Networks (CNNs) have demonstrated promising results in enhancing the efficiency of image-based brain tumor segmentation. Most current work on CNNs, however, is devoted to creating increasingly complicated convolution modules to improve performance, which in turn raises the computing cost of the model. This work proposes a simple and effective feed-forward CNN, LightNet (Light Network). Based on multi-path and multi-level, it replaces traditional convolutional methods with light operations, which reduces network parameters and redundant feature maps. In the up-sampling stage, a light channel attention module is added to achieve richer multi-scale and spatial semantic feature information extraction of brain tumor. The performance of the network is evaluated in the Multimodal Brain Tumor Segmentation Challenge (BraTS 2015) dataset, and results are presented here alongside other high-performing CNNs. Results show comparable accuracy with other methods but with increased efficiency, segmentation performance, and reduced redundancy and computational complexity. The result is a high-performing network with a balance between efficiency and accuracy, allowing, for example, better energy performance on mobile devices.
Dongyuan Wu, Junyi Tao, Zhen Qin 0002, Rao Asad Mumtaz, Linfang Yu, Jane Courtney
IEEE J. Biomed. Health Informatics3
2024 MFNet:Real-Time Motion Focus Network for Video Frame Interpolation
abstract
As a popular research topic in computer vision, video frame interpolation is widely used in video processing tasks. However, this task is often limited by slow processing speed or high memory consumption in practical applications. To address these drawbacks, a frame interpolation network focusing on motion regions named MFNet is proposed, which consists of a sampler for adaptive and efficient separation of motion regions from the background, a fine-grained module for direct approximation of intermediate streams, and a lightweight module for bi-directional optical stream fusion. Extensive experiments show that our MFNet achieves optimal accuracy on some frame interpolation tasks and is much faster than other state-of-the-art methods. In addition, transplantation of the core components of MFNet to other frame interpolation networks can significantly improve the performance.
Guosong Zhu, Zhen Qin 0002, Yi Ding 0003, Yao Liu 0019, Zhiguang Qin
IEEE Trans. Multim.2
2023 Diff-SFCT: A Diffusion Model with Spatial-Frequency Cross Transformer for Medical Image Segmentation
abstract
Most existing semantic segmentation methods primarily employ supervised learning with discriminative models. Although these methods are straightforward, they overlook the modeling of underlying data distributions. In this paper, we propose a novel medical image segmentation framework called Diff-SFCT based on Diffusion Model. We formulate semantic segmentation as a generative problem for segmentation masks, replacing the conventional pixel-wise discriminative learning with a latent prior learning process to produce more accurate segmentation results. Diff-SFCT employs a backbone network combining Convolutional Neural Network (CNN) and Transformer, and utilizes the local perception of CNN and the global information modeling capability of Transformer. In Diff-SFCT, we design a Semantic Encoder that effectively extracts fine-grained semantic features from real images. Meanwhile, we propose a novel Spatial-Frequency Cross Transformer (SFCT) framework, which can effectively model and interact the global features of the diffuse noise mask and the real semantic features, reducing the domain gap between the two and enhancing the model’s representational capacity. Additionally, to preserve spatial and frequency information in the diffusion model, we design a Spatial-Frequency Attention Module (SFAM) as part of the Convolutional Block. This module improves the model’s spatial and frequency perception abilities while incurring negligible computational overhead. Experimental results evince that our DiffSFCT substantially outperforms other segmentation methods, exhibiting remarkable performance across various medical image segmentation datasets.
Yi Ding 0003, Guobin Zhu, Zhen Qin 0002, Minghui Pang, Mingsheng Cao 0001
BIBM4
2023 FastNet: A Lightweight Convolutional Neural Network for Tumors Fast Identification in Mobile-Computer-Assisted Devices
abstract
Histopathology diagnosis is an important standard for breast tumors identifying. However, histopathology image analysis is complex, tedious and error-prone, due to the super-resolution image. In recent years, deep learning technology has been successfully applied to histopathology image analysis and made great progress. The well-known deep neural networks usually have tens of million parameters, which consume much memory to deploy the state-of-the-art model. In addition, deep neural networks rely on high-performance hardware resources, which impede the deployment of state-of-the-art model on portable equipment. In this work, a novel framework which consists of a weight accumulation method and a lightweight fast neural network (FastNet) was proposed for tumor fast identification (TFI) in mobile computer-assisted devices. The weight accumulation method was designed to obtain the tissue mask regions of interest and remove the useless background area in histopathology images, which greatly reduces the redundant computation cost. Furthermore, we proposed the lightweight FastNet to improve the computational efficiency on mobile devices. A novel attention loss function was designed and applied in FastNet. The attention loss function pays more attention on the positive samples and the indistinguishable samples, which greatly improves performance. The proposed FastNet was compared with three state-of-the-art methods commonly used for image classification and object detection. Experimental results indicated that FastNet achieves highest recall of 96.94%, highest F1 score of 97.33% and highest accuracy of 97.34%, besides least trainable parameters of 0.22M and smallest floating point operations of 210M FLOPs.
Zhen Qin 0002, Dajiang Chen, Ning Zhang 0007, Yi Ding 0003, Fuhu Deng, Zhiguang Qin, Minghui Pang
IEEE Internet Things J.2
2023 MSDP: multi-scheme privacy-preserving deep learning via differential privacy
abstract
Abstract Human activity recognition (HAR) generates a massive amount of the dataset from the Internet of Things (IoT) devices, to enable multiple data providers to jointly produce predictive models for medical diagnosis. That the accuracy of the models is greatly improved when trained on a large number of datasets from these data providers on the untrusted cloud server is very significant and raises privacy concerns. With the migration of a deep neural network (DNN) in the learning experience in HAR, we present a privacy-preserving DNN model known as Multi-Scheme Differential Privacy (MSDP) depending on the fusion of Secure Multi-party Computation (SMC) and 𝜖-differential privacy, making it very practical since existing proposals are unable to make all the fully homomorphic encryption multi-key which is very impracticable. MSDP inputs a secure multi-party alternative to the ReLU function to reduce the communication and computational cost at a minimal level. With the aid of experimental verification on the four of the most widely used human activity recognition datasets, MSDP demonstrates superior performance with very good generalization performance and is proven to be secure as compared with existing ultramodern models without breach of privacy.
Kwabena Owusu-Agyemang, Zhen Qin 0002, Hu Xiong, Yao Liu 0019, Tianming Zhuang, Zhiguang Qin
Pers. Ubiquitous Comput.2
2023 Interpreting Universal Adversarial Example Attacks on Image Classification Models
abstract
Mitigating adversarial deep learning attacks remains challenging, partly because of the ease and low cost in carrying out such attacks. Therefore, in this paper, we focus on the understanding of universal adversarial example attack on image classification models. Specifically, we seek to understand the difference(s) between adversarial examples in two adversarial datasets (DAmageNet and PGD dataset) and clean examples in ImageNet learned by the classification model, and whether we can use such findings to resist adversarial example attacks. We also seek to determine if we can retrain a discriminator to discriminate whether the input image is an adversarial example, using adversarial training. We then design a number of experiments (e.g., class activation map (CAM) analysis, feature map analysis, feature maps/filters changing, adversarial training, and binary classification model) to help us determine whether the universal adversarial dataset can be successfully used to attack the classification model. This, in turn, contributes to a better understanding of adversarial defenses over pretrained classification model from an interpretation perspective. To the best of our knowledge, this work is one of the earliest works to systematically investigate the interpretation of universal adversarial example attack on image classification models, both visually and quantitatively.
Yi Ding 0003, Fuyuan Tan, Ji Geng 0001, Zhen Qin 0002, Mingsheng Cao 0001, Kim-Kwang Raymond Choo, Zhiguang Qin
IEEE Trans. Dependable Secur. Comput.4
2022 Generic network for domain adaptation based on self-supervised learning and deep clustering
abstract
Domain adaptation methods train a model to find similar feature representations between a source and target domain. Recent methods leverage self-supervised learning to discover the analogous representations of the two domains. However, prior self-supervised methods have three significant drawbacks: (1) leveraging pretext tasks that are susceptible to learning low-level representations, (2) aligning the two domains using adversarial loss without considering if the extracted features are low-level representations, (3) the models are not flexible to accommodate various proportions of target labels, i.e., they assume target labels are always available. This paper presents a Generic Domain Adaptation Network (GDAN) to address these issues. First, we introduce a criterion based on instance discrimination to select appropriate pretext tasks to learn high-level domain invariant representations. Then, we propose a semantic neighbor cluster to align the two domain features. The semantic neighbor cluster implements a clustering technique in a feature embedding space to form clusters according to high-level semantic similarities. Finally, we present a weighted target loss function to balance the model weights according to the target labels. This loss function makes GDAN flexible for semi-supervised scenarios, i.e., partly labeled target data. We evaluate the proposed methods on four domain adaptation benchmark datasets. The experiment findings show that the proposed methods align the two domains well and achieve competitive results.
Adu Asare Baffour, Zhen Qin 0002, Ji Geng 0001, Yi Ding 0003, Fuhu Deng, Zhiguang Qin
Neurocomputing2
2022 Segmentation mask and feature similarity loss guided GAN for object-oriented image-to-image translation
Zhen Qin 0002, Qingya Chen, Yi Ding 0003, Tianming Zhuang, Zhiguang Qin, Kim-Kwang Raymond Choo
Inf. Process. Manag.1
2022 MVFusFra: A Multi-View Dynamic Fusion Framework for Multimodal Brain Tumor Segmentation
abstract
Medical practitioners generally rely on multimodal brain images, for example based on the information from the axial, coronal, and sagittal views, to inform brain tumor diagnosis. Hence, to further utilize the 3D information embedded in such datasets, this paper proposes a multi-view dynamic fusion framework (hereafter, referred to as MVFusFra) to improve the performance of brain tumor segmentation. The proposed framework consists of three key building blocks. First, a multi-view deep neural network architecture, which represents multi learning networks for segmenting the brain tumor from different views and each deep neural network corresponds to multi-modal brain images from one single view. Second, the dynamic decision fusion method, which is mainly used to fuse segmentation results from multi-views into an integrated method. Then, two different fusion methods (i.e., voting and weighted averaging) are used to evaluate the fusing process. Third, the multi-view fusion loss (comprising segmentation loss, transition loss, and decision loss) is proposed to facilitate the training process of multi-view learning networks, so as to ensure consistency in appearance and space, for both fusing segmentation results and the training of the learning network. We evaluate the performance of MVFusFra on the BRATS 2015 and BRATS 2018 datasets. Findings from the evaluations suggest that fusion results from multi-views achieve better performance than segmentation results from the single view, and also implying effectiveness of the proposed multi-view fusion loss. A comparative summary also shows that MVFusFra achieves better segmentation performance, in terms of efficiency, in comparison to other competing approaches.
Yi Ding 0003, Ji Geng 0001, Zhen Qin 0002, Kim-Kwang Raymond Choo, Zhiguang Qin, Xiaolin Hou
IEEE J. Biomed. Health Informatics4
2022 DeepKeyGen: A Deep Learning-Based Stream Cipher Generator for Medical Image Encryption and Decryption
abstract
The need for medical image encryption is increasingly pronounced, for example, to safeguard the privacy of the patients' medical imaging data. In this article, a novel deep learning-based key generation network (DeepKeyGen) is proposed as a stream cipher generator to generate the private key, which can then be used for encrypting and decrypting of medical images. In DeepKeyGen, the generative adversarial network (GAN) is adopted as the learning network to generate the private key. Furthermore, the transformation domain (that represents the "style" of the private key to be generated) is designed to guide the learning network to realize the private key generation process. The goal of DeepKeyGen is to learn the mapping relationship of how to transfer the initial image to the private key. We evaluate DeepKeyGen using three data sets, namely, the Montgomery County chest X-ray data set, the Ultrasonic Brachial Plexus data set, and the BraTS18 data set. The evaluation findings and security analysis show that the proposed key generation network can achieve a high-level security in generating the private key.
Yi Ding 0003, Fuyuan Tan, Zhen Qin 0002, Mingsheng Cao 0001, Kim-Kwang Raymond Choo, Zhiguang Qin
IEEE Trans. Neural Networks Learn. Syst.3
2021 Underwater Information Sensing Method Based on Improved Dual-Coupled Duffing Oscillator Under Lévy Noise Description
Hanwen Zhang 0009, Zhen Qin 0002, Dajiang Chen
CollaborateCom (1)2
2021 Spatial self-attention network with self-attention distillation for fine-grained image recognition
abstract
The underlining task for fine-grained image recognition captures both the inter-class and intra-class discriminate features. Existing methods generally use auxiliary data to guide the network or a complex network comprising multiple sub-networks. They have two significant drawbacks: (1) Using auxiliary data like bounding boxes requires expert knowledge and expensive data annotation. (2) Using multiple sub-networks make network architecture complex and requires complicated training or multiple training steps. We propose an end-to-end Spatial Self-Attention Network (SSANet) comprising a spatial self-attention module (SSA) and a self-attention distillation (Self-AD) technique. The SSA encodes contextual information into local features, improving intra-class representation. Then, the Self-AD distills knowledge from the SSA to a primary feature map, obtaining inter-class representation. By accumulating classification losses from these two modules enables the network to learn both inter-class and intra-class features in one training step. The experiment findings demonstrate that SSANet is effective and achieves competitive performance.
Adu Asare Baffour, Zhen Qin 0002, Yong Wang 0046, Zhiguang Qin, Kim-Kwang Raymond Choo
J. Vis. Commun. Image Represent.2
2021 A Fuzzy Authentication System Based on Neural Network Learning and Extreme Value Statistics
abstract
Internet-connected smart devices in, on, and around us, (e.g., embedded devices, wearable devices, and smart sensors) can collect human biometric features and facilitate identity authentication. Existing approaches are mainly based on pattern recognition and machine learning algorithms, which may not be capable of processing uncertain user information. Thus, focusing on the uncertainty of users' identity, this article proposes a fuzzy authentication system based on neural network and extreme value analysis. Specifically, we utilize biometric gait information of human body recognition. Our proposed authentication system is designed to implicitly authenticate users based on their gait, and can detect uncertain users and reject the authentication of unknown users. The performance is evaluated using an open dataset of 153 volunteers, where we manage to achieve a recognition accuracy rate of 98.4% and an error rate of unauthorized users at 6%.
Zhen Qin 0002, Gu Huang, Hu Xiong, Zhiguang Qin, Kim-Kwang Raymond Choo
IEEE Trans. Fuzzy Syst.1
2020 Attacking the Dialogue System at Smart Home
Erqiang Deng, Zhen Qin 0002, Yi Ding 0003, Zhiguang Qin
CollaborateCom (1)2
2020 Brain tumor segmentation with deep convolutional symmetric neural network
Hao Chen 0047, Zhiguang Qin, Yi Ding 0003, Tian Lan 0005, Zhen Qin 0002
Neurocomputing5
2020 A framework for hierarchical division of retinal vascular networks
Linfang Yu, Zhen Qin 0002, Tianming Zhuang, Yi Ding 0003, Zhiguang Qin, Kim-Kwang Raymond Choo
Neurocomputing2
2019 Learning-Aided User Identification Using Smartphone Sensors for Smart Homes
abstract
Smart homes expects to improve the convenience, comfort, and energy efficiency of the residents by connecting and controlling various appliances. As the personal information and computing hub for smart homes, smartphones allow people to monitor and control their homes anytime and anywhere. Therefore, the security and privacy of smartphones and the stored data are crucial in smart homes. To protect smartphones from potential attacks, various built-in sensors can be utilized for user authentication/identification and access control to achieve enhanced security. In this paper, we propose a framework, smartphone sensor user identification (SSUI), in order to facilitate user identification based on the relationships between different types of sensor data and smartphone users. Specifically in SSUI, the time and frequency features are extracted and learned separately using convolution neural network (CNN). The CNN outputs are then processed using recurrent neural network, according to several time bins. Using both of our own dataset (collected from 17 participants) and a publicly available dataset (i.e., Heterogeneity Dataset for Human Activity Recognition), we demonstrate the effectiveness of the proposed SSUI framework, where we achieve an accuracy rate of over 91.45% in various scenarios.
Zhen Qin 0002, Lingzhou Hu, Ning Zhang 0007, Dajiang Chen, Kuan Zhang 0001, Zhiguang Qin, Kim-Kwang Raymond Choo
IEEE Internet Things J.1
2017 S2M: A Lightweight Acoustic Fingerprints-Based Wireless Device Authentication Protocol
abstract
Device authentication is a critical and challenging issue for the emerging Internet of Things (IoT). One promising solution to authenticate IoT devices is to extract a fingerprint to perform device authentication by exploiting variations in the transmitted signal caused by hardware and manufacturing inconsistencies. In this paper, we propose a lightweight device authentication protocol [named speaker-to-microphone (S2M)] by leveraging the frequency response of a speaker and a microphone from two wireless IoT devices as the acoustic hardware fingerprint. S2M authenticates the legitimate user by matching the fingerprint extracted in the learning process and the verification process, respectively. To validate and evaluate the performance of S2M, we design and implement it in both mobile phones and PCs and the extensive experimental results show that S2M achieves both low false negative rate and low false positive rate in various scenarios under different attacks.
Dajiang Chen, Ning Zhang 0007, Zhen Qin 0002, Xufei Mao, Zhiguang Qin, Xuemin Shen, Xiang-Yang Li 0001
IEEE Internet Things J.3
2016 On the security of two identity-based signature schemes based on pairings
Zhen Qin 0002, Chen Yuan 0002, Hu Xiong
Inf. Process. Lett.1
2014 Demographic information prediction based on smartphone application usage
abstract
Demographic information is usually treated as private data (e.g., gender and age), but has been shown great values in personalized services, advertisement, behavior study and other aspects. In this paper, we propose a novel approach to make efficient demographic prediction based on smartphone application usage. Specifically, we firstly consider to characterize the data set by building a matrix to correlate users with types of categories from the log file of smartphone applications. By considering the category-unbalance problem, we predict users' demographic information and propose an optimization method to further smooth the obtained results with category neighbors and user neighbors. The evaluation is supplemented by the dataset from real world workload. The results show advantages of the proposed prediction approach compared with baseline prediction. In particular, the proposed approach can achieve 81.21% of Accuracy in gender prediction. While in dealing with a more challenging multi-class problem, the proposed approach can still achieve good performance (e.g., 73.84% of Accuracy in the prediction of age group and 66.42% of Accuracy in the prediction of phone level).
Zhen Qin 0002, Yong Xia 0001, Hongrong Cheng, Yingjie Zhou 0001, Zhengguo Sheng, Victor C. M. Leung
SMARTCOMP1