Wentian Zhang

dblp:229/0641 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
18since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Self-Supervised Cross-Level Consistency Learning For Fundus Image Classification
abstract
The rapid development of intelligent systems for eye disease diagnosis decreases the risk of people suffering from vision impairment. However, the superior discrimination ability of existing retinal disease diagnosis methods heavily relies on the large-scale high-quality annotations. In this work, we adapt the self-supervised technique for fundus image classification with the merits of bypassing the over-dependence of labeled data. Unlike most current self-supervised approaches, which only learn global pre-text representations from view-level, our method further incorporates the region-level representations into the learning process, since the pathological changes in fundus images are usually subtle and scattered. Specifically, we propose a novel self-supervised cross-level consistency learning scheme (S2C2L), which leverages both view-level and region-level representations of a vision Transformer to improve the robustness of extracted self-supervised representation. A diagnosis perception module (DPM) is constructed to enhance the activation of local pathological regions from both region and view levels, and a cross-level consistency loss is dedicated to align the representations from both levels. Extensive experiments on iChallenge-AMD, LAG and APTOS2019 datasets validate the state-of-the-art performance of our method for three common eye diseases.
Qi Bi, Hao Zheng 0008, Xu Sun 0006, Jingjun Yi, Wentian Zhang, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001
ICASSP5
2024 A uniform representation model for OCT-based fingerprint presentation attack detection and reconstruction
Wentian Zhang, Feng Liu 0013, Ramachandra Raghavendra
Pattern Recognit.1
2024 Anomaly detection via gating highway connection for retinal fundus images
Wentian Zhang, Jinheng Xie, Yawen Huang, Yu Zhang 0185, Yuexiang Li, Ramachandra Raghavendra, Yefeng Zheng 0001
Pattern Recognit.1
2024 Dual Teacher Knowledge Distillation With Domain Alignment for Face Anti-Spoofing
abstract
Face recognition systems have raised concerns due to their vulnerability to different presentation attacks, and system security has become an increasingly critical concern. Although many face anti-spoofing (FAS) methods perform well in intra-dataset scenarios, their generalization remains a challenge. To address this issue, some methods adopt domain adversarial training (DAT) to extract domain-invariant features. Differently, in this paper, we propose a domain adversarial attack (DAA) method by adding perturbations to the input images, which makes them indistinguishable across domains and enables domain alignment. Moreover, since models trained on limited data and types of attacks cannot generalize well to unknown attacks, we propose a dual perceptual and generative knowledge distillation framework for face anti-spoofing that utilizes pre-trained face-related models containing rich face priors. Specifically, we adopt two different face-related models as teachers to transfer knowledge to the target student model. The pre-trained teacher models are not from the task of face anti-spoofing but from perceptual and generative tasks, respectively, which implicitly augment the data. By combining both DAA and dual-teacher knowledge distillation, we develop a dual teacher knowledge distillation with domain alignment framework (DTDA) for face anti-spoofing. The advantage of our proposed method has been verified through extensive ablation studies and comparison with state-of-the-art methods on public datasets across multiple protocols.
Zhe Kong, Wentian Zhang, Tao Wang 0052, Kaihao Zhang, Yuexiang Li, Xiaoying Tang 0001, Wenhan Luo
IEEE Trans. Circuits Syst. Video Technol.2
2024 Taming Self-Supervised Learning for Presentation Attack Detection: De-Folding and De-Mixing
abstract
Biometric systems are vulnerable to presentation attacks (PAs) performed using various PA instruments (PAIs). Even though there are numerous PA detection (PAD) techniques based on both deep learning and hand-crafted features, the generalization of PAD for unknown PAI is still a challenging problem. In this work, we empirically prove that the initialization of the PAD model is a crucial factor for generalization, which is rarely discussed in the community. Based on such observation, we proposed a self-supervised learning-based method, denoted as DF-DM. Specifically, DF-DM is based on a global-local view coupled with de-folding and de-mixing to derive the task-specific representation for PAD. During de-folding, the proposed technique will learn region-specific features to represent samples in a local pattern by explicitly minimizing the generative loss. While de-mixing drives detectors to obtain the instance-specific features with global information for more comprehensive representation by minimizing the interpolation-based consistency. Extensive experimental results show that the proposed method can achieve significant improvements in terms of both face and fingerprint PAD in more complicated and hybrid datasets when compared with the state-of-the-art methods. When training in CASIA-FASD and Idiap Replay-Attack, the proposed method can achieve an 18.60% equal error rate (EER) in OULU-NPU and MSU-MFSD, exceeding the baseline performance by 9.54%. The source code of the proposed technique is available at https://github.com/kongzhecn/dfdm.
Zhe Kong, Wentian Zhang, Feng Liu 0013, Wenhan Luo, LinLin Shen, Ramachandra Raghavendra
IEEE Trans. Neural Networks Learn. Syst.2
2024 Numerical Differentiation From Noisy Signals: A Kernel Regularization Method to Improve Transient-State Features for the Electronic Nose
abstract
As the simplest feature extraction, traditional hand-crafted transient-state features have been widely used in the area of electronic noses (e-noses). However, the influence of noise in the calculation of numerical differentiation leads to inaccuracy and instability in extracting these features. To tackle this issue, a novel numerical differentiation algorithm is proposed, which uses kernel-based regularization. The proposed method can provide accurate and stable transient-state features by directly estimating high-order derivatives from the noise-contaminated sensor’s reading. The feature representation is a prerequisite for the good performance of e-noses. Nevertheless, it should be noted that this performance in real applications can still be affected by other factors, such as sensor drift and the disturbance of nontarget odors. These issues can be addressed by applying a framework of domain adaptation and one-class classification. The proposed method and the adopted framework are verified in a field experiment, which identifies the odor of four targets and two disturbance whiskies measured by a self-designed e-nose system. The classification accuracy with traditional features is improved from$\mathbf{71.90\%}$to$\mathbf{86.36\%}$, showing the good potential of the proposed method for application in the area of e-noses.
Taoping Liu, Wentian Zhang, Li Wang 0148, Maiken Ueland, Shari L. Forbes, Wei Xing Zheng 0001, Steven W. Su
IEEE Trans. Syst. Man Cybern. Syst.2
2023 AdaptiveMix: Improving GAN Training via Feature Space Shrinkage
abstract
Due to the outstanding capability for data generation, Generative Adversarial Networks (GANs) have attracted considerable attention in unsupervised learning. However, training GANs is difficult, since the training distribution is dynamic for the discriminator, leading to unstable image representation. In this paper, we address the problem of training GANs from a novel perspective, i.e., robust image classification. Motivated by studies on robust image representation, we propose a simple yet effective module, namely AdaptiveMix, for GANs, which shrinks the regions of training data in the image representation space of the discriminator. Considering it is intractable to directly bound feature space, we propose to construct hard samples and narrow down the feature distance between hard and easy samples. The hard samples are constructed by mixing a pair of training images. We evaluate the effectiveness of our AdaptiveMix with widely-used and state-of-the-art GAN architectures. The evaluation results demonstrate that our AdaptiveMix can facilitate the training of GANs and effectively improve the image quality of generated samples. We also show that our AdaptiveMix can be further applied to image classification and Out-Of-Distribution (OOD) detection tasks, by equipping it with state-of-the-art methods. Extensive experiments on seven publicly available datasets show that our method effectively boosts the performance of baselines. The code is publicly available at https://github.com/WentianZhang-ML/AdaptiveMix.
Wentian Zhang, Bing Li 0024, Haoqian Wu, Nanjun He, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001
CVPR2
2023 NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation
abstract
Temporal video segmentation is the get-to- go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations can help compare the semantics in the corresponding scales, but lack a wider view of larger temporal spans, especially when the video is complex and structured. Therefore, we present two abstractive levels of temporal segmentations and study their hierarchy to the existing fine-grained levels. Accordingly, we collect NewsNet, the largest news video dataset consisting of 1,000 videos in over 900 hours, associated with several tasks for hierarchical temporal video segmentation. Each news video is a collection of stories on different topics, represented as aligned audio, visual, and textual data, along with extensive frame-wise annotations in four granularities. We assert that the study on NewsNet can advance the understanding of complex structured video and benefit more areas such as short-video creation, personalized advertisement, digital instruction, and education. Our dataset and code is publicly available at https://github.com/NewsNet-Benchmark/NewsNet.
Haoqian Wu, Mingchen Zhuge, Bing Li 0024, Ruizhi Qiao, Xiujun Shu, Bei Gan, Liangsheng Xu, Bo Ren 0002, Mengmeng Xu 0006, Wentian Zhang, Ramachandra Raghavendra, Chia-Wen Lin, Bernard Ghanem
CVPR12
2023 BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion
abstract
Recent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using other modalities as conditions, considerable paired data, e.g., box/mask-image pairs, and fine-tuning time are required for nurturing models. As such paired data is time-consuming and labor-intensive to acquire and restricted to a closed set, this potentially becomes the bottleneck for applications in an open world. This paper focuses on the simplest form of user-provided conditions, e.g., box or scribble. To mitigate the aforementioned problem, we propose a training-free method to control objects and contexts in the synthesized images adhering to the given spatial conditions. Specifically, three spatial constraints, i.e., Inner-Box, Outer-Box, and Corner Constraints, are designed and seamlessly integrated into the denoising step of diffusion models, requiring no additional training and massive annotated layout data. Extensive experimental results demonstrate that the proposed constraints can control what and where to present in the images while retaining the ability of Diffusion models to synthesize with high fidelity and diverse concept coverage.
Jinheng Xie, Yuexiang Li, Yawen Huang, Wentian Zhang, Yefeng Zheng 0001, Zheng Shou 0001
ICCV5
2023 Dynamically Masked Discriminator for GANs
abstract
Training Generative Adversarial Networks (GANs) remains a challenging problem. The discriminator trains the generator by learning the distribution of real/generated data. However, the distribution of generated data changes throughout the training process, which is difficult for the discriminator to learn. In this paper, we propose a novel method for GANs from the viewpoint of online continual learning. We observe that the discriminator model, trained on historically generated data, often slows down its adaptation to the changes in the new arrival generated data, which accordingly decreases the quality of generated results. By treating the generated data in training as a stream, we propose to detect whether the discriminator slows down the learning of new knowledge in generated data. Therefore, we can explicitly enforce the discriminator to learn new knowledge fast. Particularly, we propose a new discriminator, which automatically detects its retardation and then dynamically masks its features, such that the discriminator can adaptively learn the temporally-vary distribution of generated data. Experimental results show our method outperforms the state-of-the-art approaches.
Wentian Zhang, Bing Li 0024, Jinheng Xie, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Bernard Ghanem
NeurIPS1
2022 Effective Presentation Attack Detection Driven by Face Related Task
Wentian Zhang, Feng Liu 0013, Ramachandra Raghavendra, Christoph Busch 0001
ECCV (5)1
2022 A Multi-task Network with Weight Decay Skip Connection Training for Anomaly Detection in Retinal Fundus Images
Wentian Zhang, Xu Sun 0006, Yuexiang Li, Nanjun He, Feng Liu 0013, Yefeng Zheng 0001
MICCAI (2)1
2022 Blockchain-Enabled Fish Provenance and Quality Tracking System
abstract
Accurate assessment of fish quality is difficult in practice due to the lack of trusted fish provenance and quality tracking information. Working with Sydney Fish Market (SFM), we develop a Blockchain-enabled fish provenance and quality tracking (BeFAQT) system. A multilayer Blockchain architecture based on attribute-based encryption (ABE) is proposed to tackle the privacy issue caused by applying Blockchain to secure supply chain data and achieve trusted and confidential data sharing among parties in fish supply chains. An Internet-of-Things (IoT) chain saves encrypted fish provenance and quality tracking data, and an ABE chain is specifically designed for the access control to the data in the IoT chain. Latest IoT and artificial intelligence (AI) technologies, including NarrowBand-IoT, image processing, and biosensing, are developed for fish origin proof, supply chain tracking, and objective fish quality assessment. As proven by field trials with SFM and a local fish supply chain, the BeFAQT is able to provide trusted and comprehensive fish provenance and quality tracking information in real time.
Xu Wang 0004, Guangsheng Yu, Ren Ping Liu 0001, Jian Zhang 0002, Qiang Wu 0001, Steven W. Su, Ying He 0011, Zongjian Zhang, Litao Yu, Taoping Liu, Wentian Zhang, Peter Loneragan, Eryk Dutkiewicz, Erik Poole, Nick Paton
IEEE Internet Things J.11
2022 Fingerprint Presentation Attack Detector Using Global-Local Model
abstract
The vulnerability of automated fingerprint recognition systems (AFRSs) to presentation attacks (PAs) promotes the vigorous development of PA detection (PAD) technology. However, PAD methods have been limited by information loss and poor generalization ability, resulting in new PA materials and fingerprint sensors. This article thus proposes a global-local model-based PAD (RTK-PAD) method to overcome those limitations to some extent. The proposed method consists of three modules, called: 1) the global module; 2) the local module; and 3) the rethinking module. By adopting the cut-out-based global module, a global spoofness score predicted from nonlocal features of the entire fingerprint images can be achieved. While by using the texture in-painting-based local module, a local spoofness score predicted from fingerprint patches is obtained. The two modules are not independent but connected through our proposed rethinking module by localizing two discriminative patches for the local module based on the global spoofness score. Finally, the fusion spoofness score by averaging the global and local spoofness scores is used for PAD. Our experimental results evaluated on LivDet 2017 show that the proposed RTK-PAD can achieve an average classification error (ACE) of 2.28% and a true detection rate (TDR) of 91.19% when the false detection rate (FDR) equals 1.0%, which significantly outperformed the state-of-the-art methods by ~10% in terms of TDR (91.19% versus 80.74%).
Wentian Zhang, Feng Liu 0013, Haoqian Wu, LinLin Shen
IEEE Trans. Cybern.2
2022 Fingerprint Presentation Attack Detection by Channel-Wise Feature Denoising
abstract
Due to the diversity of attack materials, fingerprint recognition systems (AFRSs) are vulnerable to malicious attacks. It is thus important to propose effective fingerprint presentation attack detection (PAD) methods for the safety and reliability of AFRSs. However, current PAD methods often exhibit poor robustness under new attack types settings. This paper thus proposes a novel channel-wise feature denoising fingerprint PAD (CFD-PAD) method by handling the redundant noise information ignored in previous studies. The proposed method learns important features of fingerprint images by weighing the importance of each channel and identifying discriminative channels and “noise” channels. Then, the propagation of “noise” channels is suppressed in the feature map to reduce interference. Specifically, a PA-Adaptation loss is designed to constrain the feature distribution to make the feature distribution of live fingerprints more aggregate and that of spoof fingerprints more disperse. Experimental results evaluated on the LivDet 2017 dataset showed that the proposed CFD-PAD can achieve 2.53% average classification error (ACE) and a 93.83% true detection rate when the false detection rate equals 1.0% (TDR@FDR=1%). Also, the proposed method markedly outperforms the best single-model-based methods in terms of ACE (2.53% vs. 4.56%) and TDR@FDR=1%(93.83% vs. 73.32%), which demonstrates its effectiveness. Although we have achieved a comparable result with the state-of-the-art multiple-model-based methods, there still is an increase in TDR@FDR=1% from 91.19% to 93.83%. In addition, the proposed model is simpler, lighter and more efficient and has achieved a 74.76% reduction in computation time compared with the state-of-the-art multiple-model-based method.The source code is available athttps://github.com/kongzhecn/cfd-pad.
Feng Liu 0013, Zhe Kong, Wentian Zhang, LinLin Shen
IEEE Trans. Inf. Forensics Secur.4
2022 A Multiscale Wavelet Kernel Regularization-Based Feature Extraction Method for Electronic Nose
abstract
In the electronic nose (e-nose), a stable feature representation of the gas sensor’s response is a key step to realize subsequent odor identification algorithms. However, the noises in gas sensors hinder the acquisition of such features. In order to solve this problem, this article proposes a stable feature extraction algorithm which takes the impulse response of the e-nose system as the feature. The impulse response is estimated from a nonparametric model constrained by a multiscale wavelet kernel regularization matrix. The kernel regularization matrix equips the proposed feature extraction method with an ability in resistance to random noise. A numerical experiment proves that compared with single-scale kernel regularization, the use of multiscale wavelet kernel helps to achieve more stable and accurate impulse response estimation. Then, a field experiment is conducted to demonstrate the performance of the proposed features. This experiment aims to identify four different whiskies measured by a self-designed e-nose with four commercial gas sensors. Under the framework of transfer learning, the classification result based on the proposed features outperforms those using other considered features. The accuracy of whisky identification reaches 92.00%, showing a good potential of applying the proposed feature representations in the area of e-noses.
Taoping Liu, Wentian Zhang, Jun Li 0010, Maiken Ueland, Shari L. Forbes, Wei Xing Zheng 0001, Steven W. Su
IEEE Trans. Syst. Man Cybern. Syst.2
2021 Finger Vein Verification using Intrinsic and Extrinsic Features
abstract
Finger vein has attracted substantial attention due to its good security. However, the variability of the finger vein data will be caused by the illumination, environment temperature, acquisition equipment, and so on, which is a great challenge for finger vein recognition. To address this problem, we propose a novel method to design an endto-end deep Convolutional Neural Network (CNN) for robust finger vein recognition. The approach mainly includes an Intrinsic Feature Learning (IFL) module using an auto-encoder network and an Extrinsic Feature Learning (EFL) module based on a Siamese network. The IFL module is designed to estimate the expectation of intra-class finger vein images with various offsets and rotation, while the EFL module is constructed to learn the inter-class feature representation. Then, robust verification is finally achieved by considering the distances of both intrinsic and extrinsic features. We conduct experiments on two public datasets (i.e. SDUMLA-HMT and MMCBNU_6000) and an in-house dataset (MultiView-FV) with more deformation finger vein images, and the equal error rate (EER) is 0.47%, 0.1%, and 1.69% respectively. The comparison against baseline and existing algorithms shows the effectiveness of our proposed method.
Liying Lin, Wentian Zhang, Feng Liu 0013, Zhihui Lai 0001
IJCB3
2021 One-Class Fingerprint Presentation Attack Detection Using Auto-Encoder Network
abstract
Automated Fingerprint Recognition Systems (AFRSs) have been threatened by Presentation Attack (PA) since its existence. It is thus desirable to develop effective presentation attack detection (PAD) methods. However, the unpredictable PAs make PAD be a challenging problem. This paper proposes a novel One-Class PAD (OCPAD) method for Optical Coherence Technology (OCT) images based fingerprint PA detection. The proposed OCPAD model is learned from a training set only consists of Bonafides (i.e. real fingerprints). The reconstruction error and latent code obtained from the trained auto-encoder network in the proposed model is taken as the basis for the following spoofness score calculation. To get more accurate reconstruction error, we propose an activation map based weighting model to further refine the accuracy of reconstruction error. We test different statistics and distance measures and finally use a decision level fusion to make the final prediction. Our experiments are performed using a dataset with 93200 bonafide scans and 48400 PA scans. The results show that the proposed OCPAD can achieve a True Positive Rate (TPR) of 99.43% when the False Positive Rate (FPR) equals to 10% and a TPR of 96.59% when FPR=5%, which significantly outperformed a feature based approach and a supervised learning based model requiring PAs for training.
Feng Liu 0013, Wentian Zhang, Guojie Liu, LinLin Shen
IEEE Trans. Image Process.3