EDBT 2026 Demo / reviewers in the wild / expert
Yuan Zhang 0023
dblp:48/2168-23
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-0032-8631ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flexible Modal Mixture-of-Experts With Inter-Modal Knowledge Distillation for Face Anti-Spoofing
Hui Ma 0018, Ajian Liu 0001, Ning Li 0035, Boyun Wang, Hang Zou 0002, Yuan Zhang 0023, Jing Huang 0017, Zhiqiang Pu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Toward Generalized Iris Presentation Attack Detection: A Mask-and-Distill Mixture of Experts ApproachabstractIris Presentation Attack Detection (PAD) is critical for securing recognition systems, yet its practical deployment is severely hindered by the poor generalization of models across different acquisition devices and diverse datasets. To address this persistent cross-domain challenge, we first introduce a comprehensive evaluation framework, the Iris Presentation Attack Detection Cross-Domain-Testing (IPAD-CDT) Protocol, designed to evaluate the model robustness in these scenarios. Our core contribution is a novel Masked Mixture-of-Experts (MMoE) method, which enhances the generalization of Transformer-based architectures. MMoE introduces a structured information asymmetry, where "student" Experts learn robust features from masked inputs by distilling knowledge from an unmasked "teacher" Expert via a cosine distance loss. This mask-and-distill mechanism effectively mitigates overfitting and guides the model to learn domain-invariant cues. By integrating MMoE into a CLIP-based model, we conduct extensive experiments on our IPAD-CDT protocol. The results demonstrate that our method sets a new state-of-the-art, significantly outperforming existing models, especially in the challenging cross-dataset and cross-device settings. Hang Zou 0002, Chenxi Du, Ajian Liu 0001, Yuan Zhang 0023, Jing Liu 0062, Jun Wan 0001, Hui Zhang 0061, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Multi-Modal BEV Enhancement Fusion for 3D Object Detection in Autonomous DrivingabstractRecent success in 3D object detection have underscored its importance in autonomous driving, particularly through the integration of diverse sensor modalities like RGB images and LiDAR point clouds. With the benefits of Lift-Splat Shift (LSS) paradigm, different data modalities can be effectively fused in Bird’s-Eye-View (BEV), significantly improving the detection performance. Although BEV-based fusion has significantly advanced 3D detection technology, the limited enhancement of image features and the inconsistency between different modalities still hinder the overall performance of the detector. In this work, we focus on effective enhancement strategies and design 3D object detection pipeline named ECL3D to further push the detection performance boundary. The first strategy, Depth-Semantic Feature Enhancement (DSE), aims to improve input features for the view transformer during training without adding computational burden during inference. This approach leverages low-resolution depth distribution supervision to maintain the accuracy of frustum generation and high-resolution depth supervision to provide richer clues for front-end features. The second strategy, Instance BEV Feature Enhancement (IBFE), introduces a mechanism to enhance instance-relevant features for multi-modal fusion. This strategy effectively suppresses background noise and enhances object-region features. Comprehensive experiments on nuScenes dataset demonstrate the effectiveness of our approach. Without any test-time-augmentation strategy, our detector achieves state-of-the-art performance in camera-LiDAR fusion 3D object detection task with mAP and NDS of 72.8% and 75.2%, respectively. The code is coming soon at: https://github.com/muchen2019/ECL3D Yuan Zhang 0023, Xinchi Li, Huaici Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | La-SoftMoE CLIP for Unified Physical-Digital Face Attack DetectionabstractFacial recognition systems are susceptible to both physical and digital attacks, posing significant security risks. Traditional approaches often treat these two attack types separately due to their distinct characteristics. Thus, when being combined attacked, almost all methods could not deal. Some studies attempt to combine the sparse data from both types of attacks into a single dataset and try to find a common feature space, which is often impractical due to the space is difficult to be found or even non-existent. To overcome these challenges, we propose a novel approach that uses the sparse model to handle sparse data, utilizing different parameter groups to process distinct regions of the sparse feature space. Specifically, we employ the Mixture of Experts (MoE) framework in our model, expert parameters are matched to tokens with varying weights during training and adaptively activated during testing. However, the traditional MoE struggles with the complex and irregular classification boundaries of this problem. Thus, we introduce a flexible self-adapting weighting mechanism, enabling the model to better fit and adapt. In this paper, we proposed La-SoftMoE CLIP, which allows for more flexible adaptation to the Unified Attack Detection (UAD) task, significantly enhancing the model’s capability to handle diversity attacks. Experiment results demonstrate that our proposed method has SOTA performance. Hang Zou 0002, Chenxi Du, Hui Zhang 0061, Yuan Zhang 0023, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 4 |
| 2024 | Style-conditional Prompt Token Learning for Generalizable Face Anti-spoofingabstractFace anti-spoofing (FAS) based on domain generalization (DG) has attracted increasing attention from researchers.The reason for the poor generalization is that the model is overfitted to salient liveness-irrelevant signals.However, the previous methods alleviate the overfitting by mapping the images from multiple domains into a common feature space or promoting the separation of image features from domain-specific features and task-related features.If the text features of vision-language pre-trained (VLP) models (e.g., CLIP) are used to dynamically adjust the image features to gain a better generalization, we can not only explore a wider feature space but also avoid the potential degradation of semantic information.Specifically, we propose a FAS method of Style-Conditional Prompt Token Learning (S-CPTL), which aims to generate generalized text features by training the introduced prompt tokens to carry visual styles and use them as weights for classifiers to improve the model's generalization.Compared to the inherently static prompt token, we propose the dynamic prompt token, which can adaptively capture live-irrelevant signals from the instance-specific styles and increase their diversity through mixed feature statistics to further reduce the overfitting of the model.Thorough experimental analysis demonstrates that S-CPTL exceeds current top-performing methods in four distinct cross-dataset benchmarks. Jiabao Guo, Huan Liu 0030, Yizhi Luo, Xueli Hu, Hang Zou 0002, Yuan Zhang 0023, Hui Liu 0018, Bo Zhao 0023 |
ACM Multimedia | 6 |
| 2024 | Fine-Grained Prompt Learning for Face Anti-SpoofingabstractThere has been an increasing focus on domain-generalized (DG) face anti-spoofing (FAS). However, existing methods aim to project a shared visual space through adversarial training, making exploring the space without losing semantic information challenging. We investigate the DG inadequacies resulting from classifier overfitting to a significantly different domain distribution. To address this issue, we propose a novel Fine-Grained Prompt Learning (FGPL) based on Vision-Language Models (VLMs), such as CLIP, which can adaptively adjust weights for classifiers with text features to mitigate overfitting. Specifically, FGPL first motivates the prompts to learn content and domain semantic information by capturing Domain-Agnostic and Domain-Specific features. Furthermore, our prompts are designed to be category-generalized by diversifying the Domain-Specific prompts. Additionally, we design an Adaptive Convolutional Adapter (AC-adapter), which is implemented through an adaptive combination of Vanilla Convolution and Central Difference Convolution, to be inserted into the image encoder for quickly bridging the gap between general image recognition and FAS task. Extensive experiments demonstrate that the proposed FGPL is effective and outperforms state-of-the-art methods on several cross-domain datasets. Xueli Hu, Huan Liu 0030, Haocheng Yuan, Zhiyang Fu, Yizhi Luo, Ning Zhang 0033, Hang Zou 0002, Jianwen Gan, Yuan Zhang 0023 |
ACM Multimedia | 9 |
| 2024 | Distance-based feature repack algorithm for video coding for machines
Yuan Zhang 0023, Xiaoli Gong, Hualong Yu, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Low-Rate Feature Compression for Collaborative Intelligence: Reducing Redundancy in Spatial and Statistical LevelsabstractTo distribute the storage and computation load caused by growing capacity of deep neural network (DNN), collaborative intelligence (CI) framework has been proposed, where a deep model is split and executed in two distributed devices respectively. Intermediate feature must be transferred from the front end to the back in order to perform distributed inference, thus transmission process is the bottleneck that influences the inference efficiency in terms of accuracy and delay. Specifically for a bandwidth-limited human-in-loop visual analysis task, feature compression approach needs exploration to reduce the data volume to be transmitted, in order to achieve low transmission delay as well as maintain analysis performance and human perception ability. In this article, the redundancy of intermediate feature both in spatial and statistical levels are firstly analyzed. A mathematical expression for the goal of feature compression is formulated, based on which a two-level redundancy removal based low-rate feature compression approach is proposed. For the front-end device, an information squeezing (IS) module is developed to squeeze the key information of input image and inject them into a low-resolution image. Then a backbone network is split into two parts with respects to the application demands of CI, and can be deployed at the front and back ends correspondingly. With a specifically designed objective function, IS module and the partitioned backbone network are optimized collaboratively to reduce the two-level redundancy, thus compressing the intermediate feature. A generative adversarial network (GAN)-based restoration module is proposed to recover an image with original resolution from the compressed feature, for satisfying human perception. Comprehensive experiments are conduct to validate the efficiency of the proposed method. Fan Li 0003, Yuan Zhang 0023 |
IEEE Trans. Multim. | 4 |
| 2023 | Frame importance and temporal memory effect-based fast video quality assessment for user-generated content
Yuan Zhang 0023, Lijun He 0001 |
Appl. Intell. | 1 |
| 2023 | HPV-RCNN: Hybrid Point-Voxel Two-Stage Network for LiDAR-Based 3-D Object DetectionabstractThe current two-stage detectors remarkably benefit from hybrid representation of points and 3-D voxels, but they have high time cost and leave room for improving the accuracy of small objects. On the contrary, 2-D voxel-based methods tend to have good efficiency and better performance for small objects. An intuitive idea of optimizing a two-stage algorithm is to use a 2-D voxel-based backbone. However, naive representation substitution cannot achieve optimal joint learning of each representation and may cause a decrease in accuracy. In this article, we propose hybrid point–voxel RCNN (HPV-RCNN), a novel point cloud detection network which combines the merits of points and 2-D voxels. First, we propose a multiattentive voxel feature encoding module (MAVFE) to exploit multilevel attention of multiscale voxels. We also present a partial fusion pyramid network (PFPN) to effectively integrate multiresolution features and generate high-quality proposals. Then, a multiscale region of interest (RoI)-grid pooling (MSRGP) module is proposed to adaptively abstract proposal-specific features from sampled keypoints in multiple receptive fields. In addition, a cascade attentive module (CAM) is adopted to achieve incrementally proposal refinement by subsequent multiple subnetworks. Our method reaches top performance among two-stage methods in Cyclist and Pedestrian categories on the KITTI dataset while achieving real-time inference speed. Extensive experiments on challenging roadside DAIR-V2X-I dataset also demonstrate that our method achieves superior detection performance. Chen Feng 0029, Chao Xiang, Xiaopo Xie, Yuan Zhang 0023, Xuesong Li 0003 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Mutual Learning Inspired Prediction Network for Video Anomaly Detection
Yuan Zhang 0023, Fan Li 0003, Lu Yu 0003 |
PRCV (3) | 1 |
| 2022 | Visual Analysis Motivated Super-Resolution Model for Image ReconstructionabstractThis paper presents a concise end-to-end visual analysis motivated super-resolution model VASR for image reconstruction. Compatible with the existing machine vision feature coding framework, the features extracted from the machine vision task model are super-resolution amplified to reconstruct the original image for human vision. The experimental results show that without additional bit-streams, VASR can well complete the task of image reconstruction based on the extracted machine features, and has achieved good results on COCO, OpenImages, TVD, and DIV2K datasets. Huifen Wang, Junda Xue, Yuan Zhang 0023 |
VCIP | 4 |