Zeyu Lei

dblp:256/4676 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0001-4963-6961ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A First Look at the Mobile Driving License (mDL) Standard and its Real-world Usage
Zeyu Lei, Güliz Seray Tuncay, Abdullah Imran, Z. Berkay Celik, Antonio Bianchi
AsiaCCS1
2025 Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding Transformer
abstract
Non-semantic features or semantic-agnostic features, which are irrelevant to image context but sensitive to image manipulations, are recognized as evidential to Image Manipulation Localization (IML). Since manual labels are impossible, existing works rely on handcrafted methods to extract non-semantic features. Handcrafted non-semantic features jeopardize IML model's generalization ability in unseen or complex scenarios. Therefore, for IML, the elephant in the room is: How to adaptively extract non-semantic features? Non-semantic features are context-irrelevant and manipulation-sensitive. That is, within an image, they are consistent across patches unless manipulation occurs. Then, spare and discrete interactions among image patches are sufficient for extracting non-semantic features. However, image semantics vary drastically on different patches, requiring dense and continuous interactions among image patches for learning semantic representations. Hence, in this paper, we propose a Sparse Vision Transformer (SparseViT), which reformulates the dense, global self-attention in ViT into a sparse, discrete manner. Such sparse self-attention breaks image semantics and forces SparseViT to adaptively extract non-semantic features for images. Besides, compared with existing IML models, the sparse self-attention mechanism largely reduced the model size (max 80% in FLOPs), achieving stunning parameter efficiency and computation reduction. Extensive experiments demonstrate that, without any handcrafted feature extractors, SparseViT is superior in both generalization and efficiency across benchmark datasets.
Xiaochen Ma 0001, Xuekang Zhu, Chaoqun Niu, Zeyu Lei, Jizhe Zhou 0001
AAAI5
2025 Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation Localization
abstract
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
Xuekang Zhu, Xiaochen Ma 0001, Zhuohang Jiang, Xiwen Wang 0002, Zeyu Lei, Wentao Feng, Chi-Man Pun, Jizhe Zhou 0001
AAAI7
2025 M3: Manipulation Mask Manufacturer for Arbitrary-Scale Super-Resolution Mask
Xiaochen Ma 0001, Xuekang Zhu, Bingkui Tong, Zeyu Lei, Jizhe Zhou 0001
CVM (1)7
2025 WDiff: Wavelet-based Diffusion Models for Surgical Endoscopic Image Low-Light Enhancement
abstract
Endoscopes, both white light and fluorescence, face challenges like insufficient illumination during surgical procedures. Deep learning, especially diffusion models, demonstrates considerable potential for low-light enhancement in the medical field. However, challenges in enhancing endoscopic images persist, including high computational resource consumption, lengthy processing times, and potential result distortion. To address these issues, we propose a wavelet-based diffusion method for fast and efficient low-light surgical endoscopic image enhancement, dubbed WDiff. It utilizes wavelet transform to reduce computational resource and enhance inference speed while preserving key features. To avoid degradation during reverse process, we introduce a Degradation Corrector Unit (DCU) to ensure accurate sampling results. Moreover, we design a Detail Coefficients Restoration Block (DCRB) to reconstruct the local sparse information in detail coefficients across horizontal, vertical, and diagonal orientations. Extensive experiments on benchmark datasets demonstrate that our method outperforms other SOTA methods both visually and quantitatively, achieving an optimal balance between complexity and efficiency.
Zeyu Lei, Lidan Fu, Anqi Xiao, Jie Tian 0001, Zhenhua Hu
ICME1
2025 ScopeVerif: Analyzing the Security of Android's Scoped Storage via Differential Analysis
Zeyu Lei, Güliz Seray Tuncay, Beatrice Carissa Williem, Z. Berkay Celik, Antonio Bianchi
NDSS1
2024 IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization
abstract
A comprehensive benchmark is yet to be established in the Image Manipulation Detection & Localization (IMDL) field. The absence of such a benchmark leads to insufficient and misleading model evaluations, severely undermining the development of this field. However, the scarcity of open-sourced baseline models and inconsistent training and evaluation protocols make conducting rigorous experiments and faithful comparisons among IMDL models challenging. To address these challenges, we introduce IMDL-BenCo, the first comprehensive IMDL benchmark and modular codebase. IMDL-BenCo: i) decomposes the IMDL framework into standardized, reusable components and revises the model construction pipeline, improving coding efficiency and customization flexibility; ii) fully implements or incorporates training code for state-of-the-art models to establish a comprehensive IMDL benchmark; and iii) conducts deep analysis based on the established benchmark and codebase, offering new insights into IMDL model architecture, dataset characteristics, and evaluation standards.Specifically, IMDL-BenCo includes common processing algorithms, 8 state-of-the-art IMDL models (1 of which are reproduced from scratch), 2 sets of standard training and evaluation protocols, 15 GPU-accelerated evaluation metrics, and 3 kinds of robustness evaluation. This benchmark and codebase represent a significant leap forward in calibrating the current progress in the IMDL field and inspiring future breakthroughs.Code is available at: https://github.com/scu-zjz/IMDLBenCo
Xiaochen Ma 0001, Xuekang Zhu, Zhuohang Jiang, Bingkui Tong, Zeyu Lei, Chi-Man Pun, Jiancheng Lv 0001, Jizhe Zhou 0001
NeurIPS7
2022 RF-DCM: Multi-Granularity Deep Convolutional Model Based on Feature Recalibration and Fusion for Driver Fatigue Detection
abstract
Fatigue driving is one of the main causes of traffic accidents. For real-world driver fatigue detection, the large pose deformations exhibited by the captured global face significantly increase the difficulty of extracting effective features. Furthermore, previous fatigue detection methods have not achieved desired results in distinguishing actions with similar appearance, such as yawning and speaking. In this article, we propose a multi-granularity Deep Convolutional Model based on feature Recalibration and Fusion for driver fatigue detection (RF-DCM). Our deep model leverages cues from partial faces to alleviate the pose variations and obtains robust feature representations from both the global face and different local parts. The core innovative techniques are as follows: A multi-granularity extraction sub-network extracts more efficient multi-granularity features while compressing the parameters of the network. In order to match multi-granularity features, a feature rectification sub-network and a feature fusion sub-network are designed to adaptively recalibrate and fuse the multi-granularity features. A long short term memory network is used to explore the relationship among sequence frames to distinguish actions with similar appearances. Extensive experimental results on the public drowsy driver dataset from NTHU Driver Drowsy competition demonstrate significant performance improvements of our model over all published state-of-the-art methods.
Rui Huang 0013, Yan Wang 0042, Zijian Li 0005, Zeyu Lei, Yu-Fan Xu
IEEE Trans. Intell. Transp. Syst.4
2022 Unsupervised Learning of Depth Estimation and Camera Pose With Multi-Scale GANs
abstract
Unsupervised learning methods have achieved remarkable performance in monocular depth estimation and camera pose, which mostly solve the multi-task learning problem by using their inner geometry consistency as the self-supervision signal. While most existing approaches mostly adopt the generative model to obtain the depth map prediction, so in the resolution of depth map there is room for improvement. To this end, we present our unsupervised learning architecture based on adversarial learning model, which is used for unsupervised learning of high-resolution single view depth and camera pose. Specifically, we present a multi-scale deep convolutional Generative Adversarial Network (GAN) based learning system, which consists of three networks (pose estimation network PCNN, Generator-D and Discriminator-D for depth map prediction). Furthermore, in order to generate high-resolution depth map, we propose a multi-scale GAN model (MSGAN) to decompose the hard high-quality image generation problem into more manageable sub-problems through a coarse-to-fine process. Then, we modify the overall generation architecture of GAN model by changing the down-sampling and up-sampling components to improve the quality and accuracy of the depth map prediction. Finally, in order to improve the rate of convergence, we use the Least Square Error to increase the penalty for outliers. Detailed quantitative and qualitative evaluations of the proposed framework on the KITTI dataset show that the proposed method provides better results for both pose estimation and depth recovery.
Yu-Fan Xu, Yan Wang 0042, Rui Huang 0013, Zeyu Lei, Junyao Yang, Zijian Li 0005
IEEE Trans. Intell. Transp. Syst.4
2021 On the Insecurity of SMS One-Time Password Messages against Local Attackers in Modern Mobile Devices
Zeyu Lei, Yuhong Nan, Yanick Fratantonio, Antonio Bianchi
NDSS1
2021 Attention based multilayer feature fusion convolutional neural network for unsupervised monocular depth estimation
Zeyu Lei, Yan Wang 0042, Zijian Li 0005, Junyao Yang
Neurocomputing1