VLDB 2026 Research / reviewers in the wild / expert
Zhi Li 0012
dblp:43/3166-12
· DBLP profile ↗
25ranked-venue papers
2as first author
22since 2021 · last 2026
0000-0001-9813-4979ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Security and privacy · 4 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E³SAM2: Entropy-Aware and Edge-Guided Adaptation of SAM2 for Echocardiography Video SegmentationabstractFoundation segmentation models, such as SAM and its video-oriented variant SAM2, have achieved remarkable success in natural image and video segmentation. However, their direct application to echocardiography video is challenged by structural uncertainty arising from severe speckle noise and blurry anatomical boundaries. To address this, we propose E³SAM2, a lightweight adaptation framework that introduces a novel entropy-based methodology to explicitly model and mitigate such uncertainty. Specifically, an entropy-guided attention mechanism is introduced to steer the model’s focus toward structurally reliable features, particularly in speckle-dominated regions. Additionally, an entropy regularization loss is introduced to further enhance target-background discrimination. To better resolve indistinct anatomical contours, an edge-aware supervision module is incorporated to inject explicit boundary priors for sharper delineation. These components are efficiently integrated through a global-local feature adapter. Experiments on CAMUS and EchoNet-Dynamic datasets demonstrate that E³SAM2 achieves state-of-the-art segmentation and clinical estimation performance, while maintaining high computational efficiency. Zhi Li 0012, Zhenyu Dai, Shuyun Li |
AAAI | 2 |
| 2026 | DNPR: Zero-shot industrial anomaly detection via dynamic normal prototype refinement
Shuyun Li, Zhi Li 0012 |
Expert Syst. Appl. | 2 |
| 2026 | Improving generalization of universal adversarial perturbation via integrating hierarchical features and spatial transformations
Zhi Li 0012 |
Expert Syst. Appl. | 2 |
| 2025 | Generate universal adversarial perturbations by shortest-distance soft maximum direction attack
Dengbo Liu, Zhi Li 0012, Daoyun Xu |
Comput. Secur. | 2 |
| 2025 | Adaptive adversarial pattern contrast algorithm for black-box model and domain attack
Yi Wang 0146, Zhi Li 0012 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | SRAD: A spatially-aware reconstruction network with anomaly suppression for multi-class anomaly detection
Shuyun Li, Zhi Li 0012, Rongxiang Wang |
Neurocomputing | 2 |
| 2025 | SRIS-Net: a robust image steganography algorithm based on feature score mapsabstractImage steganography algorithms based on deep learning are often trained using either spatial- or frequency-domain features. It is difficult for features from a single domain to comprehensively express the content of an entire image, which usually leads to poor performance because steganography is commonly multi-task. To solve this problem, this paper proposes a robust image steganography algorithm based on feature score maps, called the secure and robust image steganography network (SRIS-Net). First, instead of spatial-domain steganography, our proposed algorithm utilizes a convolutional neural network to obtain shallow spatial-domain features. These features are decomposed by Laplacian pyramid frequency-domain decomposition (LPFDD) to hide secret information in the different frequency sub-bands with a progressive assisted hiding strategy that significantly reduces the influence of the secret information on the cover image, achieving significant invisibility and robust performance. In addition, we propose a global–local embedding module (GLEM) to achieve embedding by considering the overall structure of the image and the local details, and a dual multi-scale aggregation sub-network (DMSubNet) to perform multi-scale reconstruction to improve the quality of the carrier image. For security, we propose a dual-task discriminator structure, while giving a real/fake judgment of the image, which can generate a feature score map of the cover image’s region of interest (ROI) to guide the embedding module to generate a carrier image with higher imperceptibility and undetectability. Experimental results on BOSSBase show that our SRIS-Net outperforms mainstream methods in terms of undetectability and robustness, with more than 9.2 and 3.4 dB improvement in visual quality, respectively, and the capacity can be increased up to approximately 72–96 bits per pixel. Ai Xiao, Zhi Li 0012, Guomei Wang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | Q-space-coordinate-guided neural networks for high-fidelity diffusion tensor estimation from minimal diffusion-weighted imagesabstractDiffusion tensor imaging (DTI) is a widely used imaging technique for mapping living human brain tissue’s microstructure and structural connectivity. Recently, deep learning methods have been proposed to rapidly estimate diffusion tensors (DTDs) using only a small quantity of diffusion-weighted (DW) images. However, these methods typically use the DW images obtained with fixed q-space sampling schemes as the training data, limiting the application scenarios of such methods. To address this issue, we develop a new deep neural network called q-space-coordinate-guided diffusion tensor imaging (QCG-DTI), which can efficiently and correctly estimate DTs under flexible q-space sampling schemes. First, we propose a q-space-coordinate-embedded feature consistency strategy to ensure the correspondence between q-space-coordinates and their respective DW images. Second, a q-space-coordinate fusion (QCF) module is introduced which efficiently embeds q-space-coordinates into multiscale features of the corresponding DW images by linearly adjusting the feature maps along the channel dimension, thus eliminating the dependence on fixed diffusion sampling schemes. Finally, a multiscale feature residual dense (MRD) module is proposed which enhances the network’s feature extraction and image reconstruction capabilities by using dual-branch convolutions with different kernel sizes to extract features at different scales. Compared to state-of-the-art methods that rely on a fixed sampling scheme, the proposed network can obtain high-quality diffusion tensors and derived parameters even using DW images acquired with flexible q-space sampling schemes. Compared to state-of-the-art deep learning methods, QCG-DTI reduces the mean absolute error by approximately 15% on fractional anisotropy and around 25% on mean diffusivity. Maokun Zheng, Zhi Li 0012, Guomei Wang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2025 | Deep feature clustering for multi-class industrial image anomaly detection
Rongxiang Wang, Zhi Li 0012, Shuyun Li |
Knowl. Based Syst. | 2 |
| 2025 | Leveraging Learnable Defense Prompts and One-Step Diffusion Model for Adversarial Purification in Remote SensingabstractDeep neural networks (DNNs) have demonstrated outstanding performance in remote sensing imagery (RSI) analysis. However, their sensitivity to adversarial perturbations remains a critical bottleneck for their deployment in security-sensitive scenarios. Currently, most mainstream adversarial purification methods focus on pixel space example reconstruction, neglecting the effective recovery of deep semantic features and the classification decision space, thereby failing to mitigate the adversarial attacks adequately. To address this challenge, we propose a novel multimodal adversarial purification framework, Leveraging Learnable Defense Prompts and a one-step diffusion model (LDP-Diff), which aims to achieve robust adversarial defense with semantic consistency. Firstly, we introduce a learnable defense prompt (LDP) mechanism in LDP-Diff. Through the learnable optimization of input prompts, LDP-Diff can provide precise semantic guidance for the purification process. Next, we introduce a One-Step Diffusion Model combined with the low-rank adaptation (LoRA) for lightweight fine-tuning, significantly reducing the inference complexity while effectively removing adversarial perturbations and achieving high-quality structural and semantic reconstruction. Moreover, to further enhance the stability of purified examples in the classification space, we propose a entropy-guided logit fusion (EGLF) loss, which quantifies the confidence entropy and adaptively fuses the logit distributions between examples to alleviate the class bias remaining in the classification space after reconstruction. Comprehensive experimental results show that LDP-Diff exhibits excellent defense capabilities in multiple RSI datasets. LDP-Diff achieves an accuracy improvement of 2.40% on the UCM dataset and realizes a significant performance improvement of 10.84% on the AID dataset compared to the existing state-of-the-art (SOTA) methods. Notably, LDP-Diff exhibits excellent robustness, generalization, computational efficiency, and significant advantages in cross-domain defense. Zhi Li 0012, Zhenyu Dai, Chuhua Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | TSAD: Two-Stage Separable Adversarial Distortion-Based Robust Watermarking Framework for Diffusion Tensor ImagingabstractRecent deep learning-based watermarking methods have achieved impressive results. However, they struggle with unknown distortions and often suffer from poor generalization, slow convergence, unstable training, and degraded visual quality in watermarked images. To address the above problems, this paper proposes a two-stage separable adversarial distortion (TSAD)-based robust watermarking algorithm for diffusion tensor imaging (DTI). The algorithm uses a noise-free end-to-end network in the first stage for learning and training DTI images. In the second stage, it fixes the watermark embedding network trained in the first stage, interacts the noise distortion network with the watermark extraction network to perform adversarial training for improving robustness. Experimental results show that our method achieves comparable or better robustness to seen distortions and better robustness to unseen distortions, along with enhanced stability, faster convergence, and improved visual quality in watermarked DTI images. Zhi Li 0012, Zhangyu Liu, Hong Yue, Fei Cheng 0001, Qin Mao, Xuekai Wei, Mingliang Zhou 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2024 | Feature decoupling and interaction network for defending against adversarial examples
Zhi Li 0012, Shuaiwei Liu, Yi Wang 0146 |
Image Vis. Comput. | 2 |
| 2024 | A robust tensor watermarking algorithm for diffusion-tensor imagesabstractWatermarking algorithms that use convolution neural networks have exhibited good robustness in studies of deep learning networks. However, after embedding watermark signals by convolution, the feature fusion efficiency of convolution is relatively low; this can easily lead to distortion in the embedded image. When distortion occurs in medical images, especially in diffusion tensor images (DTIs), the clinical value of the DTI is lost. To address this issue, a robust watermarking algorithm for DTIs implemented by fusing convolution with a Transformer is proposed to ensure the robustness of the watermark and the consistency of sampling distance, which enhances the quality of the reconstructed image of the watermarked DTIs after embedding the watermark signals. In the watermark-embedding network, T1-weighted (T1w) images are used as prior knowledge. The correlation between T1w images and the original DTI is proposed to calculate the most significant features from the T1w images by using the Transformer mechanism. The maximum of the correlation is used as the most significant feature weight to improve the quality of the reconstructed DTI. In the watermark extraction network, the most significant watermark features from the watermarked DTI are adequately learned by the Transformer to robustly extract the watermark signals from the watermark features. Experimental results show that the average peak signal-to-noise ratio of the watermarked DTI reaches 50.47 dB, the diffusion characteristics such as mean diffusivity and fractional anisotropy remain unchanged, and the main axis deflection angle α AC is close to 1. Our proposed algorithm can effectively protect the copyright of the DTI and barely affects the clinical diagnosis. Chengmeng Liu, Zhi Li 0012, Guomei Wang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Effective reversible data hiding scheme for interpolated images using an improved data encoding strategy
Xiangguang Xiong, Zhi Li 0012, Mengting Fan |
Soft Comput. | 2 |
| 2024 | A Robust Adversarial Defense Algorithm for Enhancing Salient Object Detection in Remote Sensing ImageabstractDeep neural networks (DNNs) have been widely applied to salient object detection in optical remote sensing images (ORSIs-SOD) with significant progress. However, the introduction of adversarial perturbations can severely degrade their detection performance. Moreover, the relatively limited labeled samples in ORSI-SOD make the model more susceptible for overfitting, increasing sensitivity to adversarial perturbations. Therefore, researching defense algorithms against adversarial samples to enhance the robustness of salient object detection (SOD) models in the field of ORSIs is of great importance. However, existing defense studies have seldom focused on the ORSI-SOD task and have not considered the unique challenges it faces to propose effective defense solutions. Thus, this article proposes a robust adversarial defense algorithm specifically for ORSI-SOD for the first time. This algorithm first proposes a fixed luminance channel multimodal data random fusion adversarial defense strategy to enhance the model’s robustness against the color and contour information of targets. Then, the algorithm proposes a sinusoidal noise-driven adversarial feature adaptive masking module, which adaptively reduces the impact of adversarial perturbations and enhances the model’s generalization ability. In addition, the algorithm proposes a global context-aware denoising perception module (GCA-DPM) into the network to strengthen the model’s understanding of the global content of images, further suppressing the impact of adversarial noise and enhancing the detection accuracy. Finally, experimental results demonstrate that the proposed defense algorithm effectively enhances the adversarial robustness of various SOD models. Against three different attack strategies, the algorithm increases the average$F_{\beta }$of four SOD models by 23.94% and 21.98% on the optical remote sensing saliency detection (ORSSD) and extended ORSSD (EORSSD) datasets, respectively. Xueli Shi, Zhi Li 0012, Yi Wang 0146 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | VSTNet: Robust watermarking scheme based on voxel space transformation for diffusion tensor imaging images
Zhi Li 0012, Ruwei Luo, Zhangyu Liu, Changhong Li |
J. Inf. Secur. Appl. | 2 |
| 2023 | Robust zero-watermarking algorithm for diffusion-weighted images based on multiscale feature fusion
Zhangyu Liu, Zhi Li 0012, Guomei Wang, Youliang Tian |
Multim. Syst. | 2 |
| 2022 | Multi-scale dehazing network via high-frequency feature fusion
Yongjun Zhang 0007, Zhi Li 0012, Zhongwei Cui, Yitong Yang |
Comput. Graph. | 3 |
| 2022 | Single image deraining using multi-stage and multi-scale joint channel coordinate attention fusion networkabstractRain streaks can seriously degrade the visual quality of an image and are detrimental to subsequent algorithms such as object detection and semantic segmentation. Therefore, removing rain streaks is a very important task. The deraining task has two main limitations: the first is to encode information about rain streaks in different densities and directions, the second is to keep the background details of the image while removing the rain streak. To address these limitations, we propose an effective algorithm, called multi-stage and multi-scale joint channel coordinate attention fusion network (MMAFN). We mainly propose a two-stage network structure, both of which use an encoder-decoder network to extract features. The first-stage network extracts coarse features and the second-stage network integrates the features of the former to further refine features. We design the joint channel coordinate attention block to encode features of rain streaks in different directions and densities. In addition, to better fuse features of different scales and enhance the generalization performance of the network, the inception attention branch block and the multi-level feature fusion block are designed. Extensive experiments substantiate the superiority of the proposed network and prove that our method outperforms the recent state-of-the-art method. The average PSNR of the five test sets is improved by 0.2dB. On the Test100 test set, the PSNR is increased by 0.93dB at most. Yitong Yang, Yongjun Zhang 0007, Zhongwei Cui, Zhi Li 0012, Haoliang Zhao, Yangtin Ou, Heliang Yang, Xihe Wang |
Int. J. Intell. Syst. | 4 |
| 2022 | An efficient robust zero watermarking scheme for diffusion tensor-Magnetic resonance imaging high-dimensional data
Zhi Li 0012, Bin Fan 0004 |
J. Inf. Secur. Appl. | 2 |
| 2022 | DwiMark: a multiscale robust deep watermarking framework for diffusion-weighted imaging images
Bin Fan 0004, Zhi Li 0012 |
Multim. Syst. | 2 |
| 2022 | An adaptive high capacity reversible data hiding algorithm in interpolation domain
Xiangguang Xiong, Lihui Wang 0002, Zhi Li 0012, Mengting Fan, Yue Min Zhu |
Signal Process. | 3 |
| 2015 | Adaptively imperceptible video watermarking based on the local motion entropy
Zhi Li 0012, Xiao-Wei Chen, Jianhua Ma 0002 |
Multim. Tools Appl. | 1 |
| 2009 | Watermarking Relational Databases for Ownership Protection Based on DWTabstractThis paper primarily researches on the feasibility of embedding watermark to relational databases in the discrete wavelet transform (DWT) domain. Watermark embedding is known to be difficult for relational databases. This paper focuses on the analysis of the wavelet high frequency coefficients of corresponding data and gives the definition of the intensive factor. By employing the linear correlation detecting method, this paper proposes the watermarking algorithm, which can embed the watermark into relational database successfully in DWT domain. The watermark can be distributed to different parts of the relational database. Experiments show that the embedded digital watermarks in the proposed algorithm are invisible and some degree of the robustness of the proposed algorithm against the commonly attacks that are used in databases processing. Chuanxian Jiang, Zhi Li 0012 |
IAS | 3 |
| 2007 | Based on Motion Characteristics to Calculate the Adaptive Embedding Tolerance for Imperceptible Video WatermarkingabstractThis scheme describes a procedure that uses the motion characteristics of the videos to calculate the embedding tolerance for imperceptible video watermarking in wavelet domain. This procedure exploits the perceptual properties of HVS and motion content of video sequence adoptively calculating the embedding tolerance of watermark in order to improve the imperceptibility of watermarking according to the content of video. The experiment results indicate that this method can effectively reduce the degradation in the video quality. Zhi Li 0012, Susumu Yamamoto, Takashima Youichi |
CAD/Graphics | 1 |