VLDB 2026 Research / reviewers in the wild / expert
Zhichao Lian
dblp:82/9614
· DBLP profile ↗
58ranked-venue papers
8as first author
44since 2021 · last 2026
0000-0002-6643-8975ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 3 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 7 since 2021Systems, architecture and hardware · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic AlignmentabstractImage-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance. Wei Zhang 0196, Yeying Jin, Xin Li 0082, Yan Zhang 0004, Xiaofeng Cong, Cong Wang 0018, Fengcai Qiao, Zhichao Lian |
AAAI | 8 |
| 2026 | FESA-CLIP: Frequency-Enhanced Semantic-Agnostic Decoupling for Generalizable AI-Generated Image Detection
Huanglei Yang, Zhichao Lian |
ICPR (7) | 3 |
| 2026 | Sonar-FERT: An Accurate Detector Based on RT-DETR for Underwater Sonar Imagery
Yumeng Sun, Zhichao Lian |
ICPR (7) | 2 |
| 2026 | Anomaly Detection Method Based on Dynamic Graph Structure and Multimodal Fusion for IoTabstractIn the era of increasingly complex and covert network attack on Internet of Things (IoT), existing anomaly detection paradigms often exhibit limitations in characterizing the dynamic evolution and cross-source dependencies inherent in multi-source threat scenarios. Most current approaches rely on static graph structures or single-modal features, failing to fully capture the temporal evolution of topology and heterogeneous information within network traffic. To address these challenges, this work has proposed a novel Anomaly Detection method based on Dynamic Graph structure and Multi-Modal Fusion (DGMF-Net) for IoT. First, we have adopted a Conditional Denoising Diffusion Probabilistic Model (CDDPM) to generate high-quality minority class samples, effectively mitigating data imbalance. Second, we have introduced a sliding time window mechanism to construct dynamic graphs, explicitly modeling the temporal dependencies and evolving correlations among traffic features. Furthermore, we have designed a multimodal feature extraction module that learns representations from three dimensions: statistical patterns (via 1D-CNN), temporal sequences (via GRU), and graph structures (via DGCN). Finally, a cross-modal attention fusion mechanism is adopted to dynamically integrate these heterogeneous features, thereby enhancing our model's perception of complex abnormal patterns. Experimental results on the UNSW-NB15 and CIC-IDS-2017 datasets have demonstrated that our proposed method has achieved superior performance, generalization, and robustness compared to state-of-the-art approaches for anomaly detection in a dynamic network environment of IoT with low false alarm rates. Zhichao Lian, Aniruddha Bhattacharjya, Qianmu Li |
IEEE Internet Things J. | 2 |
| 2025 | Latent Diffusion-based Face Anonymization with Identity and Attribute DecouplingabstractFace anonymization aims to change face identity information in images for privacy protection. However, most existing methods fail to achieve an effective balance between utility preservation and face anonymization. The key to preserving the utility of an image lies in retaining its original attributes. In this paper, we accomplish face anonymization by decoupling face attribute features and identity features using the latent diffusion model. On the one hand, the attribute preservation module we designed achieves attribute preservation through the combination of self-attention control and a controllable noise-adding mechanism. On the other hand, we separate the identity features of the face through the identity dissociation module, thus realizing controllable face anonymization. Through sufficient experiments, we show that the images generated by our method possess high utility and high anonymity. Our method can compete with the current state-of-the-art methods. Chenrui Liu, Zhichao Lian |
ICME | 2 |
| 2025 | Free Try-On: Virtual Try-On without Garment-Agnostic Images and Warped GarmentsabstractVirtual try-on focuses on transferring garment images onto target human images. Despite advancements in diffusion-based try-on models, existing methods face limitations, including the loss of human appearance details caused by garment-agnostic images and artifacts from warped garment images. To address these issues, we propose a training method based on pseudo-labeled data, which eliminates the need for garment-agnostic images by leveraging a pretrained try-on model to generate additional training data. Furthermore, we introduce a bidirectional interaction dual UNet and a Reference Fusion mechanism, enabling the generation of high-quality try-on images without using warped garments. Our model also supports try-on image generation from garment text descriptions. Additionally, we enhance Stable Diffusion’s Variational Autoencoder (VAE) with two extra encoders, significantly improving the quality of low-resolution try-on image generation. Experiments on the VITON-HD and DressCode datasets demonstrate superior qualitative and quantitative performance compared to existing methods. Code is available at https://github.com/heiheizwplus/Free-tryon. Xuekang Peng, Zhichao Lian |
ICME | 3 |
| 2025 | SLAG: A Sensitive Layer Activation-Guided Jailbreak Attack on Vision-Language ModelsabstractWe present SLAG, a Sensitive Layer ActivationGuided jailbreak attack for Large Vision-Language Models (LVLMs). SLAG identifies safety-sensitive layers whose activations differ between harmful and benign inputs, and perturbs images using a dual-objective loss to enhance harmful generation while suppressing refusals. On MiniGPT-4, SLAG achieves a 95.0% Attack Success Rate with only 2000 steps, outperforming existing image-only attacks and approaching multimodal methods. Layer analysis shows that a few mid-to-late layers suffice, revealing multiple activation pathways linked to safety failures. Yisheng Li, Zhichao Lian |
ICPADS | 4 |
| 2025 | Meta-Gradient Adversarial Attack for Weight Assignment Based on Model DifferencesabstractImproving the transferability of adversarial attacks against black-box face recognition (FR) models remains a key challenge. Although model ensemble is an effective strategy to enhance transferability, existing methods often rely on simple averaging of losses or gradients, overlooking the inherent differences among models. To overcome this limitation, we present Meta-Gradient Adversarial Attack for Weight Assignment Based on Model Differences (MWMD), a novel approach inspired by meta-learning. MWMD mitigates overfitting from fixed ensemble strategies by adopting a chunk-wise model selection mechanism that dynamically selects different model combinations. It further enhances transferability by quantifying model differences and adaptively adjusting each model's contribution, enabling the capture of intrinsic transferable information. Additionally, a spatial dynamic threshold filter is introduced to align gradient update directions. Extensive experiments on 2 mainstream face datasets and 15 FR models validate that MWMD outperforms state-of-the-art ensemble-based black-box approaches. Zhichao Lian |
ICPADS | 3 |
| 2025 | Hierarchical Integration Knowledge Distillation: Enhancing Adversarial Robustness of Student Models via Clean Data Distillation
Shidong Li, Zhichao Lian |
KSEM (1) | 2 |
| 2025 | MCFA: Multi-Choice Full-body AnonymizationabstractIn the age of data explosion, data collection is everywhere, but data may contain a lot of private information. Therefore, anonymization techniques are used to protect privacy information while preserving important data information. For human images, personal privacy information is present not only in the face; parts of the body also contain a lot of personal identification information. Thus, human anonymization should be extended from the face to the full-body. However, current human anonymization mainly focuses on face anonymization. In this paper, we propose a multi-choice full-body anonymization framework that provides different anonymization choices at the semantic level while ensuring the image diversity and face diversity of full-body anonymized data. For full-body anonymized data, we propose a comprehensive evaluation system. We conduct a detailed evaluation of the anonymized data from three aspects: data quality and diversity, anonymization performance, and data utility. We have demonstrated through extensive experiments that our method provides more anonymous choices and better diversity performance compared to state-of-theart methods while ensuring data quality and anonymization performance. Zhichao Lian |
SMC | 2 |
| 2025 | Adversarial purification of information maskingabstractAdversarial attacks meticulously generate minuscule, imperceptible perturbations that add to images to deceive neural networks . Adversarial purification methods seek to remove perturbations using generative models to achieve defense. However, residual perturbations lead to less-than-ideal results. Under the premise that perturbations are difficult to remove completely, we are the first to quantify the hazards of residual perturbations and explore how to achieve more robust defenses by reducing perturbations and resisting the impact of residual perturbations. Motivated by this, we propose a novel adversarial purification approach named Information Mask Purification (IMPure). Our method utilizes informative masks and a regional intersection reconstruction to generate images to reduce the perturbation residues. During training, we use the combination module to guide the generative model in recovering feature representations. Finally, we establish a combined constraint of pixel loss and perceptual loss to augment the model’s reconstruction adaptability. Extensive experiments on the complex dataset ImageNet with classifier models demonstrate that our approach achieves state-of-the-art results in defending against adversarial attack methods. Implementation code and pre-trained weights can be accessed at https://github.com/NoWindButRain/IMPure . Zhichao Lian, Shuangquan Zhang, Liang Xiao 0001 |
Neurocomputing | 2 |
| 2025 | Centroid-based Contrastive Consistency Learning for transferable deepfake detection
Ruiqi Zha, Zhichao Lian, Qianmu Li |
Neurocomputing | 2 |
| 2025 | InpaintingPose: Enhancing human pose transfer by image inpainting
Chenglin Zhou, Xuekang Peng, Zhichao Lian |
Image Vis. Comput. | 4 |
| 2025 | Diffusion-Based Adversarial Purification With Feature DistillationabstractAdversarial purification is a defense strategy that utilizes generative models to neutralize adversarial perturbations. Diffusion models stand out for their powerful generative ability, making them the latest choice for generative models in adversarial purification methods. It is crucial to keep the balance between robustness against attacks and the integrity of the image content. However, these diffusion models are typically trained exclusively on clean data. We experimentally demonstrate that in the adversarial purification task, the data distribution operated by the diffusion model is often contaminated by adversarial perturbations. Adjusting the diffusion length to effectively mask adversarial perturbations while preserving image label semantics proves challenging. In this paper, we propose an innovative distillation-based diffusion approach for adversarial purification, enabling the diffusion model to operate effectively on the contaminated data distribution. Our approach involves a novel training process that integrates adversarial samples into the diffusion model's training by considering adversarial perturbations as part of the diffusion noise predicted by the model. The key to resisting perturbations is to align the feature representation of the purified image more closely with that of the clean image. To accomplish this, we develop a feature distillation technique to empower the diffusion model to learn and extract clean feature representations from adversarial samples. Extensive experimentation shows that our method achieves state-of-the-art performance against various adaptive attack benchmarks. Zhichao Lian, Liang Xiao 0001 |
IEEE Trans. Big Data | 2 |
| 2025 | Understanding Convolutional Neural Networks From ExcitationsabstractSaliency maps have proven to be a highly efficacious approach for explicating the decisions of convolutional neural networks (CNNs). However, extant methodologies predominantly rely on gradients, which constrain their ability to explicate complex models. Furthermore, such approaches are not fully adept at leveraging negative gradient information to improve interpretive veracity. In this study, we present a novel concept, termed positive and negative excitation (PANE), which enables the direct extraction of PANE for each layer, thus enabling complete layer-by-layer information utilization sans gradients. To organize these excitations into final saliency maps, we introduce a double-chain backpropagation procedure. A comprehensive experimental evaluation, encompassing both binary classification and multiclassification tasks, was conducted to gauge the effectiveness of our proposed method. Encouragingly, the results evince that our approach offers a significant improvement over the state-of-the-art methods in terms of salient pixel removal, minor pixel removal, and inconspicuous adversarial perturbation generation guidance. In addition, we verify the correlation between PANEs. Zijian Ying, Qianmu Li, Zhichao Lian, Jun Hou 0002, Tao Wang 0108 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Boosting Adversarial Transferability by Uniform Scale and Mix Mask Method
Tao Wang 0108, Qianmu Li, Zhichao Lian, Zijian Ying, Shunmei Meng |
ADMA (6) | 3 |
| 2024 | Multi-texture Fusion Attack: A Robust Adversarial Camouflage in Physical World
Yisheng Li, Xuekang Peng, Zhichao Lian |
ICIC (9) | 3 |
| 2024 | Multi-target Attention Dispersion Adversarial Attack Against Aerial Object Detector
Shujuan Wang, Zhichao Lian, Shuohao Li |
ICPR (4) | 3 |
| 2024 | Towards Generalizable Forgery Detection Model via Alignment and Fine-tuningabstractExisting forgery detection models perform well on in-distribution images, but struggle with unseen images. In this paper, we explore the performance of deepfake detection models trained on the large dataset and fine-tuned on small datasets. The experiments show that the models exhibit limited generalization capabilities, and the process of fine-tuning further decreases their ability to generalize. We believe that incorrect feature mapping is key to affecting model performance. Therefore, we tested the linear probing method on pre-trained models, which reduced the loss of generalization performance and confirmed our hypothesis. Furthermore, we proposed a fine-tuning strategy that adapts the model to different datasets on the classifier, improving the performance on both in-distribution and out-of-distribution samples. Xuekang Peng, Yisheng Li, Zhichao Lian |
ISPA | 3 |
| 2024 | Deepfake Video Detection Based on the Decomposition of Spatial-Temporal Attention Mechanism in ViViTabstractWith the rapid development of Deepfake synthesis technology in recent years, our cybersecurity and individual privacy face challenges. In pursuit of robust Deepfake detection, researchers have tried to use temporal cues in videos, employing models such as RNNs and 3D convolutional networks. Despite these efforts, there remains ample opportunity for improvement within these models. In this paper, we introduced an approach utilizing the Video Vision Transformer, which is based on the decomposition of Spatial-Temporal attention mechanisms, for the detection of video face forgery. This method is designed to capture spatial artifacts and temporal inconsistencies. Besides, difference module is introduced to screen features and reduce the interference of natural factors on the model. To enable robust Deepfake detection. We have conducted extensive experiments across some datasets, including FaceForensics++, DFDC and WildDeepfake datasets. Which demonstrates the validity of the model. Guoqing Sun, Zhichao Lian |
ISPA | 2 |
| 2024 | A DDoS Detection Model Based on Feature Construction and Deep ForestabstractDDoS attacks websites and servers by disrupting network services in an attempt to drain the application's resources. With the explosive growth of electric vehicles (EVs), DDoS attacks have threatened the security of EVs and charging stations. In this work, we developed a novel machine learning model named FCDForest to detect DDoS attacks on the CICEV 2023 dataset. FCDForest employs feature construction and deep forest to detect DDoS attacks on the CICEV 2023 dataset. The feature construction is used to construct features, and the deep forest is used as the classification model to detect DDoS attacks. This study selected six existing models as comparison models of FCDForest. In light of our experiments, FCDForest achieved the highest accuracy of 0.94 on the CICEV 2023 dataset. Our experiments indicated that FCDForest is feasible for DDoS attack detection, and feature construction method can improve models’ performance on DDoS attack detection. Shuangquan Zhang, Jiahui Fei, Zhichao Lian |
ISPA | 6 |
| 2024 | A Novel Explainable Method based on Grad-CAM for Network Intrusion DetectionabstractWhen deep learning models are employed in Network Intrusion Detection Systems (NIDSs) to cope with a variety of rising attacks from network, the interpretability of these applications are not studied adequately, which result in the uncertainty of their classification basis and also can not give the warning for how to improve model decisions. In this paper, a new framework is designed to provide a NIDS with visual and quantitative analysis, including a modified ensemble Convolutional Neural Network (CNN) model and a novel explainable method. The ensemble model is used as a feature extractor and aims to make classification. The explainable method in combination with Gradient-weighted Class Activation Mapping (Grad-CAM) is made to calculate feature importance of network traffic from the aspect of spatial relations, and find out the key features for improving model performance. The results of the experiments on NSL-KDD and UNSW-NB15 datasets demonstrate that the new framework, which has a high accuracy comparing with the existing models, can explain the feature importance effectively, and also improve model performance. Zhichao Lian, Shuangquan Zhang, Zhanfeng Wang |
QRS | 2 |
| 2024 | Intrusion Detection System Based on FastICA and Multi-Grained Cascaded ForestabstractWith the advancement of big data, microproces-sors, and other applications, the Internet of Things (IoT) has seen significant development. Owing to the lack of necessary security defense mechanisms, IoT devices are susceptible to being targeted and controlled by attackers. They can manipu-late a vast array of IoT devices to launch DDoS attacks on the network infrastructure of a country or region, leading to serious economic losses and social security risks. Intrusion detection methods based on deep learning generally rely on numerous high-quality training instances, making it difficult to apply them to network traffic lacking sufficient labeled data. Traditional machine learning (ML) methods have limited capabilities in extracting and representing features from high-dimensional data, making it challenging to discover underlying structures and patterns in the data. To address the above issues, this paper proposes an intrusion detection system (IDS) that integrates the Fast Independent Component Analysis (FastICA) module and the multi-Grained Cascade Forest (GcForest). By utilizing FastICA to preprocess the raw data and extract significant features, as well as improving the model's feature extraction capabilities with the aid of cascade modules. Experiments have indicated that FastICA-GcF achieves the binary classification accuracy of 99.95% on CIC-DDoS-2019, and the multiclass classification accuracy of 98.77%, outperforming existing DDoS attack detection models. Jiahui Fei, Shuangquan Zhang, Zhichao Lian |
SMC | 3 |
| 2024 | AAFM-Net: An Ensemble CNN with Auxiliary Attention Filtering Module for Intrusion DetectionabstractIn recent years, electric vehicles (EV) have developed greatly and have begun to gradually replace traditional cars. With this development, the potential security threats faced by electric vehicle network systems are also increasing. To cope with these threats in the EV network, in the paper we propose an auxiliary attention filtering module (AAFM) that cooperates with an ensemble convolutional neural network (CNN). AAFM combines the attention of features from different dimensions, filters irrelevant features, and compensates for the output of the model. The latest CICEV2023 dataset is used for training and evaluation. Experiments demonstrate that AAFM can effectively improve model performance and deal with attacks in extreme situations compared to 8 classic and effective intrusion detection models. AAFM-Net achieves over 90% accuracy. Shuangquan Zhang, Zhichao Lian |
SMC | 3 |
| 2024 | RPID: Boosting Transferability of Adversarial Attacks on Vision TransformersabstractVision Transformers (ViTs) have achieved excellent performance on many computer vision tasks, which has attracted attention of many researchers for their adversarial robustness. As a kind of black-box attack, transfer-based at-tacks usually use adversarial examples generated by a surrogate model to attack structurally different models. It is practical and poses a certain threat to the application of ViTs in critical security areas. Existing transfer-based attacks against ViTs suffer from weak adversarial transferability and noticeable perceptibility. In this work, we propose a method called Reduce Regional Perturbation Interaction and Differentiated (RPID) attack, which employs two strategies of reducing correlation between regional perturbations and adding differentiated perturbations to produce adversarial examples. Extensive experiments demonstrate that our proposed method improves the transferability of the baseline methods for adversarial attacks against ViTs while maintaining stealthiness. Shujuan Wang, Zhichao Lian, Shuohao Li |
SMC | 4 |
| 2024 | Convergent Grey Wolf Optimizer Metaheuristics for Scheduling Crowdsourcing Applications in Mobile Edge ComputingabstractMobile crowdsourcing is a new computing paradigm that enables outsourcing computation tasks to mobile crowd nodes by means of offloading the tasks from the user to a mobile edge computing (MEC) server. This article studies the problem of scheduling security-critical tasks of crowdsourcing applications in a multiserver MEC environment. We formulate this scheduling problem as an integer program and propose a family of convergent grey wolf optimizer (CGWO) metaheuristic algorithms to seek for the best scheduling solutions. Our proposed CGWO uses a task permutation to represent a candidate solution to the formulated scheduling problem, and employs a probability-based mapping scheme to map each search agent in grey wolf optimizer (GWO) onto a valid task permutation. We introduce a new position update strategy for generating the next generation of grey wolf population after each round of search. With this strategy, we prove our proposed CGWO guarantees its convergence to the global best solution. More importantly, we provide a thorough analysis on the movement trajectories of grey wolves during the evolutionary procedure, in order to determine appropriate parameter values such that CGWO would not be trapped in local optima. Experimental results justify the superiority of CGWO metaheuristics over the standard GWO in solving the crowdsourcing task scheduling problem. Zhichao Lian, Jiangang Shu, Yi Zhang 0025, Jin Sun 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Enhancing consistency in virtual try-on: A novel diffusion-based approach
Chenglin Zhou, Zhichao Lian |
Image Vis. Comput. | 3 |
| 2024 | A novel forgery classification method based on multi-scale feature capsule network in mobile edge computingabstractAbstract Face recognition is one of the most important applications of MEC. However, there have been many fake face data that can deceive MEC devices, causing serious problems such as information leakage. Face forgery detection can effectively solve this problem. Current face forgery detection methods have achieved high accuracy. However, most of the methods are researched on the classification of face authenticity. We find that studying multi‐classification of forgery methods can not only improve the accuracy of the model to identify fake faces, but also help improve the generalization ability of fake face classification. We argue that multi‐scale features and high‐frequency features can expose more detailed forgery artifacts. So, we design four modules, which take advantage of the complementarity of RGB features and frequency features, global features and local features. The first module is a residual‐guided multi‐scale spatial attention module, which uses residuals to guide the RGB feature extractor to extract fake features from a multi‐scale perspective. The second module is a multi‐scale retinal feature extraction module. The third module is a multi‐scale channel attention‐guided local frequency statistics module, which extracts local frequency responses using sliding‐window DCT. The last module is a capsule network classification module with overall correlation to classify the fused features. The information transfer between the subject capsule and the classification capsule can increase the integrity of the model, making the model converge faster. We conduct experiments on the databases FaceForensics++, DeepfakeDetection, and FakeAVCeleb. Experimental result shows that our method performs well on forgery classification. Zhichao Lian |
Softw. Pract. Exp. | 1 |
| 2023 | Intrusion Detection System Based on Adversarial Domain Adaptation Algorithm
Jiahui Fei, Yuejin Wang, Zhichao Lian |
GPC (1) | 4 |
| 2023 | Integration Model of Deep Forgery Video Detection Based on rPPG and Spatiotemporal Signal
Lujia Yang, Wenye Shu, Yongjia Wang, Zhichao Lian |
GPC (1) | 4 |
| 2023 | Multi-convolution and Adaptive-Stride Based Transferable Adversarial Attacks
Qingfu Huang, Zhichao Lian |
ICANN (5) | 3 |
| 2023 | Learnable Snake R-CNN for Instance-Level Biomedical Image SegmentationabstractPrecisely knowing each instance’s position and extents is a critical first step in many biological applications. State-of-the-art techniques rely either on deep learning models designed to predict segmentation masks on each Region of Interest (RoI) or on classic active contour methods. The former struggles to precisely delineating boundaries and tends to output masks at low resolutions when the cells/nuclei are very irregular while the latter often needs good initialization and manual setting of parameters, thus limiting their usefulness. To bridge this gap, we introduce Snake R-CNN, a new level of the learnable active contour model that predict boundary on each RoI in a sequent way. To do so, for each RoI, we reformulate the contour deformation task in terms of a hidden state evolution problem and update the evolution process using energy minimization. We learn snake parameterizations per instance in an end-to-end manner, and demonstrate its effectiveness for contour inferences of various cell/nucleus types where consistently higher performances were obtained for comparison against state-of-the-arts. Jie Song 0014, Ziyun Cai, Yurong Song, Guoping Jiang, Zhichao Lian, Liang Xiao 0001 |
ICIP | 5 |
| 2023 | Improving Transferbility of Adversarial Attack on Face Recognition with Feature Attention
Chongyi Lv, Zhichao Lian |
ICONIP (15) | 2 |
| 2023 | MIC: An Effective Defense Against Word-Level Textual Backdoor Attacks
Shufan Yang, Qianmu Li, Zhichao Lian, Pengchuan Wang, Jun Hou 0002 |
ICONIP (6) | 3 |
| 2023 | Interpretable attribution based on concept vectorsabstractWith the development of artificial intelligence, there has been an increasing number of interpretable methods. Under the premise of ensuring the accuracy of the interpretable process, there is a growing emphasis on the comprehensibility of interpretable results. As a result, the interpretable method is gradually shifting from a pixel-level to a concept-level approach. Although the concept-level interpretable results are easier to understand, it needs to generate a lot of concepts and select important concepts from these concepts to explain the CNN model. However, these concepts frequently include duplicated or background-related concepts, which not only affect how the relevance of concepts is measured but also lead to pointless computations. To address these issues, this work, which is based on the TCAV algorithm, provides a pixel-level and concept-level interpretation framework. Given a well-trained classification model, the proposed framework helps to extract the concept of a specific class on a specific layer by adding knowledge guidance, and measures the importance of the concept using a combination of LRP and TCAV. Through extensive experiments, we verify that the proposed framework can provide potential guidance for improving the performance of neural networks. Zhichao Lian |
ICPADS | 2 |
| 2023 | Full-coverage Invisible Camouflage For Adversarial Targeted AttackabstractWith the rapid advancement of artificial intelligence technology in recent years, there has been an increasing focus on security issues. Deep learning models, despite their capabilities, are susceptible to various attacks such as counter samples, patches, and camouflage, which can lead to inaccurate outputs. However, existing physical attack methods often overlook the importance of environmental invisibility, making it easier to detect camouflaged objects. In this study, we propose a novel method for generating targeted covert camouflage to effectively attack detection models while remaining undetected. Specifically, our method incorporates targeted attack loss into the yolov3 detection model and combines style transfer and three-dimensional camouflage training to seamlessly integrate camouflage with the surrounding environment. This approach significantly enhances the camouflage’s ability to blend into the environment. Through extensive experiments, we have demonstrated the effectiveness of our targeted covert camouflage method, which outperforms other existing approaches in terms of concealment. Hongtuo Zhou, Zhichao Lian |
ICPADS | 2 |
| 2023 | A unified RGB-T crowd counting learning framework
Siqi Gu, Zhichao Lian |
Image Vis. Comput. | 2 |
| 2023 | Deep Open-Curve Snake for Discriminative 3D Neuron TrackingabstractOpen-Curve Snake (OCS) has been successfully used in three-dimensional tracking of neurites. However, it is limited when dealing with noise-contaminated weak filament signals in real-world applications. In addition, its tracking results are highly sensitive to initial seeds and depend only on image gradient-derived forces. To address these issues and boost the canonical OCS tracker to a new level of learnable deep learning algorithms, we present Deep Open-Curve Snake (DOCS), a novel discriminative 3D neuron tracking framework that simultaneously learns a 3D distance-regression discriminator and a 3D deeply-learned tracker under the energy minimization, which can promote each other. In particular, the open curve tracking process in DOCS is formed as convolutional neural network prediction procedures of new deformation fields, stretching directions, and local radii and iteratively updated by minimizing a tractable energy function containing fitting forces and curve length. By sharing the same deep learning architectures in an end-to-end trainable framework, DOCS is able to fully grasp the information available in the volumetric neuronal data to address segmentation, tracing, and reconstruction of complete neuron structures in the wild. We demonstrated the superiority of DOCS by evaluating it on both the BigNeuron and Diadem datasets where consistently state-of-the-art performances were achieved for comparison against current neuron tracing and tracking approaches. Our method improves the average overlap score and distance score about 1.7% and 17% in the BigNeuron challenge data set, respectively, and the average overlap score about 4.1% in the Diadem dataset. Jie Song 0014, Zhichao Lian, Liang Xiao 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Mutil-level Local Alignment and Semantic Matching Network for Image-Text Retrieval
Zhukai Jiang, Zhichao Lian |
ICANN (3) | 2 |
| 2022 | A Novel Contrastive Learning Framework for Self-Supervised Anomaly DetectionabstractAnomaly detection is significant in the field of computer vision and refers to identifying those samples in dataset that are different from normal samples. In practice, abnormal products are rare and anomaly detection usually calculates the difference between the inputs and the reconstructed images by reconstruction-based methods. Contrastive learning both maximizes the similarity between a sample and its augmentations, and the differences between different samples, which is suitable for improving the detection capability of the autoencoder. Inspired by this, we design a novel contrastive learning architecture for anomaly detection. In this work, we make reasonable sample pairs to simulate possible real anomalies and maximizes the distance between normal and abnormal samples. Remarkably, our approach improves the vanilla autoencoder model by 14.4% in terms of the AUROC score on the MVTec AD. Jingze Li, Zhichao Lian |
ICIP | 2 |
| 2022 | An Effective Fusion Method to Enhance the Robustness of CNNabstractWith the development of technology rapidly, applications of convolutional neural networks have improved the convenience of our life. However, in image classification field, it has been found that when some perturbations are added to images, the CNN would misclassify it. Thus various defense methods have been proposed. The previous approach only considered how to incorporate modules in the network to improve robustness, but did not focus on the way the modules were incorporated. In this paper, we design a new fusion method to enhance the robustness of CNN. We use a dot product-based approach to add the denoising module to ResNet18 and the attention mechanism to further improve the robustness of the model. The experimental results on CIFAR10 have shown that our method is effective and better than the state-of-the-art methods under the attack of FGSM and PGD. Yating Ma, Zhichao Lian |
ICIP | 2 |
| 2022 | Attention Based Adversarial Attacks with Low PerturbationsabstractDeep neural networks are vulnerable to adversarial examples generated by black-box attacks with tiny perturbations. However, black-box based transferability attacks usually add perturbation to the whole image. It is easy for the defender to detect the adversarial examples. Inspired by attention modules, we propose a method named Gradient-mask and Attention-whey (GM&AW) to reduce redundant noise and maintain attack effect of them in this work. During Gradient-mask iterations, we only choose the regions with greater gradient and update them along the direction of gradient. Then, we utilize Attention-whey optimization combining attention mechanism and query to further decrease noise. Extensive experiments on Imagenet demonstrate that our approach can achieve enormous decrease on noise in$\ell_{2}$norm and keep attack success rate for black models. Qingfu Huang, Zhichao Lian, Qianmu Li |
ICME | 2 |
| 2022 | Fused Pruning based Robust Deep Neural Network Watermark EmbeddingabstractDeep Neural Network (DNN) models are usually trained with tremendous data and computation resources. Thus, DNN models are now regarded as important assets, however facing a great risk of being stolen and illegal distribution. In recent years, watermark is introduced to protect the ownership of DNN models. The watermark can be extracted in a relatively simple way to declare the ownership of the model. However, watermark is vulnerable to be attacked. In this work, we propose a watermark defense method for DNN model based on pruning. Inspired by the pruning methods, we design a fused channel-wise pruning strategy which selects important filters for watermarks embedding. Specifically, we introduce a novel method to enhance the watermark robustness by selecting important filters as the watermark carrier based on multiple pruning methods, including network slimming, efficient filter and entropy. We conduct experiments on the VGG-19 model with the CIFAR-10 dataset. The experimental results show that this method is robust against fine-tuning attack, pruning attack and overwriting attack. In addition, our method does not significantly change the distribution of model weights so that the watermark is hard to be detected. Shuo Wang 0017, Huiyun Jing, Zhichao Lian, Shunmei Meng, Qianmu Li |
ICPR | 4 |
| 2021 | Sparse Coding Driven Deep Decision Tree Ensembles for Nucleus Segmentation in Digital Pathology ImagesabstractAutomating generalized nucleus segmentation has proven to be non-trivial and challenging in digital pathology. Most existing techniques in the field rely either on deep neural networks or on shallow learning-based cascading models. The former lacks theoretical understanding and tends to degrade performance when only limited amounts of training data are available while the latter often suffers from limitations for generalization. To address these issues, we propose sparse coding driven deep decision tree ensembles (ScD2TE), an easily trained yet powerful representation learning approach with performance highly competitive to deep neural networks in the generalized nucleus segmentation task. We explore the possibility of stacking several layers based on fast convolutional sparse coding–decision tree ensemble pairwise modules and generate a layer-wise encoder–decoder architecture with intra-decoder and inter-encoder dense connectivity patterns. Under this architecture, all the encoders share the same assumption across the different layers to represent images and interact with their decoders to give fast convergence. Compared with deep neural networks, our proposed ScD2TE does not require back-propagation computation and depends on less hyper-parameters. ScD2TE is able to achieve a fast end-to-end pixel-wise training in a layer-wise manner. We demonstrated the superiority of our segmentation method by evaluating it on the multi-disease state and multi-organ dataset where consistently higher performances were obtained for comparison against other state-of-the-art deep learning techniques and cascading methods with various connectivity patterns. Jie Song 0014, Liang Xiao 0001, Mohsen Molaei, Zhichao Lian |
IEEE Trans. Image Process. | 4 |
| 2020 | PSDNet: A Balanced Architecture of Accuracy and Parameters for Semantic SegmentationabstractIn this paper, we present our Pyramid Pooling Module (PPM) with SE1Cblock and D2SUpsample Network (PSDNet), a novel architecture for accurate semantic segmentation. Started from the known work called Pyramid Scene Parsing Network (PSPNet), PSDNet takes advantage of pyramid pooling structure with channel attention module and feature transformation module in Pyramid Pooling Module (PPM). The enhanced PPM with these two components can strengthen context information flowing in the network instead of damaging it. The channel attention module we mentioned is an improved “Squeeze and Excitation with 1D Convolution” (SE1C) block which can explicitly model interrelationship between channels with fewer number of parameters. We propose a feature transformation module named “Depth to Space Upsampling” (D2SUpsample) in the PPM which keeps integrity of features by transforming features instead of interpolating features, at the same time reducing parameters. In addition, we introduce a joint strategy in SE1Cblock which combines two variants of global pooling without increasing parameters. Compared with PSPNet, our work achieves higher accuracy on public datasets with 73.97% mIoU and 82.89% mAcc accuracy on Cityscapes Dataset based on ResNet50 backbone. Zhichao Lian |
ICPR | 2 |
| 2020 | A Novel Unsupervised Hashing Method for Image Retrieval Based on K-Reciprocal Nearest Neighbors
Zhichao Lian |
PRCV (2) | 2 |
| 2020 | A Real Time Face Tracking System based on Multiple Information Fusion
Zhichao Lian, Chanying Huang |
Multim. Tools Appl. | 1 |
| 2019 | Multi-layer boosting sparse convolutional model for generalized nuclear segmentation from histopathology images
Jie Song 0014, Liang Xiao 0001, Mohsen Molaei, Zhichao Lian |
Knowl. Based Syst. | 4 |
| 2019 | Locally Low-Rank Regularized Video Stabilization With Motion Diversity ConstraintsabstractThis paper presents a novel motion aware regularization model with diversity constraints for motion smoothing in video stabilization. Differing from the global path optimization methods, the proposed model puts an emphasis on the relations of inter-frame motions and incorporates sliding windowed low-rank and smoothness constraints. The rationale behind it is cinematography rules which assume that camera motions can be divided into diverse patterns: zero velocity, constant velocity, and acceleration motion. Firstly, a locally motion aware fidelity term is adopted in light of the local motion stationarity. Secondly, to improve the robustness of the model for different motion patterns, a locally low-rank constrained regularization term is further introduced by considering the motion correlation in a local temporal window. Moreover, to cope with the over-smoothing problem in rapid motion situations with extreme acceleration, a motion steering kernel and varying window length are employed to enhance the flexibility of the proposed model. The experimental results demonstrate the superiority of the proposed optimization model and the efficiency to suppress over-smoothing when rapid motions occur. Meanwhile, we compare the results with some state-of-the-art methods quantitatively and qualitatively, and our method can achieve a comparable or even better stabilization effect. Huicong Wu, Liang Xiao 0001, Zhichao Lian, Hiuk Jae Shim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Contour-Seed Pairs Learning-Based Framework for Simultaneously Detecting and Segmenting Various Overlapping Cells/Nuclei in Microscopy ImagesabstractIn this paper, we propose a novel contour-seed pairs learning-based framework for robust and automated cell/nucleus segmentation. Automated granular object segmentation in microscopy images has significant clinical importance for pathology grading of the cell carcinoma and gene expression. The focus of the past literature is dominated by either segmenting a certain type of cells/nuclei or simply splitting the clustered objects without contours inference of them. Our method addresses these issues by formulating the detection and segmentation tasks in terms of a unified regression problem, where a cascade sparse regression chain model is trained and then applied to return object locations and entire boundaries of clustered objects. In particular, we first learn a set of online convolutional features in each layer. Then, in the proposed cascade sparse regression chain, with the input from the learned features, we iteratively update the locations and clustered object boundaries until convergence. In this way, the boundary evidences of each individual object can be easily delineated and be further fed to a complete contour inference procedure optimized by the minimum description length principle. For any probe image, our method enables to analyze free-lying and overlapping cells with complex shapes. Experimental results show that the proposed method is very generic and performs well on contour inferences of various cell/nucleus types. Compared with the current segmentation techniques, our approach achieves state-of-the-art performances on four challenging datasets, i.e., the kidney renal cell carcinoma histopathology dataset, Drosophila Kc167 cellular dataset, differential interference contrast red blood cell dataset, and cervical cytology dataset. Jie Song 0014, Liang Xiao 0001, Zhichao Lian |
IEEE Trans. Image Process. | 3 |
| 2017 | A novel adaptive kernel correlation filter tracker with multiple feature integrationabstractRecently, correlation filters (CFs) for visual tracking present competitive performances on both accuracy and robustness, but there is still a need for improving their overall tracking capabilities. Most CF trackers learn a best filter to regress training data to a fixed target response, which might lead to drifting. In this paper, we present an appealing tracker based on the Kernelized Correlation Filter (KCF), which can adaptively change the target response. Furthermore, we utilize a fast and accurate scale estimation approach by learning an independent correlation filter instead of employing an exhaustive scale search strategy to estimate the target size. In addition, color naming integrates into the histogram of orientation gradient feature to further boost the performance for our tracker. We validate our tracker on the popular OTB50 datasets, which outperforms the state-of-the-art methods in terms of efficiency and accuracy. Zhonggeng Liu, Zhichao Lian |
ICIP | 2 |
| 2017 | Boundary-to-Marker Evidence-Controlled Segmentation and MDL-Based Contour Inference for Overlapping NucleiabstractThis paper presents a novel method for automated morphology delineation and analysis of cell nuclei in histopathology images. Combining the initial segmentation information and concavity measurement, the proposed method first segments clusters of nuclei into individual pieces, avoiding segmentation errors introduced by the scale-constrained Laplacian-of-Gaussian filtering. After that a nuclear boundary-to-marker evidence computing is introduced to delineate individual objects after the refined segmentation process. The obtained evidence set is then modeled by the periodic B-splines with the minimum description length principle, which achieves a practical compromise between the complexity of the nuclear structure and its coverage of the fluorescence signal to avoid the underfitting and overfitting results. The algorithm is computationally efficient and has been tested on the synthetic database as well as 45 real histopathology images. By comparing the proposed method with several state-of-the-art methods, experimental results show the superior recognition performance of our method and indicate the potential applications of analyzing the intrinsic features of nuclei morphology. Jie Song 0014, Liang Xiao 0001, Zhichao Lian |
IEEE J. Biomed. Health Informatics | 3 |
| 2016 | ELM-based classification of ADHD patients using a novel local feature extraction methodabstractRecently, it has been an increasing interest in modeling abnormal temporal dynamics of functional interactions in psychiatric disorders. However, the accuracy of differentiating attention-deficit/hyperactivity disorder (ADHD) children form normal children has still much space for improvement. To further improve the accuracy, the key issue is to extract more effective features from original fMRI data. In this paper, we propose a novel local feature extraction method named Local Binary Encoding Method (LBEM) that can effectively characterize functional interaction patterns (FIPs). In particular, we show that the proposed method can well discriminate the functional interaction abnormalities, which is composed of a Bayesian connectivity change point model, a local feature extraction method and a kernel Extreme Learning Machine (ELM)-based classifier. The experiment on a real dataset of 23 ADHD children and 45 normal control (NC) children has shown that our method achieved better classification performance compared to the existing methods. Zhichao Lian, Min Li 0009, Zhonggeng Liu, Liang Xiao 0001, Zhihui Wei |
BIBM | 2 |
| 2012 | Local Line Derivative Pattern for face recognitionabstractIn this paper, we propose a novel face descriptor for face recognition, named Local Line Derivative Pattern (LLDP). High-order derivative images in two directions are obtained by convolving original images with Sobel Masks. A revised binary coding function is proposed and three standards on arranging the weights are also proposed. Based on the standards, the weights of a line neighborhood in two directions are arranged. The LLDP labels in two directions are calculated with the proposed binary coding function and weights. The labeled image is divided into blocks where spatial histograms are extracted separately and concatenated into an entire histogram as features for recognition. The experiments on the FERET and Extended Yale B show superior performances of the proposed LLDP compared to other existing methods based on the LBP. The results prove that the LLDP has good robustness against expression, illumination and aging variations. Zhichao Lian, Meng Joo Er, Yang Cong |
ICIP | 1 |
| 2012 | A novel efficient local illumination compensation method based on DCT in logarithm domain
Zhichao Lian, Meng Joo Er, Yanchun Liang 0001 |
Pattern Recognit. Lett. | 1 |
| 2011 | A Novel Face Recognition Approach under Illumination Variations Based on Local Binary Pattern
Zhichao Lian, Meng Joo Er, Juekun Li |
CAIP (2) | 1 |
| 2011 | A Novel Local Illumination Normalization Approach for Face Recognition
Zhichao Lian, Meng Joo Er, Juekun Li |
ISNN (2) | 1 |
| 2010 | An efficient illumination normalization method in a transformed domainabstractThis paper proposes a novel illumination normalization approach with low computation complexity for face recognition. In this proposed method, a block-wise Walsh-Hadamard transform (WHT) is employed in the logarithm domain. An appropriate number of low-frequency WHT coefficients are zeroed to compensate for illumination variations. Experiments on different databases demonstrate that the proposed method obtains results comparable to those of conventional Discrete Cosine Transform method but with a higher efficiency. It also achieves better performances for cases with larger illumination variations. Furthermore, both analytical proof and experimental results demonstrate that principal component analysis (PCA) can be directly implemented in the WHT domain. Zhichao Lian, Meng Joo Er |
ICARCV | 1 |