Jingzhi Li 0002

dblp:17/1271-2 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
16since 2021 · last 2025
0000-0001-7054-9267ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Interpreting Object-level Foundation Models via Visual Precision Search
abstract
Advances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection. However, interpreting these models’ decisions has grown increasingly challenging. Existing interpretable attribution methods for object-level task interpretation have notable limitations: (1) gradient-based methods lack precise localization due to visual-textual fusion in foundation models, and (2) perturbation-based methods produce noisy saliency maps, limiting fine-grained interpretability. To address these, we propose a Visual Precision Search method that generates accurate attribution maps with fewer regions. Our method bypasses internal model parameters to overcome attribution issues from multimodal fusion, dividing inputs into sparse sub-regions and using consistency and collaboration scores to accurately identify critical decision-making regions. We also conducted a theoretical analysis of the boundary guarantees and scope of applicability of our method. Experiments on RefCOCO, MS COCO, and LVIS show our approach enhances object-level task interpretability over SOTA for Grounding DINO and Florence-2 across various evaluation metrics, with faithfulness gains of 23.7%, 31.6%, and 20.1% on MS COCO, LVIS, and RefCOCO for Grounding DINO, and 50.7% and 66.9% on MS COCO and RefCOCO for Florence-2. Additionally, our method can interpret failures in visual grounding and object detection tasks, surpassing existing methods across multiple evaluation metrics. The code is released at https://github.com/RuoyuChen10/VPS.
Ruoyu Chen 0001, Siyuan Liang 0004, Jingzhi Li 0002, Shiming Liu, Maosen Li, Zhen Huang 0006, Hua Zhang 0008, Xiaochun Cao
CVPR3
2025 FaceInsight: A Multimodal Large Language Model for Face Perception
abstract
Recent advances in multimodal large language models (MLLMs) have demonstrated strong capabilities in understanding general visual content. However, these general-domain MLLMs perform poorly in face perception tasks, often producing inaccurate or misleading responses to face-specific queries. To address this gap, we propose FaceInsight, a versatile face perception MLLM that provides fine-grained information. Our approach introduces visual textual alignment of facial knowledge to model both uncertain dependencies and deterministic relationships among facial information, mitigating the limitations of language-driven reasoning. Additionally, we incorporate face segmentation maps as an auxiliary perceptual modality, enriching visual input with localized structural cues to enhance semantic understanding. Comprehensive experiments show that FaceInsight consistently outperforms nine compared MLLMs under both training-free and fine-tuned settings.
Jingzhi Li 0002, Changjiang Luo, Ruoyu Chen 0001, Hua Zhang 0008, Wenqi Ren, Jianhou Gan, Xiaochun Cao
ACM Multimedia1
2025 Anti-Fake Vaccine: Safeguarding Privacy Against Face Swapping via Visual-Semantic Dual Degradation
Jingzhi Li 0002, Changjiang Luo, Hua Zhang 0008, Yang Cao 0011, Xiaochun Cao
Int. J. Comput. Vis.1
2025 Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
abstract
The objective of few-shot object detection (FSOD) is to detect novel objects with few training samples. The core challenge of this task is how to construct a generalized feature space for novel categories with limited data on the basis of the base category space, which could adapt the learned detection model to unknown scenarios. Most existing fine-tuning-based approaches tackle the challenge via pre-training a feature extractor based on the base categories and then fine-tuning the detector through the novel categories. However, limited by insufficient samples for novel categories, two issues still exist: (1) the features of the novel category are easily implicitly represented by the features of the base category, leading to inseparable classifier boundaries, (2) novel categories with fewer data are not enough to fully represent the distribution, where the model fine-tuning is prone to overfitting. To address these issues, we introduce the side information to alleviate the negative influences derived from the feature space and sample viewpoints and formulate a novel generalized feature representation learning method for FSOD. Specifically, we first utilize embedding side information to construct a knowledge matrix to quantify the semantic relationship between the base and novel categories. Then, to strengthen the discrimination between semantically similar categories, we further develop contextual semantic supervised contrastive learning which embeds side information. Furthermore, to prevent overfitting problems caused by sparse samples, a side-information guided region-aware masked module is introduced to augment the diversity of samples, which finds and abandons biased information that discriminates between similar categories via counterfactual explanation, and refines the discriminative representation space further. Finally, we theoretically analyze the generalization bound for introducing our proposed module and demonstrate that our proposed model can effectively reduce the upper bound of the generalization error. Extensive experiments using ResNet and ViT backbones on PASCAL VOC, MS COCO, LVIS V1, FSOD-1 K, and FSVOD-500 benchmarks demonstrate that our model outperforms the previous state-of-the-art methods, significantly improving the ability of FSOD in most shots/splits. The code is released athttps://github.com/RuoyuChen10/CCL-FSOD.
Ruoyu Chen 0001, Hua Zhang 0008, Jingzhi Li 0002, Li Liu 0002, Zhen Huang 0006, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Logit Standardization in Knowledge Distillation
abstract
Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relationsfrom teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic dis-tillation evaluation; nonetheless, this challenge is success-fully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github.
Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Rui Wang 0032, Xiaochun Cao
CVPR3
2024 Less is More: Fewer Interpretable Region via Submodular Subset Selection
abstract
Image attribution algorithms aim to identify important regions that are highly relevant to model decisions. Although existing attribution solutions can effectively assign importance to target elements, they still face the following challenges: 1) existing attribution methods generate inaccurate small regions thus misleading the direction of correct attribution, and 2) the model cannot produce good attribution results for samples with wrong predictions. To address the above challenges, this paper re-models the above image attribution problem as a submodular subset selection problem, aiming to enhance model interpretability using fewer regions. To address the lack of attention to local regions, we construct a novel submodular function to discover more accurate small interpretation regions. To enhance the attribution effect for all samples, we also impose four different constraints on the selection of sub-regions, i.e., confidence, effectiveness, consistency, and collaboration scores, to assess the importance of various subsets. Moreover, our theoretical analysis substantiates that the proposed function is in fact submodular. Extensive experiments show that the proposed method outperforms SOTA methods on two face datasets (Celeb-A and VGG-Face2) and one fine-grained dataset (CUB-200-2011). For correctly predicted samples, the proposed method improves the Deletion and Insertion scores with an average of 4.9\% and 2.5\% gain relative to HSIC-Attribution. For incorrectly predicted samples, our method achieves gains of 81.0\% and 18.4\% compared to the HSIC-Attribution algorithm in the average highest confidence and Insertion score respectively. The code is released at https://github.com/RuoyuChen10/SMDL-Attribution.
Ruoyu Chen 0001, Hua Zhang 0008, Siyuan Liang 0004, Jingzhi Li 0002, Xiaochun Cao
ICLR4
2024 Toward Generalized Few-Shot Open-Set Object Detection
abstract
Open-set object detection (OSOD) aims to detect the known categories and reject unknown objects in a dynamic world, which has achieved significant attention. However, previous approaches only consider this problem in data-abundant conditions, while neglecting the few-shot scenes. In this paper, we seek a solution for the generalized few-shot open-set object detection (G-FOOD), which aims to avoid detecting unknown classes as known classes with a high confidence score while maintaining the performance of few-shot detection. The main challenge for this task is that few training samples induce the model to overfit on the known classes, resulting in a poor open-set performance. We propose a new G-FOOD algorithm to tackle this issue, named Few-shOt Open-set Detector (FOOD), which contains a novel class weight sparsification classifier (CWSC) and a novel unknown decoupling learner (UDL). To prevent over-fitting, CWSC randomly sparses parts of the normalized weights for the logit prediction of all classes, and then decreases the co-adaptability between the class and its neighbors. Alongside, UDL decouples training the unknown class and enables the model to form a compact unknown decision boundary. Thus, the unknown objects can be identified with a confidence probability without any threshold, prototype, or generation. We compare our method with several state-of-the-art OSOD methods in few-shot scenes and observe that our method improves the F-score of unknown classes by 4.80%-9.08% across all shots in VOC-COCO dataset settings.
Binyi Su, Hua Zhang 0008, Jingzhi Li 0002, Zhong Zhou
IEEE Trans. Image Process.3
2023 Generating Transferable 3D Adversarial Point Cloud via Random Perturbation Factorization
abstract
Recent studies have demonstrated that existing deep neural networks (DNNs) on 3D point clouds are vulnerable to adversarial examples, especially under the white-box settings where the adversaries have access to model parameters. However, adversarial 3D point clouds generated by existing white-box methods have limited transferability across different DNN architectures. They have only minor threats in real-world scenarios under the black-box settings where the adversaries can only query the deployed victim model. In this paper, we revisit the transferability of adversarial 3D point clouds. We observe that an adversarial perturbation can be randomly factorized into two sub-perturbations, which are also likely to be adversarial perturbations. It motivates us to consider the effects of the perturbation and its sub-perturbations simultaneously to increase the transferability for sub-perturbations also contain helpful information. In this paper, we propose a simple yet effective attack method to generate more transferable adversarial 3D point clouds. Specifically, rather than simply optimizing the loss of perturbation alone, we combine it with its random factorization. We conduct experiments on benchmark dataset, verifying our method's effectiveness in increasing transferability while preserving high efficiency.
Bangyan He, Jian Liu 0012, Yiming Li 0004, Siyuan Liang 0004, Jingzhi Li 0002, Xiaojun Jia, Xiaochun Cao
AAAI5
2023 Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
abstract
Knowledge Distillation (KD) for object detection aims to train a compact detector by transferring knowledge from a teacher model. Since the teacher model perceives data in a way different from humans, existing KD methods only distill knowledge that is consistent with labels annotated by human expert while neglecting knowledge that is not consistent with human perception, which results in insufficient distillation and sub-optimal performance. In this paper, we propose inconsistent knowledge distillation (IKD), which aims to distill knowledge inherent in the teacher model's counter-intuitive perceptions. We start by considering the teacher model's counter-intuitive perceptions of frequency and non-robust features. Unlike previous works that exploit fine-grained features or introduce additional regularizations, we extract inconsistent knowledge by providing diverse input using data augmentation. Specifically, we propose a sample-specific data augmentation to transfer the teacher model's ability in capturing distinct frequency components and suggest an adversarial feature augmentation to extract the teacher model's perceptions of non-robust features in the data. Extensive experiments demonstrate the effectiveness of our method which outperforms state-of-the-art KD baselines on one-stage, two-stage and anchor-free object detectors (at most +1.0 mAP). Our codes will be made available at https://github.com/JWLiang007/IKD.git.
Siyuan Liang 0004, Aishan Liu, Ke Ma 0001, Jingzhi Li 0002, Xiaochun Cao
ACM Multimedia5
2023 Privacy-Enhancing Face Obfuscation Guided by Semantic-Aware Attribution Maps
abstract
Face recognition technology is increasingly being integrated into our daily life, e.g. Face ID. With the advancement of machine learning algorithms, the personal information such as age, gender, and race can be easily deduced from the recorded face images in these applications. This poses a serious privacy threat to individuals who do not want to be profiled, as face images are collected for biometric purposes. Existing methods mostly focus on adding the invisible adversarial perturbations into the images to make automatic inference infeasible. However, the application scenarios of these methods are limited due to the perturbations depending on the specific model. In this paper, we introduce a novel face privacy-enhancing framework by obfuscating the stored faces, which could maintain the data utility (face identity) while protecting the privacy of users (facial attributes). Specifically, we first develop a feature attribution module to discover the identity-related facial parts. Within this module, we introduce a pixel importance estimation model based on Shapley value to obtain a pixel-level attribution map, and then each pixel on the attribution map is aggregated into semantic facial parts, which are used to quantify the importance of different facial parts. Next, we design a privacy-enhancing module to generate the high-quality obfuscated images, which can modify the privacy semantic content and preserve the identity-related information. Using the proposed method, users can choose the single or multiple attributes to be obfuscated without affecting identity matching. Extensive experiments conducted on CelebA-HQ and VGGFace2-HQ benchmarks demonstrate the effectiveness and generalization ability of our method.
Jingzhi Li 0002, Hua Zhang 0008, Siyuan Liang 0004, Pengwen Dai, Xiaochun Cao
IEEE Trans. Inf. Forensics Secur.1
2023 Prediction With Visual Evidence: Sketch Classification Explanation via Stroke-Level Attributions
abstract
Sketch classification models have been extensively investigated by designing a task-driven deep neural network. Despite their successful performances, few works have attempted to explain the prediction of sketch classifiers. To explain the prediction of classifiers, an intuitive way is to visualize the activation maps via computing the gradients. However, visualization based explanations are constrained by several factors when directly applying them to interpret the sketch classifiers: (i) low-semantic visualization regions for human understanding. and (ii) neglecting of the inter-class correlations among distinct categories. To address these issues, we introduce a novel explanation method to interpret the decision of sketch classifiers with stroke-level evidences. Specifically, to achieve stroke-level semantic regions, we first develop a sketch parser that parses the sketch into strokes while preserving their geometric structures. Then, we design a counterfactual map generator to discover the stroke-level principal components for a specific category. Finally, based on the counterfactual feature maps, our model could explain the question of "why the sketch is classified as X" by providing positive and negative semantic explanation evidences. Experiments conducted on two public sketch benchmarks, Sketchy-COCO and TU-Berlin, demonstrate the effectiveness of our proposed model. Furthermore, our model could provide more discriminative and human understandable explanations compared with these existing works.
Sixuan Liu, Jingzhi Li 0002, Hua Zhang 0008, Long Xu 0001, Xiaochun Cao
IEEE Trans. Image Process.2
2023 Event-Aware Video Deraining via Multi-Patch Progressive Learning
abstract
In this paper, we address the problem of video-based rain streak removal by developing an event-aware multi-patch progressive neural network. Rain streaks in video exhibit correlations in both temporal and spatial dimensions. Existing methods have difficulties in modeling the characteristics. Based on the observation, we propose to develop a module encoding events from neuromorphic cameras to facilitate deraining. Events are captured asynchronously at pixel-level only when intensity changes by a margin exceeding a certain threshold. Due to this property, events contain considerable information about moving objects including rain streaks passing though the camera across adjacent frames. Thus we suggest that utilizing it properly facilitates deraining performance non-trivially. In addition, we develop a multi-patch progressive neural network. The multi-patch manner enables various receptive fields by partitioning patches and the progressive learning in different patch levels makes the model emphasize each patch level to a different extent. Extensive experiments show that our method guided by events outperforms the state-of-the-art methods by a large margin in synthetic and real-world datasets.
Shangquan Sun, Wenqi Ren, Jingzhi Li 0002, Kaihao Zhang, Meiyu Liang, Xiaochun Cao
IEEE Trans. Image Process.3
2023 Sim2Word: Explaining Similarity with Representative Attribute Words via Counterfactual Explanations
abstract
Recently, we have witnessed substantial success using the deep neural network in many tasks. Although there still exist concerns about the explainability of decision making, it is beneficial for users to discern the defects in the deployed deep models. Existing explainable models either provide the image-level visualization of attention weights or generate textual descriptions as post hoc justifications. Different from existing models, in this article we propose a new interpretation method that explains the image similarity models by salience maps and attribute words. Our interpretation model contains visual salience maps generation and the counterfactual explanation generation. The former has two branches: global identity relevant region discovery and multi-attribute semantic region discovery. The first branch aims to capture the visual evidence supporting the similarity score, which is achieved by computing counterfactual feature maps. The second branch aims to discover semantic regions supporting different attributes, which helps to understand which attributes in an image might change the similarity score. Then, by fusing visual evidence from two branches, we can obtain the salience maps indicating important response evidence. The latter will generate the attribute words that best explain the similarity using the proposed erasing model. The effectiveness of our model is evaluated on the classical face verification task. Experiments conducted on two benchmarks—VGGFace2 and Celeb-A—demonstrate that our model can provide convincing interpretable explanations for the similarity. Moreover, our algorithm can be applied to evidential learning cases, such as finding the most characteristic attributes in a set of face images, and we verify its effectiveness on the VGGFace2 dataset.
Ruoyu Chen 0001, Jingzhi Li 0002, Hua Zhang 0008, Changchong Sheng, Li Liu 0002, Xiaochun Cao
ACM Trans. Multim. Comput. Commun. Appl.2
2022 A Large-Scale Multiple-objective Method for Black-box Attack Against Object Detection
Siyuan Liang 0004, Longkang Li, Yanbo Fan, Xiaojun Jia, Jingzhi Li 0002, Baoyuan Wu, Xiaochun Cao
ECCV (4)5
2022 Accurate Scene Text Detection Via Scale-Aware Data Augmentation and Shape Similarity Constraint
abstract
Scene text detection has attracted increasing concerns with the rapid development of deep neural networks in recent years. However, existing scene text detectors may overfit on the public datasets due to the limited training data, or generate inaccurate localization for arbitrary-shape scene texts. This paper presents an arbitrary-shape scene text detection method that can achieve better generalization ability and more accurate localization. We first propose a Scale-Aware Data Augmentation (SADA) technique to increase the diversity of training samples. SADA considers the scale variations and local visual variations of scene texts, which can effectively relieve the dilemma of limited training data. At the same time, SADA can enrich the training minibatch, which contributes to accelerating the training process. Furthermore, a Shape Similarity Constraint (SSC) technique is exploited to model the global shape structure of arbitrary-shape scene texts and backgrounds from the perspective of the loss function. SSC encourages the segmentation of text or non-text in the candidate boxes to be similar to the corresponding ground truth, which is helpful to localize more accurate boundaries for arbitrary-shape scene texts. Extensive experiments have demonstrated the effectiveness of the proposed techniques, and state-of-the-art performances are achieved over public arbitrary-shape scene text benchmarks (e.g.,CTW1500,Total-TextandArT).
Pengwen Dai, Yang Li 0093, Hua Zhang 0008, Jingzhi Li 0002, Xiaochun Cao
IEEE Trans. Multim.4
2021 Identity-Preserving Face Anonymization via Adaptively Facial Attributes Obfuscation
abstract
With the popularity of using computer vision technology in monitoring system, there is an increasing societal concern on intruding people's privacy as the captured images/videos may contain identity-related information e.g. people's face. Existing methods on protecting such privacy focus on removing the identity-related information from faces. However, this would weaken the utility of current monitoring system. In this paper, we develop a face anonymization framework that could obfuscate visual appearance while preserving the identity discriminability. The framework is composed of two parts: an identity-aware region discovery module and an identity-aware face confusion module. The former adaptively locates the identity-independent attributes on human faces, and the latter generates the privacy-preserving faces using original faces and discovered facial attributes. To optimize the face generator, we employ a multi-task based loss function, which consists of discriminator loss, identify preserving loss, and reconstruction loss functions. Our model can achieve a balance between recognition utility and appearance anonymizing by modifying different numbers of facial attributes according to pratical demands, and provide a variety of results. Extensive experiments conducted on two public benchmarks Celeb-A and VGG-Face2 demonstrate the effectiveness of our model under distinct face recognition scenarios.
Jingzhi Li 0002, Lutong Han, Ruoyu Chen 0001, Hua Zhang 0008, Lili Wang 0006, Xiaochun Cao
ACM Multimedia1
2020 Learning Disentangled Representations for Identity Preserving Surveillance Face Camouflage
abstract
In this paper, we focus on protecting the facial privacy for people under the surveillance scenarios, by changing some visual appearances of the faces while keeping them recognizable by the current face recognition systems. This is a challenging problem because we need to retain the most important structures of the captured facial images, while modify the salient facial regions to protect personal privacy. To address this problem, we introduce a novel individual face protection model, which can camouflage the face appearance from the perspective of human visual perception and preserve the identity features of faces used for face authentication. To that end, we develop an encoder-decoder network architecture which can separately disentangle the facial feature representation into an appearance code and an identification code. Specifically, we first randomly divide the input face image into two groups, the source and target sets, where the identity and appearance codes can be correspondingly extracted. Then, we recombine the identity and appearance codes to synthesize a new face, which has the same identity as the source subject. Finally, the synthesized faces are employed to replace the original face to protect the individual privacy. Note that our model is end-to-end with a multi-task loss function, which can better preserve the identity and stabilize the training process. Experiments conducted on Cross-Age Celebrity dataset demonstrate the effectiveness of our model and validate our superiority in terms of visual quality and scalability.
Jingzhi Li 0002, Lutong Han, Hua Zhang 0008, Xiaoguang Han 0001, Jingguo Ge, Xiaochun Cao
ICPR1