Zhipu Liu

dblp:254/0480 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Debiased Dual-Invariant Defense for Adversarially Robust Person Re-Identification
abstract
Person re-identification (ReID) is a fundamental task in many real-world applications such as pedestrian trajectory tracking. However, advanced deep learning-based ReID models are highly susceptible to adversarial attacks, where imperceptible perturbations to pedestrian images can cause entirely incorrect predictions, posing significant security threats. Although numerous adversarial defense strategies have been proposed for classification tasks, their extension to metric learning tasks such as person ReID remains relatively unexplored. Moreover, the several existing defenses for person ReID fail to address the inherent unique challenges of adversarially robust ReID. In this paper, we systematically identify the challenges of adversarial defense in person ReID into two key issues: model bias and composite generalization requirements. To address them, we propose a debiased dual-invariant defense framework composed of two main phases. In the data balancing phase, we mitigate model bias using a diffusion-model-based data resampling strategy that promotes fairness and diversity in training data. In the bi-adversarial self-meta defense phase, we introduce a novel metric adversarial training approach incorporating farthest negative extension softening to overcome the robustness degradation caused by the absence of classifier. Additionally, we introduce an adversarially-enhanced self-meta mechanism to achieve dual-generalization for both unseen identities and unseen attack types. Experiments demonstrate that our method significantly outperforms existing state-of-the-art defenses.
Yanxiang Zhao, Zhongyun Hua, Zhipu Liu, Zhaoquan Gu, Qing Liao 0001, Leo Yu Zhang
AAAI4
2025 Degradation-Consistent Learning via Bidirectional Diffusion for Low-Light Image Enhancement
abstract
Low-light image enhancement aims to improve the visibility of degraded images to better align with human visual perception. While diffusion-based methods have shown promising performance due to their strong generative capabilities. However, their unidirectional modelling of degradation often struggles to capture the complexity of real-world degradation patterns, leading to structural inconsistencies and pixel misalignments. To address these challenges, we propose a bidirectional diffusion optimization mechanism that jointly models the degradation processes of both low-light and normal-light images, enabling more precise degradation parameter matching and enhancing generation quality. Specifically, we perform bidirectional diffusion-from low-to-normal light and from normal-to-low light during training and introduce an adaptive feature interaction block (AFI) to refine feature representation. By leveraging the complementarity between these two paths, our approach imposes an implicit symmetry constraint on illumination attenuation and noise distribution, facilitating consistent degradation learning and improving the model's ability to perceive illumination and detail degradation. Additionally, we design a reflection-aware correction module (RACM) to guide color restoration post-denoising and suppress overexposed regions, ensuring content consistency and generating high-quality images that align with human visual perception. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art methods in both quantitative and qualitative evaluations while generalizing effectively to diverse degradation scenarios.Code
Jinhong He, Minglong Xue, Zhipu Liu, Mingliang Zhou 0001, Aoxiang Ning, Palaiahnakote Shivakumara
ACM Multimedia3
2025 Multi-Model Synergy Perception for Open-World Person Re-Identification
abstract
Open-world person re-identification aims to train a model on source doamins and generalize well on unseen domains. Existing domain generalizable person re-identification methods primarily employ the equality training paradigm to train the model on multi-source domains. However, in open-world scenarios, domain imbalance often causes domain bias issue that leads to sub-optimal generalization ability, which is seriously overlooked. In this paper, we propose a Multi-model Synergy Perception (MSP) framework equipped with an Asynchronous Training Paradigm (ATP) on biased domains to maintain the domain balance for exploring the domain-invariant features. With the philosophy of divide and conquer, we divide the biased source domains into multiple debiased sub-source domains and employ a multi-network architecture to learn these sub-source domains in parallel. Additionally, to better generalize knowledge across these sub-source domains, we propose a Structure Synergy Perception (SSP) module that constructs the feature relationship distribution for each sub-domain and aligns them to map the unique knowledge to each other. Furthermore, considering the consistency of sub-source domains, we further propose a Synergy Distillation Perception (SDP) to improve the model both semantic and domain generalization ability. The main idea of SDP is to use the center guided soft label and the part based triplet graph to distill each submodel, which can facilitate the network to explore domain-invariant representations of images. Extensive experiments demonstrate that our method outperforms state-of-the-arts for open-domain person ReID.
Zhipu Liu, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.1
2025 Exploring Invariance Matters for Domain Generalization
abstract
Domain generalization (DG) aims to solve the problem of significant performance degradation when target domain data collected from the Out-Of-Distribution (O.O.D). Previous efforts try to exploit invariant features in the source domain through CNN networks. However, inspired by causal mechanisms, we find that the complex spurious-invariant information is still hidden in this view invariant features, and the impact of domain and class discrepancies on extracting invariance has not been effectively mitigated. To alleviate these issues, we propose a self-weighted multi-view mining invariance domain generalization framework (SMIDG). On the one hand, to make up for the insufficiency of traditional single-view convolutional feature extraction networks, we propose to mine features from another frequency view and use the self-adaptive adversarial masks to eliminate some spurious correlations, ensuring causal invariance in the coarse-grained generalization. However, due to inconsistencies in discriminative information between inter-domain and intra-domain samples, as well as inter-class and intra-class samples, the coarse-grained elimination of spurious associations does not fully resolve this issue. On the other hand, we also consider the fine-grained generalization from two aspects. Firstly, to better tackle the domain discrepancies, we propose a novel progressive contrastive learning strategy that learns the underlying specific features of samples while gradually mitigating domain discrepancies, thereby ensuring domain invariance in fine-grained generalization. Secondly, due to the issue of feature inconsistency, we adopt a self-adaptive hard sample mining method with information gain to ensure that the model pays more attention on hard disentangled samples, thus maintaining feature invariance. Extensive experiments on five benchmark datasets demonstrate that our method outperforms state-of-the-art approaches. Our code is available at https://github.com/bihhm/SMIDG.
Shanshan Wang 0008, Houmeng He, Xun Yang 0001, Zhipu Liu, Yuanhong Zhong, Xingyi Zhang 0001, Meng Wang 0001
IEEE Trans. Image Process.4
2023 Neural Image Parts Group Search for Person Re-Identification
abstract
Employing partition strategy to explore fine-grained features has been verified to be beneficial for person re-identification in recent literature. However, existing methods primarily rely on expert experience to manually design various partition strategies, which may lead to a sub-optimal solution for fine-grained features exploration. In this paper, we propose a Neural Parts Group Search (NPGS) strategy that auto-searches the optimal parts group via evolutionary algorithm (EA) to facilitate the network to exploit the local details. And during search process, designing a high-quality search space is especially crucial for an efficient optimization. Considering the human top-down structure and the semantic coherence of parts, we design a coarse-to-fine parts search space (C2F-PSP) in NPGS, which effectively reduce the search complexity without the loss of parts expressivity. Additionally, since only employing the high-level semantic features is insufficient for the NPGS to search effective parts, we further develop an efficient feature aggregation strategy named hierarchical low-rank bilinear pooling that progressively integrates the high-level semantic property and the low-level fine-grained details to facilitate the NPGS to explore the fine-grained features. Furthermore, to relieve the interference of background during parts search process, we propose a novel Relational Attention Module (RAM) by exploiting the channel and spatial structural interdependence of pixels to strengthen the discriminative regions. Extensive experiments on the mainstream evaluation datasets demonstrate that our method outperforms the recent state-of-the-art Re-ID models.
Zhipu Liu, Lei Zhang 0038, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Style Uncertainty Based Self-Paced Meta Learning for Generalizable Person Re-Identification
abstract
Domain generalizable person re-identification (DG ReID) is a challenging problem, because the trained model is often not generalizable to unseen target domains with different distribution from the source training domains. Data augmentation has been verified to be beneficial for better exploiting the source data to improve the model generalization. However, existing approaches primarily rely on pixel-level image generation that requires designing and training an extra generation network, which is extremely complex and provides limited diversity of augmented data. In this paper, we propose a simple yet effective feature based augmentation technique, named Style-uncertainty Augmentation (SuA). The main idea of SuA is to randomize the style of training data by perturbing the instance style with Gaussian noise during training process to increase the training domain diversity. And to better generalize knowledge across these augmented domains, we propose a progressive learning to learn strategy named Self-paced Meta Learning (SpML) that extends the conventional one-stage meta learning to multi-stage training process. The rationality is to gradually improve the model generalization ability to unseen target domains by simulating the mechanism of human learning. Furthermore, conventional person Re-ID loss functions are unable to leverage the valuable domain information to improve the model generalization. So we further propose a distance-graph alignment loss that aligns the feature relationship distribution among domains to facilitate the network to explore domain-invariant representations of images. Extensive experiments on four large-scale benchmarks demonstrate that our SuA-SpML achieves state-of-the-art generalization to unseen domains for person ReID.
Lei Zhang 0038, Zhipu Liu, Wensheng Zhang 0002, David Zhang 0001
IEEE Trans. Image Process.2
2020 Hierarchical Bi-Directional Feature Perception Network for Person Re-Identification
abstract
Previous Person Re-Identification (Re-ID) models aim to focus on the most discriminative region of an image, while its performance may be compromised when that region is missing caused by camera viewpoint changes or occlusion. To solve this issue, we propose a novel model named Hierarchical Bi-directional Feature Perception Network (HBFP-Net) to correlate multi-level information and reinforce each other. First, the correlation maps of cross-level feature-pairs are modeled via low-rank bilinear pooling. Then, based on the correlation maps, Bi-directional Feature Perception (BFP) module is employed to enrich the attention regions of high-level feature, and to learn abstract and specific information in low-level feature. And then, we propose a novel end-to-end hierarchical network which integrates multi-level augmented features and inputs the augmented low- and middle-level features to following layers to retrain a new powerful network. What's more, we propose a novel trainable generalized pooling, which can dynamically select any value of all locations in feature maps to be activated. Extensive experiments implemented on the mainstream evaluation datasets including Market-1501, CUHK03 and DukeMTMC-ReID show that our method outperforms the recent SOTA Re-ID models.
Zhipu Liu, Lei Zhang 0038, Yang Yang 0002
ACM Multimedia1