EDBT 2026 Demo / reviewers in the wild / expert
Jingzi Gu
dblp:138/8122
· DBLP profile ↗
14ranked-venue papers
1as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KA-CDRE: Knowledge-Augmented Cross-Document Relation Extraction
Peize Li, Jingzi Gu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005 |
ADMA (4) | 3 |
| 2025 | Corer: Concept Residue Erasing in Text-to-Image Diffusion ModelsabstractThe remarkable development of text-to-image generation models has raised notable security concerns, such as the infringement of portrait rights and the generation of inappropriate content. Concept erasure has been proposed to remove the model’s knowledge about protected or inappropriate concepts. Although many methods have tried to balance the efficacy (erasing target concepts) and specificity (retaining irrelevant concepts), they can still generate abundant erasure concepts under the steering of semantically related inputs. In this work, we propose Corer to address this "concept residue" issue. Specifically, we first introduce the mechanism of neighbor-concept mining to dig out the associated concepts and expand the erasing range. Furthermore, to mitigate the negative impact on the generation of irrelevant concepts caused by the expansion of erasure scope, Corer preserves the specificity through the beyond-concept regularization. We also employ the closed-form solution to optimize weights of U-Net, as well as the prediction noise alignment with the LoRA module. Extensive experiments on multiple benchmarks demonstrate that Corer outperforms previous concept-erasing methods in terms of superior erasing efficacy, specificity, and generality. Yufan Liu 0002, Jinyang An, Huashan Chen, Wanqian Zhang, Dayan Wu, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ICME | 7 |
| 2024 | A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- IdentificationabstractExtensive advancements have been made in person ReID through the mining of semantic information. Nevertheless, existing methods that utilize semantic-parts from a single image modality do not explicitly achieve this goal. Whiteness the impressive capabilities in multimodal understanding of Vision Language Foundation Model CLIP, a recent two-stage CLIP-based method employs automated prompt engineering to obtain specific textual labels for classifying pedestrians. However, we note that the predefined soft prompts may be inadequate in expressing the entire visual context and struggle to generalize to unseen classes. This paper presents an end-to-end Prompt-driven Semantic Guidance (PromptSG) framework that harnesses the rich semantics inherent in CLIP. Specifically, we guide the model to attend to regions that are semantically faithful to the prompt. To provide personalized language descriptions for specific individuals, we propose learning pseudo tokens that represent specific visual contexts. This design not only facilitates learning fine-grained attribute information but also can inherently leverage language prompts during inference. Without requiring additional labeling efforts, our PromptSG achieves state-of-the-art by over 10% on MSMTI7 and nearly 5% on the Market-I50I benchmark. The codes will be available at h t tps: / / gi th ub. com/ YzXian16/PromptSG Zexian Yang, Dayan Wu, Chenming Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
CVPR | 5 |
| 2024 | Prediction Exposes Your Face: Black-Box Model Inversion via Prediction Alignment
Yufan Liu 0002, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ECCV (36) | 5 |
| 2024 | SD4Privacy: Exploiting Stable Diffusion for Protecting Facial PrivacyabstractRecently, adversarial examples are introduced to protect personal images from being identified by unauthorized face recognition systems. Existing approaches follow the transfer-based adversarial attack paradigm, where local surrogate models are utilized to generate protected images. However, these surrogate models can neither be necessary nor efficient for generating adversarial examples. In this paper, we propose SD4Privacy, i.e., Stable Diffusion for Privacy, which exploits the latent space of Stable Diffusion Model to synthesize adversarial examples. First, we learn an optimal textual embedding of target image to preserve its representative semantics, directly guiding the sampling process of synthesized image. Then, we utilize the encoder of UNet in Stable Diffusion as the substitution of surrogate classification models, which enables the efficient adversarial guidance by semantic h-space of UNet for adversarial example generation. Experiments show the state-of-the-art protection performance, as well as high-quality protected images with visual naturalness and imperceptible perturbations. Jinyang An, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ICME | 5 |
| 2024 | Exploiting Vision-Language Model for Visible-Infrared Person Re-identification via Textual Modality AlignmentabstractVisible-Infrared Person Re-identification (VI-ReID) aims at matching the images of specific person captured by different modality cameras. Previous methods introduce a synthesized auxiliary modality to relieve the modality discrepancy. However, they directly fuse the raw pixels of visible and infrared images, ignoring the high-level semantic patterns. Additionally, the huge modality gap can’t be bridged up closely, which leads to the oscillations in the feature space. Thus, in this paper, we propose a novel Textual Modality Alignment Learning method, named TMAL, which tackles these two issues in a unified two-stage framework. Specifically, we first exploit the semantic alignment in CLIP model through learnable text tokens, which are then encoded to form semantic representations of each identity. In the second stage, we propose the Modality Alignment Module, empowering the image encoder with modality-shared and modality-specific features. We also introduce the Identity Enhancement module (IEM) to extract more informative modality-specific features. Experiments on two benchmarks demonstrate the efficacy of our method. Bingyu Duan, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ICME | 5 |
| 2024 | Privacy-Preserving Replay and Adaptive Relation Distillation for Camera Incremental Person Re-IdentificationabstractTraditional person re-identification (ReID) methods trained on static data are ill-suited to real-world dynamic surveillance systems. Recently, a more desirable setting "Camera Incremental Person ReID (CIPR)", has been proposed to continually adapt to new cameras and accumulate knowledge. However, prior work on relation distillation heavily constrains intra-class relations for all identities, while under-exploring the credibility of different identities in knowledge transfer. Besides, their rehearsal-free setting sidesteps privacy concerns but compromises performance. In this paper, we present a novel framework, P2-ARD, designed specifically for CIPR. Firstly, we propose an innovative Adaptive Relation Distillation loss that automatically selects more crucial identities for distillation. Additionally, we introduce the privacy-preserving replay scheme to effectively retain semantic information while ensuring the privacy of the identity. Finally, we incorporate a cycle-consistent correlation method to address the class overlap issue in CIPR. Extensive experiments demonstrate our method outperforming the state-of-the-art. Zexian Yang, Dayan Wu, Wanqian Zhang, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ICME | 4 |
| 2024 | Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based CustomizationabstractWith the development of diffusion-based customization methods like DreamBooth, individuals now have access to train the models that can generate their personalized images. Despite the convenience, malicious users have misused these techniques to create fake images, thereby triggering a privacy security crisis. In light of this, proactive adversarial attacks are proposed to protect users against customization. The adversarial examples are trained to distort the customization model's outputs and thus block the misuse. In this paper, we propose DisDiff (Disrupting Diffusion), a novel adversarial attack method to disrupt the diffusion model outputs. We first delve into the intrinsic image-text relationships, well-known as cross-attention, and empirically find that the subject-identifier token plays an important role in guiding image generation. Thus, we propose the Cross-Attention Erasure module to explicitly "erase" the indicated attention maps and disrupt the text guidance. Besides, we analyze the influence of the sampling process of the diffusion model on Projected Gradient Descent (PGD) attack and introduce a novel Merit Sampling Scheduler to adaptively modulate the perturbation updating amplitude in a step-aware manner. Our DisDiff outperforms the state-of-the-art methods by 12.75% of FDFR scores and 7.25% of ISM scores across two facial benchmarks and two commonly used prompts on average. Yisu Liu, Jinyang An, Wanqian Zhang, Dayan Wu, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ACM Multimedia | 5 |
| 2024 | Feature Refinement and Calibration for Continual Visual Search
Qinghang Su, Xiaohua Chen 0002, Jingzi Gu, Bo Li 0063 |
PRCV (9) | 3 |
| 2022 | Deep Piecewise Hashing for Efficient Hamming Space RetrievalabstractHamming space retrieval can achieve constant-time image search, which is more efficient than linear scan. In Hamming space retrieval, the data points inside the Hamming ball imply retrievable while the data points outside are irretrievable. Therefore, it is crucial to explicitly characterize the Hamming ball. However, for the existing Hamming space retrieval methods, many similar points are found close to the outside of the Hamming ball while many dissimilar points are found close to the query point, leading to the decline of both retrieval accuracy and recall. In this paper, we present a novel method named Deep Piecewise Hashing (DPH), for Efficient Hamming Space Retrieval. A piecewise loss is elaborately designed to guide the learning of hash codes. Meanwhile, a piecewise probability distribution is introduced in the proposed loss function. The piecewise probability distribution pays more attention to the learning of those "marginal" similar points. It considers both discrimination and robustness for the dissimilar points inside the Hamming ball. Comprehensive experiments on two datasets, MS-COCO and NUS-WIDE, demonstrate that DPH can yield state-of-the-art Hamming space retrieval performance. Jingzi Gu, Dayan Wu, Peng Fu 0008, Bo Li 0063, Weiping Wang 0005 |
ICASSP | 1 |
| 2019 | Adversary Guided Asymmetric Hashing for Cross-Modal RetrievalabstractCross-modal hashing has attracted considerable attention for large-scale multimodal retrieval task. A majority of hashing methods have been proposed for cross-modal retrieval. However, these methods inadequately focus on feature learning process and cannot fully preserve higher-ranking correlation of various item pairs as well as the multi-label semantics of each item, so that the quality of binary codes may be downgraded. To tackle these problems, in this paper, we propose a novel deep cross-modal hashing method, called Adversary Guided Asymmetric Hashing (AGAH). Specifically, it employs an adversarial learning guided multi-label attention module to enhance the feature learning part which can learn discriminative feature representations and keep the cross-modal invariability. Furthermore, in order to generate hash codes which can fully preserve the multi-label semantics of all items, we propose an asymmetric hashing method which utilizes a multi-label binary code map that can equip the hash codes with multi-label semantic information. In addition, to ensure higher-ranking correlation of all similar item pairs than those of dissimilar ones, we adopt a new triplet-margin constraint and a cosine quantization technique for Hamming space similarity preservation. Extensive empirical studies show that AGAH outperforms several state-of-the-art methods for cross-modal retrieval. Wen Gu, Xiaoyan Gu 0001, Jingzi Gu, Bo Li 0063, Weiping Wang 0005 |
ICMR | 3 |
| 2018 | Efficient Algorithms of Parallel Skyline Join over Data Streams
Jinchao Zhang 0002, Jingzi Gu, Shuai Cheng 0002, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ICA3PP (1) | 2 |
| 2015 | LuBase: A Search-Efficient Hybrid Storage System for Massive Text Data
Debin Jia, Zhengwei Liu, Xiaoyan Gu 0001, Bo Li 0063, Jingzi Gu, Weiping Wang 0005, Dan Meng 0002 |
ICA3PP (2) | 5 |
| 2012 | Medical Image Retrieval Method Based on Relevance Feedback
Haiwei Pan, Qilong Han, Jingzi Gu, Pengyuan Li 0001 |
ADMA | 4 |