EDBT 2026 Demo / reviewers in the wild / expert
Guquan Jing
dblp:408/4003
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0007-4709-9932ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Describing-Verifying-Scoring: A Hierarchical Reasoning Framework for Zero-Shot Composed Image RetrievalabstractZero-Shot Composed Image Retrieval (ZS-CIR) aims to identify target images using a composed query of a reference image and modification text without labeled triplets. While recent advances leverage Multimodal Large Language Models (MLLMs) for intent reasoning, they often suffer from hallucination-induced inaccuracies where misaligned descriptions degrade retrieval reliability, and insufficient reasoning due to shallow prompting strategies. To address these challenges, we propose DVSCIR, a novel training-free framework featuring a hierarchical Describing-Verifying-Scoring pipeline with MLLM. Specifically, the Describing stage generates an initial candidate caption, followed by a Verifying stage that rectifies potential hallucinations to ensure description accuracy. The Scoring stage performs a fine-grained re-ranking to identify the optimal match. Within each stage, a hierarchical Chain-of-Thought (CoT) process tailored for ZS-CIR guides the MLLM from low-level perception to deep intentional reasoning via sequential steps within structured sections. This progression ensures robust cross-modal correspondence through a hierarchical refinement of the retrieval process. Extensive experiments across four benchmarks demonstrate that DVSCIR achieves state-of-the-art performance, validating its effectiveness in ZS-CIR. Guquan Jing, Yujian Lee, Hui Zhang 0062 |
ICMR | 1 |
| 2025 | NCL-CIR: Noise-aware Contrastive Learning for Composed Image RetrievalabstractComposed Image Retrieval (CIR) seeks to find a target image using a multi-modal query, which combines an image with modification text to pinpoint the target. While recent CIR methods have shown promise, they mainly focus on exploring relationships between the query pairs (image and text) through data augmentation or model design. These methods often assume perfect alignment between queries and target images, an idealized scenario rarely encountered in practice. In reality, pairs are often partially or completely mismatched due to issues like inaccurate modification texts, low-quality target images, and annotation errors. Ignoring these mismatches leads to numerous False Positive Pair (FFPs) denoted as noise pairs in the dataset, causing the model to overfit and ultimately reducing its performance. To address this problem, we propose the Noise-aware Contrastive Learning for CIR (NCL-CIR), comprising two key components: the Weight Compensation Block (WCB) and the Noise-pair Filter Block (NFB). The WCB coupled with diverse weight maps can ensure more stable token representations of multi-modal queries and target images. Meanwhile, the NFB, in conjunction with the Gaussian Mixture Model (GMM) predicts noise pairs by evaluating loss distributions, and generates soft labels correspondingly, allowing for the design of the soft-label based Noise Contrastive Estimation (NCE) loss function. Consequently, the overall architecture helps to mitigate the influence of mismatched and partially matched samples, with experimental results demonstrating that NCL-CIR achieves exceptional performance on the benchmark datasets. Yujian Lee, Zailong Chen, Yiyang Hu, Guquan Jing |
ICASSP | 7 |
| 2025 | Face Relighting with Ratio Function for Explicit Geometric RepresentationabstractThis paper addresses the problem of face relighting under varying illumination conditions. Lighting is a fundamental element in portrait photography that shapes the mood, geometry, and overall realism of the captured characters. Most previous studies have mainly treated relighting as a 2D generation task without incorporating the geometric features of the characters. In contrast, inspired by ratio image-based methods, this paper proposes to disentangle shadow and brightness variations through geometric information and utilizes generative adversarial networks (GANs) to obtain relighted images with brightness consistency. We design a novel relighting-ratio function that integrates the Cook-Torrance reflectance model to more explicitly represent the face geometry than previous ratio image-based methods. This relighting-ratio function is derived from an image rendering formula that quantizes variables such as albedo that are affected by the lighting direction, while systematically excluding variables such as normal and viewpoint that are not affected by lighting. We conduct quantitative and qualitative experiments on the Multi-PIE and CelebA-HQ datasets and show that the proposed method outperforms existing SOTA methods using lighting directions. Yiyang Hu, Zequn Zhang, Hui Zhang 0062, Guquan Jing |
ICASSP | 4 |
| 2025 | ESTI: An Efficient Spatial-Temporal Interaction Network For Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to identify the target pedestrian from video sequences. However, redundant information exist in input frames. Extracting spatial-temporal features in whole adjacent frames can introduce additional computational overhead. Furthermore, this process leads to the loss of critical spatial and temporal details, causing suboptimal representations. To mitigate these issues, we propose an Efficient Spatial-Temporal Interaction (ESTI) network, which processes half of the input sequence separately through spatial and temporal branches, extracting high-level discriminative features across multiple layers and avoiding redundancy computations. In particular, we propose a Feature Enhancement Module (FEM) for the spatial branch to focus on enhancing spatial dependencies adaptively, and a Temporal Interaction Module (TIM) for temporal branch to capture temporal correlations effectively. Spatial-temporal interaction is performed at the final layer to generate distinctive representations. Extensive experiments on three challenging video Re-ID datasets show that our ESTI achieves competitive results while maintaining low computational complexity. Guquan Jing, Yiyang Hu, Yujian Lee, Hui Zhang 0062 |
ICME | 1 |
| 2025 | Boosting Audio-Visual Segmentation via Triple-Modalities AlignmentabstractThe Audio-Visual Segmentation (AVS) task aims to identify sound-producing objects in the visual domain using auditory cues. Enhancing segmentation efficiency by incorporating prior knowledge, such as object locations and textual prompts, has proven to be crucial. However, existing methods suffer from feature misalignment during model training, leading to ineffective integration and reduced performance. To address this, we propose Triple-modalities alignment (TM-align), which combines audio signals, visual images, and textual prompts. By leveraging prompts from a frozen multi-modal large language model (MLLM), we extract two types of semantic information: contextual semantic description (C.S.D) and prompt specific summary (P.S.S). TM-align yields three pairs of aligned features: visual and C.S.D, visual and P.S.S, visual and audio, within two of our proposed cross-modalities alignment (CMA) models. To further enhance the alignment, we employ Jensen-Shannon Divergence (JSD) to regulate the domain distribution of the latter two features. By effectively aligning the three modalities, TM-align reduces redundancy and improves the overall AVS performance. Experimental results demonstrate that TM-align outperforms the mainstream AVS models.1 Yujian Lee, Zailong Chen, Wentao Fan 0001, Guquan Jing, Yiyang Hu |
ICME | 5 |
| 2025 | Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language ModelsabstractComposed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git. Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai |
ICMR | 7 |
| 2025 | Text-Guided Realistic Single Image Relighting with Wavelet Mamba Diffusion Network
Yunting Lai, Hui Zhang 0062, Yiyang Hu, Guquan Jing |
ICMR | 6 |
| 2025 | 3D-Aided Pedestrian Representation Learning for Video-Based Person Re-IdentificationabstractVideo-based person re-identification (Re-ID) aims to match the target pedestrian from video sequences. Recent methods perform frame-level feature extraction followed by temporal aggregation to obtain video representations. However, they pay insufficient attention to the quality of frame-level features, which suffer from issues including multi-frame misalignment, partial occlusion and appearance confusion. People live in a 3D space. 3D pedestrian representations can provide rich geometric information and shape cues that offer promising solutions to these challenges in video-based Re-ID. To mitigate these issues, this paper proposes a 3D-Aid Pedestrian Representation Learning (3DAPRL) network, which introduces 3D modality to video-based Re-ID. Specifically, two novel modules are designed,i.e., the Cross-Modal Fusion (CMF) module and the Shape-aware Spatial-Temporal Interaction (SSTI) module, to enhance pedestrian representation learning. The CMF module generates discriminative fusion representations by utilizing 3D pedestrian data, while the SSTI module learns spatial-temporal 3D shape representation which are distinguishable for finding the target pedestrian in video scenarios. Both features generated from the CMF and SSTI modules contribute to the final video representation. Extensive experiments on four challenging video-based Re-ID datasets demonstrate that our 3DAPRL network reaches better performance than state-of-the-arts methods. Guquan Jing, Yujian Lee, Yiyang Hu, Hui Zhang 0062 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |