VLDB 2026 Research / reviewers in the wild / expert
Jie Qin 0004
dblp:30/5781-4
· DBLP profile ↗
7ranked-venue papers in the field
0as first author
6since 2021 · last 2026
0000-0002-0306-534XORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5Database Systems & Data Management · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-view Driver Gaze Estimation via Mutual Enhancement
Chengzheng Fu, Rong Quan, Siyu Chen 0004, Yiming Ni, Jie Qin 0004 |
ICMR | 5 |
| 2026 | PRISM: Preference-Guided Semantic Reasoning with Vision-Language Models for Object Goal NavigationabstractObject Goal Navigation (ObjectNav) requires an agent to locate a target object in unseen environments based on partial visual observations. A central challenge lies in reasoning about where the target is likely to appear under semantic uncertainty, sparse observations, and ambiguous scene cues. Existing approaches often rely on extensive training or static semantic priors, limiting adaptability and interpretability in unfamiliar environments. We introduce PRISM, a preference-guided semantic reasoning and mapping framework that integrates vision-language models (VLMs) into the navigation loop to support efficient semantic object search. PRISM constructs a dynamic Preference-guided Semantic Map (PSM) that aggregates observed semantics, predicted unobserved regions, and language-mediated semantic reasoning, enabling the agent to incrementally refine its belief over likely target locations. Based on the PSM, PRISM further employs Adaptive Preference Exploration Strategies (APES) to adjust exploration behavior using contextual semantic cues and historical feedback under semantic ambiguity. Experiments on the Gibson, Matterport3D (MP3D), and Habitat-Matterport3D (HM3D) datasets show that our method achieves consistent improvements over prior training-free approaches, yielding absolute gains of +2.7% SR on Gibson, +2.8% SR on MP3D, and +3.2% SR on HM3D. Cong Pan 0001, Chengjie Fan, Wanjie Cai, Xichen Ding, Jie Qin 0004 |
ICMR | 6 |
| 2026 | Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
Rong Quan, Yantao Lai, Dong Liang 0008, Jie Qin 0004 |
ICMR | 4 |
| 2025 | DCN: Decoupled-Coupled Network for Text-based Person SearchabstractText-based person search aims to identify a person based on textual descriptions, by simultaneously addressing person detection and cross-modal alignment between text queries and person images. Existing approaches often struggle with conflicts in exploiting proposals across these two sub-tasks. Specifically, cross-modal alignment requires highly precise proposals, while person detection can tolerate a certain degree of proposal inaccuracy but always needs a large number of proposals. In this paper, we propose the Decoupled-Coupled Network (DCN) to tackle the above conflicts. We first attempt to resolve the above conflicts by proposing a Decoupled Proposal Selection (DPS) strategy, inspired by the divide-and-conquer principle. DPS adaptively selects the most suitable proposals for each sub-task, ensuring their distinct requirements are adequately met. We further present a Coupled Cascade Refinement (CCR) module to jointly optimize both sub-tasks in a multi-stage manner, progressively improving detection accuracy and fostering cross-modal alignment between text and person image. In addition, we introduce two types of objective functions to optimize the inherently multi-positive contrastive learning challenge. Extensive experiments conducted on two benchmarks demonstrate the effectiveness and superiority of our DCN over existing competitors. Rong Quan, Liangxu Su, Wentong Li 0001, Yichao Yan, Jie Qin 0004 |
MMAsia | 7 |
| 2025 | Large Models are Good Annotators for Zero-Shot LearningabstractHuman-annotated attributes serve as effective semantic label embeddings for zero-shot learning (ZSL); however, their annotation is labor-intensive and difficult to scale. Recent studies have explored weakly supervised semantic label embeddings to reduce human effort, but these methods often fail to capture visual similarity and underperform compared to human-annotated semantics. In this work, we propose a minimally supervised yet effective approach: GPT- and CLIP-powered attributes (GCAtt). Specifically, we introduce a three-step interaction process with ChatGPT-comprising preliminary design, hierarchical refinement, and specific value determination-to generate attributes that are both category-shared and discriminative for classification. Additionally, we develop a method that encodes attributes and their values as potential text pairings, leveraging CLIP's retrieval capabilities for annotation. Experimental results on four widely used benchmarks demonstrate that GCAtt consistently outperforms human-annotated semantics. Code and data are available at https://github.com/RowenaHe/GCAtt. Qingzhi He, Wentong Li 0001, Shengcai Liao, Rong Quan, Tong Cui, Jie Qin 0004 |
SIGIR | 7 |
| 2024 | MACA: Memory-aided Coarse-to-fine Alignment for Text-based Person SearchabstractText-based person search (TBPS) aims to search for the target person in the full image through textual descriptions. The key to addressing this task is to effectively perform cross-modality alignment between text and images. In this paper, we propose a novel TBPS framework, named Memory-Aided Coarse-to-fine Alignment (MACA), to learn an accurate and reliable alignment between the two modalities. Firstly, we introduce a proposal-based alignment module, which performs contrastive learning to accurately align the textual modality with different pedestrian proposals at a coarse-grained level. Secondly, for the fine-grained alignment, we propose an attribute-based alignment module to mitigate unreliable features by aligning text-wise details with image-wise global features. Moreover, we introduce an intuitive memory bank strategy to supplement useful negative samples for more effective contrastive learning, improving the convergence and generalization ability of the model based on the learned discriminative features. Extensive experiments on CUHK-SYSU-TBPS and PRW-TBPS demonstrate the superiority of MACA over state-of-the-art approaches. The code is available at https://github.com/suliangxu/MACA. Liangxu Su, Rong Quan, Jie Qin 0004 |
SIGIR | 4 |
| 2019 | Unsupervised Nonnegative Adaptive Feature Extraction for Data RepresentationabstractIn this paper, we propose a novel unsupervised Nonnegative Adaptive Feature Extraction (NAFE) algorithm for data representation and classification. The formulation of NAFE integrates the sparsity constrained nonnegative matrix factorization (NMF), representation learning, and adaptive reconstruction weight learning into a unified model. Specifically, NAFE performs feature and weight learning over the new robust representations of NMF for more accurate measure and representation. For nonnegative adaptive feature extraction, our NAFE first utilizes the sparsity constrained NMF to obtain the new and robust representations of the original data. To preserve the manifold structures of the learnt new representations, we also incorporate a neighborhood reconstruction error over the weight matrix for joint minimization. Note that to further improve the representation power, the weights are jointly shared in the new low-dimensional nonnegative representation space, low-dimensional nonlinear manifold space, and low-dimensional projective subspace, i.e., local neighborhood information is clearly preserved in different feature spaces so that informative representations and features can be jointly obtained. To enable NAFE to extract features from new data, we also include a feature approximation error by a linear projection so that the learnt extractor can obtain features from new data efficiently. Extensive simulations show that our formulation can deliver state-of-the-art results on several public databases for feature extraction and classification, compared with several related methods. Yan Zhang 0053, Zhao Zhang 0001, Sheng Li 0001, Jie Qin 0004, Guangcan Liu, Meng Wang 0001, Shuicheng Yan |
IEEE Trans. Knowl. Data Eng. | 4 |