You Qin

dblp:330/7504 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Grounding is All You Need? Dual Temporal Grounding for Video Dialog
abstract
In the realm of video dialog response generation, capturing both the essence of video content and the temporal nuances of conversation history is crucial. While some approaches rely on large-scale pretrained visual-language models, often neglecting temporal dynamics, others emphasize spatial-temporal relationships within videos but demand intricate object trajectory pre-extractions and overlook dialog temporal dynamics. This paper introduces the Dual Temporal Grounding-enhanced Video Dialog model (DTGVD), designed to bridge the gap between these two approaches. DTGVD uniquely integrates the strengths of both by emphasizing dual temporal relationships. It achieves this by predicting dialog turn-specific temporal regions, selectively filtering video content, and grounding responses in both video and dialog contexts.A key innovation of DTGVD is its advanced handling of chronological interplay within dialogs. By effectively capturing and leveraging dependencies between dialog turns, it enables a more nuanced understanding of conversational dynamics. To further align video and dialog temporal dynamics, we introduce a list-wise contrastive learning strategy. In this framework, accurately grounded turn-clip pairings are treated as positive samples, while less precise pairings serve as negative samples. This refined classification is then seamlessly integrated into our end-to-end response generation mechanism.Evaluations using AVSD@DSTC-7 and AVSD@DSTC-8 datasets underscore the superiority of our methodology.
You Qin, Wei Ji 0008, Xinze Lan, Hao Fei 0001, Xun Yang 0001, Dan Guo 0001, Roger Zimmermann, Lizi Liao
IEEE Trans. Multim.1
2025 Secure On-Device Video OOD Detection without Backpropagation
Shawn Li, Peilin Cai, Yuxiao Zhou 0006, Zhiyu Ni, Renjie Liang, You Qin, Yi Nian, Zhengzhong Tu, Xiyang Hu, Yue Zhao 0016
ICCV6
2025 Generalized Video Moment Retrieval
abstract
In this paper, we introduce the Generalized Video Moment Retrieval (GVMR) framework, which extends traditional Video Moment Retrieval (VMR) to handle a wider range of query types. Unlike conventional VMR systems, which are often limited to simple, single-target queries, GVMR accommodates both non-target and multi-target queries. To support this expanded task, we present the NExT-VMR dataset, derived from the YFCC100M collection, featuring diverse query scenarios to enable more robust model evaluation. Additionally, we propose BCANet, a transformer-based model incorporating the novel Boundary-aware Cross Attention (BCA) module. The BCA module enhances boundary detection and uses cross-attention to achieve a comprehensive understanding of video content in relation to queries. BCANet accurately predicts temporal video segments based on natural language descriptions, outperforming traditional models in both accuracy and adaptability. Our results demonstrate the potential of the GVMR framework, the NExT-VMR dataset, and BCANet to advance VMR systems, setting a new standard for future multimedia information retrieval research.
You Qin, Yicong Li 0004, Wei Ji 0008, Li Li 0091, Pengcheng Cai, Lina Wei, Roger Zimmermann
ICLR1
2025 FSGformer: Frequency Separation and Guidance Transformer for Pansharpening
abstract
Pansharpening is a crucial task in remote sensing image processing, aiming to generate high-resolution multispectral (HRMS) images by fusing low-resolution multispectral (LRMS) images with high-resolution panchromatic (PAN) images. However, most current deep learning methods for pansharpening rarely consider the frequency differences in PAN and MS images effectively, resulting in harmful mixing of frequency information and inefficient learning of features. Furthermore, frequency separation-based methods continue to face challenges such as insufficient consideration of the relationship between frequency and spatial information, amplification of noise due to separation, and inadequate learning of frequency information. To address these problems, we propose a novel frequency separation and guidance Transformer, named FSGformer, which focuses on the differences and interactions between high- and low-frequency components. Specifically, we design an adaptive frequency separator tailored for pansharpening to effectively differentiate between distinct frequencies. Subsequently, we develop a carefully designed guidance module that enables the fusion process to benefit from the interaction of frequency information. In addition, we introduce a novel Transformer module that features a joint spatial and spectral attention mechanism and integrate it into a meticulously crafted network architecture to support the effective representation of different frequency information, thereby generating high-quality fused results. Moreover, we incorporate a hybrid frequency separation (HFS) loss to enhance overall performance. Extensive experimental evaluations have confirmed the superiority and generality of our FSGformer. The code is available athttps://github.com/lqrscode/FSGformer.
You Qin, Lanyu Li, Junmin Liu
IEEE Trans. Geosci. Remote. Sens.3
2024 Panoptic Scene Graph Generation with Semantics-Prototype Learning
abstract
Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates lead to biased predicate annotations in the dataset, i.e. different predicates for the same object pairs. Biased predicate annotations make PSG models struggle in constructing a clear decision plane among predicates, which greatly hinders the real application of PSG models. To address the intrinsic bias above, we propose a novel framework named ADTrans to adaptively transfer biased predicate annotations to informative and unified ones. To promise consistency and accuracy during the transfer process, we propose to observe the invariance degree of representations in each predicate class, and learn unbiased prototypes of predicates with different intensities. Meanwhile, we continuously measure the distribution changes between each presentation and its prototype, and constantly screen potentially biased data. Finally, with the unbiased predicate-prototype representation embedding space, biased annotations are easily identified. Experiments show that ADTrans significantly improves the performance of benchmark models, achieving a new state-of-the-art performance, and shows great generalization and effectiveness on multiple datasets. Our code is released at https://github.com/lili0415/PSG-biased-annotation.
Li Li 0091, Wei Ji 0008, Yiming Wu 0005, Mengze Li 0001, You Qin, Lina Wei, Roger Zimmermann
AAAI5
2024 Mrtnet: Multi-Resolution Temporal Network for Video Sentence Grounding
abstract
Video sentence grounding locates a specific moment in a video based on a text query. Existing methods focus on single temporal resolution, ignoring multi-scale temporal consistency. We introduce MRTNet, a multi-resolution grounding network with four key components: a feature encoder, a Multi-Resolution Temporal (MRT) module, a Query-aware Attention (QAM) module, and a predictor. The MRT module uses an encoder-decoder network and Transformers to predict start and end times. The QAM module fuses visual and text features. Both MRT and QAM modules are easily integrated into existing VSG models. We also employ a loss function for cross-modal feature supervision at multiple scales. Extensive experiments on two prevalent datasets have shown the effectiveness of MRTNet.
Wei Ji 0008, You Qin, Long Chen 0016, Yinwei Wei, Yiming Wu 0005, Roger Zimmermann
ICASSP2
2024 Domain-Wise Invariant Learning for Panoptic Scene Graph Generation
abstract
Panoptic Scene Graph Generation (PSG) involves the detection of objects and the prediction of their corresponding relationships (predicates). However, the presence of biased predicate annotations poses a significant challenge for PSG models, as it hinders their ability to establish a clear decision boundary among different predicates. This issue substantially impedes the practical utility and real-world applicability of PSG models. To address the intrinsic bias above, we propose a novel framework to infer potentially biased annotations by measuring the predicate prediction risks within each subject-object pair (domain), and adaptively transfer the biased annotations to consistent ones by learning invariant predicate representation embeddings. Experiments show that our method significantly improves the performance of benchmark models, achieving a new state-of-the-art performance, and shows great generalization and effectiveness on PSG dataset.
Li Li 0091, You Qin, Wei Ji 0008, Yuxiao Zhou 0006, Roger Zimmermann
ICASSP2
2024 Unveiling Causalities in SAR ATR: A Causal Interventional Approach for Limited Data
abstract
Synthetic aperture radar automatic target recognition (SAR ATR) methods often struggle due to inadequate training data. In this letter, we introduce a causal interventional ATR method (CIATR), specifically designed to address the challenges posed by limited synthetic aperture radar (SAR) data. This approach is key in revealing the underlying causal relationships among essential factors in ATR, enabling us to achieve the desired causal effect without altering the imaging conditions (ICs). To address the challenges in SAR ATR with limited data, we developed a structural causal model (SCM) based on causal inference principles. This model helps identify how ICs, as confounders, induce spurious correlations between SAR images and their classifications, which can be solved by standard backdoor adjustment. Our implementation of backdoor adjustment begins with data augmentation, employing a spatial-frequency domain hybrid transformation. This step is crucial in estimating the potential effects of varied ICs on SAR images. Following this, a feature discrimination strategy is introduced to incorporate a hybrid similarity measurement. This technique is essential for assessing and mitigating the impact of changing ICs on the features extracted from SAR images, focusing on both structural and vector angle influences. The CIATR method effectively uncovers the true causal relationships between SAR images and their classes, even with limited data. Tested on MSTAR and OpenSARship datasets, our method shows promising performance in limited data scenarios, achieving 75.05% for ten-way five shots.
You Qin, Siyi Luo, Yulin Huang 0001, Jifang Pei, Jianyu Yang 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Biased-Predicate Annotation Identification via Unbiased Visual Predicate Representation
abstract
Panoptic Scene Graph Generation (PSG) translates visual scenes to structured linguistic descriptions, i.e., mapping visual instances to subjects/objects, and their relationships to predicates. However, the annotators' preferences and semantic overlaps between predicates inevitably lead to the semantic mappings of multiple predicates to one relationship, i.e., biased-predicate annotations. As a result, with the contradictory mapping between visual and linguistics, PSG models are struggled to construct clear decision planes among predicates, so as to cause existing poor performances. Obviously, it is essential for the PSG task to tackle this multi-modal contradiction. Therefore, we propose a novel method that utilizes unbiased visual predicate representations for Biased-Annotation Identification (BAI) as a fundamental step for PSG/SGG tasks. Our BAI includes three main steps: predicate representation extraction, predicate representation debiasing, and biased-annotation identification. With flexible biased annotation processing methods, our BAI can act as a fundamental step of dataset debiasing. Experimental results demonstrate that our proposed BAI has achieved state-of-the-art performance, which promotes the performance of benchmark models to various degrees with ingenious biased annotation processing methods. Furthermore, our BAI shows great generalization and effectiveness on multiple datasets. Our codes are released at https://github.com/lili0415/BAI.
Li Li 0091, You Qin, Wei Ji 0008, Renjie Liang
ACM Multimedia3
2022 HRL2E: Hierarchical Reinforcement Learning with Low-level Ensemble
abstract
Goal-conditioned hierarchical reinforcement learning (HRL) is a promising approach to solve challenging tasks with sparse rewards and long horizons. However, it suffers from the non-stationary problem due to the updating and unstable low level. To stabilize the low level more quickly and accelerate the non-stationary stage, we propose a novel HRL method: Hierarchical Reinforcement Learning with Low-level Ensemble (HRL2E). In HRL2E, the high level generates goals as high-level actions based on current states. Then the low level made up of several homogeneous policies attempts to complete these goals within a specific timestep budget. The improvement of our approach to the general goal-conditioned HRL algorithms can be summarized in two aspects. First, we estimate the target value function with the ensemble, stabilizing the training process. Second, we propose the Gates module composed of several scoring machines to score each low-level policy and judge which one has the most success potential to execute a specific goal. We adopt Twin Delayed Deep Deterministic Policy Gradient (TD3) in each level. Experimental comparison between our method and state-of-the-art goal-conditioned HRL methods on challenging continuous control tasks in MuJoCo domains shows our method can significantly accelerate training.
You Qin, Zhi Wang 0001, Chunlin Chen 0001
IJCNN1