Hongyu Qu

dblp:221/0114 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Spatio-Temporal Decoupled Knowledge Compensator for Few-Shot Action Recognition
abstract
Few-Shot Action Recognition (FSAR) is a challenging task that requires recognizing novel action categories with a few labeled videos. Recent works typically apply semantically coarse category names as auxiliary contexts to guide the learning of discriminative visual features. However, such context provided by the action names is too limited to provide sufficient background knowledge for capturing novel spatial and temporal concepts in actions. In this paper, we propose DiST, an innovative Decomposition-incorporation framework for FSAR that makes use of decoupled Spatial and Temporal knowledge provided by large language models to learn expressive multi-granularity prototypes. In the decomposition stage, we decouple vanilla action names into diverse spatio-temporal attribute descriptions (action-related knowledge). Such commonsense knowledge complements semantic contexts from spatial and temporal perspectives. In the incorporation stage, we propose Spatial/Temporal Knowledge Compensators (SKC/TKC) to discover discriminative object-level and frame-level prototypes, respectively. In SKC, object-level prototypes adaptively aggregate important patch tokens under the guidance of spatial knowledge. Moreover, in TKC, frame-level prototypes utilize temporal attributes to assist in inter-frame temporal relation modeling. These learned prototypes thus provide transparency in capturing fine-grained spatial details and diverse temporal patterns. Experimental results show DiST achieves state-of-the-art results on five standard FSAR datasets.
Hongyu Qu, Xiangbo Shu, Rui Yan 0010, Hailiang Gao, Wenguan Wang, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Locality-Aware Cross-Modal Correspondence Learning for Dense Audio-Visual Events Detection
abstract
Dense-localization Audio-Visual Events (DAVE) aims to identify time boundaries and corresponding categories for events that are both audible and visible in a long video, where events may co-occur and exhibit varying durations. However, complex audio-visual scenes often involve asynchronization between modalities, making accurate localization challenging. Existing DAVE solutions extract audio and visual features through unimodal encoders, and fuse them via dense cross-modal interaction. However, independent unimodal encodingstruggles to emphasize shared semantics between modalitieswithout cross-modal guidance, while dense cross-modal attention mayover-attend to semantically unrelated audio-visual features. To address these problems, we present LOCO, a Locality-aware cross-modal Correspondence learning framework for DAVE. LOCO leverages the local temporal continuity of audio-visual events as important guidance to filter irrelevant cross-modal signals and enhance cross-modal alignment throughout both unimodal and cross-modal encoding stages. i) Specifically, LOCO applies Local Correspondence Feature (LCF) Modulation to enforce unimodal encoders to focus on modality-shared semantics by modulating agreement between audio and visual features based on local cross-modal coherence. ii) To better aggregate cross-modal relevant features, we further customize Local Adaptive Cross-modal (LAC) Interaction, which dynamically adjusts attention regions in a data-driven manner. This adaptive mechanism focuses attention on local event boundaries and accommodates varying event durations. By incorporating LCF and LAC, LOCO provides solid performance gains and outperforms existing DAVE methods. The source code will be released.
Ling Xing 0003, Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Learning Clustering-based Prototypes for Compositional Zero-Shot Learning
abstract
Learning primitive (i.e., attribute and object) concepts from seen compositions is the primary challenge of Compositional Zero-Shot Learning (CZSL). Existing CZSL solutions typically rely on oversimplified data assumptions, e.g., modeling each primitive with a single centroid primitive presentation, ignoring the natural diversities of the attribute (resp. object) when coupled with different objects (resp. attribute). In this work, we develop ClusPro, a robust clustering-based prototype mining framework for CZSL that defines the conceptual boundaries of primitives through a set of diversified prototypes. Specifically, ClusPro conducts within-primitive clustering on the embedding space for automatically discovering and dynamically updating prototypes. To learn high-quality embeddings for discriminative prototype construction, ClusPro repaints a well-structured and independent primitive embedding space, ensuring intra-primitive separation and inter-primitive decorrelation through prototype-based contrastive learning and decorrelation learning. Moreover, ClusPro effectively performs prototype clustering in a non-parametric fashion without the introduction of additional learnable parameters or computational budget during testing. Experiments on three benchmarks demonstrate ClusPro outperforms various top-leading CZSL solutions under both closed-world and open-world settings. Our code is available at CLUSPRO.
Hongyu Qu, Jianan Wei, Xiangbo Shu, Wenguan Wang
ICLR1
2025 Reliable and Diverse Hierarchical Adapter for Zero-shot Video Classification
abstract
Adapting pre-trained vision-language models to downstream tasks has emerged as a novel paradigm for zero-shot learning. Existing test-time adaptation (TTA) methods such as TPT attempt to fine-tune visual or textual representations to accommodate downstream tasks but still require expensive optimization costs. To this end, Training-free Dynamic Adapter (TDA) maintains a cache containing visual features for each category in a parameter-free manner and measures sample confidence based on prediction entropy of test samples. Inspired by TDA, this work aims to develop the first training-free adapter for zero-shot video classification. Capturing the intrinsic temporal relationships within video data to construct and maintain the video cache is key to extending TDA to the video domain. In this work, we propose a reliable and diverse Hierarchical Adapter for zero-shot video classification, which consists of Frame-level Cache Refiner and Video-level Cache Updater. Before each video sample enters the corresponding cache, it needs to be refined at frame level based on prediction entropy and temporal probability difference. Due to the limited capacity of the cache, we update the cache during inference based on the principle of diversity. Experiments on four popular video classification benchmarks demonstrate the effectiveness of Hierarchical Adapter. The code is available at https://github.com/Gwxer/Hierarchical-Adapter.
Wenxuan Ge, Rui Yan 0010, Hongyu Qu, Guosen Xie, Xiangbo Shu
IJCAI4
2025 TEST-V: TEst-time Support-set Tuning for Zero-shot Video Classification
abstract
Recently, adapting Vision Language Models (VLMs) to zero-shot visual classification by tuning class embedding with a few prompts (Test-time Prompt Tuning, TPT) or replacing class names with generated visual samples (support-set) has shown promising results. However, TPT cannot avoid the semantic gap between modalities while the support-set cannot be tuned. To this end, we draw on each other's strengths and propose a novel framework, namely TEst-time Support-set Tuning for zero-shot Video Classification (TEST-V). It first dilates the support-set with multiple prompts (Multi-prompting Support-set Dilation, MSD) and then erodes the support-set via learnable weights to mine key cues dynamically (Temporal-aware Support-set Erosion, TSE). Specifically, i) MSD expands the support samples for each class based on multiple prompts inquired from LLMs to enrich the diversity of the support-set. ii) TSE tunes the support-set with factorized learnable weights according to the temporal prediction consistency in a self-supervised manner to dig pivotal supporting cues for each class. TEST-V achieves state-of-the-art results across four benchmarks and shows good interpretability.
Rui Yan 0010, Hongyu Qu, Xiaoyu Du 0002, Jinhui Tang 0001, Tieniu Tan
IJCAI3
2025 OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild
abstract
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to $\textbf{i)}$ $\textit{the scarcity of annotated datasets}$, and $\textbf{ii)}$ $\textit{the insufficient diversity of labeled data}$. In this work, we present OmniGaze, a semi-supervised framework for 3D gaze estimation, which utilizes large-scale unlabeled data collected from diverse and unconstrained real-world environments to mitigate domain bias and generalize gaze estimation in the wild. First, we build a diverse collection of unlabeled facial images, varying in facial appearances, background environments, illumination conditions, head poses, and eye occlusions. In order to leverage unlabeled data spanning a broader distribution, OmniGaze adopts a standard pseudo-labeling strategy and devises a reward model to assess the reliability of pseudo labels. Beyond pseudo labels as 3D direction vectors, the reward model also incorporates visual embeddings extracted by an off-the-shelf visual encoder and semantic cues from gaze perspective generated by prompting a Multimodal Large Language Model to compute confidence scores. Then, these scores are utilized to select high-quality pseudo labels and weight them for loss computation. Extensive experiments demonstrate that OmniGaze achieves state-of-the-art performance on five datasets under both in-domain and cross-domain settings. Furthermore, we also evaluate the efficacy of OmniGaze as a scalable data engine for gaze estimation, which exhibits robust zero-shot generalization on four unseen datasets.
Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang 0001
NeurIPS1
2025 Test-Time Tuning for Zero-Shot Spatio-Temporal Action Localization
Hongyu Qu, Rui Yan 0010, Jinhui Tang 0001
PRCV (11)2
2025 Revisiting Few-Shot Compositional Action Recognition With Knowledge Calibration
abstract
The primary challenge of Few-shot Compositional Action Recognition (FSCAR) lies in effectively generalizing to and identifying unseen compositions (i.e., motions and objects) from only a few labeled videos. However, current approaches typically evaluate FSCAR as a subsidiary task of standard CAR, ignoring the insufficient generalization of the former in the regime of more significant distribution bias and limited data. To this end, we thoroughly revisit FSCAR, explicitly acknowledging the crucial role of fine-tuning, and propose a novel trio-tuning-testing framework to alleviate the problem of compositional generalization in few-shot scenarios. Specifically, we devise an effective inner-to-outer baseline, namely Motion-Object Composer (MOC), to hierarchically learn comprehensive representations from concepts to compositions. Furthermore, we propose the Trio-Knowledge Calibration (TKC) strategy to calibrate the inference, by transferring the prior visual and language knowledge learned from fine-tuning. Extensive experiments demonstrate the state-of-the-art performance of the approach compared to current competitive methods.
Hongyu Qu, Xiangbo Shu
IEEE Signal Process. Lett.2
2025 Hierarchical Motion-Enhanced Matching Framework for Few-Shot Action Recognition
abstract
Few-Shot Action Recognition (FSAR) aims to recognize novel class action with limited annotated training data from the same class. Most FSAR methods subconsciously follow the few-shot image classification solutions by solely focusing on appearance-level matching between support and query videos, such as part-level matching, frame-level matching, and segment-level matching. However, these methods, almost always, have two main limitations: 1) generally ignore the relationship among these part-, frame- and segment-level features and 2) may mismatch the same class actions under fast-term and slow-term dynamics. To this end, we present a novel Hierarchical Motion-enhanced Matching (HM${^{2}}$) framework to hierarchically learn the relation-aware multi-modal features, and jointly promote the multi-modal matching, including appearance-level matching on segments, frames, and parts, as well as the motion-level matching on dynamics. Specifically, we first propose a new Hierarchical Tokenizer (HT) to learn multi-modal features, namely utilizing a hierarchical Transformer to learn appearance-level features, along with a Slow-Fast Aware Motion (SFAM) strategy to learn motion-level features covering fast- and slow-term dynamics. Next, we propose a new Relation-aware Matcher (RM) to match the multi-modal features, by leveraging a Hierarchical Relational Graph Convolutional Network (H-RGCN) to capture the relationship among these appearance-level features. Further, a Dual Sample-to-Class Matching (DSCM) strategy is proposed to measure the bidirectional similarities among appearance- and motion-modal features by sample-to-class matching and class-to-sample matching. Extensive experiments on four golden FSAR datasets demonstrate significant performance improvements of HM${^{2}}$compared with the state-of-the-art methods.
Hailiang Gao, Guosen Xie, Rui Yan 0010, Qiongjie Cui, Hongyu Qu, Xiangbo Shu
IEEE Trans. Multim.5
2025 MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action Recognition
abstract
Recent few-shot action recognition (FSAR) methods typically perform semantic matching on learned discriminative features to achieve promising performance. However, most FSAR methods focus on single-scale (e.g., frame-level, segment-level,etc.) feature alignment, which ignores that human actions with the same semantic may appear at different velocities. To this end, we develop a novel Multi-Velocity Progressive-alignment (MVP-Shot) framework to progressively learn and align semantic-related action features at multi-velocity levels. Concretely, a Multi-Velocity Feature Alignment (MVFA) module is designed to measure the similarity between features from support and query videos with different velocity scales and then merge all similarity scores in a residual fashion. To avoid the multiple velocity features deviating from the underlying motion semantic, our proposed Progressive Semantic-Tailored Interaction (PSTI) module injects velocity-tailored text information into the video feature via feature interaction on channel and temporal domains at different velocities. The above two modules compensate for each other to make more accurate query sample predictions under the few-shot settings. Experimental results show our method outperforms current state-of-the-art methods on multiple standard few-shot benchmarks (i.e., HMDB51, UCF101, Kinetics, SSv2-full, and SSv2-small).
Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Hailiang Gao, Guosen Xie
IEEE Trans. Multim.1
2024 DTS-TPT: Dual Temporal-Sync Test-time Prompt Tuning for Zero-shot Activity Recognition
Rui Yan 0010, Hongyu Qu, Xiangbo Shu, Jinhui Tang 0001, Tieniu Tan
IJCAI2
2023 CLEGAN: Toward Low-Light Image Enhancement for UAVs via Self-Similarity Exploitation
abstract
Low-light remote sensing image enhancement (LLIE) for unmanned aerial vehicles (UAVs) has significant scientific and practical value because unfavorable lighting conditions make capture more difficult, resulting in undesired images. As acquiring real-world low-light/normal-light image pairs in the field of remote sensing is almost infeasible, performing LLIE in an unpaired manner is practical and valuable. However, without the paired data as the supervision, learning an LLIE network is challenging. To address the challenges, the paper proposes a novel yet effective method to unpaired LLIE, which maximizes the mutual information between low-light and restored images through self-similarity contrastive learning (SSCL) in a fully unsupervised fashion within a single deep GAN framework, named CLEGAN. Instead of supervising the learning using ground truth data, we propose to regularize the unpaired training using the information extracted from the input itself. The non-local patch sampling strategy in SSCL naturally makes the negative samples differ from the positive samples for discriminative representation. Moreover, the single GAN embeds the dual illumination perception module (DIPM) to handle the internal recurrence of information and overall uneven illumination distribution in remote sensing images. DIPM mainly consists of two cooperative blocks: spatial adaptive light adjustment module (SALAM) and global adaptive light adjustment module (GALAM). Specifically, SALAM exploits the internal recurrence of information in remote sensing images to encode a wider range of contextual information into local features and make proper light estimation. Simultaneously, GALAM enhances the most valuable illumination-related channels in the feature map to achieve better light estimation. The experiments on several datasets including low-light remote sensing image dataset and public low-light image datasets show that CLEGAN performs favorably against the existing unpaired LLIE approaches, and even outperforms several fully-supervised methods.
Ling Xing 0003, Hongyu Qu, Sheng Xu 0003
IEEE Trans. Geosci. Remote. Sens.2
2022 TCDNet: Tree Crown Detection From UAV Optical Images Using Uncertainty-Aware One-Stage Network
abstract
Tree crown detection plays a vital role in forestry management, resource statistics and yields forecasting. RGB high-resolution aerial images have emerged as a cost-effective source of data for tree crown detection. To address the challenges in the detection using UAV optical images, we propose a one-stage object detection network, TCDNet. First, the network provides an attention enhancement feature extraction module to enable the model to distinguish between tree crowns and their complex backgrounds. Second, an efficient loss is introduced to enable it to be aware of the overlap between adjacent trees, thus effectively avoiding misdetection. The experimental results on two publicly available datasets show that the proposed network outperforms state-of-art networks in terms of precision, recall and mean average precision.
Weichao Wu, Xijian Fan, Hongyu Qu, Xubing Yang, Tardi Tjahjadi
IEEE Geosci. Remote. Sens. Lett.3