Hanyu Guo

dblp:202/2521 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Vision-Language Enhancement Network Based on Decoupling-Joint Adaptation for Few-Shot Action Recognition
abstract
Learning robust and generalizable feature extractors to generate discriminative prototypes is crucial for few-shot action recognition. However, most existing methods rely on fine-tuning large pre-trained image models, easily leading to transferability and overfitting issues. In this paper, we propose a novel vision-language enhancement network based on decoupling-joint adaptation (VEDA) for few-shot action recognition, which decouples visual features into temporal and spatial branches, followed by a joint operation that integrates these two branches using an adapter-tuning paradigm. VEDA can gradually equip the model with spatio-temporal reasoning capabilities. Since relying exclusively on local frame feature matching results in inaccurate performance, we design a video-level relation module (VLR) to enhance video context awareness through global feature matching. In addition, we design a vision-language fusion module (VLF) that introduces multimodal information to alleviate the data scarcity issue. Simultaneously, we apply adapter-tuning to both visual and textual branches to enhance the generalization ability. Based on the proposed components above, our network can extract both informative and discriminative prototypes, resulting in excellent recognition performance. Experimental results on five challenging benchmarks demonstrate the effectiveness of the proposed VEDA. The code will be released soon at https://github.com/ReverseSuzhou/VEDA.
Suzhou Que, Hanyu Guo, Kaiwen Du, Yan Yan 0001, Yanwei Pang, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 TFPA: Text Features Guided Dynamic Parameter Adjustment for Few Shot Action Recognition
abstract
Most few-shot learning methods aim to train models to learn parameters that can generalize to new categories using training sets, after which the model parameters are typically fixed. However, due to limited data, models often fail to learn generalizable parameters, as they tend to overfit source domain-specific inductive biases. This can lead to catastrophic forgetting or poor adaptation to new domains. Unlike previous methods, we propose a Text Feature guided dynamic Parameter Adjustment (TFPA) method for few-shot action recognition. Inspired by basis decomposition in vector spaces, TFPA reformulates the traditional linear layer into a set of basis mapping matrices in the parameter space. Each matrix functions analogously to a basis vector in linear algebra, and their linear combinations collectively span the parameter space. To construct a domain-adaptive parameter matrix from these combinations, we propose a Coordinate Vector Computation (CVC) module, which leverages text features as semantic guidance to adaptively estimate optimal linear combination coefficients for the basis mapping matrices. Furthermore, we propose a Centroid Exclusion Loss (CEL) and a Contrastive Clustering Loss (CCL) to enhance the distinctiveness among the basis mapping matrices. These regularization terms promote functional specialization and reduce redundancy across the basis mapping matrices, thereby enhancing performance. Experimental results on five benchmark datasets demonstrate the effectiveness and strong generalization ability of our method in few-shot action recognition. The code will be released soon at https://github.com/ReverseSuzhou/TFPA.
Hanyu Guo, Suzhou Que, Junlong Gao, Hanzi Wang
ACM Multimedia1
2025 Edge Guided Network With Motion Enhancement for Few-Shot Action Recognition
abstract
Existing state-of-the-art methods for few-shot action recognition (FSAR) achieve promising performance by spatial and temporal modeling. However, most current methods ignore the importance of edge information and motion cues, leading to inferior performance. For the few-shot task, it is important to effectively explore limited data. Additionally, effectively utilizing edge information is beneficial for exploring motion cues, and vice versa. In this paper, we propose a novel edge guided network with motion enhancement (EGME) for FSAR. To the best of our knowledge, this is the first work to utilize the edge information as guidance in the FSAR task. Our EGME contains two crucial components, including an edge information extractor (EIE) and a motion enhancement module (ME). Specifically, EIE is used to obtain edge information on video frames. Afterward, the edge information is used as guidance to fuse with the frame features. In addition, ME can adaptively capture motion-sensitive features of videos. It adopts a self-gating mechanism to highlight motion-sensitive regions in videos from a large temporal receptive field. Based on the above designed components, EGME can capture edge information and motion cues, resulting in superior recognition performance. Experimental results on four challenging benchmarks show that EGME performs favorably against recent advanced methods.
Kaiwen Du, Weirong Ye, Hanyu Guo, Yan Yan 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Bi-Directional Motion Attention with Contrastive Learning for few-shot Action Recognition
abstract
In recent years, many few-shot action recognition methods have achieved competitive performance by adopting metric-based techniques. However, they suffer from two limitations: (1) Spatio-temporal relationship is modeled independently, overlooking the spatio-temporal correspondence between target objects across video frames. (2) Inter-class similarities are not well exploited in the task. As a result, their performance is significantly constrained by the presence of similar segments among different classes. In this paper, a novel BiMACL method for few-shot action recognition is presented, consisting of a Temporal Difference Spatial Attention Module (TDSAM) that uses motion attention to effectively capture the spatio-temporal correspondence between video frames, and a Contrastive Temporal-Relational CrossTransformers (CTRX) module to alleviate the adverse effects of similar subsequences of frames among distinct classes. Extensive experimental results demonstrate the superiority of our method over most methods for few-shot action recognition. Code is available at https://github.com/YWCandGHY/BiMACL.
Hanyu Guo, Wanchuan Yu, Yan Yan 0001, Hanzi Wang
ICASSP1
2023 Community Tour: An Expandable Knowledge Exploration System for Urban Migrant Children
abstract
Urban migrant children encounter difficulties in developing a sense of belonging, which compromises their living experiences and academic performance. This paper introduces Community Tour, an expandable knowledge exploration system designed to assist migrant children with their extracurricular learning, and empower social workers to systematically carry out community events. A knowledge exploration interaction process has been designed to localize STEAM education with community elements through practical tasks, learning motivation, and achievements. Prototypes of interactive installation and back-end platform have been built and partial validation experiments have been conducted, with future work focusing on collaborating with communities for field testing and design iteration. The sustainability of the system lies in the potential for education on various themes, and its compatibility from urban villages to more regular communities, contributing to the child-friendly cities.
Bo Shui, Hanyu Guo, Haoyang Li 0005, Chufan Shi, Xiaomei Nie
IDC2
2023 Wesee: Digital Cultural Heritage Interpretation for Blind and Low Vision People
Yalan Luo, Weiyue Lin, Xiaomei Nie, Xiang Qian, Hanyu Guo
INTERACT (1)6
2023 Spatio-Temporal Self-supervision for Few-Shot Action Recognition
Wanchuan Yu, Hanyu Guo, Yan Yan 0001, Jie Li 0001, Hanzi Wang
PRCV (1)2
2018 Generalized Value Iteration Networks: Life Beyond Lattices
abstract
In this paper, we introduce a generalized value iteration network (GVIN), which is an end-to-end neural network planning module. GVIN emulates the value iteration algorithm by using a novel graph convolution operator, which enables GVIN to learn and plan on irregular spatial graphs. We propose three novel differentiable kernels as graph convolution operators and show that the embedding-based kernel achieves the best performance. Furthermore, we present episodic Q-learning, an improvement upon traditional n-step Q-learning that stabilizes training for VIN and GVIN. Lastly, we evaluate GVIN on planning problems in 2D mazes, irregular graphs, and real-world street networks, showing that GVIN generalizes well for both arbitrary graphs and unseen graphs of larger scaleand outperforms a naive generalization of VIN (discretizing a spatial graph into a 2D image).
Sufeng Niu, Siheng Chen, Hanyu Guo, Colin Targonski, Melissa C. Smith, Jelena Kovacevic
AAAI3