Guyue Hu 0001

dblp:230/3820-1 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-6198-8230ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Structure and progress aware diffusion for medical image segmentation
Siyuan Song, Guyue Hu 0001, Chenglong Li 0002, Dengdi Sun, Zhe Jin 0001, Jin Tang 0001
Pattern Recognit.2
2026 Fine-Grained and Granularity-Dynamic Framework for Referring Remote Sensing Image Segmentation
Duzhi Yuan, Guyue Hu 0001, Aihua Zheng, Chenglong Li 0002, Jin Tang 0001
IEEE Signal Process. Lett.3
2026 Revisiting Frequency Domain: Spatial-Frequency Joint Tuning for Referring Image Segmentation
abstract
Referring image segmentation aims at segmenting the target object referred by a natural language expression, which requires semantic-level object understanding and pixel-level contour segmentation. Existing methods are limited to the spatial domain, thus ignoring potential discriminability from the frequency domain and missing mutual boost between the spatial and frequency domains, also facing heavy high-frequency degeneration issues. In this paper, we revisit frequency domain and propose a novel lightweight spatial-frequency joint tuning (SFJT) plugin for referring image segmentation. We decompose 4 mainstream methodology paradigms into two general stages consisting of roughly semantic-level object understanding and precisely pixel-level contour segmentation, then respectively enhance them with parameter-efficient spatial-frequency tuning strategies. Specifically, we develop a spatial-frequency joint prompting technique (SFJ-Prompt) during early object understanding stage, which mines bidirectional spatial-frequency information to realize mutually spatial-frequency boosting, facilitating more comprehensive object understanding. Besides, we introduce a LoRA-based high-frequency auxiliary branch (HF-LoRA) during latter contour segmentation stage, which compensates for heavy high-frequency degeneration issues in spatial neural networks, facilitating more precise contour segmentation. Eventually, extensive experiments of 4 mainstream methodology paradigms for referring image segmentation on 4 large-scale datasets demonstrate the effectiveness and superiority of the proposed method.
Guyue Hu 0001, Yuxing Tong, Dong Geng, Chenglong Li 0002
IEEE Trans. Circuits Syst. Video Technol.1
2026 Cross-Modal Person Retrieval With One-to-Many Relation Modeling
abstract
Existing text-image person retrieval methods are built upon fixed-point or distributional embeddings, but they typically perform single-point alignment across modalities, making it difficult to effectively capture the one-to-many cross-modal semantic associations. To address this, we propose a One-to-Many Relation modEling network (OMRE) that explicitly constructs one-to-many semantic matching structures across modalities, thereby modeling richer and more diverse semantic associations. Specifically, to achieve one-to-many matching modeling, we design a bidirectional one-to-many alignment module, which constructs cross-modal matching distributions by aggregating relations between the mean embedding and multiple sampled embeddings, and minimizes their discrepancy with the true distribution to capture complex semantic associations. To construct fine-grained one-to-many matching relationships, we propose a collaborative reconstruction-based similarity refinement module, which maximizes the semantic consistency between multiple reconstructed masked tokens and the original tokens, effectively achieving robust and precise one-to-many cross-modal fine-grained semantic alignment. Moreover, to enhance the discriminative capability of one-to-many semantic distributions, we introduce a Hard Negative Mining mechanism that focuses on semantically similar but mismatched samples, helping to refine distribution boundaries in the probabilistic space and suppress interference from hard negative samples. Extensive experiments on three public datasets demonstrate that our method not only achieves superior overall performance but also exhibits excellent generalization ability. The code will be released on https://github.com/Yifei-AHU/OMRE.
Yifei Deng, Chenglong Li 0002, Guyue Hu 0001, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Contrastive Semantic-Aware Masked Autoencoder for Point Cloud Self-Supervised Learning
abstract
Masked Autoencoder (MAE) has shown remarkable potential in self-supervised representation learning for 3D point clouds. However, these methods primarily rely on point-level or low-level feature reconstruction, forcing the model to focus on local regions while lacking enough global discriminability in the feature representation. Moreover, conventional masking strategies randomly mask some point patches, thereby neglecting the semantic structure of the point cloud and hindering the holistic understanding of global information and geometric structures. To address these challenges, we proposed a Contrastive Semantic-aware Masked Autoencoder (Point-CSMAE), which is equipped with a semantic-aware masking (SAM) strategy and a contrastive regularization (CR) mechanism. Specifically, the semantic-aware masking strategy adaptively selects patches with richer semantic information for masking and reconstruction, enhancing the understanding of global geometric structure. Furthermore, the contrastive regularization mechanism adaptively aligns the global information between the masked and visible parts, thus improving the learned global semantic representation. Meanwhile, the CR mechanism assists the SAM strategy with effective global semantic representations. Extensive experiments on various downstream tasks, including shape classification, few-shot classification, and part segmentation, demonstrate the superiority of the proposed approach.
Guyue Hu 0001
IEEE Signal Process. Lett.2
2025 Federated Client-Tailored Adapter for Medical Image Segmentation
abstract
Medical image segmentation in X-ray images is beneficial for computer-aided diagnosis and lesion localization. Existing methods mainly fall into a centralized learning paradigm, which is inapplicable in the practical medical scenario that only has access to distributed data islands. Federated Learning has the potential to offer a distributed solution but struggles with heavy training instability due to client-wise domain heterogeneity (including distribution diversity and class imbalance). In this paper, we propose a novel Federated Client-tailored Adapter (FCA) framework for medical image segmentation, which achieves stable and client-tailored adaptive segmentation without sharing sensitive local data. Specifically, the federated adapter stirs universal knowledge in off-the-shelf medical foundation models to stabilize the federated training process. In addition, we develop two client-tailored federated updating strategies that adaptively decompose the adapter into common and individual components, then globally and independently update the parameter groups associated with common client-invariant and individual client-specific units, respectively. They further stabilize the heterogeneous federated learning process and realize optimal client-tailored instead of sub-optimal global-compromised segmentation models. Extensive experiments on three large-scale datasets demonstrate the effectiveness and superiority of the proposed FCA framework for federated medical segmentation.
Guyue Hu 0001, Siyuan Song, Yukun Kang, Zhu Yin, Gangming Zhao, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Inf. Forensics Secur.1
2025 Dynamic Strip Convolution and Adaptive Morphology Perception Plugin for Medical Anatomy Segmentation
abstract
Medical anatomy segmentation is essential for computer-aided diagnosis and lesion localization in medical images. For example, segmenting individual ribs benefits localizing the lung lesions and providing vital medical measurements (such as rib spacing) for generating medical reports. Existing methods segment shape-different anatomies (such as striped ribs, bulky lungs, and angular scapula) with the same network architecture, the morphology heterogeneity is heavily overlooked. Although some shape-aware operators like deformable convolution and dynamic snake convolution have been introduced to cater to specific object morphology, they still struggle with orientation-varying strip structures, such as 24 ribs and 2 clavicles. In this paper, we propose a novel convolution plugin (DSC-AMP) for medical anatomy segmentation, which is comprised of a dynamic strip convolution (DSC) operator and an adaptive morphology perception (AMP) strategy. Specifically, the dynamic strip convolution customizes gradually varying directions and offsets for each local region, achieving dynamic striped receptive fields. Additionally, the adaptive morphology perception strategy incorporates insights from various shape-aware convolutional kernels, enabling the model to discern and integrate crucial representations corresponding to heterogeneous anatomies. Extensive experiments on two large-scale datasets demonstrate the effectiveness and superiority of the proposed approach for tackling heterogeneous medical anatomy segmentation.
Guyue Hu 0001, Yukun Kang, Gangming Zhao, Zhe Jin 0001, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Medical Imaging1
2025 Video Compressed Sensing Via Wavelet Residual Sampling and Dual-Domain Fusion
abstract
Deep learning-based compressed sensing (CS) technology attracts widespread attention owing to its remarkable reconstruction with only a few sampling measurements and low computational complexity. However, the existing video compressive sampling approaches cannot fully exploit the inherent interframe and intraframe correlations and sparsity of video sequences. To address this limitation, a novel sampling and reconstruction method for video CS (called WRDD) is proposed, which exploits the advantages of wavelet residual sampling and dual-domain fusion optimization. Specifically, in order to capture high-frequency details and achieve efficient and high-quality measurements, we propose a wavelet residual (WR) sampling strategy for the nonkeyframe sampling, which is achieved by the wavelet residuals between nonkeyframes and keyframes. Furthermore, a dual-domain (DD) fusion strategy is proposed, which fully combine intraframe and interframe to improve the reconstruction quality of nonkeyframes both in the pixel domain and multilevel feature domains. Extensive experiments demonstrate that our WRDD surpasses the state-of-the-art video and image CS methods in both subjective and objective evaluations. Besides, it exhibits outstanding antinoise capability and computational efficiency.
Zhu Yin, ZhongCheng Wu, Wuzhen Shi, Guyue Hu 0001, Weisi Lin
IEEE Trans. Multim.4
2024 Hard-Soft Pseudo Labels Guided Semi-Supervised Learning for Point Cloud Classification
abstract
Point clouds are widely applied in 3D visual sensing and perception. However, manually annotating point clouds is much more tedious and time-consuming than that for 2D images. Fortunately, semi-supervised learning can leverage massive unlabeled data to alleviate this issue, which is becoming a promising technique nowadays. In this paper, we propose a novel semi-supervised learning (SSL) framework for point cloud classification, named HPSSL. Its unsupervised learning branch performs both the representation embedding and pseudo-classification tasks. Specifically, both hard and soft pseudo labels of unlabeled samples are generated from a shared classifier to guide the class-aware contrastive learning in our SSL framework. Besides, a prediction consistency strategy is proposed to enhance the discrimination of feature representation and the exactness of pseudo labels. Furthermore, we force the supervised learning branch to interact with the unsupervised learning branch via distribution alignment, thus achieving representation consistency. Extensive experiments on three 3D shape recognition benchmarks demonstrate the effectiveness of the proposed approach.
Guyue Hu 0001
IEEE Signal Process. Lett.2
2023 Information-density Masking Strategy for Masked Image Modeling
abstract
Recent representation learning approaches mainly fall into two paradigms: contrastive learning (CL) and masked image modeling (MIM). Combining these two methods may boost the performance, but its learning process still heavily depends on the random masking strategy. We conjecture that the random masking may hinder learning the comprehensive relationship between concept and visual patches. To overcome these limitations, we propose an information-density masking (IDM) strategy for general visual transformers. Specifically, the IDM mask out the visual patches according to their activation values of attention maps. To obtain the attention maps before the reconstruction, a self-supervised training framework CAMAE is further proposed. In addition, in order to reduce the redundancy among different attention maps, we introduce a pattern-learning balance (PLB) sampling to adaptively adjust the learning progress in different attention spaces. Extensive experiments indicate that our method efficiently retains more comprehensive visual characteristics and achieves state-of-the-art performance.
Yang Chen 0071, Guyue Hu 0001
ICME3
2023 RT-Net: replay-and-transfer network for class incremental object detection
Bo Cui 0004, Guyue Hu 0001
Appl. Intell.2
2021 DeepCollaboration: Collaborative Generative and Discriminative Models for Class Incremental Learning
abstract
An important challenge for neural networks is to learn incrementally, i.e., learn new classes without catastrophic forgetting. To overcome this problem, generative replay technique has been suggested, which can generate samples belonging to learned classes while learning new ones. However, such generative models usually suffer from increased distribution mismatch between the generated and original samples along the learning process. In this work, we propose DeepCollaboration (D-Collab), a collaborative framework of deep generative and discriminative models to solve this problem effectively. We develop a discriminative learning model to incrementally update the latent feature space for continual classification. At the same time, a generative model is introduced to achieve conditional generation using the latent feature distribution produced by the discriminative model. Importantly, the generative and discriminative models are connected through bidirectional training to enforce cycle-consistency of mappings between feature and image domains. Furthermore, a domain alignment module is used to eliminate the divergence between the feature distributions of generated images and real ones. This module together with the discriminative model can perform effective sample mining to facilitate incremental learning. Extensive experiments on several visual recognition datasets show that our system can achieve state-of-the-art performance.
Bo Cui 0004, Guyue Hu 0001
AAAI2
2020 Progressive Relation Learning for Group Activity Recognition
abstract
Group activities usually involve spatio-temporal dynamics among many interactive individuals, while only a few participants at several key frames essentially define the activity. Therefore, effectively modeling the group-relevant and suppressing the irrelevant actions (and interactions) are vital for group activity recognition. In this paper, we propose a novel method based on deep reinforcement learning to progressively refine the low-level features and high-level relations of group activities. Firstly, we construct a semantic relation graph (SRG) to explicitly model the relations among persons. Then, two agents adopting policy according to two Markov decision processes are applied to progressively refine the SRG. Specifically, one feature-distilling (FD) agent in the discrete action space refines the low-level spatio-temporal features by distilling the most informative frames. Another relation-gating (RG) agent in continuous action space adjusts the high-level semantic graph to pay more attention to group-relevant relations. The SRG, FD agent, and RG agent are optimized alternately to mutually boost the performance of each other. Extensive experiments on two widely used benchmarks demonstrate the effectiveness and superiority of the proposed approach.
Guyue Hu 0001, Bo Cui 0004
CVPR1
2020 Joint Learning in the Spatio-Temporal and Frequency Domains for Skeleton-Based Action Recognition
abstract
Benefiting from its succinctness and robustness, skeleton-based action recognition has recently attracted much attention. Most existing methods utilize local networks (e.g. recurrent network, convolutional network, and graph convolutional network) to extract spatio-temporal dynamics hierarchically. As a consequence, the local and non-local dependencies, which contain more details and semantics respectively, are asynchronously captured in different level of layers. Moreover, existing methods are limited to the spatio-temporal domain and ignore information in the frequency domain. To better extract synchronous detailed and semantic information from multi-domains, we propose a residual frequency attention (rFA) block to focus on discriminative patterns in the frequency domain, and a synchronous local and non-local (SLnL) block to simultaneously capture the details and semantics in the spatio-temporal domain. In addition, to optimize the whole learning processes of the multi-branch network, we put it under a pseudo multi-task learning paradigm. During training, 1) a soft-margin focal loss (SMFL) is proposed to optimize the intra-branch separated learning process, which can automatically conduct data selection and encourage intrinsic margins in classifiers; 2) A mutual learning policy is also proposed to further facilitate the inter-branch collaborative learning process. Eventually, our approach achieves the state-of-the-art performance on several large-scale datasets for skeleton-based action recognition.
Guyue Hu 0001, Bo Cui 0004
IEEE Trans. Multim.1
2019 Skeleton-Based Action Recognition with Synchronous Local and Non-Local Spatio-Temporal Learning and Frequency Attention
abstract
Benefiting from its succinctness and robustness, skeleton-based action recognition has recently attracted much attention. Most existing methods utilize local networks (e.g recurrent, convolutional, and graph convolutional networks) to extract spatio-temporal dynamics hierarchically. As a consequence, the local and non-local dependencies, which contain more details and semantics respectively, are asynchronously captured in different level of layers. Moreover, existing methods are limited to the spatio-temporal domain and ignore information in the frequency domain. To better extract synchronous detailed and semantic information from multi-domains, we propose a residual frequency attention (rFA) block to focus on discriminative patterns in the frequency domain, and a synchronous local and non-local (SLnL) block to simultaneously capture the details and semantics in the spatio-temporal domain. Besides, a soft-margin focal loss (SMFL) is proposed to optimize the learning whole process, which automatically conducts data selection and encourages intrinsic margins in classifiers. Our approach significantly outperforms other state-of-the-art methods on several large-scale datasets.
Guyue Hu 0001, Bo Cui 0004
ICME1