Guoxi Huang

dblp:258/5002 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Bayesian Neural Networks for One-to-Many Mapping in Image Enhancement
abstract
In image enhancement tasks, such as low-light and underwater image enhancement, a degraded image can correspond to multiple plausible target images due to dynamic photography conditions. This naturally results in a one-to-many mapping problem. To address this, we propose a Bayesian Enhancement Model (BEM) that incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. To enable fast inference, we introduce a BNN-DNN framework: a BNN is first employed to model the one-to-many mapping in a low-dimensional space, followed by a Deterministic Neural Network (DNN) that refines fine-grained image details. Extensive experiments on multiple low-light and underwater image enhancement benchmarks demonstrate the effectiveness of our method.
Guoxi Huang, Ruirui Lin, Zipeng Qi, David Bull 0001, Nantheera Anantrasirichai
AAAI1
2026 Generalized adversarial feature aggregation and KAN-enhanced network for semi-supervised visible-infrared person re-identification
Rui Sun 0004, Jicheng Shen, Guoxi Huang, Jingjing Wu 0001
Image Vis. Comput.4
2026 Learning Corruption-Invariant Components and Cross-Modal Correspondence for Unsupervised Visible-Infrared Person Re-Identification
abstract
Unsupervised Visible-Infrared Person Re-Identification (US-VI-ReID) has great potential prospects because it does not require label information. However, corrupted pedestrian images collected due to corruption factors in real-world scenarios (e.g., noise, blur, and weather changes) largely limit the scalability of US-VI-ReID. In this paper, we explore the robustness of US-VI-ReID for the first time and propose a Multi-Granularity Spatial-Frequency Prototype Learning (MSPL) framework. The framework mainly consists of Multi-Channel Soft Augmentation (MSA), Robust Frequency Domain Feature Learning (RFL) module and Cross-modal Spatial-Frequency Prototype Matching (CSPM). Specifically, the MSA alleviates the sensitivity of model to color and abnormal samples through rich channel combinations and soft erasing. Subsequently, the RFL performs deep global filtering and amplitude attention compensated InstanceNorm to complete frequency and style modulation, concentrating on degradation-robust frequency content. Finally, the CSPM is designed to achieve multi-granularity prototype contrastive learning on cluster level and view level, then conduct cross-modal matching of multi-granularity spatial-frequency prototypes, thus establishing robust label association. With the above modules, our proposed framework can learn corruption-invariant feature components and generate robust cross-modal correspondence from unlabeled cross-modal images. Extensive experiments demonstrate that our MSPL outperforms other state-of-the-art methods by a large margin on the challenging SYSU-MM01-C and RegDB-C, while maintaining competitive on the SYSU-MM01 and RegDB.
Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001, Wei Jia 0001
IEEE Trans. Inf. Forensics Secur.4
2025 Implicit Alignment-Based Cross-Modal Symbiotic Network for Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification aims to utilize textual descriptions to retrieve specific person images from large image databases. The core challenge of this task lies in the significant feature differences between the abstract nature of text and the intuitiveness of images. Existing solutions primarily rely on explicit alignment of global or fine-grained local features, which lack flexibility and struggle to effectively capture and leverage subtle features and relationship information in multimodal data. Particularly, for different images of the same person, the emphasis in feature extraction should be adjusted according to the differences in text descriptions. To address these issues, this paper proposes a Cross-Modal Symbiotic Network (CMSN) based on implicit alignment. First, CMSN employs an Implicit Multi-scale Feature Integration (IMFI) module to implicitly extract and fuse multiscale features from images and text, thereby adaptively capturing the feature relationships between the two modalities. Second, a Combined Representation Learning (CRL) module is used to produce a combined representation of the text and image features, utilizing a Combined-Representation Identity Alignment (CRIA) loss to align and constrain the identity centers of the three feature vectors. Finally, we design a Semi-Positive Triplet (SPT) loss function, which defines semi-positive samples using other images and texts of the same identity, providing additional supervisory information to the model and further reducing modality heterogeneity. Extensive experiments on the CUHK-PEDES dataset demonstrate that CMSN achieves an impressive Rank-1 and mAP accuracy of 76.46% and 70.28%, respectively, significantly outperforming existing SOTA methods.
Rui Sun 0004, Guoxi Huang, Jingjing Wu 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Layered Rendering Diffusion Model for Controllable Zero-Shot Image Synthesis
Zipeng Qi, Guoxi Huang
ECCV (66)2
2024 Text-augmented Multi-Modality contrastive learning for unsupervised visible-infrared person re-identification
Rui Sun 0004, Guoxi Huang
Image Vis. Comput.2
2024 Diffusion Augmentation and Pose Generation Based Pre-Training Method for Robust Visible-Infrared Person Re-Identification
abstract
Cross-Modal Visible-Infrared Person Re-identification (VI-REID) constitutes a vital application for constructing all-time surveillance systems. However, the current VI-REID model exhibits significant performance deterioration in noisy environments. Existing algorithms endeavor to mitigate this challenge through fine-tuning stages. We contend that, in contrast to fine-tuning stages, the pre-training phase can effectively exploit the attributes of extensive unlabeled data, thereby facilitating the development of a robust VI-REID model. Therefore, in this paper, we propose a pre-training method for VI-REID based on Diffusion Augmentation and Pose Generation (DAPG), aiming to enhance the robustness and recognition rate of VI-REID models in the presence of damaged scenes. Multiple transfer experiments on the SYSU-MM01 and RegDB datasets demonstrate that our method outperforms existing self-supervised methods, as evidenced by the results.
Rui Sun 0004, Guoxi Huang, Ruirui Xie
IEEE Signal Process. Lett.2
2023 Masked Image Residual Learning for Scaling Deeper Vision Transformers
abstract
Deeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training. To ease the training of deeper ViTs, we introduce a self-supervised learning framework called $\textbf{M}$asked $\textbf{I}$mage $\textbf{R}$esidual $\textbf{L}$earning ($\textbf{MIRL}$), which significantly alleviates the degradation problem, making scaling ViT along depth a promising direction for performance upgrade. We reformulate the pre-training objective for deeper layers of ViT as learning to recover the residual of the masked image. We provide extensive empirical evidence showing that deeper ViTs can be effectively optimized using MIRL and easily gain accuracy from increased depth. With the same level of computational complexity as ViT-Base and ViT-Large, we instantiate $4.5{\times}$ and $2{\times}$ deeper ViTs, dubbed ViT-S-54 and ViT-B-48. The deeper ViT-S-54, costing $3{\times}$ less than ViT-Large, achieves performance on par with ViT-Large. ViT-B-48 achieves 86.2\% top-1 accuracy on ImageNet. On one hand, deeper ViTs pre-trained with MIRL exhibit excellent generalization capabilities on downstream tasks, such as object detection and semantic segmentation. On the other hand, MIRL demonstrates high pre-training efficiency. With less pre-training time, MIRL yields competitive performance compared to other approaches.
Guoxi Huang, Hongtao Fu, Adrian G. Bors
NeurIPS1
2022 Busy-Quiet Video Disentangling for Video Classification
abstract
In video data, busy motion details from moving regions are conveyed within a specific frequency bandwidth in the frequency domain. Meanwhile, the rest of the frequencies of video data are encoded with quiet information with substantial redundancy, which causes low processing efficiency in existing video models that take as input raw RGB frames. In this paper, we consider allocating intenser computation for the processing of the important busy information and less computation for that of the quiet information. We design a trainable Motion Band-Pass Module (MBPM) for separating busy information from quiet information in raw video data. By embedding the MBPM into a two-pathway CNN architecture, we define a Busy-Quiet Net (BQN). The efficiency of BQN is determined by avoiding redundancy in the feature space processed by the two pathways: one operating on Quiet features of low-resolution, while the other processes Busy features. The proposed BQN outperforms many recent video processing models on Something-Something V1, Kinetics400, UCF101 and HMDB51 datasets. The code is available at: https://github.com/guoxih/busy-quiet-net.
Guoxi Huang, Adrian G. Bors
WACV1
2022 BQN: Busy-Quiet Net Enabled by Motion Band-Pass Module for Action Recognition
abstract
A rich video data representation can be realized by means of spatio-temporal frequency analysis. In this research study we show that a video can be disentangled, following the learning of video characteristics according to their spatio-temporal properties, into two complementary information components, dubbed Busy and Quiet. The Busy information characterizes the boundaries of moving regions, moving objects, or regions of change in movement. Meanwhile, the Quiet information encodes global smooth spatio-temporal structures defined by substantial redundancy. We design a trainable Motion Band-Pass Module (MBPM) for separating Busy and Quiet-defined information, in raw video data. We model a Busy-Quiet Net (BQN) by embedding the MBPM into a two-pathway CNN architecture. The efficiency of BQN is determined by avoiding redundancy in the feature spaces defined by the two pathways. While one pathway processes the Busy features, the other processes Quiet features at lower spatio-temporal resolutions reducing both memory and computational costs. Through experiments we show that the proposed MBPM can be used as a plug-in module in various CNN backbone architectures, significantly boosting their performance. The proposed BQN is shown to outperform many recent video models on Something-Something V1, Kinetics400, UCF101 and HMDB51 datasets.
Guoxi Huang, Adrian G. Bors
IEEE Trans. Image Process.1
2020 Learning Spatio-Temporal Representations With Temporal Squeeze Pooling
abstract
In this paper, we propose a new video representation learn¬ing method, named Temporal Squeeze (TS) pooling, which can extract the essential movement information from a long sequence of video frames and map it into a set of few im¬ages, named Squeezed Images. By embedding the Tempo¬ral Squeeze pooling as a layer into off-the-shelf Convolution Neural Networks (CNN), we design anew video classification model, named Temporal Squeeze Network (TeSNet). The re¬sulting Squeezed Images contain the essential movement in¬formation from the video frames, corresponding to the op¬timization of the video classification task. We evaluate our architecture on two video classification benchmarks, and the results achieved are compared to the state-of-the-art.
Guoxi Huang, Adrian G. Bors
ICASSP1
2020 Region-based Non-local Operation for Video Classification
abstract
Convolutional Neural Networks (CNNs) model long-range dependencies by deeply stacking convolution operations with small window sizes, which makes the optimizations difficult. This paper presents region-based non-local (RNL) operations as a family of self-attention mechanisms, which can directly capture long-range dependencies without using a deep stack of local operations. Given an intermediate feature map, our method recalibrates the feature at a position by aggregating the information from the neighboring regions of all positions. By combining a channel attention module with the proposed RNL, we design an attention chain, which can be integrated into the off-the-shelf CNNs for end-to-end training. We evaluate our method on two video classification benchmarks. The experimental results of our method outperform other attention mechanisms, and we achieve state-of-the-art performance on the Something-Something V1 dataset. The code is available at: https://github.com/guoxih/region-based-non-local-network.
Guoxi Huang, Adrian G. Bors
ICPR1