Guang Han 0002

dblp:97/7391-2 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-scale prototype contrast and feature fusion for visible-infrared person re-identification
Qiangqiang Xie, Xudong Shi 0007, Dapeng Li 0001, Haitao Zhao 0004, Guang Han 0002
Multim. Syst.5
2026 Correction: Multi-scale prototype contrast and feature fusion for visible-infrared person re-identification
Qiangqiang Xie, Xudong Shi 0007, Dapeng Li 0001, Haitao Zhao 0004, Guang Han 0002
Multim. Syst.5
2026 A Multi-Degradation Fundus Image Restoration Network Guided by Frequency Prompt
abstract
High-quality fundus images are critical for clinical diagnosis, yet real-world acquisition challenges often introduce multi-component degradations. Current deep learning methods typically address single degradations, lacking a unified handling of complex scenarios. In this paper, we propose the Multi-degradation Fundus Image Restoration Network (MFR-Net), an all-in-one restoration framework integrating frequency-aware prompt learning. MFR-Net comprehensively extracts the frequency domain features of different degradation components, and injects them into the backbone network through designed prompt generation and interaction modules. Furthermore, to enhance the model's domain generalization capability, the unsupervised domain adaptation is incorporated into a more reliable perceptual and image quality-oriented space for domain alignment. Extensive experimental results demonstrate that the proposed method outperforms several state-of-the-art models in the restoration of degraded retinal images, especially in the restoration of complex degradations in real images, where the quantitative indicators have been improved by up to 5.42% compared with SOTA algorithms.
Guang Han 0002, Yaolong Hu, Linlin Hao, Sam Kwong
IEEE Trans. Medical Imaging1
2024 HashNeck is a Boosting Tool for Deep Learning to Hashing
abstract
The goal of hashing for image and video retrieval is to encode multimedia data into compact binary codes, allowing for efficient approximate nearest neighbor search by ensuring that similar images or videos have closely related codes in Hamming space. To improve the effectiveness of hashing, we propose introducing a classification task to assist in training the hash network and enhance the discriminability of predicted hash codes. Unlike conventional multi-task learning approaches, we propose a HashNeck structure for the classification branch that utilizes the similarities between the expected and predicted hash codes to determine whether a neuron should participate in the classification task. By only guiding neurons with correctly predicted hash codes through the classification task, we effectively resolve the conflict between the hash and classification tasks. We evaluated the effectiveness of our proposed method on benchmark image and video datasets, including ImageNet100, MS COCO, NUS-WIDE, UCF-101, and HMDB51. The experimental results on image and video retrieval tasks demonstrate that our method outperforms state-of-the-art hashing methods in terms of retrieval performance. These compelling results demonstrate the superiority of our algorithm and its potential for improving the field of deep learning to hashing.
Hua Gao, Chenchen Hu, Guang Han 0002, Jiafa Mao, Wei Huang 0015, Kaiyuan Wan
ICMR3
2024 Point-level feature learning based on vision transformer for occluded person re-identification
abstract
Person re-identification is challenging due to the presence of variations in pose and occlusion, which significantly impact the matching of visual features across different camera views and pose considerable difficulty for accurate person re-identification. This paper proposes a novel method for occluded person re-identification by introducing point-level feature learning based on vision transformers. Our approach utilizes a pose estimator to detect the keypoints of the human body and employs these points to locate intermediate features. These intermediate features of keypoints are input to a pose-based transformer branch to learn point-level features. Then, we design a part-based transformer branch to learn part-level features that capture visual features of different image parts, further enhancing the discriminative power of the learned features. Additionally, we employ a global branch to learn the global-level feature by treating the person's image as a single entity. Finally, we integrate point-level, part-level, and global-level features to represent a person's features. The experimental results on occluded and partial person re-identification datasets demonstrate the effectiveness of our proposed approach in improving re-identification. Our approach shows potential for improving person re-identification in scenarios with occlusion and pose variations.
Hua Gao, Chenchen Hu, Guang Han 0002, Jiafa Mao, Wei Huang 0015, Qiu Guan
Image Vis. Comput.3
2024 Text-to-Image Person Re-Identification Based on Multimodal Graph Convolutional Network
abstract
Text-to-image person re-identification (ReID) is a common subproblem in the field of person re-identification and image-text retrieval. Recent approaches generally follow the structure of a dual-stream network, extracting image and text features. There is no deep interaction between images and text in this approach, making it difficult for the network to learn a highly semantic feature representation. In addition, for both image data and text data, the feature extraction process is modeled in a regular way, such as using Transformer to extract sequence embeddings. However, this type of modeling disregards the inherent relationships among multimodal input embeddings. A more flexible approach to mining multimodal data, which uniformly treats the data as graphs, is proposed. In this way, the extraction and interaction of multimodal information are accomplished by means of messages passing between graph nodes. First, a unified multimodal feature extraction and fusion network is proposed based on the graph convolutional network, which enables the progression of multimodal information from ‘local’ to ‘global’. Second, an asymmetric multilevel alignment module, which focuses on more accurate ‘local’ information from a ‘global’ perspective, is proposed to progressively divide the multimodal information at each level. Last, a cross-modal representation matching strategy based on similarity distribution and mutual information is proposed to achieve cross-modal alignment. The proposed algorithm in this paper is simple and efficient, and the testing results on three public datasets (CUHK-PEDES, ICFG-PEDES and RSTPReID) show that it can achieve SOTA-level performance.
Guang Han 0002, Ziyang Li 0001, Haitao Zhao 0004, Sam Kwong
IEEE Trans. Multim.1
2023 Low-light images enhancement and denoising network based on unsupervised learning multi-stream feature modeling
Guang Han 0002, Yingfan Wang, Jixin Liu 0001, Fanyu Zeng
J. Vis. Commun. Image Represent.1
2023 Visual video evaluation association modeling based on chaotic pseudo-random multi-layer compressed sensing for visual privacy-protected keyframe extraction
Jixin Liu 0001, Yicong Li 0005, Guang Han 0002, Ning Sun 0005
J. Vis. Commun. Image Represent.3
2023 Unsupervised learning based dual-branch fusion low-light image enhancement
Guang Han 0002, Fanyu Zeng
Multim. Tools Appl.1
2023 Unsupervised Cross-View Facial Expression Image Generation and Recognition
abstract
We propose an unsupervised cross-view facial expression adaptation network (UCFEAN) to simultaneously generate and recognize cross-view facial expressions in images in an unsupervised manner. The main idea of UCFEAN is to convert the unsupervised domain adaptation between two image spaces with different appearance into semi-supervised learning (SSL) in feature spaces with the same semantic content. The cyclic image generation of cross-view facial expressions based on the generative adversarial network (GAN) is carried out to project unlabelled target images and labelled source images to the corresponding feature spaces with the same semantic content. This helps realize the unsupervised feature learning of the target image. Labels of facial expressions represented in the projected target features can then be learned using the projected source features, because the distributions of the projected features in the two domains are close enough for knowledge transfer by using SSL. Three techniques are developed to train UCFEAN in an effective and stable manner. Extensive experiments are conducted to evaluate the UCFEAN on two multi-view facial expression image databases including RaFD and Multi-PIE. The results show that the proposed method can generate realistic target images of the facial expression and recognize cross-view facial expressions with high precision.
Ning Sun 0005, Qingyi Lu, Wenming Zheng, Jixin Liu 0001, Guang Han 0002
IEEE Trans. Affect. Comput.5
2023 Multi-Stage Visual Tracking With Siamese Anchor-Free Proposal Network
abstract
The austere challenge of visual object tracking is to find the target to be tracked in various noise interference and obtain its accurate bounding box coordinates. Recently, the object tracking technology based on the Siamese network has made great breakthroughs, and more and more Siamese network trackers have been proposed with superior performance. They still have some shortcomings. To this end, a new Multi-Stage visual tracking algorithm with Siamese Anchor-Free Proposal Network (MS-SiamAFPN) is proposed in this paper. The algorithm is a three-stage Siamese network tracker composed of Feature Extraction and Fusion (FEF) sub-network, Classification and Regression (CR) sub-network, Validation and Regression (VR) sub-network in series. Firstly, the Anchor-Free Proposal Network (AFPN) module is designed in the CR stage, which can make full use of positive and negative samples for training while reducing neural network parameters. Secondly, aim to achieve better robustness and recognizability in the VR stage, on the one hand, a novel Feature Purification (FP) module is designed, which can automatically select the important channels, and extract the features of irregular regions on the input fusion features, so as to strengthen the representation ability of image features. On the other hand, the target recognition and position regression are regarded as different processing tasks, and the recognition score and position fine-tuning of candidate targets are obtained by newly designing the Dual-Branch Network (DBN) structure, thereby avoiding feature ambiguity. Due to the synergy of the above these innovations, MS-SiamAFPN has obtained a large performance improvement, and achieved SOTA performance in multiple public dataset benchmarks.
Guang Han 0002, Jinpeng Su, Yaoming Liu, Yuqiu Zhao, Sam Kwong
IEEE Trans. Multim.1
2021 Multi-stream slowFast graph convolutional networks for skeleton-based action recognition
Ning Sun 0005, Ling Leng, Jixin Liu 0001, Guang Han 0002
Image Vis. Comput.4
2021 Video action recognition with visual privacy protection based on compressed sensing
Jixin Liu 0001, Ruxue Zhang, Guang Han 0002, Ning Sun 0005, Sam Kwong
J. Syst. Archit.3
2021 Privacy-Preserving In-Home Fall Detection Using Visual Shielding Sensing and Private Information-Embedding
abstract
Falls are the main cause of accidental injuries, and even death among elderly people, especially those who live alone in their homes. The absence of a reliable fall detection system has long been a serious problem for home health monitoring. A video surveillance system can be used to monitor elderly people at home to detect falls, but the traditional implementation of such intelligent detection falls short of personal privacy-related considerations; additionally, many people do not want to be watched in their homes. To solve this problem, we propose a fall detection system with visual shielding that can ensure the safety of elderly people in their homes while preserving their personal privacy. Multilayer compressed sensing is first used to achieve visually shielded video frames. By combining low-rank sparse decomposition theory with the improved local binary pattern on the three orthogonal planes, the object features are extracted from the shielded video frames. Finally, to compensate for the information lost in the compressed video to a certain extent, a private information-embedded classification model is proposed to identify fall-related behavior. The experimental results on two public fall datasets show that the proposed method delivers impressive accuracy and a low error rate while effectively distinguishing between fall- and nonfall-related behaviors in videos.
Jixin Liu 0001, Rong Tan, Guang Han 0002, Ning Sun 0005, Sam Kwong
IEEE Trans. Multim.3
2020 Visual privacy-preserving level evaluation for multilayer compressed sensing model using contrast and salient structural features
Jixin Liu 0001, Ning Sun 0005, Guang Han 0002, Sam Kwong
Signal Process. Image Commun.4
2019 Deep spatial-temporal feature fusion for facial expression recognition in static images
Ning Sun 0005, Ruizhi Huan, Jixin Liu 0001, Guang Han 0002
Pattern Recognit. Lett.5
2019 Generalized compressed sensing with QR-based vision matrix learning for face recognition under natural scenes
Jixin Liu 0001, Guang Han 0002, Ning Sun 0005, Xiaofei Li 0002, Zhiguo Gong, Quan-Sen Sun
Signal Process. Image Commun.2
2019 Fusing Object Semantics and Deep Appearance Features for Scene Recognition
abstract
Scene images generally show the characteristics of large intra-class variety and high inter-class similarity because of complicated appearances, subtle differences, and ambiguous categorization. Hence, it is difficult to achieve satisfactory accuracy by using a single representation. For solving this issue, we present a comprehensive representation for scene recognition by fusing deep features extracted from three discriminative views, including the information of object semantics, global appearance, and contextual appearance. These views show diversity and complementarity of features. The object semantics representation of the scene image, denoted by spatial-layout-maintained object semantics features, is extracted from the output of a deep-learning-based multi-classes detector by using spatial fisher vectors, which can simultaneously encode the category and layout information of objects. A multi-direction long short-term memory-based model is built to represent contextual information of the scene image, and the activation of the fully connected layer of a convolutional neural network is used to represent the global appearance of scene image. These three kinds of deep features are then fused to draw a final conclusion for scene recognition. Extensive experiments are conducted to evaluate the proposed comprehensive representation on three benchmarks scene image database. The results show that the three deep features complement to each other strongly and are effective in improving recognition performance after fusion. The proposed method can achieve scene recognition accuracy of 89.51% on the MIT67 database, 78.93% on the SUN397 database, and 57.27% on the Places365 databases, respectively, which are better percentages than the accuracies obtained by the latest reported deep-learning-based scene recognition methods.
Ning Sun 0005, Jixin Liu 0001, Guang Han 0002, Cong Wu 0007
IEEE Trans. Circuits Syst. Video Technol.4
2016 Multi-band joint local sparse tracking via wavelet transforms
abstract
A novel multi‐band joint local sparse tracking algorithm via wavelet transforms is proposed in this study. The object image may contain rich information of different types; the authors use wavelet transforms to decompose the object image into some sub‐band images first. This will help extract the information in different frequency ranges for the object. Then same block operation is executed on all the sub‐band images. The l 2, 1 mixed‐norm is used to describe the multi‐band joint local sparse representation on each patch; it can effectively extract the structural information in different frequency ranges. Thus, more accurate object appearance model can be established. Second, the coefficients on the diagonal of coefficient matrix are extracted for the confidence degrees of the candidate objects in this band, and then the confidence degree results in all the bands are fused to determine the best candidate object in the current frame. This can effectively alleviate the object drifting. Finally, both qualitative and quantitative evaluation results on 15 challenging video sequences demonstrate that the proposed tracking algorithm in this study can achieve better tracking effects compared with the other state‐of‐the‐art algorithms.
Guang Han 0002, Jixin Liu 0001, Ning Sun 0005, Kun Du, Xiaofei Li 0002
IET Comput. Vis.1
2016 Robust object tracking based on local region sparse appearance model
Guang Han 0002, Jixin Liu 0001, Ning Sun 0005, Cailing Wang
Neurocomputing1
2015 Colour compressed sensing imaging via sparse difference and fractal minimisation recovery
abstract
In colour compressed sensing (CS) imaging, the current two bottlenecks for application are (1) high computation cost of sparse representation (SR) with over‐complete dictionary and (2) unsatisfactory imaging quality of CS recovery with l 1 ‐norm minimisation. Thus, this study proposes a novel colour CS imaging framework. In the framework, two improvements are achieved: (1) the authors present the sparse difference to reduce the computation cost of SR in RGB colour imaging; (2) the authors use fractal dimension instead of l 1 ‐norm as the object function to actualise high quality CS recovery. The feasibility of our colour CS imaging framework is proved by sseveral experiments.
Jixin Liu 0001, Xiaofei Li 0002, Guang Han 0002, Ning Sun 0005, Kun Du, Quan-Sen Sun
IET Image Process.3