VLDB 2026 Research / reviewers in the wild / expert
Hao Zhang 0110
dblp:55/2270-110
· DBLP profile ↗
20ranked-venue papers
3as first author
20since 2021 · last 2025
0000-0002-0404-6941ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CoFormer: Coupling Attentive Model for Visual Sentiment Analysis with Hierarchical Emotion Loss
Gaifang Luo, Hao Zhang 0110, Zhaoyu Xiong, Dan Xu 0001 |
CGI (1) | 2 |
| 2025 | A Brain-Inspired Multimodal Sentiment Analysis Framework via Rationale-Guided Representation
Gaifang Luo, Hao Zhang 0110, Haomin Tan, Zhijing Wu 0011, Dan Xu 0001 |
CogSci | 2 |
| 2025 | Construct a Powerful Discriminative Relationship for Few-Shot Action RecognitionabstractLearning discriminative features from very few labeled samples has gradually become a hot issue in the task of human skeleton action recognition. Most of the existing works follow the paradigms of meta-learning or contrastive learning. However, we argue that the discriminative relationship established in this way is rather simple, because the temporal and spatial features are complex and it is difficult to find the relationships of the same category. In this paper, we propose a Multi-Level Semantic Prompting Joint Contrastive Learning Head (MSJCL-Head), which consists of Joint Hard-Soft Contrastive Learning module (JCL) and Action Semantic Prompts (ASP), to obtain the discriminative representations of texts and skeletons with effect strength and discover and calibrate ambiguous samples in the feature space. A large number of experiments have been conducted on the NTU-T, NTU-S and Kinetics datasets, and the results show that our model has achieved competitive results in few-shot tasks. Qianhan Tang, Ningxin Wang, Kangjian He, Hao Zhang 0110, Dan Xu 0001 |
ICME | 5 |
| 2025 | An AI-Enhanced VR Metaverse for Ethnic Festival Culture Protection and Inheritance
Tingyu Zhu, Qianhan Tang, Hao Zhang 0110, Dan Xu 0001 |
ICXR | 7 |
| 2025 | DPA-SAM: Enhancing Medical Image Segmentation with 3D-DCAF and PGAttention
Liye Li, Kangjian He, Gaifang Luo, Hao Zhang 0110, Yijie He, Dan Xu 0001 |
PRCV (14) | 4 |
| 2025 | TSSA-Net: Transposed Sparse Self-Attention-based network for image super-resolution
Guanhao Chen, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Multi-stage image inpainting using improved partial convolutionsabstractAbstract In recent years, deep learning models have dramatically influenced image inpainting. However, many existing studies still suffer from over‐smoothed or blurred textures when missing regions are large or contain rich visual details. To restore textures at a fine‐grained level, a multi‐stage inpainting approach is proposed, which applies a series of partial inpainting modules as well as a progressive inpainting module to inpaint missing areas from their boundaries to the centre successively. Some improvements are made on the partial convolutions to reduce artifacts like blurriness, which require a convolution kernel to contain known pixels more than a certain proportion. Towards photorealistic inpainting results, the intermediate outputs from each stage are used to compute the loss. Finally, to facilitate the training process, a multi‐step training is designed that progressively adds inpainting modules to optimize the model. Experiments show that this method outperforms the current excellent techniques on the publicly available datasets: CelebA, Places2 and Paris StreetView. Cheng Li 0021, Dan Xu 0001, Hao Zhang 0110 |
IET Image Process. | 3 |
| 2024 | Affective image recognition with multi-attribute knowledge in deep neural networks
Hao Zhang 0110, Gaifang Luo, Yingying Yue, Kangjian He, Dan Xu 0001 |
Multim. Tools Appl. | 1 |
| 2024 | A multi-weight fusion framework for infrared and visible image fusion
Yiqiao Zhou, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110 |
Multim. Tools Appl. | 5 |
| 2024 | Decoupled Knowledge Embedded Graph Convolutional Network for Skeleton-Based Human Action RecognitionabstractSkeleton-based action recognition has broad prospects owing to the fact that skeleton data is more robust to scene noise and camera view changes. Recently, researchers mainly aim to explore deep-learning feature engineering with competitive recognition accuracy for skeleton actions. However, a high-performance recognition network is usually stacked by complex feature extraction modules introducing massive computational costs. In this work, we designed a powerful and universal action knowledge distillation paradigm based on decoupled knowledge distillation for transferring action knowledge from heavy teachers to lightweight students more robustly. We constructed a network architecture space consisting of the shrinking versions of outdated 2s-AGCN and searched for several robust students. On this basis, this paradigm is further developed into a powerful decoupled knowledge embedded graph convolutional network (DKE-GCN), which outperforms the teacher significantly on three public datasets and achieves the state-of-the-art. In addition, a light-DKE-GCN is designed to achieve comparable performance with teacher with 16× less parameters, 26× less FLOPs and 8× FPS. Hao Zhang 0110, Xuejie Zhang 0002, Dan Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | RegFSC-Net: Medical Image Registration via Fourier Transform With Spatial Reorganization and Channel Refinement NetworkabstractMedical image registration is crucial in medical image analysis applications. Recently, U-Net-style networks have been commonly used for unsupervised image registration, predicting dense displacement fields in full-resolution space. However, this process is resource-intensive and time-consuming for high-resolution volumetric image data. To address this challenge, this paper proposes a novel model named RegFSC-Net, which utilizes Fourier transform with spatial reorganization (SR) and channel refinement (CR) network for registration. We embed efficient feature extraction modules SR and CR modules into the encoder, and adopt a parameter-free model to drive the decoder to improve the U-shaped network. Precisely, RegFSC-Net does not directly predict the full-resolution displacement field in space but learns the low-dimensional representation of the displacement field in the bandlimited Fourier domain, which is beneficial in reducing network parameters, memory usage, and computational costs. Experimental results show that RegFSC-Net outperforms various state-of-the-art methods. Specifically, in comparison to the widely recognized Transformer-based method TransMorph, RegFSC-Net utilizes only around 8.2% of its parameters, resulting in a 1.95% higher Dice score and significantly faster inference speeds of 126.67% and 419.99% on GPU and CPU, respectively. Furthermore, we also designed three variants of RegFSC-Net and demonstrated their potential applications in computer-aided diagnosis. Chenou Liu, Kangjian He, Dan Xu 0001, Hongzhen Shi, Hao Zhang 0110, Kunyuan Zhao |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Skeleton-Based Human Action Recognition via Multi-Knowledge Flow Embedding Hierarchically Decomposed Graph Convolutional Network
Hao Zhang 0110, Shouzheng Sun, Dan Xu 0001 |
CAD/Graphics | 3 |
| 2023 | Skeleton-based Human Action Recognition via Large-kernel Attention Graph Convolutional NetworkabstractThe skeleton-based human action recognition has broad application prospects in the field of virtual reality, as skeleton data is more resistant to data noise such as background interference and camera angle changes. Notably, recent works treat the human skeleton as a non-grid representation, e.g., skeleton graph, then learns the spatio-temporal pattern via graph convolution operators. Still, the stacked graph convolution plays a marginal role in modeling long-range dependences that may contain crucial action semantic cues. In this work, we introduce a skeleton large kernel attention operator (SLKA), which can enlarge the receptive field and improve channel adaptability without increasing too much computational burden. Then a spatiotemporal SLKA module (ST-SLKA) is integrated, which can aggregate long-range spatial features and learn long-distance temporal correlations. Further, we have designed a novel skeleton-based action recognition network architecture called the spatiotemporal large-kernel attention graph convolution network (LKA-GCN). In addition, large-movement frames may carry significant action information. This work proposes a joint movement modeling strategy (JMM) to focus on valuable temporal interactions. Ultimately, on the NTU-RGBD 60, NTU-RGBD 120 and Kinetics-Skeleton 400 action datasets, the performance of our LKA-GCN has achieved a state-of-the-art level. Hao Zhang 0110, Kangjian He, Dan Xu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | An Optimized Material Point Method for Soil-Water Coupled Simulation
Zhaoyu Xiong, Hao Zhang 0110, Dan Xu 0001 |
CGI | 2 |
| 2022 | OsaMOT: Occlusion and scale-aware multi-object tracking algorithm for low viewpointabstractAbstract Multi‐object tracking (MOT), which uses the context information of image sequences to locate, maintain identities and generate trajectories of multiple targets in each frame, is key technology in the field of computer vision. To address the problems of occlusion and scale variation in low‐viewpoint MOT, OsaMOT is proposed here. First, according to the global occlusion state of each frame, OsaMOT proposes the adaptive anti‐occlusion feature to enhance the awareness and adaptability for occlusion. At the same time, OsaMOT uses the cascade screening mechanism to reduce the “virtual new target” phenomenon due to the dramatic change in target features caused by scale variation and occlusion. Finally, considering that the occluded templates will affect the tracking performance, OsaMOT proposes an adaptive anti‐noise template update mechanism according to the partial occlusion state of the target, which improves the purity of the template library and further enhances the applicability to occlusion. The experimental results show that OsaMOT can weaken the influence of scale variation, partial occlusion, short‐term full occlusion and long‐term full occlusion in the low‐viewpoint tracking scenes. Most evaluation indexes of OsaMOT under low‐viewpoint tracking scenario are superior to those of some typical algorithms proposed in recent years, and the tracking robustness is improved. Yingying Yue, Dan Xu 0001, Kangjian He, Hongzhen Shi, Hao Zhang 0110 |
IET Image Process. | 5 |
| 2022 | MAM: A multipath attention mechanism for image recognitionabstractAbstract Attention mechanism has shown excellent performance in many computer vision tasks, while the previous literature may not adequately consider different types of attention mechanisms or is individual elaborate designed for a certain network. In this paper, a general yet effective multipath attention mechanism (MAM) to explore the effect of visual attention for image recognition is proposed. In contrast with other attentions that leverage global pooling, the main advantage is that the MAM considers both the correlation of featuremaps and different scale structural information into account. The backbone representations are enhanced by adding MAM laterally along independent and separate dimensions, channel and spatial. Due to only a simple and unified calculation block is generated, MAM can be flexibly integrated into various CNNs within few parameters and trained together end‐to‐end. Furthermore, the topology structures of attention path arrangement are investigated using different connection schemes. Experimental results on several image recognition datasets show that the model outperforms various existing models. Finally, performance improvement through visualisation is intuitively discussed. The source code for the proposed attention module is publicly available. Hao Zhang 0110, Guoqin Peng, Dan Xu 0001, Hongzhen Shi |
IET Image Process. | 1 |
| 2022 | Graph transformer network with temporal kernel attention for skeleton-based action recognitionabstractSkeleton-based human action recognition has caused wide concern, as skeleton data can robustly adapt to dynamic circumstances such as camera view changes and background interference thus allowing recognition methods to focus on robust features. In recent studies, the human body is modeled as a topological graph, and the graph convolution network (GCN) is used to extract features of actions. Although GCN has a strong ability to learn spatial modes, it ignores the varying degrees of higher-order dependencies that are captured by message passing. Moreover, the joints represented by vertices are interdependent, and hence incorporating an attention mechanism to weigh dependencies is beneficial. In this work, we propose a kernel attention adaptive graph transformer network (KA-AGTN), which models the higher-order spatial dependencies between joints by the graph transformer operator based on multihead self-attention. In addition, the Temporal Kernel Attention (TKA) block in KA-AGTN generates a channel-level attention score using temporal features, which can enhance temporal motion correlation. After combining the two-stream framework and adaptive graph strategy, KA-AGTN outperforms the baseline 2s-AGCN by 1.9% and by 1% under X-Sub and X-View on the NTU-RGBD 60 dataset, by 3.2% and 3.1% under X-Sub and X-Set on the NTU-RGBD 120 dataset, and by 2% and 2.3% under Top-1 and Top-5 and achieves the state-of-the-art performance on the Kinetics-Skeleton 400 dataset. Hao Zhang 0110, Dan Xu 0001, Kangjian He |
Knowl. Based Syst. | 2 |
| 2022 | Learning multi-level representations for affective image recognitionabstractAbstract Images can convey intense affective experiences and affect people on an affective level. With the prevalence of online pictures and videos, evaluating emotions from visual content has attracted considerable attention. Affective image recognition aims to classify the emotions conveyed by digital images automatically. The existing studies using manual features or deep networks mainly focus on low-level visual features or high-level semantic representation without considering all factors. To better understand how deep networks are working for affective recognition tasks, we investigate the convolutional features by visualization them in this work. Our research shows that the hierarchical CNN model mainly relies on deep semantic information while ignoring the shallow visual details, which are essential to evoke emotions. To form a more general and discriminative representation, we propose a multi-level hybrid model that learns and integrates the deep semantics and shallow visual representations for sentiment classification. In addition, this study shows that class imbalance would affect performance as the main category of the affective dataset will overwhelm training and degenerate the deep networks. Therefore, a new loss function is introduced to optimize the deep affective model. Experimental results on several affective image recognition datasets show that our model outperforms various existing studies. The source code is publicly available. Hao Zhang 0110, Dan Xu 0001, Gaifang Luo, Kangjian He |
Neural Comput. Appl. | 1 |
| 2021 | Image Emotion Analysis Based on the Distance Relation of Emotion Categories via Deep Metric Learning
Guoqin Peng, Hao Zhang 0110, Dan Xu 0001 |
CGI | 2 |
| 2021 | Contrastive learning for a single historical painting's blind super-resolutionabstractMost of the existing blind super-resolution(SR) methods explicitly estimate the kernel in pixel space, which usually has a large deviation and results in poor SR performance. As a seminal work, DASR learns abstract representations to distinguish various degradations in the feature space, which effectively reduces degradation estimation bias. Therefore, we also employ the feature space to extract degradation representations for an ancient painting. However, most of the blind SR mehods, including DASR, are committed to removing degradations introduced by kernels, downsampling and additive noise. Among them, downsampling degradation is often accompanied by unpleasant artifacts. To address this issue, the paper designs a high-resolution(HR) representation encoder EHR based on contrastive learning to distinguish artifacts introduced by downsampling. Moreover, to optimize the ill-posed nature of blind SR, we propose a contrastive regularization(CR) to minimize the contrastive loss based on VGG-19. With the help of CR, the SR images are pulled closer to the HR images and pushed far away from bicubic LR observations. Benefiting from these improvements, our method consistently achieves higher quantitative performance and better visual quality with more natural textures than state-of-the-art approaches on a specialized painting dataset. Hongzhen Shi, Dan Xu 0001, Kangjian He, Hao Zhang 0110, Yingying Yue |
Vis. Informatics | 4 |