VLDB 2026 Research / reviewers in the wild / expert
Honghu Pan
dblp:223/1084
· DBLP profile ↗
14ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-3319-692XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adversarial flow-based generative models for visible-to-Infrared person re-Identification
Honghu Pan, Yongyong Chen, Xin Li 0034, Zhenyu He 0001 |
Pattern Recognit. | 1 |
| 2026 | Text-to-motion retrieval by text-to-motion generation
Honghu Pan, Yongyong Chen |
Pattern Recognit. | 1 |
| 2025 | Towards unified bijective image-text generation for text-to-image person re-identification
Xiaoguang Ma, Jianmin Ji, Honghu Pan |
Knowl. Based Syst. | 5 |
| 2024 | Data Generation Scheme for Thermal Modality with Edge-Guided Adversarial Conditional Diffusion ModelabstractIn challenging low-light and adverse weather conditions, thermal vision algorithms, especially object detection, have exhibited remarkable potential, contrasting with the frequent struggles encountered by visible vision algorithms. Nevertheless, the efficacy of thermal vision algorithms driven by deep learning models remains constrained by the paucity of available training data samples. To this end, this paper introduces a novel approach termed the edge-guided conditional diffusion model (ECDM). This framework aims to produce meticulously aligned pseudo thermal images at the pixel level, leveraging edge information extracted from visible images. By utilizing edges as contextual cues from the visible domain, the diffusion model achieves meticulous control over the delineation of objects within the generated images. To alleviate the impacts of those visible-specific edge information that should not appear in the thermal domain, a two-stage modality adversarial training (TMAT) strategy is proposed to filter them out from the generated images by differentiating the visible and thermal modality. Extensive experiments on LLVIP demonstrate ECDM's superiority over existing state-of-the-art approaches in terms of image generation quality. The pseudo thermal images generated by ECDM also help to boost the performance of various thermal object detectors by up to 7.1 mAP. Code is available at https://github.com/lengmo1996/ECDM. Honghu Pan, Qiang Wang 0001, Zhenyu He 0001 |
ACM Multimedia | 2 |
| 2024 | Learning diverse fine-grained features for thermal infrared tracking
Qiao Liu 0001, Gaojun Li, Honghu Pan, Zhenyu He 0001 |
Expert Syst. Appl. | 4 |
| 2024 | MIMR: Modality-Invariance Modeling and Refinement for unsupervised visible-infrared person re-identification
Zhiqi Pang, Chunyu Wang 0002, Honghu Pan, Lingling Zhao, Junjie Wang 0005, Maozu Guo 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Unified Conditional Image Generation for Visible-Infrared Person Re-IdentificationabstractThis paper proposes a unified multi-modal image generation method to address two critical challenges in visible-infrared (VI) person re-identification (ReID): the insufficiency of training samples and the large cross-modality discrepancy. To be specific, we propose to generate cross-modal and middle-modal images to explicitly reduce the modality discrepancy, and generate intra-modal images to serve as training samples for datasets augmentation. To this end, we adapt the conditional diffusion model for multi-modal image generation. The condition includes a binary modality indicator and modal-irrelative pedestrian contour to control the target modality and pedestrian identity, respectively. For the intra-modality and cross-modality image generation, we modify the structure of UNet to take as input the conditions, and estimate the conditional probability density by optimizing its variational lower bound. Furthermore, we devise modal discriminators and adversarial training strategies to achieve modality alignment. The middle-modality image generation method shares the same network architecture with intra- and cross-modality generation, but has specific training objectives. We define the middle modality as the distribution equidistant from the visible modality and infrared modality. We employ the adversarial training to measure the distance from the visible or infrared modality to the middle modality, and thus minimize the difference between these two adversarial losses, serving as an equidistant constraint. Experimental results on SYSU-MM01 and RegDB demonstrate the effectiveness and generalization of the intra-modality, cross-modality, and middle-modality image generation. Honghu Pan, Wenjie Pei, Xin Li 0034, Zhenyu He 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Pose-guided adversarial video prediction for image-to-video person re-identificationabstractAbstract The image‐to‐video (I2V) person re‐identification (Re‐ID) is a cross‐modality pedestrian retrieval task, whose crux is to reduce the large modality discrepancy between images and videos. To this end, this paper proposes to predict the following video frames from a single image. Thus, the I2V person Re‐ID can be transformed to video‐to‐video (V2V) Re‐ID. Considering that predicting video frames from a single image is an ill‐posed problem, this paper proposes two strategies to improve the quality of the predicted videos. First, a pose‐guided video prediction pipeline is proposed. The given single image and pedestrian pose are encoded via image encoder and pose encoder, respectively; then, the image feature and pose feature are concatenated as the input of the video decoder. The authors minimize the difference between the predicted video and true video, and simultaneously minimize the difference between the true pose and predicted pose. Second, the conditional adversarial training strategy is employed to generate high‐quality video frames. Specifically, the discriminator takes the source image as condition and distinguishes whether the input frames are fake or true following frames of the source image. Experimental results demonstrate that the pose‐guided adversarial video prediction can effectively improve accuracy of I2V Re‐ID. Yunqi He, Liqiu Chen, Honghu Pan |
IET Image Process. | 3 |
| 2023 | Class-guided human motion prediction via multi-spatial-temporal supervision
Honghu Pan, Lian Wu, Chao Huang 0008, Xiaoling Luo 0001, Yong Xu 0001 |
Neural Comput. Appl. | 2 |
| 2023 | Multi-granularity graph pooling for video-based person re-identification
Honghu Pan, Yongyong Chen, Zhenyu He 0001 |
Neural Networks | 1 |
| 2023 | Pose-Aided Video-Based Person Re-Identification via Recurrent Graph Convolutional NetworkabstractExisting methods for video-based person re- identification (ReID) mainly learn the appearance feature of a given pedestrian via a feature extractor and a feature aggregator. However, the appearance models would fail to learn a large inter-class variance when different pedestrians have similar appearances. Considering that different pedestrians have different walking postures and body proportions, we propose to learn the discriminative pose feature beyond the appearance feature for video retrieval. Specifically, we implement a two-branch architecture to separately learn the appearance feature and pose feature, and then concatenate them together for inference. To learn the pose feature, we first detect the pedestrian pose in each frame through an off-the-shelf pose detector, and construct a temporal graph using the pose sequence. We then exploit a recurrent graph convolutional network (RGCN) to learn the node embeddings of the temporal pose graph, which devises a global information propagation mechanism to simultaneously achieve the neighborhood aggregation of intra-frame nodes and message passing among inter-frame graphs. Finally, we propose a dual-attention method (DAM) consisting of node-attention and time-attention to obtain the temporal graph representation from the node embeddings, where the self-attention mechanism is employed to learn the importance of each node and each frame. We verify the proposed method on three video-based ReID datasets, i.e., Mars, DukeMTMC and iLIDS-VID, whose experimental results demonstrate that the learned pose feature can effectively improve the performance of existing appearance models. Honghu Pan, Qiao Liu 0001, Yongyong Chen, Yunqi He, Yuan Zheng 0002, Feng Zheng 0001, Zhenyu He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Toward Complete-View and High-Level Pose-Based Gait RecognitionabstractModel-based gait recognition methods usually adopt the pedestrian walking postures to identify human beings. However, existing methods did not explicitly resolve the large intra-class variance of human pose due to changes in camera view. In this paper, we propose a lower-upper generative adversarial network (LUGAN) to generate multi-view pose sequences for each single-view sample to reduce the cross-view variance. Based on the prior of camera imaging, we prove that the spatial coordinates between cross-view poses satisfy a linear transformation of a full-rank matrix. Hence, LUGAN employs the adversarial training to learn full-rank transformation matrices from the source pose and target views to obtain the target pose sequences. The generator of LUGAN is composed of graph convolutional (GCN) layers, fully connected (FC) layers and two-branch convolutional (CNN) layers: GCN layers and FC layers encode the source pose sequence and target view, then CNN layers take as input the encoded features to learn a lower triangular matrix and an upper one, finally the transformation matrix is formulated by multiplying the lower and upper triangular matrices. For the purpose of adversarial training, we develop a conditional discriminator that distinguishes whether the pose sequence is true or generated. Furthermore, to facilitate the high-level correlation learning, we propose a plug-and-play module, named multi-scale hypergraph convolution (HGC), to replace the spatial graph convolutional layer in baseline, which can simultaneously model the joint-level, part-level and body-level correlations. Extensive experiments on three large gait recognition datasets (i.e., CASIA-B, OUMVLP-Pose and NLPR) demonstrate that our method outperforms the baseline model by a large margin. Honghu Pan, Yongyong Chen, Tingyang Xu, Yunqi He, Zhenyu He 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | AAGCN: Adjacency-aware Graph Convolutional Network for person re-identification
Honghu Pan, Zhenyu He 0001, Chunkai Zhang |
Knowl. Based Syst. | 1 |
| 2022 | TCDesc: Learning Topology Consistent Descriptors for Image MatchingabstractThe triplet loss is widely used in learning the local descriptors for image matching. However, existing triplet loss-based methods, like HardNet and DSM, employ the point-to-point distance metric, which neglects the neighborhood information of descriptors. Considering the fact that local neighborhood structures of matching descriptors would be similar under the ideal condition, this paper aims to learn the neighborhood topology-consistent descriptors (TCDesc). To this end, we first propose the linear combination weight as the topology weight to depict the neighborhood topology for each descriptor, where the difference between the center descriptor and the linear combination of its neighbors is minimized. For the global comparison, we then define a global topology vector by using the local topology weights. Next, beyond the Euclidean distance, we define a topology distance with the topology vectors to indicate the topological difference between the matching descriptors. Furthermore, we propose an adaptive weighting strategy to jointly minimize the topology distance and Euclidean distance in triplet loss. Experimental results on four widely-used datasets, i.e., UBC PhotoTourism, HPatches, W1BS and Oxford, demonstrate that our method can effectively improve the performance of both HardNet and DSM. Honghu Pan, Yongyong Chen, Zhenyu He 0001, Fanyang Meng, Nana Fan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |