Hongchen Tan

dblp:250/2401 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FSA-GS: Fine-Structure-Aware Gaussian Splatting for sparse-view novel view synthesis
Yangeng Li, Hongchen Tan, Zili Yi, Jingchun Zhou
Comput. Vis. Image Underst.5
2026 S3CE-net: Spike-guided spatiotemporal semantic coupling and expansion network for long sequence event re-identification
Xianheng Ma, Xiuping Liu, Yi Zhang 0100, Hongchen Tan
Inf. Sci.4
2026 Semantics-aware high-frequency enhancement for event-based lip-reading
Yi Zhang 0100, Xiuping Liu, Hongchen Tan
Inf. Sci.5
2026 Spectrum-guided feature enhancement network for event person re-identification
Hongchen Tan, Yi Zhang 0100, Xiuping Liu
Pattern Recognit.1
2026 Visual-Textual Information-Driven Tactile Data Generation Method
abstract
Tactile data can enhance the environmental perception and interaction capabilities of intelligent agents, serving as a foundational component for the development of embodied intelligence. Despite its critical role, tactile data acquisition remains cost-prohibitive and labor-intensive, resulting in severe data scarcity. Cross-modal generation offers a promising solution by leveraging abundant visual and textual data. However, effectively aligning heterogeneous visual-textual modalities under data-scarce and sparsely-annotated conditions remains a significant challenge. To address these challenges, a visual-textual information-driven tactile data generation (VTTac) framework is proposed, which features three key innovations. First, a multi-granularity text enhancement strategy is introduced to mitigate annotation sparsity through hierarchical semantic enrichment. Second, a cascaded dual cross-attention mechanism is designed to ensure cross-modal alignment. Third, a condition adapter injects a low-frequency background prior, enabling the generative backbone to focus on high-frequency texture synthesis. Subsequently, a wavelet transform seamlessly fuses these synthesized details with the real background. Extensive evaluations across three datasets demonstrate that VTTac consistently outperforms representative baselines. Furthermore, downstream tasks validate the physical faithfulness of the synthesized data for material classification and semantic reasoning, and zero-shot experiments confirm generalization to unseen objects.
Zhangzheng Tu, Hongchen Tan, Huchuan Lu
IEEE Trans. Image Process.5
2026 KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality Assessment
abstract
No-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.3
2025 Spectrum-guided Spatial Feature Enhancement Network for event-based lip-reading
Yi Zhang 0100, Xiuping Liu, Hongchen Tan, Xin Li 0003
Neurocomputing3
2025 Fine-grained text and image guided point cloud completion with CLIP model
Jun Zhou 0023, Mingjie Wang 0002, Hongchen Tan, Nannan Li 0002, Xiuping Liu
Neurocomputing4
2025 Hierarchical Event-RGB Interaction Network for single-eye expression recognition
Runduo Han, Xiuping Liu, Yi Zhang 0100, Hongchen Tan, Xin Li 0003
Inf. Sci.5
2025 Adversarial perturbation and defense for generalizable person re-identification
Hongchen Tan, Kaiqiang Xu, Pingping Tao, Xiuping Liu
Neural Networks1
2025 A Bioinspired Deep Learning Framework for Saliency-Based Image Quality Assessment
abstract
Advancements in deep learning have led to significant progress in no-reference (NR) image quality assessment (NR-IQA) for evaluating the perceived quality of digital images without relying on a reference. However, existing NR-IQA models remain suboptimal in handling complex and diverse natural images. Visual saliency constitutes a critical element for enhancing the reliability of NR-IQA, but the optimal use of saliency in deep learning-based NR-IQA has not heretofore been significantly explored. In this article, we present a novel method for integrating saliency in NR-IQA, which is motivated by the saliency-based visual search mechanism that different parts of the visual input are visited by the focus of attention (FOA) in the order of decreasing saliency. By dividing saliency into the high and low levels of FOA, we build a bioinspired deep neural network-BioSIQNet-based on a multitask learning (MTL) framework. The network architecture consists of two saliency-specific tasks and one primary image quality assessment (IQA) task. The low and high saliency (HS) are separately encoded and integrated into the early and deeper layers of the IQA network, respectively, analogous to the hierarchical processing in the visual cortex of the brain that allocates low attentional resources to process the simple patterns and high resources to learn intricate representations. We demonstrate that leveraging the synergy between visual attention and image quality perception and joint learning of these interconnected visual tasks can enhance the overall learning capabilities of the primary IQA model. Experiments validate the effectiveness of our proposed BioSIQNet for NR-IQA.
Huasheng Wang, Yueran Ma, Hongchen Tan, Xiaochang Liu, Ying Chen 0011, Hantao Liu
IEEE Trans. Neural Networks Learn. Syst.3
2024 Clean and robust multi-level subspace representations learning for deep multi-view subspace clustering
Kaiqiang Xu, Kewei Tang, Zhixun Su, Hongchen Tan
Expert Syst. Appl.4
2024 Attention-Bridged Modal Interaction for Text-to-Image Generation
abstract
We propose a novel Text-to-Image Generation Network, Attention-bridged Modal Interaction Generative Adversarial Network (AMI-GAN), to better explore modal interaction and perception for high-quality image synthesis. The AMI-GAN contains two novel designs: an Attention-bridged Modal Interaction (AMI) module and a Residual Perception Discriminator (RPD). In AMI, we mainly design a multi-scale attention mechanism to exploit semantics alignment, fusion, and enhancement between text and image, to better refine details and context semantics of the synthesized image. In RPD, we design a multi-scale information perception mechanism with our proposed novel information adjustment function, to encourage the discriminator to better perceive visual differences between the real and synthesized image. Consequently, the discriminator will drive the generator to improve the visual quality of the synthesized image. Besides, based on these novel designs, we can design two versions, a single-stage generation framework (AMI-GAN-S), and a multi-stage generation framework (AMI-GAN-M), respectively. The former can synthesize high-resolution images because of its low computational cost; the latter can synthesize images with realistic detail. Experimental results on two widely used T2I datasets showed that our AMI-GANs achieve competitive performance in T2I task.
Hongchen Tan, Kaiqiang Xu, Huasheng Wang, Xiuping Liu, Xin Li 0003
IEEE Trans. Circuits Syst. Video Technol.1
2024 Blind Image Quality Assessment via Adaptive Graph Attention
abstract
Recent advancements in blind image quality assessment (BIQA) are primarily propelled by deep learning technologies. While leveraging transformers can effectively capture long-range dependencies and contextual details in images, the significance of local information in image quality assessment can be undervalued. To address this challenging problem, we propose a novel feature enhancement framework tailored for BIQA. Specifically, we devise an Adaptive Graph Attention (AGA) module to simultaneously augment both local and contextual information. It not only refines the post-transformer features into an adaptive graph, facilitating local information enhancement, but also exploits interactions amongst diverse feature channels. The proposed technique can better reduce redundant information introduced during feature updates compared to traditional convolution layers, streamlining the self-updating process for feature maps. Experimental results show that our proposed model outperforms state-of-the-art BIQA models in predicting the perceived quality of images. The code of the model will be made publicly available.
Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Hantao Liu
IEEE Trans. Circuits Syst. Video Technol.3
2024 SSPNet: Predicting Visual Saliency Shifts
abstract
When images undergo quality degradation caused by editing, compression or transmission, their saliency tends to shift away from its original position. Saliency shifts indicate visual behaviour change and therefore contain vital information regarding perception of visual content and its distortions. Given a pristine image and its distorted format, we want to be able to detect saliency shifts induced by distortions. The resulting saliency shift map (SSM) can be used to identify the region and degree of visual distraction caused by distortions, and consequently to perceptually optimise image coding or enhancement algorithms. To this end, we first create a largest-of-its-kind eye-tracking database, comprising 60 pristine images and their associated 540 distorted formats viewed by 96 subjects. We then propose a computational model to predict the saliency shift map (SSM), utilising transformers and convolutional neural networks. Experimental results demonstrate that the proposed model is highly effective in detecting distortion-induced saliency shifts in natural images.
Huasheng Wang, Jianxun Lou, Xiaochang Liu, Hongchen Tan, Roger M. Whitaker, Hantao Liu
IEEE Trans. Multim.4
2024 Global and Local Interactive Perception Network for Referring Image Segmentation
abstract
The effective modal fusion and perception between the language and the image are necessary for inferring the reference instance in the referring image segmentation (RIS) task. In this article, we propose a novel RIS network, the global and local interactive perception network (GLIPN), to enhance the quality of modal fusion between the language and the image from the local and global perspectives. The core of GLIPN is the global and local interactive perception (GLIP) scheme. Specifically, the GLIP scheme contains the local perception module (LPM) and the global perception module (GPM). The LPM is designed to enhance the local modal fusion by the correspondence between word and image local semantics. The GPM is designed to inject the global structured semantics of images into the modal fusion process, which can better guide the word embedding to perceive the whole image's global structure. Combined with the local-global context semantics fusion, extensive experiments on several benchmark datasets demonstrate the advantage of the proposed GLIPN over most state-of-the-art approaches.
Jing Liu 0059, Hongchen Tan, Yongli Hu, Huasheng Wang
IEEE Trans. Neural Networks Learn. Syst.2
2023 Deep Ordinal Regression Framework for No-Reference Image Quality Assessment
abstract
Due to the rapid development of deep learning techniques, no-reference image quality assessment (NR-IQA) has achieved significant improvement. NR-IQA aims to predict a real-valued variable for image quality, using the image in question as the sole input. Existing deep learning-based NR-IQA models are formulated as a regression problem and trained by minimising the mean squared error. The error measurement does not consider the relative ordering between different ratings on the quality scale, which consequently affects the efficacy of the model. To account for this problem, we reformulate NR-IQA learning as an ordinal regression problem and propose a simple yet effective framework using deep convolutional neural networks (DCNN) and Transformers. NR-IQA learning is achieved by a deep ordinal loss and using a soft ordinal inference to transform the predicted probabilities to a continuous variable for image quality. Experimental results demonstrate the superiority of our proposed NR-IQA model based on deep ordinal regression. In addition, this framework can be easily extended with various DCNN architectures to build advanced IQA models.
Huasheng Wang, Yulin Tu, Xiaochang Liu, Hongchen Tan, Hantao Liu
IEEE Signal Process. Lett.4
2023 ALR-GAN: Adaptive Layout Refinement for Text-to-Image Synthesis
abstract
We propose a novel Text-to-Image Generation Network, Adaptive Layout Refinement Generative Adversarial Network (ALR-GAN), to adaptively refine the layout of synthesized images without any auxiliary information. The ALR-GAN includes an Adaptive Layout Refinement (ALR) module and a Layout Visual Refinement (LVR) loss. The ALR module aligns the layout structure (which refers to locations of objects and background) of a synthesized image with that of its corresponding real image. In ALR module, we proposed an Adaptive Layout Refinement (ALR) loss to balance the matching of hard and easy features, for more efficient layout structure matching. Based on the refined layout structure, the LVR loss further refines the visual representation within the layout area. Experimental results on two widely-used datasets show that ALR-GAN performs competitively at the Text-to-Image generation task.
Hongchen Tan, Xiuping Liu, Xin Li 0003
IEEE Trans. Multim.1
2023 MHSA-Net: Multihead Self-Attention Network for Occluded Person Re-Identification
abstract
This article presents a novel person reidentification model, named multihead self-attention network (MHSA-Net), to prune unimportant information and capture key local information from person images. MHSA-Net contains two main novel components: multihead self-attention branch (MHSAB) and attention competition mechanism (ACM). The MHSAB adaptively captures key local person information and then produces effective diversity embeddings of an image for the person matching. The ACM further helps filter out attention noise and nonkey information. Through extensive ablation studies, we verified that the MHSAB and ACM both contribute to the performance improvement of the MHSA-Net. Our MHSA-Net achieves competitive performance in the standard and occluded person Re-ID tasks.
Hongchen Tan, Xiuping Liu, Xin Li 0003
IEEE Trans. Neural Networks Learn. Syst.1
2023 DR-GAN: Distribution Regularization for Text-to-Image Generation
abstract
This article presents a new text-to-image (T2I) generation model, named distribution regularization generative adversarial network (DR-GAN), to generate images from text descriptions from improved distribution learning. In DR-GAN, we introduce two novel modules: a semantic disentangling module (SDM) and a distribution normalization module (DNM). SDM combines the spatial self-attention mechanism (SSAM) and a new semantic disentangling loss (SDL) to help the generator distill key semantic information for the image generation. DNM uses a variational auto-encoder (VAE) to normalize and denoise the image latent distribution, which can help the discriminator better distinguish synthesized images from real images. DNM also adopts a distribution adversarial loss (DAL) to guide the generator to align with normalized real image distributions in the latent space. Extensive experiments on two public datasets demonstrated that our DR-GAN achieved a competitive performance in the T2I task. The code link: https://github.com/Tan-H-C/DR-GAN-Distribution-Regularization-for-Text-to-Image-Generation.
Hongchen Tan, Xiuping Liu, Xin Li 0003
IEEE Trans. Neural Networks Learn. Syst.1
2022 PMAN: Progressive Multi-Attention Network for Human Pose Transfer
abstract
This paper presents a novel approach for human pose transfer, progressive multi-attention network (PMAN), which generates a new human image by transferring the pose of a given person to a target pose. The network gradually updates the pose feature and the image feature through a series of multi-attention transfer blocks (MATBs). Each MATB consists of two attention mechanisms: pose-conditioned batch normalization (PCBN) and cooperative attention mechanism (CAM). Specifically, in low-level feature space, the PCBN layer with pose information is used to replace the BN layer of the image channel to realize the preliminary guidance of pose to image. The CAM is implemented as two gated mechanisms in high-level feature space, which reveals the essence of human pose transfer, that is, mutual guidance and dynamic control between pose and image. Gated memory writing is used to calculate the pixel-wise weight of the pose by using global image information to guide the update of the pose. Gated response utilizes an adaptive gating mechanism to dynamically control the pose information flow so as to update the image. A large number of subjective and objective experiments on DeepFashion and Market-1501 demonstrate the superiority of our method. The proposed multi-attention mechanism is well adapted to the human pose transfer task and provides a possible new idea for other cross-domain generation tasks.
Baoyu Chen, Yi Zhang 0100, Hongchen Tan, Xiuping Liu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Incomplete Descriptor Mining With Elastic Loss for Person Re-Identification
abstract
In this paper, we propose a novel person Re-ID model, Consecutive Batch DropBlock Network (CBDB-Net), to capture the attentive and robust person descriptor for the person Re-ID task. The CBDB-Net contains two novel designs: the Consecutive Batch DropBlock Module (CBDBM) and the Elastic Loss (EL). In the Consecutive Batch DropBlock Module (CBDBM), we firstly conduct uniform partition on the feature maps. And then, we independently and continuously drop each patch from top to bottom on the feature maps, which can output multiple incomplete feature maps. In the training stage, these multiple incomplete features can better encourage the Re-ID model to capture the robust person descriptor for the Re-ID task. In the Elastic Loss (EL), we design a novel weight control item to help the Re-ID model adaptively balance hard sample pairs and easy sample pairs in the whole training process. Through an extensive set of ablation studies, we verify that the Consecutive Batch DropBlock Module (CBDBM) and the Elastic Loss (EL) each contribute to the performance boosts of CBDB-Net. We demonstrate that our CBDB-Net can achieve the competitive performance on the three standard person Re-ID datasets (the Market-1501, the DukeMTMC-Re-ID, and the CUHK03 dataset), three occluded Person Re-ID datasets (the Occluded DukeMTMC, the Partial-REID, and the Partial iLIDS dataset), and a general image retrieval dataset (In-Shop Clothes Retrieval dataset).
Hongchen Tan, Xiuping Liu, Yuhao Bian, Huasheng Wang
IEEE Trans. Circuits Syst. Video Technol.1
2022 Deep Supervised Descent Method With Multiple Seeds Generation for 3-D Tracking in Point Cloud
abstract
Three-dimensional (3-D) tracking in point cloud is a core competence of autonomous robots to perceive and forecast the environment. How to initialize bounding box seeds and optimize their position and orientation are very crucial for 3-D object tracking in point clouds. Nevertheless, existing methods mainly resort to developing a powerful classifier based on the initial bounding box seeds. In this article, we propose an end-to-end deep supervised descent method (SDM), which seamlessly integrates multiple seeds generation for the initialization of seeds and sequential updates for the estimation of accurate result. Specifically, we start with transforming the SDM iterative process into a trainable recurrent module. It explicitly learns a series of descent directions in the parameter space, to gradually optimize the initial seeds. Moreover, to alleviate drifting of this process, we initialize multiple seeds based on aggregated point sets generated by the deep Hough voting. Besides, a discrimination module is introduced to determine the bounding box with the highest score as the final result. Importantly, a specific multitask loss is proposed to train our model in an end-to-end way. Experiments on KITTI, PandaSet, and Waymo datasets show that our method could achieve significant improvements (up to 11.2% in success ratio) as compared to state-of-the-art trackers.
Shengjing Tian, Bin Liu 0057, Hongchen Tan, Jun Liu 0036, Meng Liu 0006, Xiuping Liu
IEEE Trans. Ind. Informatics3
2022 Cross-Modal Semantic Matching Generative Adversarial Networks for Text-to-Image Synthesis
abstract
Synthesizing photo-realistic images based on text descriptions is a challenging image generation problem. Although many recent approaches have significantly advanced the performance of text-to-image generation, to guarantee semantic matchings between the text description and synthesized image remains very challenging. In this paper, we propose a new model, Cross-modal Semantic Matching Generative Adversarial Networks (CSM-GAN), to improve the semantic consistency between text description and synthesized image for a fine-grained text-to-image generation. Two new modules are proposed in CSM-GAN: Text Encoder Module (TEM) and Textual-Visual Semantic Matching Module (TVSMM). TVSMM is aimed at making the distance of the pairs of synthesized image and its corresponding text description closer, in global semantic embedding space, than those of mismatched pairs. This improves the semantic consistency and consequently, the generalizability of CSM-GAN. In TEM, we introduce Text Convolutional Neural Networks (Text_CNNs) to capture and highlight local visual features in textual descriptions. Thorough experiments on two public benchmark datasets demonstrated the superiority of CSM-GAN over other representative state-of-the-art methods.
Hongchen Tan, Xiuping Liu, Xin Li 0003
IEEE Trans. Multim.1
2021 Label2im: Knowledge Graph Guided Image Generation from Labels
Hewen Xiao, Yuqiu Kong, Hongchen Tan, Xiuping Liu
BMVC3
2021 KT-GAN: Knowledge-Transfer Generative Adversarial Network for Text-to-Image Synthesis
abstract
This paper presents a new framework, Knowledge-Transfer Generative Adversarial Network (KT-GAN), for fine-grained text-to-image generation. We introduce two novel mechanisms: an Alternate Attention-Transfer Mechanism (AATM) and a Semantic Distillation Mechanism (SDM), to help generator better bridge the cross-domain gap between text and image. The AATM updates word attention weights and attention weights of image sub-regions alternately, to progressively highlight important word information and enrich details of synthesized images. The SDM uses the image encoder trained in the Image-to-Image task to guide training of the text encoder in the Text-to-Image task, for generating better text features and higher-quality images. With extensive experimental validation on two public datasets, our KT-GAN outperforms the baseline method significantly, and also achieves the competive results over different evaluation metrics.
Hongchen Tan, Xiuping Liu, Meng Liu 0006, Xin Li 0003
IEEE Trans. Image Process.1
2019 Semantics-Enhanced Adversarial Nets for Text-to-Image Synthesis
abstract
This paper presents a new model, Semantics-enhanced Generative Adversarial Network (SEGAN), for fine-grained text-to-image generation. We introduce two modules, a Semantic Consistency Module (SCM) and an Attention Competition Module (ACM), to our SEGAN. The SCM incorporates image-level semantic consistency into the training of the Generative Adversarial Network (GAN), and can diversify the generated images and improve their structural coherence. A Siamese network and two types of semantic similarities are designed to map the synthesized image and the groundtruth image to nearby points in the latent semantic feature space. The ACM constructs adaptive attention weights to differentiate keywords from unimportant words, and improves the stability and accuracy of SEGAN. Extensive experiments demonstrate that our SEGAN significantly outperforms existing state-of-the-art methods in generating photo-realistic images. All source codes and models will be released for comparative study.
Hongchen Tan, Xiuping Liu, Xin Li 0003, Yi Zhang 0100
ICCV1
2019 Feature preserving GAN and multi-scale feature enhancement for domain adaption person Re-identification
Xiuping Liu, Hongchen Tan, Xin Tong 0001, Junjie Cao 0001, Jun Zhou 0023
Neurocomputing2