Jing Hu 0009

dblp:95/6046-9 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-0921-0592ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 RL-I2IT: Image-to-image translation with deep reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Chengming Feng, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Xin Li 0005, Hongtu Zhu, Siwei Lyu, Xin Wang 0045
Neural Networks1
2025 RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation
abstract
Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https://github.com/fengxiaoming520/RLMiniStyler.
Jing Hu 0009, Chengming Feng, Shu Hu 0001, Ming-Ching Chang, Xin Li 0005, Xi Wu 0004, Xin Wang 0045
IJCAI1
2025 Improving Generalization of Medical Image Registration Foundation Model
abstract
Deformable registration is a fundamental task in medical image processing, aiming to achieve precise alignment by establishing nonlinear correspondences between images. Traditional methods offer good adaptability and interpretability but are limited by computational efficiency. Although deep learning approaches have significantly improved registration speed and accuracy, they often lack flexibility and generalizability across different datasets and tasks. In recent years, foundation models have emerged as a promising direction, leveraging large and diverse datasets to learn universal features and transformation patterns for image registration, thus demonstrating strong cross-task transferability. However, these models still face challenges in generalization and robustness when encountering novel anatomical structures, varying imaging conditions, or unseen modalities. To address these limitations, this paper incorporates Sharpness-Aware Minimization (SAM) into foundation models to enhance their generalization and robustness in medical image registration. By optimizing the flatness of the loss landscape, SAM improves model stability across diverse data distributions and strengthens its ability to handle complex clinical scenarios. Experimental results show that foundation models integrated with SAM achieve significant improvements in cross-dataset registration performance, offering new insights for the advancement of medical image registration technology. Our code is available at https://github.com/Promise13/fm_sam.
Jing Hu 0009, Kaiwei Yu, Hongjiang Xian, Shu Hu 0001
IJCNN1
2025 Multi-Source Feature Fusion and Spatio-Temporal Unet for Precipitation Nowcasting
abstract
Precipitation nowcasting is an extremely critical task in the field of weather forecasting, as it facilitates advancements in meteorological observation. Nevertheless, accurate short-term precipitation forecasting remains a significant challenge at present. Traditional methods have relied on physical equations for predictions, which are often computationally consuming. Current deep learning approaches, using CNNs and RNNs, roughly extract the latent features of spatiotemporal data, but the feature extraction process usually overlooks the dynamic changes occurring between prediction image frames. Furthermore, most methods utilize a single precipitation variable as input for predicting future precipitation, neglecting that precipitation events are triggered by multiple meteorological factors. To tackle this issue, we propose a novel neural network model, Multi-Source Feature Fusion and Spatio-Temporal Unet (MFFST-Unet), which utilizes multi-source feature information to guide precipitation forecasting. Additionally, we introduce the Inter-Frame Difference Regularization(IFDR) Loss, which is combined with MSE Loss to optimize the frame stability of model predictions through adaptive weighting. We conducted training and testing on the SEVIR dataset, achieving high-resolution precipitation nowcasting for a one-hour forecast. Experimental results indicate that our MFFST-Unet model surpasses other deep learning methods, achieving a maximum improvement of 34.72% in the CSI precipitation metric, demonstrating its significant practical implications for weather forecasting applications.
Dufu Liu, Xia Yuan, Xi Wu 0004, Jing Hu 0009
IJCNN6
2024 RepMedGraf: Re-parameterization Medical Generated Radiation Field for Improved 3D Image Reconstruction
abstract
In the field of medical imaging, 3D image reconstruction has emerged as a crucial technique for accurate disease diagnosis. Neural Radiance Field (NeRF) has shown promise in generating high-quality 3D models through deep learning, but its application in medical imaging remains limited. This paper presents RepMedGraf, an improved model addressing the limitations of NeRF. RepMedGraf utilizes a generator based on lightweight RepVGG Blocks instead of MedNeRF model. By doing so, the risk of overfitting is reduced, and training efficiency is improved. The proposed method is trained on publicly available chest and knee datasets. Comparative evaluations are conducted based on the generated radiation fields, demonstrating the effectiveness and quality of the 3D models produced by RepMedGraf. The results highlight the potential of RepMedGraf as a valuable tool in medical imaging for enhanced diagnostic accuracy and improved patient care.
Ruotong Sun, Fayang Liao, Qinrui Fan, Shu Hu 0001, Xin Wang 0045, Jing Hu 0009
AVSS9
2024 High fidelity medical image super-resolution based on Medical Multi-Feature Compensation Attention GAN
abstract
MRI and CT medical images play a crucial role in clinical medicine, and high-resolution images enhance diagnostic quality. Improving image resolution through hardware upgrades is often expensive and may increase radiation exposure for patients. Common super-resolution methods used for natural images do not account for the unique relationships between colors and pixels in medical images. In this paper, we present a novel approach to medical image super-resolution using the Medical Multi-Feature Compensation Attention GAN (MMSRGAN). The proposed method integrates a feature compensation attention module, enhancing the reconstruction of high-resolution images by addressing the unique characteristics and correlations in medical data. Extensive experiments on IXI-T1 datasets demonstrate that MMSRGAN significantly outperforms existing deep learning-based methods, achieving state-of-the-art results in terms of PSNR and SSIM. This advancement highlights the potential for improved diagnostic accuracy and treatment planning in clinical settings through enhanced medical imaging.
Qinrui Fan, Xi Wu 0004, Jing Hu 0009
IJCB4
2024 Contextual Reinforcement Learning for Unsupervised Deformable Multimodal Medical Images Registration
abstract
Multimodal deformable image registration refers to the process of finding the spatial correspondence between pairs of images with multimodal and mapping them onto the same coordinate system. Most of the deep learning-based registration methods are one-shot registration, which is difficult to handle images with significant deformations or displacements. Reinforcement learning can handle these challenges by viewing registration as a strategic decision-making process which is a step-by-step registration. However, it faces challenges with high-dimensional and continuous deformation fields. To overcome this, we introduce a planner network that maps high-dimensional input state to low-dimensional plan, guiding the actor to generate continuous actions. In order to handle complex multi-modal registration, we propose a multi-frame plan module which encourages artificial agent to explicitly utilize the redundant states in the registration process and learn more accurate registration actions from the generated state frames. To facilitate the training and convergence of the model, we define an unsupervised reward function and incorporate spectral normalization layers. The entire framework is a fully unsupervised registration framework and training in an end-to-end manner. We evaluated our method on publicly available T1w and T2w brain datasets, and the results indicate that our method has excellent deformable registration capability for multimodal images.
Hongjiang Xian, Zhikun Shuai, Jing Hu 0009, Shu Hu 0001
IJCB4
2024 Efficient Image Super-Resolution via Symmetric Visual Attention Network
abstract
In recent years, efficient super-resolution research has focused on reducing model complexity and improving efficiency by leveraging deep small-kernel convolution, but it has the problem of a small receptive field, which leads to a limited ability of the network to reconstruct details. Large kernel convolution can provide a large receptive field and lead to a substantial enhancement in the quality of image reconstruction, but its computational cost is too high. To minimize the model’s parameter count and achieve efficient super-resolution reconstruction, this study introduces a symmetric visual attention network. The network decomposes the large kernel convolution into three different lightweight and efficient convolutions. It then forms a bottleneck structure by leveraging the varied receptive field sizes of these convolutions in combination. The attention mechanism is integrated to create a bottleneck attention module, enhancing the network’s feature awareness. Furthermore, the bottleneck attention modules are symmetrically arranged to construct a symmetric large kernel attention block, thereby further enhancing the network’s capability to extract deep features. The experimental results demonstrate that the proposed model achieves competitive quantitative metrics when compared to other lightweight super-resolution methods, and the details of the reconstructed images are enhanced. With only 183K parameters, the model achieves a lightweight yet high-quality super-resolution model, offering a novel solution approach for efficient super-resolution.
Qinrui Fan, Chengxu Wu, Shu Hu 0001, Xi Wu 0004, Xin Wang 0001, Jing Hu 0009
IJCNN6
2024 Dual-granularity feature fusion in visible-infrared person re-identification
abstract
Abstract Visible‐infrared person re‐identification (VI‐ReID) aims to recognize images of the same person captured in different modalities. Existing methods mainly focus on learning single‐granularity representations, which have limited discriminability and weak robustness. This paper proposes a novel dual‐granularity feature fusion network for VI‐ReID. Specifically, a dual‐branch module that extracts global and local features and then fuses them to enhance the representative ability is adopted. Furthermore, an identity‐aware modal discrepancy loss that promotes modality alignment by reducing the gap between features from visible and infrared modalities is proposed. Finally, considering the influence of non‐discriminative information in the modal‐shared features of RGB‐IR, a greyscale conversion is introduced to extract modality‐irrelevant discriminative features better. Extensive experiments on the SYSU‐MM01 and RegDB datasets demonstrate the effectiveness of the framework and superiority over state‐of‐the‐art methods.
Shuang Cai, Shanmin Yang, Jing Hu 0009, Xi Wu 0004
IET Image Process.3
2023 Controlling Neural Style Transfer with Deep Reinforcement Learning
abstract
Controlling the degree of stylization in the Neural Style Transfer (NST) is a little tricky since it usually needs hand-engineering on hyper-parameters. In this paper, we propose the first deep Reinforcement Learning (RL) based architecture that splits one-step style transfer into a step-wise process for the NST task. Our RL-based method tends to preserve more details and structures of the content image in early steps, and synthesize more style patterns in later steps. It is a user-easily-controlled style-transfer method. Additionally, as our RL-based model performs the stylization progressively, it is lightweight and has lower computational complexity than existing one-step Deep Learning (DL) based models. Experimental results demonstrate the effectiveness and robustness of our method.
Chengming Feng, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin B. Zhu, Xi Wu 0004, Hongtu Zhu, Siwei Lyu
IJCAI2
2022 Stochastic Planner-Actor-Critic for Unsupervised Deformable Image Registration
abstract
Large deformations of organs, caused by diverse shapes and nonlinear shape changes, pose a significant challenge for medical image registration. Traditional registration methods need to iteratively optimize an objective function via a specific deformation model along with meticulous parameter tuning, but which have limited capabilities in registering images with large deformations. While deep learning-based methods can learn the complex mapping from input images to their respective deformation field, it is regression-based and is prone to be stuck at local minima, particularly when large deformations are involved. To this end, we present Stochastic Planner-Actor-Critic (spac), a novel reinforcement learning-based framework that performs step-wise registration. The key notion is warping a moving image successively by each time step to finally align to a fixed image. Considering that it is challenging to handle high dimensional continuous action and state spaces in the conventional reinforcement learning (RL) framework, we introduce a new concept `Plan' to the standard Actor-Critic model, which is of low dimension and can facilitate the actor to generate a tractable high dimensional action. The entire framework is based on unsupervised training and operates in an end-to-end manner. We evaluate our method on several 2D and 3D medical image datasets, some of which contain large deformations. Our empirical results highlight that our work achieves consistent, significant gains and outperforms state-of-the-art methods.
Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Shu Hu 0001, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Siwei Lyu
AAAI2
2022 Learning a deep dual-level network for robust DeepFake detection
Wenbo Pu, Jing Hu 0009, Xin Wang 0045, Yuezun Li, Shu Hu 0001, Bin B. Zhu, Rui Song 0006, Qi Song 0001, Xi Wu 0004, Siwei Lyu
Pattern Recognit.2
2021 Stochastic Actor-Executor-Critic for Image-to-Image Translation
abstract
Training a model-free deep reinforcement learning model to solve image-to-image translation is difficult since it involves high-dimensional continuous state and action spaces. In this paper, we draw inspiration from the recent success of the maximum entropy reinforcement learning framework designed for challenging continuous control problems to develop stochastic policies over high dimensional continuous spaces including image representation, generation, and control simultaneously. Central to this method is the Stochastic Actor-Executor-Critic (SAEC) which is an off-policy actor-critic model with an additional executor to generate realistic images. Specifically, the actor focuses on the high-level representation and control policy by a stochastic latent action, as well as explicitly directs the executor to generate low-level actions to manipulate the state. Experiments on several image-to-image translation tasks have demonstrated the effectiveness and robustness of the proposed SAEC when facing high-dimensional continuous space problems.
Ziwei Luo 0002, Jing Hu 0009, Xin Wang 0045, Siwei Lyu, Bin Kong 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004
IJCAI2
2021 End-to-end multimodal image registration via reinforcement learning
Jing Hu 0009, Ziwei Luo 0002, Xin Wang 0045, Shanhui Sun, Youbing Yin, Kunlin Cao, Qi Song 0001, Siwei Lyu, Xi Wu 0004
Medical Image Anal.1
2018 Robust Multimodal Image Registration Using Deep Recurrent Reinforcement Learning
Shanhui Sun, Jing Hu 0009, Mingqing Yao, Jinrong Hu, Qi Song 0001, Xi Wu 0004
ACCV (2)2
2018 Noise Robust Single Image Super-Resolution Using a Multiscale Image Pyramid
abstract
Single image super-resolution (SR) generates a high-resolution (HR) image by estimating the mapping function between image patches of different resolutions. However, this kind of SR method cannot be directly applied to noisy images, since noise will be reinforced in the process of super-resolution. To this end, this paper presents a simultaneous super-resolution and denoising method by exploiting the noise decreasing property contained in the multiscale image pyramid. Experimental results confirm that our method is able to outperform other state-of-the-art super-resolution methods when super-resolving noisy images across differing noise levels.
Jing Hu 0009, Xi Wu 0004, Jiliu Zhou
ICIP1
2018 Noise robust single image super-resolution using a multiscale image pyramid
Jing Hu 0009, Xi Wu 0004, Jiliu Zhou
Signal Process.1