Jiahuan Ji

dblp:254/8057 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-3592-7362ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Deep bilateral learning for image interpolation
Jiahuan Ji, Kai-Kuang Ma, Baojiang Zhong, Fuhui Zhou, Qihui Wu 0001
Knowl. Based Syst.1
2025 From Static Dense to Dynamic Sparse: Vision-Radar Fusion-Based UAV Detection
abstract
Precise unmanned aerial vehicle (UAV) detection over long distances is of crucial importance for guaranteeing the airspace security. Although deep learning-based vision detectors have been developed, they still rely on a large amount of hand-crafted fixed feature priors. The existing static dense-based detectors suffer from the severe mismatch and imbalance between the small size and the high mobility of UAVs. To solve the problem, a novel multimodal fusion-based dynamic sparse UAV detection framework is proposed. The framework reformulates the feature priors in a completely dynamic sparse paradigm by using the radar data. Based on the framework, a vision-radar fusion-based dynamic sparse network (Vira-DSNet) is proposed for more balanced and robust UAV detection. The Vira-DSNet exploits our designed dynamic sparse candidate generator and radar-guided semantic feature transform to generate a small set of customized high-quality object candidates and semantic features based on the radar data. Moreover, based on Hungarian bisection matching, our Vira-DSNet eliminates the post-processing and is completely end-to-end differentiable. Furthermore, the Vira-DSNet is deployed in our developed actual vision-radar fusionbased UAV detection system to evaluate the performance in the practical applications. Experimental results demonstrate that our Vira-DSNet achieves an average precision AP50of 88.2%. It is also shown that the average recall AR1of Vira-DSNet is higher than the state-of-the-art scheme by 10.1%, while maintaining the real-time performance.
Yiyao Wan, Jiahuan Ji, Fuhui Zhou, Qihui Wu 0001, Tony Q. S. Quek
IEEE Trans. Inf. Forensics Secur.2
2025 A Multimodal Scale Normalization Framework for Vision-Radar Small UAV Positioning
abstract
Uncrewed aerial vehicles (UAVs) positioning is of crucial importance in diverse applications. However, it is extremely challenging to realize the precise UAVs positioning over long distances due to the small size and dramatic scale variations associated with the high mobility in the wide area. To tackle this issue, a multimodal scale normalization framework is proposed for the scale-robust precise pixel-level UAV positioning. The framework exploits our proposed distance-aware image slicing and distance-aware scale normalization module. Moreover, a modal fusion-based scale normalization network is proposed that can accept arbitrary low-resolution UAV patches and produce the consistent high-resolution images at a uniform UAV instance scale with a single learnable model. The proposed framework is generic and can be directly used in the existing pixel-level positioning pipelines to improve the positioning performance and scale robustness. To verify the proposed framework in the real application, a practical vision-radar UAV positioning system is developed. Experimental results on the real-world dataset demonstrate the generality and effectiveness of our framework. Moreover, the ablation experiments also confirm the contribution of each module in the framework.
Yiyao Wan, Jiahuan Ji, Wenqing Xie, Fuhui Zhou, Qihui Wu 0001
IEEE Trans. Mob. Comput.2
2024 An Image Decomposition-Guided Network for Image Interpolation
abstract
A novel image decomposition-guided network (IDGN) for image interpolation is proposed in this paper by incorporating the fundamentals of subband image decomposition into the design of our deep-learning network. In our work, a filter bank consisting of a Gaussian filter and a differenceof-Gaussian filter is designed for decomposing the low-resolution input image into multiple subbands of the same resolution without downsampling. These subbands are inherited with different low-frequency and high-frequency information and are ready to be interpolated individually in our developed network. For training our IDGN, the decomposed low-resolution subbands need to be paired up with their corresponding ground-truth high-resolution subbands. Since our human visual system is sensitive to high-frequency signals, a perception-regulated (PR) loss function is proposed to guide our IDGN by putting more emphasis on the high-frequency subbands during the training process. Extensive experimental results have shown that our IDGN can achieve superior performance when compared with a number of state-of-the-art image interpolation methods.
Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma, Fuhui Zhou, Qihui Wu 0001
ICIP1
2024 A Channel-Wise Multi-Scale Network for Single Image Super-Resolution
abstract
Existing multi-scale feature extraction methods extract image features using various convolution window sizes conducted on the spatial dimension of the feature maps. However, such an approach inevitably encounters redundant convolution operations. To address this concern, we propose to extract multiscale features on the channel dimension rather than on the spatial dimension. To demonstrate, a channel-wise multi-scale network (CMSN) is proposed for conducting single image super-resolution (SISR). In our CMSN, a sequence of channel-wise multi-scale blocks (CMSBs) is designed to extract multi-scale features at increasing levels by performing convolutions with different channel numbers (i.e., scales). To fuse the image features generated from different levels in our CMSN, a hybrid attention-aware feature fusion block (HAFFB) is proposed. Extensive experimental results have clearly shown the superiority of our CMSN to that of several state-of-the-art SISR methods on delivering superior high-resolution images, both objectively and subjectively. This reveals the potential of channel-wise, versus spatial-wise, on the effectiveness of multi-scale feature extraction.
Jiahuan Ji, Baojiang Zhong, Qihui Wu 0001, Kai-Kuang Ma
IEEE Signal Process. Lett.1
2023 A Content-Based Multi-Scale Network for Single Image Super-Resolution
abstract
A novel content-based multi-scale network (CMNet) is proposed in this paper for conducting single image super-resolution (SISR). Its core lies in a content-based multi-scale image representation (CMIR), which is motivated by the fact that the contents of real-world images normally have different scales. Thus, it is expected that individual treatments of these contents would yield superior SISR performance. In our CMIR, the difference curvature (DCurv) is first exploited to generate a primal sketch of the input image. Then, a filter bank is designed and used to obtain a set of coefficient matrices, and each matrix reflects the characteristics of the image content at the corresponding scale. Based on these coefficient matrices, the CMIR of the input image is formed. To conduct SISR, each scale of CMIR is processed individually in our CMNet, and the produced multi-scale outputs are then integrated to arrive at the final SISR image with a higher image quality. Extensive experiments have demonstrated that our developed CMNet can deliver superior performance compared with a number of state-of-the-art SISR methods.
Jiahuan Ji, Baojiang Zhong, Kai-Kuang Mu
ICASSP1
2023 Learning Multi-Scale Features for Jpeg Image Artifacts Removal
abstract
A recently proposed quantization-table convolutional network (QCN) has proven as a state-of-the-art JPEG image artifacts removal method. To reduce computational complexity, the QCN learns image features from the down-sampled version of the input image. Consequently, the performance might be compromised as some salient features that can only be learned from the original input image with full resolution will be lost. To solve this problem, a novel multi-scale feature extraction block (MFEB) is proposed in this paper, which contains a coarse-scale branch and a fine-scale branch for learning salient image features from the down-sampled input image and the original full resolutions, respectively. To avoid introducing too much additional computational complexity due to the fine-scale branch, in our MFEB, this branch uses only one layer, while the coarse-scale branch exploits multiple layers. The feature sets obtained from the two branches are then fused. With the MFEB, a novel multi-scale artifacts removal network (MARN) is then developed to remove JPEG image artifacts. Extensive experiments have clearly shown that our MARN can deliver superior performance to that of a number of state-of-the-art methods.
Jiahuan Ji, Baojiang Zhong, Weigang Song, Kai-Kuang Ma
ICIP1
2022 Reference-Based Jpeg Image Artifacts Removal
abstract
Since information loss incurred in lossy image compression is irreversible, it is extremely challenging to achieve a satisfactory effect of artifacts removal by using the compressed image itself only. To solve the problem, a novel reference-based artifacts removal network (RARN) is proposed in this paper, which exploits a high-quality reference image to provide useful information for facilitating the removal of artifacts and the reconstruction of details. In our RARN, a feature extraction module is first established, which takes both the compressed and reference images as inputs and produces multi-scale feature pairs as outputs. Then, a feature transfer non-local block (FTNB) is developed to match the feature pairs and transfer relevant features from the reference image to the compressed one in the feature space. Finally, image information is recovered from the multi-scale outputs of FTNB by using a reconstruction module. Extensive experimental results clearly show that our proposed RARN can deliver superior performance over a number of state-of-the-art methods.
Weigang Song, Jiahuan Ji, Baojiang Zhong
ICIP2
2022 A Direction-Decoupled Non-Local Attention Network for Single Image Super-Resolution
abstract
Thenon-local attentionmechanism has often been exploited in deep learning to capturelong-range dependencies(LRDs) from the same image for enhancing the performance of various image processing methods. However, the initially proposed non-local attention process inevitably yields extremely-high computation complexity, sinceallthe feature points are involved in computing the LRDs. To address this concern, a recently proposedcriss-cross network(CCNet), which has arecurrent criss-cross attention(RCCA) module, is used to compute the LRDs by involving only a small set of feature points for significantly reducing computation. Motivated by the RCCA, a noveldirection-decoupled non-local attention(DNA) module is proposed in this paper that is able to further reduce the computation complexity of RCCA by half approximately. To verify the performance of our new non-local attention module, a DNA network is developed for conducting single image super-resolution (SISR). Extensive experimental results have clearly demonstrated the superiority of using our DNA network for SISR when compared with that of state-of-the-art methods.
Zijiang Song, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma
IEEE Signal Process. Lett.3
2021 Single Image Super-Resolution Using Asynchronous Multi-Scale Network
abstract
An existing multi-scale residual network (MSRN) has demonstrated its success on conducting the single image super-resolution (SISR) task. The MSRN consists of a number of multi-scale residual blocks (MSRBs), and each MSRB performs convolutions by exploiting two different sizes of windows for conducting multi-scale feature extraction. The smaller window is used to extract image features at a low scale, while the larger one is used for a high scale. To significantly reduce the number of parameters involved in the MSRB, a new feature extraction module, called the asynchronous multi-scale block (AMB), is proposed in this paper. It is based on the fact that the larger window used in the MSRB can be replaced by two smaller windows without affecting the original MSRB's function. Consequently, by replacing each MSRB with our AMB, an asynchronous multi-scale network (AMNet) is then constructed, which can yield a significant reduction on computational complexity. This means that more AMBs can be used in our AMNet to deliver superior SISR performance, while maintaining the same or comparable computational complexity to that of the MSRN. To consolidate all image features generated from all scales, a new fusion scheme, called the adaptive feature fusion block (AFFB), is proposed that weights the extracted features according to their importance for further increasing SISR's performance. Extensive experimental results have clearly shown the superiority of our proposed AMNet when compared with multiple state-of-the-arts.
Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma
IEEE Signal Process. Lett.1
2020 Single Image Super-Resolution Via A Progressive Mixture Model
abstract
In this paper, a progressive mixture model (PMM) for single image super-resolution is proposed. Our model consists of an offline training stage and an online reconstruction stage, and both stages are conducted progressively by exploiting a uniform iterative scheme. In the training stage, the training dataset is clustered into finer and finer groups, and a set of mixture models are sequentially learned. In the reconstruction stage, residuals of the low-resolution (LR) input image are estimated by using the trained mixture models at increasing levels, and then they are progressively added to the LR image for producing the high-resolution (HR) image. Extensive experimental simulation results have clearly shown that the proposed model consistently delivers highly accurate and visually pleasant HR images, compared to that of the state-of the-art image super-resolution methods.
Run Su, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma
ICIP3
2020 Image Interpolation Using Multi-Scale Attention-Aware Inception Network
abstract
A new multi-scale deep learning (MDL) framework is proposed and exploited for conducting image interpolation in this paper. The core of the framework is a seeding network that needs to be designed for the targeted task. For image interpolation, a novel attention-aware inception network (AIN) is developed as the seeding network; it has two key stages: 1) feature extraction based on the low-resolution input image; and 2) feature-to-image mapping to enlarge image's size or resolution. Note that the designed seeding network, AIN, needs to be trained with a matched training dataset at each scale. For that, multi-scale image patches are generated using our proposed pyramid cut, which outperforms the conventional image pyramid method by completely avoiding aliasing issue. After training, the trained AINs are then combined for processing the input image in the testing stage. Extensive experimental simulation results obtained from seven image datasets (comprising 359 images in total) have clearly shown that the proposed MAIN consistently delivers highly accurate interpolated images.
Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma
IEEE Trans. Image Process.1
2019 Multi-Scale Defense of Adversarial Images
abstract
Deep learning has achieved great success in image classification. However, recent researches show that existing deep learning-based classifiers remain weak for recognizing adversarial images. In this paper, an effective multi-scale defense method is proposed to solve the problem. In our method, an input image is first evolved by Gaussian kernels of different intensities to generate a multi-scale representation of the image. These evolved images are then fed into a classifier trained with a multi-scale strategy to yield multi-scale confidences. Finally, an average confidence is exploited to generate classification result. Furthermore, by monitoring the change of confidence values during the image evolution process, our method is able to achieve an indication of the attacking risk. Experimental results show that our method performs favorably against a number of state-of-the-art methods.
Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma
ICIP1