VLDB 2026 Research / reviewers in the wild / expert
Juan Zhang 0001
dblp:14/2573-1
· DBLP profile ↗
19ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0003-1558-1656ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Continual few-shot relation extraction via multi-task balanced dual-branch network
Chenyang Shan, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014 |
Neurocomputing | 2 |
| 2026 | MAFIFusion: a multi-attention and feature interaction network for infrared and visible image fusion
Haochen Yu, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao, Bo Huang 0014, Yadong Zhu |
Multim. Syst. | 2 |
| 2025 | Multi-query network learning for hoi detection via vision-language models
Leideng Shi, Juan Zhang 0001 |
Appl. Intell. | 2 |
| 2025 | Multimodal-Aware Fusion Network for Referring Remote Sensing Image SegmentationabstractReferring remote sensing image segmentation (RRSIS) is a novel visual task in remote sensing images segmentation, which aims to segment objects based on a given text description, with great significance in practical application. Previous studies fuse visual and linguistic modalities by explicit feature interaction, which fail to effectively excavate useful multimodal information from dual-branch encoder. In this letter, we design a multimodal-aware fusion network (MAFN) to achieve fine-grained alignment and fusion between the two modalities. We propose a correlation fusion module (CFM) to enhance multiscale visual features by introducing adaptive noise in transformer, and integrate cross-modal aware features. In addition, MAFN employs multiscale refinement convolution (MSRC) to adapt to the various orientations of objects at different scales to boost their representation ability to enhances segmentation accuracy. Extensive experiments have shown that MAFN is significantly more effective than the state of the art (SOTA) on RRSIS-D datasets. The source code is available athttps://github.com/Roaxy/MAFN. Leideng Shi, Juan Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | ECKT: enhancing cross-task knowledge transfer in continual few-shot relation extraction
Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao |
J. Supercomput. | 2 |
| 2025 | Depth-guided color correction and multi-scale Retinex network for underwater image enhancement
Zhan Hu, Juan Zhang 0001, Yongbin Gao, Bo Huang 0014, Zhijun Fang 0001 |
Vis. Comput. | 2 |
| 2024 | A lightweight RGB superposition effect adjustment network for low-light image enhancement and denoising
Pei-Dong Chen, Juan Zhang 0001, Yongbin Gao, Zhijun Fang 0001, Jenq-Neng Hwang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Self-Enhanced Attention for Image CaptioningabstractAbstract Image captioning, which involves automatically generating textual descriptions based on the content of images, has garnered increasing attention from researchers. Recently, Transformers have emerged as the preferred choice for the language model in image captioning models. Transformers leverage self-attention mechanisms to address gradient accumulation issues and eliminate the risk of gradient explosion commonly associated with RNN networks. However, a challenge arises when the input features of the self-attention mechanism belong to different categories, as it may result in ineffective highlighting of important features. To address this issue, our paper proposes a novel attention mechanism called Self-Enhanced Attention (SEA), which replaces the self-attention mechanism in the decoder part of the Transformer model. In our proposed SEA, after generating the attention weight matrix, it further adjusts the matrix based on its own distribution to effectively highlight important features. To evaluate the effectiveness of SEA, we conducted experiments on the COCO dataset, comparing the results with different visual models and training strategies. The experimental results demonstrate that when using SEA, the CIDEr score is significantly higher compared to the scores obtained without using SEA. This indicates the successful addressing of the challenge of effectively highlighting important features with our proposed mechanism. Qingyu Sun, Juan Zhang 0001, Zhijun Fang 0001, Yongbin Gao |
Neural Process. Lett. | 2 |
| 2024 | Self-Supervised Learning of Depth and Ego-Motion for 3D Perception in Human Computer Interactionabstract3D perception of depth and ego-motion is of vital importance in intelligent agent and Human Computer Interaction (HCI) tasks, such as robotics and autonomous driving. There are different kinds of sensors that can directly obtain 3D depth information. However, the commonly used Lidar sensor is expensive, and the effective range of RGB-D cameras is limited. In the field of computer vision, researchers have done a lot of work on 3D perception. While traditional geometric algorithms require a lot of manual features for depth estimation, Deep Learning methods have achieved great success in this field. In this work, we proposed a novel self-supervised method based on Vision Transformer (ViT) with Convolutional Neural Network (CNN) architecture, which is referred to as ViT-Depth . The image reconstruction losses computed by the estimated depth and motion between adjacent frames are treated as supervision signal to establish a self-supervised learning pipeline. This is an effective solution for tasks that need accurate and low-cost 3D perception, such as autonomous driving, robotic navigation, 3D reconstruction, and so on. Our method could leverage both the ability of CNN and Transformer to extract deep features and capture global contextual information. In addition, we propose a cross-frame loss that could constrain photometric error and scale consistency among multi-frames, which lead the training process to be more stable and improve the performance. Extensive experimental results on autonomous driving dataset demonstrate the proposed approach is competitive with the state-of-the-art depth and motion estimation methods. Shanbao Qiao, Naixue Xiong, Yongbin Gao, Zhijun Fang 0001, Juan Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Spatial graph attention network-based object tracking with adaptive cosine window
Liuyi Fan, Bo Huang 0014, Juan Zhang 0001, Yongbin Gao |
Appl. Intell. | 4 |
| 2023 | MACFNet: multi-attention complementary fusion network for image denoising
Jiaolong Yu, Juan Zhang 0001, Yongbin Gao |
Appl. Intell. | 2 |
| 2023 | Single-image Deraining via a channel memory network
Yan Zhang 0116, Juan Zhang 0001 |
Appl. Intell. | 4 |
| 2023 | Corrigendum to "Single-image deraining via a Recurrent Memory Unit Network" [Knowl.-Based Syst. 218 (2021) 106832]
Yan Zhang 0116, Juan Zhang 0001, Bo Huang 0014, Zhijun Fang 0001 |
Knowl. Based Syst. | 2 |
| 2022 | Multiple Patch-Aware Network for Faster Real-World Image DehazingabstractThis paper proposes a Multiple Patch-aware Dehazing Network (MPADN), to remove haze in real-world images fast and efficiently. Firstly, we design a Multiple Patch-aware Module (MPAM), which utilizes joint decisions from multiple patch awareness to gain more stable local features. Compared with the multi-patch hierarchical strategy, MPAM extremely compresses the size of the model. Besides, we propose a novel data enhancement method called Concentration Sampling Enhancement (CSE), which generates new training samples by haze concentration sampling based on hazy images and clear images. This method makes use of the precious limited real-world paired data to generate more potentially usable data. It is easy to extend CSE to more models or tasks where the ground truth and the input information are similar. Experiments show that MPADN outperforms state-of-art fast dehazing methods, and it only takes 0.0065s to process a 1600x1200 image. Juan Zhang 0001, Xiaoqi Lang |
ICASSP | 2 |
| 2022 | Double-Branch Dehazing Network based on Self-Calibrated Attentional Convolution
Juan Zhang 0001, Jenq-Neng Hwang, Bo Huang 0014 |
Knowl. Based Syst. | 2 |
| 2021 | The analysis of isolation measures for epidemic control of COVID-19
Bo Huang 0014, Yongbin Gao, Guohui Zeng, Juan Zhang 0001, Jin Liu 0016 |
Appl. Intell. | 5 |
| 2021 | Single-image deraining via a Recurrent Memory Unit Network
Yan Zhang 0116, Juan Zhang 0001, Bo Huang 0014, Zhijun Fang 0001 |
Knowl. Based Syst. | 2 |
| 2019 | DD-CycleGAN: Unpaired image dehazing via Double-Discriminator Cycle-Consistent Generative Adversarial Network
Jingming Zhao, Juan Zhang 0001, Zhi Li 0049, Jenq-Neng Hwang, Yongbin Gao, Zhijun Fang 0001, Bo Huang 0014 |
Eng. Appl. Artif. Intell. | 2 |
| 2019 | Unsupervised learning of depth estimation based on attention model and global pose optimization
Renyue Dai, Yongbin Gao, Zhijun Fang 0001, Anjie Wang, Juan Zhang 0001, Cengsi Zhong |
Signal Process. Image Commun. | 6 |