Juan Wang 0012

dblp:74/3634-12 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0002-3848-9433ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Reinforcement Learning-Based Sequential Parameter Tuning for Image Signal Processing
abstract
Hardware image signal processing (ISP) transforms RAW inputs into high-quality RGB images through a series of processing modules, each with numerous tunable parameters. Traditionally, these parameters are manually tuned by imaging experts, a time-consuming and subjective process. Recent deep learning approaches predict ISP parameters, but often treat the process as a black box and overlook the intrinsic relationships among ISP modules. To address these fundamental issues, we introduce a novel ISP parameter optimization model based on single-agent reinforcement learning (RL) (i.e., SARL-ISP), formulating the hardware ISP parameter tuning as a sequential optimization problem. During the optimization process, the agent updates ISP parameter tuning strategies for different tasks through interaction with the environment. In order to explore the influence of the sequential structure of hardware ISP modules and the coupling relationships among ISP parameters on the tuning process, we further propose a sequential ISP framework based on collaborative multi-agent RL (i.e., MARL-ISP). Specifically, the serialized parameter tuning module (SPTM) realistically simulates the process of manual prediction and module pipeline. Additionally, the feature selection module (FSM) facilitates the transmission and fusion of agent features, thereby selecting more appropriate feature inputs for downstream tasks. Extensive experiments across various tasks (e.g., object detection, instance segmentation) validate the effectiveness and efficiency of our models. Even with minimal training data, our models also outperform current state-of-the-art methods in both quantitative metrics and qualitative evaluations.
Bing Li 0001, Congyan Lang, Zhikun Zhao, Juan Wang 0012, Weihua Xiong, Weiming Hu 0004, Long Cheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning
Zhikun Zhao, Congyan Lang, Bing Li 0001, Juan Wang 0012
ICCV5
2025 Accelerated Self-Supervised Multi-Illumination Color Constancy With Hybrid Knowledge Distillation
abstract
Color constancy, the human visual system's ability to perceive consistent colors under varying illumination conditions, is crucial for accurate color perception. Recently, deep learning algorithms have been introduced into this task and have achieved remarkable achievements. However, existing methods are limited by the scale of current multi-illumination datasets and model size, hindering their ability to learn discriminative features effectively and their practical value for deployment in cameras. To overcome these limitations, this paper proposes a multi-illumination color constancy approach based on self-supervised learning and knowledge distillation. This approach includes three phases: self-supervised pre-training, supervised fine-tuning, and knowledge distillation. During the pre-training phase, we train Transformer-based and U-Net based encoders by two pretext tasks: light normalization task to learn lighting color contextual representation and grayscale colorization task to acquire objects' inherent color information. For the downstream color constancy task, we fine-tune the encoders and design a lightweight decoder to obtain better illumination distributions with fewer parameters. During the knowledge distillation phase, we introduce a hybrid knowledge distillation technique to align CNN features with those of Transformer and U-Net respectively. Our proposed method outperforms state-of-the-art techniques on multi-illumination and single-illumination benchmarks. Extensive ablation studies and visualizations confirm the effectiveness of our model.
Ziyu Feng, Bing Li 0001, Congyan Lang, Zheming Xu, Haina Qin, Juan Wang 0012, Weihua Xiong
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal Processing
abstract
Hardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance metrics, which is time-consuming and biased towards human perception due to complex interaction with the output image. Since the relationship between any single parameter’s variation and the output performance metric is a complex, non-linear function, optimizing such a large number of ISP parameters is challenging. To address this challenge, we propose a novel Sequential ISP parameter optimization model, called the RL-SeqISP model, which utilizes deep reinforcement learning to jointly optimize all ISP parameters for a variety of imaging applications. Concretely, inspired by the sequential tuning process of human experts, the proposed model can progressively enhance image quality by seamlessly integrating information from both the image feature space and the parameter space. Furthermore, a dynamic parameter optimization module is introduced to avoid ISP parameters getting stuck into local optima, which is able to more effectively guarantee the optimal parameters resulting from the sequential learning strategy. These merits of the RL-SeqISP model as well as its high efficiency are substantiated by comprehensive experiments on a wide range of downstream tasks, including two visual analysis tasks (instance segmentation and object detection), and image quality assessment (IQA), as compared with representative methods both quantitatively and qualitatively. In particular, even using only 10% of the training data, our model outperforms other SOTA methods by an average of 7% mAP on two visual analysis tasks.
Zhikun Zhao, Congyan Lang, Mingxuan Cai, Longfei Han, Juan Wang 0012, Bing Li 0001
AAAI7
2024 PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
Zewen Chen, Haina Qin, Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Liang Wang 0001
ECCV (1)3
2023 Learning to Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines Optimization
abstract
The hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subtask and contains a set of tunable hyperparameters. A large number of hyperparameters form a complex mapping with the ISP output. The industry typically relies on manual and time-consuming hyperparameter tuning by image experts, biased towards human perception. Recently, several automatic ISP hyperparameter optimization methods using downstream evaluation metrics come into sight. However, existing methods for ISP tuning treat the high-dimensional parameter space as a global space for optimization and prediction all at once without inducing the structure knowledge of ISP. To this end, we propose a sequential ISP hyperparameter prediction framework that utilizes the sequential relationship within ISP modules and the similarity among parameters to guide the model sequence process. We validate the proposed method on object detection, image segmentation, and image quality tasks.
Haina Qin, Longfei Han, Weihua Xiong, Juan Wang 0012, Bing Li 0001, Weiming Hu 0004
CVPR4
2023 Hierarchical Curriculum Learning for No-Reference Image Quality Assessment
Juan Wang 0012, Zewen Chen, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004
Int. J. Comput. Vis.1
2023 DeformSg2im: Scene graph based multi-instance image generation with a deformable geometric layout
Yuxiao Li 0001, Danlan Huang, Juan Wang 0012, Ning Ge 0001, Jianhua Lu
Neurocomputing4
2023 Self-Prior Guided Pixel Adversarial Networks for Blind Image Inpainting
abstract
Blind image inpainting involves two critical aspects, i.e., "where to inpaint" and "how to inpaint". Knowing "where to inpaint" can eliminate the interference arising from corrupted pixel values; a good "how to inpaint" strategy yields high-quality inpainted results robust to various corruptions. In existing methods, these two aspects usually lack explicit and separate consideration. This paper fully explores these two aspects and proposes a self-prior guided inpainting network (SIN). The self-priors are obtained by detecting semantic-discontinuous regions and by predicting global semantic structures of the input image. On the one hand, the self-priors are incorporated into the SIN, which enables the SIN to perceive valid context information from uncorrupted regions and to synthesize semantic-aware textures for corrupted regions. On the other hand, the self-priors are reformulated to provide a pixel-wise adversarial feedback and a high-level semantic structure feedback, which can promote the semantic continuity of inpainted images. Experimental results demonstrate that our method achieves state-of-the-art performance in metric scores and in visual quality. It has an advantage over many existing methods that assume "where to inpaint" is known in advance. Extensive experiments on a series of related image restoration tasks validate the effectiveness of our method in obtaining high-quality inpainting.
Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 HiMoReNet: A Hierarchical Model for Human Motion Refinement
abstract
3D human pose estimation has a broad range of applications, including anomaly detection and animation creation. Despite that significant progress on relative research has been made during the past decades, producing precise and smooth estimations for input videos still remains challenging mainly because of its ill-posed attributes. In this paper, we propose HiMoReNet, a post-processing motion refinement neural network based on an elaborate hierarchical architecture. Firstly, we distinguish characteristic motion patterns of joints at different locations by grouping the joints and employing respective spatiotemporal processing modules for each group. In addition, by mimicking interactions among multiple body parts, global context information is leveraged to further guide the motion refinement. Quantitative and qualitative results on the 3DPW dataset demonstrate that our proposed HiMoReNet achieves the state-of-the-art performance, and excels in jitter removal and precise pose estimation.
Juan Wang 0012, Ning Ge 0001, Jianhua Lu
IEEE Signal Process. Lett.2
2022 Teacher-Guided Learning for Blind Image Quality Assessment
Zewen Chen, Juan Wang 0012, Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004
ACCV (3)2
2022 Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines
Haina Qin, Longfei Han, Juan Wang 0012, Congxuan Zhang, Bing Li 0001, Weiming Hu 0004
ECCV (19)3
2021 Semantic Perceptual Image Compression With a Laplacian Pyramid of Convolutional Networks
abstract
The existing image compression methods usually choose or optimize low-level representation manually. Actually, these methods struggle for the texture restoration at low bit rates. Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. To achieve better perceptual quality, generative models are widely used, especially generative adversarial networks (GAN). However, training GAN is intractable, especially for high-resolution images, with the challenges of unconvincing reconstructions and unstable training. To overcome these problems, we propose a novel DNN-based image compression framework in this paper. The key point is decomposing an image into multi-scale sub-images using the proposed Laplacian pyramid based multi-scale networks. For each pyramid scale, we train a specific DNN to exploit the compressive representation. Meanwhile, each scale is optimized with different aspects, including pixel, semantics, distribution and entropy, for a good "rate-distortion-perception" trade-off. By independently optimizing each pyramid scale, we make each stage manageable and make each sub-image plausible. Experimental results demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in providing improved visual quality. Additionally, a better performance in the down-stream visual analysis tasks which are conducted on the reconstructed images, validates the excellent semantics-preserving ability of the proposed method.
Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
IEEE Trans. Image Process.1
2020 Local-to-Global Semantic Supervised Learning for Image Captioning
abstract
Image captioning is a challenging problem owing to the complexity of image content and the diverse ways of describing the content in natural language. Although current methods have made substantial progress in terms of objective metrics (such as BLEU, METEOR, ROUGE-L and CIDEr), there still exist some problems. Specifically, most of these methods are trained to maximize the log-likelihood or objective metrics. As a result, these methods often generate rigid and semantically incomplete captions. In this paper, we develop a new model that aims to generate captions conforming to human evaluation. The core idea is to use local-to-global semantic supervised learning by introducing the two-level optimization objective functions. At the word level, we match each word to the image regions using the local attention objective function; at the sentence level, we align the entire sentence and the image using the global semantic objective function. Experimentally, we compare the proposed model with current methods on MSCOCO dataset. We show that either local attention supervision or global semantic supervision is the necessary component for the success of our model through ablation studies. Furthermore, combining these two supervision objective functions achieves state-of-the-art performance in terms of both standard evaluation metrics and human judgment.
Juan Wang 0012, Yiping Duan, Xiaoming Tao 0001, Jianhua Lu
ICC1
2019 Semantic Perceptual Image Compression with a Laplacian Pyramid of Convolutional Networks
abstract
Recently, deep neural network (DNN)-based image compression methods have achieved impressive results. These methods generally use thumbnail images or crop small patches from high-resolution images to train their networks. Instead of using patch-based training mode, we propose a novel DNN-based image compression framework in this paper. We apply the Laplacian pyramid to construct a multi-scale image representation. By learning the increasingly detailed representations, the proposed method is able to progressively restore an image. Furthermore, we use the adversarial networks for training to encourage the perceptual quality of the reconstructed image. Particularly at low bitrates, our model can only store the global semantics of an image and automatically synthesize the texture to achieve high subjective quality. Experimental results on demonstrate that our method achieves state-of-the-art performance, with advantages over existing methods in terms of visual quality.
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
ICIP1
2018 Boundary Objectness Network for Object Detection and Localization
abstract
In this paper, we present the boundary objectness network (BON), an effective convolutional neural network (CNN) for object detection. Its core contribution is to accurately localize the objects. Generally, the CNN-based localizers predict four bounding box coordinates by learning a regression function. This method shows a low Intersection-of-Union (IoU) with the ground truth box. In our work, the localization is formu-lated as a probabilistic problem. Specifically, the deep features inside the candidate proposal are mapped into a row and a column feature vector, which are called boundary object-ness. The boundary objectness indicates the existence of an object in the horizontal and vertical direction of the proposal, enabling us to elaborately localize the object. Moreover, the modules of object detection share the common convolution-al layers. Meanwhile, a multi-task loss function is designed for joint training strategy. Experimental results on the PAS-CAL VOC datasets demonstrate the competitive performance of our method. For the VGG16 model, we achieve 77.6 % mAP at a speed of 4 frame per second (FPS), thus having the potential for real-time processing.
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Jianhua Lu
ICASSP1
2018 Hierarchical objectness network for region proposal generation and object detection
Juan Wang 0012, Xiaoming Tao 0001, Mai Xu, Yiping Duan, Jianhua Lu
Pattern Recognit.1
2017 Prior-Information-Based Remote Sensing Image Compression with Bayesian Dictionary Learning
abstract
Requirements for higher resolution remote sensing images lead to rapid increase of data amount in space communications. However, since satellite communications capacity is suffering from great pressure, seeking for more effective compression scheme is supposed to solve existing conflict between tremendous data and limited bandwidth. For this reason, this paper proposes a prior-information-based remote sensing image compression scheme. We firstly utilize prior information contained in historical remote sensing images for incremental image extraction, which is assumed to have removed redundant information possessed both on the satellite and ground. Moreover, Bayesian dictionary serves to sparsely represent the incremental image, generating finite number of representation coefficients in place of numerous pixels. Finally, quantization and encoding schemes are further designed for efficient data transmission. Experimental results show that the proposed scheme is competitive to existing general image compression schemes.
Xiaoming Tao 0001, Shaoyang Li, Zizhuo Zhang, Xijia Liu, Juan Wang 0012, Jianhua Lu
VTC Spring5
2016 Quantization and Entropy Coding Scheme for Dictionary Learning Based Image Compression
abstract
Most recently, there has been a growing interest in the study of dictionary learning based (DL-based) image compression, which has potential in relieving the bandwith-hungry bottleneck of visual communication. All existing DL-based image compression approaches mainly focus on the effective representation of images, thus losing sight of two basic elements of image compression, i.e., quantization and entropy coding. For this reason, this paper proposes a quantization and entropy coding scheme for DL-based image compression. In our scheme, the proposed Partition-Interval K-means (PIK) quantizer adaptively maps continuous coefficients to discrete values. The arithmetic coding, combined with differential coding technique, is applied to encode the indices of nonzero coefficients as well as the labels of quantization values. In our experiments, the proposed scheme is verified to be more effective than other quantization and entropy coding schemes for DL-based image compression.
Juan Wang 0012, Xiaoming Tao 0001, Xijia Liu, Ning Ge 0001, Jianhua Lu
VTC Fall1
2015 A New Radiation Correction Method for Remote Sensing Images Based on Change Detection
Juan Wang 0012, Xijia Liu, Xiaoming Tao 0001, Ning Ge 0001
ICIG (1)1