Wanjie Sun

dblp:138/6873 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-7781-4542ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
abstract
Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a large-scale and richly annotated visual question answering (VQA) dataset specifically tailored for roadside scenarios. The dataset comprises 34,736 diverse QA pairs collected under varying weather, illumination, and traffic conditions, targeting not only object attributes but also the intent, legality, and interaction patterns of traffic participants. RoadSceneVQA challenges models to perform both explicit recognition and implicit commonsense reasoning, grounded in real-world traffic rules and contextual dependencies. To fully exploit the reasoning potential of Multi-modal Large Language Models (MLLMs), we further propose CogniAnchor Fusion (CAF), a vision-language fusion module inspired by human-like scene anchoring mechanisms. CAF enables precise and efficient cross-modal interaction. Moreover, we propose the Assisted Decoupled Chain-of-Thought (AD-CoT) to enhance the reasoned thinking via CoT prompting and multi-task learning. Experimental results on RoadSceneVQA and CODA-LM benchmark show that the pipeline consistently improves both reasoning accuracy and computational efficiency, allowing the MLLM to achieve state-of-the-art performance in structural traffic perception and reasoning tasks.
Runwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao, Ziren Tang, Ningwei Ouyang, Shaofeng Liang, Yuxuan Fan, Wanjie Sun, Yutao Yue
AAAI12
2025 Timestep-Aware Diffusion Model for Extreme Image Rescaling
Wanjie Sun, Zhenzhong Chen 0001
ICCV3
2025 Lightweight image super-resolution with sliding Proxy Attention Network
Wanjie Sun, Zhenzhong Chen 0001
Signal Process.2
2025 Unifying Dimensions: A Linear Adaptive Mixer for Lightweight Image Super-Resolution
abstract
Window-based Transformers have demonstrated outstanding performance in super-resolution due to their adaptive modeling capabilities through local self-attention (SA). However, they exhibit higher computational complexity and inference latency than convolutional neural networks. In this paper, we first identify that the adaptability of the Transformers is derived from their adaptive spatial aggregation and advanced structural design, while their high latency results from the computational costs and memory layout transformations. To address these limitations and simulate the aggregation approach, we propose an efficient convolution-based Focal Separable Attention (FSA) mechanism that enables long-range dynamic modeling with linear computational complexity. Additionally, we introduce a dual-branch structure integrated with an ultra-lightweight Information Exchange Module (IEM) to enhance information aggregation within the token mixing process. Finally, we modify the existing spatial-gate-based feedforward neural networks by incorporating a self-gate mechanism to preserve high-dimensional channel information, thereby enabling the modeling of more complex relationships. This modification is referred to as the Dual-Gated Feed-Forward Network (DGFN). With these advancements, we construct a convolution-based Transformer framework named the Linear Adaptive Mixer Network (LAMNet). Extensive experiments demonstrate that LAMNet performs better than existing Transformer-based methods while maintaining the computational efficiency of convolutional neural networks, which can achieve a speedup $3\times $ of the inference time. The code will be publicly available at: https://github.com/zononhzy/LAMNet.
Wanjie Sun
IEEE Trans. Image Process.2
2024 Learning Many-to-Many Mapping for Unpaired Real-World Image Super-Resolution and Downscaling
abstract
Learning based single image super-resolution (SISR) for real-world images has been an active research topic yet a challenging task, due to the lack of paired low-resolution (LR) and high-resolution (HR) training images. Most of the existing unsupervised real-world SISR methods adopt a two-stage training strategy by synthesizing realistic LR images from their HR counterparts first, then training the super-resolution (SR) models in a supervised manner. However, the training of image degradation and SR models in this strategy are separate, ignoring the inherent mutual dependency between downscaling and its inverse upscaling process. Additionally, the ill-posed nature of image degradation is not fully considered. In this paper, we propose an image downscaling and SR model dubbed as SDFlow, which simultaneously learns a bidirectional many-to-many mapping between real-world LR and HR images unsupervisedly. The main idea of SDFlow is to decouple image content and degradation information in the latent space, where content information distribution of LR and HR images is matched in a common latent space. Degradation information of the LR images and the high-frequency information of the HR images are fitted to an easy-to-sample conditional distribution. Experimental results on real-world image SR datasets indicate that SDFlow can generate diverse realistic LR and SR images both quantitatively and qualitatively.
Wanjie Sun, Zhenzhong Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Learned Scale-Arbitrary Image Downscaling for Non-Learnable Upscaling
abstract
In this letter, we propose a learned scale-arbitrary image downscaling method to downscale high-resolution (HR) images to a target low-resolution (LR) where it could be well upscaled by traditional simple upscaling methods. Specifically, the scale information is first fused into the feature extraction process of the input HR image by employing our proposed scale-adaptive feature enhancement module (FEM). Then the proposed scale-arbitrary feature resampler (SAFR) adaptively generates sampling weights which are applied to the resampled candidate feature maps to produce the downscaled LR image. Experimental results demonstrate that when compared to the existing learned image downscaling method for non-learnable upscaling, the reconstructed HR images upscaled by our proposed method receive better quality..
Chengrui Huang 0003, Wanjie Sun, Zhenzhong Chen 0001
IEEE Signal Process. Lett.2
2022 Joint image demosaicking and denoising with mutual guidance of color channels
Wanjie Sun
Signal Process.2
2022 Learning Discrete Representations From Reference Images for Large Scale Factor Image Super-Resolution
abstract
Image super-resolution (SR) task aims to recover high-resolution (HR) images from degraded low-resolution (LR) images, which has achieved great progress due to the recent advances of deep neural networks. Due to severe information loss of the LR images, it is more challenging to reconstruct high quality HR images at large scale factors, i. e., higher than 4× . Traditional reference image based SR methods usually perform patch matching to locate detailed texture from HR reference images which could provide fine details from similar image contents. But it suffers from difficulties in achieving good matching in the largely downscaled image space or feature space due to the ill-posed nature between LR and HR mapping. In this paper, we tackle this problem by exploiting fine details contained in reference HR images. Inspired by vector quantization (VQ), we propose a simple yet effective auto-encoder convolutional neural network (CNN) module to learn discrete representations of images. Furthermore, we propose to progressively learn pairs of cross-scale discrete feature representations using paired LR and HR reference images. The coarser scale of the discrete representation is responsible for encoding the global image structure while the paired finer scale of the discrete representation takes charge of capturing missing details in the finer image scale. During inference, continuous features of the test LR image are used as queries to retrieve finer scale discrete representations (value) by searching the nearest coarser scale discrete representations (key). Then, the queries and retrieved values are combined to progressively recover the HR image. Experimental results indicate that when compared with the state-of-the-art image SR models, the proposed method can achieve advanced performance in terms of both objective quality and subjective quality. The code will be available on URL: https://github.com/sunwj/refsr.
Wanjie Sun, Zhenzhong Chen 0001
IEEE Trans. Image Process.1
2021 Visual Scanpath Prediction Using IOR-ROI Recurrent Mixture Density Network
abstract
A visual scanpath represents the human eye movements when scanning the visual field for acquiring and receiving visual information. Predicting visual scanpaths when a certain stimulus is presented plays an important role in modeling overt human visual attention and search behavior. In this paper, we presented an 'Inhibition of Return - Region of Interest' (IOR-ROI) recurrent mixture density network based framework learning to produce human-like visual scanpaths under task-free viewing conditions. The proposed model simultaneously predicts a sequence of ordered fixation positions and their corresponding fixation durations. Our model integrates bottom-up features and semantic features extracted by convolutional neural networks. Then the integrated feature maps are fed into the IOR-ROI Long Short-Term Memory (LSTM) which is the core component of the proposed model. The IOR-ROI LSTM is a dual LSTM unit, i.e., the IOR-LSTM and the ROI-LSTM, capturing IOR dynamics and gaze shift behavior simultaneously. IOR-LSTM simulates the visual working memory to adaptively maintain and update visual information regarding previously fixated regions. ROI-LSTM is responsible for predicting the next possible ROIs given the spatially inhibited image feature maps on the feature-wise basis. Fixation duration is predicted by a regression neural network given the viewing history and image feature maps corresponding to currently fixated ROI. Considering the eye movement pattern variations among subjects, a mixture density network is adopted to model the next fixation distribution as Gaussian mixtures and the fixation duration is also modeled using Gaussian distribution. Our model is evaluated on the OSIE and MIT low resolution eye-tracking datasets and experimental results indicate that the proposed method can achieve superior performance in predicting visual scanpaths. The code will be publicly available on URL: https://github.com/sunwj/scanpath.
Wanjie Sun, Zhenzhong Chen 0001, Feng Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2020 DOVE: Decomposition Oriented Video super-rEsolution
abstract
Video super-resolution (VSR) has attracted a lot of attention that converts a low resolution (LR) video into a high resolution (HR) one. The original LR video is typically produced either by the downscaling processing or low-resolution sensor. Considering that the resolution degradation or limitation makes different impacts on different low-frequency (LF) and high-frequency (HF) components of the LR video signal, we propose a Decomposition Oriented Video super-rEsolution (DOVE) method in this paper. More specifically, a three-stream VSR network is designed in which the proposed LF and HF stream is responsible for modeling LF and HF components in the feature space. Moreover, a multi-frame refinement stream takes features of coarsely aligned frames as input and generates finely aligned counterparts progressively to guide the learning of LF and HF streams at the intermediate feature level. Furthermore, non-local channel attention is devised to capture long-range dependencies on a global scale both in the channel domain. Experimental results indicate that separating the learning of LF and HF components helps better estimate the HR frame from LR frames and superior VSR performance is achieved when compared with that of recent state-of-the-art methods.
Huairui Wang, Wanjie Sun, Daiqin Yang
VCIP2
2020 Learned Image Downscaling for Upscaling Using Content Adaptive Resampler
abstract
Deep convolutional neural network based image super-resolution (SR) models have shown superior performance in recovering the underlying high resolution (HR) images from low resolution (LR) images obtained from the predefined downscaling methods. In this paper, we propose a learned image downscaling method based on content adaptive resampler (CAR) with consideration on the upscaling process. The proposed resampler network generates content adaptive image resampling kernels that are applied to the original HR input to generate pixels on the downscaled image. Moreover, a differentiable upscaling (SR) module is employed to upscale the LR result into its underlying HR counterpart. By back-propagating the reconstruction error down to the original HR input across the entire framework to adjust model parameters, the proposed framework achieves a new state-of-the-art SR performance through upscaling guided image resamplers which adaptively preserve detailed information that is essential to the upscaling. Experimental results indicate that the quality of the generated LR image is comparable to that of the traditional interpolation based method and the significant SR performance gain is achieved by deep SR models trained jointly with the CAR model. The code is publicly available on: https://github.com/sunwj/CAR.
Wanjie Sun, Zhenzhong Chen 0001
IEEE Trans. Image Process.1
2018 Scanpath Prediction for Visual Attention using IOR-ROI LSTM
abstract
Predicting scanpath when a certain stimulus is presented plays an important role in modeling visual attention and search. This paper presents a model that integrates convolutional neural network and long short-term memory (LSTM) to generate realistic scanpaths. The core part of the proposed model is a dual LSTM unit, i.e., an inhibition of return LSTM (IOR-LSTM) and a region of interest LSTM (ROI-LSTM), capturing IOR dynamics and gaze shift behavior simultaneously. IOR-LSTM simulates the visual working memory to adaptively integrate and forget scene information. ROI-LSTM is responsible for predicting the next ROI given the inhibited image features. Experimental results indicate that the proposed architecture can achieve superior performance in predicting scanpaths.
Wanjie Sun
IJCAI2