Jingwei Xin

dblp:175/1662 · DBLP profile ↗
← Back
29ranked-venue papers
13as first author
22since 2021 · last 2026
0000-0001-9551-6007ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal
abstract
Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct structured information beneath cloud-covered areas. In this work, we propose a Visibility-guided Semantic Estimation and Reconstruction network for cloud removal (VISER-CR), which reformulates CR as a structure-guided completion problem. Specifically, VISER-CR explicitly models cloud interference via spatial masking, encouraging the model to reason beyond pixel-level appearance and enhance scene-level structural understanding. Moreover, to further improve the representation of structural information, we introduce Patch Saliency Encoding, a self-guided mechanism that implicitly models structural alignment among patches, significantly enhancing clustering consistency and semantic separability in the latent space. This adaptive mechanism guides the network to focus on learning and reconstructing structurally important regions, thereby reducing redundancy and improving overall cloud removal performance. Extensive experiments on multiple benchmark datasets demonstrate the superior effectiveness of our method.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001
AAAI1
2026 One-step diffusion-based real-world image super-resolution with visual perception distillation
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
Neurocomputing2
2026 IMEVSI: Online Adaptive Video Stream Interpolation via Inertia-Aware Motion Estimation
abstract
Recent video frame interpolation (VFI) methods rely on computationally heavy modules (e.g., global attention module) to handle large motions, incurring prohibitive costs which hinders their practical real-time deployment. In this work, we revisit the core objective of VFI: enhancing the temporal resolution of videos. We identify that previous VFI’s frame-isolated processing ignores continuous temporal modeling, introduces computational redundancy in video streaming scenarios. To address this, we propose an online learning recurrent net with inertia-aware motion estimation(IMEVSI). It consists of implicit motion propagation( IMP), explicit motion propagation(EMP) and adaptive online learning strategy(AOL). IMP and EMP are used to high order inter-frame motion modeling considering motion inertia, AOL are proposed to bridge the motion domain gap between training and deployment. For IMP, we initiate from explicit physical motion modeling, progressively integrating learnable parameters into inertia-ware motion extraction and finally unify motion propagation and extraction within our recurrent motion propagation Transformer(RMPT). For EMP, we directly inject adjacent motion into current flow estimation recognizing its inertia contribution. For AOL,we leverage cycle consistency to dynamically adjust intermediate flow estimator and maintains an adaptive threshold to control parameter update. Extensive experiments demonstrate that our method outperforms state-of-the-art (SOTA) approaches on regular, large-motion, and high-resolution benchmarks while achieving excellent inference speed and FLOPs.
Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Video Frame Interpolation via Appearance-Based Intermediate Flow Estimation
abstract
Intermediate flow estimation is an important part of video frame interpolation (VFI). Most previous works use interpolation to derive the intermediate flow assuming localized linear motion. However, this method is not effective when dealing with extreme motions. In this work, we assume that the motion trajectory of an object is determined by the appearance characteristics of this object. Based on this assumption, we propose a new intermediate flow estimation method, which obtains the motion features of intermediate frames from image appearance and inter-frame motion features. In addition, in order to fully extract the inter-frame features, we rethink the difference of VFI and previous works on using Swin-Transformer and compute the appearance features and motion features within the adaptive neighborhood by cyclically shifting the window. Experimental results show that our method achieves state-of-the-art performance on different datasets for both fixed-time and arbitrary-time interpolation. Moreover, our proposed method outperforms models that require inputting a sequence of four frames when handling videos with extremely large motion. The source code is available from https://github.com/chen12304/IFE-VFI.
Keyi Chen 0015, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.2
2026 Interpretable General Image Fusion via Scalable Autoregressive Modeling
abstract
Existing image fusion methods have developed increasingly sophisticated network architectures for exploiting modality-shared and modality-specific features. However, despite these advancements in feature extraction, most methods ultimately rely on relatively simple implicit or explicit fusion strategies, which can compromise interpretability and limit fusion accuracy. In this paper, we incorporate visual autoregressive modeling to bridge the gap between implicit feature extraction and explicit modality fusion. First, the proposed approach conducts a low-to-high resolution autoregressive objective with modality-specific features, introducing a scalable feature autoregressive mechanism. It aggregates local and global contextual dependencies while enhancing implicit cross-scale interaction. Furthermore, to promote the consistency and complementarity across modalities, we embed an explicit high-order fusion strategy within the progressive modality-specific feature extraction process. This integration facilitates a next-scale synergistic relationship between implicit learning and explicit fusion. Our High-order Feature AutoRegressive Fusion framework (HFARFusion) provides a robust and interpretable solution for general image fusion tasks, effectively balancing fusion performance and transparency through the strengths of autoregressive learning. Extensive experiments demonstrate the outstanding performance of the proposed method in several classical fusion tasks, including infrared-visible, medical, multi-focus, and multi-exposure image fusion. Our code is available at https://github.com/happysbn/HFARFusion.
Jingwei Xin, Boneng Shi, Zhen Li 0026, Xuehao Song, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 Unsupervised Face Super-Resolution via Integrating Faithful 3D Facial Priors
abstract
Recently, unsupervised face super-resolution (FSR) has attracted significant attention due to its remarkable generalization performance. However, existing methods neglect the incorporation of facial priors, which can effectively guide the restoration of face images. The root cause of this issue lies in the significant challenges associated with incorporating facial priors into unsupervised frameworks. First, unsupervised methods often face the challenge of real-world low-quality (LQ) images that are severely corrupted, making it unrealistic to extract reliable prior information from them. Second, the estimation of facial priors exponentially increases the model’s parameters and computational complexity, contradicting the purpose of unsupervised methods for practical deployment. In this work, we fundamentally address the aforementioned challenges and proposeFaith3D-FSR, a novel approach that incorporates faithful 3D facial priors into unsupervised FSR. Specifically, we introduceFaith3Dmechanism for faithful prior integration, which deconstructs super-resolution images into 3D elements and uses the 3D priors from real high-quality (HQ) images as reference for calibration solely during the training phase. This strategy enables more precise guidance on the super-resolution in a high-dimensional space, without requiring additional prior estimation during inference. It successfully overcomes the aforementioned challenges, making it more suitable for real-world applications, and offers a plug-and-play solution for incorporating 3D priors into unsupervised FSR. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on multiple benchmark datasets and across a range of evaluation metrics. The code is available here.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 MVFusion: Generative Representation Learning With Masked Variational Autoencoders for Multi-Modality Image Fusion
abstract
Creating a comprehensively representative image while maintaining the merits of various modalities is a key focus of current Multi-Modality Image Fusion research. Existing unified methods often struggle to handle varying types of degradation while extracting modality-shared and modality-specific information from source images, leading to limitations in their generative or representation capabilities under different conditions. To address the challenge, we propose MVFusion, a novel self-supervised masked variational autoencoder framework that simultaneously enhances generative training and representation learning. It is designed to cope with varying image quality and dataset composition with a unified framework while ensuring effective fusion of modality information. Specifically, MVFusion employs a self-supervised masked autoencoder to reduce the impact of redundancy and degradation in the source images, and thus learns the latent distribution of degraded input images in the generative training stage. In addition, we incorporate variational feature learning to further preserve the distinctive modality features in the representation learning stage. Extensive experiments demonstrate that our model achieves promising results in several classical fusion tasks, including infrared-visible, multi-focus, multi-exposure, and medical image fusion. The code is available at https://github.com/shiboneng/MVFusion.
Jingwei Xin, Boneng Shi, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2025 I²NQ: Inter and Intra Nonuniform Quantization for Single Image Super-Resolution
abstract
Quantizing neural network is an efficient model compression technique that converts weights and activations from floating-point to integer. However, existing model quantization methods are primarily designed for high-level visual tasks. They do not sufficiently consider the unique characteristics of feature distribution in image super-resolution (SR) reconstruction models. On the one hand, the objective of SR is to restore high-frequency and fine-detail information while preserving the overall feature distribution. Therefore, the regularization techniques are removed to maintain the original distribution. However, vanilla quantization methods often employ regularization techniques to normalize the features for stable network training, which destroys the inherent information of the feature distribution. On the other hand, the feature distribution in SR models exhibits a nonuniform bell-shaped form. Common quantization methods adopt a uniform quantization strategy with equal quantization intervals. This fails to effectively capture the nonuniform feature distribution in SR. To address the above issue, we propose a novel method named Inter and Intra Nonuniform Quantization, which takes into account the specific characteristics of the feature distribution in the context of SR reconstruction models. Additionally, we propose a weight adjustment method called flex-scale-weight-adjust (FSWA). It can maintain the diversity of weight information and reduce quantization errors. Extensive experiments demonstrate that our proposed method surpasses other quantization methods in both the evaluation of reconstruction metrics and visual reconstruction performance.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2025 Rectified Binary Network for Single-Image Super-Resolution
abstract
Binary neural network (BNN) is an effective approach to reduce the memory usage and the computational complexity of full-precision convolutional neural networks (CNNs), which has been widely used in the field of deep learning. However, there are different properties between BNNs and real-valued models, making it difficult to draw on the experience of CNN composition to develop BNN. In this article, we study the application of binary network to the single-image super-resolution (SISR) task in which the network is trained for restoring original high-resolution (HR) images. Generally, the distribution of features in the network for SISR is more complex than those in recognition models for preserving the abundant image information, e.g., texture, color, and details. To enhance the representation ability of BNN, we explore a novel activation-rectified inference (ARI) module that achieves a more complete representation of features by combining observations from different quantitative perspectives. The activations are divided into several parts with different quantification intervals and are inferred independently. This allows the binary activations to retain more image detail and yield finer inference. In addition, we further propose an adaptive approximation estimator (AAE) for gradually learning the accurate gradient estimation interval in each layer to alleviate the optimization difficulty. Experiments conducted on several benchmarks show that our approach is able to learn a binary SISR model with superior performance over the state-of-the-art methods. The code will be released at https://github.com/jwxintt/Rectified-BSR.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Large Pose Face Recognition via Facial Representation Learning
abstract
Overcoming image acquisition perspectives and face pose variations is a key problem in unconstrained face recognition tasks. One of the practical approaches is by reconstructing the face with extreme pose into a version that is more easily recognized by the discriminator, such as a frontal face. Often, existing methods attempt to balance the accuracy of downstream tasks with human visual perception, but ignore the differences in propensity between the two. Besides, large-scale datasets of profile-frontal paired face images are absent, which further hinders the training of models. In this work, we investigate a variety of face reconstruction approaches and propose a very simple, but very effective method to match face images across different scenes, named facial representation learning (FRL). The core idea of FRL is to introduce a representation generator in front of a pre-trained face recognition model, which can extract face representations from arbitrary faces that are more suitable for recognition model discrimination. In particular, the representation generator reconstructs the facial representation by minimising identity differences from the frontal face and adds pixel-level and adversarial constraints to cater for discriminator preferences. Extensive benchmark experiments show that the proposed method not only achieves better performance than state-of-the-art methods, but also can further squeeze the inference potential of existing face recognition models.
Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Inf. Forensics Secur.1
2024 Toward Pixel-Level Precision for Binary Super-Resolution With Mixed Binary Representation
abstract
Binary neural network (BNN) is an effective method for reducing model computational and memory cost, which has achieved much progress in the super-resolution (SR) field. However, there is still a noticeable performance gap between a binary SR network and its full-precision counterpart. Considering that the information density in quantization features is far lower than full-precision features, we aim to improve the precision of quantization features to produce rich-enough output activations for SR task. First, we make several observations that a multibit value could be approximated by multiple 1-bit values, and the computation power of binary convolution could be improved by approximating the multibit convolution process. Then, we propose a mixed binary representation set to approximate multibit activations, which is effective in compensating the quantization precision loss. Finally, we present a new precision-driven binary convolution (PDBC) module, which increases the convolution precision and protects image detail information without extra computation. Compared with normal binary convolution, our method could largely reduce the information loss caused by binarization. In experiments, our methods consistently show superior performance over the baseline models and can surpass state-of-the-art methods in terms of peak signal to noise ratio (PSNR) and visual quality.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Local Means Binary Networks for Image Super-Resolution
abstract
The success of modern single image super-resolution (SISR) algorithms is inspired by the development of deep convolutional neural networks (CNNs). However, these CNN-based methods require considerable computation and complexity, making it impossible for these methods to perform real-time calculations in edge devices. Thus, lightweight model design has become a development trend in the super-resolution field, including pruning, quantization, and other methods. The 1-bit quantization is an extreme lightweight method which can reduce the calculation amount of the model in an extreme manner and is friendly to hardware such as edge devices. Most existing binary quantization approaches lead to a large information loss during forward propagation, especially in detailed color information (e.g., edge, texture, and contrast). The loss of color information makes modern binary methods unsuitable for SISR tasks. We think the loss occurs because these methods typically utilize a uniform threshold to quantize the weights and activations. Thus, in this article, we thoroughly analyze the difference between normal classification tasks and SISR tasks, and present a binarization scheme based on local means. The proposed method can maintain more detailed information in feature maps using dynamic thresholds during quantization. Specifically, each value in the full precision activations has a corresponding threshold during the quantization process, and those thresholds are determined by the full precision values of the surroundings. In addition, a gradient approximator is introduced to adaptively optimize the gradient for updating binary weights. We then verify the effectiveness of our method for training binary networks on several SISR benchmarks including VDSR and SRResNet. Experimental results show that the proposed method can outperform the state-of-the-art algorithms to obtain binary networks for image super-resolution with better peak signal-to-noise ratio (PSNR) values and visual quality.
Nannan Wang 0001, Jingwei Xin, Jie Li 0001, Xinbo Gao 0001, Kai Han 0002, Yunhe Wang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Advanced Binary Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
Int. J. Comput. Vis.1
2023 Learning a High Fidelity Identity Representation for Face Frontalization
abstract
This paper considers the problem of face frontalization in the wild, which transforms a face image with profile views into a frontal face. Face frontalization provides an effective solution to the face recognition problem in uncontrolled scenes. However, the existing methods either focus on deep learning techniques as an end-to-end framework or combine other explicit facial prior estimation tasks, such as 3D representation, optical flow estimation and so on, where computation is highly redundant and facial identity cannot be well represented. In this paper, we focus on how to maximise the potential of the model for identity learning and representation, and propose an accurate and lightweight face frontalization approach, named identity-preserving model (IPM). IPM has a well-designed encoder-decoder architecture which restores input face to a frontal counterpart. The encoder is constructed to extract representation from the input face, where a contrastive loss function is applied that encourages representations to form compact clusters, while preserving their relationships across the corpora. Then a cross-domain rectification module is proposed to eliminate the representation differences between the recognition and reconstruction domains, thus improving the accuracy of the reconstructed face. Extensive experiments on benchmark datasets show that the proposed IPM approach not only outperforms the state-of-the-art on public datasets but also can cope with images in the uncontrolled scenes.
Jingwei Xin, Zikai Wei, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Unsupervised Across Domain Consistency- Difference Network for Hyperspectral Image Super-Resolution
abstract
Without reducing the spectral resolution, hyperspectral image super-resolution has achieved remarkable progress thanks to the success of deep neural networks. However, existing methods can not fully excavate the latent high-frequency details only in the single spatial domain. Different from existing methods that only achieves the super-resolution task in spatial domain, we optimize the amplitude spectrum and phase spectrum in frequency domain to obtain high resolution hyperspectral image (HR-HSI). We propose a new unsupervised framework to reconstruct HR-HSI using only the observed low resolution HSI and HR multispectral image. Based on triple-level modeling, the encoder-decoder learns abundant features including contextual information from multiple scales. In addition, we propose iterative across domain consistency-difference (ADCD) module, which is embedded between encoder and decoder. In ADCD module, three parallel convolution streams, (amplitude spectrum adjustment branch, phase spectrum adjustment branch and spatial domain branch) are used to explore the consistency-difference between each other, which is preserved by memory units within the module. Particularly, we embed the dilated causal convolution in the frequency domain processing branch, which is convenient to flexibly adjust the receptive field and adapt to different domains. Extensive experiments are conducted on widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the proposed method.
Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 FABNet: Frequency-Aware Binarized Network for Single Image Super-Resolution
abstract
Remarkable achievements have been obtained with binary neural networks (BNN) in real-time and energy-efficient single-image super-resolution (SISR) methods. However, existing approaches often adopt the Sign function to quantize image features while ignoring the influence of image spatial frequency. We argue that we can minimize the quantization error by considering different spatial frequency components. To achieve this, we propose a frequency-aware binarized network (FABNet) for single image super-resolution. First, we leverage the wavelet transformation to decompose the features into low-frequency and high-frequency components and then employ a "divide-and-conquer" strategy to separately process them with well-designed binary network structures. Additionally, we introduce a dynamic binarization process that incorporates learned-threshold binarization during forward propagation and dynamic approximation during backward propagation, effectively addressing the diverse spatial frequency information. Compared to existing methods, our approach is effective in reducing quantization error and recovering image textures. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed methods could surpass state-of-the-art approaches in terms of PSNR and visual quality with significantly reduced computational costs. Our codes are available at https://github.com/xrjiang527/FABNet-PyTorch.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Jie Li 0001, Xiaoyu Wang 0002, Xinbo Gao 0001
IEEE Trans. Image Process.3
2023 Curvature Consistent Network for Microscope Chip Image Super-Resolution
abstract
Detecting hardware Trojan (HT) from a microscope chip image (MCI) is crucial for many applications, such as financial infrastructure and transport security. It takes an inordinate cost in scanning high-resolution (HR) microscope images for HT detection. It is useful when the chip image is in low-resolution (LR), which can be acquired faster and at a lower cost than its HR counterpart. However, the lost details and noises due to the electric charge effect in LR MCIs will affect the detection performance, making the problem more challenging. In this article, we address this issue by first discussing why recovering curvature information matters for HT detection and then proposing a novel MCI super-resolution (SR) method via a curvature consistent network (CCN). It consists of a homogeneous workflow and a heterogeneous workflow, where the former learns a mapping between homogeneous images, i.e., LR and HR MCIs, and the latter learns a mapping between heterogeneous images, i.e., MCIs and curvature images. Besides, a collaborative fusion strategy is used to leverage features learned from both workflows level-by-level by recovering the HR image eventually. To mitigate the issue of lacking an MCI dataset, we construct a new benchmark consisting of realistic MCIs at different resolutions, called MCI. Experiments on MCI demonstrate that the proposed CCN outperforms representative SR methods by recovering more delicate circuit lines and yields higher HT detection performance. The dataset is available at github.com/RuiZhang97/CCN.
Mingjin Zhang, Jingwei Xin, Jing Zhang 0037, Dacheng Tao, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 Learning Deep Resonant Prior for Hyperspectral Image Super-Resolution
abstract
Hyperspectral image super-resolution (HSISR) task has been widely studied, and significant progress has been made by leveraging the deep convolution neural network (CNN) techniques. Nevertheless, the scarcity of training images hinders the research progress of HSISR task. Moreover, the differences in imaging conditions and the number of spectral bands among different datasets, make it very difficult to construct a unified deep neural network. In this paper, we first present a non-training based HSISR method based on deep prior knowledge, which captures the image prior to restore the high resolution image by using the intrinsic characteristics of CNN. Then, we append a special network input processing module onto the HSI super-resolution network to automatically adjust the structure of the input so that the choice of network structure is no longer limited, while the network design focuses on exploiting the spatial information of hyperspectral images and the correlation between spectral bands, making the method more suitable for HSISR tasks and greatly extending its applications. Extensive experiment results on the hyperspectral image datasets illustrate the effectiveness of the proposed method, and we have got comparable results with the state-of-the-art methods while requiring no training samples.
Zhaori Gong, Nannan Wang 0001, De Cheng, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 External-Internal Attention for Hyperspectral Image Super-Resolution
abstract
In recent years, hyperspectral image (HSI) super-resolution has made significant progress by leveraging convolution neural network. Existing methods with spectral or spatial attention, which only consider the spectral similarity or pixel-pixel similarity, ignore sample-sample correlations and sparsity. Therefore, based on the fusion of HSI and multispectral image, we propose a new HSI super-resolution model with external-internal attention. Instead of considering a single sample, external attention module is employed to exploit the incorporating correlations between different samples to get a better feature representation. In addition, an internal attention module based on non-local operation is designed to explore the long-range dependencies information. Particularly, oriented to high mapping precision and low computational cost inference, spherical locality sensitive hashing is used to divide features into different hash buckets so that every query point is calculated in the hash bucket assigned to it, rather than based a weight sum of features across all positions. The sequential external-internal attention greatly improves the generalization ability and robustness of the model by modeling at the dataset level and at the sample level. Extensive experiments are conducted on five widely-used datasets in comparison with state-of-the-art models, demonstrating the advantage of the method we proposed.
Zhiling Guo, Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Wavelet-Based Dual Recursive Network for Image Super-Resolution
abstract
Although remarkable progress has been made on single-image super-resolution (SISR), deep learning methods cannot be easily applied to real-world applications due to the requirement of its heavy computation, especially for mobile devices. Focusing on the fewer parameters and faster inference SISR approach, we propose an efficient and time-saving wavelet transform-based network architecture, where the image super-resolution (SR) processing is carried out in the wavelet domain. Different from the existing methods that directly infer high-resolution (HR) image with the input low-resolution (LR) image, our approach first decomposes the LR image into a series of wavelet coefficients (WCs) and the network learns to predict the corresponding series of HR WCs and then reconstructs the HR image. Particularly, in order to further enhance the relationship between WCs and image deep characteristics, we propose two novel modules [wavelet feature mapping block (WFMB) and wavelet coefficients reconstruction block (WCRB)] and a dual recursive framework for joint learning strategy, thus forming a WCs prediction model to realize the efficient and accurate reconstruction of HR WCs. Experimental results show that the proposed method can outperform state-of-the-art methods with more than a 2× reduction in model parameters and computational complexity.
Jingwei Xin, Jie Li 0001, Nannan Wang 0001, Heng Huang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Training Binary Neural Network without Batch Normalization for Image Super-Resolution
abstract
Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces floating-point calculations, which is unfriendly to low-precision hardwares. Therefore, there is still room for improvement in terms of model performance and efficiency. Focusing on this issue, in this paper, we first explore a novel binary training mechanism based on the feature distribution, allowing us to replace all BN layers with a simple training method. Then, we construct a strong baseline by combining the highlights of recent binarization methods, which already surpasses the state-of-the-arts. Next, to train highly accurate binarized SR model, we also develop a lightweight network architecture and a multi-stage knowledge distillation strategy to enhance the model representation ability. Extensive experiments demonstrate that the proposed method not only presents advantages of lower computation as compared to conventional floating-point networks but outperforms the state-of-the-art binary methods on the standard SR networks.
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Xinbo Gao 0001
AAAI3
2021 Learning lightweight super-resolution networks with weight pruning
Nannan Wang 0001, Jingwei Xin, Xiaobo Xia, Xi Yang 0011, Xinbo Gao 0001
Neural Networks3
2020 Facial Attribute Capsules for Noise Face Super Resolution
abstract
Existing face super-resolution (SR) methods mainly assume the input image to be noise-free. Their performance degrades drastically when applied to real-world scenarios where the input image is always contaminated by noise. In this paper, we propose a Facial Attribute Capsules Network (FACN) to deal with the problem of high-scale super-resolution of noisy face image. Capsule is a group of neurons whose activity vector models different properties of the same entity. Inspired by the concept of capsule, we propose an integrated representation model of facial information, which named Facial Attribute Capsule (FAC). In the SR processing, we first generated a group of FACs from the input LR face, and then reconstructed the HR face from this group of FACs. Aiming to effectively improve the robustness of FAC to noise, we generate FAC in semantic, probabilistic and facial attributes manners by means of integrated learning strategy. Each FAC can be divided into two sub-capsules: Semantic Capsule (SC) and Probabilistic Capsule (PC). Them describe an explicit facial attribute in detail from two aspects of semantic representation and probability distribution. The group of FACs model an image as a combination of facial attribute information in the semantic space and probabilistic space by an attribute-disentangling way. The diverse FACs could better combine the face prior information to generate the face images with fine-grained semantic attributes. Extensive benchmark experiments show that our method achieves superior hallucination results and outperforms state-of-the-art for very low resolution (LR) noise face image super resolution.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001
AAAI1
2020 Video Face Super-Resolution with Motion-Adaptive Feedback Cell
abstract
Video super-resolution (VSR) methods have recently achieved a remarkable success due to the development of deep convolutional neural networks (CNN). Current state-of-the-art CNN methods usually treat the VSR problem as a large number of separate multi-frame super-resolution tasks, at which a batch of low resolution (LR) frames is utilized to generate a single high resolution (HR) frame, and running a slide window to select LR frames over the entire video would obtain a series of HR frames. However, duo to the complex temporal dependency between frames, with the number of LR input frames increase, the performance of the reconstructed HR frames become worse. The reason is in that these methods lack the ability to model complex temporal dependencies and hard to give an accurate motion estimation and compensation for VSR process. Which makes the performance degrade drastically when the motion in frames is complex. In this paper, we propose a Motion-Adaptive Feedback Cell (MAFC), a simple but effective block, which can efficiently capture the motion compensation and feed it back to the network in an adaptive way. Our approach efficiently utilizes the information of the inter-frame motion, the dependence of the network on motion estimation and compensation method can be avoid. In addition, benefiting from the excellent nature of MAFC, the network can achieve better performance in the case of extremely complex motion scenarios. Extensive evaluations and comparisons validate the strengths of our approach, and the experimental results demonstrated that the proposed framework is outperform the state-of-the-art methods.
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001, Zhifeng Li 0001
AAAI1
2020 Binarized Neural Network for Single Image Super Resolution
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Heng Huang 0001, Xinbo Gao 0001
ECCV (4)1
2020 Image Super-Resolution via Deep Feature Recalibration Network
Jingwei Xin, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
PRCV (1)1
2020 Image super-resolution via multi-view information fusion networks
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Yi Yu 0001, Xinbo Gao 0001
Neurocomputing3
2019 Residual Attribute Attention Network for Face Image Super-Resolution
abstract
Facial prior knowledge based methods recently achieved great success on the task of face image super-resolution (SR). The combination of different type of facial knowledge could be leveraged for better super-resolving face images, e.g., facial attribute information with texture and shape information. In this paper, we present a novel deep end-to-end network for face super resolution, named Residual Attribute Attention Network (RAAN), which realizes the efficient feature fusion of various types of facial information. Specifically, we construct a multi-block cascaded structure network with dense connection. Each block has three branches: Texture Prediction Network (TPN), Shape Generation Network (SGN) and Attribute Analysis Network (AAN). We divide the task of face image reconstruction into three steps: extracting the pixel level representation information from the input very low resolution (LR) image via TPN and SGN, extracting the semantic level representation information by AAN from the input, and finally combining the pixel level and semantic level information to recover the high resolution (HR) image. Experiments on benchmark database illustrate that RAAN significantly outperforms state-of-the-arts for very low-resolution face SR problem, both quantitatively and qualitatively.
Jingwei Xin, Nannan Wang 0001, Xinbo Gao 0001, Jie Li 0001
AAAI1
2016 Interference migration using concurrent transmission for energy-efficient HetNets
Xiao Ma 0007, Min Sheng, Jiandong Li 0001, Jingwei Xin
Sci. China Inf. Sci.4