Xin Li 0106

dblp:09/1365-106 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
12since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 9 since 2021Computer networks · 1
YearPublicationVenuePosition
2023 LMR: A Large-Scale Multi-Reference Dataset for Reference-based Super-Resolution
abstract
It is widely agreed that reference-based super-resolution (RefSR) achieves superior results by referring to similar high quality images, compared to single image super-resolution (SISR). Intuitively, the more references, the better performance. However, previous RefSR methods have all focused on single-reference image training, while multiple reference images are often available in testing or practical applications. The root cause of such training-testing mismatch is the absence of publicly available multi-reference SR training datasets, which greatly hinders research efforts on multi-reference super-resolution. To this end, we construct a large-scale, multi-reference super-resolution dataset, named LMR. It contains 112, 142 groups of 300×300 training images, which is 10× of the existing largest RefSR dataset. The image size is also some times larger. More importantly, each group is equipped with 5 reference images with different similarity levels. Furthermore, we propose a new baseline method for multi-reference super-resolution: MRefSR, including a Multi-Reference Attention Module (MAM) for feature fusion of an arbitrary number of reference images, and a Spatial Aware Filtering Module (SAFM) for the fused feature selection. The proposed MRefSR achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations. Our code and data are available at: https://github.com/wdmwhh/MRefSR.
Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Errui Ding, Zhaoxiang Zhang 0001
ICCV2
2023 Toward Interactive Self-Supervised Denoising
abstract
Self-supervised denoising frameworks have recently been proposed to learn denoising models without noisy-clean image pairs, showing great potential in various applications. The denoising model is expected to produce visually pleasant images without noise patterns. However, it is non-trivial to achieve this goal using self-supervised methods because 1) the self-supervised model is difficult to restore the perceptual information due to the lack of clean supervision, and 2) perceptual quality is relatively subjective to users’ preferences. In this paper, we make the first attempt to build an interactive self-supervised denoising model to tackle the aforementioned problems. Specifically, we propose an interactive two-branch network to effectively restore perceptual information. The network consists of a denoising branch and an interactive branch, where the former focuses on efficient denoising, and the latter modulates the denoising branch. Based on the delicate architecture design, our network can produce various denoising outputs, allowing the user to easily select the most appealing outcome for satisfying the perceptual requirement. Moreover, to optimize the network with only noisy images, we propose a novel two-stage training strategy in a self-supervised way. Once the network is optimized, it can be interactively changed between noise reduction and texture restoration, providing more denoising choices for users. Existing self-supervised denoising methods can be integrated into our method to be user-friendly with interaction. Extensive experiments and comprehensive analyses are conducted to validate the effectiveness of the proposed method.
Mingde Yao, Dongliang He, Xin Li 0106, Fu Li 0002, Zhiwei Xiong
IEEE Trans. Circuits Syst. Video Technol.3
2023 Bidirectional Translation Between UHD-HDR and HD-SDR Videos
abstract
With the popularization of ultra high definition (UHD) high dynamic range (HDR) displays, recent works focus onupgradinghigh definition (HD) standard dynamic range (SDR) videos to UHD-HDR versions, aiming to provides richer details and higher contrasts on advanced modern displays. However, joint considering theupgrading&downgradingtranslations between two types of videos, which is practical in real applications, is generally neglected. On the one hand,downgradingtranslation is the key to showing UHD-HDR videos on HD-SDR displays. On the other hand, considering both translations enables joint optimization and results in high quality translation. To this end, we propose the bidirectional translation network (BiT-Net), which jointly considers two translations in one network for the first time. In brief, BiT-Net is elaborately designed in aninvertiblefashion that can be efficiently inferred along forward and backward directions fordowngradingandupgradingtasks, respectively. Based on this framework, we divide each direction into three sub-tasks,i.e., decomposition, structure-guided translation, and synthesis, to effectively translate the dynamic range and the high-frequency details. Benefiting from the dedicated architecture, our BiT-Net can work on 1) downgrading UHD-HDR videos, 2) upgrading existing HD-SDR videos, and 3) synthesizing UHD-HDR versions from the downgraded HD-SDR videos. Experiments show that the proposed method achieves state-of-the-art performances on all these three tasks.
Mingde Yao, Dongliang He, Xin Li 0106, Zhihong Pan 0001, Zhiwei Xiong
IEEE Trans. Multim.3
2023 Dual-Affinity Style Embedding Network for Semantic-Aligned Image Style Transfer
abstract
Image style transfer aims at synthesizing an image with the content from one image and the style from another. User studies have revealed that the semantic correspondence between style and content greatly affects subjective perception of style transfer results. While current studies have made great progress in improving the visual quality of stylized images, most methods directly transfer global style statistics without considering semantic alignment. Current semantic style transfer approaches still work in an iterative optimization fashion, which is impractically computationally expensive. Addressing these issues, we introduce a novel dual-affinity style embedding network (DaseNet) to synthesize images with style aligned at semantic region granularity. In the dual-affinity module, feature correlation and semantic correspondence between content and style images are modeled jointly for embedding local style patterns according to semantic distribution. Furthermore, the semantic-weighted style loss and the region-consistency loss are introduced to ensure semantic alignment and content preservation. With the end-to-end network architecture, DaseNet can well balance visual quality and inference efficiency for semantic style transfer. Experimental results on different scene categories have demonstrated the effectiveness of the proposed method.
Zhuoqi Ma, Xin Li 0106, Fu Li 0003, Dongliang He, Errui Ding, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Towards Bidirectional Arbitrary Image Rescaling: Joint Optimization and Cycle Idempotence
abstract
Deep learning based single image super-resolution models have been widely studied and superb results are achieved in upscaling low-resolution images with fixed scale factor and downscaling degradation kernel. To improve real world applicability of such models, there are growing interests to develop models optimized for arbitrary upscaling factors. Our proposed method is the first to treat arbitrary rescaling, both upscaling and downscaling, as one unified process. Using joint optimization of both directions, the proposed model is able to learn upscaling and downscaling simultaneously and achieve bidirectional arbitrary image rescaling. It improves the performance of current arbitrary upscaling models by a large margin while at the same time learns to maintain visual perception quality in downscaled images. The proposed model is further shown to be robust in cycle idempotence test, free of severe degradations in reconstruction accuracy when the downscaling-to-upscaling cycle is applied repetitively. This robustness is beneficial for image rescaling in the wild when this cycle could be applied to one image for multiple times. It also performs well on tests with arbitrary large scales and asymmetric scales, even when the model is not trained with such tasks. Extensive experiments are conducted to demonstrate the superior performance of our model.
Zhihong Pan 0001, Baopu Li, Dongliang He, Mingde Yao, Xin Li 0106, Errui Ding
CVPR7
2022 Neural Color Operators for Sequential Image Retouching
Yili Wang 0003, Xin Li 0106, Kun Xu 0003, Dongliang He, Qi Zhang 0029, Fu Li 0003, Errui Ding
ECCV (19)2
2022 RRSR: Reciprocal Reference-Based Image Super-Resolution with Progressive Feature Alignment and Selection
Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Yili Wang 0003, Zhaoxiang Zhang 0001
ECCV (19)2
2021 Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer
abstract
Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize complex style, especially when holistic global and local patterns exist. Inspired by the common painting process of drawing a draft and revising the details, we introduce a novel feed-forward method named Laplacian Pyramid Network (LapStyle). LapStyle first transfers global style patterns in low-resolution via a Drafting Network. It then revises the local details in high-resolution via a Revision Network, which hallucinates a residual image according to the draft and the image textures extracted by Laplacian filtering. Higher resolution details can be easily generated by stacking Revision Networks with multiple Laplacian pyramid levels. The final stylized image is obtained by aggregating outputs of all pyramid levels. Experiments demonstrate that our method can synthesize high quality stylized images in real time, where holistic style patterns are properly transferred.
Zhuoqi Ma, Fu Li 0003, Dongliang He, Xin Li 0106, Errui Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
CVPR5
2021 Learning Semantic Person Image Generation by Region-Adaptive Normalization
abstract
Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic appearance, resulting in inconsistent and less realistic textures on the reconstructed results. To address this issue, we propose a new two-stage framework to handle the pose and appearance translation. In the first stage, we predict the target semantic parsing maps to eliminate the difficulties of pose transfer and further benefit the latter translation of per-region appearance style. In the second one, with the predicted target semantic maps, we suggest a new person image generation method by incorporating the region-adaptive normalization, in which it takes the per-region styles to guide the target appearance generation. Extensive experiments show that our proposed SPGNet can generate more semantic, consistent, and photorealistic results and perform favorably against the state of the art methods in terms of quantitative and qualitative evaluation. The source code and model are available at https://github.com/cszy98/SPGNet.git.
Zhengyao Lv, Xiaoming Li 0002, Xin Li 0106, Fu Li 0003, Dongliang He, Wangmeng Zuo
CVPR3
2021 Paint Transformer: Feed Forward Neural Painting with Stroke Prediction
abstract
Neural painting refers to the procedure of producing a series of strokes for a given image and non-photo-realistically recreating it using neural networks. While reinforcement learning (RL) based agents can generate a stroke sequence step by step for this task, it is not easy to train a stable RL agent. On the other hand, stroke optimization methods search for a set of stroke parameters iteratively in a large search space; such low efficiency significantly limits their prevalence and practicality. Different from previous methods, in this paper, we formulate the task as a set prediction problem and propose a novel Transformer-based framework, dubbed Paint Transformer, to predict the parameters of a stroke set with a feed forward network. This way, our model can generate a set of strokes in parallel and obtain the final painting of size 512 × 512 in near real time. More importantly, since there is no dataset available for training the Paint Transformer, we devise a self-training pipeline such that it can be trained without any off-the-shelf dataset while still achieving excellent generalization capability. Experiments demonstrate that our method achieves better painting performance than previous ones with cheaper training and inference costs. Codes and models are available on https://github.com/wzmsltw/PaintTransformer.
Songhua Liu, Dongliang He, Fu Li 0003, Ruifeng Deng, Xin Li 0106, Errui Ding, Hao Wang 0014
ICCV6
2021 AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer
abstract
Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distributions, or adaptively normalize deep content feature according to the style such that their global statistics are matched. Although effective, leaving shallow feature unexplored and without locally considering feature statistics, they are prone to unnatural output with unpleasing local distortions. To alleviate this problem, in this paper, we propose a novel attention and normalization module, named Adaptive Attention Normalization (AdaAttN), to adaptively perform attentive normalization on per-point basis. Specifically, spatial attention score is learnt from both shallow and deep features of content and style images. Then perpoint weighted statistics are calculated by regarding a style feature point as a distribution of attention-weighted output of all style feature points. Finally, the content feature is normalized so that they demonstrate the same local feature statistics as the calculated per-point weighted style feature statistics. Besides, a novel local feature loss is derived based on AdaAttN to enhance local visual quality. We also extend AdaAttN to be ready for video style transfer with slight modifications. Experiments demonstrate that our method achieves state-of-the-art arbitrary image/video style transfer. Codes and models are available on https://github.com/wzmsltw/AdaAttN.
Songhua Liu, Dongliang He, Fu Li 0003, Xin Li 0106, Zhengxing Sun, Qian Li 0014, Errui Ding
ICCV6
2021 Image Inpainting by End-to-End Cascaded Refinement With Mask Awareness
abstract
Inpainting arbitrary missing regions is challenging because learning valid features for various masked regions is nontrivial. Though U-shaped encoder-decoder frameworks have been witnessed to be successful, most of them share a common drawback of mask unawareness in feature extraction because all convolution windows (or regions), including those with various shapes of missing pixels, are treated equally and filtered with fixed learned kernels. To this end, we propose our novel mask-aware inpainting solution. Firstly, a Mask-Aware Dynamic Filtering (MADF) module is designed to effectively learn multi-scale features for missing regions in the encoding phase. Specifically, filters for each convolution window are generated from features of the corresponding region of the mask. The second fold of mask awareness is achieved by adopting Point-wise Normalization (PN) in our decoding phase, considering that statistical natures of features at masked points differentiate from those of unmasked points. The proposed PN can tackle this issue by dynamically assigning point-wise scaling factor and bias. Lastly, our model is designed to be an end-to-end cascaded refinement one. Supervision information such as reconstruction loss, perceptual loss and total variation loss is incrementally leveraged to boost the inpainting results from coarse to fine. Effectiveness of the proposed framework is validated both quantitatively and qualitatively via extensive experiments on three public datasets including Places2, CelebA and Paris StreetView.
Manyu Zhu, Dongliang He, Xin Li 0106, Chao Li 0034, Fu Li 0003, Xiao Liu 0022, Errui Ding, Zhaoxiang Zhang 0001
IEEE Trans. Image Process.3
2020 Deep Concept-wise Temporal Convolutional Networks for Action Localization
abstract
Existing action localization approaches adopt shallow temporal convolutional networks (i.e., TCN) on 1D feature map extracted from video frames. In this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which generally are highly abstract and can be regarded as latent concepts, are excessively recombined in temporal convolution. To address this issue, we introduce a novel concept-wise temporal convolutional network (C-TCN) as an alternative to TCN for training deeper action localization networks. To address this issue, we introduce a novel concept-wise temporal convolution (CTC) layer as an alternative to conventional temporal convolution layer for training deeper action localization networks. Instead of recombining latent concepts, CTC layer deploys a number of temporal filters to each concept separately with shared filter parameters across concepts. Thus can capture common temporal patterns of different concepts and significantly enrich representation ability. Via stacking CTC layers, we proposed a deep concept-wise temporal convolutional network (C-TCN), which boosts the state-of-the-art action localization performance on THUMOS'14 from 42.8 to 52.1 in terms of mAP(%), achieving a relative improvement of 21.7%. Favorable result is also obtained on ActivityNet.
Xin Li 0106, Xiao Liu 0022, Wangmeng Zuo, Chao Li 0034, Xiang Long, Dongliang He, Fu Li 0003, Shilei Wen, Chuang Gan 0001
ACM Multimedia1
2019 BMN: Boundary-Matching Network for Temporal Action Proposal Generation
abstract
Temporal action proposal generation is an challenging and promising task which aims to locate temporal regions in real-world videos where action or event may occur. Current bottom-up proposal generation methods can generate proposals with precise boundary, but cannot efficiently generate adequately reliable confidence scores for retrieving proposals. To address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending boundaries and combine all densely distributed BM pairs into the BM confidence map. Based on BM mechanism, we propose an effective, efficient and end-to-end proposal generation method, named Boundary-Matching Network (BMN), which generates proposals with precise temporal boundaries as well as reliable confidence scores simultaneously. The two-branches of BMN are jointly trained in an unified framework. We conduct experiments on two challenging datasets: THUMOS-14 and ActivityNet-1.3, where BMN shows significant performance improvement with remarkable efficiency and generalizability. Further, combining with existing action classifier, BMN can achieve state-of-the-art temporal action detection performance.
Xiao Liu 0022, Xin Li 0106, Errui Ding, Shilei Wen
ICCV3
2019 Mobile Display Power Reduction for Video Using Standardized Metadata
abstract
The inordinate power consumption of contemporary display panels accelerates battery depletion in mobile devices. Existing display-power reduction approaches are suboptimal because they incur additional computations and latency on the mobile device or because they assume a server-client model in which a single entity controls the server and the client. To overcome these limitations, we extract standardized metadata at a server and transmit it to a client along with the video bitstream. This metadata includes quality indicators, associated power controls, and contrast-enhancement information. The metadata guarantees that specified quality levels are achieved while avoiding flicker and minimizing power consumption. Furthermore, the metadata has minimal computational, storage & power overheads and, because it is standardized, the metadata benefits servers and clients in open ecosystems wherein all entities function independently. Our research has been implemented on mobile devices and provides, on average, 32.5 percent power reduction at high quality and up to 84 percent power reduction at acceptable quality. The ISO/IEC Moving Pictures Expert Group has recently standardized our proposed metadata in the Green Metadata Standard.
Felix C. A. Fernandes, Esmaeil Faramarzi, Xin Li 0106, Zhan Ma 0001, Xavier Ducloux
IEEE Trans. Mob. Comput.3
2017 A Practical System Towards the Secure, Robust and Pervasive Mobile Workstyle
abstract
We develop an innovative PC2PC (personal computer to pervasive computing) system to enable the secure, robust and pervasive mobile workstyle. PC2PC server compresses the desktop screens of any virtualized system, and delivers the stream through any popular networks to PC2PC client remotely for stream decoding, rendering and end-user interaction (such as keyboard/mouse commands). We have implemented the overall system from the scratch, where the emerging screen content coding (SCC) extension of the High-Efficiency Video Coding (HEVC) is implemented to compress and stream the desktop screens in real-time, and three core asset channels (i.e., system, display, inputs, etc) are defined to enable systematic end-to-end communication. Compared with the commercial Red Hat SPICE virtual desktop infrastructure (VDI) scheme, our PC2PC could save the network bandwidth by a factor of 2, 7 and 4 respectively for typical video streaming, web browsing and stationary office applications at same visual quality. Meanwhile, we have also measured the delays in the system and presented the preliminary study on the user experience impact. A simple network estimation is applied to optimize the quality-bandwidth adaptation for both single user and multiuser scenarios to combat the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
VTC Spring5
2017 Interactive Screen Video Streaming-Based Pervasive Mobile Workstyle
abstract
In this paper, we develop an interactive screen video streaming-based system to enable the ubiquitous mobile workstyle, which is referred to as personal computer to pervasive computing (PC2PC). The desktop screens of virtualized systems are compressed in the PC2PC servers and delivered to remote end users for stream decoding, rendering, and interactions. We have implemented a system from the scratch, where the emerging screen content coding extension of high-efficiency video coding is implemented to compress and stream the desktop screens of the virtualized system in real time. Three core asset channels, system, display, and inputs, are defined to enable systematic end-to-end communication. Compared with Red Hat SPICE virtual desktop infrastructure scheme, the proposed PC2PC could save network bandwidth consumption by a factor of 2, 7, and 4, respectively, in terms of typical video streaming, web browsing, and stationary office applications at the same visual quality. Meanwhile, we have also measured the delays of the system and presented preliminary results on the user experience aspect. A simple network estimation is applied to optimize the quality bandwidth adaptation for both single user and multiuser scenarios to consider the network dynamics.
Zhan Ma 0001, Tao Yue 0003, Xun Cao, Yiling Xu, Xin Li 0106, Yongjin Wang
IEEE Trans. Multim.5