VLDB 2026 Research / reviewers in the wild / expert
Xiao Liu 0022
dblp:82/1364-22
· DBLP profile ↗
26ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0002-5689-9786ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DSRIR: Dynamic spatial refinement learning for progressive all-in-one image restoration
Xiao Liu 0022, Yutong Yang, Zhengyong Wang, Xiaohai He, Honggang Chen, Yi Li 0069, Pingyu Wang |
Inf. Process. Manag. | 2 |
| 2026 | D2S-RSG-SSD: Dual Double-Sampling With Random Sub-Samples Generation for Self-Supervised Real Image DenoisingabstractRecent advances in self-supervised image denoising have highlighted the potential of Blind-Spot Networks (BSNs). However, existing methods suffer from three major limitations: (1) Their effectiveness in real-world scenarios is limited by strong assumptions, such as noise independence, which rarely hold in practice. (2) While sampling-based strategies can partially improve performance, BSNs inherently suffer from information loss caused by centroid masking, and removing the blind spot leads to noise overfitting, both of which hinder denoising performance. (3) Sampling-based methods often introduce checkerboard artifacts, yet existing studies typically overlook the fundamental differences between these artifacts and real noise. To address these issues, we propose a novel self-supervised denoising framework, Dual Double-Sampling with Random Sub-samples Generation (D2S-RSG-SSD). To address Limitation 1, we introduce a sampling-based framework that breaks noise dependence by combining Random Sub-samples Generation (RSG) with a cross-paired loss $\mathcal {L}_{RSG}$LRSG. RSG generates diverse sub-samples with inherent variance, referred to as sampling differences, which serve as natural perturbations to augment training data and disrupt spatial noise correlations. The proposed loss function ensures full utilization of these sub-samples while stabilizing optimization. To address Limitation 2, we propose a Dual Double-Sampling (D2S) strategy with fixed sampling patterns and a dual-branch architecture. This design reduces reliance on pixel-level information and leverages complementary features to mitigate both noise overfitting and information loss. A key advantage is its compatibility with various advanced denoising networks, lifting the constraint of using BSNs in self-supervised settings. Additionally, we introduce a fixed sub-image sampling strategy to prevent pattern collapse during inference and ensure stability. To address Limitation 3, we explicitly differentiate checkerboard artifacts from real noise and develop a dedicated artifact remover to correct pixel discontinuities caused by sampling-based operations. This design preserves fine image details while reducing over-smoothing. Experiments on benchmark real-noise datasets and self-captured noisy images demonstrate the robustness and generalizability of our framework, achieving better performance over existing methods. Xiao Liu 0022, Xiuya Shi, Yizhong Pan, Shuhang Gu, Wei Liu 0044, Chao Ren 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Degradation-Aware Feature Perturbation for All-in-One Image RestorationabstractAll-in-one image restoration aims to recover clear images from various degradation types and levels with a unified model. Nonetheless, the significant variations among degradation types present challenges for training a universal model, often resulting in task interference, where the gradient update directions of different tasks may diverge due to shared parameters. To address this issue, motivated by the routing strategy, we propose DFPIR, a novel all-in-one image restorer that introduces Degradation-aware Feature Perturbations(DFP) to adjust the feature space to align with the unified parameter space. In this paper, the feature perturbations primarily include channel-wise perturbations and attention-wise perturbations. Specifically, channel-wise perturbations are implemented by shuffling the channels in high-dimensional space guided by degradation types, while attention-wise perturbations are achieved through selective masking in the attention space. To achieve these goals, we propose a Degradation-Guided Perturbation Block (DGPB) to implement these two functions, positioned between the encoding and decoding stages of the encoder-decoder architecture. Extensive experimental results demonstrate that DFPIR achieves state-of-the-art performance on several all-in-one image restoration tasks including image denoising, image dehazing, image deraining, motion deblurring, and low-light image enhancement. Our codes are available at https://github.com/TxpHome/DFPIR. Xiangpeng Tian, Xiangyu Liao, Xiao Liu 0022, Meng Li 0095, Chao Ren 0002 |
CVPR | 3 |
| 2025 | Efficient image super resolution via Mixed Window and Dimension Interaction
Xiao Liu 0022, Xiangyu Liao, Chao Ren 0002 |
Neurocomputing | 3 |
| 2025 | Real-world blind image super-resolution with mixed and probabilistic scheme based synthetic degradation pipeline
Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Chao Ren 0002 |
Knowl. Based Syst. | 1 |
| 2025 | Transformer-Style Convolutional Network for Efficient Natural and Industrial Image SuperresolutionabstractSingle image superresolution (SISR) is a critical task in computer vision with significant applications in both natural and industrial contexts. Although transformer-based approaches for SISR have achieved notable progress due to their exceptional representational capabilities, their quadratic computational complexity poses challenges for deployment on devices with limited resources. Conversely, convolutional networks (ConvNets) are inherently efficient but have difficulty capturing long-range pixel relationships because of their focus on spatial locality. This gives rise to a complementary relationship between the representational power of transformers and the efficiency of ConvNets, both of which are essential for practical applications. Motivated by this, in this article, we introduce TSCN, a novel transformer-style ConvNet. Our analysis highlights the strengths of transformers, including large-range dependencies modeling, two-order features interaction, input self-adaptation, and incorporating advanced components. Based on these insights, we guide the design of ConvNets to fully exploit these characteristics. Specifically, we rethink spatial convolution to enhance the modeling of spatial features and modify the macrostructure of the transformer by replacing self-attention and feed-forward network with the large-range multiorder convolution modulation (LMCM) layer and spatial awareness dynamic feature flow (SADFF) layer. The LMCM integrates reweighting into the large-range convolutional modulation technology, allowing self-adaptive recalibration of input representations using convolutional features as weight matrices and multiorder features interaction. In addition, the SADFF introduces spatial awareness, locality, and dynamic information flow modulation between layers. Experimental results demonstrate that our TSCN outperforms the state-of-the-art method SRFormer on multiple benchmarks by 0.03$\sim$0.17 dB, while using fewer parameters and computations. Xiao Liu 0022, Zhengyong Wang, Xiaohai He, Haosong Gou, Chao Ren 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | GBPG-Net: Global Background Prior-Guided Rain and Snow Image RestorationabstractThe aim of image restoration in the presence of rain and snow effects is to eliminate these disturbances while retaining the underlying background structure. Most existing methods tend to directly learn the mapping from corrupted images to clean ones, often resulting in residual rain or snow artifacts and compromised background structures. In this work, both theoretical analysis and experimental findings confirm the robustness of the hue channel in HSV color space to rain and snow disturbances, even when extracted from corrupted images. Motivated by this insight, we propose to leverage the global clean background cues inherent in the hue channel to guide the network in preserving the image background structure and removing interference. To this end, we introduce the global background prior-guided network (GBPG-Net) for restoring rain and snow-affected images, which employs a triangular formation to facilitate continuous interaction and updating of the global background prior (GBP) with the image feature within the GBPG-unit, resulting in improved interference removal and background structure preservation. Specifically, the GBPG-Net incorporates the global clean background prior injector (GCBPI) to inject the GBP into the network. Subsequently, the prior-guided local detail excavation (PGLDE) module, built on GCBPI, further refines interference removal and structure preservation to process local details intricately. Finally, the prior-guided local-global aggregation (PGLGA) module aggregates global background features with local detailed features, enabling the network to better understand the overall content and subtle interference for more accurate reconstruction. Quantitative and qualitative evaluations on synthetic and real datasets demonstrate the effectiveness of the proposed GBPG-Net in deraining and desnowing tasks, highlighting its advantages over existing methods. The code and supplementary documentation are available at https://github.com/liux520/GBPG-Net. Xiao Liu 0022, Haosong Gou, Zhengyong Wang, Chao Ren 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Efficient Information Modulation Network for Image Super-ResolutionabstractRecent researches have shown that the success of Transformers comes from their macro-level framework and advanced components, not just their self-attention (SA) mechanism. Comparable results can be obtained by replacing SA with spatial pooling, shifting, MLP, fourier transform and constant matrix, all of which have spatial information encoding capability like SA. In light of these findings, this work focuses on combining efficient spatial information encoding technology with superior macro architectures in Transformers. We rethink spatial convolution to achieve more efficient encoding of spatial features and dynamic modulation value representations by convolutional modulation techniques. The large-kernel convolution and Hadamard product are utilizated in the proposed Multi-orders Long-range convolutional modulation (MOLRCM) layer to imitate the implementation of SA. Moreover, MOLRCM layer also achieve long-range correlations and self-adaptation behavior, similar to SA, with linear complexity. On the other hand, we also address the sub-optimality of vanilla feed-forward networks (FFN) by introducing spatial awareness and locality, improving feature diversity, and regulating information flow between layers in the proposed Spatial Awareness Dynamic Feature Flow Modulation (SADFFM) layer. Experiment results show that our proposed efficient information modulation network (EIMN) performs better both quantitatively and qualitatively. Codes and supplementary materials link: https://github.com/liux520/EIMN. Xiao Liu 0022, Xiangyu Liao, Xiuya Shi, Linbo Qing, Chao Ren 0002 |
ECAI | 1 |
| 2023 | Unsupervised Image Denoising in Real-World Scenarios via Self-Collaboration Parallel Generative Adversarial BranchesabstractDeep learning methods have shown remarkable performance in image denoising, particularly when trained on large-scale paired datasets. However, acquiring such paired datasets for real-world scenarios poses a significant challenge. Although unsupervised approaches based on generative adversarial networks (GANs) offer a promising solution for denoising without paired datasets, they are difficult in surpassing the performance limitations of conventional GAN-based unsupervised frameworks without significantly modifying existing structures or increasing the computational complexity of denoisers. To address this problem, we propose a self-collaboration (SC) strategy for multiple denoisers. This strategy can achieve significant performance improvement without increasing the inference complexity of the GAN-based denoising framework. Its basic idea is to iteratively replace the previous less powerful denoiser in the filter-guided noise extraction module with the current powerful denoiser. This process generates better synthetic clean-noisy image pairs, leading to a more powerful denoiser for the next iteration. In addition, we propose a baseline method that includes parallel generative adversarial branches with complementary "self-synthesis" and "unpaired-synthesis" constraints. This baseline ensures the stability and effectiveness of the training network. The experimental results demonstrate the superiority of our method over state-of-the-art unsupervised methods. https://github.com/linxin0/SCPGabNet Chao Ren 0002, Xiao Liu 0022, Jie Huang 0036, Yinjie Lei |
ICCV | 3 |
| 2023 | Random Sub-Samples Generation for Self-Supervised Real Image DenoisingabstractWith sufficient paired training samples, the supervised deep learning methods have attracted much attention in image denoising because of their superior performance. However, it is still very challenging to widely utilize the supervised methods in real cases due to the lack of paired noisy-clean images. Meanwhile, most self-supervised denoising methods are ineffective as well when applied to the real-world denoising tasks because of their strict assumptions in applications. For example, as a typical method for self-supervised denoising, the original blind spot network (BSN) assumes that the noise is pixel-wise independent, which is much different from the real cases. To solve this problem, we propose a novel self-supervised real image denoising framework named Sampling Difference As Perturbation (SDAP) based on Random Sub-samples Generation (RSG) with a cyclic sample difference loss. Specifically, we dig deeper into the properties of BSN to make it more suitable for real noise. Surprisingly, we find that adding an appropriate perturbation to the training images can effectively improve the performance of BSN. Further, we propose that the sampling difference can be considered as perturbation to achieve better results. Finally we propose a new BSN framework in combination with our RSG strategy. The results show that it significantly outperforms other state-of-the-art self-supervised denoising methods on real-world datasets. The code is available at https://github.com/p1y2z3/SDAP. Yizhong Pan, Xiao Liu 0022, Xiangyu Liao, Yuanzhouhan Cao, Chao Ren 0002 |
ICCV | 2 |
| 2023 | Efficient Parallel Multi-Scale Detail and Semantic Encoding Network for Lightweight Semantic SegmentationabstractIn this work, we propose PMSDSEN, a parallel multi-scale encoder-decoder network architecture for semantic segmentation, inspired by the human visual perception system's ability to aggregate contextual information in various contexts and scales. Our approach introduces the efficient Parallel Multi-Scale Detail and Semantic Encoding (PMSDSE) unit to extract detailed local information and coarse large-range relationships in parallel, enabling the recognition of object boundaries and object-level areas. By stacking multiple PMSDSEs, our network learns fine-grained details and textures along with abstract category and semantic information, effectively utilizing a larger range of surrounding context information for robust segmentation. To further enhance the network's receptive field without increasing computational complexity, the Multi-Scale Semantic Extractor (MSSE) at the end of the encoder is utilized for multi-scale semantic context extraction and detailed information encoding. Additionally, the Dynamic Weighted Feature Fusion (DWFF) strategy is employed to integrate shallow layer detail information and deep layer semantic information during the decoder stage. Our method can obtain multi-scale context from local to global, achieving efficiently low-level feature extraction to high-level semantic interpretation at different scales and in different contexts. Without bells and whistles, PMSDSEN obtains a better trade-off between accuracy and complexity on popular benchmarks, including Cityscapes and Camvid. Specifically, PMSDSEN attains 73.2% mIoU with only 0.9M parameters on the Cityscapes test set. Codes and supplementary materials link: https://github.com/liux520/PMSDSEN. Xiao Liu 0022, Xiuya Shi, Lufei Chen, Linbo Qing, Chao Ren 0002 |
ACM Multimedia | 1 |
| 2023 | CasaPuNet: Channel Affine Self-Attention- Based Progressively Updated Network for Real Image DenoisingabstractRecently, the popularity of deep learning has brought broad applications of computer vision technology in industrial information systems. However, the process of image acquisition will inevitably introduce noise, which may heavily degrade image visual quality. Most of the proposed denoising methods are nonblind and they have limited performance in removing real noise with different noise levels. To overcome this problem, we propose a deep convolutional neural network (CNN)-based blind model, i.e., channel affine self-attention (CASA) based progressively updated network (CasaPuNet) for real image denoising. First, we introduce degradation mapping module (DMM) to extract degradation information, which makes the remaining subnetwork of CasaPuNet perform nonblind denoising. Then, CasaPuNet adopts a multistage architecture, which resolves the large gap between the noisy input and clean output into several small gaps and eliminates these small gaps step by step through progressive inference. Finally, a novel CASA is designed to adaptively fuse the features from multiple stages according to input statistics. Specifically, CASA extracts channel information from different features and converts them into channel weights through an affine structure for adaptive adjustment. CASA brings a significant performance gain with a small number of parameters. Extensive experiments demonstrate that CasaPuNet outperforms state-of-the-art denoising methods both quantitatively and visually. Jie Huang 0036, Xiao Liu 0022, Yizhong Pan, Xiaohai He, Chao Ren 0002 |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Image Inpainting by End-to-End Cascaded Refinement With Mask AwarenessabstractInpainting arbitrary missing regions is challenging because learning valid features for various masked regions is nontrivial. Though U-shaped encoder-decoder frameworks have been witnessed to be successful, most of them share a common drawback of mask unawareness in feature extraction because all convolution windows (or regions), including those with various shapes of missing pixels, are treated equally and filtered with fixed learned kernels. To this end, we propose our novel mask-aware inpainting solution. Firstly, a Mask-Aware Dynamic Filtering (MADF) module is designed to effectively learn multi-scale features for missing regions in the encoding phase. Specifically, filters for each convolution window are generated from features of the corresponding region of the mask. The second fold of mask awareness is achieved by adopting Point-wise Normalization (PN) in our decoding phase, considering that statistical natures of features at masked points differentiate from those of unmasked points. The proposed PN can tackle this issue by dynamically assigning point-wise scaling factor and bias. Lastly, our model is designed to be an end-to-end cascaded refinement one. Supervision information such as reconstruction loss, perceptual loss and total variation loss is incrementally leveraged to boost the inpainting results from coarse to fine. Effectiveness of the proposed framework is validated both quantitatively and qualitatively via extensive experiments on three public datasets including Places2, CelebA and Paris StreetView. Manyu Zhu, Dongliang He, Xin Li 0106, Chao Li 0034, Fu Li 0003, Xiao Liu 0022, Errui Ding, Zhaoxiang Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Dynamic Instance Normalization for Arbitrary Style TransferabstractPrior normalization methods rely on affine transformations to produce arbitrary image style transfers, of which the parameters are computed in a pre-defined way. Such manually-defined nature eventually results in the high-cost and shared encoders for both style and content encoding, making style transfer systems cumbersome to be deployed in resource-constrained environments like on the mobile-terminal side. In this paper, we propose a new and generalized normalization module, termed as Dynamic Instance Normalization (DIN), that allows for flexible and more efficient arbitrary style transfers. Comprising an instance normalization and a dynamic convolution, DIN encodes a style image into learnable convolution parameters, upon which the content image is stylized. Unlike conventional methods that use shared complex encoders to encode content and style, the proposed DIN introduces a sophisticated style encoder, yet comes with a compact and lightweight content encoder for fast inference. Experimental results demonstrate that the proposed approach yields very encouraging results on challenging style patterns and, to our best knowledge, for the first time enables an arbitrary style transfer using MobileNet-based lightweight architecture, leading to a reduction factor of more than twenty in computational cost as compared to existing approaches. Furthermore, the proposed DIN provides flexible support for state-of-the-art convolutional operations, and thus triggers novel functionalities, such as uniform-stroke placement for non-natural images and automatic spatial-stroke control. Yongcheng Jing, Xiao Liu 0022, Yukang Ding, Xinchao Wang, Errui Ding, Mingli Song, Shilei Wen |
AAAI | 2 |
| 2020 | Deep Concept-wise Temporal Convolutional Networks for Action LocalizationabstractExisting action localization approaches adopt shallow temporal convolutional networks (i.e., TCN) on 1D feature map extracted from video frames. In this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which generally are highly abstract and can be regarded as latent concepts, are excessively recombined in temporal convolution. To address this issue, we introduce a novel concept-wise temporal convolutional network (C-TCN) as an alternative to TCN for training deeper action localization networks. To address this issue, we introduce a novel concept-wise temporal convolution (CTC) layer as an alternative to conventional temporal convolution layer for training deeper action localization networks. Instead of recombining latent concepts, CTC layer deploys a number of temporal filters to each concept separately with shared filter parameters across concepts. Thus can capture common temporal patterns of different concepts and significantly enrich representation ability. Via stacking CTC layers, we proposed a deep concept-wise temporal convolutional network (C-TCN), which boosts the state-of-the-art action localization performance on THUMOS'14 from 42.8 to 52.1 in terms of mAP(%), achieving a relative improvement of 21.7%. Favorable result is also obtained on ActivityNet. Xin Li 0106, Xiao Liu 0022, Wangmeng Zuo, Chao Li 0034, Xiang Long, Dongliang He, Fu Li 0003, Shilei Wen, Chuang Gan 0001 |
ACM Multimedia | 3 |
| 2019 | StNet: Local and Global Spatial-Temporal Modeling for Action RecognitionabstractDespite the success of deep learning for static image understanding, it remains unclear what are the most effective network architectures for spatial-temporal modeling in videos. In this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatialtemporal network (StNet) architecture for both local and global modeling in videos. Particularly, StNet stacks N successive video frames into a super-image which has 3N channels and applies 2D convolution on super-images to capture local spatial-temporal relationship. To model global spatialtemporal structure, we apply temporal convolution on the local spatial-temporal feature maps. Specifically, a novel temporal Xception block is proposed in StNet, which employs a separate channel-wise and temporal-wise convolution over the feature sequence of a video. Extensive experiments on the Kinetics dataset demonstrate that our framework outperforms several state-of-the-art approaches in action recognition and can strike a satisfying trade-off between recognition accuracy and model complexity. We further demonstrate the generalization performance of the leaned video representations on the UCF101 dataset. Dongliang He, Chuang Gan 0001, Fu Li 0003, Xiao Liu 0022, Yandong Li, Limin Wang 0002, Shilei Wen |
AAAI | 5 |
| 2019 | Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in VideosabstractThe task of video grounding, which temporally localizes a natural language description in a video, plays an important role in understanding videos. Existing studies have adopted strategies of sliding window over the entire video or exhaustively ranking all possible clip-sentence pairs in a presegmented video, which inevitably suffer from exhaustively enumerated candidates. To alleviate this problem, we formulate this task as a problem of sequential decision making by learning an agent which regulates the temporal grounding boundaries progressively based on its policy. Specifically, we propose a reinforcement learning based framework improved by multi-task learning and it shows steady performance gains by considering additional supervised boundary information during training. Our proposed framework achieves state-of-the-art performance on ActivityNet’18 DenseCaption dataset (Krishna et al. 2017) and Charades-STA dataset (Sigurdsson et al. 2016; Gao et al. 2017) while observing only 10 or less clips per video. Dongliang He, Jizhou Huang, Fu Li 0003, Xiao Liu 0022, Shilei Wen |
AAAI | 5 |
| 2019 | STGAN: A Unified Selective Transfer Network for Arbitrary Image Attribute EditingabstractArbitrary attribute editing generally can be tackled by incorporating encoder-decoder and generative adversarial networks. However, the bottleneck layer in encoder-decoder usually gives rise to blurry and low quality editing result. And adding skip connections improves image quality at the cost of weakened attribute manipulation ability. Moreover, existing methods exploit target attribute vector to guide the flexible translation to desired target domain. In this work, we suggest to address these issues from selective transfer perspective. Considering that specific editing task is certainly only related to the changed attributes instead of all target attributes, our model selectively takes the difference between target and source attribute vectors as input. Furthermore, selective transfer units are incorporated with encoder-decoder to adaptively select and modify encoder feature for enhanced attribute editing. Experiments show that our method (i.e., STGAN) simultaneously improves attribute manipulation accuracy as well as perception quality, and performs favorably against state-of-the-arts in arbitrary face attribute editing and season translation. Ming Liu 0018, Yukang Ding, Xiao Liu 0022, Errui Ding, Wangmeng Zuo, Shilei Wen |
CVPR | 4 |
| 2019 | BMN: Boundary-Matching Network for Temporal Action Proposal GenerationabstractTemporal action proposal generation is an challenging and promising task which aims to locate temporal regions in real-world videos where action or event may occur. Current bottom-up proposal generation methods can generate proposals with precise boundary, but cannot efficiently generate adequately reliable confidence scores for retrieving proposals. To address these difficulties, we introduce the Boundary-Matching (BM) mechanism to evaluate confidence scores of densely distributed proposals, which denote a proposal as a matching pair of starting and ending boundaries and combine all densely distributed BM pairs into the BM confidence map. Based on BM mechanism, we propose an effective, efficient and end-to-end proposal generation method, named Boundary-Matching Network (BMN), which generates proposals with precise temporal boundaries as well as reliable confidence scores simultaneously. The two-branches of BMN are jointly trained in an unified framework. We conduct experiments on two challenging datasets: THUMOS-14 and ActivityNet-1.3, where BMN shows significant performance improvement with remarkable efficiency and generalizability. Further, combining with existing action classifier, BMN can achieve state-of-the-art temporal action detection performance. Xiao Liu 0022, Xin Li 0106, Errui Ding, Shilei Wen |
ICCV | 2 |
| 2019 | Image Inpainting With Learnable Bidirectional Attention MapsabstractMost convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been suggested to address this issue, but it adopts handcrafted feature re-normalization, and only considers forward mask-updating. In this paper, we present a learnable attention map module for learning feature re-normalization and mask-updating in an end-to-end manner, which is effective in adapting to irregular holes and propagation of convolution layers. Furthermore, learnable reverse attention maps are introduced to allow the decoder of U-Net to concentrate on filling in irregular holes instead of reconstructing both holes and known regions, resulting in our learnable bidirectional attention maps. Qualitative and quantitative experiments show that our method performs favorably against state-of-the-arts in generating sharper, more coherent and visually plausible inpainting results. The source code and pre-trained models will be available at: https://github.com/Vious/LBAM_inpainting/. Chaohao Xie, Shaohui Liu, Chao Li 0034, Ming-Ming Cheng, Wangmeng Zuo, Xiao Liu 0022, Shilei Wen, Errui Ding |
ICCV | 6 |
| 2018 | Multimodal Keyless Attention Fusion for Video ClassificationabstractThe problem of video classification is inherently sequential and multimodal, and deep neural models hence need to capture and aggregate the most pertinent signals for a given input video. We propose Keyless Attention as an elegant and efficient means to more effectively account for the sequential nature of the data. Moreover, comparing a variety of multimodal fusion methods, we find that Multimodal Keyless Attention Fusion is the most successful at discerning interactions between modalities. We experiment on four highly heterogeneous datasets, UCF101, ActivityNet, Kinetics, and YouTube-8M to validate our conclusion, and show that our approach achieves highly competitive results. Especially on large-scale data, our method has great advantages in efficiency and performance. Most remarkably, our best single model can achieve 77.0% in terms of the top-1 accuracy and 93.2% in terms of the top-5 accuracy on the Kinetics validation set, and achieve 82.2% in terms of GAP@20 on the official YouTube-8M test set. Xiang Long, Chuang Gan 0001, Gerard de Melo, Xiao Liu 0022, Yandong Li, Fu Li 0003, Shilei Wen |
AAAI | 4 |
| 2018 | Attention Clusters: Purely Attention Based Local Feature Integration for Video ClassificationabstractRecently, substantial research effort has focused on how to apply CNNs or RNNs to better capture temporal patterns in videos, so as to improve the accuracy of video classification. In this paper, however, we show that temporal information, especially longer-term patterns, may not be necessary to achieve competitive results on common trimmed video classification datasets. We investigate the potential of a purely attention based local feature integration. Accounting for the characteristics of such features in video classification, we propose a local feature integration framework based on attention clusters, and introduce a shifting operation to capture more diverse signals. We carefully analyze and compare the effect of different attention mechanisms, cluster sizes, and the use of the shifting operation, and also investigate the combination of attention clusters for multimodal integration. We demonstrate the effectiveness of our framework on three real-world video classification datasets. Our model achieves competitive results across all of these. In particular, on the large-scale Kinetics dataset, our framework obtains an excellent single model accuracy of 79.4% in terms of the top-1 and 94.0% in terms of the top-5 accuracy on the validation set. Xiang Long, Chuang Gan 0001, Gerard de Melo, Jiajun Wu 0001, Xiao Liu 0022, Shilei Wen |
CVPR | 5 |
| 2018 | Fine-Grained Video Categorization with Redundancy Reduction Attention
Xiao Tan 0001, Feng Zhou 0002, Xiao Liu 0022, Kaiyu Yue, Errui Ding |
ECCV (5) | 4 |
| 2017 | Localizing by Describing: Attribute-Guided Attention Localization for Fine-Grained RecognitionabstractA key challenge in fine-grained recognition is how to find and represent discriminative local regions.Recent attention models are capable of learning discriminative region localizers only from category labels with reinforcement learning. However, not utilizing any explicit part information, they are not able to accurately find multiple distinctive regions.In this work, we introduce an attribute-guided attention localization scheme where the local region localizers are learned under the guidance of part attribute descriptions.By designing a novel reward strategy, we are able to learn to locate regions that are spatially and semantically distinctive with reinforcement learning algorithm. The attribute labeling requirement of the scheme is more amenable than the accurate part location annotation required by traditional part-based fine-grained recognition methods.Experimental results on the CUB-200-2011 dataset demonstrate the superiority of the proposed scheme on both fine-grained recognition and attribute recognition. Xiao Liu 0022, Shilei Wen, Errui Ding, Yuanqing Lin |
AAAI | 1 |
| 2017 | Kernel Pooling for Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) with Bilinear Pooling, initially in their full form and later using compact representations, have yielded impressive performance gains on a wide range of visual tasks, including fine-grained visual categorization, visual question answering, face recognition, and description of texture and style. The key to their success lies in the spatially invariant modeling of pairwise (2ndorder) feature interactions. In this work, we propose a general pooling framework that captures higher order interactions of features in the form of kernels. We demonstrate how to approximate kernels such as Gaussian RBF up to a given order using compact explicit feature maps in a parameter-free manner. Combined with CNNs, the composition of the kernel can be learned from data in an end-to-end fashion via error back-propagation. The proposed kernel pooling scheme is evaluated in terms of both kernel approximation error and visual recognition accuracy. Experimental evaluations demonstrate state-of-the-art performance on commonly used fine-grained recognition datasets. Yin Cui, Feng Zhou 0002, Xiao Liu 0022, Yuanqing Lin, Serge J. Belongie |
CVPR | 4 |
| 2017 | Deep Metric Learning with Angular LossabstractThe modern image search system requires semantic understanding of image, and a key yet under-addressed problem is to learn a good metric for measuring the similarity between images. While deep metric learning has yielded impressive performance gains by extracting high level abstractions from image data, a proper objective loss function becomes the central issue to boost the performance. In this paper, we propose a novel angular loss, which takes angle relationship into account, for learning better similarity metric. Whereas previous metric learning methods focus on optimizing the similarity (contrastive loss) or relative similarity (triplet loss) of image pairs, our proposed method aims at constraining the angle at the negative point of triplet triangles. Several favorable properties are observed when compared with conventional methods. First, scale invariance is introduced, improving the robustness of objective against feature variance. Second, a third-order geometric constraint is inherently imposed, capturing additional local structure of triplet triangles than contrastive loss or triplet loss. Third, better convergence has been demonstrated by experiments on three publicly available datasets. Feng Zhou 0002, Shilei Wen, Xiao Liu 0022, Yuanqing Lin |
ICCV | 4 |