Fu Li 0003

dblp:37/4556-3 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
19since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 22 · 2 first-author · 15 since 2021
YearPublicationVenuePosition
2023 AdaCM: Adaptive ColorMLP for Real-Time Universal Photo-Realistic Style Transfer
abstract
Photo-realistic style transfer aims at migrating the artistic style from an exemplar style image to a content image, producing a result image without spatial distortions or unrealistic artifacts. Impressive results have been achieved by recent deep models. However, deep neural network based methods are too expensive to run in real-time. Meanwhile, bilateral grid based methods are much faster but still contain artifacts like overexposure. In this work, we propose the Adaptive ColorMLP (AdaCM), an effective and efficient framework for universal photo-realistic style transfer. First, we find the complex non-linear color mapping between input and target domain can be efficiently modeled by a small multi-layer perceptron (ColorMLP) model. Then, in AdaCM, we adopt a CNN encoder to adaptively predict all parameters for the ColorMLP conditioned on each input content and style image pair. Experimental results demonstrate that AdaCM can generate vivid and high-quality stylization results. Meanwhile, our AdaCM is ultrafast and can process a 4K resolution image in 6ms on one V100 GPU.
Honglin Lin, Fu Li 0003, Dongliang He
AAAI3
2023 Master: Meta Style Transformer for Controllable Zero-Shot and Few-Shot Artistic Style Transfer
abstract
Transformer-based models achieve favorable performance in artistic style transfer recently thanks to its global receptive field and powerful multi-head/layer attention operations. Nevertheless, the over-paramerized multi-layer structure increases parameters significantly and thus presents a heavy burden for training. Moreover, for the task of style transfer, vanilla Transformer that fuses content and style features by residual connections is prone to content-wise distortion. In this paper, we devise a novel Transformer model termed as Master specifically for style transfer. On the one hand, in the proposed model, different Transformer layers share a common group of parameters, which (1) reduces the total number of parameters, (2) leads to more robust training convergence, and (3) is readily to control the degree of stylization via tuning the number of stacked layers freely during inference. On the other hand, different from the vanilla version, we adopt a learnable scaling operation on content features before content-style feature interaction, which better preserves the original similarity between a pair of content features while ensuring the stylization quality. We also propose a novel meta learning scheme for the proposed model so that it can not only work in the typical setting of arbitrary style transfer, but also adaptable to the few-shot setting, by only fine-tuning the Transformer encoder layer in the few-shot stage for one specific style. Text-guided few-shot style transfer is firstly achieved with the proposed framework. Extensive experiments demonstrate the superiority of Master under both zero-shot and few-shot style transfer settings.
Hao Tang 0005, Songhua Liu, Shaoli Huang, Fu Li 0003, Dongliang He, Xinchao Wang
CVPR5
2023 LMR: A Large-Scale Multi-Reference Dataset for Reference-based Super-Resolution
abstract
It is widely agreed that reference-based super-resolution (RefSR) achieves superior results by referring to similar high quality images, compared to single image super-resolution (SISR). Intuitively, the more references, the better performance. However, previous RefSR methods have all focused on single-reference image training, while multiple reference images are often available in testing or practical applications. The root cause of such training-testing mismatch is the absence of publicly available multi-reference SR training datasets, which greatly hinders research efforts on multi-reference super-resolution. To this end, we construct a large-scale, multi-reference super-resolution dataset, named LMR. It contains 112, 142 groups of 300×300 training images, which is 10× of the existing largest RefSR dataset. The image size is also some times larger. More importantly, each group is equipped with 5 reference images with different similarity levels. Furthermore, we propose a new baseline method for multi-reference super-resolution: MRefSR, including a Multi-Reference Attention Module (MAM) for feature fusion of an arbitrary number of reference images, and a Spatial Aware Filtering Module (SAFM) for the fused feature selection. The proposed MRefSR achieves significant improvements over state-of-the-art approaches on both quantitative and qualitative evaluations. Our code and data are available at: https://github.com/wdmwhh/MRefSR.
Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Errui Ding, Zhaoxiang Zhang 0001
ICCV4
2023 Adversarial Dual-Student With Differentiable Spatial Warping for Semi-Supervised Semantic Segmentation
abstract
A common challenge posed to robust semantic segmentation is the expensive data annotation cost. Existing semi-supervised solutions show great potential for solving this problem. Their key idea is constructing consistency regularization with unsupervised data augmentation from unlabeled data for model training. The perturbations for unlabeled data enable the consistency training loss, which benefits semi-supervised semantic segmentation. However, these perturbations destroy image context and introduce unnatural boundaries, which is harmful for semantic segmentation. Besides, the widely adopted semi-supervised learning framework, i.e. mean-teacher, suffers performance limitation since the student model finally converges to the teacher model. In this paper, first of all, we propose a context friendly differentiable geometric warping to conduct unsupervised data augmentation; secondly, a novel adversarial dual-student framework is proposed to improve the Mean-Teacher from the following two aspects: (1) dual student models are learned independently except for a stabilization constraint to encourage exploiting model diversities; (2) adversarial training scheme is applied to both students and the discriminators are resorted to distinguish reliable pseudo-label of unlabeled data for self-training. Effectiveness is validated via extensive experiments on PASCAL VOC2012 and Cityscapes. Our solution significantly improves the performance and state-of-the-art results are achieved on both datasets. Remarkably, compared with fully supervision, our solution achieves comparable mIoU of 73.4% using only 12.5% annotated data on PASCAL VOC2012. Our codes and models are available athttps://github.com/cao-cong/ADS-SemiSeg.
Cong Cao 0005, Dongliang He, Fu Li 0003, Huanjing Yue, Jing-Yu Yang 0002, Errui Ding
IEEE Trans. Circuits Syst. Video Technol.4
2023 Dual-Affinity Style Embedding Network for Semantic-Aligned Image Style Transfer
abstract
Image style transfer aims at synthesizing an image with the content from one image and the style from another. User studies have revealed that the semantic correspondence between style and content greatly affects subjective perception of style transfer results. While current studies have made great progress in improving the visual quality of stylized images, most methods directly transfer global style statistics without considering semantic alignment. Current semantic style transfer approaches still work in an iterative optimization fashion, which is impractically computationally expensive. Addressing these issues, we introduce a novel dual-affinity style embedding network (DaseNet) to synthesize images with style aligned at semantic region granularity. In the dual-affinity module, feature correlation and semantic correspondence between content and style images are modeled jointly for embedding local style patterns according to semantic distribution. Furthermore, the semantic-weighted style loss and the region-consistency loss are introduced to ensure semantic alignment and content preservation. With the end-to-end network architecture, DaseNet can well balance visual quality and inference efficiency for semantic style transfer. Experimental results on different scene categories have demonstrated the effectiveness of the proposed method.
Zhuoqi Ma, Xin Li 0106, Fu Li 0003, Dongliang He, Errui Ding, Nannan Wang 0001, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.4
2022 Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model
abstract
To achieve disentangled image manipulation, previous works depend heavily on manual annotation. Meanwhile, the available manipulations are limited to a pre-defined set the models were trainedfor. We propose a novelframework, i.e., Predict, Prevent, and Evaluate (PPE), for disentangled text-driven image manipulation that requires little manual annotation while being applicable to a wide variety of ma-nipulations. Our method approaches the targets by deeply exploiting the power of the large-scale pre-trained vision-language model CLIP [32]. Concretely, we firstly Predict the possibly entangled attributes for a given text command. Then, based on the predicted attributes, we introduce an entanglement loss to Prevent entanglements during training. Finally, we propose a new evaluation metric to Evaluate the disentangled image manipulation. We verify the effectiveness of our method on the challenging face editing task. Extensive experiments show that the proposed PPE frame-work achieves much better quantitative and qualitative re-sults than the up-to-date StyleCLIP [31] baseline. Code is available at https://github.com/zipengxuc/PPE.
Zipeng Xu, Hao Tang 0005, Fu Li 0003, Dongliang He, Nicu Sebe, Radu Timofte, Luc Van Gool, Errui Ding
CVPR4
2022 CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
Haoran Wang 0004, Dongliang He, Boyang Xia, Fu Li 0003, Zhong Ji, Errui Ding, Jingdong Wang 0001
ECCV (36)6
2022 Neural Color Operators for Sequential Image Retouching
Yili Wang 0003, Xin Li 0106, Kun Xu 0003, Dongliang He, Qi Zhang 0029, Fu Li 0003, Errui Ding
ECCV (19)6
2022 RRSR: Reciprocal Reference-Based Image Super-Resolution with Progressive Feature Alignment and Selection
Lin Zhang 0013, Xin Li 0106, Dongliang He, Fu Li 0003, Yili Wang 0003, Zhaoxiang Zhang 0001
ECCV (19)4
2022 Boosting Video-Text Retrieval with Explicit High-Level Semantics
abstract
Video-text retrieval (VTR) is an attractive yet challenging task for multi-modal understanding, which aims to search for relevant video (text) given a query (video). Existing methods typically employ completely heterogeneous visual-textual information to align video and text, whilst lacking the awareness of homogeneous high-level semantic information residing in both modalities. To fill this gap, in this work, we propose a novel visual-linguistic aligning model named HiSE for VTR, which improves the cross-modal representation by incorporating explicit high-level semantics. First, we explore the hierarchical property of explicit high-level semantics, and further decompose it into two levels, i.e. discrete semantics and holistic semantics. Specifically, for visual branch, we exploit an off-the-shelf semantic entity predictor to generate discrete high-level semantics. In parallel, a trained video captioning model is employed to output holistic high-level semantics. As for the textual modality, we parse the text into three parts including occurrence, action and entity. In particular, the occurrence corresponds to the holistic high-level semantics, meanwhile both action and entity represent the discrete ones. Then, different graph reasoning techniques are utilized to promote the interaction between holistic and discrete high-level semantics. Extensive experiments demonstrate that, with the aid of explicit high-level semantics, our method achieves the superior performance over state-of-the-art methods on three benchmark datasets, including MSR-VTT, MSVD and DiDeMo.
Haoran Wang 0004, Dongliang He, Fu Li 0003, Zhong Ji, Jungong Han, Errui Ding
ACM Multimedia4
2022 Purely Attention Based Local Feature Integration for Video Classification
abstract
Recently, substantial research effort has focused on how to apply CNNs or RNNs to better capture temporal patterns in videos, so as to improve the accuracy of video classification. In this paper, we investigate the potential of a purely attention based local feature integration. Accounting for the characteristics of such features in video classification, we first propose Basic Attention Clusters (BAC), which concatenates the output of multiple attention units applied in parallel, and introduce a shifting operation to capture more diverse signals. Experiments show that BAC can achieve excellent results on multiple datasets. However, BAC treats all feature channels as an indivisible whole, which is suboptimal for achieving a finer-grained local feature integration over the channel dimension. Additionally, it treats the entire local feature sequence as an unordered set, thus ignoring the sequential relationships. To improve over BAC, we further propose the channel pyramid attention schema by splitting features into sub-features at multiple scales for coarse-to-fine sub-feature interaction modeling, and propose the temporal pyramid attention schema by dividing the feature sequences into ordered sub-sequences of multiple lengths to account for the sequential order. Our final model pyramid×pyramid attention clusters (PPAC) combines both channel pyramid attention and temporal pyramid attention to focus on the most important sub-features, while also preserving the temporal information of the video. We demonstrate the effectiveness of PPAC on seven real-world video classification datasets. Our model achieves competitive results across all of these, showing that our proposed framework can consistently outperform the existing local feature integration methods across a range of different scenarios.
Xiang Long, Gerard de Melo, Dongliang He, Fu Li 0003, Zhizhen Chi, Shilei Wen, Chuang Gan 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Semi-Supervised Temporal Action Proposal Generation via Exploiting 2-D Proposal Map
abstract
Temporal action proposal generation aims to generate temporal video segments containing human actions in untrimmed videos, which is always a preliminary for such video understanding tasks as action localization and temporally description grounding,etc. Fully-supervised solutions, though proven to be effective, suffer much from heavy data annotation overhead. To address this problem, this paper focuses on a rarely investigated yet practical problem of semi-supervised learning for temporal action proposal generation. Firstly, we propose aProposal Map oriented Mean-Teacher(PM-MT) model, which can use both labeled and unlabeled data for end-to-end model training. Secondly, aSuppression-and-Re-Generation(SRG) strategy is designed to generate high-quality pseudo labels for unlabeled data, which are then used to finetune the model. Extensive experiments demonstrate the effectiveness of our proposed method, by achieving the state-of-the-art results on two public benchmark datatsets on the task of semi-supervised action proposal generation and outperforming fully-supervised learning methods with only a portion of labeled data.
Weining Wang 0001, Dongliang He, Fu Li 0003, Shilei Wen, Liang Wang 0001, Jing Liu 0001
IEEE Trans. Multim.4
2021 MVFNet: Multi-View Fusion Network for Efficient Video Recognition
abstract
Conventionally, spatiotemporal modeling network and its complexity are the two most concentrated research topics in video action recognition. Existing state-of-the-art methods have achieved excellent accuracy regardless of the complexity meanwhile efficient spatiotemporal modeling solutions are slightly inferior in performance. In this paper, we attempt to acquire both efficiency and effectiveness simultaneously. First of all, besides traditionally treating H x W x T video frames as space-time signal (viewing from the Height-Width spatial plane), we propose to also model video from the other two Height-Time and Width-Time planes, to capture the dynamics of video thoroughly. Secondly, our model is designed based on 2D CNN backbones and model complexity is well kept in mind by design. Specifically, we introduce a novel multi-view fusion (MVF) module to exploit video dynamics using separable convolution for efficiency. It is a plug-and-play module and can be inserted into off-the-shelf 2D CNNs to form a simple yet effective model called MVFNet. Moreover, MVFNet can be thought of as a generalized video modeling framework and it can specialize to be existing methods such as C2D, SlowOnly, and TSM under different settings. Extensive experiments are conducted on popular benchmarks (i.e., Something-Something V1 & V2, Kinetics, UCF-101, and HMDB-51) to show its superiority. The proposed MVFNet can achieve state-of-the-art performance with 2D CNN's complexity.
Dongliang He, Fu Li 0003, Chuang Gan 0001, Errui Ding
AAAI4
2021 Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer
abstract
Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize complex style, especially when holistic global and local patterns exist. Inspired by the common painting process of drawing a draft and revising the details, we introduce a novel feed-forward method named Laplacian Pyramid Network (LapStyle). LapStyle first transfers global style patterns in low-resolution via a Drafting Network. It then revises the local details in high-resolution via a Revision Network, which hallucinates a residual image according to the draft and the image textures extracted by Laplacian filtering. Higher resolution details can be easily generated by stacking Revision Networks with multiple Laplacian pyramid levels. The final stylized image is obtained by aggregating outputs of all pyramid levels. Experiments demonstrate that our method can synthesize high quality stylized images in real time, where holistic style patterns are properly transferred.
Zhuoqi Ma, Fu Li 0003, Dongliang He, Xin Li 0106, Errui Ding, Nannan Wang 0001, Jie Li 0001, Xinbo Gao 0001
CVPR3
2021 Learning Semantic Person Image Generation by Region-Adaptive Normalization
abstract
Human pose transfer has received great attention due to its wide applications, yet is still a challenging task that is not well solved. Recent works have achieved great success to transfer the person image from the source to the target pose. However, most of them cannot well capture the semantic appearance, resulting in inconsistent and less realistic textures on the reconstructed results. To address this issue, we propose a new two-stage framework to handle the pose and appearance translation. In the first stage, we predict the target semantic parsing maps to eliminate the difficulties of pose transfer and further benefit the latter translation of per-region appearance style. In the second one, with the predicted target semantic maps, we suggest a new person image generation method by incorporating the region-adaptive normalization, in which it takes the per-region styles to guide the target appearance generation. Extensive experiments show that our proposed SPGNet can generate more semantic, consistent, and photorealistic results and perform favorably against the state of the art methods in terms of quantitative and qualitative evaluation. The source code and model are available at https://github.com/cszy98/SPGNet.git.
Zhengyao Lv, Xiaoming Li 0002, Xin Li 0106, Fu Li 0003, Dongliang He, Wangmeng Zuo
CVPR4
2021 Paint Transformer: Feed Forward Neural Painting with Stroke Prediction
abstract
Neural painting refers to the procedure of producing a series of strokes for a given image and non-photo-realistically recreating it using neural networks. While reinforcement learning (RL) based agents can generate a stroke sequence step by step for this task, it is not easy to train a stable RL agent. On the other hand, stroke optimization methods search for a set of stroke parameters iteratively in a large search space; such low efficiency significantly limits their prevalence and practicality. Different from previous methods, in this paper, we formulate the task as a set prediction problem and propose a novel Transformer-based framework, dubbed Paint Transformer, to predict the parameters of a stroke set with a feed forward network. This way, our model can generate a set of strokes in parallel and obtain the final painting of size 512 × 512 in near real time. More importantly, since there is no dataset available for training the Paint Transformer, we devise a self-training pipeline such that it can be trained without any off-the-shelf dataset while still achieving excellent generalization capability. Experiments demonstrate that our method achieves better painting performance than previous ones with cheaper training and inference costs. Codes and models are available on https://github.com/wzmsltw/PaintTransformer.
Songhua Liu, Dongliang He, Fu Li 0003, Ruifeng Deng, Xin Li 0106, Errui Ding, Hao Wang 0014
ICCV4
2021 AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer
abstract
Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distributions, or adaptively normalize deep content feature according to the style such that their global statistics are matched. Although effective, leaving shallow feature unexplored and without locally considering feature statistics, they are prone to unnatural output with unpleasing local distortions. To alleviate this problem, in this paper, we propose a novel attention and normalization module, named Adaptive Attention Normalization (AdaAttN), to adaptively perform attentive normalization on per-point basis. Specifically, spatial attention score is learnt from both shallow and deep features of content and style images. Then perpoint weighted statistics are calculated by regarding a style feature point as a distribution of attention-weighted output of all style feature points. Finally, the content feature is normalized so that they demonstrate the same local feature statistics as the calculated per-point weighted style feature statistics. Besides, a novel local feature loss is derived based on AdaAttN to enhance local visual quality. We also extend AdaAttN to be ready for video style transfer with slight modifications. Experiments demonstrate that our method achieves state-of-the-art arbitrary image/video style transfer. Codes and models are available on https://github.com/wzmsltw/AdaAttN.
Songhua Liu, Dongliang He, Fu Li 0003, Xin Li 0106, Zhengxing Sun, Qian Li 0014, Errui Ding
ICCV4
2021 DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features
abstract
Image Retrieval is a fundamental task of obtaining images similar to the query one from a database. A common image retrieval practice is to firstly retrieve candidate images via similarity search using global image features and then re-rank the candidates by leveraging their local features. Previous learning-based studies mainly focus on either global or local image representation learning to tackle the retrieval task. In this paper, we abandon the two-stage paradigm and seek to design an effective single-stage solution by integrating local and global information inside images into compact image representations. Specifically, we propose a Deep Orthogonal Local and Global (DOLG) information fusion framework for end-to-end image retrieval. It attentively extracts representative local information with multi-atrous convolutions and self-attention at first. Components orthogonal to the global image representation are then extracted from the local information. At last, the orthogonal components are concatenated with the global representation as a complementary, and then aggregation is performed to generate the final representation. The whole framework is end-to-end differentiable and can be trained with image-level labels. Extensive experimental results validate the effectiveness of our solution and show that our model achieves state-of-the-art image retrieval performances on Revisited Oxford and Paris datasets.1
Dongliang He, Baorong Shi, Xuetong Xue, Fu Li 0003, Errui Ding, Jizhou Huang
ICCV6
2021 Image Inpainting by End-to-End Cascaded Refinement With Mask Awareness
abstract
Inpainting arbitrary missing regions is challenging because learning valid features for various masked regions is nontrivial. Though U-shaped encoder-decoder frameworks have been witnessed to be successful, most of them share a common drawback of mask unawareness in feature extraction because all convolution windows (or regions), including those with various shapes of missing pixels, are treated equally and filtered with fixed learned kernels. To this end, we propose our novel mask-aware inpainting solution. Firstly, a Mask-Aware Dynamic Filtering (MADF) module is designed to effectively learn multi-scale features for missing regions in the encoding phase. Specifically, filters for each convolution window are generated from features of the corresponding region of the mask. The second fold of mask awareness is achieved by adopting Point-wise Normalization (PN) in our decoding phase, considering that statistical natures of features at masked points differentiate from those of unmasked points. The proposed PN can tackle this issue by dynamically assigning point-wise scaling factor and bias. Lastly, our model is designed to be an end-to-end cascaded refinement one. Supervision information such as reconstruction loss, perceptual loss and total variation loss is incrementally leveraged to boost the inpainting results from coarse to fine. Effectiveness of the proposed framework is validated both quantitatively and qualitatively via extensive experiments on three public datasets including Places2, CelebA and Paris StreetView.
Manyu Zhu, Dongliang He, Xin Li 0106, Chao Li 0034, Fu Li 0003, Xiao Liu 0022, Errui Ding, Zhaoxiang Zhang 0001
IEEE Trans. Image Process.5
2020 Multi-Label Classification with Label Graph Superimposing
abstract
Images or videos always contain multiple objects or actions. Multi-label recognition has been witnessed to achieve pretty performance attribute to the rapid development of deep learning technologies. Recently, graph convolution network (GCN) is leveraged to boost the performance of multi-label recognition. However, what is the best way for label correlation modeling and how feature learning can be improved with label system awareness are still unclear. In this paper, we propose a label graph superimposing framework to improve the conventional GCN+CNN framework developed for multi-label recognition in the following two aspects. Firstly, we model the label correlations by superimposing label graph built from statistical co-occurrence information into the graph constructed from knowledge priors of labels, and then multi-layer graph convolutions are applied on the final superimposed graph for label embedding abstraction. Secondly, we propose to leverage embedding of the whole label system for better representation learning. In detail, lateral connections between GCN and CNN are added at shallow, middle and deep layers to inject information of label system into backbone CNN for label-awareness in the feature learning process. Extensive experiments are carried out on MS-COCO and Charades datasets, showing that our proposed solution can greatly improve the recognition performance and achieves new state-of-the-art recognition performance.
Ya Wang 0002, Dongliang He, Fu Li 0003, Xiang Long, Jinwen Ma, Shilei Wen
AAAI3
2020 Deep Concept-wise Temporal Convolutional Networks for Action Localization
abstract
Existing action localization approaches adopt shallow temporal convolutional networks (i.e., TCN) on 1D feature map extracted from video frames. In this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which generally are highly abstract and can be regarded as latent concepts, are excessively recombined in temporal convolution. To address this issue, we introduce a novel concept-wise temporal convolutional network (C-TCN) as an alternative to TCN for training deeper action localization networks. To address this issue, we introduce a novel concept-wise temporal convolution (CTC) layer as an alternative to conventional temporal convolution layer for training deeper action localization networks. Instead of recombining latent concepts, CTC layer deploys a number of temporal filters to each concept separately with shared filter parameters across concepts. Thus can capture common temporal patterns of different concepts and significantly enrich representation ability. Via stacking CTC layers, we proposed a deep concept-wise temporal convolutional network (C-TCN), which boosts the state-of-the-art action localization performance on THUMOS'14 from 42.8 to 52.1 in terms of mAP(%), achieving a relative improvement of 21.7%. Favorable result is also obtained on ActivityNet.
Xin Li 0106, Xiao Liu 0022, Wangmeng Zuo, Chao Li 0034, Xiang Long, Dongliang He, Fu Li 0003, Shilei Wen, Chuang Gan 0001
ACM Multimedia8
2019 StNet: Local and Global Spatial-Temporal Modeling for Action Recognition
abstract
Despite the success of deep learning for static image understanding, it remains unclear what are the most effective network architectures for spatial-temporal modeling in videos. In this paper, in contrast to the existing CNN+RNN or pure 3D convolution based approaches, we explore a novel spatialtemporal network (StNet) architecture for both local and global modeling in videos. Particularly, StNet stacks N successive video frames into a super-image which has 3N channels and applies 2D convolution on super-images to capture local spatial-temporal relationship. To model global spatialtemporal structure, we apply temporal convolution on the local spatial-temporal feature maps. Specifically, a novel temporal Xception block is proposed in StNet, which employs a separate channel-wise and temporal-wise convolution over the feature sequence of a video. Extensive experiments on the Kinetics dataset demonstrate that our framework outperforms several state-of-the-art approaches in action recognition and can strike a satisfying trade-off between recognition accuracy and model complexity. We further demonstrate the generalization performance of the leaned video representations on the UCF101 dataset.
Dongliang He, Chuang Gan 0001, Fu Li 0003, Xiao Liu 0022, Yandong Li, Limin Wang 0002, Shilei Wen
AAAI4
2019 Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
abstract
The task of video grounding, which temporally localizes a natural language description in a video, plays an important role in understanding videos. Existing studies have adopted strategies of sliding window over the entire video or exhaustively ranking all possible clip-sentence pairs in a presegmented video, which inevitably suffer from exhaustively enumerated candidates. To alleviate this problem, we formulate this task as a problem of sequential decision making by learning an agent which regulates the temporal grounding boundaries progressively based on its policy. Specifically, we propose a reinforcement learning based framework improved by multi-task learning and it shows steady performance gains by considering additional supervised boundary information during training. Our proposed framework achieves state-of-the-art performance on ActivityNet’18 DenseCaption dataset (Krishna et al. 2017) and Charades-STA dataset (Sigurdsson et al. 2016; Gao et al. 2017) while observing only 10 or less clips per video.
Dongliang He, Jizhou Huang, Fu Li 0003, Xiao Liu 0022, Shilei Wen
AAAI4
2019 Predicting Rate Control Target Through A Learning Based Content Adaptive Model
abstract
Rate Control (RC) plays an important role in video encoding. Traditional solutions are using fixed rate or fixed quantization parameters as the unified rate-control targets for all videos in one given video application. However, unified rate-control targets tend to have some bad encoding cases because of applying wrong rate for the video content. In this paper, we propose one content-adaptive rate control solution. We employ one neural-network based model which can end-to-end learn the optimal rate-control target appropriate to the content characteristics. The experimental results show that the proposed model can predict the optimal rate-factor value with the accuracy up to 77.637%. With this model, the proposed video-encoding method can significantly decrease the encoding quality fluctuation.
Huaifei Xing, Huifeng Shen, Dongliang He, Fu Li 0003
PCS6
2018 Multimodal Keyless Attention Fusion for Video Classification
abstract
The problem of video classification is inherently sequential and multimodal, and deep neural models hence need to capture and aggregate the most pertinent signals for a given input video. We propose Keyless Attention as an elegant and efficient means to more effectively account for the sequential nature of the data. Moreover, comparing a variety of multimodal fusion methods, we find that Multimodal Keyless Attention Fusion is the most successful at discerning interactions between modalities. We experiment on four highly heterogeneous datasets, UCF101, ActivityNet, Kinetics, and YouTube-8M to validate our conclusion, and show that our approach achieves highly competitive results. Especially on large-scale data, our method has great advantages in efficiency and performance. Most remarkably, our best single model can achieve 77.0% in terms of the top-1 accuracy and 93.2% in terms of the top-5 accuracy on the Kinetics validation set, and achieve 82.2% in terms of GAP@20 on the official YouTube-8M test set.
Xiang Long, Chuang Gan 0001, Gerard de Melo, Xiao Liu 0022, Yandong Li, Fu Li 0003, Shilei Wen
AAAI6
2018 Subspace Clustering With K-Support Norm
abstract
Subspace clustering aims to cluster a collection of data points lying in a union of subspaces. Based on the assumption that each point can be approximately represented as a linear combination of other points, extensive efforts have been made to compute an affinity matrix in a self-expressive framework for describing the similarity between points. However, the existing clustering methods consider the average feature solutions, which would not be powerful enough to capture the intrinsic relationship between points. In this paper, we present the k-support norm subspace clustering (KSC) method by utilizing k-support norm regularization. The k-support norm trades off the sparsity of ℓ1norm and the uniform shrinkage of ℓ2norm to yield better predictive performance on the data connection. The theoretical analysis of KSC makes up a large proportion of paper. In the noise-free case, we provide the kEBD condition, which ensures the coefficient matrix to be block diagonal. If the data are corrupted, we prove the incompletion-grouping effect for KSC. Moreover, we provide the statistical recovery guarantee for both noise-free and noise cases. The theory analyses show the validity and feasibility of KSC, and the experimental results on multiple challenging databases demonstrate the effectiveness of the proposed algorithm.
Baohua Li, Huchuan Lu, Fu Li 0003, Wei Wu 0010
IEEE Trans. Circuits Syst. Video Technol.3
2017 Visual tracking with structured patch-based model
Fu Li 0003, Xu Jia 0012, Cheng Xiang 0001, Huchuan Lu
Image Vis. Comput.1
2017 Visual Tracking via Joint Discriminative Appearance Learning
abstract
In this paper, we present a discriminative tracking method based on dictionary learning and support vector machine (SVM) classification, where the dictionary and the classifier are jointly learned within a unified objective function. A discriminative differential tracking method is proposed, which estimates the motion parameters iteratively by the gradient-based method to maximize the SVM classification score, leading the bounding box to move purposively. As the target appearance may change across frames, an online update scheme is exploited, which not only reserves the discriminative information, but also adaptively accounts for the appearance changes in the dynamic scenes. We examine the proposed method on the benchmark challenging image sequences, including heavy occlusion, pose change, illumination variation, and so on. Extensive evaluations demonstrate that the proposed tracker performs favorably against other state-of-the-art algorithms.
Fu Li 0003, Huchuan Lu, Gang Hua 0001
IEEE Trans. Circuits Syst. Video Technol.2
2016 Human running detection: Benchmark and baseline
Shihong Lao, Dong Wang 0004, Fu Li 0003, Haihong Zhang
Comput. Vis. Image Underst.3
2016 Dual Group Structured Tracking
abstract
The sparse representation (SR)-based tracking framework generally considers the testing candidates and dictionary atoms individually, thus failing to model the structured information within data. In this paper, we present a robust tracking framework by exploiting the dual group structure of both candidate samples and dictionary templates, and formulate the SR at group level. The similar samples are encoded simultaneously by a few atom groups, which induces the inter-group sparsity, and also each group enjoys different internal sparsity. In this way, not only the potential commonality shared by the related candidates is taken into account but also the individual differences between samples are reflected. Then, we provide two effective optimization methods to solve our formulation by block-coordinate gradient descent and alternating direction method of multipliers, respectively, and make a comparison between them in terms of both effectiveness and efficiency. Finally, we embed the dual group structure model into the particle filter framework for visual tracking. Extensive experimental results demonstrate that our tracker achieves favorable performance against the state-of-the-art tracking methods.
Fu Li 0003, Huchuan Lu, Dong Wang 0004, Yi Wu 0001, Kaihua Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2015 Visual tracking via guided filter
abstract
In this paper, we propose a novel tracking algorithm based on an explicit image filter - guided filter. The guided filter utilizes the structure in the guidance image and performs as an edge-preserving smoothing operator. In our work, we treat the target as the guidance and the incoming candidates are filtered depending on the similarity between the guidance image and each input. The edge-preserving smoothing property depending on the guidance image is a critical advantage for object tracking. First, the guided filter can help to pick out the valuable candidates and make the inaccurate ones blurry so that the tracker can distinguish the target from numerous bad candidates easily. Besides, the filtering process can recover the content of the target being occluded according to the guidance image, which can help to alleviate the drifting problem effectively. Eventually, to generate a robust tracker, we take advantage of the combination of positive and negative templates to conduct effective sparse representation. Experimental results show that our algorithm outperforms relative trackers.
Dandan Du, Huchuan Lu, Lihe Zhang, Fu Li 0003
ICIP4
2014 Robust Visual Tracking with Dual Group Structure
Fu Li 0003, Huchuan Lu, Dong Wang 0004
ACCV (4)1