Li Shen 0005

dblp:91/3680-5 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
7since 2021 · last 2025
0000-0002-2283-4976ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
19 papers
3D vision · 19% Deep learning architectures and training · 16% Efficient and distributed learning · 12%
Theoretical computer science
2 papers
Mathematical optimization · 97% Algorithms and data structures · 3%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%

Topics — the 30 heaviest of 53, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
convolutional neural network
1.342020
Squeeze-and-Excitation Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks · NeurIPS 2018
Squeeze-and-Excitation Networks · CVPR 2018
Machine learning › Generative modeling › face synthesis
talking face generation
1.322024
DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Depth-Aware Generative Adversarial Network for Talking Head Video Generation · CVPR 2022
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
1.022022
AlphaGAN: Fully Differentiable Architecture Search for Generative Adversarial Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation · CVPR 2020
Computer vision › Image recognition and object detection
image classification
0.942018
Squeeze-and-Excitation Networks · CVPR 2018
Multi-Level Discriminative Dictionary Learning With Application to Large Scale Image Classification · IEEE Trans. Image Process. 2015
Adaptive Sharing for Image Classification · IJCAI 2015
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.912025
Image-to-video Adaptation with Outlier Modeling and Robust Self-learning · AAAI 2025
Machine learning › Transfer learning and domain adaptation › domain adaptation › visual domain adaptation
image-to-video adaptation
0.912025
Image-to-video Adaptation with Outlier Modeling and Robust Self-learning · AAAI 2025
Computer vision › Video understanding and tracking
video classification
0.912025
Image-to-video Adaptation with Outlier Modeling and Robust Self-learning · AAAI 2025
Machine learning › Deep learning architectures and training › attention mechanism › attention module
channel attention
0.822020
Squeeze-and-Excitation Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Squeeze-and-Excitation Networks · CVPR 2018
Computer vision › 3D vision
depth estimation
0.812024
DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Computer vision › 3D vision › depth estimation
self-supervised depth estimation
0.812024
DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Machine learning › Efficient and distributed learning
distributed training
0.722022
Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration · J. Mach. Learn. Res. 2022
Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks · ICML 2020
Computer vision › 3D vision › 3d human reconstruction
clothed human reconstruction
0.712023
High-Resolution Volumetric Reconstruction for Clothed Humans · ACM Trans. Graph. 2023
Computer vision › 3D vision › motion capture
human performance capture
0.712023
High-Resolution Volumetric Reconstruction for Clothed Humans · ACM Trans. Graph. 2023
Computer vision › 3D vision › 3d reconstruction
volumetric reconstruction
0.712023
High-Resolution Volumetric Reconstruction for Clothed Humans · ACM Trans. Graph. 2023
Mathematical optimization › continuous optimization › convex optimization
operator splitting
0.622018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
GSOS: Gauss-Seidel Operator Splitting Algorithm for Multi-Term Nonsmooth Convex Composite Optimization · ICML 2017
Machine learning › Optimization for machine learning › stochastic optimization › adaptive gradient methods
adam convergence
0.612022
Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration · J. Mach. Learn. Res. 2022
Machine learning › Optimization for machine learning › adaptive optimization
adaptive stochastic optimization
0.612022
Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › one-shot neural architecture search
differentiable architecture search
0.612022
AlphaGAN: Fully Differentiable Architecture Search for Generative Adversarial Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Computer vision › 3D vision › depth estimation
facial depth estimation
0.612022
Depth-Aware Generative Adversarial Network for Talking Head Video Generation · CVPR 2022
Machine learning › Generative modeling
generative adversarial network
0.612022
AlphaGAN: Fully Differentiable Architecture Search for Generative Adversarial Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Generative modeling › generative adversarial network
neural architecture search for GANs
0.612022
AlphaGAN: Fully Differentiable Architecture Search for Generative Adversarial Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Machine learning › Optimization for machine learning
bilevel optimization
0.412020
MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation · CVPR 2020
Machine learning › Optimization for machine learning › optimization › metric optimization › AUC maximization
stochastic AUC maximization
0.412020
Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks · ICML 2020
Distributed systems › distributed machine learning › distributed training
communication-efficient distributed training
0.412020
Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks · ICML 2020
Distributed systems
distributed optimization
0.412020
Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks · ICML 2020
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
discriminative dictionary learning
0.422015
Multi-Level Discriminative Dictionary Learning With Application to Large Scale Image Classification · IEEE Trans. Image Process. 2015
Multi-level Discriminative Dictionary Learning towards Hierarchical Visual Categorization · CVPR 2013
Machine learning › Optimization for machine learning
stochastic optimization
0.412019
A Sufficient Condition for Convergences of Adam and RMSProp · CVPR 2019
Machine learning › Deep learning architectures and training
attention mechanism
0.312018
Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks · NeurIPS 2018
Natural language and speech › Language models and text generation › language modeling › long-context language modeling
context utilization
0.312018
Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks · NeurIPS 2018
Machine learning › Deep learning architectures and training
feature aggregation
0.312018
Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks · NeurIPS 2018

Methods — techniques the papers use, named apart from their topics

convergence analysis · 1.5self-supervised learning · 1.3generative adversarial network · 1.3cross-modal attention · 1.3first-order optimization · 1.0outlier modeling · 0.9label propagation · 0.9batch nuclear norm maximization · 0.9coarse-to-fine strategy · 0.73d convolution · 0.7stochastic gradient descent · 0.4primal-dual optimization · 0.4non-convex concave reformulation · 0.4variable metric · 0.3over-relaxation · 0.3hybrid proximal extragradient · 0.3barzilai-borwein line search · 0.3operator splitting · 0.3
YearPublicationVenuePosition
2025 Image-to-video Adaptation with Outlier Modeling and Robust Self-learning
abstract
The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Existing methods reduce the domain discrepancy via close-set domain adaptation techniques, resulting in inaccurate domain alignment as there exist outlier target frames. To tackle this issue, we extend the vanilla classifier with outlier classes, where each outlier class responsible for capturing outlier frames for a specific class via batch nuclear norm maximization loss. We further propose a new loss by treating the source images apart from class c as instances from outlier class specific for c. As for the modality gap, existing methods usually utilize the pseudo labels obtained from an image-level adapted model to learn a video-level model. Rare efforts are dedicated to handling the noise in pseudo labels. We proposed a new metric based on label propagation consistency to select samples for training a better video-level model. Experiments on 3 benchmarks validating the effectiveness of our method.
Junbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen 0005, Qingming Huang, Huimin Ma 0001
AAAI4
2024 DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation
abstract
Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in constructing accurate 3D facial structures and suppressing complex background noises for generation. However, dense 3D annotations for facial videos is prohibitively costly to obtain. In this paper, first, we present a novel self-supervised method for learning dense 3D facial geometry (i.e., depth) from face videos, without requiring camera parameters and 3D geometry annotations in training. We further propose a strategy to learn pixel-level uncertainties to perceive more reliable rigid-motion pixels for geometry learning. Second, we design an effective geometry-guided facial keypoint estimation module, providing accurate keypoints for generating motion fields. Lastly, we develop a 3D-aware cross-modal (i.e., appearance and depth) attention mechanism, which can be applied to each generation layer, to capture facial geometries in a coarse-to-fine manner. Extensive experiments are conducted on three challenging benchmarks (i.e., VoxCeleb1, VoxCeleb2, and HDTF). The results demonstrate that our proposed framework can generate highly realistic-looking reenacted talking videos, with new state-of-the-art performances established on these benchmarks.
Fa-Ting Hong, Li Shen 0005, Dan Xu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 High-Resolution Volumetric Reconstruction for Clothed Humans
abstract
We present a novel method for reconstructing clothed humans from a sparse set of, e.g., 1–6 RGB images. Despite impressive results from recent works employing deep implicit representation, we revisit the volumetric approach and demonstrate that better performance can be achieved with proper system design. The volumetric representation offers significant advantages in leveraging 3D spatial context through 3D convolutions, and the notorious quantization error is largely negligible with a reasonably large yet affordable volume resolution, e.g., 512. To handle memory and computation costs, we propose a sophisticated coarse-to-fine strategy with voxel culling and subspace sparse convolution. Our method starts with a discretized visual hull to compute a coarse shape and then focuses on a narrow band nearby the coarse shape for refinement. Once the shape is reconstructed, we adopt an image-based rendering approach, which computes the colors of surface points by blending input images with learned weights. Extensive experimental results show that our method significantly reduces the mean point-to-surface (P2S) precision of state-of-the-art methods by more than 50% to achieve approximately 2mm accuracy with a 512 volume resolution. Additionally, images rendered from our textured model achieve a higher peak signal-to-noise ratio (PSNR) compared to state-of-the-art methods.
Sicong Tang, Guangyuan Wang, Qing Ran, Lingzhi Li 0002, Li Shen 0005, Ping Tan 0002
ACM Trans. Graph.5
2022 Depth-Aware Generative Adversarial Network for Talking Head Video Generation
abstract
Talking head video generation aims to produce a synthetic human face video that contains the identity and pose information respectively from a given source image and a driving video. Existing works for this task heavily rely on 2D representations (e.g. appearance and motion) learned from the input images. However, dense 3D facial geometry (e.g. pixel-wise depth) is extremely important for this task as it is particularly beneficial for us to essentially generate accurate 3D face structures and distinguish noisy information from the possibly cluttered background. Nevertheless, dense 3D geometry annotations are prohibitively costly for videos and are typically not available for this video generation task. In this paper, we introduce a self-supervised face-depth learning method to automatically recover dense 3D facial geometry (i.e. depth) from the face videos without the requirement of any expensive 3D annotation data. Based on the learned dense depth maps, we further propose to leverage them to estimate sparse facial keypoints that capture the critical movement of the human head. In a more dense way, the depth is also utilized to learn 3D-aware cross-modal (i.e. appearance and depth) attention to guide the generation of motion fields for warping source image representations. All these contributions compose a novel depth-aware generative adversarial network (DaGAN) for talking head generation. Extensive experiments conducted demonstrate that our proposed method can generate highly realistic faces, and achieve significant results on the unseen human faces.11https://github.com/harlanhong/CVPR2022-DaGAN
Fa-Ting Hong, Longhao Zhang, Li Shen 0005, Dan Xu 0002
CVPR3
2022 Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration
abstract
Adam is one of the most influential adaptive stochastic algorithms for training deep neural networks, which has been pointed out to be divergent even in the simple convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporating a temporal decorrelation technique, seeking an analogous surrogate, etc., have been tried to promote Adam-type algorithms to converge. In contrast with existing approaches, we introduce an alternative easy-to-check sufficient condition, which merely depends on the parameters of the base learning rate and combinations of historical second-order moments, to guarantee the global convergence of generic Adam for solving large-scale non-convex stochastic optimization. This observation, coupled with this sufficient condition, gives much deeper interpretations on the divergence of Adam. On the other hand, in practice, mini-Adam and distributed-Adam are widely used without any theoretical guarantee. We further give an analysis on how the batch size or the number of nodes in the distributed system affects the convergence of Adam, which theoretically shows that mini-batch and distributed Adam can be linearly accelerated by using a larger mini-batch size or a larger number of nodes. At last, we apply the generic Adam and mini-batch Adam with the sufficient condition for solving the counterexample and training several neural networks on various real-world datasets. Experimental results are exactly in accord with our theoretical analysis.
Congliang Chen, Li Shen 0005, Fangyu Zou, Wei Liu 0005
J. Mach. Learn. Res.2
2022 AlphaGAN: Fully Differentiable Architecture Search for Generative Adversarial Networks
abstract
Generative Adversarial Networks (GANs) are formulated as minimax game problems that generative networks attempt to approach real data distributions by adversarial learning against discriminators which learn to distinguish generated samples from real ones, of which the intrinsic problem complexity poses challenges to performance and robustness. In this work, we aim to boost model learning from the perspective of network architectures, by incorporating recent progress on automated architecture search into GANs. Specially we propose a fully differentiable search framework, dubbedalphaGAN, where the searching process is formalized as solving a bi-level minimax optimization problem. The outer-level objective aims for seeking an optimal network architecture towards pure Nash Equilibrium conditioned on the network parameters of generators and discriminators optimized with a traditional adversarial loss within inner level. The entire optimization performs a first-order approach by alternately minimizing the two-level objective in a fully differentiable manner that enables obtaining a suitable architecture efficiently from an enormous search space. Extensive experiments on CIFAR-10 and STL-10 datasets show that our algorithm can obtain high-performing architectures only with 3-GPU hours on a single GPU in the search space comprised of approximate$2\times 10^{11}$possible configurations. We further validate the method on the state-of-the-art network StyleGAN2, and push the score of Fréchet Inception Distance (FID) further, i.e., achieving 1.94 on CelebA, 2.86 on LSUN-church and 2.75 on FFHQ, with relative improvements$3\%{\sim} 26\%$over the baseline architecture. We also provide a comprehensive analysis of the behavior of the searching process and the properties of searched architectures, which would benefit further research on architectures for generative models. Codes and models are available athttps://github.com/yuesongtian/AlphaGAN.
Yuesong Tian, Li Shen 0008, Li Shen 0005, Guinan Su, Zhifeng Li 0001, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Communication Efficient Primal-Dual Algorithm for Nonconvex Nonsmooth Distributed Optimization
abstract
Decentralized optimization problems frequently appear in the large scale machine learning problems. However, few works work on the difficult nonconvex nonsmooth case. In this paper, we propose a decentralized primal-dual algorithm to solve this type of problem in a decentralized manner and the proposed algorithm can achieve an $\mathcal{O}(1/\epsilon^2)$ iteration complexity to attain an $\epsilon-$solution, which is the well-known lower iteration complexity bound for nonconvex optimization. To our knowledge, it is the first algorithm achieving this rate under a nonconvex, nonsmooth decentralized setting. Furthermore, to reduce communication overhead, we also modifying our algorithm by compressing the vectors exchanged between agents. The iteration complexity of the algorithm with compression is still $\mathcal{O}(1/\epsilon^2)$. Besides, we apply the proposed algorithm to solve nonconvex linear regression problem and train deep learning model, both of which demonstrate the efficiency and efficacy of the proposed algorithm.
Congliang Chen, Jiawei Zhang 0007, Li Shen 0005, Peilin Zhao, Zhi-Quan Luo
AISTATS3
2020 MiLeNAS: Efficient Neural Architecture Search via Mixed-Level Reformulation
abstract
Many recently proposed methods for Neural Architecture Search (NAS) can be formulated as bilevel optimization. For efficient implementation, its solution requires approximations of second-order methods. In this paper, we demonstrate that gradient errors caused by such approximations lead to suboptimality, in the sense that the optimization procedure fails to converge to a (locally) optimal solution. To remedy this, this paper proposes MiLeNAS, a mixed-level reformulation for NAS that can be optimized efficiently and reliably. It is shown that even when using a simple first-order method on the mixed-level formulation, MiLeNAS can achieve a lower validation error for NAS problems. Consequently, architectures obtained by our method achieve consistently higher accuracies than those obtained from bilevel optimization. Moreover, MiLeNAS proposes a framework beyond DARTS. It is upgraded via model size-based search and early stopping strategies to complete the search process in around 5 hours. Extensive experiments within the convolutional architecture search space validate the effectiveness of our approach.
Chaoyang He 0001, Haishan Ye, Li Shen 0005, Tong Zhang 0001
CVPR3
2020 Communication-Efficient Distributed Stochastic AUC Maximization with Deep Neural Networks
abstract
In this paper, we study distributed algorithms for large-scale AUC maximization with a deep neural network as a predictive model. Although distributed learning techniques have been investigated extensively in deep learning, they are not directly applicable to stochastic AUC maximization with deep neural networks due to its striking differences from standard loss minimization problems (e.g., cross-entropy). Towards addressing this challenge, we propose and analyze a communication-efficient distributed optimization algorithm based on a \emph{non-convex concave} reformulation of the AUC maximization, in which the communication of both the primal variable and the dual variable between each worker and the parameter server only occurs after multiple steps of gradient-based updates in each worker. Compared with the naive parallel version of an existing algorithm that computes stochastic gradients at individual machines and averages them for updating the model parameter, our algorithm requires a much less number of communication rounds and still achieves linear speedup in theory. To the best of our knowledge, this is the \textbf{first} work that solves the \emph{non-convex concave min-max} problem for AUC maximization with deep neural networks in a communication-efficient distributed manner while still maintaining the linear speedup property in theory. Our experiments on several benchmark datasets show the effectiveness of our algorithm and also confirm our theory.
Zhishuai Guo, Zhuoning Yuan, Li Shen 0005, Wei Liu 0005, Tianbao Yang
ICML4
2020 Squeeze-and-Excitation Networks
abstract
The central building block of convolutional neural networks (CNNs) is the convolution operator, which enables networks to construct informative features by fusing both spatial and channel-wise information within local receptive fields at each layer. A broad range of prior research has investigated the spatial component of this relationship, seeking to strengthen the representational power of a CNN by enhancing the quality of spatial encodings throughout its feature hierarchy. In this work, we focus instead on the channel relationship and propose a novel architectural unit, which we term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We show that these blocks can be stacked together to form SENet architectures that generalise extremely effectively across different datasets. We further demonstrate that SE blocks bring significant improvements in performance for existing state-of-the-art CNNs at slight additional computational cost. Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.251 percent, surpassing the winning entry of 2016 by a relative improvement of ∼ 25 percent. Models and code are available at https://github.com/hujie-frank/SENet.
Jie Hu 0019, Li Shen 0005, Samuel Albanie, Gang Sun 0005, Enhua Wu
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 A Sufficient Condition for Convergences of Adam and RMSProp
abstract
Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporating a temporal decorrelation technique, seeking an analogous surrogate, etc., have been tried to promote Adam/RMSProp-type algorithms to converge. In contrast with existing approaches, we introduce an alternative easy-to-check sufficient condition, which merely depends on the parameters of the base learning rate and combinations of historical second-order moments, to guarantee the global convergence of generic Adam/RMSProp for solving large-scale non-convex stochastic optimization. Moreover, we show that the convergences of several variants of Adam, such as AdamNC, AdaEMA, etc., can be directly implied via the proposed sufficient condition in the non-convex setting. In addition, we illustrate that Adam is essentially a specifically weighted AdaGrad with exponential moving average momentum, which provides a novel perspective for understanding Adam and RMSProp. This observation coupled with this sufficient condition gives much deeper interpretations on their divergences. At last, we validate the sufficient condition by applying Adam and RMSProp to tackle a certain counterexample and train deep neural networks. Numerical results are exactly in accord with our theoretical analysis.
Fangyu Zou, Li Shen 0005, Zequn Jie, Wei Liu 0005
CVPR2
2018 Squeeze-and-Excitation Networks
abstract
Convolutional neural networks are built upon the convolution operation, which extracts informative features by fusing spatial and channel-wise information together within local receptive fields. In order to boost the representational power of a network, several recent approaches have shown the benefit of enhancing spatial encoding. In this work, we focus on the channel relationship and propose a novel architectural unit, which we term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We demonstrate that by stacking these blocks together, we can construct SENet architectures that generalise extremely well across challenging datasets. Crucially, we find that SE blocks produce significant performance improvements for existing state-of-the-art deep architectures at minimal additional computational cost. SENets formed the foundation of our ILSVRC 2017 classification submission which won first place and significantly reduced the top-5 error to 2.251%, achieving a ~25% relative improvement over the winning entry of 2016. Code and models are available at https://github.com/hujie-frank/SENet.
Jie Hu 0019, Li Shen 0005, Gang Sun 0005
CVPR2
2018 Comparator Networks
Weidi Xie, Li Shen 0005, Andrew Zisserman
ECCV (11)2
2018 VGGFace2: A Dataset for Recognising Faces across Pose and Age
abstract
In this paper, we introduce a new large-scale face dataset named VGGFace2. The dataset contains 3.31 million images of 9131 subjects, with an average of 362.6 images for each subject. Images are downloaded from Google Image Search and have large variations in pose, age, illumination, ethnicity and profession (e.g. actors, athletes, politicians). The dataset was collected with three goals in mind: (i) to have both a large number of identities and also a large number of images for each identity; (ii) to cover a large range of pose, age and ethnicity; and (iii) to minimise the label noise. We describe how the dataset was collected, in particular the automated and manual filtering stages to ensure a high accuracy for the images of each identity. To assess face recognition performance using the new dataset, we train ResNet-50 (with and without Squeeze-and-Excitation blocks) Convolutional Neural Networks on VGGFace2, on MS-Celeb-1M, and on their union, and show that training on VGGFace2 leads to improved recognition performance over pose and age. Finally, using the models trained on these datasets, we demonstrate state-of-the-art performance on the IJB-A and IJB-B face recognition benchmarks, exceeding the previous state-of-the-art by a large margin. The dataset and models are publicly available.
Qiong Cao, Li Shen 0005, Weidi Xie, Omkar M. Parkhi, Andrew Zisserman
FG2
2018 An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method
abstract
We propose a novel algorithmic framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-gradient (VMOR-HPE) method with a global convergence guarantee for the maximal monotone operator inclusion problem. Its iteration complexities and local linear convergence rate are provided, which theoretically demonstrate that a large over-relaxed step-size contributes to accelerating the proposed VMOR-HPE as a byproduct. Specifically, we find that a large class of primal and primal-dual operator splitting algorithms are all special cases of VMOR-HPE. Hence, the proposed framework offers a new insight into these operator splitting algorithms. In addition, we apply VMOR-HPE to the Karush-Kuhn-Tucker (KKT) generalized equation of linear equality constrained multi-block composite convex optimization, yielding a new algorithm, namely nonsymmetric Proximal Alternating Direction Method of Multipliers with a preconditioned Extra-gradient step in which the preconditioned metric is generated by a blockwise Barzilai-Borwein line search technique (PADMM-EBB). We also establish iteration complexities of PADMM-EBB in terms of the KKT residual. Finally, we apply PADMM-EBB to handle the nonnegative dual graph regularized low-rank representation problem. Promising results on synthetic and real datasets corroborate the efficacy of PADMM-EBB.
Li Shen 0005, Peng Sun 0011, Wei Liu 0005, Tong Zhang 0001
ICML1
2018 Gather-Excite: Exploiting Feature Context in Convolutional Neural Networks
abstract
While the use of bottom-up local operators in convolutional neural networks (CNNs) matches well some of the statistics of natural images, it may also prevent such models from capturing contextual long-range feature interactions. In this work, we propose a simple, lightweight approach for better context exploitation in CNNs. We do so by introducing a pair of operators: gather, which efficiently aggregates feature responses from a large spatial extent, and excite, which redistributes the pooled information to local features. The operators are cheap, both in terms of number of added parameters and computational complexity, and can be integrated directly in existing architectures to improve their performance. Experiments on several datasets show that gather-excite can bring benefits comparable to increasing the depth of a CNN at a fraction of the cost. For example, we find ResNet-50 with gather-excite operators is able to outperform its 101-layer counterpart on ImageNet with no additional learnable parameters. We also propose a parametric gather-excite operator pair which yields further performance gains, relate it to the recently-introduced Squeeze-and-Excitation Networks, and analyse the effects of these changes to the CNN feature activation statistics.
Jie Hu 0019, Li Shen 0005, Samuel Albanie, Gang Sun 0005, Andrea Vedaldi
NeurIPS2
2017 GSOS: Gauss-Seidel Operator Splitting Algorithm for Multi-Term Nonsmooth Convex Composite Optimization
abstract
In this paper, we propose a fast Gauss-Seidel Operator Splitting (GSOS) algorithm for addressing multi-term nonsmooth convex composite optimization, which has wide applications in machine learning, signal processing and statistics. The proposed GSOS algorithm inherits the advantage of the Gauss-Seidel technique to accelerate the optimization procedure, and leverages the operator splitting technique to reduce the computational complexity. In addition, we develop a new technique to establish the global convergence of the GSOS algorithm. To be specific, we first reformulate the iterations of GSOS as a two-step iterations algorithm by employing the tool of operator optimization theory. Subsequently, we establish the convergence of GSOS based on the two-step iterations algorithm reformulation. At last, we apply the proposed GSOS algorithm to solve overlapping group Lasso and graph-guided fused Lasso problems. Numerical experiments show that our proposed GSOS algorithm is superior to the state-of-the-art algorithms in terms of both efficiency and effectiveness.
Li Shen 0005, Wei Liu 0005, Ganzhao Yuan, Shiqian Ma
ICML1
2016 Co-Occurrence Feature Learning for Skeleton Based Action Recognition Using Regularized Deep LSTM Networks
abstract
Skeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term Memory (LSTM) can learn feature representations and model long-term temporal dependencies automatically, we propose an end-to-end fully connected deep LSTM network for skeleton based action recognition. Inspired by the observation that the co-occurrences of the joints intrinsically characterize human actions, we take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. To train the deep LSTM network effectively, we propose a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons. Experimental results on three human action recognition datasets consistently demonstrate the effectiveness of the proposed model.
Wentao Zhu 0001, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Yanghao Li, Li Shen 0005, Xiaohui Xie
AAAI6
2016 Relay Backpropagation for Effective Learning of Deep Convolutional Neural Networks
Li Shen 0005, Zhouchen Lin, Qingming Huang
ECCV (7)1
2015 Adaptive Sharing for Image Classification
Li Shen 0005, Gang Sun 0005, Zhouchen Lin, Qingming Huang, Enhua Wu
IJCAI1
2015 Multi-Level Discriminative Dictionary Learning With Application to Large Scale Image Classification
abstract
The sparse coding technique has shown flexibility and capability in image representation and analysis. It is a powerful tool in many visual applications. Some recent work has shown that incorporating the properties of task (such as discrimination for classification task) into dictionary learning is effective for improving the accuracy. However, the traditional supervised dictionary learning methods suffer from high computation complexity when dealing with large number of categories, making them less satisfactory in large scale applications. In this paper, we propose a novel multi-level discriminative dictionary learning method and apply it to large scale image classification. Our method takes advantage of hierarchical category correlation to encode multi-level discriminative information. Each internal node of the category hierarchy is associated with a discriminative dictionary and a classification model. The dictionaries at different layers are learnt to capture the information of different scales. Moreover, each node at lower layers also inherits the dictionary of its parent, so that the categories at lower layers can be described with multi-scale information. The learning of dictionaries and associated classification models is jointly conducted by minimizing an overall tree loss. The experimental results on challenging data sets demonstrate that our approach achieves excellent accuracy and competitive computation cost compared with other sparse coding methods for large scale image classification.
Li Shen 0005, Gang Sun 0005, Qingming Huang, Shuhui Wang, Zhouchen Lin, Enhua Wu
IEEE Trans. Image Process.1
2014 Sharing model with multi-level feature representations
abstract
Hierarchical classification models have been proposed to achieve high accuracy by transferring effective information across the categories. One important challenge for this paradigm is to design what can be transferred across the categories. In this paper, we propose a novel method to learn a sharing model by taking advantage of multi-level feature representations. Unlike many of the existing methods which learn the sharing model based on identical feature space, multi-level feature detectors enable our model to capture rich visual information in hierarchical category structure. Moreover, hierarchical classifier parameters associated with multi-level feature representations are learned to model the visual correlation in the hierarchy. The experimental results on Caltech-256 dataset and ImageNet subset demonstrate that our method achieves excellent performance compared with some state-of-the-art methods, and shows the advantage of multi-level information transfer.
Li Shen 0005, Gang Sun 0005, Shuhui Wang, Enhua Wu, Qingming Huang
ICIP1
2013 Multi-level Discriminative Dictionary Learning towards Hierarchical Visual Categorization
abstract
For the task of visual categorization, the learning model is expected to be endowed with discriminative visual feature representation and flexibilities in processing many categories. Many existing approaches are designed based on a flat category structure, or rely on a set of pre-computed visual features, hence may not be appreciated for dealing with large numbers of categories. In this paper, we propose a novel dictionary learning method by taking advantage of hierarchical category correlation. For each internode of the hierarchical category structure, a discriminative dictionary and a set of classification models are learnt for visual categorization, and the dictionaries in different layers are learnt to exploit the discriminative visual properties of different granularity. Moreover, the dictionaries in lower levels also inherit the dictionary of ancestor nodes, so that categories in lower levels are described with multi-scale visual information using our dictionary learning approach. Experiments on Image Net object data subset and SUN397 scene dataset demonstrate that our approach achieves promising performance on data with large numbers of classes compared with some state-of-the-art methods, and is more efficient in processing large numbers of categories.
Li Shen 0005, Shuhui Wang, Gang Sun 0005, Shuqiang Jiang, Qingming Huang
CVPR1