EDBT 2026 Demo / reviewers in the wild / expert
Anbang Yao
dblp:37/7198
· DBLP profile ↗
47ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-3878-8679ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Grid Convolution for 3D Human Pose Estimationabstract3D human pose estimation from 2D keypoint observation has been used in many human-centered computer vision applications. In this work, we tackle the task by formulating a novel grid representation learning paradigm that relies on grid convolution (GridConv), mimicking the wisdom of regular convolution operations in image space. GridConv is defined based on Semantic Grid Transformation (SGT) which leverages a binary assignment matrix to map standard skeleton 2D pose onto a regular weave-like grid pose joint by joint. We provide two ways to implement SGT: handcrafted and learnable SGT. Surprisingly, both designs turn out to achieve promising results and the learnable one is better, demonstrating the great potential of this new lifting representation learning formulation. To improve the ability of GridConv to encode contextual cues, we introduce an attention module over the convolutional kernel, making grid convolution operations input-dependent, spatial-aware and grid-specific. Besides our spatial grid lifting network for single-frame input, we also present a spatial-temporal grid lifting network for video-based input, which relies on an efficient multi-scale grid learning strategy to encode spatial and temporal joint variations. Extensive experiments demonstrate that the proposed grid lifting network outperforms existing approaches by remarkable margins on Human3.6M and MPI-INF-3DHP datasets. Our grid lifting networks also exhibit good generalization ability across three other keypoint-based tasks: 3D hand pose estimation, head pose estimation, and action recognition. Yangyuxuan Kang, Anbang Yao, Shandong Wang, Yurong Chen 0001, Enhua Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Ace-of-Spades: Accelerating Spatially Sparse Convolution for 3D Scene UnderstandingabstractSemantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU/GPU implementations. The existing CNN accelerators designed for weight/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present Ace-of-Spade s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator Ss p nna with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The Ss p nna core together with the 64 KB of L1 memory requires 0.31 mm 2 of area in 10nm process at 1 GHz. Overall, AoS achieves speedup/energy savings of 19.9x / 49.9x and 2.2x / 7.1x over the state-of-the-art CPU and GPU implementations respectively. Om Ji Omer, Prashant Laddha, Gurpreet S. Kalsi, Kamlesh R. Pillai, Anirud Thyagharajan, Ahimanyu Kulkarni, Anbang Yao, Yurong Chen 0001, Sreenivas Subramoney |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2025 | Morse: Dual-Sampling for Lossless Acceleration of Diffusion ModelsabstractIn this paper, we present $Morse$, a simple dual-sampling framework for accelerating diffusion models losslessly. The key insight of Morse is to reformulate the iterative generation (from noise to data) process via taking advantage of fast jump sampling and adaptive residual feedback strategies. Specifically, Morse involves two models called $Dash$ and $Dot$ that interact with each other. The Dash model is just the pre-trained diffusion model of any type, but operates in a jump sampling regime, creating sufficient space for sampling efficiency improvement. The Dot model is significantly faster than the Dash model, which is learnt to generate residual feedback conditioned on the observations at the current jump sampling point on the trajectory of the Dash model, lifting the noise estimate to easily match the next-step estimate of the Dash model without jump sampling. By chaining the outputs of the Dash and Dot models run in a time-interleaved fashion, Morse exhibits the merit of flexibly attaining desired image generation performance while improving overall runtime efficiency. With our proposed weight sharing strategy between the Dash and Dot models, Morse is efficient for training and inference. Our method shows a lossless speedup of 1.78$\times$ to 3.31$\times$ on average over a wide range of sampling step budgets relative to 9 baseline diffusion models on 6 image generation tasks. Furthermore, we show that our method can be also generalized to improve the Latent Consistency Model (LCM-SDXL, which is already accelerated with consistency distillation technique) tailored for few-step text-to-image synthesis. The code and models are available at https://github.com/deep-optimization/Morse. Anbang Yao |
ICML | 3 |
| 2024 | Small Scale Data-Free Knowledge DistillationabstractData-free knowledge distillation is able to utilize the knowledge learned by a large teacher network to augment the training of a smaller student network without accessing the original training data, avoiding privacy, security, and proprietary risks in real applications. In this line of research, existing methods typically follow an inversion-and-distillation paradigm in which a generative adversarial network on-the-fly trained with the guidance of the pre-trained teacher network is used to synthesize a large-scale sample set for knowledge distillation. In this paper, we reexam-ine this common data-free knowledge distillation paradigm, showing that there is considerable room to improve the overall training efficiency through a lens of “small-scale inverted data for knowledge distillation”. In light of three empirical observations indicating the importance of how to balance class distributions in terms of synthetic sample di-versity and difficulty during both data inversion and distillation processes, we propose Small Scale Data-free Knowledge Distillation (SSD-KD). Informulation, SSD-KD introduces a modulating function to balance synthetic samples and a priority sampling function to select proper samples, facilitated by a dynamic replay buffer and a reinforcement learning strategy. As a result, SSD-KD can perform dis-tillation training conditioned on an extremely small scale of synthetic samples (e.g., 10x less than the original training data scale), making the overall training efficiency one or two orders of magnitude faster than many mainstream methods while retaining superior or competitive model performance, as demonstrated on popular image classification and semantic segmentation benchmarks. The code is available at https://github.com/OSVAI/SSD-KD. Yikai Wang 0001, Huaping Liu 0001, Fuchun Sun 0001, Anbang Yao |
CVPR | 5 |
| 2024 | KernelWarehouse: Rethinking the Design of Dynamic ConvolutionabstractDynamic convolution learns a linear mixture of $n$ static kernels weighted with their input-dependent attentions, demonstrating superior performance than normal convolution. However, it increases the number of convolutional parameters by $n$ times, and thus is not parameter efficient. This leads to no research progress that can allow researchers to explore the setting $n > 100$ (an order of magnitude larger than the typical setting $n < 10$) for pushing forward the performance boundary of dynamic convolution while enjoying parameter efficiency. To fill this gap, in this paper, we propose KernelWarehouse, a more general form of dynamic convolution, which redefines the basic concepts of “kernels”, “assembling kernels” and “attention function” through the lens of exploiting convolutional parameter dependencies within the same layer and across neighboring layers of a ConvNet. We testify the effectiveness of KernelWarehouse on ImageNet and MS-COCO datasets using various ConvNet architectures. Intriguingly, KernelWarehouse is also applicable to Vision Transformers, and it can even reduce the model size of a backbone while improving the model accuracy. For instance, KernelWarehouse ($n = 4$) achieves 5.61%|3.90%|4.38% absolute top-1 accuracy gain on the ResNet18|MobileNetV2|DeiT-Tiny backbone, and KernelWarehouse ($n = 1/4$) with 65.10% model size reduction still achieves 2.29% gain on the ResNet18 backbone. The code and models are available at https://github.com/OSVAI/KernelWarehouse. Anbang Yao |
ICML | 2 |
| 2024 | ScaleKD: Strong Vision Transformers Could Be Excellent TeachersabstractIn this paper, we question if well pre-trained vision transformer (ViT) models could be used as teachers that exhibit scalable properties to advance cross architecture knowledge distillation research, in the context of adopting mainstream large-scale visual recognition datasets for evaluation. To make this possible, our analysis underlines the importance of seeking effective strategies to align (1) feature computing paradigm differences, (2) model scale differences, and (3) knowledge density differences. By combining three closely coupled components namely *cross attention projector*, *dual-view feature mimicking* and *teacher parameter perception* tailored to address the alignment problems stated above, we present a simple and effective knowledge distillation method, called *ScaleKD*. Our method can train student backbones that span across a variety of convolutional neural network (CNN), multi-layer perceptron (MLP), and ViT architectures on image classification datasets, achieving state-of-the-art knowledge distillation performance. For instance, taking a well pre-trained Swin-L as the teacher model, our method gets 75.15\%|82.03\%|84.16\%|78.63\%|81.96\%|83.93\%|83.80\%|85.53\% top-1 accuracies for MobileNet-V1|ResNet-50|ConvNeXt-T|Mixer-S/16|Mixer-B/16|ViT-S/16|Swin-T|ViT-B/16 models trained on ImageNet-1K dataset from scratch, showing 3.05\%|3.39\%|2.02\%|4.61\%|5.52\%|4.03\%|2.62\%|3.73\% absolute gains to the individually trained counterparts. Intriguingly, when scaling up the size of teacher models or their pre-training datasets, our method showcases the desired scalable properties, bringing increasingly larger gains to student models. We also empirically show that the student backbones trained by our method transfer well on downstream MS-COCO and ADE20K datasets. More importantly, our method could be used as a more efficient alternative to the time-intensive pre-training paradigm for any target student model on large-scale datasets if a strong pre-trained ViT is available, reducing the amount of viewed training samples up to 195$\times$. The code is available at *https://github.com/deep-optimization/ScaleKD*. Anbang Yao |
NeurIPS | 4 |
| 2023 | 3D Human Pose Lifting with Grid ConvolutionabstractExisting lifting networks for regressing 3D human poses from 2D single-view poses are typically constructed with linear layers based on graph-structured representation learning. In sharp contrast to them, this paper presents Grid Convolution (GridConv), mimicking the wisdom of regular convolution operations in image space. GridConv is based on a novel Semantic Grid Transformation (SGT) which leverages a binary assignment matrix to map the irregular graph-structured human pose onto a regular weave-like grid pose representation joint by joint, enabling layer-wise feature learning with GridConv operations. We provide two ways to implement SGT, including handcrafted and learnable designs. Surprisingly, both designs turn out to achieve promising results and the learnable one is better, demonstrating the great potential of this new lifting representation learning formulation. To improve the ability of GridConv to encode contextual cues, we introduce an attention module over the convolutional kernel, making grid convolution operations input-dependent, spatial-aware and grid-specific. We show that our fully convolutional grid lifting network outperforms state-of-the-art methods with noticeable margins under (1) conventional evaluation on Human3.6M and (2) cross-evaluation on MPI-INF-3DHP. Code is available at https://github.com/OSVAI/GridConv. Yangyuxuan Kang, Anbang Yao, Shandong Wang, Enhua Wu |
AAAI | 3 |
| 2023 | Compacting Binary Neural Networks by Sparse Kernel SelectionabstractBinary Neural Network (BNN) represents convolution weights with 1-bit values, which enhances the efficiency of storage and computation. This paper is motivated by a previously revealed phenomenon that the binary kernels in successful BNNs are nearly power-law distributed: their values are mostly clustered into a small number of codewords. This phenomenon encourages us to compact typical BNNs and obtain further close performance through learning non-repetitive kernels within a binary kernel subspace. Specifically, we regard the binarization process as kernel grouping in terms of a binary codebook, and our task lies in learning to select a smaller subset of codewords from the full code-book. We then leverage the Gumbel-Sinkhorn technique to approximate the codeword selection process, and develop the Permutation Straight-Through Estimator (PSTE) that is able to not only optimize the selection process end-to-end but also maintain the non-repetitive occupancy of selected codewords. Experiments verify that our method reduces both the model size and bit-wise computational costs, and achieves accuracy improvements compared with state-of-the-art BNNs under comparable budgets. Yikai Wang 0001, Wenbing Huang 0001, Yinpeng Dong, Fuchun Sun 0001, Anbang Yao |
CVPR | 5 |
| 2023 | NORM: Knowledge Distillation via N-to-One Representation Matching
Lujun Li 0001, Chao Li 0009, Anbang Yao |
ICLR | 4 |
| 2023 | Ske2Grid: Skeleton-to-Grid Representation Learning for Action RecognitionabstractThis paper presents Ske2Grid, a new representation learning framework for improved skeleton-based action recognition. In Ske2Grid, we define a regular convolution operation upon a novel grid representation of human skeleton, which is a compact image-like grid patch constructed and learned through three novel designs. Specifically, we propose a graph-node index transform (GIT) to construct a regular grid patch through assigning the nodes in the skeleton graph one by one to the desired grid cells. To ensure that GIT is a bijection and enrich the expressiveness of the grid representation, an up-sampling transform (UPT) is learned to interpolate the skeleton graph nodes for filling the grid patch to the full. To resolve the problem when the one-step UPT is aggressive and further exploit the representation capability of the grid patch with increasing spatial size, a progressive learning strategy (PLS) is proposed which decouples the UPT into multiple steps and aligns them to multiple paired GITs through a compact cascaded design learned progressively. We construct networks upon prevailing graph convolution networks and conduct experiments on six mainstream skeleton-based action recognition datasets. Experiments show that our Ske2Grid significantly outperforms existing GCN-based solutions under different benchmark settings, without bells and whistles. Code and models are available at https://github.com/OSVAI/Ske2Grid. Yangyuxuan Kang, Anbang Yao, Yurong Chen 0001 |
ICML | 3 |
| 2023 | Augmentation-free Dense Contrastive Distillation for Efficient Semantic Segmentation
Meina Song, Anbang Yao |
NeurIPS | 5 |
| 2022 | Efficient Meta-Tuning for Content-Aware Neural Video Delivery
Xiaoqi Li 0009, Jiaming Liu 0003, Shizun Wang, Ming Lu 0002, Yurong Chen 0001, Anbang Yao, Yandong Guo, Shanghang Zhang |
ECCV (18) | 7 |
| 2022 | Omni-Dimensional Dynamic Convolution
Aojun Zhou, Anbang Yao |
ICLR | 3 |
| 2022 | OANet: Learning Two-View Correspondences and Geometry Using Order-Aware NetworkabstractEstablishing correct correspondences between two images should consider both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential or fundamental matrix. Specifically, this proposed network is built hierarchically and comprises three operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in canonical order and invariant to input permutations. Next, the clusters are spatially correlated to encode the global context of correspondences. After that, the context-encoded clusters are interpolated back to the original size and position to build a hierarchical architecture. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts. Besides, based on the proposed method and advanced local feature, we won the first place in CVPR 2019 image matching workshop challenge and also achieve state-of-the-art results in the Visual Localization benchmark. Code is available at https://github.com/zjhthu/OANet. Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Long Quan, Hongen Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Sub-bit Neural Networks: Learning to Compress and Accelerate Binary Neural NetworksabstractIn the low-bit quantization field, training Binarized Neural Networks (BNNs) is the extreme solution to ease the deployment of deep models on resource-constrained devices, having the lowest storage cost and significantly cheaper bit-wise operations compared to 32-bit floating-point counterparts. In this paper, we introduce Sub-bit Neural Networks (SNNs), a new type of binary quantization design tailored to compress and accelerate BNNs. SNNs are inspired by an empirical observation, showing that binary kernels learnt at convolutional layers of a BNN model are likely to be distributed over kernel subsets. As a result, unlike existing methods that binarize weights one by one, SNNs are trained with a kernel-aware optimization framework, which exploits binary quantization in the fine-grained convolutional kernel space. Specifically, our method includes a random sampling step generating layer-specific subsets of the kernel space, and a refinement step learning to adjust these subsets of binary kernels via optimization. Experiments on visual recognition benchmarks and the hardware deployment on FPGA validate the great potentials of SNNs. For instance, on ImageNet, SNNs of ResNet-18/ResNet-34 with 0.56-bit weights achieve 3.13/3.33× runtime speedup and 1.8× compression over conventional BNNs with moderate drops in recognition accuracy. Promising results are also obtained when applying SNNs to binarize both weights and activations. Our code is available at https://github.com/yikaiw/SNN. Yikai Wang 0001, Yi Yang 0001, Fuchun Sun 0001, Anbang Yao |
ICCV | 4 |
| 2021 | Dynamic Normalization and Relay for Video Action RecognitionabstractConvolutional Neural Networks (CNNs) have been the dominant model for video action recognition. Due to the huge memory and compute demand, popular action recognition networks need to be trained with small batch sizes, which makes learning discriminative spatial-temporal representations for videos become a challenging problem. In this paper, we present Dynamic Normalization and Relay (DNR), an improved normalization design, to augment the spatial-temporal representation learning of any deep action recognition model, adapting to small batch size training settings. We observe that state-of-the-art action recognition networks usually apply the same normalization parameters to all video data, and ignore the dependencies of the estimated normalization parameters between neighboring frames (at the same layer) and between neighboring layers (with all frames of a video clip). Inspired by this, DNR introduces two dynamic normalization relay modules to explore the potentials of cross-temporal and cross-layer feature distribution dependencies for estimating accurate layer-wise normalization parameters. These two DNR modules are instantiated as a light-weight recurrent structure conditioned on the current input features, and the normalization parameters estimated from the neighboring frames based features at the same layer or from the whole video clip based features at the preceding layers. We first plug DNR into prevailing 2D CNN backbones and test its performance on public action recognition datasets including Kinetics and Something-Something. Experimental results show that DNR brings large performance improvements to the baselines, achieving over 4.4% absolute margins in top-1 accuracy without training bells and whistles. More experiments on 3D backbones and several latest 2D spatial-temporal networks further validate its effectiveness. Code will be available at https://github.com/caidonkey/dnr. Anbang Yao, Yurong Chen 0001 |
NeurIPS | 2 |
| 2020 | Explicit Residual Descent for 3D Human Pose Estimation from 2D Joint Locations
Yangyuxuan Kang, Anbang Yao, Shandong Wang, Ming Lu 0002, Yurong Chen 0001, Enhua Wu |
BMVC | 2 |
| 2020 | Learning to Learn Parameterized Classification Networks for Scalable Input Images
Anbang Yao, Qifeng Chen 0001 |
ECCV (29) | 2 |
| 2020 | PSConv: Squeezing Feature Pyramid into One Compact Poly-Scale Convolutional Layer
Anbang Yao, Qifeng Chen 0001 |
ECCV (21) | 2 |
| 2020 | Resolution Switchable Networks for Runtime Efficient Image Recognition
Yikai Wang 0001, Fuchun Sun 0001, Anbang Yao |
ECCV (15) | 4 |
| 2020 | Knowledge Transfer via Dense Cross-Layer Mutual-Distillation
Anbang Yao, Dawei Sun 0007 |
ECCV (15) | 1 |
| 2020 | CASNet: Common Attribute Support Network for image instance and panoptic segmentationabstractInstance segmentation and panoptic segmentation is being paid more and more attention in recent years. In comparison with bounding box based object detection and semantic segmentation, instance segmentation can provide more analytical results at pixel level. Given the insight that pixels belonging to one instance have one or more common attributes of current instance, we bring up an one-stage instance segmentation network named Common Attribute Support Network (CASNet), which realizes instance segmentation by predicting and clustering common attributes. CASNet is designed in the manner of fully convolutional and can implement training and inference from end to end. And CASNet manages predicting the instance without overlaps and holes, which problem exists in most of current instance segmentation algorithms. Furthermore, it can be easily extended to panoptic segmentation through minor modifications with little computation overhead. CASNet builds a bridge between semantic and instance segmentation from finding pixel class ID to obtaining class and instance ID by operations on common attribute. Through experiment for instance and panoptic segmentation, CASNet gets mAP 32.8% and PQ 59.0% on Cityscapes validation dataset by joint training, and mAP 36.3% and PQ 66.1% by separated training mode. For panoptic segmentation, CASNet gets state-of-the-art performance on the Cityscapes validation dataset. Yuqing Hou, Anbang Yao, Yurong Chen 0001, Keqiang Li 0002 |
ICPR | 3 |
| 2020 | Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer FusionabstractWe propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, unlike existing multimodal methods that necessitate individual encoders for different modalities, we verify that multimodal features can be learnt within a shared single network by merely maintaining modality-specific batch normalization layers in the encoder, which also enables implicit fusion via joint feature representation learning. Secondly, we propose a bidirectional multi-layer fusion scheme, where multimodal features can be exploited progressively. To take advantage of such scheme, we introduce two asymmetric fusion operations including channel shuffle and pixel shift, which learn different fused features with respect to different fusion directions. These two operations are parameter-free and strengthen the multimodal feature interactions across channels as well as enhance the spatial feature discrimination within channels. We conduct extensive experiments on semantic segmentation and image translation tasks, based on three publicly available datasets covering diverse modalities. Results indicate that our proposed framework is general, compact and is superior to state-of-the-art fusion frameworks. Yikai Wang 0001, Fuchun Sun 0001, Ming Lu 0002, Anbang Yao |
ACM Multimedia | 4 |
| 2020 | Pointly-supervised scene parsing with uncertainty mixture
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023 |
Comput. Vis. Image Underst. | 3 |
| 2020 | Learning to Draw Sight Lines
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023 |
Int. J. Comput. Vis. | 3 |
| 2019 | Deeply-Supervised Knowledge SynergyabstractConvolutional Neural Networks (CNNs) have become deeper and more complicated compared with the pioneering AlexNet. However, current prevailing training scheme follows the previous way of adding supervision to the last layer of the network only and propagating error information up layer-by-layer. In this paper, we propose Deeply-supervised Knowledge Synergy (DKS), a new method aiming to train CNNs with improved generalization ability for image classification tasks without introducing extra computational cost during inference. Inspired by the deeply-supervised learning scheme, we first append auxiliary supervision branches on top of certain intermediate network layers. While properly using auxiliary supervision can improve model accuracy to some degree, we go one step further to explore the possibility of utilizing the probabilistic knowledge dynamically learnt by the classifiers connected to the backbone network as a new regularization to improve the training. A novel synergy loss, which considers pairwise knowledge matching among all supervision branches, is presented. Intriguingly, it enables dense pairwise knowledge matching operations in both top-down and bottom-up directions at each training iteration, resembling a dynamic synergy process for the same task. We evaluate DKS on image classification datasets using state-of-the-art CNN architectures, and show that the models trained with it are consistently better than the corresponding counterparts. For instance, on the ImageNet classification benchmark, our ResNet-152 model outperforms the baseline model with a 1.47% margin in Top-1 accuracy. Code is available at https://github.com/sundw2014/DKS. Dawei Sun 0007, Anbang Yao, Aojun Zhou, Hao Zhao 0002 |
CVPR | 2 |
| 2019 | HBONet: Harmonious Bottleneck on Two Orthogonal DimensionsabstractMobileNets, a class of top-performing convolutional neural network architectures in terms of accuracy and efficiency trade-off, are increasingly used in many resource-aware vision applications. In this paper, we present Harmonious Bottleneck on two Orthogonal dimensions (HBO), a novel architecture unit, specially tailored to boost the accuracy of extremely lightweight MobileNets at the level of less than 40 MFLOPs. Unlike existing bottleneck designs that mainly focus on exploring the interdependencies among the channels of either groupwise or depthwise convolutional features, our HBO improves bottleneck representation while maintaining similar complexity via jointly encoding the feature interdependencies across both spatial and channel dimensions. It has two reciprocal components, namely spatial contraction-expansion and channel expansion-contraction, nested in a bilaterally symmetric structure. The combination of two interdependent transformations performing on orthogonal dimensions of feature maps enhances the representation and generalization ability of our proposed module, guaranteeing compelling performance with limited computational resource and power. By replacing the original bottlenecks in MobileNetV2 backbone with HBO modules, we construct HBONets which are evaluated on ImageNet classification, PASCAL VOC object detection and Market-1501 person re-identification. Extensive experiments show that with the severe constraint of computational budget our models outperform MobileNetV2 counterparts by remarkable margins of at most 6.6%, 6.3% and 5.0% on the above benchmarks respectively. Code and pretrained models are available at https://github.com/d-li14/HBONet. Aojun Zhou, Anbang Yao |
ICCV | 3 |
| 2019 | A Closed-Form Solution to Universal Style TransferabstractUniversal style transfer tries to explicitly minimize the losses in feature space, thus it does not require training on any pre-defined styles. It usually uses different layers of VGG network as the encoders and trains several decoders to invert the features into images. Therefore, the effect of style transfer is achieved by feature transform. Although plenty of methods have been proposed, a theoretical analysis of feature transform is still missing. In this paper, we first propose a novel interpretation by treating it as the optimal transport problem. Then, we demonstrate the relations of our formulation with former works like Adaptive Instance Normalization (AdaIN) and Whitening and Coloring Transform (WCT). Finally, we derive a closed-form solution named Optimal Style Transfer (OST) under our formulation by additionally considering the content loss of Gatys. Comparatively, our solution can preserve better structure and achieve visually pleasing results. It is simple yet effective and we demonstrate its advantages both quantitatively and qualitatively. Besides, we hope our theoretical analysis can inspire future works in neural style transfer. Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Feng Xu 0005, Li Zhang 0023 |
ICCV | 3 |
| 2019 | Learning Two-View Correspondences and Geometry Using Order-Aware NetworkabstractEstablishing correspondences between two images requires both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential matrix. Specifically, this proposed network is built hierarchically and comprises three novel operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in a canonical order and invariant to input permutations. Next, the clusters are spatially correlated to form the global context of correspondences. After that, the context-encoded clusters are recovered back to the original size through a proposed upsampling operator. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts. Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Hongen Liao, Long Quan |
ICCV | 4 |
| 2018 | Explicit Loss-Error-Aware Quantization for Low-Bit Deep Neural NetworksabstractBenefiting from tens of millions of hierarchically stacked learnable parameters, Deep Neural Networks (DNNs) have demonstrated overwhelming accuracy on a variety of artificial intelligence tasks. However reversely, the large size of DNN models lays a heavy burden on storage, computation and power consumption, which prohibits their deployments on the embedded and mobile systems. In this paper, we propose Explicit Loss-error-aware Quantization (ELQ), a new method that can train DNN models with very low-bit parameter values such as ternary and binary ones to approximate 32-bit floating-point counterparts without noticeable loss of predication accuracy. Unlike existing methods that usually pose the problem as a straightforward approximation of the layer-wise weights or outputs of the original full-precision model (specifically, minimizing the error of the layer-wise weights or inner products of the weights and the inputs between the original and respective quantized models), our ELQ elaborately bridges the loss perturbation from the weight quantization and an incremental quantization strategy to address DNN quantization. Through explicitly regularizing the loss perturbation and the weight approximation error in an incremental way, we show that such a new optimization method is theoretically reasonable and practically effective. As validated with two mainstream convolutional neural network families (i.e., fully convolutional and non-fully convolutional), our ELQ shows better results than state-of-the-art quantization methods on the large scale ImageNet classification dataset. Code will be made publicly available. Aojun Zhou, Anbang Yao, Yurong Chen 0001 |
CVPR | 2 |
| 2018 | Efficient Semantic Scene Completion Network with Spatial Group Convolution
Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023, Hongen Liao |
ECCV (12) | 3 |
| 2017 | Network Sketching: Exploiting Binary Structure in Deep CNNs
Yiwen Guo, Anbang Yao, Hao Zhao 0002, Yurong Chen 0001 |
CVPR | 2 |
| 2017 | RON: Reverse Connection with Objectness Prior Networks for Object DetectionabstractWe present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problems: (a) multi-scale object localization and (b) negative sample mining. To address (a), we design the reverse connection, which enables the network to detect objects on multi-levels of CNNs. To deal with (b), we propose the objectness prior to significantly reduce the searching space of objects. We optimize the reverse connection, objectness prior and object detector jointly by a multi-task loss function, thus RON can directly predict final detection results from all locations of various feature maps. Extensive experiments on the challenging PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO benchmarks demonstrate the competitive performance of RON. Specifically, with VGG-16 and low resolution 384×384 input size, the network gets 81.3% mAP on PASCAL VOC 2007, 80.7% mAP on PASCAL VOC 2012 datasets. Its superiority increases when datasets become larger and more difficult, as demonstrated by the results on the MS COCO dataset. With 1.5G GPU memory at test phase, the speed of the network is 15 FPS, 3 times faster than the Faster R-CNN counterpart. Code will be made publicly available. Tao Kong, Fuchun Sun 0001, Anbang Yao, Huaping Liu 0001, Ming Lu 0002, Yurong Chen 0001 |
CVPR | 3 |
| 2017 | Physics Inspired Optimization on Semantic Transfer Features: An Alternative Method for Room Layout EstimationabstractIn this paper, we propose an alternative method to estimate room layouts of cluttered indoor scenes. This method enjoys the benefits of two novel techniques. The first one is semantic transfer (ST), which is: (1) a formulation to integrate the relationship between scene clutter and room layout into convolutional neural networks, (2) an architecture that can be end-to-end trained, (3) a practical strategy to initialize weights for very deep networks under unbalanced training data distribution. ST allows us to extract highly robust features under various circumstances, and in order to address the computation redundance hidden in these features we develop a principled and efficient inference scheme named physics inspired optimization (PIO). PIOs basic idea is to formulate some phenomena observed in ST features into mechanics concepts. Evaluations on public datasets LSUN and Hedau show that the proposed method is more accurate than state-of-the-art methods. Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023 |
CVPR | 3 |
| 2017 | Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style TransferabstractRecently, the community of style transfer is trying to incorporate semantic information into traditional system. This practice achieves better perceptual results by transferring the style between semantically-corresponding regions. Yet, few efforts are invested to address the computation bottleneck of back-propagation. In this paper, we propose a new framework for fast semantic style transfer. Our method decomposes the semantic style transfer problem into feature reconstruction part and feature decoder part. The reconstruction part tactfully solves the optimization problem of content loss and style loss in feature space by particularly reconstructed feature. This significantly reduces the computation of propagating the loss through the whole network. The decoder part transforms the reconstructed feature into the stylized image. Through a careful bridging of the two modules, the proposed approach not only achieves competitive results as backward optimization methods but also is about two orders of magnitude faster. Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Feng Xu 0005, Yurong Chen 0001, Li Zhang 0023 |
ICCV | 3 |
| 2017 | Incremental Network Quantization: Towards Lossless CNNs with Low-precision Weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Yurong Chen 0001 |
ICLR (Poster) | 2 |
| 2017 | Learning supervised scoring ensemble for emotion recognition in the wildabstractState-of-the-art approaches for the previous emotion recognition in the wild challenges are usually built on prevailing Convolutional Neural Networks (CNNs). Although there is clear evidence that CNNs with increased depth or width can usually bring improved predication accuracy, existing top approaches provide supervision only at the output feature layer, resulting in the insufficient training of deep CNN models. In this paper, we present a new learning method named Supervised Scoring Ensemble (SSE) for advancing this challenge with deep CNNs. We first extend the idea of recent deep supervision to deal with emotion recognition problem. Benefiting from adding supervision not only to deep layers but also to intermediate layers and shallow layers, the training of deep CNNs can be well eased. Second, we present a new fusion structure in which class-wise scoring activations at diverse complementary feature layers are concatenated and further used as the inputs for second-level supervision, acting as a deep feature ensemble within a single CNN architecture. We show our proposed learning method brings large accuracy gains over diverse backbone networks consistently. On this year's audio-video based emotion recognition task, the average recognition rate of our best submission is 60.34%, forming a new envelop over all existing records. Shandong Wang, Anbang Yao, Yurong Chen 0001 |
ICMI | 4 |
| 2016 | HyperNet: Towards Accurate Region Proposal Generation and Joint Object DetectionabstractAlmost all of the current top-performing object detection networks employ region proposals to guide the search for object instances. State-of-the-art region proposal methods usually need several thousand proposals to get high recall, thus hurting the detection efficiency. Although the latest Region Proposal Network method gets promising detection accuracy with several hundred proposals, it still struggles in small-size object detection and precise localization (e.g., large IoU thresholds), mainly due to the coarseness of its feature maps. In this paper, we present a deep hierarchical network, namely HyperNet, for handling region proposal generation and object detection jointly. Our HyperNet is primarily based on an elaborately designed Hyper Feature which aggregates hierarchical feature maps first and then compresses them into a uniform space. The Hyper Features well incorporate deep but highly semantic, intermediate but really complementary, and shallow but naturally high-resolution features of the image, thus enabling us to construct HyperNet by sharing them both in generating proposals and detecting objects via an end-to-end joint training strategy. For the deep VGG16 model, our method achieves completely leading recall and state-of-the-art object detection accuracy on PASCAL VOC 2007 and 2012 using only 100 proposals per image. It runs with a speed of 5 fps (including all steps) on a GPU, thus having the potential for real-time processing. Tao Kong, Anbang Yao, Yurong Chen 0001, Fuchun Sun 0001 |
CVPR | 2 |
| 2016 | HoloNet: towards robust emotion recognition in the wildabstractIn this paper, we present HoloNet, a well-designed Convolutional Neural Network (CNN) architecture regarding our submissions to the video based sub-challenge of the Emotion Recognition in the Wild (EmotiW) 2016 challenge. In contrast to previous related methods that usually adopt relatively simple and shallow neural network architectures to address emotion recognition task, our HoloNet has three critical considerations in network design. (1) To reduce redundant filters and enhance the non-saturated non-linearity in the lower convolutional layers, we use a modified Concatenated Rectified Linear Unit (CReLU) instead of ReLU. (2) To enjoy the accuracy gain from considerably increased network depth and maintain efficiency, we combine residual structure and CReLU to construct the middle layers. (3) To broaden network width and introduce multi-scale feature extraction property, the topper layers are designed as a variant of inception-residual structure. The main benefit of grouping these modules into the HoloNet is that both negative and positive phase information implicitly contained in the input data can flow over it in multiple paths, thus deep multi-scale features explicitly capturing emotion variation can be well extracted from multi-path sibling layers, and then can be further concatenated for robust recognition. We obtain competitive results in this year’s video based emotion recognition sub-challenge using an ensemble of two HoloNet models trained with given data only. Specifically, we obtain a mean recognition rate of 57.84%, outperforming the baseline accuracy with an absolute margin of 17.37%, and yielding 4.04% absolute accuracy gain compared to the result of last year’s winner team. Meanwhile, our method runs with a speed of several thousands of frames per second on a GPU, thus it is well applicable to real-time scenarios. Anbang Yao, Shandong Wang, Liang Sha, Yurong Chen 0001 |
ICMI | 1 |
| 2016 | Dynamic Network Surgery for Efficient DNNsabstractDeep learning has become a ubiquitous technology to improve machine intelligence. However, most of the existing deep models are structurally very complex, making them difficult to be deployed on the mobile platforms with limited computational power. In this paper, we propose a novel network compression method called dynamic network surgery, which can remarkably reduce the network complexity by making on-the-fly connection pruning. Unlike the previous methods which accomplish this task in a greedy way, we properly incorporate connection splicing into the whole process to avoid incorrect pruning and make it as a continual network maintenance. The effectiveness of our method is proved with experiments. Without any accuracy loss, our method can efficiently compress the number of parameters in LeNet-5 and AlexNet by a factor of $\bm{108}\times$ and $\bm{17.7}\times$ respectively, proving that it outperforms the recent pruning method by considerable margins. Code and some models are available at https://github.com/yiwenguo/Dynamic-Network-Surgery. Yiwen Guo, Anbang Yao, Yurong Chen 0001 |
NIPS | 2 |
| 2015 | Capturing AU-Aware Facial Features and Their Latent Relations for Emotion Recognition in the WildabstractThe Emotion Recognition in the Wild (EmotiW) Challenge has been held for three years. Previous winner teams primarily focus on designing specific deep neural networks or fusing diverse hand-crafted and deep convolutional features. They all neglect to explore the significance of the latent relations among changing features resulted from facial muscle motions. In this paper, we study this recognition challenge from the perspective of analyzing the relations among expression-specific facial features in an explicit manner. Our method has three key components. First, we propose a pair-wise learning strategy to automatically seek a set of facial image patches which are important for discriminating two particular emotion categories. We found these learnt local patches are in part consistent with the locations of expression-specific Action Units (AUs), thus the features extracted from such kind of facial patches are named AU-aware facial features. Second, in each pair-wise task, we use an undirected graph structure, which takes learnt facial patches as individual vertices, to encode feature relations between any two learnt facial patches. Finally, a robust emotion representation is constructed by concatenating all task-specific graph-structured facial feature relations sequentially. Extensive experiments on the EmotiW 2015 Challenge testify the efficacy of the proposed approach. Without using additional data, our final submissions achieved competitive results on both sub-challenges including the image based static facial expression recognition (we got 55.38% recognition accuracy outperforming the baseline 39.13% with a margin of 16.25%) and the audio-video based emotion recognition (we got 53.80% recognition accuracy outperforming the baseline 39.33% and the 2014 winner team's final result 50.37% with the margins of 14.47% and 3.43%, respectively). Anbang Yao, Junchao Shao, Ningning Ma, Yurong Chen 0001 |
ICMI | 1 |
| 2013 | Robust Face Representation Using Hybrid Spatial Feature Interdependence MatrixabstractA key issue in face recognition is to seek an effective descriptor for representing face appearance. In the context of considering the face image as a set of small facial regions, this paper presents a new face representation approach coined spatial feature interdependence matrix (SFIM). Unlike classical face descriptors which usually use a hierarchically organized or a sequentially concatenated structure to describe the spatial layout features extracted from local regions, SFIM is attributed to the exploitation of the underlying feature interdependences regarding local region pairs inside a class specific face. According to SFIM, the face image is projected onto an undirected connected graph in a manner that explicitly encodes feature interdependence-based relationships between local regions. We calculate the pair-wise interdependence strength as the weighted discrepancy between two feature sets extracted in a hybrid feature space fusing histograms of intensity, local binary pattern and oriented gradients. To achieve the goal of face recognition, our SFIM-based face descriptor is embedded in three different recognition frameworks, namely nearest neighbor search, subspace-based classification, and linear optimization-based classification. Extensive experimental results on four well-known face databases and comprehensive comparisons with the state-of-the-art results are provided to demonstrate the efficacy of the proposed SFIM-based descriptor. Anbang Yao |
IEEE Trans. Image Process. | 1 |
| 2012 | A compact association of particle filtering and kernel based object tracking
Anbang Yao, Xinggang Lin, Guijin Wang |
Pattern Recognit. | 1 |
| 2011 | Spatial Feature Interdependence Matrix (SFIM): A Robust Descriptor for Face Recognition
Anbang Yao |
PSIVT (1) | 1 |
| 2010 | An incremental Bhattacharyya dissimilarity measure for particle filtering
Anbang Yao, Guijin Wang, Xinggang Lin, Xiujuan Chai |
Pattern Recognit. | 1 |
| 2009 | Hand posture recognition in video using multiple cuesabstractHand posture conveys profound information for computer vision applications, but the articulated hand structure and restraint capture condition cast a tough obstacle on practical implementation, especially in real time video. This paper presents a framework to recognize hand postures in consecutive video frames. Mixture of Gaussian skin/non skin models is constructed for hand region detection, followed by particle filter to track hand. Then a soft-decision scheme based on extended Histogram of Orientated gradient is proposed to refine the best posture region and recognize it from pre-defined posture set. Experimental result shows promising performance under various capture conditions. Liang Sha, Guijin Wang, Anbang Yao, Xinggang Lin, Xiujuan Chai |
ICME | 3 |
| 2008 | Kernel based articulated object tracking with scale adaptation and model updateabstractKernel based object tracking (KBOT) is one of the most popular and effective techniques for tracking task. However the constancy of the target model and unsound scale adaptation method are two main limitations. In this paper, we present a kernel based approach incorporated with scale estimation and target model update for articulated object tracking task. After predicating the object center with scale fixed KBOT, we extend scale selection theory to estimate the local optimal object scale. Once the object scale has been estimated, a kernel density estimation based strategy is developed to update the target model. Experimental results show that our approach is superior to traditional KBOT in the following two aspects: 1) it is less affected by the object scale change; 2) it is less prone to appearance variation. Anbang Yao, Guijin Wang, Xinggang Lin |
ICASSP | 1 |