Yurong Chen 0001

dblp:02/41-1 · DBLP profile ↗
← Back
51ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-9333-1746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 since 2021Systems, architecture and hardware · 10 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Grid Convolution for 3D Human Pose Estimation
abstract
3D human pose estimation from 2D keypoint observation has been used in many human-centered computer vision applications. In this work, we tackle the task by formulating a novel grid representation learning paradigm that relies on grid convolution (GridConv), mimicking the wisdom of regular convolution operations in image space. GridConv is defined based on Semantic Grid Transformation (SGT) which leverages a binary assignment matrix to map standard skeleton 2D pose onto a regular weave-like grid pose joint by joint. We provide two ways to implement SGT: handcrafted and learnable SGT. Surprisingly, both designs turn out to achieve promising results and the learnable one is better, demonstrating the great potential of this new lifting representation learning formulation. To improve the ability of GridConv to encode contextual cues, we introduce an attention module over the convolutional kernel, making grid convolution operations input-dependent, spatial-aware and grid-specific. Besides our spatial grid lifting network for single-frame input, we also present a spatial-temporal grid lifting network for video-based input, which relies on an efficient multi-scale grid learning strategy to encode spatial and temporal joint variations. Extensive experiments demonstrate that the proposed grid lifting network outperforms existing approaches by remarkable margins on Human3.6M and MPI-INF-3DHP datasets. Our grid lifting networks also exhibit good generalization ability across three other keypoint-based tasks: 3D hand pose estimation, head pose estimation, and action recognition.
Yangyuxuan Kang, Anbang Yao, Shandong Wang, Yurong Chen 0001, Enhua Wu
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Ace-of-Spades: Accelerating Spatially Sparse Convolution for 3D Scene Understanding
abstract
Semantic understanding of 3D scenes is fundamental to many applications like robotics, autonomous driving, AR/VR. State-of-the-art methods for different 3D scene understanding tasks use 3D convolutional neural networks (CNNs) operating on point clouds. Convolution on spatially sparse data like point cloud involves irregular data accesses and compute patterns leading to poor utilization and energy efficiency in CPU/GPU implementations. The existing CNN accelerators designed for weight/activation sparsity cannot be efficiently repurposed for 3D spatially sparse CNNs given the fundamental differences in locating non-zero operands and granularity of work-dispatches. To address the dataflow challenges due to spatial sparsity and the need for specialized microarchitecture for spatially sparse convolution we present Ace-of-Spade s (AoS), an algorithm-dataflow-architecture co-designed system. AoS enables the data reuse among spatially proximate points using a locality-aware metadata structure along with a surface orientation aware point cloud reordering algorithm. AoS uses a novel technique for spatial sparsity aware selection of optimal data tiles by modelling the sparsity induced variations in the point cloud with a near-zero latency overheads. To accelerate computation on spatially sparse data, we propose a novel hardware accelerator Ss p nna with a front-end to convert varying number of operations per point into a stream of dense work dispatches to the backend compute engine. The compute engine further exploits weight and input feature data reuse through dynamic systolic grouping and multicast interconnects. The Ss p nna core together with the 64 KB of L1 memory requires 0.31 mm 2 of area in 10nm process at 1 GHz. Overall, AoS achieves speedup/energy savings of 19.9x / 49.9x and 2.2x / 7.1x over the state-of-the-art CPU and GPU implementations respectively.
Om Ji Omer, Prashant Laddha, Gurpreet S. Kalsi, Kamlesh R. Pillai, Anirud Thyagharajan, Ahimanyu Kulkarni, Anbang Yao, Yurong Chen 0001, Sreenivas Subramoney
ACM Trans. Embed. Comput. Syst.8
2024 ECT: Fine-grained edge detection with learned cause tokens
Shaocong Xu, Xiaoxue Chen, Yuhang Zheng 0004, Guyue Zhou, Yurong Chen 0001, Hongbin Zha, Hao Zhao 0002
Image Vis. Comput.5
2023 CABM: Content-Aware Bit Mapping for Single Image Super-Resolution Network with Large Input
abstract
With the development of high-definition display devices, the practical scenario of Super-Resolution (SR) usually needs to super-resolve large input like 2K to higher resolution (4K/8K). To reduce the computational and memory cost, current methods first split the large input into local patches and then merge the SR patches into the output. These methods adaptively allocate a subnet for each patch. Quantization is a very important technique for network acceleration and has been used to design the subnets. Current methods train an MLP bit selector to determine the propoer bit for each layer. However, they uniformly sample subnets for training, making simple subnets overfitted and complicated subnets underfitted. Therefore, the trained bit selector fails to determine the optimal bit. Apart from this, the introduced bit selector brings additional cost to each layer of the$SR$network. In this paper, we propose a novel method named Content-Aware Bit Mapping (CABM), which can remove the bit selector without any performance loss. CABM also learns a bit selector for each layer during training. After training, we analyze the relation between the edge information of an input patch and the bit of each layer. We observe that the edge information can be an effective metric for the selected bit. Therefore, we design a strategy to build an Edge-to-Bit lookup table that maps the edge score of a patch to the bit of each layer during inference. The bit configuration of SR network can be determined by the lookup tables of all layers. Our strategy can find better bit configuration, resulting in more efficient mixed precision networks. We conduct detailed experiments to demonstrate the generalization ability of our method. The code will be released.
Senmao Tian, Ming Lu 0002, Jiaming Liu 0003, Yandong Guo, Yurong Chen 0001, Shunli Zhang 0005
CVPR5
2023 Ske2Grid: Skeleton-to-Grid Representation Learning for Action Recognition
abstract
This paper presents Ske2Grid, a new representation learning framework for improved skeleton-based action recognition. In Ske2Grid, we define a regular convolution operation upon a novel grid representation of human skeleton, which is a compact image-like grid patch constructed and learned through three novel designs. Specifically, we propose a graph-node index transform (GIT) to construct a regular grid patch through assigning the nodes in the skeleton graph one by one to the desired grid cells. To ensure that GIT is a bijection and enrich the expressiveness of the grid representation, an up-sampling transform (UPT) is learned to interpolate the skeleton graph nodes for filling the grid patch to the full. To resolve the problem when the one-step UPT is aggressive and further exploit the representation capability of the grid patch with increasing spatial size, a progressive learning strategy (PLS) is proposed which decouples the UPT into multiple steps and aligns them to multiple paired GITs through a compact cascaded design learned progressively. We construct networks upon prevailing graph convolution networks and conduct experiments on six mainstream skeleton-based action recognition datasets. Experiments show that our Ske2Grid significantly outperforms existing GCN-based solutions under different benchmark settings, without bells and whistles. Code and models are available at https://github.com/OSVAI/Ske2Grid.
Yangyuxuan Kang, Anbang Yao, Yurong Chen 0001
ICML4
2023 From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds
abstract
Room layout estimation is a long-existing robotic vision task that benefits both environment sensing and motion planning. However, layout estimation using point clouds (PCs) still suffers from data scarcity due to annotation difficulty. As such, we address the semi-supervised setting of this task based upon the idea of model exponential moving averaging. But adapting this scheme to the state-of-the-art (SOTA) solution for PC-based layout estimation is not straightforward. To this end, we define a quad set matching strategy and several consistency losses based upon metrics tailored for layout quads. Besides, we propose a new online pseudo-label harvesting algorithm that decomposes the distribution of a hybrid distance measure between quads and PC into two components. This technique does not need manual threshold selection and intuitively encourages quads to align with reliable layout points. Surprisingly, this framework also works for the fully-supervised setting, achieving a new SOTA on the ScanNet benchmark. Last but not least, we also push the semi-supervised setting to the realistic omni-supervised setting, demonstrating significantly promoted performance on a newly annotated ARKitScenes testing set. Our codes, data and models are made publicly available**Code: https://github.com/AIR-DISCOVER/Omni-PQ.
Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Yurong Chen 0001, Hongbin Zha
ICRA7
2022 Efficient Meta-Tuning for Content-Aware Neural Video Delivery
Xiaoqi Li 0009, Jiaming Liu 0003, Shizun Wang, Ming Lu 0002, Yurong Chen 0001, Anbang Yao, Yandong Guo, Shanghang Zhang
ECCV (18)6
2022 OANet: Learning Two-View Correspondences and Geometry Using Order-Aware Network
abstract
Establishing correct correspondences between two images should consider both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential or fundamental matrix. Specifically, this proposed network is built hierarchically and comprises three operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in canonical order and invariant to input permutations. Next, the clusters are spatially correlated to encode the global context of correspondences. After that, the context-encoded clusters are interpolated back to the original size and position to build a hierarchical architecture. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts. Besides, based on the proposed method and advanced local feature, we won the first place in CVPR 2019 image matching workshop challenge and also achieve state-of-the-art results in the Visual Localization benchmark. Code is available at https://github.com/zjhthu/OANet.
Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Long Quan, Hongen Liao
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 Overfitting the Data: Compact Neural Video Delivery via Content-aware Feature Modulation
abstract
Internet video delivery has undergone a tremendous explosion of growth over the past few years. However, the quality of video delivery system greatly depends on the Internet bandwidth. Deep Neural Networks (DNNs) are utilized to improve the quality of video delivery recently. These methods divide a video into chunks, and stream LR video chunks and corresponding content-aware models to the client. The client runs the inference of models to super-resolve the LR chunks. Consequently, a large number of models are streamed in order to deliver a video. In this paper, we first carefully study the relation between models of different chunks, then we tactfully design a joint training framework along with the Content-aware Feature Modulation (CaFM) layer to compress these models for neural video delivery. With our method, each video chunk only requires less than 1% of original parameters to be streamed, achieving even better SR performance. We conduct extensive experiments across various SR backbones, video time length, and scaling factors to demonstrate the advantages of our method. Besides, our method can be also viewed as a new approach of video coding. Our primary experiments achieve better video quality compared with the commercial H.264 and H.265 standard under the same storage cost, showing the great potential of the proposed method. Code is available at: https://github.com/Neural-video-delivery/ CaFM-Pytorch-ICCV2021
Jiaming Liu 0003, Ming Lu 0002, Kaixin Chen 0001, Xiaoqi Li 0009, Shizun Wang, Zhaoqing Wang, Enhua Wu, Yurong Chen 0001, Ming Wu 0001
ICCV8
2021 Dynamic Normalization and Relay for Video Action Recognition
abstract
Convolutional Neural Networks (CNNs) have been the dominant model for video action recognition. Due to the huge memory and compute demand, popular action recognition networks need to be trained with small batch sizes, which makes learning discriminative spatial-temporal representations for videos become a challenging problem. In this paper, we present Dynamic Normalization and Relay (DNR), an improved normalization design, to augment the spatial-temporal representation learning of any deep action recognition model, adapting to small batch size training settings. We observe that state-of-the-art action recognition networks usually apply the same normalization parameters to all video data, and ignore the dependencies of the estimated normalization parameters between neighboring frames (at the same layer) and between neighboring layers (with all frames of a video clip). Inspired by this, DNR introduces two dynamic normalization relay modules to explore the potentials of cross-temporal and cross-layer feature distribution dependencies for estimating accurate layer-wise normalization parameters. These two DNR modules are instantiated as a light-weight recurrent structure conditioned on the current input features, and the normalization parameters estimated from the neighboring frames based features at the same layer or from the whole video clip based features at the preceding layers. We first plug DNR into prevailing 2D CNN backbones and test its performance on public action recognition datasets including Kinetics and Something-Something. Experimental results show that DNR brings large performance improvements to the baselines, achieving over 4.4% absolute margins in top-1 accuracy without training bells and whistles. More experiments on 3D backbones and several latest 2D spatial-temporal networks further validate its effectiveness. Code will be available at https://github.com/caidonkey/dnr.
Anbang Yao, Yurong Chen 0001
NeurIPS3
2021 On Connections Between Regularizations for Improving DNN Robustness
abstract
This paper analyzes regularization terms proposed recently for improving the adversarial robustness of deep neural networks (DNNs), from a theoretical point of view. Specifically, we study possible connections between several effective methods, including input-gradient regularization, Jacobian regularization, curvature regularization, and a cross-Lipschitz functional. We investigate them on DNNs with general rectified linear activations, which constitute one of the most prevalent families of models for image classification and a host of other machine learning applications. We shed light on essential ingredients of these regularizations and re-interpret their functionality. Through the lens of our study, more principled and efficient regularizations can possibly be invented in the near future.
Yiwen Guo, Yurong Chen 0001, Changshui Zhang
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Deep Likelihood Network for Image Restoration With Multiple Degradation Levels
abstract
Convolutional neural networks have been proven effective in a variety of image restoration tasks. Most state-of-the-art solutions, however, are trained using images with a single particular degradation level, and their performance deteriorates drastically when applied to other degradation settings. In this paper, we propose deep likelihood network (DL-Net), aiming at generalizing off-the-shelf image restoration networks to succeed over a spectrum of degradation levels. We slightly modify an off-the-shelf network by appending a simple recursive module, which is derived from a fidelity term, for disentangling the computation for multiple degradation levels. Extensive experimental results on image inpainting, interpolation, and super-resolution show the effectiveness of our DL-Net.
Yiwen Guo, Ming Lu 0002, Wangmeng Zuo, Changshui Zhang, Yurong Chen 0001
IEEE Trans. Image Process.5
2020 Explicit Residual Descent for 3D Human Pose Estimation from 2D Joint Locations
Yangyuxuan Kang, Anbang Yao, Shandong Wang, Ming Lu 0002, Yurong Chen 0001, Enhua Wu
BMVC5
2020 CASNet: Common Attribute Support Network for image instance and panoptic segmentation
abstract
Instance segmentation and panoptic segmentation is being paid more and more attention in recent years. In comparison with bounding box based object detection and semantic segmentation, instance segmentation can provide more analytical results at pixel level. Given the insight that pixels belonging to one instance have one or more common attributes of current instance, we bring up an one-stage instance segmentation network named Common Attribute Support Network (CASNet), which realizes instance segmentation by predicting and clustering common attributes. CASNet is designed in the manner of fully convolutional and can implement training and inference from end to end. And CASNet manages predicting the instance without overlaps and holes, which problem exists in most of current instance segmentation algorithms. Furthermore, it can be easily extended to panoptic segmentation through minor modifications with little computation overhead. CASNet builds a bridge between semantic and instance segmentation from finding pixel class ID to obtaining class and instance ID by operations on common attribute. Through experiment for instance and panoptic segmentation, CASNet gets mAP 32.8% and PQ 59.0% on Cityscapes validation dataset by joint training, and mAP 36.3% and PQ 66.1% by separated training mode. For panoptic segmentation, CASNet gets state-of-the-art performance on the Cityscapes validation dataset.
Yuqing Hou, Anbang Yao, Yurong Chen 0001, Keqiang Li 0002
ICPR4
2020 Pointly-supervised scene parsing with uncertainty mixture
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023
Comput. Vis. Image Underst.5
2020 Learning to Draw Sight Lines
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023
Int. J. Comput. Vis.4
2020 Object Detection from Scratch with Deep Supervision
abstract
In this paper, we propose Deeply Supervised Object Detectors (DSOD), an object detection framework that can be trained from scratch. Recent advances in object detection heavily depend on the off-the-shelf models pre-trained on large-scale classification datasets like ImageNet and OpenImage. However, one problem is that adopting pre-trained models from classification to detection task may incur learning bias due to the different objective function and diverse distributions of object categories. Techniques like fine-tuning on detection task could alleviate this issue to some extent but are still not fundamental. Furthermore, transferring these pre-trained models across discrepant domains will be more difficult (e.g., from RGB to depth images). Thus, a better solution to handle these critical problems is to train object detectors from scratch, which motivates our proposed method. Previous efforts on this direction mainly failed by reasons of the limited training data and naive backbone network structures for object detection. In DSOD, we contribute a set of design principles for learning object detectors from scratch. One of the key principles is the deep supervision, enabled by layer-wise dense connections in both backbone networks and prediction layers, plays a critical role in learning good detectors from scratch. After involving several other principles, we build our DSOD based on the single-shot detection framework (SSD). We evaluate our method on PASCAL VOC 2007, 2012 and COCO datasets. DSOD achieves consistently better results than the state-of-the-art methods with much more compact models. Specifically, DSOD outperforms baseline method SSD on all three benchmarks, while requiring only 1/2 parameters. We also observe that DSOD can achieve comparable/slightly better results than Mask RCNN [1] + FPN [2] (under similar input size) with only 1/3 parameters, using no extra data or pre-trained models.
Zhuang Liu 0003, Yu-Gang Jiang 0001, Yurong Chen 0001, Xiangyang Xue 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 A Closed-Form Solution to Universal Style Transfer
abstract
Universal style transfer tries to explicitly minimize the losses in feature space, thus it does not require training on any pre-defined styles. It usually uses different layers of VGG network as the encoders and trains several decoders to invert the features into images. Therefore, the effect of style transfer is achieved by feature transform. Although plenty of methods have been proposed, a theoretical analysis of feature transform is still missing. In this paper, we first propose a novel interpretation by treating it as the optimal transport problem. Then, we demonstrate the relations of our formulation with former works like Adaptive Instance Normalization (AdaIN) and Whitening and Coloring Transform (WCT). Finally, we derive a closed-form solution named Optimal Style Transfer (OST) under our formulation by additionally considering the content loss of Gatys. Comparatively, our solution can preserve better structure and achieve visually pleasing results. It is simple yet effective and we demonstrate its advantages both quantitatively and qualitatively. Besides, we hope our theoretical analysis can inspire future works in neural style transfer.
Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Feng Xu 0005, Li Zhang 0023
ICCV4
2019 Learning Two-View Correspondences and Geometry Using Order-Aware Network
abstract
Establishing correspondences between two images requires both local and global spatial context. Given putative correspondences of feature points in two views, in this paper, we propose Order-Aware Network, which infers the probabilities of correspondences being inliers and regresses the relative pose encoded by the essential matrix. Specifically, this proposed network is built hierarchically and comprises three novel operations. First, to capture the local context of sparse correspondences, the network clusters unordered input correspondences by learning a soft assignment matrix. These clusters are in a canonical order and invariant to input permutations. Next, the clusters are spatially correlated to form the global context of correspondences. After that, the context-encoded clusters are recovered back to the original size through a proposed upsampling operator. We intensively experiment on both outdoor and indoor datasets. The accuracy of the two-view geometry and correspondences are significantly improved over the state-of-the-arts.
Dawei Sun 0007, Zixin Luo, Anbang Yao, Lei Zhou 0011, Tianwei Shen, Yurong Chen 0001, Hongen Liao, Long Quan
ICCV7
2019 Stochastic Quantization for Learning Accurate Low-Bit Deep Neural Networks
Yinpeng Dong, Renkun Ni, Yurong Chen 0001, Hang Su 0006, Jun Zhu 0001
Int. J. Comput. Vis.4
2018 Network Decoupling: From Regular to Depthwise Separable Convolutions
Jianbo Guo, Yuxi Li 0009, Weiyao Lin, Yurong Chen 0001
BMVC4
2018 Learning Visual Knowledge Memory Networks for Visual Question Answering
abstract
Visual question answering (VQA) requires joint comprehension of images and natural language questions, where many questions can't be directly or clearly answered from visual content but require reasoning from structured human knowledge with confirmation from visual content. This paper proposes visual knowledge memory network (VKMN) to address this issue, which seamlessly incorporates structured human knowledge and deep visual features into memory networks in an end-to-end learning framework. Comparing to existing methods for leveraging external knowledge for supporting VQA, this paper stresses more on two missing mechanisms. First is the mechanism for integrating visual contents with knowledge facts. VKMN handles this issue by embedding knowledge triples (subject, relation, target) and deep visual features jointly into the visual knowledge features. Second is the mechanism for handling multiple knowledge facts expanding from question and answer pairs. VKMN stores joint embedding using key-value pair structure in the memory networks so that it is easy to handle multiple facts. Experiments show that the proposed method achieves promising results on both VQA v1.0 and v2.0 benchmarks, while outperforms state-of-the-art methods on the knowledge-reasoning related questions.
Yinpeng Dong, Yurong Chen 0001
CVPR5
2018 Explicit Loss-Error-Aware Quantization for Low-Bit Deep Neural Networks
abstract
Benefiting from tens of millions of hierarchically stacked learnable parameters, Deep Neural Networks (DNNs) have demonstrated overwhelming accuracy on a variety of artificial intelligence tasks. However reversely, the large size of DNN models lays a heavy burden on storage, computation and power consumption, which prohibits their deployments on the embedded and mobile systems. In this paper, we propose Explicit Loss-error-aware Quantization (ELQ), a new method that can train DNN models with very low-bit parameter values such as ternary and binary ones to approximate 32-bit floating-point counterparts without noticeable loss of predication accuracy. Unlike existing methods that usually pose the problem as a straightforward approximation of the layer-wise weights or outputs of the original full-precision model (specifically, minimizing the error of the layer-wise weights or inner products of the weights and the inputs between the original and respective quantized models), our ELQ elaborately bridges the loss perturbation from the weight quantization and an incremental quantization strategy to address DNN quantization. Through explicitly regularizing the loss perturbation and the weight approximation error in an incremental way, we show that such a new optimization method is theoretically reasonable and practically effective. As validated with two mainstream convolutional neural network families (i.e., fully convolutional and non-fully convolutional), our ELQ shows better results than state-of-the-art quantization methods on the large scale ImageNet classification dataset. Code will be made publicly available.
Aojun Zhou, Anbang Yao, Yurong Chen 0001
CVPR4
2018 Efficient Semantic Scene Completion Network with Spatial Group Convolution
Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023, Hongen Liao
ECCV (12)4
2018 Sparse DNNs with Improved Adversarial Robustness
abstract
Deep neural networks (DNNs) are computationally/memory-intensive and vulnerable to adversarial attacks, making them prohibitive in some real-world applications. By converting dense models into sparse ones, pruning appears to be a promising solution to reducing the computation/memory cost. This paper studies classification models, especially DNN-based ones, to demonstrate that there exists intrinsic relationships between their sparsity and adversarial robustness. Our analyses reveal, both theoretically and empirically, that nonlinear DNN-based classifiers behave differently under $l_2$ attacks from some linear ones. We further demonstrate that an appropriately higher model sparsity implies better robustness of nonlinear DNNs, whereas over-sparsified models can be more difficult to resist adversarial examples.
Yiwen Guo, Changshui Zhang, Yurong Chen 0001
NeurIPS4
2017 Network Sketching: Exploiting Binary Structure in Deep CNNs
Yiwen Guo, Anbang Yao, Hao Zhao 0002, Yurong Chen 0001
CVPR4
2017 RON: Reverse Connection with Objectness Prior Networks for Object Detection
abstract
We present RON, an efficient and effective framework for generic object detection. Our motivation is to smartly associate the best of the region-based (e.g., Faster R-CNN) and region-free (e.g., SSD) methodologies. Under fully convolutional architecture, RON mainly focuses on two fundamental problems: (a) multi-scale object localization and (b) negative sample mining. To address (a), we design the reverse connection, which enables the network to detect objects on multi-levels of CNNs. To deal with (b), we propose the objectness prior to significantly reduce the searching space of objects. We optimize the reverse connection, objectness prior and object detector jointly by a multi-task loss function, thus RON can directly predict final detection results from all locations of various feature maps. Extensive experiments on the challenging PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO benchmarks demonstrate the competitive performance of RON. Specifically, with VGG-16 and low resolution 384×384 input size, the network gets 81.3% mAP on PASCAL VOC 2007, 80.7% mAP on PASCAL VOC 2012 datasets. Its superiority increases when datasets become larger and more difficult, as demonstrated by the results on the MS COCO dataset. With 1.5G GPU memory at test phase, the speed of the network is 15 FPS, 3 times faster than the Faster R-CNN counterpart. Code will be made publicly available.
Tao Kong, Fuchun Sun 0001, Anbang Yao, Huaping Liu 0001, Ming Lu 0002, Yurong Chen 0001
CVPR6
2017 Weakly Supervised Dense Video Captioning
abstract
This paper focuses on a novel and challenging vision task, dense video captioning, which aims to automatically describe a video clip with multiple informative and diverse caption sentences. The proposed method is trained without explicit annotation of fine-grained sentence to video region-sequence correspondence, but is only based on weak video-level sentence annotations. It differs from existing video captioning systems in three technical aspects. First, we propose lexical fully convolutional neural networks (Lexical-FCN) with weakly supervised multi-instance multi-label learning to weakly link video regions with lexical labels. Second, we introduce a novel submodular maximization scheme to generate multiple informative and diverse region-sequences based on the Lexical-FCN outputs. A winner-takes-all scheme is adopted to weakly associate sentences to region-sequences in the training phase. Third, a sequence-to-sequence learning based language model is trained with the weakly supervised information obtained through the association process. We show that the proposed method can not only produce informative and diverse dense captions, but also outperform state-of-the-art single video captioning methods by a large margin.
Minjun Li, Yurong Chen 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001
CVPR5
2017 Physics Inspired Optimization on Semantic Transfer Features: An Alternative Method for Room Layout Estimation
abstract
In this paper, we propose an alternative method to estimate room layouts of cluttered indoor scenes. This method enjoys the benefits of two novel techniques. The first one is semantic transfer (ST), which is: (1) a formulation to integrate the relationship between scene clutter and room layout into convolutional neural networks, (2) an architecture that can be end-to-end trained, (3) a practical strategy to initialize weights for very deep networks under unbalanced training data distribution. ST allows us to extract highly robust features under various circumstances, and in order to address the computation redundance hidden in these features we develop a principled and efficient inference scheme named physics inspired optimization (PIO). PIOs basic idea is to formulate some phenomena observed in ST features into mechanics concepts. Evaluations on public datasets LSUN and Hedau show that the proposed method is more accurate than state-of-the-art methods.
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023
CVPR5
2017 Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style Transfer
abstract
Recently, the community of style transfer is trying to incorporate semantic information into traditional system. This practice achieves better perceptual results by transferring the style between semantically-corresponding regions. Yet, few efforts are invested to address the computation bottleneck of back-propagation. In this paper, we propose a new framework for fast semantic style transfer. Our method decomposes the semantic style transfer problem into feature reconstruction part and feature decoder part. The reconstruction part tactfully solves the optimization problem of content loss and style loss in feature space by particularly reconstructed feature. This significantly reduces the computation of propagating the loss through the whole network. The decoder part transforms the reconstructed feature into the stylized image. Through a careful bridging of the two modules, the proposed approach not only achieves competitive results as backward optimization methods but also is about two orders of magnitude faster.
Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Feng Xu 0005, Yurong Chen 0001, Li Zhang 0023
ICCV5
2017 DSOD: Learning Deeply Supervised Object Detectors from Scratch
abstract
We present Deeply Supervised Object Detector (DSOD), a framework that can learn object detectors from scratch. State-of-the-art object objectors rely heavily on the off the-shelf networks pre-trained on large-scale classification datasets like Image Net, which incurs learning bias due to the difference on both the loss functions and the category distributions between classification and detection tasks. Model fine-tuning for the detection task could alleviate this bias to some extent but not fundamentally. Besides, transferring pre-trained models from classification to detection between discrepant domains is even more difficult (e.g. RGB to depth images). A better solution to tackle these two critical problems is to train object detectors from scratch, which motivates our proposed DSOD. Previous efforts in this direction mostly failed due to much more complicated loss functions and limited training data in object detection. In DSOD, we contribute a set of design principles for training object detectors from scratch. One of the key findings is that deep supervision, enabled by dense layer-wise connections, plays a critical role in learning a good detector. Combining with several other principles, we develop DSOD following the single-shot detection (SSD) framework. Experiments on PASCAL VOC 2007, 2012 and MS COCO datasets demonstrate that DSOD can achieve better results than the state-of-the-art solutions with much more compact models. For instance, DSOD outperforms SSD on all three benchmarks with real-time detection speed, while requires only 1/2 parameters to SSD and 1/10 parameters to Faster RCNN.
Zhuang Liu 0003, Yu-Gang Jiang 0001, Yurong Chen 0001, Xiangyang Xue 0001
ICCV5
2017 Incremental Network Quantization: Towards Lossless CNNs with Low-precision Weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Yurong Chen 0001
ICLR (Poster)5
2017 Learning supervised scoring ensemble for emotion recognition in the wild
abstract
State-of-the-art approaches for the previous emotion recognition in the wild challenges are usually built on prevailing Convolutional Neural Networks (CNNs). Although there is clear evidence that CNNs with increased depth or width can usually bring improved predication accuracy, existing top approaches provide supervision only at the output feature layer, resulting in the insufficient training of deep CNN models. In this paper, we present a new learning method named Supervised Scoring Ensemble (SSE) for advancing this challenge with deep CNNs. We first extend the idea of recent deep supervision to deal with emotion recognition problem. Benefiting from adding supervision not only to deep layers but also to intermediate layers and shallow layers, the training of deep CNNs can be well eased. Second, we present a new fusion structure in which class-wise scoring activations at diverse complementary feature layers are concatenated and further used as the inputs for second-level supervision, acting as a deep feature ensemble within a single CNN architecture. We show our proposed learning method brings large accuracy gains over diverse backbone networks consistently. On this year's audio-video based emotion recognition task, the average recognition rate of our best submission is 60.34%, forming a new envelop over all existing records.
Shandong Wang, Anbang Yao, Yurong Chen 0001
ICMI5
2016 Regional Gating Neural Networks for Multi-label Image Classification
Yurong Chen 0001, Jia-Ming Liu, Yu-Gang Jiang 0001, Xiangyang Xue 0001
BMVC3
2016 HyperNet: Towards Accurate Region Proposal Generation and Joint Object Detection
abstract
Almost all of the current top-performing object detection networks employ region proposals to guide the search for object instances. State-of-the-art region proposal methods usually need several thousand proposals to get high recall, thus hurting the detection efficiency. Although the latest Region Proposal Network method gets promising detection accuracy with several hundred proposals, it still struggles in small-size object detection and precise localization (e.g., large IoU thresholds), mainly due to the coarseness of its feature maps. In this paper, we present a deep hierarchical network, namely HyperNet, for handling region proposal generation and object detection jointly. Our HyperNet is primarily based on an elaborately designed Hyper Feature which aggregates hierarchical feature maps first and then compresses them into a uniform space. The Hyper Features well incorporate deep but highly semantic, intermediate but really complementary, and shallow but naturally high-resolution features of the image, thus enabling us to construct HyperNet by sharing them both in generating proposals and detecting objects via an end-to-end joint training strategy. For the deep VGG16 model, our method achieves completely leading recall and state-of-the-art object detection accuracy on PASCAL VOC 2007 and 2012 using only 100 proposals per image. It runs with a speed of 5 fps (including all steps) on a GPU, thus having the potential for real-time processing.
Tao Kong, Anbang Yao, Yurong Chen 0001, Fuchun Sun 0001
CVPR3
2016 HoloNet: towards robust emotion recognition in the wild
abstract
In this paper, we present HoloNet, a well-designed Convolutional Neural Network (CNN) architecture regarding our submissions to the video based sub-challenge of the Emotion Recognition in the Wild (EmotiW) 2016 challenge. In contrast to previous related methods that usually adopt relatively simple and shallow neural network architectures to address emotion recognition task, our HoloNet has three critical considerations in network design. (1) To reduce redundant filters and enhance the non-saturated non-linearity in the lower convolutional layers, we use a modified Concatenated Rectified Linear Unit (CReLU) instead of ReLU. (2) To enjoy the accuracy gain from considerably increased network depth and maintain efficiency, we combine residual structure and CReLU to construct the middle layers. (3) To broaden network width and introduce multi-scale feature extraction property, the topper layers are designed as a variant of inception-residual structure. The main benefit of grouping these modules into the HoloNet is that both negative and positive phase information implicitly contained in the input data can flow over it in multiple paths, thus deep multi-scale features explicitly capturing emotion variation can be well extracted from multi-path sibling layers, and then can be further concatenated for robust recognition. We obtain competitive results in this year’s video based emotion recognition sub-challenge using an ensemble of two HoloNet models trained with given data only. Specifically, we obtain a mean recognition rate of 57.84%, outperforming the baseline accuracy with an absolute margin of 17.37%, and yielding 4.04% absolute accuracy gain compared to the result of last year’s winner team. Meanwhile, our method runs with a speed of several thousands of frames per second on a GPU, thus it is well applicable to real-time scenarios.
Anbang Yao, Shandong Wang, Liang Sha, Yurong Chen 0001
ICMI6
2016 Dynamic Network Surgery for Efficient DNNs
abstract
Deep learning has become a ubiquitous technology to improve machine intelligence. However, most of the existing deep models are structurally very complex, making them difficult to be deployed on the mobile platforms with limited computational power. In this paper, we propose a novel network compression method called dynamic network surgery, which can remarkably reduce the network complexity by making on-the-fly connection pruning. Unlike the previous methods which accomplish this task in a greedy way, we properly incorporate connection splicing into the whole process to avoid incorrect pruning and make it as a continual network maintenance. The effectiveness of our method is proved with experiments. Without any accuracy loss, our method can efficiently compress the number of parameters in LeNet-5 and AlexNet by a factor of $\bm{108}\times$ and $\bm{17.7}\times$ respectively, proving that it outperforms the recent pruning method by considerable margins. Code and some models are available at https://github.com/yiwenguo/Dynamic-Network-Surgery.
Yiwen Guo, Anbang Yao, Yurong Chen 0001
NIPS3
2016 A Bayesian Hashing approach and its application to face recognition
Qi Dai 0001, Jun Wang 0006, Yurong Chen 0001, Yu-Gang Jiang 0001
Neurocomputing4
2015 Capturing AU-Aware Facial Features and Their Latent Relations for Emotion Recognition in the Wild
abstract
The Emotion Recognition in the Wild (EmotiW) Challenge has been held for three years. Previous winner teams primarily focus on designing specific deep neural networks or fusing diverse hand-crafted and deep convolutional features. They all neglect to explore the significance of the latent relations among changing features resulted from facial muscle motions. In this paper, we study this recognition challenge from the perspective of analyzing the relations among expression-specific facial features in an explicit manner. Our method has three key components. First, we propose a pair-wise learning strategy to automatically seek a set of facial image patches which are important for discriminating two particular emotion categories. We found these learnt local patches are in part consistent with the locations of expression-specific Action Units (AUs), thus the features extracted from such kind of facial patches are named AU-aware facial features. Second, in each pair-wise task, we use an undirected graph structure, which takes learnt facial patches as individual vertices, to encode feature relations between any two learnt facial patches. Finally, a robust emotion representation is constructed by concatenating all task-specific graph-structured facial feature relations sequentially. Extensive experiments on the EmotiW 2015 Challenge testify the efficacy of the proposed approach. Without using additional data, our final submissions achieved competitive results on both sub-challenges including the image based static facial expression recognition (we got 55.38% recognition accuracy outperforming the baseline 39.13% with a margin of 16.25%) and the audio-video based emotion recognition (we got 53.80% recognition accuracy outperforming the baseline 39.33% and the 2014 winner team's final result 50.37% with the margins of 14.47% and 3.43%, respectively).
Anbang Yao, Junchao Shao, Ningning Ma, Yurong Chen 0001
ICMI4
2015 Optimal Bayesian Hashing for Efficient Face Recognition
Qi Dai 0001, Jun Wang 0006, Yurong Chen 0001, Yu-Gang Jiang 0001
IJCAI4
2010 Bundled depth-map merging for multi-view stereo
abstract
Depth-map merging is one typical technique category for multi-view stereo (MVS) reconstruction. To guarantee accuracy, existing algorithms usually require either sub-pixel level stereo matching precision or continuous depth-map estimation. The merging of inaccurate depth-maps remains a challenging problem. This paper introduces a bundle optimization method for robust and accurate depth-map merging. In the method, depth-maps are generated using DAISY feature, followed by two stages of bundle optimization. The first stage optimizes the track of connected stereo matches to generate initial 3D points. The second stage optimizes the position and normals of 3D points. High quality point cloud is then meshed as geometric models. The proposed method can be easily parallelizable on multi-core processors. Middlebury evaluation shows that it is one of the most efficient methods among non-GPU algorithms, yet still keeps very high accuracy. We also demonstrate the effectiveness of the proposed algorithm on various real-world, high-resolution, self-calibrated data sets including objects with complex details, objects with large area of highlight, and objects with non-Lambertian surface.
Eric Q. Li, Yurong Chen 0001, Yimin Zhang 0002
CVPR3
2010 A general texture mapping framework for image-based 3D modeling
abstract
This paper presents a general texture mapping framework for image-based 3D modeling. It aims to generating seamless texture map for 3D model created by real-world photos under uncontrolled environment. Our proposed method addresses two challenging problems: 1) texture discontinuity due to system error in 3D modeling from self-calibration; 2) color/lighting difference among images due to real-world uncontrolled environments. The general framework contains two stages to resolve these problems. The first stage globally optimizes the registration of texture patches and triangle faces with Markov Random Field (MRF) to optimize texture mosaic. The second stage does local radiometric correction to adjust color difference between texture patches and then blend texture boundaries to improve color continuity. The proposed method is evaluated on several 3D models by image-based 3D modeling, and demonstrates promising results.
Eric Q. Li, Yurong Chen 0001, Yimin Zhang 0002
ICIP4
2009 Parallelization and optimization of a CBVIR system on multi-core architectures
abstract
Technique advances have made image capture and storage very convenient, which results in an explosion of the amount of visual information. It becomes difficult to find useful information from these tremendous data. Content-based Visual Information Retrieval (CBVIR) is emerging as one of the best solutions to this problem. Unfortunately, CBVIR is a very compute-intensive task. Nowadays, with the boom of multi-core processors, CBVIR can be accelerated by exploiting multi-core processing capability. In this paper, we propose a parallelization implementation of a CBVIR system facing to server application and use some serial and parallel optimization techniques to improve its performance on an 8-core and on a 16-core systems. Experimental results show that optimized implementation can achieve very fast retrieval on the two multi-core systems.We also compare the performance of the application on the two multi-core systems and give an explanation of the performance difference between the two systems. Furthermore, we conduct detailed scalability and memory performance analysis to identify possible bottlenecks in the application. Based on these experimental results and performance analysis, we gain many insights into developing efficient applications on future multi-core architectures.
Qiankun Miao, Yurong Chen 0001, Yimin Zhang 0002, Guoliang Chen 0001
IPDPS2
2008 Parallelization and Characterization of Probabilistic Latent Semantic Analysis
abstract
Probabilistic Latent Semantic Analysis (PLSA) is one of the most popular statistical techniques for the analysis of two-model and co-occurrence data. It has applications in information retrieval and filtering, nature language processing, machine learning from text, and other related areas. However, PLSA is rarely applied to large datasets due to its high computational complexity.This paper presents an optimized and parallelized implementation of PLSA which is capable of processing datasets with 10000 documents in seconds. Compared to the baseline program, our parallelized program can achieve speedup of more than six on an eight-processor machine. The characterization of the parallel program is also presented. The performance analysis of the parallel program indicates that this program is memory intensive and the limited memory bandwidth is the bottleneck for better speedup.
Chuntao Hong, Jiulong Shan, Yurong Chen 0001, Yimin Zhang 0002
ICPP5
2008 SIFT implementation and optimization for multi-core systems
abstract
Scale invariant feature transform (SIFT) is an approach for extracting distinctive invariant features from images, and it has been successfully applied to many computer vision problems (e.g. face recognition and object detection). However, the SIFT feature extraction is compute-intensive, and a real-time or even super-real-time processing capability is required in many emerging scenarios. Nowadays, with the multi- core processor becoming mainstream, SIFT can be accelerated by fully utilizing the computing power of available multi-core processors. In this paper, we propose two parallel SIFT algorithms and present some optimization techniques to improve the implementation 's performance on multi-core systems. The result shows our improved parallel SIFT implementation can process general video images in super-real-time on a dual-socket, quad-core system, and the speed is much faster than the implementation on GPUs. We also conduct a detailed scalability and memory performance analysison the 8-core system and on a 32-core chip multiprocessor (CMP) simulator. The analysis helps us identify possible causes of bottlenecks, and we suggest avenues for scalability improvement to make this application more powerful on future large-scale multi- core systems.
Yurong Chen 0001, Yimin Zhang 0002
IPDPS2
2007 Parallelization and Performance Analysis of Video Feature Extractions on Multi-Core Based Systems
abstract
Content-based video information retrieval (CBVIR) has becoming one of the best solutions for retrieving useful information from today's video information explosion. And with the rapid development of modern technologies, CBVIR is emerging as a mass market desktop application. There is evidence that visual feature extraction is the most time-consuming part in a CBVIR system. In this paper, we implement three video visual feature extractions in parallel by exploring different kinds of thread-level parallelism. We also conduct detailed scalability and memory performance analysis on two multi-core based systems, in order to gain more insights into video-analysis related applications on future multi-core systems. From our analysis we identify the likely causes of bottlenecks in these kinds of applications and suggest ways to improve scalability.
Yurong Chen 0001, Yimin Zhang 0002
ICPP2
2007 Parallel Audio Quick Search on Shared-Memory Multiprocessor Systems
abstract
Audio search plays an important role in analyzing audio data and retrieving useful audio information. In this paper, a partially overlapping block-parallel active search method (POBPAS) is proposed to perform audio quick search on shared-memory multiprocessor systems (SMPs). This method uses a proper data segmentation to achieve parallelism and performs a high level of parallelism with little additional work. Several techniques including I/O optimization, proper data partition and dynamic scheduling are also introduced to maximize its scalability performance. In addition, we conduct a detailed performance characterization analysis of the parallel implementation of the POBPAS for three data sets on two Intel Xeon SMPs. Experimental results indicate that there are no obvious parallel limiting factors in the implementation except memory bandwidth. As a result, it can achieve 11.3X speedup for a larger data set (searching a 15 seconds' clip in a 27 hours' audio stream) on the 16-way processor system.
Yurong Chen 0001, Yimin Zhang 0002
IPDPS1
2007 Understanding the Memory Performance of Data-Mining Workloads on Small, Medium, and Large-Scale CMPs Using Hardware-Software Co-simulation
abstract
With the amount of data continuing to grow, extracting "data of interest" is becoming popular, pervasive, and more important than ever. Data mining, as this process is known as, seeks to draw meaningful conclusions, extract knowledge, and acquire models from vast amounts of data. These compute-intensive data-mining applications, where thread-level parallelism can be effectively exploited, are the design targets of future multi-core systems. As a result, future multi-core systems will be required to process terabyte-level workloads. To understand the memory system performance of data-mining applications, this paper presents the use of hardware-software co-simulation to explore the cache design space of several multi-threaded data mining applications. Our study reveals that the workloads are memory intensive, have large working-set sizes, and exhibit good data locality. We find that large DRAM caches can be useful to address their large working-set sizes
Wenlong Li 0003, Eric Q. Li, Aamer Jaleel, Jiulong Shan, Yurong Chen 0001, Qigang Wang, Ravi R. Iyer 0001, Ramesh Illikkal, Yimin Zhang 0002, Michael Liao, Jinhua Du
ISPASS5
2006 Parallel Information Extraction on Shared Memory Multi-processor System
abstract
Text mining is one of the best solutions for today and the future's information explosion. With the development of modern processor technologies, it will be a mass market desktop application in the many-core era. In text mining system, information extraction is a representative module and is the most compute intensive part. In this paper, we study the performance of parallel information extraction on shared memory multi-processor systems in order to gain some insights of such applications on the future's many-core architecture. In implementation, conditional random fields (CRFs) algorithm is selected as the core of module information extraction. Based on the newest CRFs toolkit FlexCRFs, we make several serial optimizations and then parallelize it with MPI and System V. IPC/shm. We also conduct a detailed performance analysis of this parallel application on the target system
Jiulong Shan, Yurong Chen 0001, Qian Diao, Yimin Zhang 0002
ICPP2
2006 Parallelization of module network structure learning and performance tuning on SMP
abstract
As an extension of Bayesian network, module network is an appropriate model for inferring causal network of a mass of variables from insufficient evidences. However learning such a model is still a time-consuming process. In this paper, we propose a parallel implementation of module network learning algorithm using OpenMP. We propose a static task partitioning strategy which distributes sub-search-spaces over worker threads to get the tradeoff between load-balance and software-cache-contention. To overcome performance penalties derived from shared-memory contention, we adopt several optimization techniques such as memory pre-allocation, memory alignment and static function usage. These optimizations have different patterns of influence on the sequential performance and the parallel speedup. Experiments validate the effectiveness of these optimizations. For a 2,200 nodes dataset, they enhance the parallel speedup up to 88%, together with a 2X sequential performance improvement. With resource contentions reduced, workload imbalance becomes the main hurdle to parallel scalability and the program behaviors more stable in various platforms.
Hongshan Jiang, Chunrong Lai, Yurong Chen 0001, Wei Hu 0002, Yimin Zhang 0002
IPDPS4
2006 Parallelization and performance characterization of protein 3D structure prediction of Rosetta
abstract
The prediction of protein 3D structure has become a hot research area in the post-genome era, through which people can understand a protein's function in health and disease, explore ways to control its actions and assist drug design. Many protein structure prediction approaches have been proposed in past decades. Among them, Rosetta is one of the best systems. However, the huge time complexity of Rosetta, e.g. a few days to predict a protein, limits its wide use in practice. To accelerate the prediction of protein 3D structure in Rosetta, this paper presents three different approaches, i.e., non-interactive, periodic interactive and asynchronous dynamic interactive scheme, to parallelize Rosetta. The asynchronous interactive scheme, with the adaptation of dynamic solution interaction, outperforms the other two, delivering much faster convergence speed and better solution quality. Detailed measurements and performance analysis also indicate that parallel Rosetta with asynchronous dynamic interactive scheme scales well
Wenlong Li 0003, Tao Wang 0003, Eric Q. Li, David Baker 0001, Steven Ge, Yurong Chen 0001, Yimin Zhang 0002
IPDPS7