Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shengtao Xiao

dblp:120/8229 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
1since 2021 · last 2027
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorArtificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Face, body and person analysis · 57% Video understanding and tracking · 17% Deep learning architectures and training · 17%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 13 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Face, body and person analysis
face alignment
0.632017
Recurrent 3D-2D Dual Learning for Large-Pose Facial Landmark Detection · ICCV 2017
Robust Facial Landmark Detection via Recurrent Attentive-Refinement Networks · ECCV (1) 2016
Integrated Face Analytics Networks through Cross-Dataset Hybrid Training · ACM Multimedia 2017
Computer vision › Face, body and person analysis › face alignment
cascaded regression
0.412019
Recurrent Shape Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Computer vision › Face, body and person analysis
human pose estimation
0.412019
Hierarchical Contextual Refinement Networks for Human Pose Estimation · IEEE Trans. Image Process. 2019
Machine learning › Deep learning architectures and training
recurrent neural network
0.412019
Recurrent Shape Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Machine learning › Deep learning architectures and training
hybrid learning
0.312017
Integrated Face Analytics Networks through Cross-Dataset Hybrid Training · ACM Multimedia 2017
Computer vision › Video understanding and tracking › object tracking › discriminative tracking
correlation filter tracking
0.212016
Recurrently Target-Attending Tracking · CVPR 2016
Computer vision › Face, body and person analysis
face detection
0.212016
A Live Face Swapper · ACM Multimedia 2016
Computer vision › Face, body and person analysis › face tracking
facial feature tracking
0.212016
A Live Face Swapper · ACM Multimedia 2016
Computer vision › Video understanding and tracking
object tracking
0.212016
Recurrently Target-Attending Tracking · CVPR 2016
Visual content generation and editing › face editing
face swapping
0.212016
A Live Face Swapper · ACM Multimedia 2016
Computer vision › Video understanding and tracking › object tracking
occlusion handling
0.222019
Recurrent Shape Regression · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Recurrently Target-Attending Tracking · CVPR 2016
Computer vision › Segmentation and scene understanding › part parsing
face parsing
0.112017
Integrated Face Analytics Networks through Cross-Dataset Hybrid Training · ACM Multimedia 2017
Computer vision › Face, body and person analysis › facial expression analysis
facial expression recognition
0.112017
Integrated Face Analytics Networks through Cross-Dataset Hybrid Training · ACM Multimedia 2017

Methods — techniques the papers use, named apart from their topics

recurrent neural network · 0.5progressive initialization · 0.5facial landmark tracking · 0.5deep learning-based face detection · 0.5virtual occlusion · 0.4tree-structured hierarchy · 0.4recurrent network · 0.4hierarchical contextual refinement · 0.4feature learning · 0.4contextual refinement unit · 0.4cascaded regression · 0.4recurrent dual learning · 0.33d-2d projection · 0.3
YearPublicationVenuePosition
2027 NestStruct-Net: A structure-aware 3D segmentation network for nested tumor subregions
Bin Ruan, Tiancai Yi, Shengtao Xiao, Yebin Huang, Lifang Wei
Expert Syst. Appl.3
2019 Recurrent Shape Regression
abstract
An end-to-end network architecture, the Recurrent Shape Regression (RSR), is presented to deal with the task of facial shape detection, a crucial step in many computer vision problems. The RSR generalizes the conventional cascaded regression into a recurrent dynamic network through abstracting common latent models with stage-to-stage operations. Instead of invariant regression transformation, we construct shape-dependent dynamic regressors to attain the recurrence of regression action itself. The regressors can be stacked into a high-order regression network to represent more complex shape regression. By further integrating feature learning as well as global shape constraint, the RSR becomes more controllable in entire optimization of shape regression, where the gradient computation can be efficiently back-propagated through time. To handle the possible partial occlusions of shapes, we propose a mimic virtual occlusion strategy by randomly disturbing certain point cliques without the requirement of any annotations of occlusion information or even occluded training data. Extensive experiments on five face datasets demonstrate that the proposed RSR outperforms the recent state-of-the-art cascaded approaches.
Zhen Cui 0001, Shengtao Xiao, Zhiheng Niu, Shuicheng Yan, Wenming Zheng
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Hierarchical Contextual Refinement Networks for Human Pose Estimation
abstract
Predicting human pose in the wild is a challenging problem due to high flexibility of joints and possible occlusion. Existing approaches generally tackle the difficulties either by holistic prediction or multi-stage processing, which suffer from poor performance for locating challenging joints or high computational cost. In this paper, we propose a new Hierarchical Contextual Refinement Network (HCRN) to robustly predict human poses in an efficient manner, where human body joints of different complexities are processed at different layers in a context hierarchy. Different from existing approaches, our proposed model predicts positions of joints from easy to difficult in a single stage through effectively exploiting informative contexts provided in the previous layer. Such approach offers two appealing advantages over state-of-the-arts: (1) more accurate than predicting all the joints together and (2) more efficient than multi-stage processing methods. We design a Contextual Refinement Unit (CRU) to implement the proposed model, which enables auto-diffusion of joint detection results to effectively transfer informative context from easy joints to difficult ones. In this way, difficult joints can be reliably detected even in presence of occlusion or severe distracting factors. Multiple CRUs are organized into a tree-structured hierarchy which is end-to-end trainable and does not require processing joints for multiple iterations. Comprehensive experiments evaluate the efficacy and efficiency of the proposed HCRN model to improve well-established baselines and achieve new state-of-the-art on multiple human pose estimation benchmarks.
Xuecheng Nie, Jiashi Feng, Junliang Xing, Shengtao Xiao, Shuicheng Yan
IEEE Trans. Image Process.4
2018 Deep Recurrent Regression for Facial Landmark Detection
abstract
We propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists of three parts. Through the first part, we encode an input face image to resolution-preserved deconvolutional feature maps via a deep network with stacked convolutional and deconvolutional layers. Then, in the second part, we estimate the initial coordinates of the facial key points by an additional convolutional layer on top of these deconvolutional feature maps. In the last part, by using the deconvolutional feature maps and the initial facial key points as input, we refine the coordinates of the facial key points by a recurrent network that consists of multiple long short-term memory components. Extensive evaluations on several benchmark data sets show that the proposed deep architecture has superior performance against the state-of-the-art methods.
Hanjiang Lai, Shengtao Xiao, Yan Pan 0002, Zhen Cui 0001, Jiashi Feng, Chunyan Xu, Jian Yin 0001, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.2
2017 Recurrent 3D-2D Dual Learning for Large-Pose Facial Landmark Detection
abstract
Despite remarkable progress of face analysis techniques, detecting landmarks on large-pose faces is still difficult due to self-occlusion, subtle landmark difference and incomplete information. To address these challenging issues, we introduce a novel recurrent 3D-2D dual learning model that alternatively performs 2D-based 3D face model refinement and 3D-to-2D projection based 2D landmark refinement to reliably reason about self-occluded landmarks, precisely capture the subtle landmark displacement and accurately detect landmarks even in presence of extremely large poses. The proposed model presents the first loop-closed learning framework that effectively exploits the informative feedback from the 3D-2D learning and its dual 2D-3D refinement tasks in a recurrent manner. Benefiting from these two mutual-boosting steps, our proposed model demonstrates appealing robustness to large poses (up to profile pose) and outstanding ability to capture fine-scale landmark displacement compared with existing 3D models. It achieves new state-of-the-art on the challenging AFLW benchmark. Moreover, our proposed model introduces a new architectural design that economically utilizes intermediate features and achieves 4× faster speed than its deep learning based counterparts.
Shengtao Xiao, Jiashi Feng, Luoqi Liu, Xuecheng Nie, Wei Wang 0108, Shuicheng Yan, Ashraf A. Kassim
ICCV1
2017 Integrated Face Analytics Networks through Cross-Dataset Hybrid Training
abstract
Face analytics benefits many multimedia applications. It consists of a number of tasks, such as facial emotion recognition and face parsing, and most existing approaches generally treat these tasks independently, which limits their deployment in real scenarios. In this paper we propose an integrated Face Analytics Network (iFAN), which is able to perform multiple tasks jointly for face analytics with a novel carefully designed network architecture to fully facilitate the informative interaction among different tasks. The proposed integrated network explicitly models the interactions between tasks so that the correlations between tasks can be fully exploited for performance boost. In addition, to solve the bottleneck of the absence of datasets with comprehensive training data for various tasks, we propose a novel cross-dataset hybrid training strategy. It allows "plug-in and play'' of multiple datasets annotated for different tasks without the requirement of a fully labeled common dataset for all the tasks. We experimentally show that the proposed iFAN achieves state-of-the-art performance on multiple face analytics tasks using a single integrated model. Specifically, iFAN achieves an overall F-score of 91.15% on the Helen dataset for face parsing, a normalized mean error of 5.81% on the MTFL dataset for facial landmark localization and an accuracy of 45.73% on the BNU dataset for emotion recognition with a single model.
Jianshu Li, Shengtao Xiao, Fang Zhao 0006, Jian Zhao 0006, Jianan Li 0001, Jiashi Feng, Shuicheng Yan, Terence Sim
ACM Multimedia2
2016 Recurrently Target-Attending Tracking
abstract
Robust visual tracking is a challenging task in computer vision. Due to the accumulation and propagation of estimation error, model drifting often occurs and degrades the tracking performance. To mitigate this problem, in this paper we propose a novel tracking method called Recurrently Target-attending Tracking (RTT). RTT attempts to identify and exploit those reliable parts which are beneficial for the overall tracking process. To bypass occlusion and discover reliable components, multi-directional Recurrent Neural Networks (RNNs) are employed in RTT to capture long-range contextual cues by traversing a candidate spatial region from multiple directions. The produced confidence maps from the RNNs are employed to adaptively regularize the learning of discriminative correlation filters by suppressing clutter background noises while making full use of the information from reliable parts. To solve the weighted correlation filters, we especially derive an efficient closedform solution with a sharp reduction in computation complexity. Extensive experiments demonstrate that our proposed RTT is more competitive over those correlation filter based methods.
Zhen Cui 0001, Shengtao Xiao, Jiashi Feng, Shuicheng Yan
CVPR2
2016 Robust Facial Landmark Detection via Recurrent Attentive-Refinement Networks
Shengtao Xiao, Jiashi Feng, Junliang Xing, Hanjiang Lai, Shuicheng Yan, Ashraf A. Kassim
ECCV (1)1
2016 A Live Face Swapper
abstract
In this technical demonstration, we propose a face swapping framework, which is able to interactively change the appearance of a face in the wild to a different person/creature's face in real time on a mobile device. To realize this objective, we develop a deep learning-based face detector which is able to accurately detect faces in the wild. Our face feature points tracking system based on progressive initialization ensures accurate and robust localization of facial landmarks under extreme poses and expressions in real time. Relying on the advances of our face detector and face feature points tracker, we construct the Face Swapper which can smoothly replace the face appearance of a user in real time.
Shengtao Xiao, Luoqi Liu, Xuecheng Nie, Jiashi Feng, Ashraf A. Kassim, Shuicheng Yan
ACM Multimedia1
2016 Constrained Multilegged Robot System Modeling and Fuzzy Control With Uncertain Kinematics and Dynamics Incorporating Foot Force Optimization
abstract
This paper studies the optimal distribution of feet forces and control of multilegged robots with uncertainties in both kinematics and dynamics. First, a constrained dynamics for multilegged robots and the constrained environment model are established by considering both kinematic and dynamic uncertainties. Under an external wrench for multilegged robots, the foot forces and moments of the supporting legs can be formulated as quadratic programming problems subject to linear and nonlinear constraints. The neurodynamics of recurrent neural network is developed for foot force optimization. For the obtained optimized tip-point force and the motion of legs, we propose a hybrid task-space trajectory and force tracking based on fuzzy system and adaptive mechanism that are used to compensate for the external perturbation, kinematics, and dynamics uncertainties. The tracking of task-space trajectory and constraint force is achieved under unknown dynamical parameters, constraints, and disturbances. Extensive simulations have been provided to verify the effectiveness of the proposed scheme.
Zhijun Li 0001, Shengtao Xiao, Shuzhi Sam Ge, Hang Su 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2014 Fuzzy approximation adaptive control of quadruped robots with kinematics and dynamics uncertainties
abstract
This paper investigates optimal feet forces distribution and control of quadruped robots with uncertainties in both kinematics and dynamics. First, a constrained dynamics of quadruped robots is established. The distribution of required forces and moments on the supporting legs of a quadruped robot can be formulated as a problem for minimizing an objective function subject to form-closure constraints and balance constraints of external force. The dynamics of recurrent neural network for realtime force optimization are proposed. For the obtained optimized tip-point force and the motion of legs, we propose the hybrid motion/force control based on adaptive fuzzy system to compensate for the external perturbation and the task-space tracking errors in the environment. The proposed control can confront the uncertainties including approximation task space error and external perturbation. The verification of the proposed control is conducted using the extensive simulations.
Zhijun Li 0001, Shengtao Xiao, Shuzhi Sam Ge
FUZZ-IEEE2