VLDB 2026 Research / reviewers in the wild / expert
Fei Jiang 0006
dblp:80/4567-6
· DBLP profile ↗
37ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0001-9495-9903ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIP2Pose: Frozen CLIP as Semantic Guide for Domain Adaptive Pose EstimationabstractUnsupervised domain adaptive pose estimation is a fundamental yet challenging task due to the need to transfer from labeled synthetic data to unlabeled real data. Nevertheless, the underlying pose semantics, which are governed by spatial structure, remain largely consistent across domains. This observation motivates the use of vision-language models, which provide domain-invariant representations that align well with high-level semantic concepts. Motivated by this, we propose CLIP2Pose, a novel framework that leverages the semantic robustness of frozen CLIP encoders to facilitate cross-domain generalization. We first introduce a semantic-driven prompt mechanism that encodes structural priors, domain-specific appearance, and instance-level context into the image representation. This guides the model to focus on semantically meaningful and structurally relevant features. Next, we propose a semantic modulation module that adaptively refines visual features by conditioning them on prompt-derived embeddings, enhancing alignment between semantics and visual patterns. To further bridge the modality and domain gaps, we design a directional alignment loss that encourages consistent structural reasoning across both vision and language representations. Extensive experiments on domain adaptive human body and hand pose benchmarks show that CLIP2Pose achieves state-of-the-art performance. Fei Jiang 0006, Dandan Zhu 0001, Jinxin Shi, Aimin Zhou |
AAAI | 2 |
| 2026 | URSD: A uncertainty-resistant semi-supervised approach for object detection
Fei Jiang 0006, Dandan Zhu 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Semi-Supervised Unconstrained Head Pose Estimation in the WildabstractExisting research on unconstrained in-the-wild head pose estimation suffers from the flaws of its datasets, which consist of either numerous samples by non-realistic synthesis or constrained collection, or small-scale natural images yet with plausible manual annotations. This makes fully-supervised solutions compromised due to the reliance on generous labels. To alleviate it, we propose the first semi-supervised unconstrained head pose estimation method SemiUHPE, which can leverage abundant easily available unlabeled head images. Technically, we choose semi-supervised rotation regression and adapt it to the error-sensitive and label-scarce problem of unconstrained head pose. Our method is based on the observation that the aspect-ratio invariant cropping of wild heads is superior to previous landmark-based affine alignment given that landmarks of unconstrained human heads are usually unavailable, especially for underexplored non-frontal heads. Instead of using a pre-fixed threshold to filter out pseudo labeled heads, we propose dynamic entropy based filtering to adaptively remove unlabeled outliers as training progresses by updating the threshold in multiple stages. We then revisit the design of weak-strong augmentations and improve it by devising two novel head-oriented strong augmentations, termed pose-irrelevant cut-occlusion and pose-altering rotation consistency respectively. Extensive experiments and ablation studies show that SemiUHPE outperforms its counterparts greatly on public benchmarks under both the front-range and full-range settings. Furthermore, our proposed method is also beneficial for solving other closely related problems, including generic object rotation regression and 3D head reconstruction, demonstrating good versatility and extensibility. Huayi Zhou 0001, Fei Jiang 0006, Yong Rui, Hongtao Lu 0001, Kui Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | MEScan360: A Memory-Enhanced Scanpath Prediction Model for Omnidirectional ImagesabstractScanpath prediction for omnidirectional images (ODIs) aims to capture the dynamic human visual attention. However, the complicated gaze behavior and inevitable projection distortion make scanpath prediction in ODIs extremely challenging. Most existing models neither capture the long-term dependencies across visual states nor fully incorporate historical memory information, leading to limited performance. To this end, we propose MEScan360, a memory-enhanced scanpath prediction model for ODIs. We introduce two key innovations: long-term memory storage unit and memory interaction module. These two components establish a more explicit link between past visual information and current visual inputs, thereby significantly enhancing the performance of scanpath prediction. Furthermore, a robust feature extraction module is designed to extract semantic feature precisely from distorted ODIs with a more lightweight structure. Extensive experiments on several benchmark datasets demonstrate that our proposed model achieves competitive performance in both accuracy and efficiency. Dandan Zhu 0001, Kaiwei Zhang, Fei Jiang 0006, Guangtao Zhai |
ICME | 4 |
| 2025 | HandNet: Occlusion-robust 3D hand mesh reconstruction with prior information
Fei Jiang 0006, Dandan Zhu 0001, Aimin Zhou |
Knowl. Based Syst. | 2 |
| 2025 | OSA-PD: Open-Set Active Learning Exploiting Predictive DiscrepancyabstractActive learning with closed-set annotation has achieved significant success. However, real-world data often comprises numerous unknown classes irrelevant to the task which often confuses the query strategies to select unknown class data for annotation. To address such challenge in the open-set annotation (OSA), we propose a novel open-set active learning framework based on predictive discrepancy, named OSA-PD, with consideration that valuable data for model improvement is often with high predictive discrepancy. Specifically, two algorithms based on different discrepancy measurements are presented under OSA-PD framework, i.e., OSA-PRD with predictive results discrepancy, and OSA-PDD with decoupled predictive distribution discrepancy. Experimental results on CIFAR100 and TinyImageNet demonstrate the proposed OSA-PD can effectively select known class data and achieve higher classification accuracy with the same amount of annotated sample in comparison with existing active learning algorithms. Fei Jiang 0006, Jiaxin Si, Lili Xiong |
IEEE Signal Process. Lett. | 3 |
| 2024 | BPJDet: Extended Object Representation for Generic Body-Part Joint DetectionabstractDetection of human body and its parts has been intensively studied. However, most of CNNs-based detectors are trained independently, making it difficult to associate detected parts with body. In this paper, we focus on the joint detection of human body and its parts. Specifically, we propose a novel extended object representation integrating center-offsets of body parts, and construct an end-to-end generic Body-Part Joint Detector (BPJDet). In this way, body-part associations are neatly embedded in a unified representation containing both semantic and geometric contents. Therefore, we can optimize multi-loss to tackle multi-tasks synergistically. Moreover, this representation is suitable for anchor-based and anchor-free detectors. BPJDet does not suffer from error-prone post matching, and keeps a better trade-off between speed and accuracy. Furthermore, BPJDet can be generalized to detect body-part or body-parts of either human or quadruped animals. To verify the superiority of BPJDet, we conduct experiments on datasets of body-part (CityPersons, CrowdHuman and BodyHands) and body-parts (COCOHumanParts and Animals5C). While keeping high detection accuracy, BPJDet achieves state-of-the-art association performance on all datasets. Besides, we show benefits of advanced body-part association capability by improving performance of two representative downstream applications: accurate crowd head detection and hand contact estimation. Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Yue Ding 0001, Hongtao Lu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Stuart: Individualized Classroom Observation of Students with Automatic Behavior Recognition And TrackingabstractEach student matters, but it is hardly for instructors to observe all the students during the courses and provide helps to the needed ones immediately. In this paper, we present StuArt, a novel automatic system designed for the individualized classroom observation, which empowers instructors to concern the learning status of each student. StuArt can recognize five representative student behaviors (hand-raising, standing, sleeping, yawning, and smiling) that are highly related to the engagement and track their variation trends during the course. To protect the privacy of students, all the variation trends are indexed by the seat numbers without any personal identification information. Furthermore, StuArt adopts various user-friendly visualization designs to help instructors quickly understand the individual and whole learning status. Experimental results on real classroom videos have demonstrated the superiority and robustness of the embedded algorithms. We expect our system promoting the development of large-scale individualized guidance of students. More information is in https://github.com/hnuzhy/StuArt. Huayi Zhou 0001, Fei Jiang 0006, Jiaxin Si, Lili Xiong, Hongtao Lu 0001 |
ICASSP | 2 |
| 2023 | Body-Part Joint Detection and Association via Extended Object RepresentationabstractThe detection of human body and its related parts (e.g., face, head or hands) have been intensively studied and greatly improved since the breakthrough of deep CNNs. However, most of these detectors are trained independently, making it a challenging task to associate detected body parts with people. This paper focuses on the problem of joint detection of human body and its corresponding parts. Specifically, we propose a novel extended object representation that integrates the center location offsets of body or its parts, and construct a dense single-stage anchor-based Body-Part Joint Detector (BPJDet). Body-part associations in BPJDet are embedded into the unified representation which contains both the semantic and geometric information. Therefore, BPJDet does not suffer from error-prone association post-matching, and has a better accuracy-speed trade-off. Furthermore, BPJDet can be seamlessly generalized to jointly detect any body part. To verify the effectiveness and superiority of our method, we conduct extensive experiments on the CityPersons, CrowdHuman and BodyHands datasets. The proposed BPJDet detector achieves state-of-the-art association performance on these three benchmarks while maintains high accuracy of detection. Code is released in https://github.com/hnuzhy/BPJDet. Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001 |
ICME | 2 |
| 2023 | Landmark-Assisted Facial Action Unit Detection with Optimal Attention and Contrastive Learning
Qiaoping Hu, Hongtao Lu 0001, Fei Jiang 0006, Yaoyi Li |
ICONIP (12) | 4 |
| 2023 | Multi-scale Local Region-Based Facial Action Unit Detection with Graph Convolutional Network
Zhenchang Zhang, Hongtao Lu 0001, Fei Jiang 0006 |
ICONIP (12) | 4 |
| 2023 | SSDA-YOLO: Semi-supervised domain adaptive YOLO for cross-domain object detection
Huayi Zhou 0001, Fei Jiang 0006, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2022 | RFAU: A Database for Facial Action Unit Analysis in Real ClassroomsabstractEmotion analysis of students plays an important role in teaching effect evaluation. To develop robust algorithms for emotion analysis of students, a database from real classrooms is required. However, most existing databases were collected from adults and constructed in laboratory settings. In this article, we present a manually-annotated facial action unit database from juveniles in real classrooms. Our database has three main characteristics: (1) it provides numerous education-related action units data from primary and high schools, complementing the vacancy of the publicly available educational action unit databases; (2) it contains 256,220 manually-annotated facial images of 1,796 juveniles, frame-by-frame annotated with 12 action units and 6-level intensities for each action unit; (3) it covers many challenges in the wild, including various head poses, low facial resolution, illuminations, and occlusions, supplementing action unit databases in the wild for research. The baselines for action unit detection and action unit intensity estimation are provided for future references. Especially, we apply the weighted balance loss to solve imbalanceswithinandbetweenlabels. Our database will be available to the research community:http://www.dlc.sjtu.edu.cn/rfau. Qiaoping Hu, Chuanneng Mei, Fei Jiang 0006, Ruimin Shen |
IEEE Trans. Affect. Comput. | 3 |
| 2021 | Laplacian Regularized Tensor Low-Rank Minimization for Hyperspectral Snapshot Compressive ImagingabstractSnapshot Compressive Imaging (SCI) systems, including hyperspectral compressive imaging and video compressive imaging, are designed to depict high-dimensional signals with limited data by mapping multiple images into one. One key module of SCI systems is a high quality reconstruction algorithm for original frames. However, most existing decoding algorithms are based on vectorization representation and fail to capture the intrinsic structural information of high dimensional signals. In this paper, we propose a tensor-based low-rank reconstruction algorithm with hyper-Laplacian constraint for hyperspectral SCI systems. First, we integrate the non-local self-similarity and tensor low-rank minimization approach to explore the intrinsic structural correlations along spatial and spectral domains. Then, we introduce a hyper-Laplacian constraint to model the global spectral structures, alleviating the ringing artifacts in the spatial domain. Experimental results on hyperspectral image corpus demonstrate the proposed algorithm achieves average 0.8~2.9 dB improvement in PSNR over state-of-the-art work. Fei Jiang 0006, Hongtao Lu 0001 |
ICASSP | 2 |
| 2021 | ShallowNet: An Efficient Lightweight Text Detection Network Based on Instance Count-Aware Supervision Information
Xingfei Hu, Deyang Wu, Fei Jiang 0006, Hongtao Lu 0001 |
ICONIP (1) | 4 |
| 2021 | Small and accurate heatmap-based face alignment via distillation strategy and cascaded architecture
Jiaxin Si, Fei Jiang 0006, Ruimin Shen, Hongtao Lu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2020 | Intelligent Student Behavior Analysis System for Real ClassroomsabstractIn this paper, we design an intelligent student behavior analysis system for recorded classrooms, which automatically detects hand-raising, standing, and sleeping behaviors of students. Detecting these behaviors is quite challenging mainly due to various scale behaviors, low resolution, and imbalanced behavior samples. To overcome the above-mentioned challenges, we first build a large-scale student behavior corpus from thirty schools, labeling these behaviors using bounding boxes frame-by-frame, which changes the behavior recognition problem into object detections. Then, we propose an improved Faster R-CNN, a classical object detection model, for student behavior analysis. Specifically, we first present a novel scale-aware detection head to overcome scale variations. Secondly, we propose a new feature fusion strategy to detect low-resolution behaviors while introduces little computation overhead. Thirdly, we utilize OHEM (Online Hard Example Mining) to alleviate severe class imbalances. Experiment results on our real corpus are increased by 3.4% mAP while maintaining a fast speed. Fei Jiang 0006, Ruimin Shen |
ICASSP | 2 |
| 2020 | Fast and Accurate Hand-Raising Gesture Detection in Classroom
Tao Liu 0031, Fei Jiang 0006, Ruimin Shen |
ICONIP (4) | 2 |
| 2020 | Facial Action Units Intensity Estimation via Graph Relation Network
Fei Jiang 0006, Ruimin Shen |
ICONIP (2) | 2 |
| 2020 | GestureDet: Real-time Student Gesture Analysis with Multi-dimensional Attention-based DetectorabstractStudents’ gestures, hand-raising, stand-up, and sleeping, indicates the engagement of students in classrooms and partially reflects teaching quality. Therefore, fast and automatically recognizing these gestures are of great importance. Due to limited computational resources in primary and secondary schools, we propose a real-time student behavior detector based on light-weight MobileNetV2-SSD to reduce the dependency of GPUs. Firstly, we build a large-scale corpus from real schools to capture various behavior gestures. Based on such a corpus, we transfer the gesture recognition task into object detections. Secondly, we design a multi-dimensional attention-based detector, named GestureDet, for real-time and accurate gesture analysis. The multi-dimensional attention mechanisms simultaneously consider all the dimensions of the training set, aiming to pay more attention to discriminative features and samples that are important for the final performance. Specifically, the spatial attention is constructed with stacked dilated convolution layers to generate a soft and learnable mask for re-weighting foreground and background features; the channel attention introduces the context modeling and squeeze-and-excitation module to focus on discriminative features; the batch attention discriminates important samples with a new designed reweight strategy. Experimental results demonstrate the effectiveness and versatility of GestureDet, which achieves 75.2% mAP on real student behavior dataset, and 74.5% on public PASCAL VOC dataset at 20fps on embedding device Nvidia Jetson TX2. Code will be made publicly available. Fei Jiang 0006, Ruimin Shen |
IJCAI | 2 |
| 2019 | Sleep Gesture Detection in Classroom Monitor SystemabstractThis paper proposes a novel method to detect sleep persons in real classroom scenes, which is useful for detecting the attention of students. There are several challenges for sleep gesture detection, including occlusion, various sleep gestures, interference terms with similar features like writing, and small sleep gesture targets. To solve these challenges, we first build a sleep gesture dataset from hundreds of real classes among schools in Shanghai. Second, to detect sleep gestures we propose a modified R-FCN integrated with feature pyramid and deformable convolution. Moreover, we design an efficient local multiscale testing algorithm to address small sleep gesture detection. Experiments on our sleep dataset have shown that our approach significantly outperforms the basic R-FCN and reaches 0.74 [email protected]. Especially, for small-size sleep gestures, our method gets an impressive improvement at 90% on our hard-samples dataset with just little additional time consumption. Besides, experiments on a large number of real video streams have proven our algorithm is applicable in real classroom scenes. Fei Jiang 0006, Ruimin Shen |
ICASSP | 2 |
| 2019 | Large-pose Face Alignment via Shape-aware HeatmapabstractIn this paper, we focus on dealing with problems of large-pose face alignment. Recently proposed heatmap-based algorithms have made promising performance on this problem. However, the traditional heatmap is constructed based on Gaussian model with fixed variance, which is inconsistent with the local shape of faces. In this paper, we propose a shape-aware heatmap to efficiently solve the problems of large-pose face alignment. Specifically, we design a novel heatmap based on Gaussian mixture model, where positions of several adjacent landmarks are utilized to construct different components. Thus the probability distribution is modified to fit the shape of the local region. The experimental results on Menpo-3D and AFLW2000-3D databases show that the proposed method outperforms the state-of-the-art algorithms. Jiaxin Si, Fei Jiang 0006, Ruimin Shen |
ICASSP | 2 |
| 2019 | Fast Sparse Coding Inference with Historical InformationabstractRecently, time-unfolded recurrent neural network (RNN) based algorithms are successfully applied for fast sparse coding (SC) inference, such as LISTA and SLSTM. However, these methods do not properly exploit the historical information which is proved to speed up the convergence. In this paper, we propose a novel RNN-based SC inference framework with attention mechanism to efficiently incorporate the related historical information. The proposed framework consists of an attention network and a time-unfolded RNN, where the RNN generates the historical information and the attention network automatically determines the importance of these historical values for the current updating. The final sparse code is obtained by passing the context vectors generated from the attention network to a soft shrinkage function. Extensive experiments on convergence analysis and image classification tasks demonstrate that our network achieves impressive improvements on SC inference in terms of the quality of estimated sparse codes and the inference time. Moreover, the proposed network can be easily extended into a supervised version to further improve the classification accuracy. Zhenchang Zhang, Fei Jiang 0006, Ruimin Shen |
ICDM | 2 |
| 2019 | An Effective Yawn Behavior Detection Method in Classroom
Zexian Wang, Fei Jiang 0006, Ruimin Shen |
ICONIP (1) | 2 |
| 2019 | Hand-raising gesture detection in real classrooms using improved R-FCN
Jiaxin Si, Jiaojiao Lin, Fei Jiang 0006, Ruimin Shen |
Neurocomputing | 3 |
| 2018 | Efficient Multi-Dimensional Tensor Sparse Coding Using t-Linear CombinationabstractIn this paper, we propose two novel multi-dimensional tensor sparse coding (MDTSC) schemes using the t-linear combination. Based on the t-linear combination, the shifted versions of the bases are used for the data approximation, but without need to store them. Therefore, the dictionaries of the proposed schemes are more concise and the coefficients have richer physical explanations. Moreover, we propose an efficient alternating minimization algorithm, including the tensor coefficient learning and the tensor dictionary learning, to solve the proposed problems. For the tensor coefficient learning, we design a tensor-based fast iterative shrinkage algorithm. For the tensor dictionary learning, we first divide the problem into several nearly-independent subproblems in the frequency domain, and then utilize the Lagrange dual to further reduce the number of optimization variables. Experimental results on multi-dimensional signals denoising and reconstruction (3DTSC, 4DTSC, 5DTSC) show that the proposed algorithms are more efficient and outperform the state-of-the-art tensor-based sparse coding models. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
AAAI | 1 |
| 2018 | Who Are Raising Their Hands? Hand-Raiser Seeking Based on Object Detection and Pose EstimationabstractIn this paper, we propose an automatic hand-raiser recognition algorithm to show who raise their hands in real classroom scenarios, which is of great importance for further analyzing the learning states of individuals. To recognize the hand-raisers, we divide the hand-raiser recognition into three subproblems, including hand-raising detection, pose estimation, and matching the raised hands to students. Several challenges exist while dealing with the above-mentioned subproblems, such as low resolution of the back row for keypoints detection, the motion distortion caused by hand raising in pose estimation, and various complex situations for matching. To solve these challenges, we first adopt an improved R-FCN algorithm for hand-raising detection, whose effectiveness has been demonstrated. Secondly, we present a novel PAF-based pose estimation algorithm for detecting keypoints of human bodies. The proposed PAF adds scale search and modified weight metric to adapt to the real and complex scenarios. Specifically, scale search improves the detection effect at low resolution by pooling human characteristics in different sizes of pictures, and modified weight metric reasonably utilizes the directional vectors of possible limb connections to optimize the case of motion distortion. Thirdly, a heuristic matching strategy based on the location of hand-raising and keypoints information is proposed to recognize the hand-raisers. Experimental results on six teaching videos in real classrooms have demonstrated the efficiency of the proposed algorithm, and 83% recognition accuracy indicates the potential applications in real classrooms. Huayi Zhou 0001, Fei Jiang 0006, Ruimin Shen |
ACML | 2 |
| 2018 | Anisotropic Total Variation Regularized Low-Rank Tensor Completion Based On Tensor Nuclear Norm for Color Image InpaintingabstractIn this paper, we propose a novel low-rank tensor completion (LRTC) model under the circulant algebra for color image inpainting, which simultaneously preserves the low-rank structures of images, and also explore the local smooth and piecewise priors of the images in the spatial domain. First, color images are naturally represented by 3-order tensors which preserve the intrinsic structures of color images. Second, we preserve the low-rank structures of these tensors with tensor nuclear norm, which can simultaneously exploit the correlations among the spatial and channel domains. Third, we integrate an anisotropic total variation into our low-rank tensor completion model, which preserve the local smooth and piecewise priors of color images. Then, an efficient alternating direction method of multipliers (ADMM) is proposed to solve the resulting optimization problem. Experimental results on eight widely used color images demonstrate the effectiveness and superiority of the proposed algorithm. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
ICASSP | 1 |
| 2018 | Hand-Raising Gesture Detection in Real ClassroomabstractThis paper proposes a novel method for hand-raising detection in the real classroom environment. Different from traditional motion detection, the hand-raising detection is quite challenging in the real classroom due to complex scenarios, various gestures, and low resolutions. To solve these challenges, we first build up a large-scale hand-raising data set from thirty primary schools and middle schools of Shanghai, China. Then we propose an improved R-FCN to solve the above-mentioned challenges. Specifically, we first design an automatic detection templates algorithm for various gestures of hand-raising detection. Second, for better detection of the small-size hands, we present a feature pyramid to simultaneously capture the detail and highly semantic features. Incorporating these two strategies into a basic R-FCN architecture, our model achieves impressive results on real classroom scenarios. After a wide test, the accuracy of the hand-raising detection achieves 85% on average, which can satisfy the real application. Jiaojiao Lin, Fei Jiang 0006, Ruimin Shen |
ICASSP | 2 |
| 2018 | Total Variation Regularized Reweighted Low-rank Tensor Completion for Color Image InpaintingabstractRecent low-rank based tensor completion (LRTC) algorithms have been successfully applied into color image inpainting. However, most of existing LRTC algorithms treat each dimension of tensors equally, which ignores the differences of the intrinsic structure correlations among dimensions. In this paper, we make a detailed analysis about the rank properties of each dimension and design a simple yet effective reweighted low-rank tensor completion model that truthfully capture the intrinsic structure correlations with reduced computational burden. Moreover, to capture the local smooth and piecewise priors of tensors, we integrate total variation into our model. Considering two formulations of LRTC, tensor unfolding and tensor decomposition, we propose corresponding two algorithms for color image recovery. Extensive experimental results on color image recovery show the efficiency and effectiveness of the proposed two algorithms against state-of-the-art competitors. Lingwei Li, Fei Jiang 0006, Ruimin Shen |
ICIP | 2 |
| 2018 | CCT: A Cross-Concat and Temporal Neural Network for Multi-Label Action Unit DetectionabstractAction Unit (AU) detection is essential for facial expression analysis. However, most existing AU detection algorithms only focus on physical features, e.g., temporal feature and AU correlations, without considering various distributions of AUs, i.e., some AUs are quite less than others. In this work, we propose a novel cross-concat and temporal (CCT) neural network, which simultaneously consider physical features and the distribution differences. First, we design a cross-concat block (CCB) to adapt to the various distributions of AUs. CCB is based on the idea of skip connections since skip connections can reuse features from different layers and capture abundant features of AUs even with relatively small-size training samples. Second, LSTM layers are utilized to capture the temporal dependencies and multi-label learning is utilized for capturing AU correlations. Experimental results on three popular AU detection datasets, BP4D, DISFA, and GFT, show that the proposed algorithm outperforms the state-of-the-art ones. Qiaoping Hu, Fei Jiang 0006, Chuanneng Mei, Ruimin Shen |
ICME | 2 |
| 2018 | Region and Temporal Dependency Fusion for Multi-label Action Unit DetectionabstractAutomatic Facial Action Unit (AU) detection from videos increases numerous interests over the past years due to its importance for analyzing facial expressions. Many proposed methods face challenges in detecting sparse face regions for different AUs, in the fusion of temporal dependency, and in learning multiple AUs simultaneously. In this paper, we propose a novel deep neural network architecture for AU detection to model above-mentioned challenges jointly. Firstly, to capture the region sparsity, we design a region pooling layer after a fully convolutional network to extract per-region features for each AU. Secondly, in order to integrate temporal dependency, Long Short Term Memory (LSTM) is stacked on the top of regional features. Finally, the regional features and outputs of LSTMs are utilized together to produce per-frame multi-label predictions. Experimental results on three large spontaneous AU datasets, BP4D, GFT and DISFA, have demonstrated our work outperforms state-of-the-art methods. On three datasets, our work has highest average F1 and AUC scores with an average F1 score improvement of 4.8% on BP4D, 12.7% on GFT and 14.3% on DISFA, and an average AUC score improvement of 27.4% on BP4D and 33.5% on DISFA. Chuanneng Mei, Fei Jiang 0006, Ruimin Shen, Qiaoping Hu |
ICPR | 2 |
| 2018 | Multi-object Detection Based on Deep Learning in Real Classrooms
Benchi Shao, Fei Jiang 0006, Ruimin Shen |
PRICAI | 2 |
| 2017 | Graph regularized tensor sparse coding for image representationabstractSparse coding (SC) is a unsupervised learning scheme that has received an increasing amount of interests in recent years. However, conventional SC vectorizes the input images, which destructs the intrinsic spatial structures of the images. In this paper, we propose a novel graph regularized tensor sparse coding (GTSC) for image representation. GTSC preserves the local proximity of elementary structures in the image by adopting the newly proposed tubal-tensor representation. Simultaneously, it considers the intrinsic geometric properties by imposing graph regularization that has been successfully applied to uncover the geometric distribution for the image data. Moreover, the learned sparse representations by GTSC have better physical explanations as the key operation (i.e., circular convolution) in the tubal-tensor model preserves the shifting invariance property. Experimental results on image clustering demonstrate the effectiveness of the proposed scheme. Fei Jiang 0006, Xiao-Yang Liu, Hongtao Lu 0001, Ruimin Shen |
ICME | 1 |
| 2017 | Highly Occluded Face Detection: An Improved R-FCN Approach
Fei Jiang 0006, Ruimin Shen |
ICONIP (6) | 2 |
| 2017 | Region-Based Face Alignment with Convolution Neural Network Cascade
Fei Jiang 0006, Ruimin Shen |
ICONIP (3) | 2 |
| 2017 | Abdominal adipose tissues extraction using multi-scale deep neural network
Fei Jiang 0006, Huating Li, Xuhong Hou, Bin Sheng 0001, Ruimin Shen, Xiao-Yang Liu, Weiping Jia, Ping Li 0016, Ruogu Fang |
Neurocomputing | 1 |