Jingying Chen 0001

dblp:64/1845-1 · DBLP profile ↗
← Back
56ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0002-1523-3478ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 11 since 2021Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GD-Head: Reconstructing High-quality 3D Avatar Head from a Single Image Using Geometry-guided Diffusion Models
abstract
Existing single-image 3D head reconstruction methods often fall short of the demanding requirements for high-quality digital avatars in immersive virtual reality applications. To bridge this gap, we propose GD-Head—a novel diffusion-based framework for high-fidelity and detail-preserving 3D head texture reconstruction from a single image. GD-Head operates in three coherent stages: first, it employs an advanced geometric model to apply weak perspective projection to the input image, through which realistic facial texture details are directly extracted and occluded regions are identified and masked in gray for subsequent inpainting. Second, the 3D head model is parameterized into 2D UV space via cylindrical UV mapping. Third, a geometry-guided diffusion model is utilized to realistically inpaint the missing texture regions. By reframing 3D texture reconstruction as a 2D image inpainting task, GD-Head consistently produces complete and seamless texture maps that accurately reflect the subject’s facial expressions and surface details, essential for believable avatar representation in VR environments. Extensive experiments on six benchmark datasets and an additional in-the-wild collection demonstrate that GD-Head achieves state-of-the-art performance in reconstruction fidelity and robustness. To support reproducibility and further research, both the code and pre-trained models will be made publicly available.
Leyuan Liu 0001, Yufei Qian, Jingying Chen 0001
ICMR4
2026 Reliable decision making on clinical EEG: Trusted multi-view learning with subjective logic for uncertainty quantification
Yiping Zuo, Dan Chen 0001, Tengfei Gao, Jingying Chen 0001
Expert Syst. Appl.5
2026 A gan inversion-based multimodal framework for micro-expression recognition
Pianpian Ma, Jingying Chen 0001, Yanfeng Ji, Xiaodi Liu
Image Vis. Comput.2
2026 Robust facial expression recognition by simultaneously addressing hard and mislabeled samples
Yuandong Min, Ruyi Xu, Jingying Chen 0001, Yanfeng Ji, Xiaodi Liu
Pattern Recognit.3
2026 Rethinking the ambiguity in Facial Expression Recognition
Yuandong Min, Ruyi Xu, Shutong Wang, Zhiyi Yang, Jingying Chen 0001
Pattern Recognit.5
2026 Attention-Emotion Assessment of ASD Children via Representation Learning Based on Cross-Modal Disentanglement and Attention Alignment
abstract
Comprehensive cognitive-affective profiling for precision ASD assessment is hindered by inadequate computational modeling of emotion-behavior dynamics. This gap persists due to limited methods for reconciling EEG, eye-tracking, and facial expression modalities-critical for capturing ASD heterogeneity drivers. This study proposes the Cognitive-Affective Ability Assessment (CA3) framework, focusing on attention control in conjunction with valence-arousal, to address these limitations through: (1)Ability-Specific Neural Disentanglementextracting attention- and emotion-specific representations from EEG using ET and facial expression recordings as anchors; (2)Symmetrical Cross-Ability Alignmentmodeling contextual dependencies via a symmetric cross-attention mechanism; and (3)Uncertainty-Aware Ability Integrationfusing per-ability predictions using Dirichlet modeling and Dempster-Shafer theory while quantifying confidence. Extensive experiments have been conducted to evaluate CA$^{3}$vs. state-of-the-arts counterparts on the multimodel datasets (EEG, ET and facial expression recordings) with CCNU (27 ASD vs. 30 typically developing (TD) children) and BNU (49 ASD vs. 48 TD), and the results demonstrate that: 1) In the ASD cognitive-affective abilities assessment task, accuracy is improved by 2.7% for attention control and 2.5% for emotion perception, and by 3.5% for cognitive-affective abilities assessment, compared with current state-of-the-art methods. 2) For attention control and valence-arousal, composite ability gaps are measured between ASD and typically developing children: 15.6% and 19.9% for mild ASD, 28.2% and 39.1% for moderate ASD, and 44.4% and 61.3% for severe ASD, respectively. Overall, the framework effectively bridges computational assessment with clinically actionable insights, enabling robust decision-making through personalized cognitive-affective ASD profiles.
Zhuo Zhou, Dan Chen 0001, Zhiyi Yang, Tengfei Gao, Jingying Chen 0001
IEEE Trans. Affect. Comput.6
2026 Enhancing facial action unit intensity estimation with ordinal regression-enhanced transformer
Ruyi Xu, Shiyuan Su, Chenglin Xie, Jingying Chen 0001
Vis. Comput.4
2025 Biomarker Discovery for ASD via HMM-Based EEG Microstate Analysis
abstract
Discovering biomarkers for Autism Spectrum Disorder (ASD) is essential for elucidating its etiology, enabling early diagnosis, and refining treatment strategies. Electroencephalogram (EEG) microstates reflect the brain's overall dynamic changes, aiding in exploring differences in brain function patterns between ASD and Typically Developing (TD) groups. To this end, this study proposes an adaptive EEG microstate analysis approach based on Hidden Markov Models (HMMs) for the discovery of ASD biomarkers. Specifically, the proposed method, within the HMM framework, adaptively extracts millisecondscale transient brain microstate patterns that recur over time and models microstates using a multivariate Gaussian distribution rather than static topological structures. Resting-state EEG from 178 children aged 3 to 6 are used to validate the proposed approach. The analysis of the four microstates (#1, #2, #3, and #4) reveals significant differences between ASD and TD groups. Temporally, the ASD group shows difficulty in microstate transitions, primarily between microstates #2 and #3. In the frequency and spatial domains, TD individuals exhibit stronger brain region activation and interaction in microstates #2 and #4, whereas the ASD group shows reduced activity. Notably, during microstate #3, the ASD group demonstrates higher spectral power and channel coherence. Additionally, Ttests on intergroup feature differences and the results of the ASD discrimination task (accuracy: 88.89%) further confirm the potential of microstate features in assessing ASD tendencies. Overall, this study holds the potential to reveal novel insights into the neural mechanisms underlying ASD and identify valuable biomarkers for clinical assessment and diagnosis.
Dan Chen 0001, Meiqi Zhou, Tengfei Gao, Jingying Chen 0001, Naiqian Mao, Leyuan Liu 0001
BIBM4
2025 ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single Image
abstract
With 3D data rapidly emerging as an important form of multimedia information, 3D human mesh recovery technology has also advanced accordingly. However, current methods mainly focus on handling humans wearing tight clothing and perform poorly when estimating body shapes and poses under diverse clothing, especially loose garments. To this end, we make two key insights: (1) tailoring clothing to fit the human body can mitigate the adverse impact of clothing on 3D human mesh recovery, and (2) utilizing human visual information from large foundational models can enhance the generalization ability of the estimation. Based on these insights, we propose ClothHMR, to accurately recover 3D meshes of humans in diverse clothing. ClothHMR primarily consists of two modules: clothing tailoring (CT) and FHVM-based mesh recovering (MR). The CT module employs body semantic estimation and body edge prediction to tailor the clothing, ensuring it fits the body silhouette. The MR module optimizes the initial parameters of the 3D human mesh by continuously aligning the intermediate representations of the 3D mesh with those inferred from the foundational human visual model (FHVM). ClothHMR can accurately recover 3D meshes of humans wearing diverse clothing, precisely estimating their body shapes and poses. Experimental results demonstrate that ClothHMR significantly outperforms existing state-of-the-art methods across benchmark datasets and in-the-wild images. Additionally, a web application for online fashion and shopping powered by ClothHMR is developed, illustrating that ClothHMR can effectively serve real-world usage scenarios. The code and model for ClothHMR are available at: https://github.com/starVisionTeam/ClothHMR.
Yunqi Gao, Leyuan Liu 0001, Yuhan Li 0009, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001
ICMR6
2025 HumanPrinter: Reconstructing 3D Human from a Single Image Like a 3D Printer
abstract
High-fidelity 3D human reconstruction is essential for numerous applications. Existing reconstruction methods still suffer from several limitations. Implicit-function-based methods often produce artifacts, particularly when handling complex poses and loose-fitting clothing. Existing deformation-based methods require the entire body mesh to be input to the network for deformation, resulting in reconstructed results that are not ideal in detail. We propose HumanPrinter, a novel method for reconstructing high-fidelity 3D clothed human models from a single RGB image. Drawing inspiration from 3D printing, HumanPrinter reconstructs the human mesh layer by layer. HumanPrinter slices the coarse mesh resulting from deformation based on the estimated SMPL-X mesh into multiple vertically stacked polygons. The network then regresses vertex offsets from visual cues extracted from the input image to deform these polygons. Finally, these deformed polygons are stitched together and further refined to achieve a complete and detailed 3D human mesh. Polygon-based deformations associate the deformations of each vertex with its adjacent vertex so that the HumanPrinter produces fewer artifacts. By reducing the number of input polygons and increasing the number of deformable vertices, the layered reconstruction method can make the network more focused on local details. Through experiments on three datasets and visual results on in-the-wild data, we demonstrate that HumanPrinter performs competitive reconstruction quality compared to current state-of-the-art methods.
Leyuan Liu 0001, Shen Chen 0005, Jingying Chen 0001
ACM Multimedia3
2025 Self-training EEG discrimination model with weakly supervised sample construction: An age-based perspective on ASD evaluation
Tengfei Gao, Dan Chen 0001, Meiqi Zhou, Yiping Zuo, Weiping Tu, Xiaoli Li 0002, Jingying Chen 0001
Neural Networks8
2025 Dual-stream network with coordinate attention for multi-view micro-expression recognition using 3D face reconstruction
Pianpian Ma, Jingying Chen 0001, Xiaodi Liu
Pattern Anal. Appl.2
2025 Functional Connectivity Analysis of Children With Autism Under Emotional Clips
abstract
Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder with marked impairments in neural system functioning. Electroencephalography (EEG) offers a promising approach to investigate the neurophysiological basis of ASD, however, most EEG studies in ASD focus on spontaneous brain activity. Emotional processing deficits are a core feature of ASD, but related connectivity patterns remain underexplored due to challenges in data collection and analysis. This study investigates functional brain connectivity differences between children with ASD (n = 32) and typically developing (TD) children (n = 32) across five frequency bands and four connectivity indices. We designed an SVM-MRMR pipeline to classify ASD and TD children using these features. Our findings reveal that ASD children exhibit more coordinated intra-brain networks and oscillatory patterns in the high-frequency range. Additionally, they show an increased number of long-range connections in the Theta band, particularly between the left and right hemispheres. ASD children also demonstrate increased frontal lobe connectivity during positive emotions and heightened temporal lobe activity during negative emotions. Functional connectivity under positive and negative emotional clips achieved a classification accuracy exceeding 85%. These findings suggest that functional connectivity derived from portable EEG devices may serve as a potential biomarker for diagnosing and classifying ASD in real-world applications.
Huicong Kang, Jingying Chen 0001, Wei Wu 0022
IEEE Trans. Affect. Comput.6
2025 Hugging Rain Man: A Novel Facial Action Units Dataset for Analyzing Atypical Facial Expressions in Children With Autism Spectrum Disorder
abstract
Children with Autism Spectrum Disorder (ASD) often exhibit atypical facial expressions. However, the specific objective facial features that underlie this subjective perception remain unclear. In this paper, we introduce a novel dataset, Hugging Rain Man (HRM), which includes facial action units (AUs) manually annotated by FACS experts for both children with ASD and typically developing (TD) children. The dataset comprises a rich collection of posed and spontaneous facial expressions, totaling approximately 130,000 frames, along with 22 AUs, 10 Action Descriptors (ADs), and atypicality ratings. A statistical analysis of static images from the HRM reveals significant differences between the ASD and TD groups across multiple AUs and ADs when displaying the same emotional expressions, confirming that participants with ASD tend to demonstrate more irregular and diverse expression patterns. Subsequently, a temporal regression method was employed to analyze atypicality of dynamic sequences, thereby bridging the gap between subjective perception and objective facial characteristics. Furthermore, baseline results for AU detection are provided for future research reference. This work not only contributes to our understanding of the unique facial expression characteristics associated with ASD but also provides potential tools for ASD early screening. Portions of the dataset and pretrained models are accessible at:https://github.com/Jonas-DL/Hugging-Rain-Man.
Yanfeng Ji, Shutong Wang, Ruyi Xu, Jingying Chen 0001, Yuxuan Quan, Xinzhou Jiang, Zhengyu Deng
IEEE Trans. Affect. Comput.4
2025 Leveraging Eye Movement for Instructing Robust Video-Based Facial Expression Recognition
abstract
Video-based facial expression recognition (VFER) is challenging due to variations caused by cultural background and expression camouflage. To tackle these problems, researchers introduced eye movement signals to complement visual information. However, existing methods either require expensive devices to capture high-quality eye movements or can only extract low-quality eye movements visually, making them ineffective in the real world. To address this, we propose an eye movement-instructed VFER (EM-VFER) that leverages high-quality eye movements to instruct the visual learning, obtaining robust performance without requiring costly devices during inference. Specifically, our EM-VFER operates in two stages: the high-quality eye movement pre-training stage and the eye movement-instructed video fine-tuning stage. In the pre-training, we compile an Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset and use it to train a multimodal Transformer. During the fine-tuning, we propose a novel progressive eye movement-instructed learning to take better advantage of the prior knowledge about high-quality eye movement signals from EMER. The instructed fine-tuning model could then make more robust predictions on downstream facial expression datasets. We evaluate our approach on three macroexpression datasets (DFEW, MAFW and Aff-wild2) and two micro-expression datasets (CASME III and CASME II). The results demonstrate that EM-VFER significantly outperforms existing methods. The code will be available.
Yuanyuan Liu 0004, Kejun Liu, Zijing Chen, Zhe Chen 0013, Chang Tang, Jingying Chen 0001, Shiguang Shan
IEEE Trans. Affect. Comput.7
2024 VS: Reconstructing Clothed 3D Human from Single Image via Vertex Shift
abstract
Various applications require high-fidelity and artifact free 3D human reconstructions. However, current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper, we propose a two-stage deformation method named Vertex Shift (VS) for reconstructing clothed 3D humans from single images. Specifically, VS first stretches the estimated SMPL-X mesh into a coarse 3D human model using shift fields inferred from normal maps, then refines the coarse 3D human model into a detailed 3D human model via a graph convolutional network embedded with implicit-function-learned features. This “stretch-refine” strategy addresses large deformations required for reconstructing loose clothing and delicate deformations for recovering intricate and detailed surfaces, achieving high-fidelity reconstructions that faithfully convey the pose, clothing, and surface details from the input images. The graph convolutional network's ability to exploit neighborhood vertices coupled with the advantages inherited from the deformation methods ensure VS rarely produces artifacts like distortions and non-human shapes and never produces artifacts like holes, broken parts, and dismembered limbs. As a result, VS can reconstruct highfidelity and artifact-less clothed 3D humans from single images, even under scenarios of challenging poses and loose clothing. Experimental results on three benchmarks and two in-the-wild datasets demonstrate that VS significantly outperforms current state-of-the-art methods. The code and models of VS are available for research purposes at https://github.com/starVisionTeam/VS.
Leyuan Liu 0001, Yuhan Li 0009, Yunqi Gao, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001
CVPR6
2024 End-to-End Dense Video Captioning Model Based on Multimodal Feature Fusion
abstract
The goal of dense video captioning is to generate multiple text descriptions related to different timestamps from a video. Traditional methods adopt a two-stage "localize-then-describe" approach but overlook the correlation between localization and captioning, leading to inaccurate predictions of event boundaries. Additionally, existing models often rely solely on visual information, neglecting audio and textual subtitle cues, which results in less accurate text descriptions. To address these issues, this paper proposes an end-to-end dense video captioning model. In this model, visual, audio, and textual features of the video are used as inputs to the encoding layer, processed through three structurally identical Deformable Transformer encoders. Three parallel tasks—localization head, captioning head, and event counter—are designed to facilitate end-to-end training. A multimodal feature fusion module is introduced to integrate the three modal features into a temporally coherent feature vector. Finally, the model’s performance is further enhanced through reinforcement learning. Extensive experiments on the ActivityNet dataset demonstrate that the proposed model outperforms PDVC and MDVC in generating higher-quality descriptions.
Shixin Peng, Jingying Chen 0001
HPCC3
2024 An Avatar-Based Intervention System for Children with Autism Spectrum Disorder
Leyuan Liu 0001, Yuanjian You, Zhichen He, Jingying Chen 0001
PRCV (9)4
2024 Evaluation of the gross motor abilities of autistic children with a computerised evaluation method
abstract
To effectively evaluate the gross motor ability of autistic children, we proposed a method of computerised evaluation of gross motor skills (CEGM). The CEGM integrates Dynamic Time Warping (DTW) method and OpenPose technology to automatically detect key joints and return a score. Ten items were selected for evaluation based on the gross motor subtest of the Psychoeducational Profile – Third Edition (PEP-3) scale, including upper limb movement, lower limb movement, and body coordination performance. 30 autistic participants (males: 23, female: 7) with an average age of 5.00 years were recruited in this study. Then we compared the results of evaluation using CEGM and the original PEP-3 gross motor subtest in autistic children. The results showed that in the evaluations using CEGM and PEP-3, Cronbach’s α coefficients and Spearman-rank correlation coefficients were all greater than 0.80, intraclass correlation coefficient (ICC) were all greater than 0.90, indicating good agreement in evaluating the gross motor ability of autistic children. Moreover, compared to the PEP-3, the evaluation using CEGM provided precise quantitative indicators (trajectory, velocity, and angle of joint). Therefore, our findings demonstrate that CEGM can be used in the initial evaluation of the gross motor ability of autistic children.
Xiaodi Liu, Jingying Chen 0001, Guangshuai Wang, Kun Zhang 0031, Jianchi Sun, Pianpian Ma, Rujing Zhang
Behav. Inf. Technol.2
2024 Facial expression intensity estimation using label-distribution-learning-enhanced ordinal regression
Ruyi Xu, Zhun Wang, Jingying Chen 0001, Longpu Zhou
Multim. Syst.3
2024 V2IED: Dual-view learning framework for detecting events of interictal epileptiform discharges
Zhekai Ming, Dan Chen 0001, Tengfei Gao, Yunbo Tang, Weiping Tu, Jingying Chen 0001
Neural Networks6
2024 Dual subspace manifold learning based on GCN for intensity-invariant facial expression recognition
Jingying Chen 0001, Jinxin Shi, Ruyi Xu
Pattern Recognit.1
2024 SeIF: Semantic-Constrained Deep Implicit Function for Single-Image 3D Head Reconstruction
abstract
Various applications require realistic, artifact-free, and animatable 3D avatars. However, traditional 3D morphable models (3DMMs) produce animatable 3D heads but fail to capture accurate geometries and details, while existing deep implicit functions have been shown to achieve realistic reconstructions but suffer from artifacts and struggle to yield 3D heads that are easy to animate. To reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images, we leverage semantics to bridge the best properties of 3DMMs and deep implicit functions and propose SeIF—a semantic-constrained deep implicit function. First, SeIF derives fine-grained semantics from a standard 3DMM (e.g., FLAME) and samples a semantic code for each query point in the query space to provide a soft constraint to the deep implicit function. The reconstruction results show that this semantic constraint does not weaken the powerful representation ability of the deep implicit function while significantly suppressing artifacts. Second, SeIF predicts a more accurate semantic code for each query point and utilizes the semantic codes to uniformize the structure of reconstructed 3D head meshes with the standard 3DMM. Since our reconstructed 3D head meshes have the same structure as the 3DMM, 3DMM-based animation approaches can be easily transferred to animate our reconstructed 3D heads. As a result, SeIF can reconstruct high-fidelity, artifact-less, and animatable 3D heads from single-view images of individuals with diverse ages, genders, races, and facial expressions. Quantitative and qualitative experimental results on seven datasets show that SeIF outperforms existing state-of-the-art methods by a large margin. The code and models of SeIF are available for research purposes athttps://github.com/starVisionTeam/SeIF.
Leyuan Liu 0001, Jianchi Sun, Changxin Gao, Jingying Chen 0001
IEEE Trans. Multim.5
2023 Single-image clothed 3D human reconstruction guided by a well-aligned parametric body model
Leyuan Liu 0001, Yunqi Gao, Jianchi Sun, Jingying Chen 0001
Multim. Syst.4
2023 Channel decoupling network for cross-modality person re-identification
Jingying Chen 0001, Shixin Peng
Multim. Tools Appl.1
2022 Orthogonal channel attention-based multi-task learning for multi-view facial expression recognition
Jingying Chen 0001, Ruyi Xu
Pattern Recognit.1
2022 Toward Children's Empathy Ability Analysis: Joint Facial Expression Recognition and Intensity Estimation Using Label Distribution Learning
abstract
Empathy ability is one of the most important social communication skills in early childhood development. To analyze the children's empathy ability, facial expression analysis (FEA) is an effective way due to its ability to understand children's emotional states. Previous works mainly focus on recognizing the facial expression categories yet fail to estimate expression intensity, the latter of which is more important for fine-grained emotion analysis. To this end, this article first proposes to analyze children's empathy ability with both the categories and the intensities of facial expressions. A novel FEA method based on intensity label distribution learning is presented, which aims to recognize expression categories and estimate their intensity levels in an end-to-end framework. First, the intensity label distribution is generated for each frame in the expression sequence using a linear interpolation estimation and a Gaussian function to address the lack of reasonable annotations for expression intensity. Then, the extended intensity label distribution is presented to automatically encode the expression intensity in a multidimensional expression space, which aims to integrate the expression recognition and intensity estimation into a unified framework as well as boost the expression recognition performance by suppressing the variations in appearance caused by intensity and by emphasizing those variations among weak expressions. Finally, a Siamese-like convolutional neural network is presented to learn the expression model from a pair of frames that includes an expressive frame and its corresponding neutral frame using the extended intensity label distribution as the supervised information, thus effectively eliminating the expression-unrelated information's influence on FEA. Numerous experiments validate that the proposed method is promising in analysis of the differences in empathy ability between typically developing children and children with autism spectrum disorder.
Jingying Chen 0001, Ruyi Xu, Kun Zhang 0031, Zongkai Yang, Honghai Liu 0001
IEEE Trans. Ind. Informatics1
2021 Design and application of facial expression analysis system in empathy ability of children with autism spectrum disorder
abstract
Empathy is an important social ability in the early childhood development.One of the significant characteristics of children with autism spectrum disorder (ASD) is their lack of empathy, which makes it difficult for them to feel and understand other people's emotions and to judge other people's behavioral intentions, leading to social disorders.This research designs and implements a facial expression analysis system that could obtain and analyze the real-time facial expressions of children when viewing stimulus materials, and then evaluate the differences of empathy ability between ASD children and typical development (TD) children.The results of this research provide new ideas for the evaluation of ASD children, and also help to develop empathy intervention plans for ASD children.
Kun Zhang 0031, Jingying Chen 0001, Ruyi Xu
FedCSIS3
2021 HEI-Human: A Hybrid Explicit and Implicit Method for Single-View 3D Clothed Human Reconstruction
Leyuan Liu 0001, Jianchi Sun, Yunqi Gao, Jingying Chen 0001
PRCV (2)4
2020 PlugNet: Degradation Aware Scene Text Recognition Supervised by a Pluggable Super-Resolution Unit
Yongqiang Mou, Jingying Chen 0001, Leyuan Liu 0001, Yaohong Huang
ECCV (15)4
2019 Progressive Pose Normalization Generative Adversarial Network for Frontal Face Synthesis and Face Recognition Under Large Pose
abstract
This paper proposes a Progressive Pose-Normalization Generative Adversarial Network (PPN-GAN) for frontal face synthesis and face recognition. The key idea is to normalize a profile face progressively: starting from inferring an intermediate face that has a small view difference to the profile face, and then increasing the view difference step by step, until the frontal view of the profile face is recovered. In addition to the progressive strategy, an additional identity discriminator and identity-aware losses in both the image and feature spaces are also incorporated into the GAN for identity preserving. Experimental results show that our method not only produces compelling perceptual results but also outperforms the state-of-the-art methods on face recognition under large-pose.
Leyuan Liu 0001, Jingying Chen 0001
ICIP3
2019 Head pose estimation with soft labels using regularized convolutional neural network
Luhui Xu, Jingying Chen 0001, Yanling Gan
Neurocomputing2
2019 Visual Focus of Attention and Spontaneous Smile Recognition Based on Continuous Head Pose Estimation by Cascaded Multi-Task Learning
abstract
Multi-person Visual focus of attention (M-VFOA) and spontaneous smile (SS) recognition are important for persons’ behavior understanding and analysis in class. Recently, promising results have been reported using special hardware in constrained environment. However, M-VFOA and SS remain challenging problems in natural and crowd classroom environment, e.g. various poses, occlusion, expressions, illumination and poor image quality, etc. In this study, a robust and un-invasive M-VFOA and SS recognition system has been developed based on continuous head pose estimation in the natural classroom. A novel cascaded multi-task Hough forest (CM-HF) combined with weighted Hough voting and multi-task learning is proposed for continuous head pose estimation, tip of the nose location and SS recognition, which improves accuracies of recognition and reduces the training time. Then, M-VFOA can be recognized based on estimated head poses, environmental cues and prior states in the natural classroom. Meanwhile, SS is classified using CM-HF with local cascaded mouth-eyes areas normalized by the estimated head poses. The method is rigorously evaluated for continuous head pose estimation, multi-person VFOA recognition, and SS recognition on some public available datasets and real-class video sequences. Experimental results show that our method reduces training time greatly and outperforms the state-of-the-art methods for both performance and robustness with an average accuracy of 83.5% on head pose estimation, 67.8% on M-VFOA recognition and 97.1% on SS recognition in challenging environments.
Yuanyuan Liu 0004, Xingmei Li, Fang Fang 0008, Fayong Zhang, Jingying Chen 0001, Zhizhong Zeng
Int. J. Pattern Recognit. Artif. Intell.5
2019 Automatic social signal analysis: Facial expression recognition using difference convolution neural network
Jingying Chen 0001, Yongqiang Lv, Ruyi Xu
J. Parallel Distributed Comput.1
2019 Person re-identification with multiple similarity probabilities using deep metric learning for efficient smart security applications
Mingfu Xiong, Dan Chen 0001, Jun Chen 0001, Jingying Chen 0001, Benyun Shi, Chao Liang 0001, Ruimin Hu
J. Parallel Distributed Comput.4
2019 Head pose estimation using improved label distribution learning with fewer annotations
Luhui Xu, Jingying Chen 0001, Yanling Gan
Multim. Tools Appl.2
2019 Facial expression recognition boosted by soft label with a diverse ensemble
Yanling Gan, Jingying Chen 0001, Luhui Xu
Pattern Recognit. Lett.2
2018 Semi-supervised Learning of Deep Difference Features for Facial Expression Recognition
Ruyi Xu, Jingying Chen 0001, Leyuan Liu 0001
PRCV (3)3
2018 Bayesian tensor factorization for multi-way analysis of multi-dimensional EEG
Yunbo Tang, Dan Chen 0001, Lizhe Wang 0001, Albert Y. Zomaya, Jingying Chen 0001, Honghai Liu 0001
Neurocomputing5
2018 Deep peak-neutral difference feature for facial expression recognition
Jingying Chen 0001, Ruyi Xu, Leyuan Liu 0001
Multim. Tools Appl.1
2018 Student engagement study based on multi-cue detection and recognition in an intelligent learning environment
Yuanyuan Liu 0004, Jingying Chen 0001, Mulan Zhang, Chuan Rao
Multim. Tools Appl.2
2017 Design and Implementation of Fire Safety Education System on Campus based on Virtual Reality Technology
abstract
Fire safety education is essential to every student on campus.Fire safety knowledge learning and operational practice are both important.There is evidence that the virtual reality (VR) based educational method can be a novel and effective approach to learning and practice.However, the existing VR-based system for fire safety education has some shortcomings such as lack of interactivity and high equipment complexity, resulting in low practicability.In order to improve the effect of fire safety education on campus, this paper establishes the model and architecture of fire safety education system based on VR technology.The framework and various elements of fire safety education system are designed and implemented according to the combination of relevant fire safety education theory and VR technology.Finally the prototype version of fire safety education system based on VR technology is built on the HTC VIVE helmet equipment.Through the usability test and comparative analysis of the application experiment, the experiment results prove the feasibility and effectiveness of the proposed approach.
Kun Zhang 0031, Jintao Suo, Jingying Chen 0001, Xiaodi Liu
FedCSIS3
2017 Book Page Identification Using Convolutional Neural Networks Trained by Task-Unrelated Dataset
Leyuan Liu 0001, Huabing Zhou, Jingying Chen 0001
ICIG (1)4
2017 An Action Unit based Hierarchical Random Forest Model to Facial Expression Recognition
Jingying Chen 0001, Mulan Zhang, Xianglong Xue, Ruyi Xu, Kun Zhang 0031
ICPRAM1
2017 Multi-level structured hybrid forest for joint head detection and pose estimation
Yuanyuan Liu 0004, Zhong Xie, Xiaohui Yuan 0001, Jingying Chen 0001, Wu Song
Neurocomputing4
2017 Comprehensive Association Rules Mining of Health Examination Data with an Extended FP-Growth Method
Bowei Wang, Dan Chen 0001, Benyun Shi, Yifu Duan, Jingying Chen 0001, Ruimin Hu
Mob. Networks Appl.6
2017 A low-cost real-time face tracking system for ITSs and SDASs
abstract
Summary It is important to track people's face efficiently and accurately in many Intelligent Transportation Systems (ITSs) and Safety Driving Assistant Systems (SDASs). This paper presents a high‐performance and low‐cost real‐time face tracking system, which runs on general onboard computer with very low CPU consumption. The proposed face tracking system is composed of four modules: the motion detector, face detector, face tracker, and face validator. The motion detector extracts motion areas by using a spatial‐temporal bi‐differential method with a very low computational cost. The face detector integrates motions into a cascade face detection framework to reject most of non‐face scanning‐windows to ensure efficient face localization. The face tracker fuses motion feature with color feature to alleviate the drifting problem during tracking. The face validator builds face appearance models online and identifies each specific tracked face to avoid confusion. Experimental results on three challenging video sequences show that the proposed face tracking system outperforms the state‐of‐the‐art face tracker and consumes only 5–13% CPU resources of a low‐spec onboard computer while processing in real time. Copyright © 2016 John Wiley & Sons, Ltd.
Leyuan Liu 0001, Jingying Chen 0001, Changxin Gao, Nong Sang
Softw. Pract. Exp.2
2016 Multi-person Visual Focus of Attention from Head Pose on a Natural Classroom
Yuanyuan Liu 0004, Leyuan Liu 0001, Jingying Chen 0001, Chunyan Su, Kun Zhang 0031
ICPRAM3
2016 Robust head pose estimation using Dirichlet-tree distribution enhanced random forests
Yuanyuan Liu 0004, Jingying Chen 0001, Zhiming Su, Zhenzhen Luo, Nan Luo, Leyuan Liu 0001, Kun Zhang 0031
Neurocomputing2
2015 Parallel Simulation of Complex Evacuation Scenarios with Adaptive Agent Models
abstract
Simulation study on evacuation scenarios has gained tremendous attention in recent years. Two major research challenges remain along this direction: (1) how to portray the effect of individuals' adaptive behaviors under various situations in the evacuation procedures and (2) how to simulate complex evacuation scenarios involving huge crowds at the individual level due to the ultrahigh complexity of these scenarios. In this study, a simulation framework for general evacuation scenarios has been developed. Each individual in the scenario is modeled as an adaptable and autonomous agent driven by a weight-based decision-making mechanism. The simulation is intended to characterize the individuals' adaptable behaviors, the interactions among individuals, among small groups of individuals, and between the individuals and the environment. To handle the second challenge, this study adopts GPGPU to sustain massively parallel modeling and simulation of an evacuation scenario. An efficient scheme has been proposed to minimize the overhead to access the global system state of the simulation process maintained by the GPU platform. The simulation results indicate that the “adaptability” in individual behaviors has a significant influence on the evacuation procedure. The experimental results also exhibit the proposed approach's capability to sustain complex scenarios involving a huge crowd consisting of tens of thousands of individuals.
Dan Chen 0001, Lizhe Wang 0001, Albert Y. Zomaya, Minggang Dou, Jingying Chen 0001, Ze Deng, Salim Hariri
IEEE Trans. Parallel Distributed Syst.5
2014 Dirichlet-tree Distribution Enhanced Random Forests for Head Pose Estimation
abstract
Head pose estimation is important in human-machine interfaces. However, illumination variation, occlusion and low image resolution make the estimation task difficult. Hence, a Dirichlet-tree distribution enhanced Random Forests approach (D-RF) is proposed in this paper to estimate head pose efficiently and robustly under various conditions. First, PCA based sub-features space from Gabor features and histogram distributions of the facial patches are extracted to eliminate the influence of occlusion and noise. Then, the D-RF is proposed to estimate the head pose in a coarse-to-fine way. In order to improve the discrimination capability of the approach, an adaptive Gaussian mixture model is introduced in the tree distribution. The proposed method has been evaluated with different data sets spanning from -90° to 90° in vertical and horizontal directions under various conditions. The experimental results demonstrate the approach’s robustness and efficiency.
Yuanyuan Liu 0004, Jingying Chen 0001, Leyuan Liu 0001, Yujiao Gong, Nan Luo
ICPRAM2
2014 Modeling and simulation for natural disaster contingency planning driven by high-resolution remote sensing images
Minggang Dou, Jingying Chen 0001, Dan Chen 0001, Xiaodao Chen, Ze Deng, Jian Wang 0079
Future Gener. Comput. Syst.2
2014 Towards Improving Social Communication Skills With Multimodal Sensory Information
abstract
How to improve social communication skills for children, especially those with social communication difficulties such as attention deficit/hyperactivity disorder, has long been a challenge faced by researchers and therapists. Recent research indicates that computer-assisted approaches may be effective in addressing this issue. This study aimed to understand children's behaviors and then provide appropriate support to improve their social communication skills. We have established an intelligent system, inside which a child can freely play interactive social skills games with virtual characters. The virtual characters can adjust their own behaviors by adapting to the child's cognitive state (e.g., focus of attention) and affective state (e.g., happiness or surprise). The child's behavior is identified in real-time by recognition of multimodal sensory information, which includes head pose and eye gaze estimation, gesture detection, and affective state detection supported by a series of algorithms proposed in this study. Furthermore, this intelligent system has been enabled in a nonintrusive manner using a novel approach of multicamera surveillance to provide the child with natural interaction with the system. Experimental results show the system can estimate a user's attention and affective states with correctness rates of 93% and 91.3%, respectively. The results obtained suggest that the methods have strong potential as alternative methods for sensing human behavior and providing appropriate support.
Jingying Chen 0001, Dan Chen 0001, Xiaoli Li 0002, Kun Zhang 0031
IEEE Trans. Ind. Informatics1
2013 Hybrid modelling and simulation of huge crowd over a hierarchical Grid architecture
Dan Chen 0001, Lizhe Wang 0001, Jingying Chen 0001, Samee Ullah Khan, Joanna Kolodziej, Mingwei Tian, Fang Huang 0001, Wangyang Liu
Future Gener. Comput. Syst.4
2013 G-Hadoop: MapReduce across distributed data centers for data-intensive computing
Lizhe Wang 0001, Jie Tao 0001, Rajiv Ranjan 0001, Holger Marten, Achim Streit, Jingying Chen 0001, Dan Chen 0001
Future Gener. Comput. Syst.6
2013 Natural Disaster Monitoring with Wireless Sensor Networks: A Case Study of Data-intensive Applications upon Low-Cost Scalable Systems
Dan Chen 0001, Zhixin Liu 0001, Lizhe Wang 0001, Minggang Dou, Jingying Chen 0001
Mob. Networks Appl.5