EDBT 2026 Demo / reviewers in the wild / expert
Hui Chen 0020
dblp:12/417-20
· DBLP profile ↗
31ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-2312-4298ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 9 · 4 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationabstractTalking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS)-based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation, with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Our method supports applications such as VR telepresence, social VR, and embodied conversational agents. The socially-aware conditioning can also be extended to other human motion generation tasks. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released. Peng Chen 0046, Xiaobao Wei, Yi Yang 0060, Naiming Yao, Hui Chen 0020, Feng Tian 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | GraphAvatar: Compact Head Avatars with GNN-Generated 3D GaussiansabstractRendering photorealistic head avatars from arbitrary viewpoints is crucial for various applications like virtual reality. Although previous methods based on Neural Radiance Fields (NeRF) can achieve impressive results, they lack fidelity and efficiency. Recent methods using 3D Gaussian Splatting (3DGS) have improved rendering quality and real-time performance but still require significant storage overhead. In this paper, we introduce a method called GraphAvatar that utilizes Graph Neural Networks (GNN) to generate 3D Gaussians for the head avatar. Specifically, GraphAvatar trains a geometric GNN and an appearance GNN to generate the attributes of the 3D Gaussians from the tracked mesh. Therefore, our method can store the GNN models instead of the 3D Gaussians, significantly reducing the storage overhead to just 10MB. To reduce the impact of face-tracking errors, we also present a novel graph-guided optimization module to refine face-tracking parameters during training. Finally, we introduce a 3D-aware enhancer for post-processing to enhance the rendering quality. We conduct comprehensive experiments to demonstrate the advantages of GraphAvatar, surpassing existing methods in visual fidelity and storage consumption. The ablation study sheds light on the trade-offs between rendering quality and model size. Xiaobao Wei, Peng Chen 0046, Ming Lu 0002, Hui Chen 0020, Feng Tian 0001 |
AAAI | 4 |
| 2025 | CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze EstimationabstractGaze estimation methods typically encounter significant performance degradation in generalized tasks due to the domain mismatch between the source and target domains. Existing approaches attempt to utilize various domain generalization techniques. However, their generalization capabilities are limited since they are constrained to a single visual modality. Notably, large-scale contrastive language-image pre-training (CLIP) models have been widely applied to downstream visual tasks for their robust generalization capabilities, but the potential of CLIP for regression tasks has not been fully explored. To bridge this gap, we introduce a novel framework called CR-CLIP, which endows CLIP with the capability to generalize gaze estimation. Specifically, we convert gaze labels into textual descriptions and achieve alignment between images and text signals with gaze cues, thereby extracting generalized gaze-related features. To enhance the model’s understanding of the numerical relationships of gaze directions, we propose a novel regression loss function based on image-text similarity. Additionally, we fine-tune the model on the original gaze dataset, achieving high precision in generalized gaze estimation. Experimental results show that our proposed method achieves state-of-the-art performance on four generalized gaze estimation tasks. Yitong Zhu, Xurong Xie, Naiming Yao, Hui Chen 0020, Feng Tian 0001 |
ICASSP | 4 |
| 2025 | GazeGaussian: High-Fidelity Gaze Redirection with 3D Gaussian SplattingabstractGaze estimation encounters generalization challenges when dealing with out-of-distribution data. To address this problem, recent methods use neural radiance fields (NeRF) to generate augmented data. However, existing methods based on NeRF are computationally expensive and lack facial details. 3D Gaussian Splatting (3DGS) has become the prevailing representation of neural fields. While 3DGS has been extensively examined in head avatars, it faces challenges with accurate gaze control and generalization across different subjects. In this work, we propose GazeGaussian, the first high-fidelity gaze redirection method that uses a two-stream 3DGS model to represent the face and eye regions separately. Leveraging the unstructured nature of 3DGS, we develop a novel representation of the eye for rigid eye rotation based on the target gaze direction. To enable synthesis generalization across various subjects, we integrate an expression-guided module to inject subject-specific information into the neural renderer. Comprehensive experiments show that GazeGaussian outperforms existing methods in rendering speed, gaze redirection accuracy, and facial synthesis across multiple datasets. The code is available at: https://ucwxb.github.io/GazeGaussian. Xiaobao Wei, Peng Chen 0046, Ming Lu 0002, Hui Chen 0020, Feng Tian 0001 |
ICCV | 5 |
| 2025 | DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided DistillationabstractReal-time speech-driven 3D facial animation has been attractive in academia and industry. Traditional methods mainly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the nondeterministic fact of speech-driven 3D face animation and employ the diffusion model for the task. Existing diffusion-based methods can improve the diversity of facial animation. However, personalized speaking styles conveying accurate lip language is still lacking, besides, efficiency and compactness still need to be improved. In this work, we propose DiffusionTalker to address the above limitations via personalizer-guided distillation. In terms of personalization, we introduce a contrastive personalizer that learns identity and emotion embeddings to capture speaking styles from audio. We further propose a personalizer enhancer during distillation to enhance the influence of embeddings on facial animation. For efficiency, we use iterative distillation to reduce the steps required for animation generation and achieve more than 8x speedup in inference. To achieve compactness, we distill the large teacher model into a smaller student model, reducing our model’s storage by 86.4% while minimizing performance loss. After distillation, users can derive their identity and emotion embeddings from audio to quickly create personal-ized animations that reflect specific speaking styles. Extensive experiments are conducted to demonstrate that our method outperforms state-of-the-art methods. The code is released at: https://github.com/ChenVoid/DiffusionTalker. Peng Chen 0046, Xiaobao Wei, Ming Lu 0002, Hui Chen 0020, Feng Tian 0001 |
ICME | 4 |
| 2024 | LLM-Empowered Few-Shot Node Classification on Incomplete Graphs with Real Node DegreesabstractGraphs constructed from real-world scenarios are often incomplete due to privacy restrictions or resource limitations, posing significant challenges for node classification, especially when labeled data are scarce. In many scenarios of incomplete graphs, the real node degrees, such as the number of followers in social networks or publications' references in citation networks, are easily accessible and informative, which could indicate the degree of incompleteness. However, most of existing researches of incomplete graphs focus on edge completion, but ignore the node completion with known node degrees. In this paper, we propose a new few-shot node classification problem on incomplete graphs with real node degrees. To deal with node completion, edge completion and label completion of this problem, we develop an effective Large Language Models (LLMs) empowered Graph Convolutional Network (GCN) model utilizing the real node Degrees, namely LLMDGCN. First, we leverage LLMs to initially fill in the missing nodes and labels. Next, we design an edge prediction module that employs the real node degrees and inter-category probability matrix to recover the missing edges for each node. We then iteratively train the GCN and the edge prediction module. The GCN generates pseudo labels, which the edge prediction module uses to restore edges, and these edges are fed back into the GCN to improve accuracy. Extensive experiments on four benchmark datasets demonstrate the effectiveness and robustness of our proposed method for the few-shot node classification on incomplete graphs with real node degrees. Yi Yang 0060, Jiaqi Zhu 0001, Hui Chen 0020, Hongan Wang |
CIKM | 4 |
| 2024 | Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
Yicong Jiang, Tianzi Wang, Xurong Xie, Juan Liu 0008, Wei Sun 0050, Hui Chen 0020, Xunying Liu, Feng Tian 0001 |
INTERSPEECH | 7 |
| 2024 | Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized CharactersabstractAnimating virtual characters has always been a fundamental research problem in virtual reality (VR). Facial animations play a crucial role as they effectively convey emotions and attitudes of virtual humans. However, creating such facial animations can be challenging, as current methods often involve utilization of expensive motion capture devices or significant investments of time and effort from human animators in tuning animation parameters. In this paper, we propose a holistic solution to automatically animate virtual human faces. In our solution, a deep learning model was first trained to retarget the facial expression from input face images to virtual human faces by estimating the blendshape coefficients. This method offers the flexibility of generating animations with characters of different appearances and blendshape topologies. Second, a practical toolkit was developed using Unity 3D, making it compatible with the most popular VR applications. The toolkit accepts both image and video as input to animate the target virtual human faces and enables users to manipulate the animation results. Furthermore, inspired by the spirit of Human-in-the-loop (HITL), we leveraged user feedback to further improve the performance of the model and toolkit, thereby increasing the customization properties to suit user preferences. The whole solution, for which we will make the code public, has the potential to accelerate the generation of facial animations for use in VR applications. https://github.com/showlab/BYOC Zechen Bai, Peng Chen 0046, Xiaolan Peng, Naiming Yao, Hui Chen 0020 |
VR | 6 |
| 2024 | Development and validation of a deep interpretable network for continuous acute kidney injury prediction in critically ill patients
Meicheng Yang, Songqiao Liu, Caiyun Ma, Hui Chen 0020, Yuwen Li 0002, Changde Wu, Jianfeng Xie, Haibo Qiu, Jianqing Li 0002, Yi Yang 0060, Chengyu Liu 0001 |
Artif. Intell. Medicine | 5 |
| 2024 | Optical Character Recognition (OCR)-Based and Gaussian Mixture Modeling-OCR-Based Slide-Level "With-Me-Ness": Automated Measurement and Feedback of Learners' Attention State during Video LecturesabstractAs video lectures are gaining more popularity, determining their effectiveness and obtaining valuable feedback have become necessary. To measure the learners’ attention state during video lectures, we specified the conceptual “with-me-ness” (WMN) as slide-level WMN (SL-WMN). The content domain on each slide was automatically extracted via an optical character recognition (OCR)-based method, while the eye gazing behaviors were analyzed through a Gaussian mixture modeling (GMM) fixation clustering method. Both domain-specific WMN and behavior-enriched WMN were then computed via OCR- and GMM-OCR-based methods to measure the learners’ attention levels. We conducted an experiment to collect in-lecture eye-tracking data, video recordings, and post-lecture test scores from 50 Grade 8 students. The results demonstrated that both OCR- and GMM-OCR-based SL-WMNs are reliable and compatible automatic measurements of learners’ attention states during video lectures. A survey from participating learners and lecturers also revealed highly favorable feedback for the developed SL-WMNs. Chengchen Lyu, Hui Chen 0020, Xiaolan Peng, Juntao Ye, Hongan Wang |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | DailyConnect: Piloting Interventions of Situation-Based Emotional Understanding in Naturalistic Home Settings for Children with Autism Spectrum DisorderabstractDailyConnect is a visual-based mobile application that supports children with autism spectrum disorder (ASD) in recalling memories by reviewing photos through discrete trial training (DTT) to understand situation-based emotions. To assess DailyConnect and its adaptability to a child’s characteristics and emotional situations, a pilot study was conducted that included 15 children with ASD and their parents and teachers. The DTT steps—memory recall and situation recognition, emotion recognition, emotion cues, facial expression recognition, and response behavior—were reliable in assessing the understanding of emotional situations when compared before and after the intervention in four categories of emotional situations, namely, happiness, sadness, anger, and fear. The results revealed that DailyConnect improves the understanding of situation-based emotions, particularly negative emotions (e.g., sadness: mean diff = −.687, sig. < .01, T = −3.866, d = 1.006; anger: mean diff = −.952, sig. < .01, T = −6.187, d = .705; fear: mean diff = −.961, sig. < .01, T = −5.522, d = .627); however, its effectiveness varied for children in different emotional situations. Furthermore, subjective feedback from participants and users (parents and teachers) provided insights into design considerations for similar mobile aids. Chengchen Lyu, Hui Chen 0020, Tong Xu 0008, Xiaolan Peng, Faliang Huang, Hongan Wang |
Int. J. Hum. Comput. Interact. | 2 |
| 2024 | Automatic detection of breast lesions in automated 3D breast ultrasound with cross-organ transfer learningabstractDeep convolutional neural networks have garnered considerable attention in numerous machine learning applications, particularly in visual recognition tasks such as image and video analyses. There is a growing interest in applying this technology to diverse applications in medical image analysis. Automated three-dimensional Breast Ultrasound is a vital tool for detecting breast cancer, and computer-assisted diagnosis software, developed based on deep learning, can effectively assist radiologists in diagnosis. However, the network model is prone to overfitting during training, owing to challenges such as insufficient training data. This study attempts to solve the problem caused by small datasets and improve model detection performance. We propose a breast cancer detection framework based on deep learning (a transfer learning method based on cross-organ cancer detection) and a contrastive learning method based on breast imaging reporting and data systems (BI-RADS). When using cross organ transfer learning and BIRADS based contrastive learning, the average sensitivity of the model increased by a maximum of 16.05%. Our experiments have demonstrated that the parameters and experiences of cross-organ cancer detection can be mutually referenced, and contrastive learning method based on BI-RADS can improve the detection performance of the model. B. A. O. Lingyun, Zhengrui Huang, Yue Sun 0001, Hui Chen 0020, Xiaochen Yuan, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 5 |
| 2023 | ChallengeDetect: Investigating the Potential of Detecting In-Game Challenge Experience from Physiological MeasuresabstractChallenge is the core element of digital games. The wide spectrum of physical, cognitive, and emotional challenge experiences provided by modern digital games can be evaluated subjectively using a questionnaire, the CORGIS, which allows for a post hoc evaluation of the overall experience that occurred during game play. Measuring this experience dynamically and objectively, however, would allow for a more holistic view of the moment-to-moment experiences of players. This study, therefore, explored the potential of detecting perceived challenge from physiological signals. For this, we collected physiological responses from 32 players who engaged in three typical game scenarios. Using perceived challenge ratings from players and extracted physiological features, we applied multiple machine learning methods and metrics to detect challenge experiences. Results show that most methods achieved a detection accuracy of around 80%. We discuss in-game challenge perception, challenge-related physiological indicators and AI-supported challenge detection to inform future work on challenge evaluation. Xiaolan Peng, Xurong Xie, Jin Huang 0009, Chutian Jiang, Haonian Wang, Alena Denisova, Hui Chen 0020, Feng Tian 0001, Hongan Wang |
CHI | 7 |
| 2023 | Unsupervised Model-Based Speaker Adaptation of End-To-End Lattice-Free MMI Model for Speech RecognitionabstractModeling the speaker variability is a key challenge for automatic speech recognition (ASR) systems. In this paper, the learning hidden unit contributions (LHUC) based adaptation techniques with compact speaker dependent (SD) parameters are used to facilitate both speaker adaptive training (SAT) and unsupervised test-time speaker adaptation for end-to-end (E2E) lattice-free MMI (LF-MMI) models. An unsupervised model-based adaptation framework is proposed to estimate the SD parameters in E2E paradigm using LF-MMI and cross entropy (CE) criterions. Various regularization methods of the standard LHUC adaptation, e.g., the Bayesian LHUC (BLHUC) adaptation, are systematically investigated to mitigate the risk of overfitting, on E2E LF-MMI CNN-TDNN and CNN-TDNN- BLSTM models. Lattice-based confidence score estimation is used for adaptation data selection to reduce the supervision label uncertainty. Experiments on the 300-hour Switchboard task suggest that, applying BLHUC in the proposed unsupervised E2E adaptation framework to byte pair encoding (BPE) based E2E LF-MMI systems consistently outperformed the baseline systems by relative word error rate (WER) reductions up to 10.5% and 14.7% on the NIST Hub5’00 and RT03 evaluation sets, and achieved the best performance in WERs of 9.0% and 9.7%, respectively. These results are comparable to the results of state-of-the-art adapted LF-MMI hybrid systems and adapted Conformer-based E2E systems. Xurong Xie, Xunying Liu, Hui Chen 0020, Hongan Wang |
ICASSP | 3 |
| 2021 | Enhancing Emotional Experience by Building Emotional Virtual Characters in VR Volleyball GamesabstractAbstract Virtual reality (VR) volleyball games can provide an immersive entertainment experience and facilitate users to get familiar with game rules. In real world, players often experience emotions from scores, strokes, passes, and the influence from their teammates or opponents, and so forth. However, seldom studies are carried out for this emotional experience in a virtual environment, as most VR volleyball games mainly concentrate on the game playing, such as body movements. In this article, we propose to enhance emotional experience by building emotional virtual characters in VR volleyball games. The virtual characters cannot only arouse their emotions but also express their facial expressions according to the game situation. Emotion patterns are learned from real‐world volleyball matches. The results demonstrate that our framework has great potential to enhance the user's emotional experience and engagement. Zechen Bai, Nai-Ming Yao, Nidhi Mishra, Hui Chen 0020, Hongan Wang, Nadia Magnenat-Thalmann |
Comput. Animat. Virtual Worlds | 4 |
| 2021 | Deep Learning Methods for Lung Cancer Segmentation in Whole-Slide Histopathology Images - The ACDC@LungHP Challenge 2019abstractAccurate segmentation of lung cancer in pathology slides is a critical step in improving patient care. We proposed the ACDC@LungHP (Automatic Cancer Detection and Classification in Whole-slide Lung Histopathology) challenge for evaluating different computer-aided diagnosis (CADs) methods on the automatic diagnosis of lung cancer. The ACDC@LungHP 2019 focused on segmentation (pixel-wise detection) of cancer tissue in whole slide imaging (WSI), using an annotated dataset of 150 training images and 50 test images from 200 patients. This paper reviews this challenge and summarizes the top 10 submitted methods for lung cancer segmentation. All methods were evaluated using metrics using the precision, accuracy, sensitivity, specificity, and DICE coefficient (DC). The DC ranged from 0.7354 ±0.1149 to 0.8372 ±0.0858. The DC of the best method was close to the inter-observer agreement (0.8398 ±0.0890). All methods were based on deep learning and categorized into two groups: multi-model method and single model method. In general, multi-model methods were significantly better (p 0.01) than single model methods, with mean DC of 0.7966 and 0.7544, respectively. Deep learning based methods could potentially help pathologists find suspicious regions for further analysis of lung cancer in WSI. Tao Tan 0002, Xichao Teng, Xiaoliang Sun, Lihong Liu, Byungjae Lee, Yilong Li 0002, Qianni Zhang, Shujiao Sun, Yushan Zheng, Junyu Yan, Yiyu Hong, Junsu Ko, Hyun Jung, Ching-Wei Wang, Vladimir Yurovskiy, Pavel Maevskikh, Vahid Khanagha, Daiqiang Li, Peter J. Schüffler, Hui Chen 0020, Yuling Tang, Geert Litjens 0001 |
IEEE J. Biomed. Health Informatics | 31 |
| 2020 | A Palette of Deepened Emotions: Exploring Emotional Challenge in Virtual Reality GamesabstractRecent work introduced the notion of 'emotional challenge' promising for understanding more unique and diverse player experiences (PX). Although emotional challenge has immediately attracted HCI researchers' attention, the concept has not been experimentally explored, especially in virtual reality (VR), one of the latest gaming environments. We conducted two experiments to investigate how emotional challenge affects PX when separately from or jointly with conventional challenge in VR and PC conditions. We found that relatively exclusive emotional challenge induced a wider range of different emotions in both conditions, while the adding of emotional challenge broadened emotional responses only in VR. In both experiments, VR significantly enhanced the measured PX of emotional responses, appreciation, immersion and presence. Our findings indicate that VR may be an ideal medium to present emotional challenge and also extend the understanding of emotional (and conventional) challenge in video games. Xiaolan Peng, Jin Huang 0009, Alena Denisova, Hui Chen 0020, Feng Tian 0001, Hongan Wang |
CHI | 4 |
| 2020 | Talking Head-based L2 Pronunciation Training: Impact on Achievement Emotions, Cognitive Load, and Their Relationships with Learning PerformanceabstractSecond language (L2) pronunciation training has been a worldwide task. Although computer technology makes it possible to develop a talking head to teach pronunciation like a real language teacher, little is known about how a talking head may act on L2 learners’ emotional and cognitive learning process. We investigate L2 learners’ achievement emotions, cognitive load, and pronunciation learning performance in a computer-assisted pronunciation training (CAPT) system embedded with four conditions: audio only (AU), a human face (HF), a 3D talking head with front view (3Df), and a 3D talking head with both front and profile views (3D). Results showed that, with learning time went on, participants’ perceived anxiety, boredom, and pride increased while shame and hopelessness decreased and enjoyment kept stable. With 3D, participants’ anxiety increased the most and boredom increased the least. Moreover, 3D group also perceived the highest germane load and got the highest pronunciation learning performance. Furthermore, anxiety and shame correlated with learning performance positively while boredom correlated with it negatively; enjoyment and pride correlated positively with performance on Mandarin tones. These findings significantly contribute to the efforts to design or select virtual characters for computer-aided language learning (CALL) and also provide a valuable reference to study achievement emotions in HCI systems. Xiaolan Peng, Hui Chen 0020, Feng Tian 0001, Hongan Wang |
Int. J. Hum. Comput. Interact. | 2 |
| 2018 | Evaluating a 3-D virtual talking head on pronunciation learning
Xiaolan Peng, Hui Chen 0020, Hongan Wang |
Int. J. Hum. Comput. Stud. | 2 |
| 2018 | Emotional facial expression transfer from a single image via generative adversarial netsabstractAbstract Facial expression transfer from a single image is a challenging task and has drawn sustained attention in the fields of computer vision and computer graphics. Recently, generative adversarial nets (GANs) have provided a new approach to facial expression transfer from a single image toward target facial expressions. However, it is still difficult to obtain a sequence of smoothly changed facial expressions. We present a novel GAN‐based method for generating emotional facial expression animations given a single image and several facial landmarks for the in‐between stages. In particular, landmarks of other subjects are incorporated into a GAN model to control the generated facial expression from a latent space. With the trained model, high‐quality face images and a smoothly changed facial expression sequence can be effectively obtained, which are showed qualitatively and quantitatively in our experiments on the Multi‐PIE and CK+ data sets. Fengchun Qiao, Nai-Ming Yao, Zirui Jiao, Hui Chen 0020, Hongan Wang |
Comput. Animat. Virtual Worlds | 5 |
| 2017 | Non-Frontal Facial Expression Recognition Using a Depth-Patch Based Deep Neural Network
Nai-Ming Yao, Hui Chen 0020, Qingpei Guo, Hongan Wang |
J. Comput. Sci. Technol. | 2 |
| 2017 | Investigations on Mandarin Aspiratory Animations Using an Airflow ModelabstractVarious three-dimensional (3-D) talking heads have been developed lately for language learning, with both external and internal articulatory movements being visualized to guide learning. Mandarin pronunciation animation is challenging due to its confusable stops and affricates with similar places of articulation. Until now, less attention has been paid to the biosignal information of aspiratory airflow, which is essential in distinguishing Mandarin consonants. This study fills a research gap by presenting the quantitative analyses of airflow, and then designing an airflow model for a 3-D pronunciation system. The airflow information was collected by Phonatory Aerodynamic System, so that confusable consonants in Mandarin could be discerned by mean airflow rate, peak airflow rate, airflow duration, and peak time. Based on the airflow parameters, an airflow model using the physical equation of fluid flow was proposed and solved, which was then combined and synchronized with the existing 3-D articulatory model. Therefore, the new multimodal system was implemented to synchronously exhibit the airflow motions and articulatory movements of uttering Mandarin syllables. Both an audio-visual perception test and a pronunciation training study were conducted to assess the effectiveness of our system. Perceptual results indicated that identification accuracy was improved for both native and nonnative groups with the help of airflow motions, while native perceivers exhibited higher accuracy due to long-term language experience. Moreover, our system helped Japanese learners of Mandarin enhance their production skills of Mandarin aspirated consonants, reflected by higher gain values of voice onset time after training. Fei Chen 0005, Hui Chen 0020, Gang Peng 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Intelligible enhancement of 3D articulation animation by incorporating airflow informationabstractThe 3D talking head has been developed fast, in which both external and internal articulators were demonstrated. For Mandarin pronunciation, the aspiration airflow is crucial to discriminate confusable Mandarin consonants. In this paper, we present a 3D talking head system for articulatory and aspiration animation with the use of EMA articulation data and airflow data simultaneously. The quantitative analyses of airflow data indicated confusable Mandarin consonants could be distinguished from each other by the mean airflow during voicing, peak expiratory airflow, and airflow duration. An airflow model was then incorporated into the 3D articulatory model to produce the airflow in accordance with articulator movements of Mandarin pronunciation. An audio-visual test was designed to evaluate the current 3D articulation and aspiration system, where minimal pairs were used to recognize the animation. The identification accuracy was significantly improved from 43.9% without airflow to 84.8% with airflow-incorporated information. Fei Chen 0005, Hui Chen 0020, Jiaying He, Gang Peng 0001 |
ICASSP | 2 |
| 2015 | 3D model-based continuous emotion recognitionabstractWe propose a real-time 3D model-based method that continuously recognizes dimensional emotions from facial expressions in natural communications. In our method, 3D facial models are restored from 2D images, which provide crucial clues for the enhancement of robustness to overcome large changes including out-of-plane head rotations, fast head motions and partial facial occlusions. To accurately recognize the emotion, a novel random forest-based algorithm which simultaneously integrates two regressions for 3D facial tracking and continuous emotion estimation is constructed. Moreover, via the reconstructed 3D facial model, temporal information and user-independent emotion presentations are also taken into account through our image fusion process. The experimental results show that our algorithm can achieve state-of-the-art result with higher Pearson's correlation coefficient of continuous emotion recognition in real time. Hui Chen 0020, Jiangdong Li, Fengjun Zhang, Yang Li 0058, Hongan Wang |
CVPR | 1 |
| 2015 | Emotional Tone-Based Audio Continuous Emotion Recognition
Hui Chen 0020, Yang Li 0058, Fengjun Zhang |
MMM (2) | 2 |
| 2014 | Left and right hand distinction for multi-touch tabletop interactionsabstractIn multi-touch interactive systems, it is of great significance to distinguish which hand of the user is touching the surface in real time. Left-right hand distinction is essential for recognizing the multi-finger gestures and further fully exploring the potential of bimanual interaction. However, left-right hand distinction is beyond the capability of most existing multi-touch systems. In this paper, we present a new method for left and right hand distinction based on the human anatomy, work area, finger orientation and finger position. Considering the ergonomics principles of gesture designing, the body-forearm triangle model was proposed. Furthermore, a heuristic algorithm was introduced to group multi-touch contact points and then made left-right hand distinction. A dataset of 2880 images has been set up to evaluate the proposed left-right hand distinction method. The experimental results demonstrate that our method can guarantee the high recognition accuracy and real time performance in freely bimanual multi-touch interactions. Zhensong Zhang, Fengjun Zhang, Hui Chen 0020, Jiasheng Liu, Hongan Wang, Guozhong Dai |
IUI | 3 |
| 2012 | A virtual surgical simulator for mandibular angle reduction based on patient specific dataabstractIn our work, a virtual reality-based surgical simulator for the mandibular angle reduction was designed and implemented on CUDA-based platform. High-fidelity visual and haptic feedbacks between the surgical instruments and the bone material are provided to enhance the perception in a realistic virtual surgical environment. Impulse-based dynamics haptic model was employed to simulate the contact forces generated on the high-speed instruments, including the reciprocating saw and the round burr. The validity of the simulated contact forces was verified by comparing against the actual force data measured through the constructed mechanical platform. An empirical study based on the patient specified data was conducted to evaluate the ability of the proposed system in training surgeons with various experiences. The results confirm the validity of our simulator. Qiong Wang 0001, Hui Chen 0020, Wen Wu 0001, Hai-yang Jin, Pheng-Ann Heng |
VR | 2 |
| 2012 | Phoneme-level articulatory animation in pronunciation training
Hui Chen 0020, Sheng Li 0010, Helen M. Meng |
Speech Commun. | 2 |
| 2012 | Real-Time Mandibular Angle Reduction Surgical Simulation With Haptic RenderingabstractMandibular angle reduction is a popular and efficient procedure widely used to alter the facial contour. The primary surgical instruments, the reciprocating saw and the round burr, employed in the surgery have a common feature: operating at a high-speed. Generally, inexperienced surgeons need a long-time practice to learn how to minimize the risks caused by the uncontrolled contacts and cutting motions in manipulation of instruments with high-speed reciprocation or rotation. A virtual reality-based surgical simulator for the mandibular angle reduction was designed and implemented on a CUDA-based platform in this paper. High-fidelity visual and haptic feedbacks are provided to enhance the perception in a realistic virtual surgical environment. The impulse-based haptic models were employed to simulate the contact forces and torques on the instruments. It provides convincing haptic sensation for surgeons to control the instruments under different reciprocation or rotation velocities. The real-time methods for bone removal and reconstruction during surgical procedures have been proposed to support realistic visual feedbacks. The simulated contact forces were verified by comparing against the actual force data measured through the constructed mechanical platform. An empirical study based on the patient-specific data was conducted to evaluate the ability of the proposed system in training surgeons with various experiences. The results confirm the validity of our simulator. Qiong Wang 0001, Hui Chen 0020, Wen Wu 0001, Hai-yang Jin, Pheng-Ann Heng |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2010 | Combined X-ray and facial videos for phoneme-level articulator dynamics
Hui Chen 0020, Wenxi Liu, Pheng-Ann Heng |
Vis. Comput. | 1 |
| 2009 | Evaluation of external and internal articulator dynamics for pronunciation learning
Hui Chen 0020, JianJun Ouyang |
INTERSPEECH | 2 |