Jenny Sheng

dblp:358/5028 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2025
0009-0007-7863-8409ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Exploring the Influence of Profile Picture Styles on Empathy and Identity Recognition in Social Media
abstract
Empathy and identity recognition are two core social interaction factors that greatly affect efficiency and effectiveness. The selection of a profile picture in social media is crucial as it serves as a visual representation of virtual identity. It not only reflects the user's personality but also influences the level of empathy and identification that others feel toward them. However, the potential impact of profile pictures with different styles (e.g., real faces, cartoon faces, and landscape images) on empathy and identity recognition is still unclear. To explore its effects, a controlled laboratory experiment and an ecological online experiment were conducted. Participants were shown a picture each time, informed to imagine interacting with the person using it as his/her profile picture, and instructed to rate an item from the basic empathy scale (BES) based on it. After rating all pictures, users then completed an identity recognition task. Results show that participants’ empathy scores for users with cartoon or real face profile pictures are greater than those with landscape profile pictures. In addition, participants performed better in identity recognition for users with real face or landscape images as profile pictures than for those with cartoon face profile pictures. Moreover, users of social media often make social categorizations (i.e., in-group/out-group categorization) based on the social identities expressed by their profile pictures. Our results also indicate that the affective empathy scale rating score was positively associated with the degree to which users of the corresponding profile pictures were categorized as in-group members.
Minjing Yu, Xinge Liu, Chao Zhou 0012, Xinxin Du, Jenny Sheng, Yong-Jin Liu 0001
IEEE Trans. Comput. Soc. Syst.5
2025 PCKRF: Point Cloud Completion and Keypoint Refinement With Fusion Data for 6D Pose Estimation
abstract
Some robust point cloud registration approaches with controllable pose refinement magnitude, such as ICP and its variants, are commonly used to improve 6D pose estimation accuracy. However, the effectiveness of these methods gradually diminishes with the advancement of deep learning techniques and the enhancement of initial pose accuracy, primarily due to their lack of specific design for pose refinement. In this paper, we propose Point Cloud Completion and Keypoint Refinement with Fusion Data (PCKRF), a new pose refinement pipeline for 6D pose estimation. The pipeline consists of two steps. First, it completes the input point clouds via a novel pose-sensitive point completion network. The network uses both local and global features with pose information during point completion. Then, it registers the completed object point cloud with the corresponding target point cloud by our proposed Color supported Iterative KeyPoint (CIKP) method. The CIKP method introduces color information into registration and registers a point cloud around each keypoint to increase stability. The PCKRF pipeline can be integrated with existing popular 6D pose estimation methods, such as the full flow bidirectional fusion network, to further improve their pose estimation accuracy. Experiments demonstrate that our method exhibits superior stability compared to existing approaches when optimizing initial poses with relatively high precision. Notably, the results indicate that our method effectively complements most existing pose estimation techniques, leading to improved performance in most cases. Furthermore, our method achieves promising results even in challenging scenarios involving textureless and symmetrical objects.
Yiheng Han, Irvin Haozhe Zhan, Long Zeng 0001, Yu-Ping Wang 0001, Ran Yi 0002, Minjing Yu, Matthieu Lin, Jenny Sheng, Yong-Jin Liu 0001
IEEE Trans. Vis. Comput. Graph.8
2024 Exploring Temporal Feature Correlation for Efficient and Stable Video Semantic Segmentation
abstract
This paper tackles the problem of efficient and stable video semantic segmentation. While stability has been under-explored, prevalent work in efficient video semantic segmentation uses the keyframe paradigm. They efficiently process videos by only recomputing the low-level features and reusing high-level features computed at selected keyframes. In addition, the reused features stabilize the predictions across frames, thereby improving video consistency. However, dynamic scenes in the video can easily lead to misalignments between reused and recomputed features, which hampers performance. Moreover, relying on feature reuse to improve prediction consistency is brittle; an erroneous alignment of the features can easily lead to unstable predictions. Therefore, the keyframe paradigm exhibits a dilemma between stability and performance. We address this efficiency and stability challenge using a novel yet simple Temporal Feature Correlation (TFC) module. It uses the cosine similarity between two frames’ low-level features to inform the semantic label’s consistency across frames. Specifically, we selectively reuse label-consistent features across frames through linear interpolation and update others through sparse multi-scale deformable attention. As a result, we no longer directly reuse features to improve stability and thus effectively solve feature misalignment. This work provides a significant step towards efficient and stable video semantic segmentation. On the VSPW dataset, our method significantly improves the prediction consistency of image-based methods while being as fast and accurate.
Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Yangguang Li 0001, Andrew Zhao, Gao Huang 0001, Yong-Jin Liu 0001
AAAI2
2024 ECAvatar: 3D Avatar Facial Animation with Controllable Identity and Emotion
Minjing Yu, Delong Pang, Ziwen Kang, Zhiyao Sun, Tian Lv, Jenny Sheng, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001
ACM Multimedia6
2024 VisHanfu: An Interactive System for the Promotion of Hanfu Knowledge via Cross-Shaped Flat Structure
Minjing Yu, Lingzhi Zeng, Xinxin Du, Jenny Sheng, Qiantian Liao, Yong-Jin Liu 0001
ACM Multimedia4
2024 Text-image conditioned diffusion for consistent text-to-3D generation
Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu 0001, Qi Wang 0079, Yu-Hui Wen, Yong-Jin Liu 0001
Comput. Aided Geom. Des.4
2024 Focus on Cooperation: A Face-to-Face VR Serious Game for Relationship Enhancement
abstract
Exploring effective approaches to enhance face-to-face interactions and interpersonal relationships is an important topic in the applications of affective computing. According to the co-actualization model, we propose a face-to-face co-participation serious game for relationship enhancement, with a focus on battling COVID-19. Moreover, a prototype system is developed using an immersive virtual environment and a low-cost brain-computer interface. Through this system, a dynamic flow experience enhancement tool is utilized to involve partners in the cooperative task. To evaluate the system performance, two studies are conducted with schoolmates as participants. Study 1 compares the cooperative and competitive modes, and demonstrates that the former elicited higher level of decision-making challenge and affections, which are beneficial for forming relationships. Study 2 further examines the effect of the dynamic flow enhancement tool in the cooperative task and the results show its effectiveness in promoting flow experience, perceived closeness, and intimacy in relationships. Given this short-term participation, participants felt a greater sense of closeness and intimacy than they had before the test. In conclusion, our proposed system is effective in enhancing schoolmate relationships.
Yulong Bian, Chao Zhou 0012, Yang Zhang 0116, Juan Liu 0008, Jenny Sheng, Yong-Jin Liu 0001
IEEE Trans. Affect. Comput.5
2024 Emotion Recognition From Few-Channel EEG Signals by Integrating Deep Feature Aggregation and Transfer Learning
abstract
Electroencephalogram (EEG) signals have been widely studied in human emotion recognition. The majority of existing EEG emotion recognition algorithms utilize dozens or hundreds of electrodes covering the whole scalp region (denoted as full-channel EEG devices in this paper). Nowadays, more and more portable and miniature EEG devices with only a few electrodes (denoted as few-channel EEG devices in this paper) are emerging. However, emotion recognition from few-channel EEG data is challenging because the device can only capture EEG signals from a portion of the brain area. Moreover, existing full-channel algorithms cannot be directly adapted to few-channel EEG signals due to the significant inter-variation between full-channel and few-channel EEG devices. To address these challenges, we propose a novel few-channel EEG emotion recognition framework from the perspective of knowledge transfer. We leverage full-channel EEG signals to provide supplementary information for few-channel signals via a transfer learning-based model CD-EmotionNet, which consists of a base emotion model for efficient emotional feature extraction and a cross-device transfer learning strategy. This strategy helps to enhance emotion recognition performance on few-channel EEG data by utilizing knowledge learned from full-channel EEG data. To evaluate our cross-device EEG emotion transfer learning framework, we construct an emotion dataset containing paired 18-channel and 5-channel EEG signals from 25 subjects, as well as 5-channel EEG signals from 13 other subjects. Extensive experiments show that our framework outperforms state-of-the-art EEG emotion recognition methods by a large margin.
Fang Liu 0035, Yezhi Shu, Niqi Liu, Jenny Sheng, Xiaoan Wang, Yong-Jin Liu 0001
IEEE Trans. Affect. Comput.5
2024 DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models
abstract
The generation of stylistic 3D facial animations driven by speech presents a significant challenge as it requires learning a many-to-many mapping between speech, style, and the corresponding natural facial motion. However, existing methods either employ a deterministic model for speech-to-motion mapping or encode the style using a one-hot encoding scheme. Notably, the one-hot encoding approach fails to capture the complexity of the style and thus limits generalization ability. In this paper, we propose DiffPoseTalk, a generative framework based on the diffusion model combined with a style encoder that extracts style embeddings from short reference videos. During inference, we employ classifier-free guidance to guide the generation process based on the speech and style. In particular, our style includes the generation of head poses, thereby enhancing user perception. Additionally, we address the shortage of scanned 3D talking face data by training our model on reconstructed 3DMM parameters from a high-quality, in-the-wild audio-visual dataset. Extensive experiments and user study demonstrate that our approach outperforms state-of-the-art methods. The code and dataset are at https://diffposetalk.github.io.
Zhiyao Sun, Tian Lv, Matthieu Lin, Jenny Sheng, Yu-Hui Wen, Minjing Yu, Yong-Jin Liu 0001
ACM Trans. Graph.5
2024 PVP-Recon: Progressive View Planning via Warping Consistency for Sparse-View Surface Reconstruction
abstract
Neural implicit representations have revolutionized dense multi-view surface reconstruction, yet their performance significantly diminishes with sparse input views. A few pioneering works have sought to tackle this challenge by leveraging additional geometric priors or multi-scene generalizability. However, they are still hindered by the imperfect choice of input views, using images under empirically determined viewpoints. We propose PVP-Recon , a novel and effective sparse-view surface reconstruction method that progressively plans the next best views to form an optimal set of sparse viewpoints for image capturing. PVP-Recon starts initial surface reconstruction with as few as 3 views and progressively adds new views which are determined based on a novel warping score that reflects the information gain of each newly added view. This progressive view planning progress is interleaved with a neural SDF-based reconstruction module that utilizes multi-resolution hash features, enhanced by a progressive training scheme and a directional Hessian loss. Quantitative and qualitative experiments on three benchmark datasets show that our system achieves high-quality reconstruction with a constrained input budget and outperforms existing baselines.
Matthieu Lin, Jenny Sheng, Ruoyu Fan, Yiheng Han, Yubin Hu 0001, Ran Yi 0002, Yu-Hui Wen, Yong-Jin Liu 0001, Wenping Wang 0001
ACM Trans. Graph.4