Shuai Shen

dblp:236/3044 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Arbitrary-Scale Super-Resolution With Compact Gaussian Splatting
abstract
Arbitrary-scale super-resolution is an essential image upsampling task, typically tackled with implicit neural representations. Unlike INR-based methods that rely on slow per-pixel decoding, Gaussian Splatting is promising for arbitrary-scale super-resolution when the explicit region-based nature of GS allows for highly efficient rendering via a lightweight decoder. Recently, Gaussian splatting has outperformed implicit neural methods in 3D scenes. However, existing attempts to address the super-resolution problem with Gaussian splatting face efficiency and accuracy challenges, such as redundant Gaussian primitives and discrete pixel sampling. Efficient Gaussian applications may result in continuous texture constraints due to limited feature richness in explicit fields, particularly the mismatch between the learned Gaussian fields and out-of-distribution sampling rates. Moreover, insufficient discrete sampling based on the given upscale factor may fail to accurately represent the splatted Gaussian field in screen space, causing high-frequency signal redundancy and aliasing. To address these challenges, we propose a compact Gaussian splatting method for efficient arbitrary-scale super-resolution, CGSSR. It constructs an efficient Gaussian space by distilling the image into a reduced number of primitives, each represented by a compact, low-dimensional feature embedding. To balance detail preservation and anti-aliasing, we introduce a scale-aware smoothing filter to regulate splatting frequency. Extensive experiments show that CGSSR achieves superior performance over existing Gaussian-based methods, especially at large scales, with higher efficiency.
Jingyi Zhang 0008, Jiajun Dong, Shuai Shen, Yansong Tang, Lei Chen 0069, Jiwen Lu
IEEE Trans. Circuits Syst. Video Technol.4
2026 HP-Gaussian: Head Prior-Guided Gaussian Splatting for Personalized Talking Head Synthesis From Few-Second Video
abstract
Gaussian Splatting-based talking head synthesis has made significant progress in recent years, yet existing methods often struggle with generalization beyond specific training identity. In this paper, we propose Head Prior guided Gaussian Splatting for personalized talking head synthesis (HP-Gaussian) that can generalize to new identities with only few training data. Unlike traditional optimization-based Gaussian Splatting methods, our approach directly predicts Gaussian parameters from multi-modal inputs, including audio and visual cues. This feed-forward design enables multiple identities pre-training, allowing the model to learn shared head priors from large-scale datasets, while supporting flexible speaker-specific adaptation. To further enhance Gaussian feature learning, we introduce a Spatial Gaussian Transformer that captures correlations among neighboring Gaussians, improving parameter estimation accuracy. Additionally, recognizing the critical importance of personalized speaking styles in the synthesis of high-quality talking videos, a two-stage training strategy is implemented. A base model is initially trained across diverse identities to establish the foundational head prior knowledge. Subsequently, we introduce the short-video personalized adaptation phase for more realistic customized talking video generation. Extensive experiments demonstrate that our HP-Gaussian can synthesize high-fidelity and personalized talking videos with remarkably few training examples, setting a new benchmark for efficiency and quality in talking head synthesis. We highly recommend viewing our demonstration video at https://youtu.be/RpjWdvikKhU for intuitive visual comparisons and qualitative results.
Shuai Shen, Wanhua Li 0001, Weipeng Hu, Jiwen Lu, Yap-Peng Tan
IEEE Trans. Image Process.1
2025 E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
Ronghao Lin, Shuai Shen, Weipeng Hu, Qiaolin He, Aolin Xiong, Haifeng Hu 0001, Yap-Peng Tan
ACM Multimedia2
2024 SD-NeRF: Towards Lifelike Talking Head Animation via Spatially-Adaptive Dual-Driven NeRFs
abstract
Recent years have witnessed great progress in audio-driven talking head animation. Among these methods, the 3D-based ones better preserve the 3D consistency of the generated head and produce more natural results compared with 2D-based approaches. However, most 3D-based methods employ 3D morphable face models as the intermediate representation and involve multi-stage training, which may lead to error accumulation. To alleviate this problem, in this article, we propose a fully end-to-end talking head animation method, which implicitly grasps the 3D structures by learning a conditional Neural Radiance Field (NeRF). As NeRF has proven to be an effective tool for 3D modeling, one can learn dynamic neural radiance fields conditioned on audio signals for talking head synthesis. Furthermore, we argue that audio signals cannot fully drive a lifelike talking head. When people are talking, they usually show many spontaneous facial movements like blinks and brow movements, which makes talkers natural and real. These movements cannot be fully driven by the audio signals since they are highly unrelated to the audio. Therefore, we incorporate motion information as another driving factor and develop an audio-motion dual-driven NeRF model to take a step toward more lifelike talking head synthesis. On this basis, as audio and motion mainly affect different regions of the human face, we propose a Spatially-adaptive Dual-driven NeRF (SD-NeRF), which fuses these two driven factors with a spatially-adaptive cross-attention mechanism. Quantitative and qualitative results demonstrate that, with finer facial controls, our method produces more realistic talking head videos compared with existing advanced works.
Shuai Shen, Wanhua Li 0001, Xiaoke Huang 0001, Jie Zhou 0001, Jiwen Lu
IEEE Trans. Multim.1
2023 DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits Animation
abstract
Talking head synthesis is a promising approach for the video production industry. Recently, a lot of effort has been devoted in this research area to improve the generation quality or enhance the model generalization. However, there are few works able to address both issues simultaneously, which is essential for practical applications. To this end, in this paper, we turn attention to the emerging powerful Latent Diffusion Models, and model the Talking head generation as an audio-driven temporally coherent denoising process (DiffTalk). More specifically, instead of employing audio signals as the single driving factor, we investigate the control mechanism of the talking face, and incorporate reference face images and landmarks as conditions for personality-aware generalized synthesis. In this way, the proposed DiffTalk is capable of producing high-quality talking head videos in synchronization with the source audio, and more importantly, it can be naturally generalized across different identities without further finetuning. Additionally, our DiffTalk can be gracefully tailored for higher-resolution synthesis with negligible extra computational cost. Extensive experiments show that the proposed DiffTalk efficiently synthesizes high-fidelity audio-driven talking head videos for generalized novel identities. For more video results, please refer to https://sstzal.github.io/DiffTalk/.
Shuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 0001, Jie Zhou 0001, Jiwen Lu
CVPR1
2023 CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face Clustering
abstract
One of the most important yet rarely studied challenges for supervised face clustering is the large intra-class variance caused by different face attributes such as age, pose, and expression. Images of the same identity but with different face attributes usually tend to be clustered into different sub-clusters. For the first time, we proposed an attribute hallucination framework named CLIP-Cluster to address this issue, which first hallucinates multiple representations for different attributes with the powerful CLIP model and then pools them by learning neighbor-adaptive attention. Specifically, CLIP-Cluster first introduces a text-driven attribute hallucination module, which allows one to use natural language as the interface to hallucinate novel attributes for a given face image based on the well-aligned image-language CLIP space. Furthermore, we develop a neighbor-aware proxy generator that fuses the features describing various attributes into a proxy feature to build a bridge among different sub-clusters and reduce the intra-class variance. The proxy feature is generated by adaptively attending to the hallucinated visual features and the source one based on the local neighbor information. On this basis, a graph built with the proxy representations is used for subsequent clustering operations. Extensive experiments show our proposed approach outperforms state-of-the-art face clustering methods with high inference efficiency.
Shuai Shen, Wanhua Li 0001, Dafeng Zhang, Zhezhu Jin, Jie Zhou 0001, Jiwen Lu
ICCV1
2023 STAR-FC: Structure-Aware Face Clustering on Ultra-Large-Scale Graphs
abstract
Face clustering is a promising method for annotating unlabeled face images. Recent supervised approaches have boosted the face clustering accuracy greatly, however their performance is still far from satisfactory. These methods can be roughly divided into global-based and local-based ones. Global-based methods suffer from the limitation of training data scale, while local-based ones are inefficient for inference due to the use of numerous overlapped subgraphs. Previous approaches fail to tackle these two challenges simultaneously. To address the dilemma of large-scale training and efficient inference, we propose the STructure-AwaRe Face Clustering (STAR-FC) method. Specifically, we design a structure-preserving subgraph sampling strategy to explore the power of large-scale training data, which can increase the training data scale from${10^{5}}$to${10^{7}}$. On this basis, a novel hierarchical GCN training paradigm is further proposed for better capturing the dynamic local structure. During inference, the STAR-FC performs efficient full-graph clustering with two steps: graph parsing and graph refinement. And the concept of node intimacy is introduced in the second step to mine the local structural information, where a calibration module is further proposed for fairer edge scores. The STAR-FC gets 93.21 pairwise F-score on standard partial MS1M within 312 seconds, which far surpasses the state-of-the-arts while maintaining high inference efficiency. Furthermore, we are the first to train on an ultra-large-scale graph with 20 M nodes, and achieve superior inference results on 12 M testing data. Overall, as a simple and effective method, the proposed STAR-FC provides a strong baseline for large-scale face clustering. Code is available inhttps://github.com/sstzal/STAR-FC.
Shuai Shen, Wanhua Li 0001, Jie Zhou 0001, Jiwen Lu
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis
Shuai Shen, Wanhua Li 0001, Yueqi Duan, Jie Zhou 0001, Jiwen Lu
ECCV (12)1
2022 Number and Operation Time Minimization for Multi-UAV-Enabled Data Collection System With Time Windows
abstract
In this article, we investigate multiple unmanned aerial vehicles (UAVs)-enabled data collection system in Internet of Things (IoT) networks with time windows, where multiple rotary-wing UAVs are dispatched to collect data from time-constrained terrestrial IoT devices. We aim to jointly minimize the number and the total operation time of UAVs by optimizing the UAV trajectory and hovering location. To this end, an optimization problem is formulated, considering the energy budget and cache capacity of UAVs as well as the data transmission constraint of IoT devices. To tackle this mix-integer nonconvex problem, we decompose the problem into two subproblems: 1) UAV trajectory and 2) hovering location optimization problems. To solve the first subproblem, an modified ant colony optimization (MACO) algorithm is proposed. For the second subproblem, the successive convex approximation (SCA) technique is applied. Then, an overall algorithm, termed the MACO-based algorithm, is given by leveraging the MACO algorithm and SCA technique. Simulation results demonstrate the superiority of the proposed algorithm.
Shuai Shen, Kun Yang 0001, Kezhi Wang, Guopeng Zhang, Haibo Mei
IEEE Internet Things J.1
2021 Structure-Aware Face Clustering on a Large-Scale Graph With 107 Nodes
abstract
Face clustering is a promising method for annotating un-labeled face images. Recent supervised approaches have boosted the face clustering accuracy greatly, however their performance is still far from satisfactory. These methods can be roughly divided into global-based and local-based ones. Global-based methods suffer from the limitation of training data scale, while local-based ones are difficult to grasp the whole graph structure information and usually take a long time for inference. Previous approaches fail to tackle these two challenges simultaneously. To address the dilemma of large-scale training and efficient inference, we propose the STructure-AwaRe Face Clustering (STAR-FC) method. Specifically, we design a structure-preserved subgraph sampling strategy to explore the power of large-scale training data, which can increase the training data scale from 105to 107. During inference, the STAR-FC performs efficient full-graph clustering with two steps: graph parsing and graph refinement. And the concept of node intimacy is introduced in the second step to mine the local structural information. The STAR-FC gets 91.97 pairwise F-score on partial MS1M within 310s which surpasses the state-of-the-arts. Furthermore, we are the first to train on very large-scale graph with 20M nodes, and achieve superior inference results on 12M testing data. Overall, as a simple and effective method, the proposed STAR-FC provides a strong baseline for large-scale face clustering. Code is available at https://sstzal.github.io/STAR-FC/.
Shuai Shen, Wanhua Li 0001, Guan Huang 0003, Dalong Du, Jiwen Lu, Jie Zhou 0001
CVPR1