Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shuyuan Tu

dblp:322/8490 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 58% Face, body and person analysis · 15% Video understanding and tracking · 13%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 50% Computer animation and physical simulation · 50%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.032025
MotionFollower: Editing Video Motion via Score-Guided Diffusion · ICCV 2025
StableAnimator: High-Quality Identity-Preserving Human Image Animation · CVPR 2025
MotionEditor: Editing Video Motion via Content-Aware Diffusion · CVPR 2024
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
StableAnimator: High-Quality Identity-Preserving Human Image Animation · CVPR 2025
Visual content generation and editing › video editing
video motion editing
0.912025
MotionFollower: Editing Video Motion via Score-Guided Diffusion · ICCV 2025
Computer animation and physical simulation
motion editing
0.812024
MotionEditor: Editing Video Motion via Content-Aware Diffusion · CVPR 2024
Visual content generation and editing
video editing
0.812024
MotionEditor: Editing Video Motion via Content-Aware Diffusion · CVPR 2024
Machine learning › Representation and self-supervised learning
contrastive learning
0.712023
Implicit Temporal Modeling with Learnable Alignment for Video Recognition · ICCV 2023
Computer vision › Video understanding and tracking
video classification
0.712023
Implicit Temporal Modeling with Learnable Alignment for Video Recognition · ICCV 2023
Computer vision › Face, body and person analysis › face recognition › face representation
face embedding
0.312025
StableAnimator: High-Quality Identity-Preserving Human Image Animation · CVPR 2025
Computer vision › Face, body and person analysis
face recognition
0.312025
StableAnimator: High-Quality Identity-Preserving Human Image Animation · CVPR 2025
Computer vision › Face, body and person analysis
human pose
0.212024
MotionEditor: Editing Video Motion via Content-Aware Diffusion · CVPR 2024

Methods — techniques the papers use, named apart from their topics

score-guided diffusion · 1.7hamilton-jacobi-bellman equation-based optimization · 1.7face encoder · 1.7distribution-aware ID adapter · 1.7skeleton alignment · 1.5controlnet · 1.5content-aware motion adapter · 1.5attention injection · 1.5learnable alignment · 0.7contrastive language-image pretraining · 0.7
YearPublicationVenuePosition
2025 StableAnimator: High-Quality Identity-Preserving Human Image Animation
abstract
Current diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses. Building upon a video diffusion model, StableAnimator contains carefully designed modules for both training and inference striving for identity consistency. In particular, StableAnimator begins by computing image and face embeddings with off-the-shelf extractors, respectively and face embeddings are further refined by interacting with image embeddings using a global content-aware Face Encoder. Then, StableAnimator introduces a novel distribution-aware ID Adapter that prevents interference caused by temporal layers while preserving ID via alignment. During inference, we propose a novel Hamilton-Jacobi-Bellman (HJB) equation-based optimization to further enhance the face quality. We demonstrate that solving the HJB equation can be integrated into the diffusion denoising process, and the resulting solution constrains the denoising path and thus benefits ID preservation. Experiments on multiple benchmarks show the effectiveness of StableAnimator both qualitatively and quantitatively.
Shuyuan Tu, Xintong Han, Zhi-Qi Cheng, Qi Dai 0001, Chong Luo 0001, Zuxuan Wu
CVPR1
2025 MotionFollower: Editing Video Motion via Score-Guided Diffusion
Shuyuan Tu, Qi Dai 0001, Sicheng Xie, Zhi-Qi Cheng, Chong Luo 0001, Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001
ICCV1
2024 MotionEditor: Editing Video Motion via Content-Aware Diffusion
abstract
Existing diffusion-based video editing models have made gorgeous advances for editing attributes of a source video over time but struggle to manipulate the motion information while preserving the original protagonist's appearance and background. To address this, we propose MotionEditor, the first diffusion model for video motion editing. MotionEditor incorporates a novel content-aware motion adapter into ControlNet to capture temporal motion correspondence. While ControlNet enables direct generation based on skeleton poses, it encounters challenges when modifying the source motion in the inverted noise due to contradictory signals between the noise (source) and the condition (reference). Our adapter complements Control-Net by involving source content to transfer adapted control signals seamlessly. Further, we build up a two-branch ar-chitecture (a reconstruction branch and an editing branch) with a high-fidelity attention injection mechanism facilitating branch interaction. This mechanism enables the editing branch to query the key and value from the reconstruction branch in a decoupled manner, making the editing branch retain the original background and protagonist appearance. We also propose a skeleton alignment algorithm to address the discrepancies in pose size and position. Experiments demonstrate the promising motion editing ability of MotionEditor, both qualitatively and quantitatively. To the best of our knowledge, MotionEditor is the first to use diffusion models specifically for video motion editing, considering the origin dynamic background and camera movement.
Shuyuan Tu, Qi Dai 0001, Zhi-Qi Cheng, Han Hu 0001, Xintong Han, Zuxuan Wu, Yu-Gang Jiang 0001
CVPR1
2024 ASKDetector: An AST-Semantic and Key Features Fusion based Code Comment Mismatch Detector
abstract
Code comments are essential for programming comprehension. Nevertheless, developers often neglect to update comments after modifying the source code. Wrong code comments may lead to bugs in the maintenance process, thus affecting the reliability of the software. So, timely comment mismatch detection is crucial for software development and maintenance. However, existing works have the following two limitations: 1) the lack of use of code structural and sequential information, and 2) the ignorance of existing associations between code and comments. In this paper, we propose a new model called ASKDetector (AST-Semantic and Key features fusion based mismatch Detector). For the first limitation, we encode code with an attention-based preorder traversal abstract syntax tree sequence to obtain both order and structural information. And CodeBERT is utilized to capture contextual semantic features further. For the second one, we encode extracted association information between the code snippets and comments to reduce the semantic gap. The correlations between the encoders are learned through a fusion layer and a multi-layer perceptron. The experimental results prove that our detector outperforms the state-of-the-art model in evaluation metrics, where our F1 and accuracy exceed an average of 3.4%.
Haiyang Yang, Hao Chen 0116, Zhirui Kuai, Shuyuan Tu, Li Kuang
ICPC4
2023 Implicit Temporal Modeling with Learnable Alignment for Video Recognition
abstract
Contrastive language-image pretraining (CLIP) has demonstrated remarkable success in various image tasks. However, how to extend CLIP with effective temporal modeling is still an open and crucial problem. Existing factorized or joint spatial-temporal modeling trades off between the efficiency and performance. While modeling temporal information within straight through tube is widely adopted in literature, we find that simple frame alignment already provides enough essence without temporal attention. To this end, in this paper, we proposed a novel Implicit Learnable Alignment (ILA) method, which minimizes the temporal modeling effort while achieving incredibly high performance. Specifically, for a frame pair, an interactive point is predicted in each frame, serving as a mutual information rich region. By enhancing the features around the interactive point, two frames are implicitly aligned. The aligned features are then pooled into a single token, which is leveraged in the subsequent spatial self-attention. Our method allows eliminating the costly or insufficient temporal self-attention in video. Extensive experiments on benchmarks demonstrate the superiority and generality of our module. Particularly, the proposed ILA achieves a top-1 accuracy of 88.7% on Kinetics-400 with much fewer FLOPs compared with Swin-L and ViViT-H. Code is released at https://github.com/Francis-Rings/ILA.
Shuyuan Tu, Qi Dai 0001, Zuxuan Wu, Zhi-Qi Cheng, Han Hu 0001, Yu-Gang Jiang 0001
ICCV1
2022 Multiple Biological Granularities Network for Person Re-Identification
abstract
The task of person re-identification is to retrieve images of a specific pedestrian among cross-camera person gallery captured in the wild. Previous approaches commonly concentrate on the whole person images and local pre-defined body parts, which are ineffective with diversity of person poses and occlusion. In order to alleviate the problem, researchers began to implement attention mechanisms to their model using local convolutions with limited fields. However, previous attention mechanisms focus on the local feature representations ignoring the exploration of global spatial relation knowledge. The global spatial relation knowledge contains clustering-like topological information which is helpful for overcoming the situation of diversity of person poses and occlusion. In this paper, we propose the Multiple Biological Granularities Network (MBGN) based on Global Spatial Relation Pixel Attention (GSRPA) taking the human body structure and global spatial relation pixels information into account. First, we design an adaptive adjustment algorithm (AABS) based on human body structure, which is complementary to our MBGN. Second, we propose a feature fusion strategy taking multiple biological granularities into account. Our strategy forces the model to learn diversity of person poses by balancing the local semantic human body parts and global spatial relations. Third, we propose the attention mechanism GSRPA. GSRPA enhances the weight of spatial relational pixels, which digs out the person topological information for overcoming occlusion problem. Extensive evaluations on the popular datasets Market-1501 and CUHK03 demonstrate the superiority of MBGN over the state-of-the-art methods.
Shuyuan Tu, Tianzhen Guan, Li Kuang
ICMR1