Su Zhaoen

dblp:386/3161 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 33% Face, body and person analysis · 29% Representation and self-supervised learning · 29%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 3 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
0.912025
Pippo: High-Resolution Multi-View Humans from a Single Image · CVPR 2025
Visual content generation and editing
3d content generation
0.912025
Pippo: High-Resolution Multi-View Humans from a Single Image · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised visual pre-training
0.812024
Sapiens: Foundation for Human Vision Models · ECCV (4) 2024

Methods — techniques the papers use, named apart from their topics

plücker ray encoding · 1.7pixel-aligned control · 1.7attention biasing · 1.7self-supervised learning · 0.8
YearPublicationVenuePosition
2025 Pippo: High-Resolution Multi-View Humans from a Single Image
abstract
We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs — e.g., a fitted parametric model or camera parameters of the input image. We pre-train Pippo on 3B human images without captions, and conduct multi-view mid-training and post-training on studio captured humans. During mid-training, to quickly absorb the studio dataset, we denoise several (upto 48) views at low-resolution, and encode target cameras coarsely using a shallow MLP. During post-training, we denoise fewer views at high-resolution and use pixel-aligned controls (e.g., Spatial anchor and Plucker rays) to enable 3D consistent generations. At inference, we propose an attention biasing technique that allows Pippo to simultaneously generate greater than 5× as many views as seen during training. Finally, we also introduce an improved metric to evaluate 3D consistency of multi-view generations, and show that Pippo outperforms existing works on multi-view human generation from a single image.
Yash Kant, Ethan Weber, Jin Kyu Kim, Rawal Khirodkar, Su Zhaoen, Julieta Martinez 0001, Igor Gilitschenski, Shunsuke Saito, Timur M. Bagautdinov
CVPR5
2024 Sapiens: Foundation for Human Vision Models
Rawal Khirodkar, Timur M. Bagautdinov, Julieta Martinez 0001, Su Zhaoen, Austin James, Peter Selednik, Shunsuke Saito
ECCV (4)4