EDBT 2026 Demo / reviewers in the wild / expert
Yu Sun 0030
dblp:62/3689-30
· DBLP profile ↗
10ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-3525-2753ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PromptHMR: Promptable Human Mesh RecoveryabstractHuman pose and shape (HPS) estimation presents challenges in diverse scenarios such as crowded scenes, person-person interactions, and single-view reconstruction. Existing approaches lack mechanisms to incorporate auxiliary "side information" that could enhance reconstruction accuracy in such challenging scenarios. Furthermore, the most accurate methods rely on cropped person detections and cannot exploit scene context while methods that process the whole image often fail to detect people and are less accurate than methods that use crops. While recent language-based methods explore HPS reasoning through large language or vision-language models, their metric accuracy is well below the state of the art. In contrast, we present PromptHMR, a transformer-based promptable method that reformulates HPS estimation through spatial and semantic prompts. Our method processes full images to maintain scene context and accepts multiple input modalities: spatial prompts like bounding boxes and masks, and semantic prompts like language descriptions or interaction labels. PromptHMR demonstrates robust performance across challenging scenarios: estimating people from bounding boxes as small as faces in crowded scenes, improving body shape estimation through language descriptions, modeling person-person interactions, and producing temporally coherent motions in videos. Experiments on benchmarks show that PromptHMR achieves state-of-the-art performance while offering flexible prompt-based control over the HPS estimation process. Yufu Wang, Yu Sun 0030, Priyanka Patel, Kostas Daniilidis, Michael J. Black, Muhammed Kocabas |
CVPR | 2 |
| 2024 | TokenHMR: Advancing Human Mesh Recovery with a Tokenized Pose RepresentationabstractWe address the problem of regressing 3D human pose and shape from a single image, with a focus on 3D accuracy. The current best methods leverage large datasets of 3D pseudo-ground-truth (p-GT) and 2D keypoints, leading to robust performance. With such methods, however, we observe a paradoxical decline in 3D pose accuracy with increasing 2D accuracy. This is caused by biases in the p-GT and the use of an approximate camera projection model. We quantify the error induced by current camera models and show that fitting 2D keypoints and p-GT accurately causes incorrect 3D poses. Our analysis defines the invalid distances within which minimizing 2D and p-GT losses is detrimental. We use this to formulate a new loss, “Threshold-Adaptive Loss Scaling” (TALS), that penalizes gross 2D and p-GT errors but not smaller ones. With such a loss, there are many 3D poses that could equally explain the 2D evidence. To reduce this ambiguity we need a prior over valid human poses but such priors can introduce unwanted bias. To address this, we exploit a tokenized representation of human pose and reformulate the problem as token prediction. This restricts the estimated poses to the space of valid poses, effectively improving robustness to occlusion. Exten-sive experiments on the EMDB and 3DPW datasets show that our reformulated loss and tokenization allows us to train on in-the-wild data while improving 3D accuracy over the state-of-the-art. Our models and code are available for research at https://tokenhmr.is.tue.mpg.de. Saikumar Dwivedi, Yu Sun 0030, Priyanka Patel, Yao Feng 0001, Michael J. Black |
CVPR | 2 |
| 2024 | ChatPose: Chatting about 3D Human PoseabstractWe introduce ChatPose, a framework employing Large Language Models (LLMs) to understand and reason about 3D human poses from images or textual descriptions. Our work is motivated by the human ability to intuitively under-stand postures from a single image or a brief description, a process that intertwines image interpretation, world knowl-edge, and an understanding of body language. Traditional human pose estimation and generation methods often op-erate in isolation, lacking semantic understanding and rea-soning abilities. ChatPose addresses these limitations by embedding SMPL poses as distinct signal tokens within a multimodal LLM, enabling the direct generation of 3D body poses from both textual and visual inputs. Leveraging the powerful capabilities of multimodal LLMs, ChatPose uni-fies classical 3D human pose and generation tasks while offering user interactions. Additionally, ChatPose empow-ers LLMs to apply their extensive world knowledge in rea-soning about human poses, leading to two advanced tasks: speculative pose generation and reasoning about pose esti-mation. These tasks involve reasoning about humans to generate 3D poses from subtle text queries, possibly ac-companied by images. We establish benchmarks for these tasks, moving beyond traditional 3D pose generation and estimation methods. Our results show that ChatPose out-performs existing multimodal LLMs and task-specific meth-ods on these newly proposed tasks. Furthermore, Chat-Pose's ability to understand and generate 3D human poses based on complex reasoning opens new directions in human pose analysis. Code and data are available for research at https://yfeng95.github.io/ChatPose. Yao Feng 0001, Saikumar Dwivedi, Yu Sun 0030, Priyanka Patel, Michael J. Black |
CVPR | 4 |
| 2024 | Learning Monocular Regression of 3D People in Crowds via Scene-Aware Blending and De-OcclusionabstractIn this study, we address the challenge of estimating 3D body pose, shape, and depth relationships from single RGB images in crowded scenes. The difficulty lies in the limited availability of in-the-wild training samples, which feature densely populated scenes. To mitigate this issue, we introduce a synthesis-based approach that fuses multiple human samples into a single composite scene. Our innovative scene-aware blending technique maintains human-scene consistency by positioning individuals within plausible locations and adjusting their scales to conform to 3D settings. Furthermore, our method enables flexible per-subject occlusion management during the blending process, bolstering the robustness of 3D human body representations through a novel de-occlusion training scheme. We present a one-stage model, CBD, designed to learn monocular regression of 3D people in crowds by leveraging blending and de-occlusion techniques. Our quantitative and qualitative evaluations on four benchmark datasets reveal that CBD surpasses existing state-of-the-art approaches in terms of 3D human pose and mesh regression accuracy, thereby establishing it as a promising solution for monocular 3D human mesh recovery in densely populated scenes. Yu Sun 0030, Lubing Xu, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D EnvironmentsabstractAlthough the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera motion. To address these issues, we adopt a novel 5D representation (space, time, and identity) that enables end-to-end reasoning about people in scenes. Our method, called TRACE, introduces several novel architectural components. Most importantly, it uses two new “maps” to reason about the 3D trajectory of people over time in camera, and world, coordinates. An additional memory unit enables persistent tracking of people even during long occlusions. TRACE is the first one-stage method to jointly recover and track 3D humans in global coordinates from dynamic cameras. By training it end-to-end, and using full image information, TRACE achieves state-of-the-art performance on tracking and HPS benchmarks. The code11https://www.yusun.work/TRACE/TRACE.html and dataset22https://github.com/Arthur151/DynaCam are released for research purposes. Yu Sun 0030, Qian Bao, Wu Liu 0005, Tao Mei 0001, Michael J. Black |
CVPR | 1 |
| 2022 | Putting People in their Place: Monocular Regression of 3D People in DepthabstractGiven an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene contains people of very different sizes, e.g. from infants to adults. To solve this, we need several things. First, we develop a novel method to infer the poses and depth of multiple people in a single image. While previous work that estimates multiple people does so by reasoning in the image plane, our method, called BEV, adds an additional imaginary Bird's-Eye-View representation to explicitly reason about depth. BEV reasons simultaneously about body centers in the image and in depth and, by combing these, estimates 3D body position. Unlike prior work, BEV is a single-shot method that is end-to-end differentiable. Second, height varies with age, making it impossible to resolve depth without also estimating the age of people in the image. To do so, we exploit a 3D body model space that lets BEV infer shapes from infants to adults. Third, to train BEV, we need a new dataset. Specifically, we create a “Relative Human” (RH) dataset that includes age labels and relative depth relationships between the people in the images. Extensive experiments on RH and AGORA demonstrate the effectiveness of the model and training scheme. BEV out-performs existing methods on depth reasoning, child shape estimation, and robustness to occlusion. The code11https://github.com/Arthur151/ROMP and dataset22https://github.com/Arthur151/Relative_Human are released for research purposes. Yu Sun 0030, Wu Liu 0005, Qian Bao, Yili Fu 0001, Tao Mei 0001, Michael J. Black |
CVPR | 1 |
| 2022 | Learning Monocular Mesh Recovery of Multiple Body Parts Via SynthesisabstractIn this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Existing optimization-based methods iteratively fit the 3D mesh to the 2d pose, which is very time-consuming. To address these limitations, we propose to integrate multiple 3D single-body-part datasets to create a highly diverse whole-body 3D motion space for learning from controllable synthetics. Compared with the learning-based approaches, the proposed method greatly alleviates the reliance on training data. Compared with the optimization-based approaches, the proposed method is a hundred times faster. Our proposed method also outperforms previous state-of-the-art methods on CMU Panoptic dataset. Yu Sun 0030, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001 |
ICASSP | 1 |
| 2022 | WOC: A Handy Webcam-based 3D Online ChatroomabstractWe develop WOC, a webcam-based 3D virtual online chatroom for multi-person interaction, which captures the 3D motion of users and drives their individual 3D virtual avatars in real-time. Compared to the existing wearable equipment-based solution, WOC offers convenient and low-cost 3D motion capture with a single camera. To promote the immersive chat experience, WOC provides high-fidelity virtual avatar manipulation, which also supports the user-defined characters. With the distributed data flow service, the system delivers highly synchronized motion and voice for all users. Deployed on the website and no installation required, users can freely experience the virtual online chat at https://yanch.cloud/. Chuanhang Yan, Yu Sun 0030, Qian Bao, Jinhui Pang, Wu Liu 0005, Tao Mei 0001 |
ACM Multimedia | 2 |
| 2021 | Monocular, One-stage, Regression of Multiple 3D PeopleabstractThis paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a One-stage fashion for Multiple 3D People (termed ROMP). The approach is conceptually simple, bounding box-free, and able to learn a per-pixel representation in an end-to-end manner. Our method simultaneously predicts a Body Center heatmap and a Mesh Parameter map, which can jointly describe the 3D body mesh on the pixel level. Through a body-center-guided sampling process, the body mesh parameters of all people in the image are easily extracted from the Mesh Parameter map. Equipped with such a fine-grained representation, our one-stage framework is free of the complex multi-stage process and more robust to occlusion. Compared with state-of-the-art methods, ROMP achieves superior performance on the challenging multi-person benchmarks, including 3DPW and CMU Panoptic. Experiments on crowded/occluded datasets demonstrate the robustness under various types of occlusion. The code, released at https://github.com/Arthur151/ROMP, is the first real-time implementation of monocular multi-person 3D mesh regression. Yu Sun 0030, Qian Bao, Wu Liu 0005, Yili Fu 0001, Michael J. Black, Tao Mei 0001 |
ICCV | 1 |
| 2019 | Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled RepresentationabstractWe describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature, we propose a skeleton-disentangling based framework, which divides this task into multi-level spatial and temporal granularity in a decoupling manner. In spatial, we propose an effective and pluggable “disentangling the skeleton from the details” (DSD) module. It reduces the complexity and decouples the skeleton, which lays a good foundation for temporal modeling. In temporal, the self-attention based temporal convolution network is proposed to efficiently exploit the short and long-term temporal cues. Furthermore, an unsupervised adversarial training strategy, temporal shuffles and order recovery, is designed to promote the learning of motion dynamics. The proposed method outperforms the state-of-the-art 3D human mesh recovery methods by 15.4% MPJPE and 23.8% PA-MPJPE on Human3.6M. State-of-the-art results are also achieved on the 3D pose in the wild (3DPW) dataset without any fine-tuning. Especially, ablation studies demonstrate that skeleton-disentangled representation is crucial for better temporal modeling and generalization. Yu Sun 0030, Yun Ye 0001, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001, Tao Mei 0001 |
ICCV | 1 |