EDBT 2026 Demo / reviewers in the wild / expert
Rongyu Chen
dblp:279/0280
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
3D vision · 32% Video understanding and tracking · 14% Face, body and person analysis · 13% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d human pose estimation |
1.7 | 2 | 2025 | Semantics-aware Test-time Adaptation for 3D Human Pose Estimation · ICML 2025 ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Computer vision › Video understanding and tracking
action segmentation |
0.9 | 1 | 2025 | Condensing Action Segmentation Datasets via Generative Network Inversion · CVPR 2025 |
Machine learning › Efficient and distributed learning
dataset distillation |
0.9 | 1 | 2025 | Condensing Action Segmentation Datasets via Generative Network Inversion · CVPR 2025 |
Machine learning › Generative modeling › inverse problem
generative network inversion |
0.9 | 1 | 2025 | Condensing Action Segmentation Datasets via Generative Network Inversion · CVPR 2025 |
Computer vision › 3D vision
pose estimation |
0.9 | 1 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Computer vision › Video understanding and tracking
temporal modeling |
0.9 | 1 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation |
0.9 | 1 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering |
0.8 | 1 | 2024 | Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement · ICLR 2024 |
Machine learning › Trustworthy machine learning › calibration
confidence calibration |
0.8 | 1 | 2024 | On the Calibration of Human Pose Estimation · ICML 2024 |
Computer vision › Face, body and person analysis
human pose estimation |
0.8 | 1 | 2024 | On the Calibration of Human Pose Estimation · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.8 | 1 | 2024 | Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement · ICLR 2024 |
Computer vision › 3D vision
human mesh recovery |
0.7 | 1 | 2023 | MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape Recovery · ICCV 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
multi-hypothesis estimation |
0.7 | 1 | 2023 | MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape Recovery · ICCV 2023 |
Computer vision › 3D vision › pose estimation
probabilistic pose estimation |
0.7 | 1 | 2023 | MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape Recovery · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 0.9test-time adaptation · 0.9network inversion · 0.9latent code · 0.9generative prior · 0.9diffusion model · 0.9attention · 0.92d pose evidence · 0.9energy score · 0.8activation shaping · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Condensing Action Segmentation Datasets via Generative Network InversionabstractThis work presents the first condensation approach for procedural video datasets used in temporal action segmentation. We propose a condensation framework that leverages generative prior learned from the dataset and network inversion to condense data into compact latent codes with significant storage reduced across temporal and channel aspects. Orthogonally, we propose sampling diverse and representative action sequences to minimize video-wise redundancy. Our evaluation on standard benchmarks demonstrates consistent effectiveness in condensing TAS datasets and achieving competitive performances. Specifically, on the Breakfast dataset, our approach reduces storage by over 500× while retaining 83% of the performance compared to training with the full dataset. Furthermore, when applied to a downstream incremental learning task, it yields superior performance compared to the state-of-the-art. Guodong Ding, Rongyu Chen, Angela Yao |
CVPR | 2 |
| 2025 | Multi-Hypothesis 3D Hand Mesh Recovering from a Single Blurry ImageabstractRecovery of 3D hand mesh from blurry hand images is challenging due to the ambiguity. Most existing works attempt to solve this issue by exploiting physical and temporal constraints. However, those works ignore the fact that multiple feasible solutions exist. In this paper, we propose a two-stage Multi-Hypothesis Hand Mesh Recovery network, consisting of a generation and selection model. In the first stage, the generation model explicitly extracts the temporal information with an unfolder. Then, a multi-hypothesis Transformer generates multiple diverse hypotheses with a lightweight hypothesis embedding set. In the second stage, the selection model selects a subset of good-quality hypotheses. We additionally combine the classifying and ranking loss to better align with the target of the selection model. Extensive experiments show that the proposed method produces much more accurate results on blurry images. Source code is available at https://github.com/RandSF/Multi_Hypothesis_BlurHandNet. Rongyu Chen, Zhongqun Zhang, Yihua Cheng, Hyung Jin Chang |
ICME | 2 |
| 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTsabstractVision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel alignment with the original images. To address these issues, we propose a systematic framework for 3D pose estimation, called ExtPose. ExtPose extends image ViT to the challenging scenario and video setting by taking in additional 2D pose evidence and capturing temporal information in a full attention-based manner. We use 2D human skeleton images to integrate structured 2D pose information. By sharing parameters and attending across modalities and frames, we enhance the consistency between 3D poses and 2D videos without introducing additional parameters. We achieve state-of-the-art (SOTA) performance on multiple human and hand pose estimation benchmarks with substantial improvements to 34.0mm (-23%) on 3DPW and 4.9mm (-18%) on FreiHAND in PA-MPJPE over the other ViT-based methods respectively. Rongyu Chen, Lian Zhuo, Linlin Yang 0001, Qi Wang 0148, Liefeng Bo, Bang Zhang, Angela Yao |
ICML | 1 |
| 2025 | Semantics-aware Test-time Adaptation for 3D Human Pose EstimationabstractThis work highlights a semantics misalignment in 3D human pose estimation. For the task of test-time adaptation, the misalignment manifests as overly smoothed and unguided predictions. The smoothing settles predictions towards some average pose. Furthermore, when there are occlusions or truncations, the adaptation becomes fully unguided. To this end, we pioneer the integration of a semantics-aware motion prior for the test-time adaptation of 3D pose estimation. We leverage video understanding and a well-structured motion-text space to adapt the model motion prediction to adhere to video semantics during test time. Additionally, we incorporate a missing 2D pose completion based on the motion-text similarity. The pose completion strengthens the motion prior’s guidance for occlusions and truncations. Our method significantly improves state-of-the-art 3D human pose estimation TTA techniques, with more than 12% decrease in PA-MPJPE on 3DPW and 3DHP. Qiuxia Lin, Rongyu Chen, Kerui Gu, Angela Yao |
ICML | 2 |
| 2024 | Scaling for Training Time and Post-hoc Out-of-distribution Detection EnhancementabstractActivation shaping has proven highly effective for identifying out-of-distribution (OOD) samples post-hoc. Activation shaping prunes and scales network activations before estimating the OOD energy score; such an extremely simple approach achieves state-of-the-art OOD detection with minimal in-distribution (ID) accuracy drops. This paper analyzes the working mechanism behind activation shaping. We directly show that the benefits for OOD detection derive only from scaling, while pruning is detrimental. Based on our analysis, we propose SCALE, an even simpler yet more effective post-hoc network enhancement method for OOD detection. SCALE attains state-of-the-art OOD detection performance without any compromises on ID accuracy. Furthermore, we integrate scaling concepts into learning and propose Intermediate Tensor SHaping (ISH) for training-time OOD detection enhancement. ISH achieves significant AUROC improvements for both near- and far-OOD, highlighting the importance of activation distributions in emphasizing ID data characteristics. Our code and models are available at https://github.com/kai422/SCALE. Rongyu Chen, Gianni Franchi, Angela Yao |
ICLR | 2 |
| 2024 | On the Calibration of Human Pose Estimationabstract2D human pose estimation predicts keypoint locations and the corresponding confidence. Calibration-wise, the confidence should be aligned with the pose accuracy. Yet existing pose estimation methods tend to estimate confidence with heuristics such as the maximum value of heatmaps. This work shows, through theoretical analysis and empirical verification, a calibration gap in current pose estimation frameworks. Our derivations directly lead to closed-form adjustments in the confidence based on additionally inferred instance size and visibility. Given the black-box nature of deep neural networks, however, it is not possible to close the gap with only closed-form adjustments. We go one step further and propose a Calibrated ConfidenceNet (CCNet) to explicitly learn network-specific adjustments with a confidence prediction branch. The proposed CCNet, as a lightweight post-hoc addition, improves the calibration of standard off-the-shelf pose estimation frameworks. Kerui Gu, Rongyu Chen, Xuanlong Yu, Angela Yao |
ICML | 2 |
| 2023 | MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape RecoveryabstractFor monocular RGB-based 3D pose and shape estimation, multiple solutions are often feasible due to factors like occlusions and truncations. This work presents a multi-hypothesis probabilistic framework by optimizing the Kullback–Leibler divergence (KLD) between the data and model distribution. Our formulation reveals a connection between the pose entropy and diversity in the multiple hypotheses that has been neglected by previous works. For a comprehensive evaluation, besides the best hypothesis (BH) metric, we factor in visibility for evaluating diversity. Additionally, our framework is label-friendly – it can be learned from only partial 2D keypoints, such as visible keypoints. Experiments on both ambiguous and real-world benchmarks demonstrate that our method outperforms other state-of-the-art multi-hypothesis methods. The project page is at https://gloryyrolg.github.io/MHEntropy. Rongyu Chen, Linlin Yang 0001, Angela Yao |
ICCV | 1 |