VLDB 2026 Research / reviewers in the wild / expert
Yongtao Ge
dblp:289/0822
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
0000-0003-1265-3204ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
3D vision · 46% Image recognition and object detection · 25% Generative modeling · 8% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d reconstruction |
0.9 | 1 | 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model training
diffusion model fine-tuning |
0.9 | 1 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 |
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction |
0.9 | 1 | 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction · ICCV 2025 |
Computer vision › Segmentation and scene understanding
image segmentation |
0.9 | 1 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.9 | 1 | 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks? · ICLR 2025 |
Computer vision › 3D vision › camera calibration
camera model |
0.7 | 1 | 2023 | Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction · ICCV 2023 |
Computer vision › 3D vision
human mesh recovery |
0.7 | 1 | 2023 | Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction · ICCV 2023 |
Computer vision › Image recognition and object detection
object detection |
0.7 | 1 | 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations · AAAI 2023 |
Computer vision › Image recognition and object detection › image annotation
point labeling |
0.7 | 1 | 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations · AAAI 2023 |
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection |
0.7 | 1 | 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations · AAAI 2023 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
0.7 | 1 | 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations · AAAI 2023 |
Computer vision › Face, body and person analysis
human pose estimation |
0.6 | 1 | 2022 | Poseur: Direct Human Pose Regression with Transformers · ECCV (6) 2022 |
Machine learning › Deep learning architectures and training
transformer |
0.6 | 1 | 2022 | Poseur: Direct Human Pose Regression with Transformers · ECCV (6) 2022 |
Machine learning › Deep learning architectures and training
data augmentation |
0.2 | 1 | 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point Annotations · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
temporal motion modeling · 0.9one-step fine-tuning · 0.9diffusion prior · 0.9point-guided copy-paste · 0.7multiple instance learning · 0.7hungarian-based point matching · 0.7transformer · 0.6direct regression · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | POMATO: Marrying Pointmap Matching with Temporal Motions for Dynamic 3D Reconstruction
Songyan Zhang, Yongtao Ge, Jinyuan Tian, Guangkai Xu, Hao Chen 0041, Chunhua Shen |
ICCV | 2 |
| 2025 | What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?abstractExtensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for a few dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) As a special case of the diffusion scheduler by setting its hyper-parameters, the multi-step generation can be simplified to a one-step fine-tuning paradigm without any loss of performance, while significantly speeding up inference. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific supervision can be beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailored for dense visual perception tasks exploiting diffusion priors. Different from the previous multi-step methods, our paradigm offers a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for task-specific supervision, which can be critical for improving the fine-grained details of predictions. Comprehensive experiments on a diverse set of dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method. Code: https://github.com/aim-uofa/GenPercept Guangkai Xu, Yongtao Ge, Chengxiang Fan, Kangyang Xie, Zhiyue Zhao, Hao Chen 0041, Chunhua Shen |
ICLR | 2 |
| 2023 | Point-Teaching: Weakly Semi-supervised Object Detection with Point AnnotationsabstractPoint annotations are considerably more time-efficient than bounding box annotations. However, how to use cheap point annotations to boost the performance of semi-supervised object detection is still an open question. In this work, we present Point-Teaching, a weakly- and semi-supervised object detection framework to fully utilize the point annotations. Specifically, we propose a Hungarian-based point-matching method to generate pseudo labels for point-annotated images. We further propose multiple instance learning (MIL) approaches at the level of images and points to supervise the object detector with point annotations. Finally, we propose a simple data augmentation, named Point-Guided Copy-Paste, to reduce the impact of those unmatched points. Experiments demonstrate the effectiveness of our method on a few datasets and various data regimes. In particular, Point-Teaching outperforms the previous best method Group R-CNN by 3.1 AP with 5% fully labeled data and 2.3 AP with 30% fully labeled data on the MS COCO dataset. We believe that our proposed framework can largely lower the bar of learning accurate object detectors and pave the way for its broader applications. The code is available at https://github.com/YongtaoGe/Point-Teaching. Yongtao Ge, Qiang Zhou 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
AAAI | 1 |
| 2023 | Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh ReconstructionabstractAs it is hard to calibrate single-view RGB images in the wild, existing 3D human mesh reconstruction (3DHMR) methods either use a constant large focal length or estimate one based on the background environment context, which can not tackle the problem of the torso, limb, hand or face distortion caused by perspective camera projection when the camera is close to the human body. The naive focal length assumptions can harm this task with the incorrectly formulated projection matrices. To solve this, we propose Zolly, the first 3DHMR method focusing on perspective-distorted images. Our approach begins with analysing the reason for perspective distortion, which we find is mainly caused by the relative location of the human body to the camera center. We propose a new camera model and a novel 2D representation, termed distortion image, which describes the 2D dense distortion scale of the human body. We then estimate the distance from distortion scale features rather than environment context features. Afterwards, We integrate the distortion feature with image features to reconstruct the body mesh. To formulate the correct projection matrix and locate the human body position, we simultaneously use perspective and weak-perspective projection loss. Since existing datasets could not handle this task, we propose the first synthetic dataset PDHuman and extend two real-world datasets tailored for this task, all containing perspective-distorted human images. Extensive experiments show that Zolly outperforms existing state-of-the-art methods on both perspective-distorted datasets and the standard benchmark (3DPW). Code and dataset will be released at https://wenjiawang0312.github.io/projects/zolly/. Wenjia Wang 0009, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Chunhua Shen, Lei Yang 0059, Taku Komura |
ICCV | 2 |
| 2022 | Poseur: Direct Human Pose Regression with Transformers
Weian Mao, Yongtao Ge, Chunhua Shen, Zhi Tian, Zhibin Wang 0004, Anton van den Hengel |
ECCV (6) | 2 |