EDBT 2026 Demo / reviewers in the wild / expert
Qi Wang 0148
dblp:19/1924-148
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0009-0005-2599-2696ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
3D vision · 56% Face, body and person analysis · 30% Video understanding and tracking · 12% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 46% Visual content generation and editing · 23% Geometric modeling and processing · 23% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d human pose estimation |
1.5 | 2 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization · CVPR 2023 |
Computer vision › 3D vision
pose estimation |
1.5 | 2 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose Estimation · ICCV 2023 |
Computer vision › Video understanding and tracking
temporal modeling |
0.9 | 1 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Computer vision › Face, body and person analysis › human pose estimation
video pose estimation |
0.9 | 1 | 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs · ICML 2025 |
Computer vision › Face, body and person analysis › human pose estimation › articulated pose estimation
hand pose estimation |
0.7 | 1 | 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose Estimation · ICCV 2023 |
Computer vision › Face, body and person analysis
human pose estimation |
0.7 | 1 | 2023 | Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization · CVPR 2023 |
Computer vision › 3D vision › pose estimation › 3d hand pose estimation
interacting hand pose estimation |
0.7 | 1 | 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose Estimation · ICCV 2023 |
Geometric modeling and processing › 3d reconstruction › 3d human reconstruction
3d face reconstruction |
0.5 | 1 | 2021 | A Virtual Character Generation and Animation System for E-Commerce Live Streaming · ACM Multimedia 2021 |
Computer animation and physical simulation
character animation |
0.5 | 1 | 2021 | A Virtual Character Generation and Animation System for E-Commerce Live Streaming · ACM Multimedia 2021 |
Computer animation and physical simulation › motion capture
motion capture animation |
0.5 | 1 | 2021 | A Virtual Character Generation and Animation System for E-Commerce Live Streaming · ACM Multimedia 2021 |
Visual content generation and editing
virtual character generation |
0.5 | 1 | 2021 | A Virtual Character Generation and Animation System for E-Commerce Live Streaming · ACM Multimedia 2021 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.2 | 1 | 2023 | Towards Stable Human Pose Estimation via Cross-View Fusion and Foot Stabilization · CVPR 2023 |
Computer vision › 3D vision › geometric optimization
pose optimization |
0.2 | 1 | 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose Estimation · ICCV 2023 |
Machine learning › Generative modeling
synthetic data generation |
0.2 | 1 | 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose Estimation · ICCV 2023 |
Multimedia systems and quality of experience › video streaming
live streaming |
0.1 | 1 | 2021 | A Virtual Character Generation and Animation System for E-Commerce Live Streaming · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
vision transformer · 1.5attention · 0.92d pose evidence · 0.9transformer-based pose estimation network · 0.7pose optimization · 0.7optimization-based reconstruction · 0.7cross-view fusion · 0.7weakly supervised 3d face reconstruction · 0.5text-driven animation · 0.5motion capture · 0.5differentiable neural rendering · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ExtPose: Robust and Coherent Pose Estimation by Extending ViTsabstractVision Transformers (ViT) are remarkable at 3D pose estimation, yet they still encounter certain challenges. One issue is that the popular ViT architecture for pose estimation is limited to images and lacks temporal information. Another challenge is that the prediction often fails to maintain pixel alignment with the original images. To address these issues, we propose a systematic framework for 3D pose estimation, called ExtPose. ExtPose extends image ViT to the challenging scenario and video setting by taking in additional 2D pose evidence and capturing temporal information in a full attention-based manner. We use 2D human skeleton images to integrate structured 2D pose information. By sharing parameters and attending across modalities and frames, we enhance the consistency between 3D poses and 2D videos without introducing additional parameters. We achieve state-of-the-art (SOTA) performance on multiple human and hand pose estimation benchmarks with substantial improvements to 34.0mm (-23%) on 3DPW and 4.9mm (-18%) on FreiHAND in PA-MPJPE over the other ViT-based methods respectively. Rongyu Chen, Lian Zhuo, Linlin Yang 0001, Qi Wang 0148, Liefeng Bo, Bang Zhang, Angela Yao |
ICML | 4 |
| 2024 | Cloth2Tex: A Customized Cloth Texture Generation Pipeline for 3D Virtual Try-OnabstractFabricating and designing 3D garments has become extremely demanding with the increasing need for synthesizing realistic dressed persons for a variety of applications, e.g. 3D virtual try-on, digitalization of 2D clothes into 3D apparel, and cloth animation. It thus necessitates a simple and straightforward pipeline to obtain high-quality texture from simple input, such as 2D reference images. Since traditional warping-based texture generation methods require a significant number of control points to be manually selected for each type of garment, which can be a time-consuming and tedious process. We propose a novel method, called Cloth2Tex, which eliminates the human burden in this process. Cloth2Tex is a self-supervised method that generates texture maps with reasonable layout and structural consistency. Another key feature of Cloth2Tex is that it can be used to support high-fidelity texture inpainting. This is done by combining Cloth2Tex with a prevailing latent diffusion model. We evaluate our approach both qualitatively and quantitatively and demonstrate that Cloth2Tex can generate high-quality texture maps and achieve the best visual effects in comparison to other methods. For more details and animated results, please see https://tomguluson92. github.io/projects/cloth2tex/. Daiheng Gao, Xindi Zhang 0003, Qi Wang 0148, Bang Zhang, Liefeng Bo, Qixing Huang |
3DV | 4 |
| 2023 | Towards Stable Human Pose Estimation via Cross-View Fusion and Foot StabilizationabstractTowards stable human pose estimation from monocular images, there remain two main dilemmas. On the one hand, the different perspectives, i.e., front view, side view, and top view, appear the inconsistent performances due to the depth ambiguity. On the other hand, foot posture plays a significant role in complicated human pose estimation, i.e., dance and sports, and foot-ground interaction, but unfortunately, it is omitted in most general approaches and datasets. In this paper, we first propose the Cross-View Fusion (CVF) module to catch up with better 3D intermediate representation and alleviate the view inconsistency based on the vision transformer encoder. Then the optimization-based method is introduced to reconstruct the foot pose and foot-ground contact for the general multi-view datasets including AIST++ and Human3.6M. Besides, the reversible kinematic topology strategy is innovated to utilize the contact information into the full-body with foot pose regressor. Extensive experiments on the popular benchmarks demonstrate that our method outperforms the state-of-the-art approaches by achieving 40.1mm PA-MPJPE on the 3DPW test set and 43.8mm on the AIST++ test set. Lian Zhuo, Qi Wang 0148, Bang Zhang, Liefeng Bo |
CVPR | 3 |
| 2023 | RenderIH: A Large-scale Synthetic Dataset for 3D Interacting Hand Pose EstimationabstractThe current interacting hand (IH) datasets are relatively simplistic in terms of background and texture, with hand joints being annotated by a machine annotator, which may result in inaccuracies, and the diversity of pose distribution is limited. However, the variability of background, pose distribution, and texture can greatly influence the generalization ability. Therefore, we present a large-scale synthetic dataset –RenderIH– for interacting hands with accurate and diverse pose annotations. The dataset contains 1M photo-realistic images with varied backgrounds, perspectives, and hand textures. To generate natural and diverse interacting poses, we propose a new pose optimization algorithm. Additionally, for better pose estimation accuracy, we introduce a transformer-based pose estimation network, TransHand, to leverage the correlation between interacting hands and verify the effectiveness of RenderIH in improving results. Our dataset is model-agnostic and can improve more accuracy of any hand pose estimation method in comparison to other real or synthetic datasets. Experiments have shown that pretraining on our synthetic data can significantly decrease the error from 6.76mm to 5.79mm, and our Transhand surpasses contemporary methods. Our dataset and code are available at https://github.com/adwardlee/RenderIH. Linrui Tian, Xindi Zhang 0003, Qi Wang 0148, Bang Zhang, Liefeng Bo, Chen Chen 0001 |
ICCV | 4 |
| 2021 | A Virtual Character Generation and Animation System for E-Commerce Live StreamingabstractVirtual character has been widely adopted in many areas, such as virtual assistant, virtual customer service, robotics and etc. In this paper, we focus on its application in e-commerce live streaming. Particularly, we propose a virtual character generation and animation system that supports e-commerce live streaming with virtual characters as anchors. The system offers a virtual character face generation tool based on a weakly supervised 3D face reconstruction method. The method takes a single photo as input and generates a 3D face model with both similarity and aesthetics considered. It does not require 3D face annotation data due to the assist of differentiable neural rendering technique which seamlessly integrates rendering into a deep learning based 3D face reconstruction framework. Moreover, the system provides two animation approaches which support two different ways of live stream respectively. The first approach is based on real-time motion capture. An actor's performance is captured in real-time via a monocular camera, and then utilized for animating a virtual anchor. The second approach is text driven animation, in which the human-like animation is automatically generated based on a text script. The relationship between text script and animation is learned based on the training data which can be accumulated via the motion capture based animation. To our best knowledge, the presented work is the first sophisticated virtual character generation and animation system that is designed for e-commerce live streaming and actually deployed on an online shopping platform with millions of daily audiences. Bang Zhang, Peng Zhang 0080, Jinwei Qi, Daiheng Gao, Haiming Zhao, Xiaoduan Feng, Qi Wang 0148, Lian Zhuo |
ACM Multimedia | 9 |