VLDB 2026 Research / reviewers in the wild / expert
Wentao Wang 0009
dblp:95/5409-9
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6626-3319ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Quality Full-Head 3D Avatar Generation from Any Single Portrait ImageabstractIn this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modeling fine-grained facial textures and maintaining identity information. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring high-quality facial texture details. To further improve performance, we propose a novel multi-view diffusion named ID-TS diffusion model, which integrate identity and expression information into the two-stage multi-view diffusion process. The low-resolution stage ensures structural consistency of heads across multiple views, while the high-resolution stage preserves facial detail fidelity and coherence. Finally, we propose an enhanced feed-forward Gaussian avatar reconstruction method that optimizes the network on multi-view images of each single subject, significantly improving 3D facial texture details. Extensive experiments show that our method demonstrates robust performance across challenging scenarios, while showcasing broad applicability across numerous downstream tasks. Yujie Gao 0001, Chencheng Wang, Xianbing Sun, Jiahui Zhan, Wentao Wang 0009, Yiyi Zhang 0002, Haohua Zhao 0001, Liqing Zhang 0001, Jianfu Zhang 0003 |
AAAI | 5 |
| 2025 | GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataabstractGiven a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images. Wentao Wang 0009, Hang Ye 0002, Fangzhou Hong, Xue Yang 0005, Jianfu Zhang 0003, Yizhou Wang 0001, Ziwei Liu 0002, Liang Pan |
NeurIPS | 1 |
| 2023 | The KFIoU Loss for Rotated Object Detection
Xue Yang 0005, Yue Zhou 0005, Gefan Zhang, Jirui Yang, Wentao Wang 0009, Junchi Yan, Xiaopeng Zhang 0008, Qi Tian 0001 |
ICLR | 5 |
| 2023 | Detecting Rotated Objects as Gaussian Distributions and its 3-D GeneralizationabstractExisting detection methods commonly use a parameterized bounding box (BBox) to model and detect (horizontal) objects and an additional rotation angle parameter is used for rotated objects. We argue that such a mechanism has fundamental limitations in building an effective regression loss for rotation detection, especially for high-precision detection with high IoU (e.g., 0.75). Instead, we propose to model the rotated objects as Gaussian distributions. A direct advantage is that our new regression loss regarding the distance between two Gaussians e.g., Kullback-Leibler Divergence (KLD), can well align the actual detection performance metric, which is not well addressed in existing methods. Moreover, the two bottlenecks i.e., boundary discontinuity and square-like problem also disappear. We also propose an efficient Gaussian metric-based label assignment strategy to further boost the performance. Interestingly, by analyzing the BBox parameters' gradients under our Gaussian-based KLD loss, we show that these parameters are dynamically updated with interpretable physical meaning, which help explain the effectiveness of our approach, especially for high-precision detection. We extend our approach from 2-D to 3-D with a tailored algorithm design to handle the heading estimation, and experimental results on twelve public datasets (2-D/3-D, aerial/text/face images) with various base detectors show its superiority. Xue Yang 0005, Gefan Zhang, Xiaojiang Yang, Yue Zhou 0005, Wentao Wang 0009, Jin Tang 0001, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Diverse image inpainting with disentangled uncertainty
Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Haoyu Ling, Liqing Zhang 0001 |
Pattern Recognit. | 1 |
| 2022 | Dual-path Image Inpainting with Auxiliary GAN InversionabstractDeep image inpainting can inpaint a corrupted image using a feed-forward inference, but still fails to handle large missing area or complex semantics. Recently, GAN inversion based inpainting methods propose to leverage semantic information in pretrained generator (e.g., StyleGAN) to solve the above issues. Different from feed-forward methods, they seek for a closest latent code to the corrupted image and feed it to a pretrained generator. However, inferring the latent code is either time-consuming or inaccurate. In this paper, we develop a dual-path inpainting network with inversion path and feed-forward path, in which inversion path provides auxiliary information to help feed-forward path. We also design a novel deformable fusion module to align the feature maps in two paths. Experiments on FFHQ and LSUN demonstrate that our method is effective in solving the aforementioned problems while producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Xue Yang 0005, Liqing Zhang 0001 |
CVPR | 1 |
| 2021 | Dense Label Encoding for Boundary Discontinuity Free Rotation DetectionabstractRotation detection serves as a fundamental building block in many visual applications involving aerial image, scene text, and face etc. Differing from the dominant regression-based approaches for orientation estimation, this paper explores a relatively less-studied methodology based on classification. The hope is to inherently dismiss the boundary discontinuity issue as encountered by the regression-based detectors. We propose new techniques to push its frontier in two aspects: i) new encoding mechanism: the design of two Densely Coded Labels (DCL) for angle classification, to replace the Sparsely Coded Label (SCL) in existing classification-based detectors, leading to three times training speed increase as empirically observed across benchmarks, further with notable improvement in detection accuracy; ii) loss re-weighting: we propose Angle Distance and Aspect Ratio Sensitive Weighting (ADARSW), which improves the detection accuracy especially for square-like objects, by making DCL-based detectors sensitive to angular distance and object’s aspect ratio. Extensive experiments and visual analysis on large-scale public datasets for aerial images i.e. DOTA, UCAS-AOD, HRSC2016, as well as scene text dataset ICDAR2015 and MLT, show the effectiveness of our approach. The source code is available at $\color{Red}{\text{DCL}}$ and is also integrated in our open source rotation detection benchmark:$\color{Red}{\text{RotationDetection}}$. Xue Yang 0005, Liping Hou, Yue Zhou 0005, Wentao Wang 0009, Junchi Yan |
CVPR | 4 |
| 2021 | Parallel Multi-Resolution Fusion Network for Image InpaintingabstractConventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow layers of the auto-encoder architecture can not be well integrated. Differing from the conventional image inpainting architecture, we design a parallel multi-resolution inpainting network with multi-resolution partial convolution, in which low-resolution branches focus on the global structure while high-resolution branches focus on the local texture details. All these high- and low-resolution streams are in parallel and fused repeatedly with multi-resolution masked representation fusion so that the reconstructed images are semantically robust and textually plausible. Experimental results show that our method can effectively fuse structure and texture information, producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Jianfu Zhang 0003, Li Niu 0002, Haoyu Ling, Xue Yang 0005, Liqing Zhang 0001 |
ICCV | 1 |
| 2021 | Rethinking Rotated Object Detection with Gaussian Wasserstein Distance LossabstractBoundary discontinuity and its inconsistency to the final detection metric have been the bottleneck for rotating detection regression loss design. In this paper, we propose a novel regression loss based on Gaussian Wasserstein distance as a fundamental approach to solve the problem. Specifically, the rotated bounding box is converted to a 2-D Gaussian distribution, which enables to approximate the indifferentiable rotational IoU induced loss by the Gaussian Wasserstein distance (GWD) which can be learned efficiently by gradient back-propagation. GWD can still be informative for learning even there is no overlapping between two rotating bounding boxes which is often the case for small object detection. Thanks to its three unique properties, GWD can also elegantly solve the boundary discontinuity and square-like problem regardless how the bounding box is defined. Experiments on five datasets using different detectors show the effectiveness of our approach, and codes are available at https://github.com/yangxue0827/RotationDetection. Xue Yang 0005, Junchi Yan, Qi Ming, Wentao Wang 0009, Xiaopeng Zhang 0008, Qi Tian 0001 |
ICML | 4 |
| 2021 | Video Semantic Segmentation via Sparse Temporal TransformerabstractCurrently, video semantic segmentation mainly faces two challenges: 1) the demand of temporal consistency; 2) the balance between segmentation accuracy and inference efficiency. For the first challenge, existing methods usually use optical flow to capture the temporal relation in consecutive frames and maintain the temporal consistency, but the low inference speed by means of optical flow limits the real-time applications. For the second challenge, flow based key frame warping is one mainstream solution. However, the unbalanced inference latency of flow-based key frame warping makes it unsatisfactory for real-time applications. Considering the segmentation accuracy and inference efficiency, we propose a novel Sparse Temporal Transformer (STT) to bridge temporal relation among video frames adaptively, which is also equipped with query selection and key selection. The key selection and query selection strategies are separately applied to filter out temporal and spatial redundancy in our temporal transformer. Specifically, our STT can reduce the time complexity of temporal transformer by a large margin without harming the segmentation accuracy and temporal consistency. Experiments on two benchmark datasets, Cityscapes and Camvid, demonstrate that our method achieves the state-of-the-art segmentation accuracy and temporal consistency with comparable inference speed. Jiangtong Li, Wentao Wang 0009, Junjie Chen 0008, Li Niu 0002, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 2 |
| 2021 | Learning High-Precision Bounding Box for Rotated Object Detection via Kullback-Leibler DivergenceabstractExisting rotated object detectors are mostly inherited from the horizontal detection paradigm, as the latter has evolved into a well-developed area. However, these detectors are difficult to perform prominently in high-precision detection due to the limitation of current regression loss design, especially for objects with large aspect ratios. Taking the perspective that horizontal detection is a special case for rotated object detection, in this paper, we are motivated to change the design of rotation regression loss from induction paradigm to deduction methodology, in terms of the relation between rotation and horizontal detection. We show that one essential challenge is how to modulate the coupled parameters in the rotation regression loss, as such the estimated parameters can influence to each other during the dynamic joint optimization, in an adaptive and synergetic way. Specifically, we first convert the rotated bounding box into a 2-D Gaussian distribution, and then calculate the Kullback-Leibler Divergence (KLD) between the Gaussian distributions as the regression loss. By analyzing the gradient of each parameter, we show that KLD (and its derivatives) can dynamically adjust the parameter gradients according to the characteristics of the object. For instance, it will adjust the importance (gradient weight) of the angle parameter according to the aspect ratio. This mechanism can be vital for high-precision detection as a slight angle error would cause a serious accuracy drop for large aspect ratios objects. More importantly, we have proved that KLD is scale invariant. We further show that the KLD loss can be degenerated into the popular Ln-norm loss for horizontal detection. Experimental results on seven datasets using different detectors show its consistent superiority, and codes are available at https://github.com/yangxue0827/RotationDetection. Xue Yang 0005, Xiaojiang Yang, Jirui Yang, Qi Ming, Wentao Wang 0009, Qi Tian 0001, Junchi Yan |
NeurIPS | 5 |
| 2020 | Image Editing via Segmentation Guided Self-Attention NetworkabstractImage editing is one of the most popular directions in computer vision. Recently, many methods have benefited from the advances in deep learning, showing promising performance in the image editing task by inpainting the editing areas. These methods take advantage of edge information as user guidance to generate the desired content. However, they are suffering from generating color discrepancy and inconsistent boundaries. In this letter, we propose a deep image editing method based on a self-attention network which copies information for each of the small patches from distant spatial locations. The proposed method smooths the image, computes segmentation maps, and utilizes the segmentation information for guiding the self-attention layers to explicitly leverage image features from surrounding areas with similar appearances. Experimental results show that the proposed method achieves better performance, is flexible for different purposes, and is fast for implementation. Jianfu Zhang 0003, Peiming Yang, Wentao Wang 0009, Yan Hong 0001, Liqing Zhang 0001 |
IEEE Signal Process. Lett. | 3 |