Wenjia Wang 0009

dblp:44/6297-9 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0121-3852ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 SceneShine: Illumination-aware Human Scene Gaussian Re-Splatting from Mobile Device Video
abstract
Standard 3DGS falls short in precise relighting and shadowing needed to realistically integrate humans into novel environments. We bridge this gap with SceneShine, an illumination-aware framework designed for seamless composition through physically-based avatar relighting and shadow casting. Relighting human surfaces in in-the-wild videos is inherently ill-posed, often making the simultaneous disentanglement of scene lighting and BRDF properties difficult. We overcome this ambiguity by utilizing a pseudo-global light map prior to guide BRDF parameter decomposition, significantly reducing relighting artifacts. Additionally, we implement point-based ray tracing to manage human-scene occlusions and dynamically update scene colors for accurate shadow casting. We also introduce a new synthetic dataset for evaluation. Extensive experiments show that our method surpasses existing approaches in reconstruction fidelity and identity preservation while achieving highly convincing illumination-aware integration1.
Xuqian Ren, Wenjia Wang 0009, Mai Ngoc Nguyen, Juho Kannala, Esa Rahtu
WACV2
2025 TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization
abstract
Synthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate controllers, each specialized for a specific interaction task. This significantly hinders the ability to tackle a wide variety of challenging HSI tasks that require the integration of multiple skills, e.g. sitting down while carrying an object (see Fig. 1). To address this issue, we present TokenHSI, a single, unified transformer-based policy capable of multi-skill unification and flexible adaptation. The key insight is to model the humanoid proprioception as a separate shared token and combine it with distinct task tokens via a masking mechanism. Such a unified policy enables effective knowledge sharing across skills, thereby facilitating the multi-task training. Moreover, our policy architecture supports variable length inputs, enabling flexible adaptation of learned skills to new scenarios. By training additional task tokenizers, we can not only modify the geometries of interaction targets but also coordinate multiple skills to address complex tasks. The experiments demonstrate that our approach can significantly improve versatility, adaptability, and extensibility in various HSI tasks.
Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang 0009, Buzhen Huang, Bo Dai 0002, Taku Komura, Jingbo Wang 0003
CVPR4
2025 SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Wenjia Wang 0009, Liang Pan, Zhiyang Dou, Jidong Mei, Zhouyingcheng Liao, Yuke Lou, Yifan Wu 0039, Lei Yang 0045, Jingbo Wang 0003, Taku Komura
ICCV1
2024 AiOS: All-in-One-Stage Expressive Human Pose and Shape Estimation
abstract
Expressive human pose and shape estimation (a.k.a. 3D whole-body mesh recovery) involves the human body, hand, and expression estimation. Most existing methods have tack-led this task in a two-stage manner, first detecting the human body part with an off-the-shelf detection model and then in-ferring the different human body parts individually. Despite the impressive results achieved, these methods suffer from 1) loss of valuable contextual information via cropping, 2) introducing distractions, and 3) lacking inter-association among different persons and body parts, inevitably causing performance degradation, especially for crowded scenes. To address these issues, we introduce a novel ali-in-one-stage framework, AiOS, for multiple expressive human pose and shape recovery without an additional human detection step. Specifically, our method is built upon DETR, which treats multi-person whole-body mesh recovery task as a progressive set prediction problem with various sequential detection. We devise the decoder tokens and extend them to our task. Specifically, we first employ a human token to probe a hu-man location in the image and encode global features for each instance, which provides a coarse location for the later transformer block. Then, we introduce a joint-related token to probe the human joint in the image and encoder a fine-grained local feature, which collaborates with the global feature to regress the whole-body mesh. This straightfor-ward but effective model outperforms previous state-of-the-art methods by a 9% reduction in NMVE on AGORA, a 30% reduction in PVE on EHF, a 10% reduction in PVE on ARCTIC, and a 3% reduction in PVE on EgoBody.
Qingping Sun, Ailing Zeng, Wanqi Yin, Wenjia Wang 0009, Haiyi Mei, Andrew Chi-Sing Leung, Ziwei Liu 0002, Lei Yang 0059, Zhongang Cai
CVPR6
2024 EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang 0003, Wenjia Wang 0009, Yuan Liu 0025, Taku Komura, Wenping Wang 0001, Lingjie Liu
ECCV (2)6
2024 MuSHRoom: Multi-Sensor Hybrid Room Dataset for Joint 3D Reconstruction and Novel View Synthesis
abstract
Metaverse technologies demand accurate, real-time, and immersive modeling on consumer-grade hardware for both non-human perception (e.g., drone/robot/autonomous car navigation) and immersive technologies like AR/VR, requiring both structural accuracy and photorealism. However, there exists a knowledge gap in how to apply geometric reconstruction and photorealism modeling (novel view synthesis) in a unified framework. To address this gap and promote the development of robust and immersive modeling and rendering with consumer-grade devices, first, we propose a real-world Multi-Sensor Hybrid Room Dataset (MuSHRoom). Our dataset presents exciting challenges and requires state-of-the-art methods to be cost-effective, robust to noisy data and devices, and can jointly learn 3D reconstruction and novel view synthesis instead of treating them as separate tasks, making them ideal for realworld applications. Second, we benchmark several famous pipelines on our dataset for joint 3D mesh reconstruction and novel view synthesis. Finally, in order to further improve the overall performance, we propose a new method that achieves a good trade-off between the two tasks. Our dataset and benchmark show great potential in promoting the improvements for fusing 3D reconstruction and highquality rendering in a robust and computationally efficient end-to-end fashion. The dataset and code are available at the project website: https://xuqianren.github.io/publications/MuSHRoom/.
Xuqian Ren, Wenjia Wang 0009, Dingding Cai, Tuuli Tuominen, Juho Kannala, Esa Rahtu
WACV2
2023 Zolly: Zoom Focal Length Correctly for Perspective-Distorted Human Mesh Reconstruction
abstract
As it is hard to calibrate single-view RGB images in the wild, existing 3D human mesh reconstruction (3DHMR) methods either use a constant large focal length or estimate one based on the background environment context, which can not tackle the problem of the torso, limb, hand or face distortion caused by perspective camera projection when the camera is close to the human body. The naive focal length assumptions can harm this task with the incorrectly formulated projection matrices. To solve this, we propose Zolly, the first 3DHMR method focusing on perspective-distorted images. Our approach begins with analysing the reason for perspective distortion, which we find is mainly caused by the relative location of the human body to the camera center. We propose a new camera model and a novel 2D representation, termed distortion image, which describes the 2D dense distortion scale of the human body. We then estimate the distance from distortion scale features rather than environment context features. Afterwards, We integrate the distortion feature with image features to reconstruct the body mesh. To formulate the correct projection matrix and locate the human body position, we simultaneously use perspective and weak-perspective projection loss. Since existing datasets could not handle this task, we propose the first synthetic dataset PDHuman and extend two real-world datasets tailored for this task, all containing perspective-distorted human images. Extensive experiments show that Zolly outperforms existing state-of-the-art methods on both perspective-distorted datasets and the standard benchmark (3DPW). Code and dataset will be released at https://wenjiawang0312.github.io/projects/zolly/.
Wenjia Wang 0009, Yongtao Ge, Haiyi Mei, Zhongang Cai, Qingping Sun, Chunhua Shen, Lei Yang 0059, Taku Komura
ICCV1
2022 HuMMan: Multi-modal 4D Human Dataset for Versatile Sensing and Modeling
Zhongang Cai, Daxuan Ren, Ailing Zeng, Zhengyu Lin, Tao Yu 0007, Wenjia Wang 0009, Xiangyu Fan 0002, Yang Gao 0042, Yifan Yu 0003, Liang Pan, Fangzhou Hong, Chen Change Loy, Lei Yang 0045, Ziwei Liu 0002
ECCV (7)6
2021 Segmenting Transparent Objects in the Wild with Transformer
abstract
This work presents a new fine-grained transparent object segmentation dataset, termed Trans10K-v2, extending Trans10K-v1, the first large-scale transparent object segmentation dataset. Unlike Trans10K-v1 that only has two limited categories, our new dataset has several appealing benefits. (1) It has 11 fine-grained categories of transparent objects, commonly occurring in the human domestic environment, making it more practical for real-world application. (2) Trans10K-v2 brings more challenges for the current advanced segmentation methods than its former version. Furthermore, a novel Transformer-based segmentation pipeline termed Trans2Seg is proposed. Firstly, the Transformer encoder of Trans2Seg provides the global receptive field in contrast to CNN's local receptive field, which shows excellent advantages over pure CNN architectures. Secondly, by formulating semantic segmentation as a problem of dictionary look-up, we design a set of learnable prototypes as the query of Trans2Seg's Transformer decoder, where each prototype learns the statistics of one category in the whole dataset. We benchmark more than 20 recent semantic segmentation methods, demonstrating that Trans2Seg significantly outperforms all the CNN-based methods, showing the proposed algorithm's potential ability to solve transparent object segmentation.Code is available in https://github.com/xieenze/Trans2Seg.
Enze Xie, Wenjia Wang 0009, Wenhai Wang, Peize Sun, Hang Xu 0004, Ding Liang, Ping Luo 0002
IJCAI2
2020 Scene Text Image Super-Resolution in the Wild
Wenjia Wang 0009, Enze Xie, Xuebo Liu 0001, Wenhai Wang, Ding Liang, Chunhua Shen, Xiang Bai
ECCV (10)1
2020 Segmenting Transparent Objects in the Wild
Enze Xie, Wenjia Wang 0009, Wenhai Wang, Mingyu Ding, Chunhua Shen, Ping Luo 0002
ECCV (13)2
2019 Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation Network
abstract
Scene text detection, an important step of scene text reading systems, has witnessed rapid development with convolutional neural networks. Nonetheless, two main challenges still exist and hamper its deployment to real-world applications. The first problem is the trade-off between speed and accuracy. The second one is to model the arbitrary-shaped text instance. Recently, some methods have been proposed to tackle arbitrary-shaped text detection, but they rarely take the speed of the entire pipeline into consideration, which may fall short in practical applications. In this paper, we propose an efficient and accurate arbitrary-shaped text detector, termed Pixel Aggregation Network (PAN), which is equipped with a low computational-cost segmentation head and a learnable post-processing. More specifically, the segmentation head is made up of Feature Pyramid Enhancement Module (FPEM) and Feature Fusion Module (FFM). FPEM is a cascadable U-shaped module, which can introduce multi-level information to guide the better segmentation. FFM can gather the features given by the FPEMs of different depths into a final feature for segmentation. The learnable post-processing is implemented by Pixel Aggregation (PA), which can precisely aggregate text pixels by predicted similarity vectors. Experiments on several standard benchmarks validate the superiority of the proposed PAN. It is worth noting that our method can achieve a competitive F-measure of 79.9% at 84.2 FPS on CTW1500.
Wenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang, Wenjia Wang 0009, Tong Lu 0002, Gang Yu 0002, Chunhua Shen
ICCV5