Yuewen Ma

dblp:145/5320 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
12since 2021 · last 2026
0009-0003-7734-7053ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies
abstract
Automating immersive VR scene creation remains a primary research challenge. Existing methods typically rely on complex geometry with post-simplification, resulting in inefficient pi pelines or li mited re alism. In th is paper, we introduce Im merseGen, a novel agent-guided framework for compact and photorealistic world generation that decouples realism from exhaustive geometric modeling. ImmerseGen represents scenes as hierarchical compositions of lightweight geometric proxies with synthesized RGBA textures, facilitating real-time rendering on mobile VR headsets. We propose terrain-conditioned texturing for base world generation, combined with context-aware texturing for scenery, to produce diverse and visually coherent worlds. VLM-based agents employ semantic grid-based analysis for precise asset placement and enrich scenes with multimodal enhancements such as visual dynamics and ambient sound. Experiments and real-time VR applications demonstrate that ImmerseGen achieves superior photorealism, spatial coherence, and rendering efficiency compared to existing methods.
Jinyan Yuan, Bangbang Yang, Panwang Pan, Xuehai Zhang, Xiao Liu 0040, Zhaopeng Cui, Yuewen Ma
IEEE Trans. Vis. Comput. Graph.9
2025 Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
Jiaxin Huang 0012, Sheng Miao, Bangbang Yang, Yuewen Ma, Yiyi Liao
ICCV4
2025 InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction From Cluttered Scenes
abstract
Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes as undifferentiated wholes and fails to recognize complete object from partial observations. In this paper, we propose InstaScene, a new paradigm towards holistic 3D perception of complex scenes with a primary goal: decomposing arbitrary instances while ensuring complete reconstruction. To achieve precise decomposition, we develop a novel spatial contrastive learning by tracing rasterization of each instance across views, significantly enhancing semantic supervision in cluttered scenes. To overcome incompleteness from limited observations, we introduce in-situ generation that harnesses valuable observations and geometric cues, effectively guiding 3D generative models to reconstruct complete instances that seamlessly align with the real world. Experiments on scene decomposition and object completion across complex real-world and synthetic scenes demonstrate that our method achieves superior decomposition accuracy while producing geometrically faithful and visually intact objects.
Zesong Yang, Bangbang Yang, Liyuan Cui, Yuewen Ma, Wenqi Dong, Zhaopeng Cui, Chenxuan Cao, Hujun Bao
ICCV4
2025 HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
abstract
Scene-level 3D generation represents a critical frontier in multimedia and computer graphics. While existing approaches have achieved encouraging progress, they still face challenges such as constrained object diversity and limited support for interactive editing. In this paper, we present HiScene, a novel hierarchical framework that bridges the gap between 2D image generation and 3D object generation and delivers high-fidelity scenes with compositional identities and aesthetic scene content. Our key insight is treating scenes as hierarchical ''objects'' under isometric views, where a room functions as a complex object that can be further decomposed into manipulatable items. This hierarchical approach enables us to generate 3D content that aligns with 2D representations while maintaining compositional structure. To ensure completeness and spatial alignment of each decomposed instance, we develop a video-diffusion-based amodal completion technique that effectively handles occlusions and shadows between objects, and introduce shape prior injection to ensure spatial coherence within the scene. Experimental results demonstrate that our method produces more natural object arrangements and complete object instances suitable for interactive applications, while maintaining physical plausibility and alignment with user inputs.
Wenqi Dong, Bangbang Yang, Zesong Yang, Tao Hu 0011, Hujun Bao, Yuewen Ma, Zhaopeng Cui
ACM Multimedia7
2025 TexPro: Text-Guided PBR Texturing with Procedural Material Modeling
abstract
In this paper, we present TexPro, a novel method for high-fidelity material generation for input 3D meshes given text prompts. Unlike existing text-conditioned texture generation methods that typically generate RGB textures with baked lighting, TexPro is able to produce diverse texture maps via procedural material modeling, which enables physically-based rendering, relighting, and additional benefits inherent to procedural materials. Specifically, we first generate multi-view reference images given the input textual prompt by employing the latest text-to-image model. We then derive texture maps through rendering-based optimization with recent differentiable procedural materials. To this end, we design several techniques to handle the misalignment between the generated multiview images and 3D meshes, and introduce a novel material agent that enhances material classification and matching by exploring both part-level understanding and object-aware material reasoning. Experiments demonstrate the superiority of the proposed method over existing SOTAs, and its capability of relighting.
Ziqiang Dang, Wenqi Dong, Zesong Yang, Bangbang Yang, Yuewen Ma, Zhaopeng Cui
Comput. Vis. Media6
2025 DeferredGS: Decoupled and Relightable Gaussian Splatting With Deferred Shading
abstract
Reconstructing and editing 3D objects and scenes both play crucial roles in computer graphics and computer vision. Neural radiance fields (NeRFs) can achieve realistic reconstruction and editing results but suffer from inefficiency in rendering. Gaussian splatting significantly accelerates rendering by rasterizing Gaussian ellipsoids. However, Gaussian splatting utilizes a single Spherical Harmonic (SH) function to model both texture and lighting, limiting independent editing capabilities of these components. Recently, attempts have been made to decouple texture and lighting with the Gaussian splatting representation but may fail to produce plausible geometry and decomposition results on reflective scenes. Additionally, the forward shading technique they employ introduces noticeable blending artifacts during relighting, as the geometry attributes of Gaussians are optimized under the original illumination and may not be suitable for novel lighting conditions. To address these issues, we introduce DeferredGS, a method for decoupling and relighting the Gaussian splatting representation using deferred shading. To achieve successful decoupling, we model the illumination with a learnable environment map and define additional attributes such as texture parameters and normal direction on Gaussians, where the normal is distilled from a jointly trained signed distance function. More importantly, we apply deferred shading, resulting in more realistic relighting effects compared to previous methods. Both qualitative and quantitative experiments demonstrate the superior performance of DeferredGSin novel view synthesis and relighting tasks.
Tong Wu 0009, Jia-Mu Sun, Yukun Lai, Yuewen Ma, Leif Kobbelt, Lin Gao 0004
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 PGT-NeuS: Progressive-Growing Tri-Plane Representation for Neural Surface Reconstruction
abstract
3D reconstruction from multi-view images is a long-standing problem in computer graphic. Neural 3D reconstruction, especially NeuS and its variants, has improved reconstruction quality compared to traditional methods. However, it is still a challenge for these methods to reconstruct fine-grained geometric details since the spherical harmonic positional encoding lacks the ability to express high-frequency signals. In this paper, we propose a multi-resolution tri-plane feature encoding that leverages the detail reconstruction capabilities of high-resolution tri-plane while using the smoothness of low-resolution tri-plane to suppress high-frequency artifacts. Additionally, a progressive training strategy is introduced, gradually merging scene details from coarse to fine granularity, enhancing reconstruction quality while maintaining training stability and reducing difficulty. Furthermore, to address reconstruction challenges arising from sparse viewpoints and inconsistent lighting in image datasets, we introduce normal priors as supervision and propose consistency verification for multi-view normal priors, which assesses the accuracy of normal priors and effectively supervise the reconstructed surfaces. Moreover, we propose a perturbing and fine-tuning strategy on regions of unreliable normal priors to further improve the quality of geometric surface reconstruction.
Xue-Kun Xiang, Yu-Jie Yuan, Wenbo Hu 0002, Yuewen Ma, Lin Gao 0004
IEEE Trans. Vis. Comput. Graph.5
2024 DreamSpace: Dreaming Your Room Space with Text-Driven Panoramic Texture Propagation
abstract
Diffusion-based methods have achieved prominent success in generating 2D media. However, accomplishing similar proficiencies for scene-level mesh texturing in 3D spatial applications, e.g., XR/VR, remains constrained, primarily due to the intricate nature of 3D geometry and the necessity for immersive free-viewpoint rendering. In this paper, we propose a novel indoor scene texturing framework, which delivers text-driven texture generation with enchanting details and authentic spatial coherence. The key insight is to first imagine a stylized 360° panoramic texture from the central viewpoint of the scene, and then propagate it to the rest areas with inpainting and imitating techniques. To ensure meaningful and aligned textures to the scene, we develop a novel coarse-to-fine panoramic texture generation approach with dual texture alignment, which both considers the geometry and texture cues of the captured scenes. To survive cluttered geometries during texture propagation, we design a separated strategy, which conducts texture inpainting in visible regions and then learns an implicit imitating network to synthesize textures in occluded and tiny structural areas. Extensive experiments and the immersive VR application on real-world indoor scenes demonstrate the high quality of the generated textures and the engaging experience on VR headsets. Project webpage: https://ybbbbt.com/publication/dreamspace.
Bangbang Yang, Wenqi Dong, Wenbo Hu 0002, Xiao Liu 0040, Zhaopeng Cui, Yuewen Ma
VR7
2024 AvatarWild: Fully controllable head avatars in the wild
abstract
Recent advancements in the field have resulted in significant progress in achieving realistic head reconstruction and manipulation using neural radiance fields (NeRF). Despite these advances, capturing intricate facial details remains a persistent challenge. Moreover, casually captured input, involving both head poses and camera movements, introduces additional difficulties to existing methods of head avatar reconstruction. To address the challenge posed by video data captured with camera motion, we propose a novel method, AvatarWild, for reconstructing head avatars from monocular videos taken by consumer devices. Notably, our approach decouples the camera pose and head pose, allowing reconstructed avatars to be visualized with different poses and expressions from novel viewpoints. To enhance the visual quality of the reconstructed facial avatar, we introduce a view-dependent detail enhancement module designed to augment local facial details without compromising viewpoint consistency. Our method demonstrates superior performance compared to existing approaches, as evidenced by reconstruction and animation results on both multi-view and single-view datasets. Remarkably, our approach stands out by exclusively relying on video data captured by portable devices, such as smartphones. This not only underscores the practicality of our method but also extends its applicability to real-world scenarios where accessibility and ease of data capture are crucial.
Shaoxu Meng, Tong Wu 0009, Yuewen Ma, Wenbo Hu 0002, Lin Gao 0004
Vis. Informatics5
2023 Tri-MipRF: Tri-Mip Representation for Efficient Anti-Aliasing Neural Radiance Fields
abstract
Despite the tremendous progress in neural radiance fields (NeRF), we still face a dilemma of the trade-off between quality and efficiency, e.g., MipNeRF [3] presents fine-detailed and anti-aliased renderings but takes days for training, while Instant-ngp [36] can accomplish the reconstruction in a few minutes but suffers from blurring or aliasing when rendering at various distances or resolutions due to ignoring the sampling area. To this end, we propose a novel Tri-Mip encoding (à la "mipmap") that enables both instant reconstruction and anti-aliased high-fidelity rendering for neural radiance fields. The key is to factorize the pre-filtered 3D feature spaces in three orthogonal mipmaps. In this way, we can efficiently perform 3D area sampling by taking advantage of 2D pre-filtered feature maps, which significantly elevates the rendering quality without sacrificing efficiency. To cope with the novel Tri-Mip representation, we propose a cone-casting rendering technique to efficiently sample anti-aliased 3D features with the Tri-Mip encoding considering both pixel imaging and observing distance. Extensive experiments on both synthetic and real-world datasets demonstrate our method achieves state-of-the-art rendering quality and reconstruction speed while maintaining a compact representation that reduces 25% model size compared against Instant-ngp. Code is available at the project webpage: https://wbhu.github.io/projects/Tri-MipRF
Wenbo Hu 0002, Bangbang Yang, Lin Gao 0004, Xiao Liu 0040, Yuewen Ma
ICCV7
2023 Interactive NeRF Geometry Editing With Shape Priors
abstract
Neural Radiance Fields (NeRFs) have shown great potential for tasks like novel view synthesis of static 3D scenes. Since NeRFs are trained on a large number of input images, it is not trivial to change their content afterwards. Previous methods to modify NeRFs provide some control but they do not support direct shape deformation which is common for geometry representations like triangle meshes. In this paper, we present a NeRF geometry editing method that first extracts a triangle mesh representation of the geometry inside a NeRF. This mesh can be modified by any 3D modeling tool (we use ARAP mesh deformation). The mesh deformation is then extended into a volume deformation around the shape which establishes a mapping between ray queries to the deformed NeRF and the corresponding queries to the original NeRF. The basic shape editing mechanism is extended towards more powerful and more meaningful editing handles by generating box abstractions of the NeRF shapes which provide an intuitive interface to the user. By additionally assigning semantic labels, we can even identify and combine parts from different objects. We demonstrate the performance and quality of our method in a number of experiments on synthetic data as well as real captured scenes.
Yu-Jie Yuan, Yang-Tian Sun, Yukun Lai, Yuewen Ma, Rongfei Jia, Leif Kobbelt, Lin Gao 0004
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 NeRF-Editing: Geometry Editing of Neural Radiance Fields
abstract
Implicit neural rendering, especially Neural Radiance Field (NeRF), has shown great potential in novel view synthesis of a scene. However, current NeRF-based methods cannot enable users to perform user-controlled shape deformation in the scene. While existing works have proposed some approaches to modify the radiance field according to the user's constraints, the modification is limited to color editing or object translation and rotation. In this paper, we propose a method that allows users to perform controllable shape deformation on the implicit representation of the scene, and synthesizes the novel view images of the edited scene without re-training the network. Specifically, we establish a correspondence between the extracted explicit mesh representation and the implicit neural representation of the target scene. Users can first utilize well-developed mesh-based deformation methods to deform the mesh representation of the scene. Our method then utilizes user edits from the mesh representation to bend the camera rays by introducing a tetrahedra mesh as a proxy, obtaining the rendering results of the edited scene. Extensive experiments demonstrate that our framework can achieve ideal editing results not only on synthetic data, but also on real scenes captured by users.
Yu-Jie Yuan, Yang-Tian Sun, Yukun Lai, Yuewen Ma, Rongfei Jia, Lin Gao 0004
CVPR4
2020 Tetrahedral mesh deformation with positional constraints
Wenjing Zhang 0009, Yuewen Ma, Jianmin Zheng, William J. Allen
Comput. Aided Geom. Des.2
2020 Robust Computation of 3D Apollonius Diagrams
abstract
Abstract Apollonius diagrams, also known as additively weighted Voronoi diagrams, are an extension of Voronoi diagrams, where the weighted distance is defined by the Euclidean distance minus the weight. The bisectors of Apollonius diagrams have a hyperbolic form, which is fundamentally different from traditional Voronoi diagrams and power diagrams. Though robust solvers are available for computing 2D Apollonius diagrams, there is no practical approach for the 3D counterpart. In this paper, we systematically analyze the structural features of 3D Apollonius diagrams, and then develop a fast algorithm for robustly computing Apollonius diagrams in 3D. Our algorithm consists of vertex location, edge tracing and face extraction, among which the key step is to adaptively subdivide the initial large box into a set of sufficiently small boxes such that each box contains at most one Apollonius vertex. Finally, we use centroidal Voronoi tessellation (CVT) to discretize the curved bisectors with well‐tessellated triangle meshes. We validate the effectiveness and robustness of our algorithm through extensive evaluation and experiments. We also demonstrate an application on computing centroidal Apollonius diagram.
Peihui Wang, Yuewen Ma, Shi-Qing Xin, Ying He 0001, Shuang-Min Chen, Jian Xu 0023, Wenping Wang 0001
Comput. Graph. Forum3
2015 Foldover-Free Mesh Warping for Constrained Texture Mapping
abstract
Mapping texture onto 3D meshes with positional constraints is a popular technique that can effectively enhance the visual realism of geometric models. Such a process usually requires constructing a valid mesh embedding satisfying a set of positional constraints, which is known to be a challenging problem. This paper presents a novel algorithm for computing a foldover-free piecewise linear mapping with exact positional constraints. The algorithm begins with an unconstrained planar embedding, followed by iterative constrained mesh transformations. At the heart of the algorithm are radial basis function (RBF)-based warping and the longest edge bisection (LEB)-based refinement. A delicate integration of the RBF-based warping and the LEB-based refinement provides a provably-foldover-free, smooth constrained mesh warping, which can handle a large number of constraints and output a visually pleasing mapping result without extra smoothing optimization. The experiments demonstrate the effectiveness of the proposed algorithm.
Yuewen Ma, Jianmin Zheng
IEEE Trans. Vis. Comput. Graph.1
2013 Inversion Free and Topology Compatible Tetrahedral Mesh Warping Driven by Boundary Surface Deformation
abstract
Warping a tetrahedral mesh driven by boundary surface deformation is useful in many applications. Although some methods have been developed to transform the mesh to conform to the deformed boundary surface, it is still a challenging problem to construct an inversion free warped mesh maintaining a compatible topology. In this paper, these two problems are solved by a novel method that combines radial basis function (RBF)-based warping and adaptive mesh refinement. We iteratively transform the mesh using RBF-based warping with a safe step size to ensure that no element is inverted. The use of the RBF-based warping ensures a smooth warping and thus generates a high-quality warped volumetric mesh. To avoid too small step sizes, we refine the elements that are potentially inverted. The refinement is performed on the original and the warped meshes in the same way so as to maintain compatible topology between them. The results of our method can be used in many areas such as finite element simulation and shape interpolation. We demonstrate the effectiveness of our method with a set of examples.
Wenjing Zhang 0009, Yuewen Ma, Jianmin Zheng
CAD/Graphics2