Yuchi Huo

dblp:138/3450 · DBLP profile ↗
← Back
68ranked-venue papers
5as first author
61since 2021 · last 2026
0000-0003-3296-7999ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 5 first-author · 53 since 2021Artificial intelligence and machine learning · 21 · 20 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiDAR-GS++: Improving LiDAR Gaussian Reconstruction via Diffusion Priors
abstract
Recent GS-based rendering has made significant progress for LiDAR, surpassing Neural Radiance Fields (NeRF) in both quality and speed. However, these methods exhibit artifacts in extrapolated novel view synthesis due to the incomplete reconstruction from single traversal scans. To address this limitation, we present LiDAR-GS++, a LiDAR Gaussian Splatting reconstruction method enhanced by diffusion priors for real-time and high-fidelity re-simulation on public urban roads. Specifically, we introduce a controllable LiDAR generation model conditioned on coarsely extrapolated rendering to produce extra geometry-consistent scans and employ an effective distillation mechanism for expansive LiDAR Gaussian reconstruction. By extending reconstruction to under-fitted regions, our approach ensures global geometric consistency for extrapolative novel views while preserving detailed scene surfaces captured by sensors. Experiments on multiple public datasets demonstrate that LiDAR-GS++ achieves state-of-the-art performance for both interpolated and extrapolated viewpoints, surpassing existing GS and NeRF-based methods.
Jiarun Liu, Rengan Xie, Sicong Du, Yiru Zhao, Yuchi Huo, Sheng Yang 0007
AAAI7
2026 PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos
abstract
We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g. garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatars support downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach.
Dianbing Xi, Guoyuan An, Jingsen Zhu, Ruiyuan Zhang, Jiayuan Lu, Yuchi Huo, Rui Wang 0004
AAAI8
2026 OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
abstract
In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff , aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.
Dianbing Xi, Jiepeng Wang 0005, Yuanzhi Liang, Xi Qiu, Yuchi Huo, Rui Wang 0004, Chi Zhang 0012, Xuelong Li 0001
AAAI5
2026 Corrigendum to "LDM: Large tensorial SDF model for textured mesh generation" [Graphical Models, Volume 140, August 2025, 101271]
Rengan Xie, Xiaoliang Luo, Lvchun Wang, Qi Wang 0111, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Yuchi Huo
Graph. Model.10
2025 HR Human: Modeling Human Avatars with Triangular Mesh and High-Resolution Textures from Videos
Yuchi Huo, Qi Wang 0111, Wenting Zheng, Rengan Xie
CVM (2)3
2025 Hand-held Object Reconstruction from RGB Video with Dynamic Interaction
abstract
This work aims to reconstruct the 3D geometry of a rigid object manipulated by one or both hands using monocular RGB video. Previous methods rely on Structure-from-Motion or hand priors to estimate relative motion between the object and camera, which typically assume textured objects or single-hand interactions. To accurately recover object geometry in dynamic interactions, we incorporate priors from 3D generation model into object pose estimation and propose semantic consistency constraints to solve the challenge of shape and texture discrepancy between the generated priors and observations. The poses are initialized, followed by joint optimization of the object poses and implicit neural representation. During optimization, a novel pose outlier voting strategy with inter-view consistency is proposed to correct large pose errors. Experiments on three datasets demonstrate that our method significantly outperforms the state-of-the-art in reconstruction quality for both single- and two-hand scenarios. Our project page: https://east-j.github.io/dynhor/
Shijian Jiang, Qi Ye 0001, Rengan Xie, Yuchi Huo, Jiming Chen 0001
CVPR4
2025 A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting
Zhiyuan Fang, Rengan Xie, Xuancheng Jin, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Rui Wang 0004, Yuchi Huo
ICCV8
2025 IntrinsicControlNet: Cross-Distribution Image Generation with Real and Unreal
Jiayuan Lu, Rengan Xie, Zhizhen Wu, Dianbing Xi, Qi Ye 0001, Rui Wang 0004, Hujun Bao, Yuchi Huo
ICCV9
2025 Inverse Rendering using Multi-Bounce Path Tracing and Reservoir Sampling
abstract
We introduce MIRReS, a novel two-stage inverse rendering framework that jointly reconstructs and optimizes explicit geometry, materials, and lighting from multi-view images. Unlike previous methods that rely on implicit irradiance fields or oversimplified ray tracing, our method begins with an initial stage that extracts an explicit triangular mesh. In the second stage, we refine this representation using a physically-based inverse rendering model with multi-bounce path tracing and Monte Carlo integration. This enables our method to accurately estimate indirect illumination effects, including self-shadowing and internal reflections, leading to a more precise intrinsic decomposition of shape, material, and lighting. To address the noise issue in Monte Carlo integration, we incorporate reservoir sampling, improving convergence and enabling efficient gradient-based optimization with low sample counts. Through both qualitative and quantitative assessments across various scenarios, especially those with complex shadows, we demonstrate that our method achieves state-of-the-art decomposition performance. Furthermore, our optimized explicit geometry seamlessly integrates with modern graphics engines supporting downstream applications such as scene editing, relighting, and material editing.
Yuxin Dai, Qi Wang 0111, Jingsen Zhu, Dianbing Xi, Yuchi Huo, Chen Qian 0006, Ying He 0001
ICLR5
2025 Leveraging Pretrained Diffusion Models for Zero-Shot Part Assembly
abstract
3D part assembly aims to understand part relationships and predict their 6-DoF poses to construct realistic 3D shapes, addressing the growing demand for autonomous assembly, which is crucial for robots. Existing methods mainly estimate the transformation of each part by training neural networks under supervision, which requires a substantial quantity of manually labeled data. However, the high cost of data collection and the immense variability of real-world shapes and parts make traditional methods impractical for large-scale applications. In this paper, we propose first a zero-shot part assembly method that utilizes pre-trained point cloud diffusion models as discriminators in the assembly process, guiding the manipulation of parts to form realistic shapes. Specifically, we theoretically demonstrate that utilizing a diffusion model for zero-shot part assembly can be transformed into an Iterative Closest Point (ICP) process. Then, we propose a novel pushing-away strategy to address the overlap parts, thereby further enhancing the robustness of the method. To verify our work, we conduct extensive experiments and quantitative comparisons to several strong baseline methods, demonstrating the effectiveness of the proposed approach, which even surpasses the supervised learning method. The code has been released on https://github.com/Ruiyuan-Zhang/Zero-Shot-Assembly.
Ruiyuan Zhang, Qi Wang 0111, Yuchi Huo, Chao Wu 0001
IJCAI4
2025 Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion
abstract
Recently, generating 3D assets with the control of condition images has achieved impressive quality. However, existing 3D generation methods are limited to handling a single control objective and lack the ability to utilize multiple images to independently control different regions of a 3D asset, which hinders their flexibility in applications. We propose Fuse3D, a novel method that enables generating 3D assets under the control of multiple images, allowing for the seamless fusion of multi-level regional controls from global views to intricate local details. First, we introduce a Multi-Condition Fusion Module to integrate the visual features from multiple image regions. Then, we propose a method to automatically align user-selected 2D image regions with their associated 3D regions based on semantic cues. Finally, to resolve control conflicts and enhance local control features from multi-condition images, we introduce a Local Attention Enhancement Strategy that flexibly balances region-specific feature fusion. Overall, we introduce the first method capable of controllable 3D asset generation from multiple condition images. The experimental results indicate that Fuse3D can flexibly fuse multiple 2D image regions into coherent 3D structures, resulting in high-quality 3D assets. Code and data for this paper are at https://jinnmnm.github.io/Fuse3d.github.io/.
Xuancheng Jin, Rengan Xie, Wenting Zheng, Rui Wang 0004, Hujun Bao, Yuchi Huo
SIGGRAPH Asia6
2025 NeLiF: Neural Lighting Function Generation for Real-Time Indoor Rendering
abstract
Recent advances in neural rendering have mainly focused on modeling radiance fields with neural representations, often overlooking the underlying mechanisms for producing various lighting effects, and consequently leading to the limited adaptability to dynamic scenes. Lighting effects, such as highlights, shadows, and indirect illuminations, are typically computed using physically-based rendering methods like path tracing, which can be computationally intensive for complex indoor luminaires. Although several recent studies have aimed to model global illumination effects with neural representations, they commonly suffer from long training times or poor generalizability to new scenes. Addressing these challenges, this work presents a novel neural lighting function generation model capable of synthesizing diverse lighting effects in real time for unseen dynamic scenes and complex indoor luminaires, achieving results comparable to state-of-the-art rendering pipelines. Our model operates in two stages. First, multi-view observation images of the luminaire are captured to encode a compact, scene-independent 3D neural lighting field. Subsequently, light information is sampled from this neural lighting field and integrated with G-buffers and shadow clues to produce the shading results. In parallel, we employ a state-of-the-art generative model together with our training-free Inverse HDR Splatting module to generate HDR 3D Gaussians representing the luminaire. This strategy capitalizes on the powerful generalization capabilities of advanced generative models, enabling efficient and accurate appearance reconstruction for a diverse range of complex luminaires. In our experiments, the model trained on a dataset of 10,000 modern indoor scenes and thousands of illuminations demonstrates strong generalizability, high efficiency, and visually convincing results across a wide range of test scenes, highlighting its potential as a practical and flexible solution for high-fidelity, real-time neural indoor rendering.
Hongtao Sheng, Yuchi Huo, Chuankun Zheng, Guangzhi Han, Yifan Peng 0001, Bin Zang, Hao Zhu 0004, Rui Tang 0015, Rui Wang 0004, Hujun Bao
SIGGRAPH Asia2
2025 AniTex: Light-Geometry Consistent PBR Material Generation for Animatable Objects
abstract
High-quality Physically-Based Rendering (PBR) materials are crucial for visual realism in 3D asset creation, yet existing methods primarily target static objects, leading to challenges in maintaining multi-frame consistency for animatable entities. To tackle this issue, we introduce AniTex, the first generative pipeline that utilizes diffusion models to synthesize high-quality PBR materials for animatable objects based on text prompts. The pipeline consists of three key stages: First, sequences of RGB images are generated using a video diffusion model conditioned on depth, normals, irradiance, and motion vectors to ensure temporal coherence and geometric alignment across multiple frames and viewpoints. Second, these RGB image sequences are decomposed into per-view, per-frame PBR material maps (albedo, roughness, metallic) by a specialized Intrinsic Diffusion Model (IDM), which is conditioned on the RGB images along with consistent geometry and lighting cues to disentangle material from illumination. Finally, these per-view, per-frame PBR maps are hierarchically blended. This process first ensures temporal coherence within each view’s frame sequence, then amalgamates these into globally consistent PBR materials for the animatable object, maintaining overall temporal coherence and visual consistency throughout its animation. Extensive experiments show that AniTex produces more realistic PBR materials for both static and animated objects, outperforming baseline methods in visual appeal.
Jieting Xu, Guoyuan An, Rengan Xie, Dianbing Xi, Wenjun Song, Rui Wang 0004, Yuchi Huo
SIGGRAPH Asia10
2025 StereoFG: Generating Stereo Frames from Centered Feature Stream
abstract
In recent years, the community has seen the emergence of neural-based super-resolution and frame generation techniques. These methods have effectively sped up high-resolution rendering by exploiting the spatial and temporal coherence between sequential frames, but none of them are designed specifically for improving the rendering performance in VR applications, where stereo rendering doubles the rendering cost.
Chenyu Zuo, Yazhen Yuan, Zhizhen Wu, Jingzhen Lan, Ming Fu, Yuchi Huo, Rui Wang 0004
SIGGRAPH Asia7
2025 LDM: Large tensorial SDF model for textured mesh generation
abstract
Previous efforts have managed to generate production-ready 3D assets from text or images. However, these methods primarily employ NeRF or 3D Gaussian representations, which are not adept at producing smooth, high-quality geometries required by modern rendering pipelines. In this paper, we propose LDM, a L arge tensorial S D F M odel, which introduces a novel feed-forward framework capable of generating high-fidelity, illumination-decoupled textured mesh from a single image or text prompts. We firstly utilize a multi-view diffusion model to generate sparse multi-view inputs from single images or text prompts, and then a transformer-based model is trained to predict a tensorial SDF field from these sparse multi-view image inputs. Finally, we employ a gradient-based mesh optimization layer to refine this model, enabling it to produce an SDF field from which high-quality textured meshes can be extracted. Extensive experiments demonstrate that our method can generate diverse, high-quality 3D mesh assets with corresponding decomposed RGB textures within seconds. The project code is available at https://github.com/rgxie/LDM .
Rengan Xie, Xiaoliang Luo, Lvchun Wang, Qi Wang 0111, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Yuchi Huo
Graph. Model.10
2025 Ultra-High Resolution Facial Texture Reconstruction from a Single Image
abstract
Advances in mobile cameras have made it easier to capture ultra-high resolution (UHR) portraits. However, existing face reconstruction methods lack specific adaptations for UHR input (e.g., 4096 × 4096), leading to under-use of high-frequency details that are crucial for achieving photorealistic rendering. Our method supports 4096 × 4096 UHR input and utilizes a divide-and-conquer approach for end-to-end 4K albedo, micronormal, and specular texture reconstruction at the original resolution. We employ a two-stage strategy to capture both global distributions and local high-frequency details, effectively mitigating mosaic and seam artifacts common in patch-based prediction. Additionally, we innovatively apply hash encoding to facial U-V coordinates to boost the model’s ability to learn regional high-frequency feature distributions. Our method can be easily incorporated in state-of-the-art facial geometry reconstruction pipelines, significantly improving the texture reconstruction quality, facilitating artistic creation workflows.
Hongxiang Huang, Guoyuan An, Jingzhen Lan, Qi Wang 0111, Rui Wang 0004, Yuchi Huo
Comput. Vis. Media7
2025 A Biophysical-Based Skin Model for Heterogeneous Volume Rendering
abstract
Realistic human skin rendering has been a long-standing challenge in computer graphics. Recently, biophysical-based skin rendering has received increasing attention, as it provides a more realistic skin-rendering and a more intuitive way to adjust the skin style. In this work, we present a novel heterogeneous biophysical-based volume rendering method for human skin that improves the realism of skin appearance while easily simulating various types of skin effects, including skin diseases, by modifying biological coefficient textures. Specifically, we introduce a two-layer skin representation by mesh deformation that explicitly models the epidermis and dermis with heterogeneous volumetric medium layers containing the corresponding spatially varying melanin and hemoglobin, respectively. Furthermore, to better facilitate skin acquisition, we introduced a learning-based framework that automatically estimates spatially varying biological coefficients from an albedo texture, enabling biophysical-based and intuitive editing, such as tanning, pathological vitiligo, and freckles. We illustrated the effects of multiple skin-editing applications and demonstrated superior quality to the commonly used random walk skin-rendering method, with more convincing skin details regarding subsurface scattering.
Qi Wang 0111, Fujun Luan, Yuxin Dai, Yuchi Huo, Hujun Bao, Rui Wang 0004
Comput. Vis. Media4
2025 Toward Weather-Robust 3D Human Body Reconstruction: Millimeter-Wave Radar-Based Dataset, Benchmark, and Multi-Modal Fusion
abstract
3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementarily, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for weather-robust 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. The limited research about the impact of missing points and sparsity features of mmWave data on reconstruction performance, as well as the lack of available datasets for paired mmWave-RGB data, further complicates the process of fusing the two modalities. To fill these gaps, we build up an automatic 3D body annotation system with multiple sensors to collect a large-scale mmWave dataset. The dataset consists of synchronized and calibrated mmWave radar point clouds and RGB(D) images under different weather conditions and skeleton/mesh annotations for humans in these scenes. With this dataset, we conduct a comprehensive analysis about the limitations of single-modality reconstruction and the impact of missing points and sparsity on the reconstruction performance. Based on the guidance of this analysis, we design ImmFusion, the first mmWave-RGB fusion solution to robustly reconstruct 3D human bodies in various weather conditions. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve robust 3D human body reconstruction in various weather environments. In addition, our method achieves superior accuracy compared to that of the state-of-the-art Transformer-based LiDAR-camera fusion methods.
Anjun Chen, Kun Shi 0003, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 CT-NeRF: Incremental Optimization of Neural Radiance Field and Camera Poses With Complex Trajectory
abstract
Neural radiance field (NeRF) has achieved impressive results in high-quality 3D scene reconstruction. However, NeRF heavily relies on precise camera poses. While recent works like BARF have introduced camera pose optimization within NeRF, their applicability is limited to simple trajectory scenes. Existing methods struggle while tackling complex trajectories involving large rotations. To address this limitation, we propose CT-NeRF, an incremental reconstruction and optimization pipeline using only RGB images without pose and depth input. In this pipeline, we first propose a local-global bundle adjustment under a pose graph connecting neighboring frames to enforce the consistency between poses to escape the local minima caused by only pose consistency with the scene structure. Further, we instantiate the consistency between poses as a reprojection error constraint resulting from pixel-level correspondences between input image pairs. Through the incremental reconstruction, CT-NeRF enables the recovery of both camera poses and scene structure and is capable of handling scenes with complex trajectories. We evaluate the performance of CT-NeRF on two real-world datasets, NeRF-Buster and Free-Dataset, which feature complex trajectories. Results show CT-NeRF outperforms existing methods in novel view synthesis and pose estimation accuracy.
Yunlong Ran, Yanxu Li, Qi Ye 0001, Yuchi Huo, Zhaopeng Cui, Zechun Bai, Jiming Chen 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 An Approach to Multi-AAV Ship Detection Based on Mobile Edge Computing Scenarios
abstract
Autonomous aerial vehicles (AAVs) are widely used for ship tracking and detection tasks. However, the real-time detection performance is limited by AAV battery capacity and computing power, resulting in a short operational duration. To address this challenge, this paper proposes a AAV ship detection system that focuses on two key aspects: algorithm improvement and computational resource allocation. Specifically, we introduce a lightweight ship detection method tailored for multi-AAV scenarios in a mobile edge computing environment. The proposed method first designs a multi-disentangled knowledge distillation approach based on an information decoupling framework and utilizes a newly designed teacher network to enhance the lightweight detection model. The teacher network disentangles two key types of entanglements: the relationship between the convolutional filters and target categories, and the relationship between the foreground and background regions in the feature maps. Additionally, a proximal policy optimization (PPO) reinforcement learning algorithm is designed to enable real-time decision-making for AAV motion, detection accuracy, and computational offloading. Finally, we validate the superiority of the proposed knowledge distillation method and demonstrate the robustness and effectiveness of the AAV path planning algorithm in various scenarios through a series of experiments. Compared to the improved student models YOLOv8-N and YOLOv10-N, our method improves [email protected] by 1.2% and 1.1% on the SeaShips7000 and FVessel validation sets. Furthermore, compared to the existing methods K-Means and DBSCAN, our approach achieves reward values approximately 2.0 times and 1.4 times higher, respectively.
Tao Liu 0016, Zhengling Lei, Yuchi Huo, Xiaocai Zhang, Gaoqi He, Huafeng Wu
IEEE Trans. Intell. Transp. Syst.6
2025 AdaptiveFusion: Adaptive Multi-Modal Multi-View Fusion for 3D Human Body Reconstruction
abstract
Recent advancements in sensor technology and deep learning have led to significant progress in 3D human body reconstruction. However, most existing approaches rely on data from a specific sensor, which can be unreliable due to the inherent limitations of individual sensing modalities. Additionally, existing multi-modal fusion methods generally require customized designs based on the specific sensor combinations or setups, which limits the flexibility and generality of these methods. Furthermore, conventional point-image projection-based and Transformer-based fusion networks are susceptible to the influence of noisy modalities and sensor poses. To address these limitations and achieve robust 3D human body reconstruction in various conditions, we propose AdaptiveFusion, a generic adaptive multi-modal multi-view fusion framework that can effectively incorporate arbitrary combinations of uncalibrated sensor inputs. By treating different modalities from various viewpoints as equal tokens, and our handcrafted modality sampling module by leveraging the inherent flexibility of Transformer models, AdaptiveFusion is able to cope with arbitrary numbers of inputs and accommodate noisy modalities with only a single training network. Extensive experiments on large-scale human datasets demonstrate the effectiveness of AdaptiveFusion in achieving high-quality 3D human body reconstruction in various environments. In addition, our method achieves superior accuracy compared to state-of-the-art fusion methods.
Anjun Chen, Kun Shi 0003, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001
IEEE Trans. Multim.6
2025 OpenSlot: Mixed Open-Set Recognition With Object-Centric Learning
abstract
Existing open-set recognition (OSR) studies typically assume that each image contains only one class label, with the unknown test set (negative) having a disjoint label space from the known test set (positive), a scenario referred to as full-label shift. This paper introduces the mixed OSR problem, where test images contain multiple class semantics, with both known and unknown classes co-occurring in the negatives, leading to a more complex super-label shift that better reflects real-world scenarios. To tackle this challenge, we propose the OpenSlot framework, based on object-centric learning, which uses slot features to represent diverse class semantics and generate class predictions. The proposed anti-noise slot (ANS) technique helps mitigate the impact of noise (invalid or background) slots during classification training, addressing the semantic misalignment between class predictions and ground truth. We evaluate OpenSlot on both mixed and conventional OSR benchmarks. Without elaborate designs, our method not only excels existing approaches in detecting super-label shifts across OSR tasks, but also achieves state-of-the-art performance on conventional benchmarks. Meanwhile, OpenSlot can localize class objects without using bounding boxes during training, demonstrating competitive performance in open-set object detection and potential for generalization.
Xu Yin, Guoyuan An, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Multim.4
2025 Streaming-Aware Neural Monte Carlo Rendering Framework with Unified Denoising-Compression and Client Collaboration
abstract
Recent advances in cloud rendering have brought us a promising alternative for interactive photorealistic rendering on lightweight devices, which used to be only available on high-end platforms equipped with powerful graphic cards. This technique enables users to perform rendering-related creative tasks, such as 3D product visualization and lighting design, from the comfort of any location using handheld devices, rather than being confined to the front of a noisy heat-generating workstation. However, existing large-scale cloud rendering systems that stream path-traced frames from the server to the client present extremely high rendering costs and transmission bandwidth requirements, even with advanced path-tracing acceleration and video compression techniques. To alleviate these problems, we propose a novel streaming-aware rendering framework that is able to learn a joint optimal model integrating two path-tracing acceleration techniques (adaptive sampling and denoising) and video compression technique. Our joint model can fully exploit the inherent connections between these techniques and thus achieve substantially reduced rendering costs and enhanced compression quality. We also introduce the collaboration of client rendering ability to assist the frame decoding by rendering G-buffers as the shared side information. We demonstrate that appropriately incorporating the geometry and material priors from G-buffers into a neural compression pipeline can significantly reduce the streaming bandwidth in a cloud rendering system, and lighten the compression module design for computation efficiency. Our experiments show that our method delivers the best quality at various bitrates compared to existing Monte Carlo rendering streaming schemes, while remaining lightweight and efficient for cross-platform thin clients, including mobiles and tablets.
Hangming Fan, Yuchi Huo, Chuankun Zheng, Chonghao Hu, Yazhen Yuan, Rui Wang 0004
ACM Trans. Graph.2
2025 A Fully-statistical Wave Scattering Model for Heterogeneous Surfaces
abstract
Heterogeneous surfaces exhibit spatially varying geometry and material, and therefore admit diverse appearances. Existing computer graphics works can only model heterogeneity using explicit structures or statistical parameters that describe a coarser level of detail. We extend the boundary by introducing a new model that describes the heterogeneous surfaces fully statistically at the microscopic level, with rich geometry and material details that are comparable to the wavelengths of light. We treat the heterogeneous surfaces as a mixture of stochastic vector processes. We adapt the well-known generalized Harvey-Shack theory to quantify the mean scattered intensity, i.e., the BRDF of these surfaces. We further explore the covariance statistic of the scattered field and derive its rank-1 decomposition. This leads to a practical algorithm that samples the speckles (fluctuating intensities) from the statistics, enriching the appearance without explicit definition of heterogeneous surfaces. The formulations are analytic, and we validate the quantities by comprehensive numerical simulations. Our heterogeneous surface model demonstrates various applications including corrosion (natural), particle deposition (man-made), and height-correlated mixture (artistic). Code for this paper is available at https://github.com/Rendering-at-ZJU/HeteroSurface.
Zhengze Liu, Yuchi Huo, Yifan Peng 0001, Rui Wang 0004
ACM Trans. Graph.2
2025 Consecutive Frame Extrapolation with Predictive Sparse Shading
abstract
The demand for high-frame-rate rendering keeps increasing in modern displays. Existing frame generation and super-resolution techniques accelerate rendering by reducing rendering samples across space or time. However, they rely on a uniform sampling reduction strategy, which undersamples areas with complex details or dynamic shading. To address this, we propose to sparsely shade critical areas while reusing generated pixels in low-variation areas for neural extrapolation. Specifically, we introduce the Predictive Error-Flow-eXtrapolation Network (EFXNet)-an architecture that predicts extrapolation errors, estimates flows, and extrapolates frames at once. Firstly, EFXNet leverages temporal coherence to predict extrapolation error and guide the sparse shading of dynamic areas. In addition, EFXNet employs a target-grid correlation module to estimate robust optical flows from pixel correlations rather than pixel values. Finally, EFXNet uses dedicated motion representations for the historical geometric and lighting components, respectively, to extrapolate temporally stable frames. Extensive experimental results show that, compared with state-of-the-art methods, our frame extrapolation method exhibits superior visual quality and temporal stability under a low rendering budget.
Zhizhen Wu, Yazhen Yuan, Zhilong Yuan, Rui Wang 0004, Yuchi Huo
ACM Trans. Graph.6
2025 MoFlow: Motion-Guided Flows for Recurrent Rendered Frame Prediction
abstract
Rendering realistic images in real-time on high-frame-rate display devices poses considerable challenges, even with advanced graphics cards. This stimulates a demand for frame prediction technologies to boost frame rates. The key to these algorithms is to exploit spatiotemporal coherence by warping rendered pixels with motion representations. However, existing motion estimation methods can suffer from low precision, high overhead, and incomplete support for visual effects. In this article, we present a rendered frame prediction framework with a novel motion representation, dubbed motion-guided flow (MoFlow) , aiming at overcoming the intrinsic limitations of optical flow and motion vectors and precisely capture the dynamics of intricate geometries, lighting, and translucent objects. Notably, we construct MoFlows using a recurrent feature streaming network, which specializes in learning latent motion features from multiple frames. The results of extensive experiments demonstrate that, compared to state-of-the-art methods, our method achieves superior visual quality and temporal stability with lower latency. The recurrent mechanism allows our method to predict single or multiple consecutive frames, increasing the frame rate by over 2×. The proposed approach represents a flexible pipeline to meet the demands of various graphics applications, devices, and scenarios.
Zhizhen Wu, Zhilong Yuan, Chenyu Zuo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Yuchi Huo
ACM Trans. Graph.8
2024 In-Hand 3D Object Reconstruction from a Monocular RGB Video
abstract
Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view images achieved compelling results in the visible part of the object. However, these methods falter in accurately capturing the shape within the hand-object contact region due to occlusion. In this paper, we propose a novel method that deals with surface reconstruction under occlusion by incorporating priors of 2D occlusion elucidation and physical contact constraints. For the former, we introduce an object amodal completion network to infer the 2D complete mask of objects under occlusion. To ensure the accuracy and view consistency of the predicted 2D amodal masks, we devise a joint optimization method for both amodal mask refinement and 3D reconstruction. For the latter, we impose penetration and attraction constraints on the local geometry in contact regions. We evaluate our approach on HO3D and HOD datasets and demonstrate that it outperforms the state-of-the-art methods in terms of reconstruction surface quality, with an improvement of 52% on HO3D and 20% on HOD. Project webpage: https://east-j.github.io/ihor.
Shijian Jiang, Qi Ye 0001, Rengan Xie, Yuchi Huo, Jiming Chen 0001
AAAI4
2024 TPGP: Temporal-Parametric Optimization with Deep Grasp Prior for Dexterous Motion Planning
abstract
Grasping motion planning aims to find a feasible grasping trajectory in the configuration space given an input target grasp. While optimizing grasp motion with two or three-fingered grippers has been well studied, the study on natural grasp motion planning with a dexterous hand remains a very challenging problem due to the high dimensional working space. In this work, we propose a novel temporal-parametric grasp prior (TPGP) optimization method to simplify the difficulty of grasping trajectory optimization for the dexterous hand while maintaining smooth and natural properties of the grasping motion. Specifically, we formulate the discrete trajectory parameters into a temporal-based parameterization, where the prior constraint provided by a hand poser network, is introduced to ensure that hand pose is natural and reasonable throughout the trajectory. Finally, we present a joint target optimization strategy to enhance the target pose for more feasible trajectories. Extensive validations on two public datasets show that our method outperforms state-of-the-art methods regarding grasp motion on various metrics.
Haoming Li 0004, Qi Ye 0001, Yuchi Huo, Qingtao Liu, Shijian Jiang, Jiming Chen 0001
ICRA3
2024 Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction
Qi Wang 0111, Yuchi Huo, Qi Ye 0001, Rui Wang 0004, Hujun Bao
IJCAI2
2024 Neural Global Illumination via Superposed Deformable Feature Fields
Chuankun Zheng, Yuchi Huo, Hongxiang Huang, Hongtao Sheng, Junrong Huang, Rui Tang 0015, Hao Zhu 0004, Rui Wang 0004, Hujun Bao
SIGGRAPH Asia2
2024 Real-Time Polygonal Lighting of Iridescence Effect using Precomputed Monomial-Gaussians
abstract
Abstract The real world consists of mass phenomena, such as iridescence on thin film and metal oxide layers, that is only explicable by wave optics. Existing research can reproduce such effects with simple point lights or low‐frequency environmental lighting. However, it remains a difficult task to efficiently rendering these effects when near‐field, high‐frequency area lights are involved. This paper presents a high‐fidelity, real‐time rendering algorithm for the iridescence effect under polygonal lights. We introduce a novel set of spherical functions, Monomial‐Gaussians, to accurately fit iridescent materials' reflectance. With a precomputed lookup table, the Monomial‐Gaussians are easily integrated over spherical polygons in linear time. Importance sampling of Monomial‐Gaussians is also supported to efficiently reduce Monte‐Carlo error. Our approach produces accurate renderings of the iridescence effect while still preserving high frame rates.
Zhengze Liu, Yuchi Huo, Yinhui Yang, Rui Wang 0004
Comput. Graph. Forum2
2024 Adaptive sampling and reconstruction for gradient-domain rendering
abstract
Gradient-domain rendering estimates finite difference gradients of image intensities and reconstructs the final result by solving a screened Poisson problem, which shows improvements over merely sampling pixel intensities. Adaptive sampling is another orthogonal research area that focuses on distributing samples adaptively in the primal domain. However, adaptive sampling in the gradient domain with low sampling budget has been less explored. Our idea is based on the observation that signals in the gradient domain are sparse, which provides more flexibility for adaptive sampling. We propose a deep-learning-based end-to-end sampling and reconstruction framework in gradient-domain rendering, enabling adaptive sampling gradient and the primal maps simultaneously. We conducted extensive experiments for evaluation and showed that our method produces better reconstruction quality than other methods in the test dataset.
Yuzhi Liang, Tao Liu 0016, Yuchi Huo, Rui Wang 0004, Hujun Bao
Comput. Vis. Media3
2024 An approach to ship target detection based on combined optimization model of dehazing and detection
Tao Liu 0016, Zhengling Lei, Yuchi Huo, Jiansen Zhao, Xiaocai Zhang
Eng. Appl. Artif. Intell.4
2024 Fine-Grained Background Representation for Weakly Supervised Semantic Segmentation
abstract
Generating reliable pseudo masks from image-level labels is challenging in the weakly supervised semantic segmentation (WSSS) task due to the lack of spatial information. Prevalent class activation map (CAM)-based solutions are challenged to discriminate the foreground (FG) objects from the suspicious background (BG) pixels (a.k.a. co-occurring) and learn the integral object regions. This paper proposes a simple fine-grained background representation (FBR) method to discover and represent diverse BG semantics and address the co-occurring problems. We abandon using the class prototype or pixel-level features for BG representation. Instead, we develop a novel primitive, negative region of interest (NROI), to capture the fine-grained BG semantic information and conduct the pixel-to-NROI contrast to distinguish the confusing BG pixels. We also present an active sampling strategy to mine the FG negatives on-the-fly, enabling efficient pixel-to-pixel intra-foreground contrastive learning to activate the entire object region. Thanks to the simplicity of design and convenience in use, our proposed method can be seamlessly plugged into various models, yielding new state-of-the-art results under various WSSS settings across benchmarks. Leveraging solely image-level (I) labels as supervision, our method achieves 73.2 mIoU and 45.6 mIoU segmentation results on Pascal Voc and MS COCO test sets, respectively. Furthermore, by incorporating saliency maps as an additional supervision signal (I+S), we attain 74.9 mIoU on Pascal Voc test set. Concurrently, our FBR approach demonstrates meaningful performance gains in weakly-supervised instance segmentation (WSIS) tasks, showcasing its robustness and strong generalization capabilities across diverse domains.
Xu Yin, Woobin Im, Dongbo Min, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Circuits Syst. Video Technol.4
2024 Neural Kernel Regression for Consistent Monte Carlo Denoising
abstract
Unbiased Monte Carlo path tracing that is extensively used in realistic rendering produces undesirable noise, especially with low samples per pixel (spp). Recently, several methods have coped with this problem by importing unbiased noisy images and auxiliary features to neural networks to either predict a fixed-sized kernel for convolution or directly predict the denoised result. Since it is impossible to produce arbitrarily high spp images as the training dataset, the network-based denoising fails to produce high-quality images under high spp. More specifically, network-based denoising is inconsistent and does not converge to the ground truth as the sampling rate increases. On the other hand, the post-correction estimators yield a blending coefficient for a pair of biased and unbiased images influenced by image errors or variances to ensure the consistency of the denoised image. As the sampling rate increases, the blending coefficient of the unbiased image converges to 1, that is, using the unbiased image as the denoised results. However, these estimators usually produce artifacts due to the difficulty of accurately predicting image errors or variances with low spp. To address the above problems, we take advantage of both kernel-predicting methods and post-correction denoisers. A novel kernel-based denoiser is proposed based on distribution-free kernel regression consistency theory, which does not explicitly combine the biased and unbiased results but constrain the kernel bandwidth to produce consistent results under high spp. Meanwhile, our kernel regression method explores bandwidth optimization in the robust auxiliary feature space instead of the noisy image space. This leads to consistent high-quality denoising at both low and high spp. Experiment results demonstrate that our method outperforms existing denoisers in accuracy and consistency.
Pengju Qiao, Qi Wang 0111, Yuchi Huo, Shiji Zhai, Wei Hua 0002, Hujun Bao, Tao Liu 0016
ACM Trans. Graph.3
2024 LightFormer: Light-Oriented Global Neural Rendering in Dynamic Scene
abstract
The generation of global illumination in real time has been a long-standing challenge in the graphics community, particularly in dynamic scenes with complex illumination. Recent neural rendering techniques have shown great promise by utilizing neural networks to represent the illumination of scenes and then decoding the final radiance. However, incorporating object parameters into the representation may limit their effectiveness in handling fully dynamic scenes. This work presents a neural rendering approach, dubbed LightFormer , that can generate realistic global illumination for fully dynamic scenes, including dynamic lighting, materials, cameras, and animated objects, in real time. Inspired by classic many-lights methods, the proposed approach focuses on the neural representation of light sources in the scene rather than the entire scene, leading to the overall better generalizability. The neural prediction is achieved by leveraging the virtual point lights and shading clues for each light. Specifically, two stages are explored. In the light encoding stage, each light generates a set of virtual point lights in the scene, which are then encoded into an implicit neural light representation, along with screen-space shading clues like visibility. In the light gathering stage, a pixel-light attention mechanism composites all light representations for each shading point. Given the geometry and material representation, in tandem with the composed light representations of all lights, a lightweight neural network predicts the final radiance. Experimental results demonstrate that the proposed LightFormer can yield reasonable and realistic global illumination in fully dynamic scenes with real-time performance.
Haocheng Ren, Yuchi Huo, Yifan Peng 0001, Hongtao Sheng, Weidong Xue, Hongxiang Huang, Jingzhen Lan, Rui Wang 0004, Hujun Bao
ACM Trans. Graph.2
2024 ReN Human: Learning Relightable Neural Implicit Surfaces for Animatable Human Rendering
abstract
Recently, implicit neural representation has been widely used to learn the appearance of human bodies in the canonical space, which can be further animated using a parametric human model. However, how to decompose the material properties from the implicit representation for relighting has not yet been investigated thoroughly. We propose to address this problem with a novel framework, ReN Human, that takes sparse or even monocular input videos collected in unconstrained lighting to produce a 3D human representation that can be rendered with novel views, poses, and lighting. Our method represents humans as deformable implicit neural representation and decomposes the geometry, material of humans as well as environment illumination for capturing a relightable and animatable human model. Moreover, we introduce a volumetric lighting grid consisting of spherical Gaussian mixtures to learn the spatially varying illumination and animatable visibility probes to model the dynamic self-occlusion caused by human motion. Specifically, we learn the material property fields and illumination using a physically-based rendering layer that uses Monte Carlo importance sampling to facilitate differentiation of the complex rendering integral. We demonstrate that our approach outperforms recent novel views and poses synthesis methods in a challenging benchmark with sparse videos, enabling high-fidelity human relighting.
Rengan Xie, In-Young Cho, Sen Yang 0008, Wei Chen 0001, Hujun Bao, Wenting Zheng, Yuchi Huo
ACM Trans. Graph.9
2024 Refined tri-directional path tracing with generated light portal
Xuchen Wei, Guiyang Pu, Yuchi Huo, Hujun Bao, Rui Wang 0004
Vis. Comput.3
2023 I2-SDF: Intrinsic Indoor Scene Reconstruction and Editing via Raytracing in Neural SDFs
abstract
In this work, we present I2-SDF, a new method for intrinsic indoor scene reconstruction and editing using differentiable Monte Carlo raytracing on neural signed distance fields (SDFs). Our holistic neural SDF-based frame-work jointly recovers the underlying shapes, incident radiance and materials from multi-view images. We introduce a novel bubble loss for fine-grained small objects and error-guided adaptive sampling scheme to largely improve the reconstruction quality on large-scale indoor scenes. Further, we propose to decompose the neural radiance field into spatially-varying material of the scene as a neural field through surface-based, differentiable Monte Carlo raytracing and emitter semantic segmentations, which enables physically based and photorealistic scene relighting and editing applications. Through a number of qualitative and quantitative experiments, we demonstrate the superior quality of our method on indoor scene reconstruction, novel view synthesis, and scene editing compared to state-of-the-art baselines. Our project page is at https://jingsenzhu.github.io/i2-sdf.
Jingsen Zhu, Yuchi Huo, Qi Ye 0001, Fujun Luan, Jifan Li, Dianbing Xi, Lisha Wang, Rui Tang 0015, Wei Hua 0002, Hujun Bao, Rui Wang 0004
CVPR2
2023 Towards Content-based Pixel Retrieval in Revisited Oxford and Paris
abstract
This paper introduces the first two landmark pixel retrieval benchmarks. Pixel retrieval is segmented instance retrieval. Like semantic segmentation extends classification to the pixel level, pixel retrieval is an extension of image retrieval and offers information about which pixels are related to the query object. In addition to retrieving images for the given query, it helps users quickly identify the query object in true positive images and exclude false positive images by denoting the correlated pixels. Our user study results show pixel-level annotation can significantly improve the user experience. Compared with semantic and instance segmentation, pixel retrieval requires a fine-grained recognition capability for variable-granularity targets. To this end, we propose pixel retrieval benchmarks named PROxford and PRParis, which are based on the widely used image retrieval datasets, ROxford and RParis. Three professional annotators label 5,942 images with two rounds of double-checking and refinement. Furthermore, we conduct extensive experiments and analysis on the SOTA methods in image search, image matching, detection, segmentation, and dense matching using our pixel retrieval benchmarks. Results show that the pixel retrieval task is challenging to these approaches and distinctive from existing problems, suggesting that further research can advance the content-based pixel-retrieval and thus user search experience. The datasets can be downloaded from this link.
Guoyuan An, Woo Jae Kim, Saelyne Yang, Yuchi Huo, Sung-Eui Yoon
ICCV5
2023 Seal-3D: Interactive Pixel-Level Editing for Neural Radiance Fields
abstract
With the popularity of implicit neural representations, or neural radiance fields (NeRF), there is a pressing need for editing methods to interact with the implicit 3D models for tasks like post-processing reconstructed scenes and 3D content creation. While previous works have explored NeRF editing from various perspectives, they are restricted in editing flexibility, quality, and speed, failing to offer direct editing response and instant preview. The key challenge is to conceive a locally editable neural representation that can directly reflect the editing instructions and update instantly. To bridge the gap, we propose a new interactive editing method and system for implicit representations, called Seal-3D1, which allows users to edit NeRF models in a pixel-level and free manner with a wide range of NeRF-like backbone and preview the editing effects instantly. To achieve the effects, the challenges are addressed by our proposed proxy function mapping the editing instructions to the original space of NeRF models in the teacher model and a two-stage training strategy for the student model with local pretraining and global finetuning. A NeRF editing system is built to showcase various editing types. Our system can achieve compelling editing effects with an interactive speed of about 1 second.
Jingsen Zhu, Qi Ye 0001, Yuchi Huo, Yunlong Ran, Jiming Chen 0001
ICCV4
2023 ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions
abstract
3D human reconstruction from RGB images achieves decent results in good weather conditions but degrades dramatically in rough weather. Complementary, mmWave radars have been employed to reconstruct 3D human joints and meshes in rough weather. However, combining RGB and mmWave signals for robust all-weather 3D human reconstruction is still an open challenge, given the sparse nature of mmWave and the vulnerability of RGB images. In this paper, we present ImmFusion, the first mmWave-RGB fusion solution to reconstruct 3D human bodies in all weather conditions robustly. Specifically, our ImmFusion consists of image and point backbones for token feature extraction and a Transformer module for token fusion. The image and point backbones refine global and local features from original data, and the Fusion Transformer Module aims for effective information fusion of two modalities by dynamically selecting informative tokens. Extensive experiments on a large-scale dataset, mmBody, captured in various environments demonstrate that ImmFusion can efficiently utilize the information of two modalities to achieve a robust 3D human body reconstruction in all weather conditions. In addition, our method's accuracy is significantly superior to that of state-of-the-art Transformer-based LiDAR-camera fusion methods.
Anjun Chen, Kun Shi 0003, Shaohao Zhu, Jiming Chen 0001, Yuchi Huo, Qi Ye 0001
ICRA8
2023 Contact2Grasp: 3D Grasp Synthesis via Hand-Object Contact Constraint
abstract
3D grasp synthesis generates grasping poses given an input object. Existing works tackle the problem by learning a direct mapping from objects to the distributions of grasping poses. However, because the physical contact is sensitive to small changes in pose, the high-nonlinear mapping between 3D object representation to valid poses is considerably non-smooth, leading to poor generation efficiency and restricted generality. To tackle the challenge, we introduce an intermediate variable for grasp contact areas to constrain the grasp generation; in other words, we factorize the mapping into two sequential stages by assuming that grasping poses are fully constrained given contact maps: 1) we first learn contact map distributions to generate the potential contact maps for grasps; 2) then learn a mapping from the contact maps to the grasping poses. Further, we propose a penetration-aware optimization with the generated contacts as a consistency constraint for grasp refinement. Extensive validations on two public datasets show that our method outperforms state-of-the-art methods regarding grasp generation on various metrics.
Haoming Li 0004, Xinzhuo Lin, Yuchi Huo, Jiming Chen 0001, Qi Ye 0001
IJCAI5
2023 Topological RANSAC for instance verification and retrieval without fine-tuning
abstract
This paper presents an innovative approach to enhancing explainable image retrieval, particularly in situations where a fine-tuning set is unavailable. The widely-used SPatial verification (SP) method, despite its efficacy, relies on a spatial model and the hypothesis-testing strategy for instance recognition, leading to inherent limitations, including the assumption of planar structures and neglect of topological relations among features. To address these shortcomings, we introduce a pioneering technique that replaces the spatial model with a topological one within the RANSAC process. We propose bio-inspired saccade and fovea functions to verify the topological consistency among features, effectively circumventing the issues associated with SP's spatial model. Our experimental results demonstrate that our method significantly outperforms SP, achieving state-of-the-art performance in non-fine-tuning retrieval. Furthermore, our approach can enhance performance when used in conjunction with fine-tuned features. Importantly, our method retains high explainability and is lightweight, offering a practical and adaptable solution for a variety of real-world applications.
Guoyuan An, Juhyeong Seon, Inkyu An, Yuchi Huo, Sung-Eui Yoon
NeurIPS4
2023 Adaptive Recurrent Frame Prediction with Learnable Motion Vectors
abstract
The utilization of dedicated ray tracing graphics cards has revolutionized the production of stunning visual effects in real-time rendering. However, the demand for high frame rates and high resolutions remains a challenge. The pixel warping approach is a crucial technique for increasing frame rate and resolution by exploiting the spatio-temporal coherence. To this end, existing super-resolution and frame prediction methods rely heavily on motion vectors from rendering engine pipelines to track object movements. This work builds upon state-of-the-art heuristic approaches by exploring a novel adaptive recurrent frame prediction framework that integrates learnable motion vectors. Our framework supports the prediction of transparency, particles, and texture animations, with improved motion vectors that capture shading, reflections, and occlusions, in addition to geometry movements. In addition, we introduce a feature streaming neural network, dubbed FSNet, that allows for the adaptive prediction of one or multiple sequential frames. Extensive experiments against state-of-the-art methods demonstrate that FSNet can operate at lower latency with significant visual enhancements and can upscale frame rates by at least two times. This approach offers a flexible pipeline to improve the rendering frame rates of various graphics applications and devices.
Zhizhen Wu, Chenyu Zuo, Yuchi Huo, Yazhen Yuan, Yifan Peng 0001, Guiyang Pu, Rui Wang 0004, Hujun Bao
SIGGRAPH Asia3
2023 FuseSR: Super Resolution for Real-time Rendering through Efficient Multi-resolution Fusion
abstract
The workload of real-time rendering is steeply increasing as the demand for high resolution, high refresh rates, and high realism rises, overwhelming most graphics cards. To mitigate this problem, one of the most popular solutions is to render images at a low resolution to reduce rendering overhead, and then manage to accurately upsample the low-resolution rendered image to the target resolution, a.k.a. super-resolution techniques. Most existing methods focus on exploiting information from low-resolution inputs, such as historical frames. The absence of high frequency details in those LR inputs makes them hard to recover fine details in their high-resolution predictions. In this paper, we propose an efficient and effective super-resolution method that predicts high-quality upsampled reconstructions utilizing low-cost high-resolution auxiliary G-Buffers as additional input. With LR images and HR G-buffers as input, the network requires to align and fuse features at multi resolution levels. We introduce an efficient and effective H-Net architecture to solve this problem and significantly reduce rendering overhead without noticeable quality deterioration. Experiments show that our method is able to produce temporally consistent reconstructions in 4 × 4 and even challenging 8 × 8 upsampling cases at 4K resolution with real-time performance, with substantially improved quality and significant performance boost compared to existing works.Project page: https://isaac-paradox.github.io/FuseSR/
Jingsen Zhu, Yuxin Dai, Chuankun Zheng, Yuchi Huo, Hujun Bao, Rui Wang 0004
SIGGRAPH Asia6
2023 Neural Super-Resolution in Real-Time Rendering Using Auxiliary Feature Enhancement
abstract
As the demand for high quality and high resolution in real-time rendering grows, superresolution is on its way to becoming a necessary component in modern real-time rendering applications (e.g., video games). The superresolution technique allows graphic applications to save computational costs by rendering at a lower resolution and reconstructing a high-resolution result. Nvidia introduced DLSS to the market as the first superresolution application in 2020, and NSRR was published on Siggraph the same year. Each of these approaches has shown powerful capabilities and is well suited to the needs of the industrial sector. In this paper, the authors propose the optimization potential of superresolution algorithms by introducing feature enhancement and feature caching modules and attempt to improve the current algorithms.
Rui Wang 0004, Yuchi Huo
J. Database Manag.4
2023 Contour-Aware Equipotential Learning for Semantic Segmentation
abstract
With increasing demands for high-quality semantic segmentation in the industry, hard-distinguishing semantic boundaries have posed a significant threat to existing solutions. Inspired by real-life experience, i.e., combining varied observations contributes to higher visual recognition confidence, we present the equipotential learning (EPL) method. This novel module transfers the predicted/ground-truth semantic labels to a self-defined potential domain to learn and infer decision boundaries along customized directions. The conversion to the potential domain is implemented via a lightweight differentiable anisotropic convolution without incurring any parameter overhead. Besides, the designed two loss functions, the point loss and the equipotential line loss implement anisotropic field regression and category-level contour learning, respectively, enhancing prediction consistencies in the inter/intra-class boundary areas. More importantly, EPL is agnostic to network architectures, and thus it can be plugged into most existing segmentation models. This paper is the first attempt to address the boundary segmentation problem with field regression and contour learning. Meaningful performance improvements on Pascal Voc 2012 and Cityscapes demonstrate that the proposed EPL module can benefit the off-the-shelf fully convolutional network models when recognizing semantic boundary areas. Besides, intensive comparisons and analysis show the favorable merits of EPL for distinguishing semantically-similar and irregular-shaped categories.
Xu Yin, Dongbo Min, Yuchi Huo, Sung-Eui Yoon
IEEE Trans. Multim.3
2023 Data-driven Digital Lighting Design for Residential Indoor Spaces
abstract
Conventionally, interior lighting design is technically complex yet challenging and requires professional knowledge and aesthetic disciplines of designers. This article presents a new digital lighting design framework for virtual interior scenes, which allows novice users to automatically obtain lighting layouts and interior rendering images with visually pleasing lighting effects. The proposed framework utilizes neural networks to retrieve and learn underlying design guidelines and the principles beneath the existing lighting designs, e.g., a newly constructed dataset of 6k 3D interior scenes from professional designers with dense annotations of lights. With a 3D furniture-populated indoor scene as the input, the framework takes two stages to perform lighting design: (1) lights are iteratively placed in the room; (2) the colors and intensities of the lights are optimized by an adversarial scheme, resulting in lighting designs with aesthetic lighting effects. Quantitative and qualitative experiments show that the proposed framework effectively learns the guidelines and principles and generates lighting designs that are preferred over the rule-based baseline and comparable to those of professional human designers.
Haocheng Ren, Hangming Fan, Rui Wang 0004, Yuchi Huo, Rui Tang 0015, Hujun Bao
ACM Trans. Graph.4
2023 NeLT: Object-Oriented Neural Light Transfer
abstract
This article presents object-oriented neural light transfer (NeLT), a novel neural representation of the dynamic light transportation between an object and the environment. Our method disentangles the global illumination of a scene into individual objects’ light transportation represented via neural networks, then composes them explicitly. It therefore enables flexible rendering with dynamic lighting, cameras, materials, and objects. Our rendering features various important global illumination effects, such as diffuse illumination, glossy illumination, dynamic shadowing, and indirect illumination, which completes the capability of existing neural object representation. Experiments show that NeLT does not require path tracing or shading results as input but achieves rendering quality comparable to state-of-the-art rendering frameworks, including the recent deep learning based denoisers.
Chuankun Zheng, Yuchi Huo, Shaohua Mo, Zhizhen Wu, Wei Hua 0002, Rui Wang 0004, Hujun Bao
ACM Trans. Graph.2
2023 Automatic Mesh and Shader Level of Detail
abstract
The level of detail (LOD) technique has been widely exploited as a key rendering optimization in many graphics applications. Numerous approaches have been proposed to automatically generate different kinds of LODs, such as geometric LOD or shader LOD. However, none of them have considered simplifying the geometry and shader at the same time. In this paper, we explore the observation that simplifications of geometric and shading details can be combined to provide a greater variety of tradeoffs between performance and quality. We present a new discrete multiresolution representation of objects, which consists of mesh and shader LODs. Each level of the representation could contain both simplified representations of shader and mesh. To create such LODs, we propose two automatic algorithms that pursue the best simplifications of meshes and shaders at adaptively selected distances. The results show that our mesh and shader LOD achieves better performance-quality tradeoffs than prior LOD representations, such as those that only consider simplified meshes or shaders.
Yuzhi Liang, Rui Wang 0004, Yuchi Huo, Hujun Bao
IEEE Trans. Vis. Comput. Graph.4
2022 SGW-Based Multi-task Learning in Vision Tasks
Ruiyuan Zhang, Yuyao Chen, Dianbing Xi, Yuchi Huo, Chao Wu 0001
ACCV (4)5
2022 Learning-based Inverse Rendering of Complex Indoor Scenes with Differentiable Monte Carlo Raytracing
abstract
Indoor scenes typically exhibit complex, spatially-varying appearance from global illumination, making inverse rendering a challenging ill-posed problem. This work presents an end-to-end, learning-based inverse rendering framework incorporating differentiable Monte Carlo raytracing with importance sampling. The framework takes a single image as input to jointly recover the underlying geometry, spatially-varying lighting, and photorealistic materials. Specifically, we introduce a physically-based differentiable rendering layer with screen-space ray tracing, resulting in more realistic specular reflections that match the input photo. In addition, we create a large-scale, photorealistic indoor scene dataset with significantly richer details like complex furniture and dedicated decorations. Further, we design a novel out-of-view lighting network with uncertainty-aware refinement leveraging hypernetwork-based neural radiance fields to predict lighting outside the view of the input photo. Through extensive evaluations on common benchmark datasets, we demonstrate superior inverse rendering quality of our method compared to state-of-the-art baselines, enabling various applications such as complex object insertion and material editing with high fidelity. Code and data will be made available at https://jingsenzhu.github.io/invrend
Jingsen Zhu, Fujun Luan, Yuchi Huo, Zihao Lin 0007, Dianbing Xi, Rui Wang 0004, Hujun Bao, Jiaxiang Zheng, Rui Tang 0015
SIGGRAPH Asia3
2022 MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis
abstract
Abstract With the rapid development of data‐driven techniques, data has played an essential role in various computer vision tasks. Many realistic and synthetic datasets have been proposed to address different problems. However, there are lots of unresolved challenges: (1) the creation of dataset is usually a tedious process with manual annotations, (2) most datasets are only designed for a single specific task, (3) the modification or randomization of the 3D scene is difficult, and (4) the release of commercial 3D data may encounter copyright issue. This paper presents MINERVAS, a Massive INterior EnviRonments VirtuAl Synthesis system, to facilitate the 3D scene modification and the 2D image synthesis for various vision tasks. In particular, we design a programmable pipeline with Domain‐Specific Language, allowing users to select scenes from the commercial indoor scene database, synthesize scenes for different tasks with customized rules, and render various types of imagery data, such as color images, geometric structures, semantic labels. Our system eases the difficulty of customizing massive scenes for different tasks and relieves users from manipulating fine‐grained scene configurations by providing user‐controllable randomness using multilevel samplers. Most importantly, it empowers users to access commercial scene databases with millions of indoor scenes and protects the copyright of core data assets, e.g., 3D CAD models. We demonstrate the validity and flexibility of our system by using our synthesized data to improve the performance on different kinds of computer vision tasks. The project page is at https://coohom.github.io/MINERVAS .
Haocheng Ren, Jia Zheng 0002, Jiaxiang Zheng, Rui Tang 0015, Yuchi Huo, Hujun Bao, Rui Wang 0004
Comput. Graph. Forum6
2022 PowerNet: Learning-Based Real-Time Power-Budget Rendering
abstract
With the prevalence of embedded GPUs on mobile devices, power-efficient rendering has become a widespread concern for graphics applications. Reducing the power consumption of rendering applications is critical for extending battery life. In this paper, we present a new real-time power-budget rendering system to meet this need by selecting the optimal rendering settings that maximize visual quality for each frame under a given power budget. Our method utilizes two independent neural networks trained entirely by synthesized datasets to predict power consumption and image quality under various workloads. This approach spares time-consuming precomputation or runtime periodic refitting and additional error computation. We evaluate the performance of the proposed framework on different platforms, two desktop PCs and two smartphones. Results show that compared to the previous state of the art, our system has less overhead and better flexibility. Existing rendering engines can integrate our system with negligible costs.
Yunjin Zhang, Rui Wang 0004, Yuchi Huo, Wei Hua 0002, Hujun Bao
IEEE Trans. Vis. Comput. Graph.3
2021 MeshChain: Secure 3D Model and Intellectual Property management Powered by Blockchain Technology
Hunmin Park, Yuchi Huo, Sung-Eui Yoon
CGI2
2021 Hypergraph Propagation and Community Selection for Objects Retrieval
abstract
Spatial verification is a crucial technique for particular object retrieval. It utilizes spatial information for the accurate detection of true positive images. However, existing query expansion and diffusion methods cannot efficiently propagate the spatial information in an ordinary graph with scalar edge weights, resulting in low recall or precision. To tackle these problems, we propose a novel hypergraph-based framework that efficiently propagates spatial information in query time and retrieves an object in the database accurately. Additionally, we propose using the image graph's structure information through community selection technique, to measure the accuracy of the initial search result and to provide correct starting points for hypergraph propagation without heavy spatial verification computations. Experiment results on ROxford and RParis show that our method significantly outperforms the existing query expansion and diffusion methods.
Guoyuan An, Yuchi Huo, Sung-Eui Yoon
NeurIPS2
2021 Multi-resolution terrain rendering using summed-area tables
Chuankun Zheng, Rui Wang 0004, Yuchi Huo, Wenting Zheng, Hai Lin 0003, Hujun Bao
Comput. Graph.4
2021 Real-time Monte Carlo Denoising with Weight Sharing Kernel Prediction Network
abstract
Abstract Real‐time Monte Carlo denoising aims at removing severe noise under low samples per pixel (spp) in a strict time budget. Recently, kernel‐prediction methods use a neural network to predict each pixel's filtering kernel and have shown a great potential to remove Monte Carlo noise. However, the heavy computation overhead blocks these methods from real‐time applications. This paper expands the kernel‐prediction method and proposes a novel approach to denoise very low spp (e.g., 1‐spp) Monte Carlo path traced images at real‐time frame rates. Instead of using the neural network to directly predict the kernel map, i.e., the complete weights of each per‐pixel filtering kernel, we predict an encoding of the kernel map, followed by a high‐efficiency decoder with unfolding operations for a high‐quality reconstruction of the filtering kernels. The kernel map encoding yields a compact single‐channel representation of the kernel map, which can significantly reduce the kernel‐prediction network's throughput. In addition, we adopt a scalable kernel fusion module to improve denoising quality. The proposed approach preserves kernel prediction methods’ denoising quality while roughly halving its denoising time for 1‐spp noisy inputs. In addition, compared with the recent neural bilateral grid‐based real‐time denoiser, our approach benefits from the high parallelism of kernel‐based reconstruction and produces better denoising results at equal time.
Hangming Fan, Rui Wang 0004, Yuchi Huo, Hujun Bao
Comput. Graph. Forum3
2021 A survey on deep learning-based Monte Carlo denoising
abstract
Monte Carlo (MC) integration is used ubiquitously in realistic image synthesis because of its flexibility and generality. However, the integration has to balance estimator bias and variance, which causes visually distracting noise with low sample counts. Existing solutions fall into two categories, in-process sampling schemes and post-processing reconstruction schemes. This report summarizes recent trends in the post-processing reconstruction scheme. Recent years have seen increasing attention and significant progress in denoising MC rendering with deep learning, by training neural networks to reconstruct denoised rendering results from sparse MC samples. Many of these techniques show promising results in real-world applications, and this report aims to provide an assessment of these approaches for practitioners and researchers.
Yuchi Huo, Sung-Eui Yoon
Comput. Vis. Media1
2021 Weakly-supervised contrastive learning in path manifold for Monte Carlo image reconstruction
abstract
Image-space auxiliary features such as surface normal have significantly contributed to the recent success of Monte Carlo (MC) reconstruction networks. However, path-space features, another essential piece of light propagation, have not yet been sufficiently explored. Due to the curse of dimensionality, information flow between a regression loss and high-dimensional path-space features is sparse, leading to difficult training and inefficient usage of path-space features in a typical reconstruction framework. This paper introduces a contrastive manifold learning framework to utilize path-space features effectively. The proposed framework employs weakly-supervised learning that converts reference pixel colors to dense pseudo labels for light paths. A convolutional path-embedding network then induces a low-dimensional manifold of paths by iteratively clustering intra-class embeddings, while discriminating inter-class embeddings using gradient descent. The proposed framework facilitates path-space exploration of reconstruction networks by extracting low-dimensional yet meaningful embeddings within the features. We apply our framework to the recent image- and sample-space models and demonstrate considerable improvements, especially on the sample space. The source code is available at https://github.com/Mephisto405/WCMC.
In-Young Cho, Yuchi Huo, Sung-Eui Yoon
ACM Trans. Graph.2
2020 Single Image Reflection Removal With Physically-Based Training Images
abstract
Recently, deep learning-based single image reflection separation methods have been exploited widely. To benefit the learning approach, a large number of training image pairs (i.e., with and without reflections) were synthesized in various ways, yet they are away from a physically-based direction. In this paper, physically based rendering is used for faithfully synthesizing the required training images, and a corresponding network structure and loss term are proposed. We utilize existing RGBD/RGB images to estimate meshes, then physically simulate the light transportation between meshes, glass, and lens with path tracing to synthesize training data, which successfully reproduce the spatially variant anisotropic visual effect of glass reflection. For guiding the separation better, we additionally consider a module, backtrack network (BT-net) for backtracking the reflections, which removes complicated ghosting, attenuation, blurred and defocused effect of glass/lens. This enables obtaining a priori information before having the distortion. The proposed method considering additional a priori information with physically simulated training data is validated with various real reflection images and shows visually pleasant and numerical advantages compared with state-of-the-art techniques.
Soomin Kim 0004, Yuchi Huo, Sung-Eui Yoon
CVPR2
2020 Spherical Gaussian-based Lightcuts for Glossy Interreflections
abstract
Abstract It is still challenging to render directional but non‐specular reflections in complex scenes. The SG‐based (Spherical Gaussian) many‐light framework provides a scalable solution but still requires a large number of glossy virtual lights to avoid spikes as well as reduce clamping errors. Directly gathering contributions from these glossy virtual lights to each pixel in a pairwise way is very inefficient. In this paper, we propose an adaptive algorithm with tighter error bounds to efficiently compute glossy interreflections from glossy virtual lights. This approach is an extension of the Lightcuts that builds hierarchies on both lights and pixels with new error bounds and new GPU‐based traversal methods between light and pixel hierarchies. Results demonstrate that our method is able to faithfully and efficiently compute glossy interreflections in scenes with highly glossy and spatial varying reflectance. Compared with the conventional Lightcuts method, our approach generates lightcuts with only one‐fourth to one‐fifth light nodes therefore exhibits better scalability. Additionally, after being implemented on GPU, our algorithms achieve a magnitude of faster performance than the previous method.
Yuchi Huo, Shihao Jin, Tao Liu 0016, Wei Hua 0002, Rui Wang 0004, Hujun Bao
Comput. Graph. Forum1
2020 Automatic Band-Limited Approximation of Shaders Using Mean-Variance Statistics in Clamped Domain
abstract
Abstract In this paper, we present a new shader smoothing method to improve the quality and generality of band‐limiting shader programs. Previous work [YB18] treats intermediate values in the program as random variables, and utilizes mean and variance statistics to smooth shader programs. In this work, we extend such a band‐limiting framework by exploring the observation that one intermediate value in the program is usually computed by a complex composition of functions, where the domain and range of composited functions heavily impact the statistics of smoothed programs. Accordingly, we propose three new shader smoothing rules for specific composition of functions by considering the domain and range, enabling better mean and variance statistics of approximations. Aside from continuous functions, the texture, such as color texture or normal map, is treated as a discrete function with limited domain and range, thereby can be processed similarly in the newly proposed framework. Experiments show that compared with previous work, our method is capable of generating better smoothness of shader programs as well as handling a broader set of shader programs.
Rui Wang 0004, Yuchi Huo, Wenting Zheng, Wei Hua 0002, Hujun Bao
Comput. Graph. Forum3
2020 Adaptive Incident Radiance Field Sampling and Reconstruction Using Deep Reinforcement Learning
abstract
Serious noise affects the rendering of global illumination using Monte Carlo (MC) path tracing when insufficient samples are used. The two common solutions to this problem are filtering noisy inputs to generate smooth but biased results and sampling the MC integrand with a carefully crafted probability distribution function (PDF) to produce unbiased results. Both solutions benefit from an efficient incident radiance field sampling and reconstruction algorithm. This study proposes a method for training quality and reconstruction networks (Q- and R-networks, respectively) with a massive offline dataset for the adaptive sampling and reconstruction of first-bounce incident radiance fields. The convolutional neural network (CNN)-based R-network reconstructs the incident radiance field in a 4D space, whereas the deep reinforcement learning (DRL)-based Q-network predicts and guides the adaptive sampling process. The approach is verified by comparing it with state-of-the-art unbiased path guiding methods and filtering methods. Results demonstrate improvements for unbiased path guiding and competitive performance in biased applications, including filtering and irradiance caching.
Yuchi Huo, Rui Wang 0004, Ruzahng Zheng, Hualin Xu, Hujun Bao, Sung-Eui Yoon
ACM Trans. Graph.1
2016 Adaptive matrix column sampling and completion for rendering participating media
abstract
Several scalable many-light rendering methods have been proposed recently for the efficient computation of global illumination. However, gathering contributions of virtual lights in participating media remains an inefficient and time-consuming task. In this paper, we present a novel sparse sampling and reconstruction method to accelerate the gathering step of the many-light rendering for participating media. Our technique explores the observation that the scattered lightings are usually locally coherent and of low rank even in heterogeneous media. In particular, we first introduce a matrix formation with light segments as columns and eye ray segments as rows, and formulate the gathering step into a matrix sampling and reconstruction problem. We then propose an adaptive matrix column sampling and completion algorithm to efficiently reconstruct the matrix by only sampling a small number of elements. Experimental results show that our approach greatly improves the performance, and obtains up to one order of magnitude speedup compared with other state-of-the-art methods of many-light rendering for participating media.
Yuchi Huo, Rui Wang 0004, Tianlei Hu, Wei Hua 0002, Hujun Bao
ACM Trans. Graph.1
2015 A matrix sampling-and-recovery approach for many-lights rendering
abstract
Instead of computing on a large number of virtual point lights (VPLs), scalable many-lights rendering methods effectively simulate various illumination effects only using hundreds or thousands of representative VPLs. However, gathering illuminations from these representative VPLs, especially computing the visibility, is still a tedious and time-consuming task. In this paper, we propose a new matrix sampling-and-recovery scheme to efficiently gather illuminations by only sampling a small number of visibilities between representative VPLs and surface points. Our approach is based on the observation that the lighting matrix used in manylights rendering is of low-rank, so that it is possible to sparsely sample a small number of entries, and then numerically complete the entire matrix. We propose a three-step algorithm to explore this observation. First, we design a new VPL clustering algorithm to slice the rows and group the columns of the full lighting matrix into a number of reduced matrices, which are sampled and recovered individually. Second, we propose a novel prediction method that predicts visibility of matrix entries from sparsely and randomly sampled entries. Finally, we adapt the matrix separation technique to recover the entire reduced matrix and compute final shadings. Experimental results show that our method heavily reduces the required visibility sampling in the final gathering and achieves 3--7 times speedup compared with the state-of-the-art methods on test scenes.
Yuchi Huo, Rui Wang 0004, Shihao Jin, Xinguo Liu, Hujun Bao
ACM Trans. Graph.1
2013 GPU-based out-of-core many-lights rendering
abstract
In this paper, we present a GPU-based out-of-core rendering approach under the many-lights rendering framework. Many-lights rendering is an efficient and scalable rendering framework for a large number of lights. But when the data sizes of lights and geometry are both beyond the in-core memory storage size, the data management of these two out-of-core data becomes critical and challenging. In our approach, we formulate such a data management as a graph traversal optimization problem that first builds out-of-core lights and geometry data into a graph, and then guides shading computations by finding a shortest path to visit all vertices in the graph. Based on the proposed data management, we develop a GPU-based out-of-GPU-core rendering algorithm that manages data between the CPU host memory and the GPU device memory. Two main steps are taken in the algorithm: the out-of-core data preparation to pack data into optimal data layouts for the many-lights rendering, and the out-of-core shading using graph-based data management. We demonstrate our algorithm on scenes with out-of-core detailed geometry and out-of-core lights. Results show that our approach generates complex global illumination effects with increased data access coherence and has one order of magnitude performance gain over the CPU-based approach.
Rui Wang 0004, Yuchi Huo, Yazhen Yuan, Kun Zhou 0001, Wei Hua 0002, Hujun Bao
ACM Trans. Graph.2