EDBT 2026 Demo / reviewers in the wild / expert
Jie Guo 0001
dblp:77/2751-1
· DBLP profile ↗
96ranked-venue papers
19as first author
69since 2021 · last 2026
0000-0002-4176-7617ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 19 first-author · 57 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Development of a Tongue Image-Based Machine Learning Tool for the Diagnosis of Colorectal Cancer: A Prospective Multicentre Clinical Cohort StudyabstractColorectal cancer (CRC) remains a persistent major global health burden, with traditional diagnostic methods like colonoscopy suffering from suboptimal patient compliance rates. This study develops an intelligent diagnostic model based on tongue images to assist in CRC diagnosis, leveraging the integrative potential of traditional tongue diagnosis and modern machine learning. Between June 2023 and July 2024, we collected and processed 1,389 tongue images from CRC patients and 1,543 from non-colorectal cancer (NCRC) participants. Our methodology combines innovative image segmentation using the Segment Anything Model (SAM) with Grounding DINO, extracts both hand-crafted features (color, texture, shape) and deep learning features via Swin-Transformer, and employs feature fusion and selection techniques. The diagnostic model achieves an accuracy of 87.93% (F1-score: 0.9072) in internal validation. In an independent external cohort of 119 CRC patients and 221 NCRC participants, it demonstrates 85.18% precision (recall: 85%, F1-score: 0.8507). This non-invasive, cost-effective approach demonstrates significant potential as a complementary screening tool for CRC, particularly in regions with limited access to conventional diagnostic resources. Xiaohe Sun, Letian Huang, Libo Qu, Xing Zeng, Zuojian Zhou, Xufeng Lang, Jie Guo 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2026 | Bounding Stratified Bernoulli Impulses for Ray Marching Gaussian Process Implicit SurfacesabstractThe theory of light transport on Gaussian process implicit surface (GPIS) provides a unified framework for rendering surfaces, participating media, and the intermediate spectrum. However, previous approaches rely on brute-force ray marching for surface intersections, requiring full noise evaluations at each marching point, whether using multivariate Gaussian sampling or sparse convolution noise approximation. This imposes a severe limitation on the rendering efficiency. In this paper, we derive bounds to significantly reduce the total number of full noise evaluations, leading to efficient ray marching for ray-surface intersections. We introduce stratified Bernoulli impulses, enabling a fast point-level bound for individual realizations to replace unnecessary full noise evaluations. To further reduce the number of point-level bound evaluations, we propose a region-level bound, leveraging a spatial acceleration structure to prune probabilistically empty regions, thereby avoiding unnecessary marching points in advance. By combining these two bounds, our bounded ray marching accelerates ray-surface intersections in GPIS, and consequently significantly improves overall GPIS rendering efficiency. Code for this paper are at https://github.com/Cchen-77/bounded-gpis. Zhimin Fan 0001, Lingqi Yan 0001, Junqiu Zhu, Yanwen Guo 0001, Kun Zhou 0001, Jie Guo 0001 |
ACM Trans. Graph. | 7 |
| 2026 | Efficient Fur and Hair Multiple Scattering Using Volumetric ApproximationabstractIn this paper, we propose an efficient method for hair and fur rendering that approximates multiple scattering as volumetric light transport, achieving near path-traced visual fidelity at a drastically reduced cost. Multiple scattering among hair fibers is crucial for a realistic appearance, especially for light-colored hair, but it is extremely expensive to simulate directly. Prior approximations, such as dual scattering and fur BSSRDF, improve performance but often fail for dense hair/fur, yielding an overly dry or blurred appearance. Unlike previous volumetric hair rendering solutions, we treat dense hair/fur as a highly anisotropic participating medium to capture soft volumetric illumination, while still using explicit fiber geometry for direct lighting to preserve fine details. We accumulate hair fiber coverage in screen space and stochastically sample a single-scattering event to approximate higher-order scattering. Our method reproduces the rich appearance of hair and fur resulting from multiple scattering, while running about 7–10× faster than path tracing, making it suitable for use in production. We also demonstrate that our approach is robust under a variety of lighting conditions. Ruike Hu, Junqiu Zhu, Minghao Lin, Ruian Zhang, Lu Wang 0007, Jie Guo 0001, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 6 |
| 2026 | Weakly-Supervised Shape Multi-Completion of Point Clouds by Structural DecompositionabstractThe challenge of transforming partial point clouds into complete meshes still persists, with current methods facing issues like data accessibility constraint, shape preservation failure and poor robustness on real-scan data. Drawing inspiration from the structural information of objects to enhance the completion, we introduce an innovative weakly-supervised shape completion method leveraging structural decomposition without the necessity of SDFs during training. By representing objects as abstract structural frameworks and part details, our method initiates by forecasting the structure of the input partial point clouds, and individually restore each component through part decomposition completion and generation. Extracted part details are represented in images, which are porous and incomplete. Hence, we utilize a completion network to complete such details. For multiple results generation, a diffusion-based generation network is employed to generate a variety of details for the missing areas. The predicted structure and details are subsequently converted back into meshes, yielding the complete results. Since the details are depicted in images, our approach eliminates the need for SDFs during the training phase, achieving weakly-supervision. We conduct extensive comparisons on both artificial and real-scan datasets, demonstrating an average improvement of over 38.1% compared to the prior method, and achieving SOTA performance. Changfeng Ma, Pengxiao Guo, Shuangyu Yang, Yuanqi Li, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | GSReuse: Temporally Adaptive Screen-Space Reuse for Accelerating 3D Gaussian SplattingabstractRecent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis. However, rendering each frame independently in a video sequence leads to redundant computations, especially when adjacent frames share significant visual overlaps. This inefficiency is particularly problematic in VR applications, where high frame rates and stereoscopic rendering amplify the per-frame cost. Existing frame interpolation or reuse strategies typically rely on image-domain information and are thus not directly applicable to 3DGS rendering, which is fundamentally point-based. To address this runtime inefficiency, we propose GSReuse, a lightweight and drop-in accelerator that speeds up 3DGS rendering by reusing computations across consecutive frames. GSReuse operates in screen space and introduces only minimal modifications to existing 3DGS rendering pipelines. It also eliminates the need for retraining scene representations. Given the rendered image, depth map, and camera parameters of the current frame, GSReuse estimates reliable Gaussian splatting motion vectors for all pixels and warps reusable contents to the new view. A tile-based filtering and masking strategy is then applied to determine which regions can be safely reused, allowing the 3DGS renderer to skip redundant rendering operations. We evaluate GSReuse on multiple benchmark datasets, showing that GSReuse significantly improves rendering, while maintaining high visual fidelity. Compared to state-of-the-art video frame reuse/generation methods, GSReuse delivers better image quality with much lower latency, facilitating practical deployment of 3DGS in VR applications. Chengzhi Tao, Jie Guo 0001, Letian Huang, Junqiu Zhu, Daoheng Wang, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | Surface Reconstruction From Point Clouds via Image-Free Point-to-Gaussian InferenceabstractThe task of surface reconstruction from point clouds is to produce high-quality meshes using sampled 3D points (no images available). Traditional methods primarily focus on geometric accuracy but often produce meshes without texture colors. In this paper, we present a brand-new perspective in point cloud reconstruction task-Imagining points as more informative Gaussian splats and obtaining colored surfaces through free-form Gaussian-rendering reconstruction. We train a universal Point-to-Gaussian model to infer the attributes of Gaussian splats for any given pointcloud with merely point coordinates (and color optionally) as input, without requiring any image. Significant technical designs are applied on initialization, regularization and loss functions, making the whole learning process stable. The inferred Gaussian splats can faithfully recover the original appearance of objects or scenes (capable of quick rendering from any viewpoint, like human's imagination ability), meanwhile closely adhering to the input shape. After obtaining a sufficient number of virtually rendered images and depth maps, we employ the truncated signed distance function (TSDF) fusion to get the reconstruction results, producing high-quality and colored meshes. Extensive experiments demonstrate that our approach surpasses state-of-the-art methods in surface reconstruction metrics while maintaining high efficiency and simplicity. Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | 360-GS: Layout-Guided Panoramic Gaussian Splatting for Indoor Roamingabstract3D Gaussian Splatting (3D-GS) has recently attracted great attention with real-time and photo-realistic renderings. This technique typically takes perspective images as input and optimizes a set of 3D elliptical Gaussians by splatting them onto the image planes, resulting in$2 D$Gaussians. However, applying 3D-GS to panoramic inputs presents challenges in effectively modeling the projection onto the spherical surface of 360° images using 2D Gaussians. In practical applications, input panoramas are often sparse, leading to unreliable initialization of 3D Gaussians and subsequent degradation of 3D-GS quality. In addition, due to the under-constrained geometry of texture-less planes (e.g., walls and floors), 3D-GS struggles to model these flat regions with elliptical Gaussians, resulting in significant floaters in novel views. To address these issues, we propose 360-GS, a novel layout-guided 360° Gaussian splatting for a limited set of panoramic inputs. Instead of splatting 3D Gaussians directly onto the spherical surface, 360-GS projects them onto the tangent plane of the unit sphere and then maps them to the spherical projections. Jiayang Bai, Letian Huang, Jie Guo 0001, Wen Gong, Yuanqi Li, Yanwen Guo 0001 |
3DV | 3 |
| 2025 | Real-Time Neural Denoising with Render-Aware Knowledge DistillationabstractReal-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we present a render-aware knowledge distillation (RAKD) framework, specifically designed for Monte Carlo denoising. We meticulously delineate the Knowledge Distillation (KD) process within RAKD, emphasizing three pivotal techniques: the strategic incorporation of an auxiliary unlabeled dataset, the integration of adversarial learning through generative adversarial network (GAN), and the application of parameter transfer for robust model initialization. These approaches are harmoniously combined to distill knowledge effectively, enabling our student model to adeptly strike a balance between preserving high-frequency details and reducing low-frequency noise. Finally, our results demonstrate that RAKD achieves state-of-the-art quality while upholding real-time performance, successfully tackling the computational constraints faced by resource-limited devices. Mengxun Kong, Jie Guo 0001, Chen Wang 0149, Yanwen Guo 0001 |
AAAI | 2 |
| 2025 | High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance EncodingabstractThe proliferation of Light Detection and Ranging (LiDAR) technology has facilitated the acquisition of three-dimensional point clouds, which are integral to applications in VR, AR, and Digital Twin. Oriented normals, critical for 3D reconstruction and scene analysis, cannot be directly extracted from scenes using LiDAR due to its operational principles. Previous traditional or learning-based methods are prone to inaccuracies due to uneven distribution and noise due to the dependence on local geometry features. This paper addresses the challenge of estimating oriented point normals by introducing a point cloud normal estimation framework via hybrid angular and Euclidean distance encoding (HAE). Our method overcomes the limitations of local geometric information by combining angular and Euclidean spaces to extract features from both point cloud coordinates and light rays, leading to more accurate normal estimation. The core of our network consists of an angular distance encoding module, which leverages both ray directions and point coordinates for unoriented normal refinement, and a ray feature fusion module for normal orientation, that is robust to noise. We also provide a point cloud dataset with ground truth normals, generated a virtual scanner, which reflects real scanning distributions and noise profiles. Yuanqi Li, Jingcheng Huang, Hongshen Wang, Peiyuan Lv, Jiuming Zheng, Jie Guo 0001, Yanwen Guo 0001 |
CVPR | 7 |
| 2025 | Sparse Point Cloud Patches Rendering via Splitting 2D GaussiansabstractCurrent learning-based methods predict NeRF or 3D Gaussians from point clouds to achieve photo-realistic rendering but still depend on categorical priors, dense point clouds, or additional refinements. Hence, we introduce a novel point cloud rendering method by predicting 2D Gaussians from point clouds. Our method incorporates two identical modules with an entire-patch architecture enabling the network to be generalized to multiple datasets. The module normalizes and initializes the Gaussians utilizing the point cloud information including normals, colors and distances. Then, splitting decoders are employed to refine the initial Gaussians by duplicating them and predicting more accurate results, making our methodology effectively accommodate sparse point clouds as well. Once trained, our approach exhibits direct generalization to point clouds across different categories. The predicted Gaussians are employed directly for rendering without additional refinement on the rendered images, retaining the benefits of 2D Gaussians. We conduct extensive experiments on various datasets, and the results demonstrate the superiority and generalization of our method, which achieves SOTA performance. The code is available at https://github.com/murcherful/GauPCRender. Changfeng Ma, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001 |
CVPR | 3 |
| 2025 | SGCR: Spherical Gaussians for Efficient 3D Curve ReconstructionabstractNeural rendering techniques have made substantial progress in generating photo-realistic 3D scenes. The latest 3D Gaussian Splatting technique has achieved high quality novel view synthesis as well as fast rendering speed. However, 3D Gaussians lack proficiency in defining accurate 3D geometric structures despite their explicit primitive representations. This is due to the fact that Gaussian’s attributes are primarily tailored and fine-tuned for rendering diverse 2D images by their anisotropic nature. To pave the way for efficient 3D reconstruction, we present Spherical Gaussians, a simple and effective representation for 3D geometric boundaries, from which we can directly reconstruct 3D feature curves from a set of calibrated multi-view images. Spherical Gaussians is optimized from grid initialization with a view-based rendering loss, where a 2D edge map is rendered at a specific view and then compared to the ground-truth edge map extracted from the corresponding image, without the need for any 3D guidance or supervision. Given Spherical Gaussians serve as intermedia for the robust edge representation, we further introduce a novel optimization-based algorithm called SGCR to directly extract accurate parametric curves from aligned Spherical Gaussians. We demonstrate that SGCR outperforms existing state-of-the-art methods in 3D edge reconstruction while enjoying great efficiency. Code is available at https://github.com/Martinyxr/SGCR. Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Junyuan Xie |
CVPR | 4 |
| 2025 | EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry FeaturesabstractPoint cloud reconstruction is a critical process in 3D representation and reverse engineering. When it comes to CAD models, edges are significant features that play a crucial role in characterizing the geometry of 3D shapes. However, few points are exactly sampled on edges during acquisition, resulting in apparent artifacts for the reconstruction task. Upsampling point cloud is a direct technical route, but there is a main challenge that the upsampled points may not align with the model edge accurately. To overcome this, we develop an integrated framework to estimate edges by joint regression of three geometry features—point-to-edge direction, point-to-edge distance and point normal. Benefiting these features, we implement a novel refinement process to move and produce more points which lie accurately on edges of the model, allowing for high-quality edge-preserving reconstruction. Experiments and comparisons against previous methods demonstrate our method’s effectiveness and superiority. Xinran Yang, Donghao Ji, Yuanqi Li, Junyuan Xie, Jie Guo 0001, Yanwen Guo 0001 |
CVPR | 5 |
| 2025 | GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections
Haiyang Bai, Songru Jiang, Tao Lu 0005, Yuanqi Li, Jie Guo 0001, Runze Fu, Yanwen Guo 0001, Lijun Chen 0006 |
ICCV | 7 |
| 2025 | Decoupled Motion Prediction for Real-time G-buffer Free Frame ExtrapolationabstractFrame extrapolation, as a typical low-latency frame generation method, improves the frame rate of real-time rendering by predicting future frames based solely on historical data. To guarantee the high quality of predictions, existing methods rely heavily on G-buffers of the target frames. However, these G-buffers are not always accessible, and enabling them in certain rendering engines can incur considerable costs. To tackle this challenge, we introduce a G-buffer free frame extrapolation framework that can achieve comparable quality with state-of-the-art G-buffer based methods. In contrast to existing learning-based approaches that handle motions of new frames implicitly and jointly, we design a decoupled strategy that predicts explicit motions for geometry, shading and disoccluded regions separately. In our framework, we first extract the geometric motion using a dual-space method, and then leverage a lightweight motion inpainting network (OccNet) to fill in the disoccluded regions. The shading motion is extracted between two historical frames and then used to propagate shading variations to new frames. Through extensive experiments across various scenes, we demonstrate that our decoupled approach can generate high-quality motions for a wide range of geometric and shading variations in a scene, thereby significantly improving the accuracy of extrapolated frames at a very low computational expense. Liang Pu, Zesen Feng, Jie Guo 0001 |
ACM Multimedia | 6 |
| 2025 | Actial: Activate Spatial Reasoning Ability of Multimodal Large Language ModelsabstractRecent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding. Xiaoyu Zhan, Wenxuan Huang 0001, Xinyu Fu 0009, Changfeng Ma, Shaosheng Cao, Bohan Jia, Shaohui Lin, Zhenfei Yin, Lei Bai 0001, Wanli Ouyang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001 |
NeurIPS | 13 |
| 2025 | Spectral-GS: Taming 3D Gaussian Splatting with Spectral EntropyabstractRecently, 3D Gaussian Splatting (3DGS) has achieved impressive results in novel view synthesis, demonstrating high fidelity and efficiency. However, it easily exhibits needle-like artifacts, especially when increasing the sampling rate. Mip-Splatting tries to remove these artifacts with a 3D smoothing filter for frequency constraints and a 2D Mip filter for approximated supersampling. Unfortunately, it tends to produce over-blurred results, and sometimes needle-like Gaussians still persist. Our spectral analysis of the covariance matrix during optimization and densification reveals that current 3DGS lacks shape awareness, relying instead on spectral radius and view positional gradients to determine splitting. As a result, needle-like Gaussians with small positional gradients and low spectral entropy fail to split and overfit high-frequency details. Furthermore, both the filters used in 3DGS and Mip-Splatting reduce the spectral entropy and increase the condition number during zooming in to synthesize novel view, causing view inconsistencies and more pronounced artifacts. Our Spectral-GS, based on spectral analysis, introduces 3D shape-aware splitting and 2D view-consistent filtering strategies, effectively addressing these issues, enhancing 3DGS’s capability to represent high-frequency details without noticeable artifacts, and achieving high-quality realistic rendering. Letian Huang, Jie Guo 0001, Jialin Dan, Ruoyu Fu, Yuanqi Li, Yanwen Guo 0001 |
SIGGRAPH Asia | 2 |
| 2025 | ANIR: Adaptive Neural Implicit Representation for 3D shape reconstruction and generation
Kun Liu 0021, Yan Zhang 0057, Yanwen Guo 0001, Jie Guo 0001 |
Comput. Aided Des. | 4 |
| 2025 | Detail-Preserving Real-Time Hair Strand Linking and FilteringabstractAbstract Realistic hair rendering remains a significant challenge in computer graphics due to the intricate microstructure of hair fibers and their anisotropic scattering properties, which make them highly sensitive to noise. Although recent advancements in image‐space and 3D‐space denoising and antialiasing techniques have facilitated real‐time rendering in simple scenes, existing methods still struggle with excessive blurring and artifacts, particularly in fine hair details such as flyaway strands. These issues arise because current techniques often fail to preserve sub‐pixel continuity and lack directional sensitivity in the filtering process. To address these limitations, we introduce a novel real‐time hair filtering technique that effectively reconstructs fine fiber details while suppressing noise. Our method improves visual quality by maintaining strand‐level details and ensuring computational efficiency, making it well‐suited for real‐time applications in video games and virtual reality (VR) and augmented reality (AR) environments. Tao Huang 0026, J. Yuan, Ruike Hu, Lu Wang 0007, Yanwen Guo 0001, Bin Chen 0019, Jie Guo 0001, Junqiu Zhu |
Comput. Graph. Forum | 7 |
| 2025 | Realistic Simulation of Underwater Scene for Image EnhancementabstractIn recent years, learning-based methods have performed remarkably well in underwater image enhancement, but their performance is limited by the lack of high-quality, diverse training datasets. Current underwater image datasets are unable to address the following three issues: intra-domain gaps in underwater environments, inter-domain gaps between synthetic and real data, and domain inaccuracies. To overcome these limitations, we construct a realistic underwater scene using 3D graphics engine through a three-step approach: 1) integrate a simulation-specific underwater light propagation models to create volumetric fog; 2) employ physical model-based rendering for accurate light field simulation; 3) configure scenes with parameters extracted from real underwater images. Based on this framework, we develop an underwater image enhancement dataset (MUSE). Experiments demonstrate that models trained on MUSE outperform those trained on conventional datasets, highlighting the effectiveness of our approach. Tingyu Liu, Qunyan Jiang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Zhonghua Ni |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Bernstein Bounds for CausticsabstractSystematically simulating specular light transport requires an exhaustive search for triangle tuples containing admissible paths. Given the extreme inefficiency of enumerating all combinations, we significantly reduce the search domain by stochastically sampling such tuples. The challenge is to design proper sampling probabilities that keep the noise level controllable. Our key insight is that by bounding the irradiance contributed by each triangle tuple at a given position, we can sample a subset of triangle tuples with potentially high contributions. Although low-contribution tuples are assigned a negligible probability, the overall variance remains low. Therefore, we derive position and irradiance bounds for caustics casted by each triangle tuple, introducing a bounding property of rational functions on a Bernstein basis. When formulating position and irradiance expressions into rational functions, we handle non-rational parts through remainder variables to maintain bounding validity. Finally, we carefully design the sampling probabilities by optimizing the upper bound of the variance, expressed only using the position and irradiance bounds. The bound-driven sampling of triangle tuples is intrinsically unbiased even without defensive sampling. It can be combined with various unbiased and biased root-finding techniques within a local triangle domain. Extensive evaluations show that our method enables the fast and reliable rendering of complex caustics effects. Yet, our method is efficient for no more than two specular vertices, where complexity grows sublinearly to the number of triangles and linearly to that of emitters, and does not consider the Fresnel and visibility terms. We also rely on parameters to control subdivisions. Zhimin Fan 0001, Chen Wang 0149, Boxuan Li, Lingqi Yan 0001, Yanwen Guo 0001, Jie Guo 0001 |
ACM Trans. Graph. | 8 |
| 2025 | Multiple Importance Reweighting for Path GuidingabstractContemporary path guiding employs an iterative training scheme to fit radiance distributions. However, existing methods combine the estimates generated in each iteration merely within image space, overlooking differences in the convergence of distribution fitting over individual light paths. This paper formulates the estimation combination task as a path reweighting process. To compute spatio-directional varying combination weights, we propose multiple importance reweighting , leveraging the importance distributions from multiple guiding iterations. We demonstrate that our proposed path-level reweighting makes guiding algorithms less sensitive to noise and overfitting in distributions. This facilitates a finer subdivision of samples both spatially and temporally (i.e., over iterations), which leads to additional improvements in the accuracy of distributions and samples. Inspired by adaptive multiple importance sampling (AMIS), we introduce a simple yet effective mixture-based weighting scheme with theoretically guaranteed consistency, demonstrating good practical performance compared to alternative weighting schemes. To further foster usage with high sample rates, we introduce a hyperparameter that controls the size of sample storage. When this size limit is exceeded, low-valued samples are splatted during rendering and reweighted using a partial mixture of distributions. We found limiting the storage size reduces memory overhead and keeps variance reduction and bias comparable to the unlimited ones. Our method is largely agnostic to the underlying guiding method and compatible with conventional pixel reweighting techniques. Extensive evaluations underscore the feasibility of our approach in various scenes, achieving variance reduction with negligible bias over state-of-the-art solutions within equal sample rates and rendering time. Zhimin Fan 0001, Lingqi Yan 0001, Yanwen Guo 0001, Jie Guo 0001 |
ACM Trans. Graph. | 6 |
| 2025 | TransparentGS: Fast Inverse Rendering of Transparent Objects with GaussiansabstractThe emergence of neural and Gaussian-based radiance field methods has led to considerable advancements in novel view synthesis and 3D object reconstruction. Nonetheless, specular reflection and refraction continue to pose significant challenges due to the instability and incorrect overfitting of radiance fields to high-frequency light variations. Currently, even 3D Gaussian Splatting (3D-GS), as a powerful and efficient tool, falls short in recovering transparent objects with nearby contents due to the existence of apparent secondary ray effects. To address this issue, we propose TransparentGS, a fast inverse rendering pipeline for transparent objects based on 3D-GS. The main contributions are three-fold. Firstly, an efficient representation of transparent objects, transparent Gaussian primitives, is designed to enable specular refraction through a deferred refraction strategy. Secondly, we leverage Gaussian light field probes (GaussProbe) to encode both ambient light and nearby contents in a unified framework. Thirdly, a depth-based iterative probes query (IterQuery) algorithm is proposed to reduce the parallax errors in our probe-based framework. Experiments demonstrate the speed and accuracy of our approach in recovering transparent objects from complex environments, as well as several applications in computer graphics and vision. Letian Huang, Dongwei Ye, Jialin Dan, Chengzhi Tao, Kun Zhou 0001, Bo Ren 0003, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001 |
ACM Trans. Graph. | 10 |
| 2025 | DSCombiner: Double Shrinkage for Combining Biased and Unbiased Monte Carlo RenderingsabstractMonte Carlo rendering often faces a dilemma, namely, whether to choose an unbiased estimator or a biased one. Although different integrators have been developed to address various scenarios, no single method can effectively manage all situations. Thus, finding a good approach to combine different integrators has always been a topic that warrants exploration. This work proposes DSCombiner, a new shrinkage estimator that flexibly combines unbiased and biased estimators (typically generated by different integrators) in image space into a single estimating procedure, strategically utilizing the strengths of different integrators while minimizing their weaknesses. DSCombiner overcomes the limitation of single shrinkage combiners by introducing a two-step shrinkage towards a noise-free radiance prior. We derive optimal shrinkage factors for the two steps within a hierarchical Bayesian framework, and provide a deep learning-based method to improve the results. Comprehensive qualitative and quantitative validations across diverse scenes demonstrate visible improvements in image quality, as compared with previous image-space and path-space combiners. Keheng Xu, Mufan Guo, Xianhao Yu, Zhimin Fan 0001, Guihuan Feng, Yanwen Guo 0001, Jie Guo 0001 |
ACM Trans. Graph. | 8 |
| 2025 | Taming High-Resolution Auxiliary G-Buffers for Deep Supersampling of Rendered ContentabstractHigh-resolution images come with rich color information and texture details. Due to the rapid upgrading of display devices and rendering technologies, high-resolution real-time rendering faces the computational overhead challenge. To address this, the current mainstream solution is to render at a lower resolution and then upsample to the target resolution by supersampling techniques. However, while many prior supersampling approaches have attempted to exploit rich rendered data such as color, depth, motion vectors at low resolution, there is little discussion on how to harness high-frequency information that is readily available in the high-resolution (HR) G-buffers of modern renders. In this article, we seek to investigate how to fully leverage information from HR G-buffers to maximize the visual quality of supersampling results. We propose a neural network for real-time supersampling of rendered content, which is based on several core designs, including gated G-buffers encoder, G-buffers attended encoder and reflection-aware loss. These designs are especially made for the sake of effectively using HR G-buffers, enabling faithful recovery of a variety of high-frequency scene details from low-resolution, highly aliased inputs. Furthermore, a simple occlusion-aware blender is proposed to efficiently rectify dis-occluded features in the warped previous frame, allowing us to better exploit history information to improve temporal stability. The experiments show that our method, equipped with strong ability to harness HR G-buffer information, significantly improves the visual fidelity of high-resolution reconstructions upon previous state-of-the-art methods, even for challenging $4 \times 4$4×4 upsampling, while still being compute-efficient. Pengjie Wang 0001, Chengzhi Yuan, Jie Guo 0001, Xiaosong Yang, Houjie Li, Ian Stephenson, Jian Chang 0001, Ying Cao 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | MixRF: Universal Mixed Radiance Fields With Points and Rays AggregationabstractRecent advancements in neural rendering methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D-GS), have significantly revolutionized photo-realistic novel view synthesis of scenes with multiple photos or videos as input. However, existing approaches within the NeRF and 3D-GS frameworks often assume the independence of point sampling and ray casting, which are intrinsic to volume rendering and alpha-blending techniques. These underlying assumptions limit the ability to aggregate context within subspaces, such as densities and colors in the radiance fields and pixels on the image plane, leading to synthesized images that lack fine details and smoothness. To overcome this, we propose a universal framework, MixRF, comprising a Radiance Field Mixer (RF-mixer) and a Color Domain Mixer (CD-mixer), to sufficiently aggregate and fully explore information in neighboring sampled points and casting rays, separately. The RF-mixer treats sampled points as an explicit point cloud, enabling the aggregation of density and color attributes from neighboring points to better capture local geometry and appearance. Meanwhile, the CD-mixer rearranges rendered pixels on the sub-image plane, improving smoothness and recovering fine details and textures. Both mixers employ a kernel-based mixing strategy to facilitate effective and controllable attribute aggregation, ensuring a more comprehensive exploration of radiance values and pixel information. Extensive experiments demonstrate that our MixRF framework is compatible with radiance field-based methods, including NeRF and 3D-GS designs. The proposed framework dramatically enhances performance in both qualitative and quantitative evaluations, with less than a $ 25\%$25% increase in computational overhead during inference. Haiyang Bai, Tao Lu 0005, Chang Gou, Jie Guo 0001, Lijun Chen 0006, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | GlossyGS: Inverse Rendering of Glossy Objects With 3D Gaussian SplattingabstractReconstructing objects from posed images is a crucial and complex task in computer graphics and computer vision. While NeRF-based neural reconstruction methods have exhibited impressive reconstruction ability, they tend to be time-comsuming. Recent strategies have adopted 3D Gaussian Splatting (3D-GS) for inverse rendering, which have led to quick and effective outcomes. However, these techniques generally have difficulty in producing believable geometries and materials for glossy objects, a challenge that stems from the inherent ambiguities of inverse rendering. To address this, we introduce GlossyGS, an innovative 3D-GS-based inverse rendering framework that aims to precisely reconstruct the geometry and materials of glossy objects by integrating material priors. The key idea is the use of micro-facet geometry segmentation prior, which helps to reduce the intrinsic ambiguities and improve the decomposition of geometries and materials. Additionally, we introduce a normal map prefiltering strategy to more accurately simulate the normal distribution of reflective surfaces. These strategies are integrated into a hybrid geometry and material representation that employs both explicit and implicit methods to depict glossy objects. We demonstrate through quantitative analysis and qualitative visualization that the proposed method is effective to reconstruct high-fidelity geometries and materials of glossy objects, and performs favorably against State-of-the-Arts. Shuichang Lai, Letian Huang, Jie Guo 0001, Bowen Pan, Xiaoxiao Long, Jiangjing Lyu, Chengfei Lv, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Deep Point Cloud Edge Reconstruction via Surface Patch SegmentationabstractParametric edge reconstruction for point cloud data is a fundamental problem in computer graphics. Existing methods first classify points as either edge points (including corners) or non-edge points, and then fit parametric edges to the edge points. However, few points are exactly sampled on edges in practical scenarios, leading to significant fitting errors in the reconstructed edges. Prominent deep learning-based methods also primarily emphasize edge points, overlooking the potential of non-edge areas. Given that sparse and non-uniform edge points cannot provide adequate information, we address this challenge by leveraging neighboring segmented patches to supply additional cues. We introduce a novel two-stage framework that reconstructs edges precisely and completely via surface patch segmentation. First, we propose PCER-Net, a Point Cloud Edge Reconstruction Network that segments surface patches, detects edge points, and predicts normals simultaneously. Second, a joint optimization module is designed to reconstruct a complete and precise 3D wireframe by fully utilizing the predicted results of the network. Concretely, the segmented patches enable accurate fitting of parametric edges, even when sparse points are not precisely distributed along the model's edges. Corners can also be naturally detected from the segmented patches. Benefiting from fitted edges and detected corners, a complete and precise 3D wireframe model with topology connections can be reconstructed by geometric optimization. Finally, we present a versatile patch-edge dataset, including CAD and everyday models (furniture), to generalize our method. Extensive experiments and comparisons against previous methods demonstrate our effectiveness and superiority. We will release the code and dataset to facilitate future research. Yuanqi Li, Hongshen Wang, Jingcheng Huang, Jianwei Guo 0003, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | Parameterize Structure With Differentiable Template for 3D Shape GenerationabstractStructural representation is crucial for reconstructing and generating editable 3D shapes with part semantics. Recent 3D shape generation works employ complicated networks and structure definitions relying on hierarchical annotations and pay less attention to the details inside parts. In this paper, we propose the method that parameterizes the shared structure in the same category using a differentiable template and corresponding fixed-length parameters. Specific parameters are fed into the template to calculate cuboids that indicate a concrete shape. We utilize the boundaries of three-view renderings of each cuboid to further describe the inside details. Shapes are represented with the parameters and three-view details inside cuboids, from which the SDF can be calculated to recover the object. Benefiting from our fixed-length parameters and three-view details, our networks for reconstruction and generation are simple and effective to learn the latent space. Our method can reconstruct or generate diverse shapes with complicated details, and interpolate them smoothly. Extensive evaluations demonstrate the superiority of our method on reconstruction from point cloud, generation, and interpolation. Changfeng Ma, Pengxiao Guo, Shuangyu Yang, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Make PBR Materials Tileable With Latent Diffusion InpaintingabstractPhysically-based-rendering (PBR) materials are crucial in modern rendering pipelines, and many studies have focused on acquiring these materials from reality or images. However, existing methods may result in non-tileable results, since the realistic inputs usually have seams. Compared to non-tileable materials, tileable PBR materials have more universal application scenarios. To address this issue, we introduce MaTi, a novel pipeline that converts non-tileable PBR materials into tileable ones with minimal distortion. MaTi rearranges material patches to align boundaries at the center of the image, and then uses a diffusion model to inpaint the seams. We use scaled gamma correction to reduce the occurrence of collapse when processing special material maps. The color correction and triangular blending are adopt to preserve the original material information. Additionally, we design a division and blending strategy to efficiently handle high resolution materials. Our experiments demonstrate that MaTi can seamlessly modify PBR materials while preserving the original information, outperforming existing synthesis methods. Xiaoyu Zhan, Jianxin Yang, Jun Wang 0039, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | FASSET: Frame Supersampling and Extrapolation Using Implicit Neural Representations of Rendering Contents
Haoyu Qin, Jie Guo 0001, Wenyang Bai, Yanwen Guo 0001 |
CVM (1) | 3 |
| 2024 | Practical Measurements of Translucent Materials with Inter-Pixel Translucency PriorabstractMaterial appearance is a key component of photorealism, with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g., consumer-level cameras), little attention is paid on acquiring the optical properties of translucent materials which are also quite common in nature. In this paper, we present a practical method for acquiring scattering properties of translucent materials, based solely on ordinary images captured with unknown lighting and camera parameters. The key to our method is an inter-pixel translucency prior which states that image pixels of a given homogeneous translucent material typically form curves (dubbed translucent curves) in the RGB space, of which the shapes are determined by the parameters of the material. We leverage this prior in a specially-designed convolutional neural network comprising multiple encoders, a translucency-aware feature fusion module and a cascaded decoder. We demonstrate, through both visual comparisons and quantitative evaluations, that high accuracy can be achieved on a wide range of real-world translucent materials. Zhenyu Chen 0001, Jie Guo 0001, Shuichang Lai, Ruoyu Fu, Mengxun Kong, Chen Wang 0149, Hongyu Sun 0001, Zhebin Zhang, Chen Li 0062, Yanwen Guo 0001 |
CVPR | 2 |
| 2024 | LiDAR-Net: A Real-Scanned 3D Point Cloud Dataset for Indoor ScenesabstractIn this paper, we present LiDAR-Net, a new real-scanned indoor point cloud dataset, containing nearly 3.6 billion precisely point-level annotated points, covering an expansive area of 30,000m2. It encompasses three prevalent daily environments, including learning scenes, working scenes, and living scenes. LiDAR-Net is characterized by its non-uniform point distribution, e.g., scanning holes and scanning lines. Additionally, it meticulously records and an-notates scanning anomalies, including reflection noise and ghost. These anomalies stem from specular reflections on glass or metal, as well as distortions due to moving persons. LiDAR-Net's realistic representation of non-uniform distribution and anomalies significantly enhances the training of deep learning models, leading to improved generalization in practical applications. We thoroughly evaluate the performance of state-of-the-art algorithms on LiDAR-Net and provide a detailed analysis of the results. Crucially, our research identifies several fundamental challenges in understanding indoor point clouds, contributing essential insights to future explorations in this field. Our dataset can be found online: http://lidar-net.njumeta.com. Yanwen Guo 0001, Yuanqi Li, Dayong Ren, Xiaohong Zhang 0009, Liang Pu, Changfeng Ma, Xiaoyu Zhan, Jie Guo 0001, Mingqiang Wei, Yan Zhang 0057, Piaopiao Yu, Shuangyu Yang, Donghao Ji, Huisheng Ye |
CVPR | 9 |
| 2024 | Semantic Human Mesh Reconstruction with TexturesabstractThe field of 3D detailed human mesh reconstruction has made significant progress in recent years. However, current methods still face challenges when used in industrial applications due to unstable results, low-quality meshes, and a lack of UV unwrapping and skinning weights. In this paper, we present SHERT, a novel pipeline that can reconstruct semantic human meshes with textures and high-precision details. SHERT applies semantic- and normal-based sampling between the detailed surface (e.g. mesh and SDF) and the corresponding SMPL-X model to obtain a partially sampled semantic mesh and then generates the complete semantic mesh by our specifically designed self-supervised completion and refinement networks. Using the complete semantic mesh as a basis, we employ a texture diffusion model to create human textures that are driven by both images and texts. Our reconstructed meshes have stable UV unwrapping, high-quality triangle meshes, and consistent semantic information. The given SMPL-X model provides semantic information and shape priors, allowing SHERT to perform well even with incorrect and incomplete inputs. The semantic information also makes it easy to substitute and animate different body parts such as the face, body, and hands. Quantitative and qualitative experiments demonstrate that SHERT is capable of producing high-fidelity and robust semantic meshes that outperform state-of-the-art methods. Xiaoyu Zhan, Jianxin Yang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Wenping Wang 0001 |
CVPR | 4 |
| 2024 | Prompt3D: Random Prompt Assisted Weakly-Supervised 3D Object DetectionabstractThe prohibitive cost of annotations for fully supervised 3D indoor object detection limits its practicality. In this work, we propose Random Prompt Assisted Weakly-supervised 3D Object Detection, termed as Prompt3D, a weakly-supervised approach that leverages position-level labels to overcome this challenge. Explicitly, our method focuses on enhancing labeling using synthetic scenes crafted from 3D shapes generated via random prompts. First, a Synthetic Scene Generation (SSG) module is introduced to assemble synthetic scenes with a curated collection of 3D shapes, created via random prompts for each category. These scenes are enriched with automatically generated point-level annotations, providing a robust supervisory frame-work for training the detection algorithm. To enhance the transfer of knowledge from virtual to real datasets, we then introduce a Prototypical Proposal Feature Alignment (PPFA) module. This module effectively alleviates the domain gap by directly minimizing the distance between feature prototypes of the same class proposals across two domains. Compared with sota BR, our method improves by 5.4% and 8.7% on mAP with VoteNet and GroupFree3D serving as detectors respectively, demonstrating the effectiveness of our proposed method. Code is available at: https://github.com/huishengye/prompt3d. Xiaohong Zhang 0009, Huisheng Ye, Qinyu Tang, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001 |
CVPR | 7 |
| 2024 | On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
Letian Huang, Jiayang Bai, Jie Guo 0001, Yuanqi Li, Yanwen Guo 0001 |
ECCV (17) | 3 |
| 2024 | GLPanoDepth: Global-to-Local Panoramic Depth EstimationabstractDepth estimation is a fundamental task in many vision applications. With the popularity of omnidirectional cameras, it becomes a new trend to tackle this problem in the spherical space. In this paper, we propose a learning-based method for predicting dense depth values of a scene from a monocular omnidirectional image. An omnidirectional image has a full field-of-view, providing much more complete descriptions of the scene than perspective images. However, fully-convolutional networks that most current solutions rely on fail to capture rich global contexts from the panorama. To address this issue and also the distortion of equirectangular projection in the panorama, we propose Cubemap Vision Transformers (CViT), a new transformer-based architecture that can model long-range dependencies and extract distortion-free global features from the panorama. We show that cubemap vision transformers have a global receptive field at every stage and can provide globally coherent predictions for spherical signals. As a general architecture, it removes any restriction that has been imposed on the panorama in many other monocular panoramic depth estimation methods. To preserve important local features, we further design a convolution-based branch in our pipeline (dubbed GLPanoDepth) and fuse global features from cubemap vision transformers at multiple scales. This global-to-local strategy allows us to fully exploit useful global and local features in the panorama, achieving state-of-the-art performance in panoramic depth estimation. Jiayang Bai, Haoyu Qin, Shuichang Lai, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Specular PolynomialsabstractFinding valid light paths that involve specular vertices in Monte Carlo rendering requires solving many non-linear, transcendental equations in high-dimensional space. Existing approaches heavily rely on Newton iterations in path space, which are limited to obtaining at most a single solution each time and easily diverge when initialized with improper seeds. We propose specular polynomials , a Newton iteration-free methodology for finding a complete set of admissible specular paths connecting two arbitrary endpoints in a scene. The core is a reformulation of specular constraints into polynomial systems, which makes it possible to reduce the task to a univariate root-finding problem. We first derive bivariate systems utilizing rational coordinate mapping between the coordinates of consecutive vertices. Subsequently, we adopt the hidden variable resultant method for variable elimination, converting the problem into finding zeros of the determinant of univariate matrix polynomials. This can be effectively solved through Laplacian expansion for one bounce and a bisection solver for more bounces. Our solution is generic, completely deterministic, accurate for the case of one bounce, and GPU-friendly. We develop efficient CPU and GPU implementations and apply them to challenging glints and caustic rendering. Experiments on various scenarios demonstrate the superiority of specular polynomial-based solutions compared to Newton iteration-based counterparts. Our implementation is available at https://github.com/mollnn/spoly. Zhimin Fan 0001, Jie Guo 0001, Zhenyu Chen 0001, Pengpei Hong, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 2 |
| 2024 | Conditional Mixture Path Guiding for Differentiable RenderingabstractThe efficiency of inverse optimization in physically based differentiable rendering heavily depends on the variance of Monte Carlo estimation. Despite recent advancements emphasizing the necessity of tailored differential sampling strategies, the general approaches remain unexplored. In this paper, we investigate the interplay between local sampling decisions and the estimation of light path derivatives. Considering that modern differentiable rendering algorithms share the same path for estimating differential radiance and ordinary radiance, we demonstrate that conventional guiding approaches, conditioned solely on the last vertex, cannot attain this density. Instead, a mixture of different sampling distributions is required, where the weights are conditioned on all the previously sampled vertices in the path. To embody our theory, we implement a conditional mixture path guiding that explicitly computes optimal weights on the fly. Furthermore, we show how to perform positivization to eliminate sign variance and extend to scenes with millions of parameters. To the best of our knowledge, this is the first generic framework for applying path guiding to differentiable rendering. Extensive experiments demonstrate that our method achieves nearly one order of magnitude improvements over state-of-the-art methods in terms of variance reduction in gradient estimation and errors of inverse optimization. The implementation of our proposed method is available at https://github.com/mollnn/conditional-mixture. Zhimin Fan 0001, Mufan Guo, Ruoyu Fu, Yanwen Guo 0001, Jie Guo 0001 |
ACM Trans. Graph. | 6 |
| 2024 | Collaborative Completion and Segmentation for Partial Point Clouds With OutliersabstractOutliers will inevitably creep into the captured point cloud during 3D scanning, degrading cutting-edge models on various geometric tasks heavily. This paper looks at an intriguing question that whether point cloud completion and segmentation can promote each other to defeat outliers. To answer it, we propose a collaborative completion and segmentation network, termed CS-Net, for partial point clouds with outliers. Unlike most of existing methods, CS-Net does not need any clean (or say outlier-free) point cloud as input or any outlier removal operation. CS-Net is a new learning paradigm that makes completion and segmentation networks work collaboratively. With a cascaded architecture, our method refines the prediction progressively. Specifically, after the segmentation network, a cleaner point cloud is fed into the completion network. We design a novel completion network which harnesses the labels obtained by segmentation together with farthest point sampling to purify the point cloud and leverages KNN-grouping for better generation. Benefited from segmentation, the completion module can utilize the filtered point cloud which is cleaner for completion. Meanwhile, the segmentation module is able to distinguish outliers from target objects more accurately with the help of the clean and complete shape inferred by completion. Besides the designed collaborative mechanism of CS-Net, we establish a benchmark dataset of partial point clouds with outliers. Extensive experiments show clear improvements of our CS-Net over its competitors, in terms of outlier robustness and completion accuracy. Changfeng Ma, Yang Yang 0092, Jie Guo 0001, Mingqiang Wei, Chong-Jun Wang, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | MFFNet: multimodal feature fusion network for point cloud semantic segmentation
Dayong Ren, Zhengyi Wu, Jie Guo 0001, Mingqiang Wei, Yanwen Guo 0001 |
Vis. Comput. | 4 |
| 2023 | Symmetric Shape-Preserving Autoencoder for Unsupervised Real Scene Point Cloud CompletionabstractUnsupervised completion of real scene objects is of vital importance but still remains extremely challenging in preserving input shapes, predicting accurate results, and adapting to multi-category data. To solve these problems, we propose in this paper an Unsupervised Symmetric Shape-Preserving Autoencoding Network, termed USSPA, to predict complete point clouds of objects from real scenes. One of our main observations is that many natural and manmade objects exhibit significant symmetries. To accommodate this, we devise a symmetry learning module to learn from those objects and to preserve structural symmetries. Starting from an initial coarse predictor, our autoencoder refines the complete shape with a carefully designed upsampling refinement module. Besides the discriminative process on the latent space, the discriminators of our USSPA also take predicted point clouds as direct guidance, enabling more detailed shape prediction. Clearly different from previous methods which train each category separately, our USSPA can be adapted to the training of multi-category data in one pass through a classifier-guided discriminator, with consistent performance on single category. For more accurate evaluation, we contribute to the community a real scene dataset with paired CAD models as ground truth. Extensive experiments and comparisons demonstrate our superiority and generalization and show that our method achieves state-of-the-art performance on unsupervised completion of real scene objects. Changfeng Ma, Pengxiao Guo, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001 |
CVPR | 4 |
| 2023 | UHDNeRF: Ultra-High-Definition Neural Radiance FieldsabstractWe propose UHDNeRF, a new framework for novel view synthesis on the challenging ultra-high-resolution (e.g., 4K) real-world scenes. Previous NeRF methods are not specifically designed for rendering on extremely high resolutions, leading to burry results with notable detail-losing problems even though trained on 4K images. This is mainly due to the mismatch between the high-resolution inputs and the low-dimensional volumetric representation. To address this issue, we introduce an adaptive implicit-explicit scene representation with which an explicit sparse point cloud is used to boost the performance of an implicit volume on modeling subtle details. Specifically, we reconstruct the complex real-world scene with a frequency separation strategy that the implicit volume learns to represent the low-frequency properties of the whole scene, and the sparse point cloud is used for reproducing high-frequency details. To better explore the information embedded in the point cloud, we extract a global structure feature and a local point-wise feature from the point cloud for each sample located in the high-frequency regions. Furthermore, a patch-based sampling strategy is introduced to reduce the computational cost. The high-fidelity rendering results demonstrate the superiority of our method for retaining high-frequency details at 4K ultra-high-resolution scenarios against state-of-the-art NeRF-based solutions. Quewei Li, Feichao Li, Jie Guo 0001, Yanwen Guo 0001 |
ICCV | 3 |
| 2023 | Convolutional Self-attention Guided Graph Neural Network for Few-Shot Action Recognition
Jie Guo 0001, Yanwen Guo 0001 |
ICIC (2) | 2 |
| 2023 | Few-Shot Action Recognition with A Transductive Maximum Margin ClassifierabstractFew-shot action recognition aims to train a classifier that can generalize well when just a small number of labeled videos per class are given. We introduce a transductive maximum margin classifier for few-shot action recognition, which leverages the unlabeled query videos to improve the recognition performance in the test task. The basic idea of the classical maximum margin classifier is to search for a classifier with the largest geometric margin so that training data can be correctly classified. Due to the insufficient number of labeled videos in the support set, it is challenging to find such a classifier with good generalization ability. We observe that exploring the geometric relationship between the separating hyperplane of the classifier and the feature vectors of the query videos can bring improvements to the classifier. In order to improve data utilization efficiency in the few-shot setting, the class prototypes are also treated as examples, which participate in the iterative training process of the model. Experimental results on two action recognition datasets including Kinetics and Something-Something V2 show that our method achieves state-of-the-art performance. Jie Guo 0001, Yanwen Guo 0001 |
IJCNN | 2 |
| 2023 | SVBRDF Reconstruction by Transferring Lighting KnowledgeabstractAbstract The problem of reconstructing spatially‐varying BRDFs from RGB images has been studied for decades. Researchers found themselves in a dilemma: opting for either higher quality with the inconvenience of camera and light calibration, or greater convenience at the expense of compromised quality without complex setups. We address this challenge by introducing a two‐branch network to learn the lighting effects in images. The two branches, referred to as Light‐known and Light‐aware, diverge in their need for light information. The Light‐aware branch is guided by the Light‐known branch to acquire the knowledge of discerning light effects and surface reflectance properties, but without the reliance of light positions. Both branches are trained using the synthetic dataset, but during testing on real‐world cases without calibration, only the Light‐aware branch is activated. To facilitate a more effective utilization of various light conditions, we employ gated recurrent units (GRUs) to fuse the features extracted from different images. The two modules mutually benefit when multiple inputs are provided. We present our reconstructed results on both synthetic and real‐world examples, demonstrating high quality while maintaining a lightweight characteristic in comparison to previous methods. Pengfei Zhu 0001, Shuichang Lai, Mufan Chen, Jie Guo 0001, Yanwen Guo 0001 |
Comput. Graph. Forum | 4 |
| 2023 | Deep graph learning for spatially-varying indoor lighting prediction
Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Piaopiao Yu, Yan Zhang 0057, Yanwen Guo 0001 |
Sci. China Inf. Sci. | 2 |
| 2023 | Reflectance edge guided networks for detail-preserving intrinsic image decomposition
Quewei Li, Jie Guo 0001, Zhengyi Wu, Yanwen Guo 0001 |
Sci. China Inf. Sci. | 2 |
| 2023 | Elastic temporal alignment for few-shot action recognitionabstractAbstract Few‐shot action recognition aims to learn a classification model with good generalisation ability when trained with only a few labelled videos. However, it is difficult to learn discriminative feature representations for videos in such a setting. The Elastic Temporal Alignment (ETA) for few‐shot action recognition is proposed. First, a convolutional neural network is employed to extract feature representations of video frames sparsely sampled from videos. In order to obtain the similarity of two videos, a temporal alignment estimation function is utilised to estimate the matching score between each pair of frames from the two videos through an elastic alignment mechanism. The analysis shows that when we judge whether two frames from respective videos are matched, multiple adjacent frames in the videos should be considered, so as to embody the temporal information. Thus, before feeding per‐frame feature vectors of videos into the temporal alignment estimation function, a temporal message passing function is leveraged to propagate the information of per‐frame features in the temporal domain. The method has been evaluated on four action recognition datasets, including Kinetics, Something‐Something V2, HMDB51, and UCF101. The experimental results verify the effectiveness of ETA and show its superiority over state‐of‐the‐art methods. Chunlei Xu, Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001 |
IET Comput. Vis. | 4 |
| 2023 | Improving Open Set Domain Adaptation Using Image-to-Image Translation and Instance-Weighted Adversarial Learning
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001 |
J. Comput. Sci. Technol. | 3 |
| 2023 | Ultra-High Resolution SVBRDF Recovery from a Single ImageabstractExisting convolutional neural networks have achieved great success in recovering Spatially Varying Bidirectional Surface Reflectance Distribution Function (SVBRDF) maps from a single image. However, they mainly focus on handling low-resolution (e.g., 256 × 256) inputs. Ultra-High Resolution (UHR) material maps are notoriously difficult to acquire by existing networks because (1) finite computational resources set bounds for input receptive fields and output resolutions, and (2) convolutional layers operate locally and lack the ability to capture long-range structural dependencies in UHR images. We propose an implicit neural reflectance model and a divide-and-conquer solution to address these two challenges simultaneously. We first crop a UHR image into low-resolution patches, each of which are processed by a local feature extractor to extract important details. To fully exploit long-range spatial dependency and ensure global coherency, we incorporate a global feature extractor and several coordinate-aware feature assembly modules into our pipeline. The global feature extractor contains several lightweight material vision transformers that have a global receptive field at each scale and have the ability to infer long-term relationships in the material. After decoding globally coherent feature maps assembled by coordinate-aware feature assembly modules, the proposed end-to-end method is able to generate UHR SVBRDF maps from a single image with fine spatial details and consistent global structures. Jie Guo 0001, Shuichang Lai, Qinghao Tu, Chengzhi Tao, Changqing Zou, Yanwen Guo 0001 |
ACM Trans. Graph. | 1 |
| 2023 | Manifold Path Guiding for Importance Sampling Specular ChainsabstractComplex visual effects such as caustics are often produced by light paths containing multiple consecutive specular vertices (dubbed specular chains) , which pose a challenge to unbiased estimation in Monte Carlo rendering. In this work, we study the light transport behavior within a sub-path that is comprised of a specular chain and two non-specular separators. We show that the specular manifolds formed by all the sub-paths could be exploited to provide coherence among sub-paths. By reconstructing continuous energy distributions from historical and coherent sub-paths, seed chains can be generated in the context of importance sampling and converge to admissible chains through manifold walks. We verify that importance sampling the seed chain in the continuous space reaches the goal of importance sampling the discrete admissible specular chain. Based on these observations and theoretical analyses, a progressive pipeline, manifold path guiding , is designed and implemented to importance sample challenging paths featuring long specular chains. To our best knowledge, this is the first general framework for importance sampling discrete specular chains in regular Monte Carlo rendering. Extensive experiments demonstrate that our method outperforms state-of-the-art unbiased solutions with up to 40 × variance reduction, especially in typical scenes containing long specular chains and complex visibility. Zhimin Fan 0001, Pengpei Hong, Jie Guo 0001, Changqing Zou, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 3 |
| 2023 | MetaLayer: A Meta-Learned BSDF Model for Layered MaterialsabstractReproducing the appearance of arbitrary layered materials has long been a critical challenge in computer graphics, with regard to the demanding requirements of both physical accuracy and low computation cost. Recent studies have demonstrated promising results by learning-based representations that implicitly encode the appearance of complex (layered) materials by neural networks. However, existing generally-learned models often struggle between strong representation ability and high runtime performance, and also lack physical parameters for material editing. To address these concerns, we introduce MetaLayer , a new methodology leveraging meta-learning for modeling and rendering layered materials. MetaLayer contains two networks: a BSDFNet that compactly encodes layered materials into implicit neural representations, and a MetaNet that establishes the mapping between the physical parameters of each material and the weights of its corresponding implicit neural representation. A new positional encoding method and a well-designed training strategy are employed to improve the performance and quality of the neural model. As a new learning-based representation, the proposed MetaLayer model provides both fast responses to material editing and high-quality results for a wide range of layered materials, outperforming existing layered BSDF models. Jie Guo 0001, Zeru Li, Xueyan He, Beibei Wang 0002, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2023 | Local-to-Global Panorama Inpainting for Locale-Aware Indoor Lighting PredictionabstractPredicting panoramic indoor lighting from a single perspective image is a fundamental but highly ill-posed problem in computer vision and graphics. To achieve locale-aware and robust prediction, this problem can be decomposed into three sub-tasks: depth-based image warping, panorama inpainting and high-dynamic-range (HDR) reconstruction, among which the success of panorama inpainting plays a key role. Recent methods mostly rely on convolutional neural networks (CNNs) to fill the missing contents in the warped panorama. However, they usually achieve suboptimal performance since the missing contents occupy a very large portion in the panoramic space while CNNs are plagued by limited receptive fields. The spatially-varying distortion in the spherical signals further increases the difficulty for conventional CNNs. To address these issues, we propose a local-to-global strategy for large-scale panorama inpainting. In our method, a depth-guided local inpainting is first applied on the warped panorama to fill small but dense holes. Then, a transformer-based network, dubbed PanoTransformer, is designed to hallucinate reasonable global structures in the large holes. To avoid distortion, we further employ cubemap projection in our design of PanoTransformer. The high-quality panorama recovered at any locale helps us to capture spatially-varying indoor illumination with physically-plausible global structures and fine details. Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Yan Zhang 0057, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | ShadowMover: Automatically Projecting Real Shadows onto Virtual ObjectabstractInserting 3D virtual objects into real-world images has many applications in photo editing and augmented reality. One key issue to ensure the reality of the composite whole scene is to generate consistent shadows between virtual and real objects. However, it is challenging to synthesize visually realistic shadows for virtual and real objects without any explicit geometric information of the real scene or manual intervention, especially for the shadows on the virtual objects projected by real objects. In view of this challenge, we present, to our knowledge, the first end-to-end solution to fully automatically project real shadows onto virtual objects for outdoor scenes. In our method, we introduce the Shifted Shadow Map, a new shadow representation that encodes the binary mask of shifted real shadows after inserting virtual objects in an image. Based on the shifted shadow map, we propose a CNN-based shadow generation model named ShadowMover which first predicts the shifted shadow map for an input image and then automatically generates plausible shadows on any inserted virtual object. A large-scale dataset is constructed to train the model. Our ShadowMover is robust to various scene configurations without relying on any geometric information of the real scene and is free of manual intervention. Extensive experiments validate the effectiveness of our method. Piaopiao Yu, Jie Guo 0001, Zhenyu Chen 0001, Chen Wang 0149, Yan Zhang 0057, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Unsupervised Point Cloud Completion and Segmentation by Generative Adversarial Autoencoding NetworkabstractMost existing point cloud completion methods assume the input partial point cloud is clean, which is not practical in practice, and are Most existing point cloud completion methods assume the input partial point cloud is clean, which is not the case in practice, and are generally based on supervised learning. In this paper, we present an unsupervised generative adversarial autoencoding network, named UGAAN, which completes the partial point cloud contaminated by surroundings from real scenes and cutouts the object simultaneously, only using artificial CAD models as assistance. The generator of UGAAN learns to predict the complete point clouds on real data from both the discriminator and the autoencoding process of artificial data. The latent codes from generator are also fed to discriminator which makes encoder only extract object features rather than noises. We also devise a refiner for generating better complete cloud with a segmentation module to separate the object from background. We train our UGAAN with one real scene dataset and evaluate it with the other two. Extensive experiments and visualization demonstrate our superiority, generalization and robustness. Comparisons against the previous method show that our method achieves the state-of-the-art performance on unsupervised point cloud completion and segmentation on real data. Changfeng Ma, Yang Yang 0092, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001 |
NeurIPS | 3 |
| 2022 | NeuLighting: Neural Lighting for Free Viewpoint Outdoor Scene Relighting with Unconstrained Photo CollectionsabstractWe propose NeuLighting, a new framework for free viewpoint outdoor scene relighting from a sparse set of unconstrained in-the-wild photo collections. Our framework represents all the scene components as continuous functions parameterized by MLPs that take a 3D location and the lighting condition as input and output reflectance and necessary outdoor illumination properties. Unlike object-level relighting methods which often leverage training images with controllable and consistent indoor illumination, we concentrate on the more challenging outdoor situation where all the images are captured under arbitrary unknown illumination. The key to our method includes a neural lighting representation that compresses the per-image illumination into a disentangled latent vector, and a new free viewpoint relighting scheme that is robust to arbitrary lighting variations across images. The lighting representation is compressive to explain a wide range of illumination and can be easily fed into the query-based NeuLighting framework, enabling efficient shading effect evaluation under any kind of novel illumination. Furthermore, to produce high-quality cast shadows, we estimate the sun visibility map to indicate the shadow regions according to the scene geometry and the sun direction. Thanks to the flexible and explainable neural lighting representation, our system supports outdoor relighting with many different illumination sources, including natural images, environment maps, and time-lapse videos. The high-fidelity renderings under novel views and illumination prove the superiority of our method against state-of-the-art relighting solutions. Quewei Li, Jie Guo 0001, Feichao Li, Yanwen Guo 0001 |
SIGGRAPH Asia | 2 |
| 2022 | Real-time Deep Radiance Reconstruction from Imperfect CachesabstractAbstract Real‐time global illumination is a highly desirable yet challenging task in computer graphics. Existing works well solving this problem are mostly based on some kind of precomputed data (caches), while the final results depend significantly on the quality of the caches. In this paper, we propose a learning‐based pipeline that can reproduce a wide range of complex light transport phenomena, including high‐frequency glossy interreflection, at any viewpoint in real time (> 90 frames per‐second), using information from imperfect caches stored at the barycentre of every triangle in a 3D scene. These caches are generated at a precomputation stage by a physically‐based offline renderer at a low sampling rate (e.g., 32 samples per‐pixel) and a low image resolution (e.g., 64×16). At runtime, a deep radiance reconstruction method based on a dedicated neural network is then involved to reconstruct a high‐quality radiance map of full global illumination at any viewpoint from these imperfect caches, without introducing noise and aliasing artifacts. To further improve the reconstruction accuracy, a new feature fusion strategy is designed in the network to better exploit useful contents from cheap G‐buffers generated at runtime. The proposed framework ensures high‐quality rendering of images for moderate‐sized scenes with full global illumination effects, at the cost of reasonable precomputation time. We demonstrate the effectiveness and efficiency of the proposed pipeline by comparing it with alternative strategies, including real‐time path tracing and precomputed radiance transfer. Tao Huang 0026, Yadong Song, Jie Guo 0001, Chengzhi Tao, Zijing Zong, Xihao Fu, Hongshan Li, Yanwen Guo 0001 |
Comput. Graph. Forum | 3 |
| 2022 | SVBRDF Recovery from a Single Image with Highlights Using a Pre-trained Generative Adversarial NetworkabstractAbstract Spatially varying bi‐directional reflectance distribution functions (SVBRDFs) are crucial for designers to incorporate new materials in virtual scenes, making them look more realistic. Reconstruction of SVBRDFs is a long‐standing problem. Existing methods either rely on an extensive acquisition system or require huge datasets, which are non‐trivial to acquire. We aim to recover SVBRDFs from a single image, without any datasets. A single image contains incomplete information about the SVBRDF, making the reconstruction task highly ill‐posed. It is also difficult to separate between the changes in colour that are caused by the material and those caused by the illumination, without the prior knowledge learned from the dataset. In this paper, we use an unsupervised generative adversarial neural network (GAN) to recover SVBRDFs maps with a single image as input. To better separate the effects due to illumination from the effects due to the material, we add the hypothesis that the material is stationary and introduce a new loss function based on Fourier coefficients to enforce this stationarity. For efficiency, we train the network in two stages: reusing a trained model to initialize the SVBRDFs and fine‐tune it based on the input image. Our method generates high‐quality SVBRDFs maps from a single input photograph, and provides more vivid rendering results compared to the previous work. The two‐stage training boosts runtime performance, making it eight times faster than the previous work. Beibei Wang 0002, Jie Guo 0001, Nicolas Holzschuch |
Comput. Graph. Forum | 4 |
| 2022 | Point attention network for point cloud semantic segmentation
Dayong Ren, Zhengyi Wu, Piaopiao Yu, Jie Guo 0001, Mingqiang Wei, Yanwen Guo 0001 |
Sci. China Inf. Sci. | 5 |
| 2022 | Rendering discrete participating media using geometrical optics approximationabstractWe consider the scattering of light in participating media composed of sparsely and randomly distributed discrete particles. The particle size is expected to range from the scale of the wavelength to several orders of magnitude greater, resulting in an appearance with distinct graininess as opposed to the smooth appearance of continuous media. One fundamental issue in the physically-based synthesis of such appearance is to determine the necessary optical properties in every local region. Since these properties vary spatially, we resort to geometrical optics approximation (GOA), a highly efficient alternative to rigorous Lorenz—Mie theory, to quantitatively represent the scattering of a single particle. This enables us to quickly compute bulk optical properties for any particle size distribution. We then use a practical Monte Carlo rendering solution to solve energy transfer in the discrete participating media. Our proposed framework is the first to simulate a wide range of discrete participating media with different levels of graininess, converging to the continuous media case as the particle concentration increases. Jie Guo 0001, Bingyang Hu, Yuanqi Li, Yanwen Guo 0001, Lingqi Yan 0001 |
Comput. Vis. Media | 1 |
| 2022 | Efficient Light Probes for Real-Time Global IlluminationabstractReproducing physically-based global illumination (GI) effects has been a long-standing demand for many real-time graphical applications. In pursuit of this goal, many recent engines resort to some form of light probes baked in a precomputation stage. Unfortunately, the GI effects stemming from the precomputed probes are rather limited due to the constraints in the probe storage, representation or query. In this paper, we propose a new method for probe-based GI rendering which can generate a wide range of GI effects, including glossy reflection with multiple bounces, in complex scenes. The key contributions behind our work include a gradient-based search algorithm and a neural image reconstruction method. The search algorithm is designed to reproject the probes' contents to any query viewpoint, without introducing parallax errors, and converges fast to the optimal solution. The neural image reconstruction method, based on a dedicated neural network and several G-buffers, tries to recover high-quality images from low-quality inputs due to limited resolution or (potential) low sampling rate of the probes. This neural method makes the generation of light probes efficient. Moreover, a temporal reprojection strategy and a temporal loss are employed to improve temporal stability for animation sequences. The whole pipeline runs in realtime (>30 frames per second) even for high-resolution (1920×1080) outputs, thanks to the fast convergence rate of the gradient-based search algorithm and a light-weight design of the neural network. Extensive experiments on multiple complex scenes have been conducted to show the superiority of our method over the state-of-the-arts. Jie Guo 0001, Zijing Zong, Yadong Song, Xihao Fu, Chengzhi Tao, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2022 | Video Vectorization via Bipartite Diffusion Curves Propagation and OptimizationabstractWe propose a new video vectorization approach for converting videos in the raster format to vector representation with the benefits of resolution independence and compact storage. Through classifying extracted curves in each video frame into salient ones and non-salient ones, we introduce a novel bipartite diffusion curves (BDCs) representation in order to preserve both important image features such as sharp boundaries and regions with smooth color variation. This bipartite representation allows us to propagate non-salient curves across frames such that the propagation, in conjunction with geometry optimization and color optimization of salient curves, ensures the preservation of fine details within each frame and across different frames, and meanwhile, achieves good spatial-temporal coherence. Thorough experiments on a variety of videos show that our method is capable of converting videos to the vector representation with low reconstruction errors, low computational cost, and fine details, demonstrating our superior performance over the state of the art. We also show that, when used for video upsampling, our method produces results comparable to video super-resolution. Yuanqi Li, Chuan Wang 0001, Jie Guo 0001, Jue Wang 0001, Yanwen Guo 0001, Wenping Wang 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material RecognitionabstractIn this paper, we aim to recognize materials with combined use of auditory and visual perception. To this end, we construct a new dataset named GLAudio that consists of both the geometry of the object being struck and the sound captured from either modal sound synthesis (for virtual objects) or real measurements (for real objects). Besides global geometries, our dataset also takes local geometries around different hitpoints into consideration. This local information is less explored in existing datasets. We demonstrate that local geometry has a greater impact on the sound than the global geometry and offers more cues in material recognition. To extract features from different modalities and perform proper fusion, we propose a new deep neural network GLAVNet that comprises multiple branches and a well-designed fusion module. Once trained on GLAudio, our GLAVNet provides state-of-the-art performance on material identification and supports fine-grained material categorization. Fengmin Shi, Jie Guo 0001, Yanwen Guo 0001 |
CVPR | 2 |
| 2021 | Hierarchical Disentangled Representation Learning for Outdoor Illumination Estimation and EditingabstractData-driven sky models have gained much attention in outdoor illumination prediction recently, showing superior performance against analytical models. However, naively compressing an outdoor panorama into a low-dimensional latent vector, as existing models have done, causes two major problems. One is the mutual interference between the HDR intensity of the sun and the complex textures of the surrounding sky, and the other is the lack of fine-grained control over independent lighting factors due to the entangled representation. To address these issues, we propose a hierarchical disentangled sky model (HDSky) for outdoor illumination prediction. With this model, any outdoor panorama can be hierarchically disentangled into several factors based on three well-designed autoencoders. The first autoencoder compresses each sunny panorama into a sky vector and a sun vector with some constraints. The second autoencoder and the third autoencoder further disentangle the sun intensity and the sky intensity from the sun vector and the sky vector with several customized loss functions respectively. Moreover, a unified framework is designed to predict all-weather sky information from a single outdoor image. Through extensive experiments, we demonstrate that the proposed model significantly improves the accuracy of outdoor illumination prediction. It also allows users to intuitively edit the predicted panorama (e.g., changing the position of the sun while preserving others), without sacrificing physical plausibility. Piaopiao Yu, Jie Guo 0001, Hongwei Che, Yanwen Guo 0001 |
ICCV | 2 |
| 2021 | Dual attention autoencoder for all-weather outdoor lighting estimation
Piaopiao Yu, Jie Guo 0001, Longhai Wu, Yanwen Guo 0001 |
Sci. China Inf. Sci. | 2 |
| 2021 | Image Stitching Based on Semantic Planar Region ConsensusabstractImage stitching for two images without a global transformation between them is notoriously difficult. In this paper, noticing the importance of semantic planar structures under perspective geometry, we propose a new image stitching method which stitches images by allowing for the alignment of a set of matched dominant semantic planar regions. Clearly different from previous methods resorting to plane segmentation, the key to our approach is to utilize rich semantic information directly from RGB images to extract semantic planar image regions with a deep Convolutional Neural Network (CNN). We specifically design a module implementing our newly proposed clustering loss to make full use of existing semantic segmentation networks to accommodate region segmentation. To train the network, a dataset for semantic planar region segmentation is constructed. With the prior of semantic planar region, a set of local transformation models can be obtained by constraining matched regions, enabling more precise alignment in the overlapping area. We also use this prior to estimate a transformation field over the whole image. The final mosaic is obtained by mesh-based optimization which maintains high alignment accuracy and relaxes similarity transformation at the same time. Extensive experiments with both qualitative and quantitative comparisons show that our method can deal with different situations and outperforms the state-of-the-arts on challenging scenes. Aocheng Li, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Highlight-aware two-stream network for single-image SVBRDF acquisitionabstractThis paper addresses the task of estimating spatially-varying reflectance (i.e., SVBRDF) from a single, casually captured image. Central to our method is a highlight-aware (HA) convolution operation and a two-stream neural network equipped with proper training losses. Our HA convolution, as a novel variant of standard (ST) convolution, directly modulates convolution kernels under the guidance of automatically learned masks representing potentially overexposed highlight regions. It helps to reduce the impact of strong specular highlights on diffuse components and at the same time, hallucinates plausible contents in saturated regions. Considering that variation of saturated pixels also contains important cues for inferring surface bumpiness and specular components, we design a two-stream network to extract features from two different branches stacked by HA convolutions and ST convolutions, respectively. These two groups of features are further fused in an attention-based manner to facilitate feature selection of each SVBRDF map. The whole network is trained end to end with a new perceptual adversarial loss which is particularly useful for enhancing the texture details. Such a design also allows the recovered material maps to be disentangled. We demonstrate through quantitative analysis and qualitative visualization that the proposed method is effective to recover clear SVBRDFs from a single casually captured image, and performs favorably against state-of-the-arts. Since we impose very few constraints on the capture process, even a non-expert user can create high-quality SVBRDFs that cater to many graphical applications. Jie Guo 0001, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2021 | Volumetric appearance stylization with stylizing kernel prediction networkabstractThis paper aims to efficiently construct the volume of heterogeneous single-scattering albedo for a given medium that would lead to desired color appearance. We achieve this goal by formulating it as a volumetric style transfer problem in which an input 3D density volume is stylized using color features extracted from a reference 2D image. Unlike existing algorithms that require cumbersome iterative optimizations, our method leverages a feed-forward deep neural network with multiple well-designed modules. At the core of our network is a stylizing kernel predictor (SKP) that extracts multi-scale feature maps from a 2D style image and predicts a handful of stylizing kernels as a highly non-linear combination of the feature maps. Each group of stylizing kernels represents a specific style. A volume autoencoder (VolAE) is designed and jointly learned with the SKP to transform a density volume to an albedo volume based on these stylizing kernels. Since the autoencoder does not encode any style information, it can generate different albedo volumes with a wide range of appearance once training is completed. Additionally, a hybrid multi-scale loss function is used to learn plausible color features and guarantee temporal coherence for time-evolving volumes. Through comprehensive experiments, we validate the effectiveness of our method and show its superiority by comparing against state-of-the-arts. We show that with our method a novice user can easily create a diverse set of realistic translucent effects for 3D models (either static or dynamic), neglecting any cumbersome process of parameter tuning. Jie Guo 0001, Zijing Zong, Jingwu He, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2021 | ExtraNet: real-time extrapolated rendering for low-latency temporal supersamplingabstractBoth the frame rate and the latency are crucial to the performance of realtime rendering applications such as video games. Spatial supersampling methods, such as the Deep Learning SuperSampling (DLSS), have been proven successful at decreasing the rendering time of each frame by rendering at a lower resolution. But temporal supersampling methods that directly aim at producing more frames on the fly are still not practically available. This is mainly due to both its own computational cost and the latency introduced by interpolating frames from the future. In this paper, we present ExtraNet, an efficient neural network that predicts accurate shading results on an extrapolated frame, to minimize both the performance overhead and the latency. With the help of the rendered auxiliary geometry buffers of the extrapolated frame, and the temporally reliable motion vectors, we train our ExtraNet to perform two tasks simultaneously: irradiance in-painting for regions that cannot find historical correspondences, and accurate ghosting-free shading prediction for regions where temporal information is available. We present a robust hole-marking strategy to automate the classification of these tasks, as well as the data generation from a series of high-quality production-ready scenes. Finally, we use lightweight gated convolutions to enable fast inference. As a result, our ExtraNet is able to produce plausibly extrapolated frames without easily noticeable artifacts, delivering a 1.5× to near 2× increase in frame rates with minimized latency in practice. Jie Guo 0001, Xihao Fu, Liqiang Lin, Hengjun Ma, Yanwen Guo 0001, Shiqiu Liu, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2020 | Hybrid Models for Open Set Recognition
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001 |
ECCV (3) | 3 |
| 2020 | Deep Surface Normal Estimation on the 2-Sphere with Confidence Guided Semantic Attention
Quewei Li, Jie Guo 0001, Qinyu Tang, Wenxiu Sun, Jin Zeng 0004, Yanwen Guo 0001 |
ECCV (24) | 2 |
| 2020 | Two-Stage Depth Video Recovery with Spatiotemporal CoherenceabstractThis paper proposes a practical two-stage method to enhance the depth channel of low-quality RGB-D videos. At the first stage of the proposed method, we select several key frames from an input RGB-D video and recover them with a new MRF-regularized low-rank matrix completion method. At the second stage, we introduce a novel Laplacian smoothing algorithm to smoothly expand the recovered key depth frames to other in-between frames. With associated weights extracted from color images and a depth confidence constraint, we are able to recover each in-between frame faithfully and guarantee long-range temporal consistency. Experiments show that our method can generate high-quality and spatiotemporally coherent RGB-D videos for a wide range of scene configurations and achieve the state-of-the-art performance. Several applications further validate the effectiveness of the proposed method. Quewei Li, Jie Guo 0001, Qinyu Tang, Yanwen Guo 0001, Jinghui Qian |
ICME | 2 |
| 2020 | DeepBRDF: A Deep Representation for Manipulating Measured BRDFabstractAbstract Effective compression of densely sampled BRDF measurements is critical for many graphical or vision applications. In this paper, we present DeepBRDF, a deep‐learning‐based representation that can significantly reduce the dimensionality of measured BRDFs while enjoying high quality of recovery. We consider each measured BRDF as a sequence of image slices and design a deep autoencoder with a masked L 2 loss to discover a nonlinear low‐dimensional latent space of the high‐dimensional input data. Thorough experiments verify that the proposed method clearly outperforms PCA‐based strategies in BRDF data compression and is more robust. We demonstrate the effectiveness of DeepBRDF with two applications. For BRDF editing, we can easily create a new BRDF by navigating on the low‐dimensional manifold of DeepBRDF, guaranteeing smooth transitions and high physical plausibility. For BRDF recovery, we design another deep neural network to automatically generate the full BRDF data from a single input image. Aided by our DeepBRDF learned from real‐world materials, a wide range of reflectance behaviors can be recovered with high accuracy. Bingyang Hu, Jie Guo 0001, Yanwen Guo 0001 |
Comput. Graph. Forum | 2 |
| 2020 | BRDF Analysis with Directional Statistics and its ApplicationsabstractData-driven BRDF models using real material measurements have become increasingly prevalent due to the development of novel gonioreflectometers, but efficient use of these models in many graphical applications remains challenging due to the few functionalities the raw data could provide. To ameliorate this issue, we propose to analyze BRDFs using directional statistics for better handling and exploring measured materials, especially isotropic materials, with efficient computation and compact storage. We conduct a thorough statistical analysis on both analytical BRDF models and measured materials from the MERL database. We show that different aspects of visual appearance can be characterized by different spherical moments, from which several descriptive measures can be derived to further facilitate their usage. We demonstrate how these measures are best leveraged in some graphical applications including gamut mapping using a new BRDF similarity measure, BRDF or SVBRDF reconstruction based on material clustering, and importance sampling for measured materials based on fast extracted GGX distributions. We finally show the potential of our approach in the categorization of surface reflectance types which is common for traditional photon mapping. Jie Guo 0001, Yanwen Guo 0001, Jingui Pan, Wenzhou Lu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Data-Driven Indoor Scene Modeling from a Single Color Image with Iterative Object Segmentation and Model RetrievalabstractWe propose a new method for modeling the indoor scene from a single color image. With our system, the user only needs to drag a few semantic bounding boxes surrounding the objects of interest. Our system then automatically finds the most similar 3D models from the ShapeNet model repository and aligns them with the corresponding objects of interest. To achieve this, each 3D model is represented as a group of view-dependent representations generated from a set of synthesized views. We iteratively conduct object segmentation and 3D model retrieval, based on the observation that good segmentation of the objects of interest can significantly improve the accuracy of model retrieval and make it robust to cluttered background and occlusions, and in turn, the retrieved 3D models can be used to assist with object segmentation. Segmentation of all objects of interest is achieved simultaneously under a unified multi-labeling framework which fully utilizes the correspondences between the objects of interest and retrieved model images. Besides, we propose a new method to estimate the scene layout of the input image with the segmentation masks, which helps compose the resulting scene and further improves the modeling result remarkably. We verify the effectiveness of our approach through experimenting with a variety of indoor images and comparing against the relevant methods. Mingming Liu 0004, Jun Wang 0039, Jie Guo 0001, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | Learning Actor Relation Graphs for Group Activity RecognitionabstractModeling relation between actors is important for recognizing group activity in a multi-person scene. This paper aims at learning discriminative relation between actors efficiently using deep models. To this end, we propose to build a flexible and efficient Actor Relation Graph (ARG) to simultaneously capture the appearance and position relation between actors. Thanks to the Graph Convolutional Network, the connections in ARG could be automatically learned from group activity videos in an end-to-end manner, and the inference on ARG could be efficiently performed with standard matrix operations. Furthermore, in practice, we come up with two variants to sparsify ARG for more effective modeling in videos: spatially localized ARG and temporal randomized ARG. We perform extensive experiments on two standard group activity recognition datasets: the Volleyball dataset and the Collective Activity dataset, where state-of-the-art performance is achieved on both datasets. We also visualize the learned actor graphs and relation features, which demonstrate that the proposed ARG is able to capture the discriminative relation information for group activity recognition. Jianchao Wu, Limin Wang 0002, Jie Guo 0001, Gangshan Wu |
CVPR | 4 |
| 2019 | A Data-Driven Framework for Appearance Editing of Measured MaterialsabstractIn this paper, we present an efficient framework for editing the appearance of measured materials. First, we separate the material data into diffuse and specular parts, which is important for measured BRDFs since editing without distinguishing the two parts makes the results less predictive. Additionally, the albedo of each part is calculated for later color editing. Next, we use a novel method to cluster the materials by their gloss levels and select the representative material of every level. With this at hand, the roughness, i.e., shape of the specular highlights can be adjusted among different levels. Finally, we reconstruct the edited materials which can be directly used in both online and offline rendering. We test the framework on the MERL dataset to validate the effectiveness of our method. Jie Guo 0001, Bingyang Hu, Yanwen Guo 0001, Jingui Pan |
ICME | 2 |
| 2019 | Temporal Segment Convolutional Kernel Networks for Sequence Modeling of VideosabstractSequence modeling is crucial for video action recognition. In this paper, we propose temporal segment convolutional kernel networks (TS-CKN), where we take advantage of convolutional neural networks to facilitate the extraction of appearance features, while time sequence is modeled with deep kernel networks. We employ the kernel methods to capture time-varying information of videos and propose a training method for kernel map approximation by matrix backpropagation. This leads to the model named deep kernel networks which can be easily integrated with existing deep learning models such as Resnet. Our approach also samples several video clips sparsely in the video and unifies class predictions from all clips. More importantly, all parameters of our model can be learned by stochastic optimization in an end-to-end manner. We evaluate our method on two standard action recognition datasets including HMDB-51 and UCF-101, achieving the state-of-the-art results. Yanwen Guo 0001, Zhicheng Yan 0001, Jie Guo 0001 |
ICME | 4 |
| 2019 | Deep Spherical Gaussian Illumination Estimation for Indoor SceneabstractIn this paper, we propose a learning-based method to estimate high dynamic range (HDR) indoor illumination from only a single low dynamic range (LDR) photograph of limited field-of-view. Considering the extreme complexity of indoor illumination that is virtually impossible to reconstruct perfectly, we choose to encode the environmental illumination in Spherical Gaussian (SG) functions with fixed centering directions and bandwidth and only allow the weights vary. An end-to-end convolutional neural network (CNN) is designed and trained to build the complex relationship between a photograph and its illumination represented by SG functions. Moreover, we employ a masked L2 loss instead of naive L2 loss to avoid the loss of high frequency information, and propose a glossy loss to improve the rendering quality. Our experiments demonstrate that the proposed approach outperforms the state-of-the-arts both qualitatively and quantitatively. Jie Guo 0001, Xiufen Cui, Yanwen Guo 0001, Piaopiao Yu |
MMAsia | 2 |
| 2019 | Label transfer between images and 3D shapes via local correspondence encoding
Jie Guo 0001, Huikun Liu, Mingming Liu 0004, Yang Liu 0014, Yanwen Guo 0001 |
Comput. Aided Geom. Des. | 3 |
| 2019 | 3D model retrieval and pose estimation for indoor images by simulating scene context
Mingming Liu 0004, Jie Guo 0001, Yanwen Guo 0001 |
Graph. Model. | 2 |
| 2019 | Fractional gaussian fields for modeling and rendering of spatially-correlated mediaabstractTransmission of radiation through spatially-correlated media has demonstrated deviations from the classical exponential law of the corresponding uncorrelated media. In this paper, we propose a general, physically-based method for modeling such correlated media with non-exponential decay of transmittance. We describe spatial correlations by introducing the Fractional Gaussian Field (FGF), a powerful mathematical tool that has proven useful in many areas but remains under-explored in graphics. With the FGF, we study the effects of correlations in a unified manner, by modeling both high-frequency, noise-like fluctuations and k -th order fractional Brownian motion (fBm) with a stochastic continuity property. As a result, we are able to reproduce a wide variety of appearances stemming from different types of spatial correlations. Compared to previous work, our method is the first that addresses both short-range and long-range correlations using physically-based fluctuation models. We show that our method can simulate different extents of randomness in spatially-correlated media, resulting in a smooth transition in a range of appearances from exponential falloff to complete transparency. We further demonstrate how our method can be integrated into an energy-conserving RTE framework with a well-designed importance sampling scheme and validate its ability compared to the classical transport theory and previous work. Jie Guo 0001, Bingyang Hu, Lingqi Yan 0001, Yanwen Guo 0001 |
ACM Trans. Graph. | 1 |
| 2019 | GradNet: unsupervised deep screened poisson reconstruction for gradient-domain renderingabstractMonte Carlo (MC) methods for light transport simulation are flexible and general but typically suffer from high variance and slow convergence. Gradientdomain rendering alleviates this problem by additionally generating image gradients and reformulating rendering as a screened Poisson image reconstruction problem. To improve the quality and performance of the reconstruction, we propose a novel and practical deep learning based approach in this paper. The core of our approach is a multi-branch auto-encoder, termed GradNet, which end-to-end learns a mapping from a noisy input image and its corresponding image gradients to a high-quality image with low variance. Once trained, our network is fast to evaluate and does not require manual parameter tweaking. Due to the difficulty in preparing ground-truth images for training, we design and train our network in a completely unsupervised manner by learning directly from the input data. This is the first solution incorporating unsupervised deep learning into the gradient-domain rendering framework. The loss function is defined as an energy function including a data fidelity term and a gradient fidelity term. To further reduce the noise of the reconstructed image, the loss function is reinforced by adding a regularizer constructed from selected rendering-specific features. We demonstrate that our method improves the reconstruction quality for a diverse set of scenes, and reconstructing a high-resolution image takes far less than one second on a recent GPU. Jie Guo 0001, Quewei Li, Yuting Qiang, Bingyang Hu, Yanwen Guo 0001, Lingqi Yan 0001 |
ACM Trans. Graph. | 1 |
| 2018 | Single Image Highlight Removal with a Sparse and Low-Rank Reflection Model
Jie Guo 0001, Zuojian Zhou, Limin Wang 0002 |
ECCV (4) | 1 |
| 2018 | Image-Based 3D Model Retrieval for Indoor Scenes by Simulating Scene ContextabstractWe propose a single image-based 3D model retrieval method for indoor scenes. By simulating the scene context of the input image, our method is able to handle several challenging scenarios featuring cluttered backgrounds and severe occlusions. To use our system, the user only needs to drag a few semantic bounding boxes for the query objects. The proposed approach then retrieves the most similar 3D models from the ShapeNet model repository, and aligns them with the corresponding objects automatically. This requires that the 3D models are represented by calibrated view-dependent visual elements learned from the rendered views. With the estimated occlusion relationships, the rendered model images are stacked at the corresponding locations to simulate the scene context. By conducting matching between these synthesized scenes and the input image, the most similar 3D models under the approximate poses are retrieved. Moreover, we show that the retrieving time can be significantly reduced based on a novel greedy algorithm. Experimental results demonstrate the effectiveness of our proposed method. Mingming Liu 0004, Jingwu He, Jie Guo 0001, Yanwen Guo 0001 |
ICIP | 4 |
| 2018 | Depth Restoration with Normal-Guided Multiresolution SuperpixelabstractIn this paper, we propose a depth restoration method using a novel superpixel technique. Guided by a normal map reconstructed from the raw depth data, this technique over-segments RGB-D images into many small regions where their depth is assumed to be smooth. As the raw depth data is incomplete, we further introduce a depth confidence map to identify the regions which are more reliable. With the produced superpixels, we can restore the incomplete depth map using a per-superpixel linear regression. A multiresolution superpixel strategy is employed when some superpixels do not contain enough valid data. Experiments show that the proposed depth restoration method can effectively fill the wide gaps along depth discontinuities without blurring the object boundaries and the depth discontinuities. Jinghui Qian, Jie Guo 0001, Jingui Pan |
ICME | 2 |
| 2018 | A Physically-based Appearance Model for Special Effect PigmentsabstractAbstract An appearance model for materials adhered with massive collections of special effect pigments has to take both high‐frequency spatial details (e.g., glints) and wave‐optical effects (e.g., iridescence) due to thin‐film interference into account. However, either phenomenon is challenging to characterize and simulate in a physically accurate way. Capturing these fascinating effects in a unified framework is even harder as the normal distribution function and the reflectance term are highly correlated and cannot be treated separately. In this paper, we propose a multi‐scale BRDF model for reproducing the main visual effects generated by the discrete assembly of special effect pigments, enabling a smooth transition from fine‐scale surface details to large‐scale iridescent patterns. We demonstrate that the wavelength‐dependent reflectance inside the pixel's footprint follows a Gaussian distribution according to the central limit theorem, and is closely related to the distribution of the thin‐film's thickness. We efficiently determine the mean and the variance of this Gaussian distribution for each pixel whose closed‐form expressions can be derived by assuming that the thin‐film's thickness is uniformly distributed. To validate its effectiveness, the proposed model is compared against some previous methods and photographs of actual materials. Furthermore, since our method does not require any scene‐dependent precomputation, the distribution of thickness is allowed to be spatially‐varying. Jie Guo 0001, Yanwen Guo 0001, Jingui Pan |
Comput. Graph. Forum | 1 |
| 2018 | A retroreflective BRDF model based on prismatic sheeting and microfacet theory
Jie Guo 0001, Yanwen Guo 0001, Jingui Pan |
Graph. Model. | 1 |
| 2017 | A physically-based BRDF model for retroreflectionabstractWe introduce an analytical BRDF model tailored to the realistic rendering of retroreflective materials. Though these specially-designed materials are frequently encountered in our daily life, such as high-visibility clothing and traffic signs, their reflectance behaviors are barely analyzed and simulated in a physically meaningful manner in computer graphics. To understand how retroreflection is actually generate, we conduct a detailed optical analysis on a typical prismatic sheet which is the most widely used retroreflector in the material industry. In particular, the effective retroreflective area of the sheet is studied theoretically and experimentally. Based on the study, we derive a practical BRDF model that takes into account all types of reflection from the sheet, including mirror reflection, retroreflection, and diffuse reflection. Experiments confirm that our BRDF can produce visually pleasing rendered images with high fidelity retroreflection effects which closely match those generated by the prismatic sheet. Jie Guo 0001, Jingui Pan |
CGI | 1 |
| 2017 | Rendering Thin Transparent Layers with Extended Normal Distribution FunctionsabstractRealistic Rendering of thin transparent layers bounded by rough surfaces involves substantial expense of computation time to account for multiple internal reflections. Resorting to Monte Carlo rendering for such material is usually impractical since recursive importance sampling is inevitable. To reduce the burden of sampling for simulating subsurface scattering and hence improve rendering performance, we adapt the microfacet model to the material with a single thin layer by introducing the extended normal distribution function (ENDF), a new representation of this model, to express visually perceived roughness due to multiple bounces of reflections and refractions. With such a representation, both surface reflection and subsurface scattering can be treated in the same microfacet framework, and the sampling process can be reduced to only once for each bounce of scattering. We derive analytical expressions of the ENDF for several cases using joint spherical warping. We also show how to choose proper shadowing-masking and Fresnel terms to make the proposed bidirectional scattering distribution function (BSDF) model energy-conserving. Experiments demonstrate that our model can be easily incorporated into a Monte Carlo path tracer with little extra computational and storage overhead, enabling some real-time applications. Jie Guo 0001, Jinghui Qian, Yanwen Guo 0001, Jingui Pan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Real-time rendering of refracting transmissive objects with multi-scale rough surfaces
Jie Guo 0001, Jingui Pan |
Vis. Comput. | 1 |
| 2015 | QoS-Aware Web Service Recommendation Using Collaborative Filtering with PGraphabstractWeb service recommendation plays an important role in building reliable service-oriented systems for both the service providers and the active users. However, with the proliferation of web services on the World Wide Web, traditional service recommendation is hard to accurately provide customized services to active users. In this paper, we propose a novel web service recommender model using collaborative filtering to improve the prediction of Quality-of-Services. Benefiting from the accuracy of hybrid recommenders, we extend the idea of optimized predicting order and design the Graph to describe the neighborhood. Furthermore, a new algorithm using adjusted topological sorting for Graph is proposed to generate the optimized order while predicting. Finally, we conduct extensive experiments to evaluate our proposed model, in which a real data set with 1.5 million invocation information is taken as input. The experiment results show that our model achieves higher prediction accuracy than other models. Zuojian Zhou, Jie Guo 0001, Jingui Pan |
ICWS | 3 |
| 2014 | Real-time simulating and rendering of layered dust
Jie Guo 0001, Jingui Pan |
Vis. Comput. | 1 |
| 2013 | Real-Time Multi-scale Refraction under All-Frequency Environmental LightingabstractWe present a novel approach for real-time rendering of transparent objects with multi-scale rough refraction under all-frequency environmental lighting. In our rendering system, the normal distribution function (NDF) of the surface bumps is expressed via mixtures of von Mises-Fisher (vMF) distributions, while the corresponding perturbation of refracted rays (RDF) can be approximated as a spherical warp of NDF, taking into account correction factors. Under the assumption of two-sided refraction, we first generate the vMF lobes of the front surface refraction. Based on these lobes, we then estimate the RDF of back surface as a warped convolution of macroscopic and microscopic NDFs within the footprint of each lobe. Moreover, summed-area tables and dual-paraboloid maps are adopted to efficiently represent the environmental lighting in order to support fully dynamic lighting conditions. We finally get a coherent rendering at the different scales that reproduces high-quality meso- and micro-geometric effects under real-time performance. Jie Guo 0001, Jingui Pan |
CAD/Graphics | 1 |
| 2013 | Line segment sampling with blue-noise propertiesabstractLine segment sampling has recently been adopted in many rendering algorithms for better handling of a wide range of effects such as motion blur, defocus blur and scattering media. A question naturally raised is how to generate line segment samples with good properties that can effectively reduce variance and aliasing artifacts observed in the rendering results. This paper studies this problem and presents a frequency analysis of line segment sampling. The analysis shows that the frequency content of a line segment sample is equivalent to the weighted frequency content of a point sample. The weight introduces anisotropy that smoothly changes among point samples, line segment samples and line samples according to the lengths of the samples. Line segment sampling thus makes it possible to achieve a balance between noise (point sampling) and aliasing (line sampling) under the same sampling rate. Based on the analysis, we propose a line segment sampling scheme to preserve blue-noise properties of samples which can significantly reduce noise and aliasing artifacts in reconstruction results. We demonstrate that our sampling scheme improves the quality of depth-of-field rendering, motion blur rendering, and temporal light field reconstruction. Xin Sun 0014, Kun Zhou 0001, Jie Guo 0001, Guofu Xie, Jingui Pan, Baining Guo |
ACM Trans. Graph. | 3 |
| 2011 | Real-time rendering with complex natural illuminationabstractComplex natural lighting provides better realism compared to traditional artificial lights, which has been increasing in real-time graphic applications, such as video games and virtual reality. However, rendering with complex natural illumination proves quite costly owing to a spherical space integral. This paper addresses the problem of real-time rendering scenes under complex environment lighting that captures soft shadows and glossy reflections. The environment lighting is separated into low-frequency lighting component and highlight component based on a rejection-based sampling method. Scenes lit under low-frequency lighting are approximated via spherical harmonic lighting, while highlight illumination effects such as soft shadows and glossy reflections are simulated by a small set of sampled virtual point lights (VPL). Experimental results demonstrate that our method can be run in real-time under complex natural lighting without precomputation. Jie Guo 0001, Jingui Pan |
ICME | 1 |