Yanwen Guo 0001

dblp:44/185-1 · also Yan-Wen Guo 0001 · DBLP profile ↗
← Back
157ranked-venue papers
10as first author
81since 2021 · last 2026
0000-0002-7605-5206ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 133 · 9 first-author · 65 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 9 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Adaptive Structural Transition and Residual-Guided Local Swap for Post-Training Pruning of Large Language Models
Keqin Chen, Yanwen Guo 0001
ICIC (15)7
2026 Bounding Stratified Bernoulli Impulses for Ray Marching Gaussian Process Implicit Surfaces
abstract
The theory of light transport on Gaussian process implicit surface (GPIS) provides a unified framework for rendering surfaces, participating media, and the intermediate spectrum. However, previous approaches rely on brute-force ray marching for surface intersections, requiring full noise evaluations at each marching point, whether using multivariate Gaussian sampling or sparse convolution noise approximation. This imposes a severe limitation on the rendering efficiency. In this paper, we derive bounds to significantly reduce the total number of full noise evaluations, leading to efficient ray marching for ray-surface intersections. We introduce stratified Bernoulli impulses, enabling a fast point-level bound for individual realizations to replace unnecessary full noise evaluations. To further reduce the number of point-level bound evaluations, we propose a region-level bound, leveraging a spatial acceleration structure to prune probabilistically empty regions, thereby avoiding unnecessary marching points in advance. By combining these two bounds, our bounded ray marching accelerates ray-surface intersections in GPIS, and consequently significantly improves overall GPIS rendering efficiency. Code for this paper are at https://github.com/Cchen-77/bounded-gpis.
Zhimin Fan 0001, Lingqi Yan 0001, Junqiu Zhu, Yanwen Guo 0001, Kun Zhou 0001, Jie Guo 0001
ACM Trans. Graph.5
2026 Efficient Fur and Hair Multiple Scattering Using Volumetric Approximation
abstract
In this paper, we propose an efficient method for hair and fur rendering that approximates multiple scattering as volumetric light transport, achieving near path-traced visual fidelity at a drastically reduced cost. Multiple scattering among hair fibers is crucial for a realistic appearance, especially for light-colored hair, but it is extremely expensive to simulate directly. Prior approximations, such as dual scattering and fur BSSRDF, improve performance but often fail for dense hair/fur, yielding an overly dry or blurred appearance. Unlike previous volumetric hair rendering solutions, we treat dense hair/fur as a highly anisotropic participating medium to capture soft volumetric illumination, while still using explicit fiber geometry for direct lighting to preserve fine details. We accumulate hair fiber coverage in screen space and stochastically sample a single-scattering event to approximate higher-order scattering. Our method reproduces the rich appearance of hair and fur resulting from multiple scattering, while running about 7–10× faster than path tracing, making it suitable for use in production. We also demonstrate that our approach is robust under a variety of lighting conditions.
Ruike Hu, Junqiu Zhu, Minghao Lin, Ruian Zhang, Lu Wang 0007, Jie Guo 0001, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.7
2026 Weakly-Supervised Shape Multi-Completion of Point Clouds by Structural Decomposition
abstract
The challenge of transforming partial point clouds into complete meshes still persists, with current methods facing issues like data accessibility constraint, shape preservation failure and poor robustness on real-scan data. Drawing inspiration from the structural information of objects to enhance the completion, we introduce an innovative weakly-supervised shape completion method leveraging structural decomposition without the necessity of SDFs during training. By representing objects as abstract structural frameworks and part details, our method initiates by forecasting the structure of the input partial point clouds, and individually restore each component through part decomposition completion and generation. Extracted part details are represented in images, which are porous and incomplete. Hence, we utilize a completion network to complete such details. For multiple results generation, a diffusion-based generation network is employed to generate a variety of details for the missing areas. The predicted structure and details are subsequently converted back into meshes, yielding the complete results. Since the details are depicted in images, our approach eliminates the need for SDFs during the training phase, achieving weakly-supervision. We conduct extensive comparisons on both artificial and real-scan datasets, demonstrating an average improvement of over 38.1% compared to the prior method, and achieving SOTA performance.
Changfeng Ma, Pengxiao Guo, Shuangyu Yang, Yuanqi Li, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.7
2026 GSReuse: Temporally Adaptive Screen-Space Reuse for Accelerating 3D Gaussian Splatting
abstract
Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis. However, rendering each frame independently in a video sequence leads to redundant computations, especially when adjacent frames share significant visual overlaps. This inefficiency is particularly problematic in VR applications, where high frame rates and stereoscopic rendering amplify the per-frame cost. Existing frame interpolation or reuse strategies typically rely on image-domain information and are thus not directly applicable to 3DGS rendering, which is fundamentally point-based. To address this runtime inefficiency, we propose GSReuse, a lightweight and drop-in accelerator that speeds up 3DGS rendering by reusing computations across consecutive frames. GSReuse operates in screen space and introduces only minimal modifications to existing 3DGS rendering pipelines. It also eliminates the need for retraining scene representations. Given the rendered image, depth map, and camera parameters of the current frame, GSReuse estimates reliable Gaussian splatting motion vectors for all pixels and warps reusable contents to the new view. A tile-based filtering and masking strategy is then applied to determine which regions can be safely reused, allowing the 3DGS renderer to skip redundant rendering operations. We evaluate GSReuse on multiple benchmark datasets, showing that GSReuse significantly improves rendering, while maintaining high visual fidelity. Compared to state-of-the-art video frame reuse/generation methods, GSReuse delivers better image quality with much lower latency, facilitating practical deployment of 3DGS in VR applications.
Chengzhi Tao, Jie Guo 0001, Letian Huang, Junqiu Zhu, Daoheng Wang, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.8
2026 Surface Reconstruction From Point Clouds via Image-Free Point-to-Gaussian Inference
abstract
The task of surface reconstruction from point clouds is to produce high-quality meshes using sampled 3D points (no images available). Traditional methods primarily focus on geometric accuracy but often produce meshes without texture colors. In this paper, we present a brand-new perspective in point cloud reconstruction task-Imagining points as more informative Gaussian splats and obtaining colored surfaces through free-form Gaussian-rendering reconstruction. We train a universal Point-to-Gaussian model to infer the attributes of Gaussian splats for any given pointcloud with merely point coordinates (and color optionally) as input, without requiring any image. Significant technical designs are applied on initialization, regularization and loss functions, making the whole learning process stable. The inferred Gaussian splats can faithfully recover the original appearance of objects or scenes (capable of quick rendering from any viewpoint, like human's imagination ability), meanwhile closely adhering to the input shape. After obtaining a sufficient number of virtually rendered images and depth maps, we employ the truncated signed distance function (TSDF) fusion to get the reconstruction results, producing high-quality and colored meshes. Extensive experiments demonstrate that our approach surpasses state-of-the-art methods in surface reconstruction metrics while maintaining high efficiency and simplicity.
Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.5
2025 360-GS: Layout-Guided Panoramic Gaussian Splatting for Indoor Roaming
abstract
3D Gaussian Splatting (3D-GS) has recently attracted great attention with real-time and photo-realistic renderings. This technique typically takes perspective images as input and optimizes a set of 3D elliptical Gaussians by splatting them onto the image planes, resulting in$2 D$Gaussians. However, applying 3D-GS to panoramic inputs presents challenges in effectively modeling the projection onto the spherical surface of 360° images using 2D Gaussians. In practical applications, input panoramas are often sparse, leading to unreliable initialization of 3D Gaussians and subsequent degradation of 3D-GS quality. In addition, due to the under-constrained geometry of texture-less planes (e.g., walls and floors), 3D-GS struggles to model these flat regions with elliptical Gaussians, resulting in significant floaters in novel views. To address these issues, we propose 360-GS, a novel layout-guided 360° Gaussian splatting for a limited set of panoramic inputs. Instead of splatting 3D Gaussians directly onto the spherical surface, 360-GS projects them onto the tangent plane of the unit sphere and then maps them to the spherical projections.
Jiayang Bai, Letian Huang, Jie Guo 0001, Wen Gong, Yuanqi Li, Yanwen Guo 0001
3DV6
2025 Real-Time Neural Denoising with Render-Aware Knowledge Distillation
abstract
Real-time Monte Carlo (MC) ray tracing with low sampling rates demands a denoising algorithm that adeptly balances the trade-off between quality and efficiency. Previous works have paid much attention on designing delicate denoising architecture while ignoring model compression. In this work, we present a render-aware knowledge distillation (RAKD) framework, specifically designed for Monte Carlo denoising. We meticulously delineate the Knowledge Distillation (KD) process within RAKD, emphasizing three pivotal techniques: the strategic incorporation of an auxiliary unlabeled dataset, the integration of adversarial learning through generative adversarial network (GAN), and the application of parameter transfer for robust model initialization. These approaches are harmoniously combined to distill knowledge effectively, enabling our student model to adeptly strike a balance between preserving high-frequency details and reducing low-frequency noise. Finally, our results demonstrate that RAKD achieves state-of-the-art quality while upholding real-time performance, successfully tackling the computational constraints faced by resource-limited devices.
Mengxun Kong, Jie Guo 0001, Chen Wang 0149, Yanwen Guo 0001
AAAI5
2025 High-quality Point Cloud Oriented Normal Estimation via Hybrid Angular and Euclidean Distance Encoding
abstract
The proliferation of Light Detection and Ranging (LiDAR) technology has facilitated the acquisition of three-dimensional point clouds, which are integral to applications in VR, AR, and Digital Twin. Oriented normals, critical for 3D reconstruction and scene analysis, cannot be directly extracted from scenes using LiDAR due to its operational principles. Previous traditional or learning-based methods are prone to inaccuracies due to uneven distribution and noise due to the dependence on local geometry features. This paper addresses the challenge of estimating oriented point normals by introducing a point cloud normal estimation framework via hybrid angular and Euclidean distance encoding (HAE). Our method overcomes the limitations of local geometric information by combining angular and Euclidean spaces to extract features from both point cloud coordinates and light rays, leading to more accurate normal estimation. The core of our network consists of an angular distance encoding module, which leverages both ray directions and point coordinates for unoriented normal refinement, and a ray feature fusion module for normal orientation, that is robust to noise. We also provide a point cloud dataset with ground truth normals, generated a virtual scanner, which reflects real scanning distributions and noise profiles.
Yuanqi Li, Jingcheng Huang, Hongshen Wang, Peiyuan Lv, Jiuming Zheng, Jie Guo 0001, Yanwen Guo 0001
CVPR8
2025 Sparse Point Cloud Patches Rendering via Splitting 2D Gaussians
abstract
Current learning-based methods predict NeRF or 3D Gaussians from point clouds to achieve photo-realistic rendering but still depend on categorical priors, dense point clouds, or additional refinements. Hence, we introduce a novel point cloud rendering method by predicting 2D Gaussians from point clouds. Our method incorporates two identical modules with an entire-patch architecture enabling the network to be generalized to multiple datasets. The module normalizes and initializes the Gaussians utilizing the point cloud information including normals, colors and distances. Then, splitting decoders are employed to refine the initial Gaussians by duplicating them and predicting more accurate results, making our methodology effectively accommodate sparse point clouds as well. Once trained, our approach exhibits direct generalization to point clouds across different categories. The predicted Gaussians are employed directly for rendering without additional refinement on the rendered images, retaining the benefits of 2D Gaussians. We conduct extensive experiments on various datasets, and the results demonstrate the superiority and generalization of our method, which achieves SOTA performance. The code is available at https://github.com/murcherful/GauPCRender.
Changfeng Ma, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001
CVPR5
2025 SGCR: Spherical Gaussians for Efficient 3D Curve Reconstruction
abstract
Neural rendering techniques have made substantial progress in generating photo-realistic 3D scenes. The latest 3D Gaussian Splatting technique has achieved high quality novel view synthesis as well as fast rendering speed. However, 3D Gaussians lack proficiency in defining accurate 3D geometric structures despite their explicit primitive representations. This is due to the fact that Gaussian’s attributes are primarily tailored and fine-tuned for rendering diverse 2D images by their anisotropic nature. To pave the way for efficient 3D reconstruction, we present Spherical Gaussians, a simple and effective representation for 3D geometric boundaries, from which we can directly reconstruct 3D feature curves from a set of calibrated multi-view images. Spherical Gaussians is optimized from grid initialization with a view-based rendering loss, where a 2D edge map is rendered at a specific view and then compared to the ground-truth edge map extracted from the corresponding image, without the need for any 3D guidance or supervision. Given Spherical Gaussians serve as intermedia for the robust edge representation, we further introduce a novel optimization-based algorithm called SGCR to directly extract accurate parametric curves from aligned Spherical Gaussians. We demonstrate that SGCR outperforms existing state-of-the-art methods in 3D edge reconstruction while enjoying great efficiency. Code is available at https://github.com/Martinyxr/SGCR.
Xinran Yang, Donghao Ji, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Junyuan Xie
CVPR5
2025 EdgeMovingNet: Edge-preserving Point Cloud Reconstruction via Joint Geometry Features
abstract
Point cloud reconstruction is a critical process in 3D representation and reverse engineering. When it comes to CAD models, edges are significant features that play a crucial role in characterizing the geometry of 3D shapes. However, few points are exactly sampled on edges during acquisition, resulting in apparent artifacts for the reconstruction task. Upsampling point cloud is a direct technical route, but there is a main challenge that the upsampled points may not align with the model edge accurately. To overcome this, we develop an integrated framework to estimate edges by joint regression of three geometry features—point-to-edge direction, point-to-edge distance and point normal. Benefiting these features, we implement a novel refinement process to move and produce more points which lie accurately on edges of the model, allowing for high-quality edge-preserving reconstruction. Experiments and comparisons against previous methods demonstrate our method’s effectiveness and superiority.
Xinran Yang, Donghao Ji, Yuanqi Li, Junyuan Xie, Jie Guo 0001, Yanwen Guo 0001
CVPR6
2025 GaRe: Relightable 3D Gaussian Splatting for Outdoor Scenes from Unconstrained Photo Collections
Haiyang Bai, Songru Jiang, Tao Lu 0005, Yuanqi Li, Jie Guo 0001, Runze Fu, Yanwen Guo 0001, Lijun Chen 0006
ICCV9
2025 Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have significantly improved 2D visual understanding, prompting interest in their application to complex 3D reasoning tasks. However, it remains unclear whether these models can effectively capture the detailed spatial information required for robust real-world performance, especially cross-view consistency, a key requirement for accurate 3D reasoning. Considering this issue, we introduce Viewpoint Learning, a task designed to evaluate and improve the spatial reasoning capabilities of MLLMs. We present the Viewpoint-100K dataset, consisting of 100K object-centric image pairs with diverse viewpoints and corresponding question-answer pairs. Our approach employs a two-stage fine-tuning strategy: first, foundational knowledge is injected to the baseline MLLM via Supervised Fine-Tuning (SFT) on Viewpoint-100K, resulting in significant improvements across multiple tasks; second, generalization is enhanced through Reinforcement Learning using the Group Relative Policy Optimization (GRPO) algorithm on a broader set of questions. Additionally, we introduce a hybrid cold-start initialization method designed to simultaneously learn viewpoint representations and maintain coherent reasoning thinking. Experimental results show that our approach significantly activates the spatial reasoning ability of MLLM, improving performance on both in-domain and out-of-domain reasoning tasks. Our findings highlight the value of developing foundational spatial skills in MLLMs, supporting future progress in robotics, autonomous systems, and 3D scene understanding.
Xiaoyu Zhan, Wenxuan Huang 0001, Xinyu Fu 0009, Changfeng Ma, Shaosheng Cao, Bohan Jia, Shaohui Lin, Zhenfei Yin, Lei Bai 0001, Wanli Ouyang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
NeurIPS14
2025 Spectral-GS: Taming 3D Gaussian Splatting with Spectral Entropy
abstract
Recently, 3D Gaussian Splatting (3DGS) has achieved impressive results in novel view synthesis, demonstrating high fidelity and efficiency. However, it easily exhibits needle-like artifacts, especially when increasing the sampling rate. Mip-Splatting tries to remove these artifacts with a 3D smoothing filter for frequency constraints and a 2D Mip filter for approximated supersampling. Unfortunately, it tends to produce over-blurred results, and sometimes needle-like Gaussians still persist. Our spectral analysis of the covariance matrix during optimization and densification reveals that current 3DGS lacks shape awareness, relying instead on spectral radius and view positional gradients to determine splitting. As a result, needle-like Gaussians with small positional gradients and low spectral entropy fail to split and overfit high-frequency details. Furthermore, both the filters used in 3DGS and Mip-Splatting reduce the spectral entropy and increase the condition number during zooming in to synthesize novel view, causing view inconsistencies and more pronounced artifacts. Our Spectral-GS, based on spectral analysis, introduces 3D shape-aware splitting and 2D view-consistent filtering strategies, effectively addressing these issues, enhancing 3DGS’s capability to represent high-frequency details without noticeable artifacts, and achieving high-quality realistic rendering.
Letian Huang, Jie Guo 0001, Jialin Dan, Ruoyu Fu, Yuanqi Li, Yanwen Guo 0001
SIGGRAPH Asia6
2025 ANIR: Adaptive Neural Implicit Representation for 3D shape reconstruction and generation
Kun Liu 0021, Yan Zhang 0057, Yanwen Guo 0001, Jie Guo 0001
Comput. Aided Des.3
2025 Detail-Preserving Real-Time Hair Strand Linking and Filtering
abstract
Abstract Realistic hair rendering remains a significant challenge in computer graphics due to the intricate microstructure of hair fibers and their anisotropic scattering properties, which make them highly sensitive to noise. Although recent advancements in image‐space and 3D‐space denoising and antialiasing techniques have facilitated real‐time rendering in simple scenes, existing methods still struggle with excessive blurring and artifacts, particularly in fine hair details such as flyaway strands. These issues arise because current techniques often fail to preserve sub‐pixel continuity and lack directional sensitivity in the filtering process. To address these limitations, we introduce a novel real‐time hair filtering technique that effectively reconstructs fine fiber details while suppressing noise. Our method improves visual quality by maintaining strand‐level details and ensuring computational efficiency, making it well‐suited for real‐time applications in video games and virtual reality (VR) and augmented reality (AR) environments.
Tao Huang 0026, J. Yuan, Ruike Hu, Lu Wang 0007, Yanwen Guo 0001, Bin Chen 0019, Jie Guo 0001, Junqiu Zhu
Comput. Graph. Forum5
2025 Rethinking mixture of rain removal via depth-guided adversarial learning
Yongzhen Wang 0001, Xuefeng Yan 0001, Yanbiao Niu, Lina Gong, Yanwen Guo 0001, Mingqiang Wei
Neural Networks5
2025 Realistic Simulation of Underwater Scene for Image Enhancement
abstract
In recent years, learning-based methods have performed remarkably well in underwater image enhancement, but their performance is limited by the lack of high-quality, diverse training datasets. Current underwater image datasets are unable to address the following three issues: intra-domain gaps in underwater environments, inter-domain gaps between synthetic and real data, and domain inaccuracies. To overcome these limitations, we construct a realistic underwater scene using 3D graphics engine through a three-step approach: 1) integrate a simulation-specific underwater light propagation models to create volumetric fog; 2) employ physical model-based rendering for accurate light field simulation; 3) configure scenes with parameters extracted from real underwater images. Based on this framework, we develop an underwater image enhancement dataset (MUSE). Experiments demonstrate that models trained on MUSE outperform those trained on conventional datasets, highlighting the effectiveness of our approach.
Tingyu Liu, Qunyan Jiang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Zhonghua Ni
IEEE Trans. Geosci. Remote. Sens.7
2025 Bernstein Bounds for Caustics
abstract
Systematically simulating specular light transport requires an exhaustive search for triangle tuples containing admissible paths. Given the extreme inefficiency of enumerating all combinations, we significantly reduce the search domain by stochastically sampling such tuples. The challenge is to design proper sampling probabilities that keep the noise level controllable. Our key insight is that by bounding the irradiance contributed by each triangle tuple at a given position, we can sample a subset of triangle tuples with potentially high contributions. Although low-contribution tuples are assigned a negligible probability, the overall variance remains low. Therefore, we derive position and irradiance bounds for caustics casted by each triangle tuple, introducing a bounding property of rational functions on a Bernstein basis. When formulating position and irradiance expressions into rational functions, we handle non-rational parts through remainder variables to maintain bounding validity. Finally, we carefully design the sampling probabilities by optimizing the upper bound of the variance, expressed only using the position and irradiance bounds. The bound-driven sampling of triangle tuples is intrinsically unbiased even without defensive sampling. It can be combined with various unbiased and biased root-finding techniques within a local triangle domain. Extensive evaluations show that our method enables the fast and reliable rendering of complex caustics effects. Yet, our method is efficient for no more than two specular vertices, where complexity grows sublinearly to the number of triangles and linearly to that of emitters, and does not consider the Fresnel and visibility terms. We also rely on parameters to control subdivisions.
Zhimin Fan 0001, Chen Wang 0149, Boxuan Li, Lingqi Yan 0001, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.7
2025 Multiple Importance Reweighting for Path Guiding
abstract
Contemporary path guiding employs an iterative training scheme to fit radiance distributions. However, existing methods combine the estimates generated in each iteration merely within image space, overlooking differences in the convergence of distribution fitting over individual light paths. This paper formulates the estimation combination task as a path reweighting process. To compute spatio-directional varying combination weights, we propose multiple importance reweighting , leveraging the importance distributions from multiple guiding iterations. We demonstrate that our proposed path-level reweighting makes guiding algorithms less sensitive to noise and overfitting in distributions. This facilitates a finer subdivision of samples both spatially and temporally (i.e., over iterations), which leads to additional improvements in the accuracy of distributions and samples. Inspired by adaptive multiple importance sampling (AMIS), we introduce a simple yet effective mixture-based weighting scheme with theoretically guaranteed consistency, demonstrating good practical performance compared to alternative weighting schemes. To further foster usage with high sample rates, we introduce a hyperparameter that controls the size of sample storage. When this size limit is exceeded, low-valued samples are splatted during rendering and reweighted using a partial mixture of distributions. We found limiting the storage size reduces memory overhead and keeps variance reduction and bias comparable to the unlimited ones. Our method is largely agnostic to the underlying guiding method and compatible with conventional pixel reweighting techniques. Extensive evaluations underscore the feasibility of our approach in various scenes, achieving variance reduction with negligible bias over state-of-the-art solutions within equal sample rates and rendering time.
Zhimin Fan 0001, Lingqi Yan 0001, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.5
2025 TransparentGS: Fast Inverse Rendering of Transparent Objects with Gaussians
abstract
The emergence of neural and Gaussian-based radiance field methods has led to considerable advancements in novel view synthesis and 3D object reconstruction. Nonetheless, specular reflection and refraction continue to pose significant challenges due to the instability and incorrect overfitting of radiance fields to high-frequency light variations. Currently, even 3D Gaussian Splatting (3D-GS), as a powerful and efficient tool, falls short in recovering transparent objects with nearby contents due to the existence of apparent secondary ray effects. To address this issue, we propose TransparentGS, a fast inverse rendering pipeline for transparent objects based on 3D-GS. The main contributions are three-fold. Firstly, an efficient representation of transparent objects, transparent Gaussian primitives, is designed to enable specular refraction through a deferred refraction strategy. Secondly, we leverage Gaussian light field probes (GaussProbe) to encode both ambient light and nearby contents in a unified framework. Thirdly, a depth-based iterative probes query (IterQuery) algorithm is proposed to reduce the parallax errors in our probe-based framework. Experiments demonstrate the speed and accuracy of our approach in recovering transparent objects from complex environments, as well as several applications in computer graphics and vision.
Letian Huang, Dongwei Ye, Jialin Dan, Chengzhi Tao, Kun Zhou 0001, Bo Ren 0003, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.9
2025 DSCombiner: Double Shrinkage for Combining Biased and Unbiased Monte Carlo Renderings
abstract
Monte Carlo rendering often faces a dilemma, namely, whether to choose an unbiased estimator or a biased one. Although different integrators have been developed to address various scenarios, no single method can effectively manage all situations. Thus, finding a good approach to combine different integrators has always been a topic that warrants exploration. This work proposes DSCombiner, a new shrinkage estimator that flexibly combines unbiased and biased estimators (typically generated by different integrators) in image space into a single estimating procedure, strategically utilizing the strengths of different integrators while minimizing their weaknesses. DSCombiner overcomes the limitation of single shrinkage combiners by introducing a two-step shrinkage towards a noise-free radiance prior. We derive optimal shrinkage factors for the two steps within a hierarchical Bayesian framework, and provide a deep learning-based method to improve the results. Comprehensive qualitative and quantitative validations across diverse scenes demonstrate visible improvements in image quality, as compared with previous image-space and path-space combiners.
Keheng Xu, Mufan Guo, Xianhao Yu, Zhimin Fan 0001, Guihuan Feng, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.7
2025 MixRF: Universal Mixed Radiance Fields With Points and Rays Aggregation
abstract
Recent advancements in neural rendering methods, such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3D-GS), have significantly revolutionized photo-realistic novel view synthesis of scenes with multiple photos or videos as input. However, existing approaches within the NeRF and 3D-GS frameworks often assume the independence of point sampling and ray casting, which are intrinsic to volume rendering and alpha-blending techniques. These underlying assumptions limit the ability to aggregate context within subspaces, such as densities and colors in the radiance fields and pixels on the image plane, leading to synthesized images that lack fine details and smoothness. To overcome this, we propose a universal framework, MixRF, comprising a Radiance Field Mixer (RF-mixer) and a Color Domain Mixer (CD-mixer), to sufficiently aggregate and fully explore information in neighboring sampled points and casting rays, separately. The RF-mixer treats sampled points as an explicit point cloud, enabling the aggregation of density and color attributes from neighboring points to better capture local geometry and appearance. Meanwhile, the CD-mixer rearranges rendered pixels on the sub-image plane, improving smoothness and recovering fine details and textures. Both mixers employ a kernel-based mixing strategy to facilitate effective and controllable attribute aggregation, ensuring a more comprehensive exploration of radiance values and pixel information. Extensive experiments demonstrate that our MixRF framework is compatible with radiance field-based methods, including NeRF and 3D-GS designs. The proposed framework dramatically enhances performance in both qualitative and quantitative evaluations, with less than a $ 25\%$25% increase in computational overhead during inference.
Haiyang Bai, Tao Lu 0005, Chang Gou, Jie Guo 0001, Lijun Chen 0006, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.8
2025 GlossyGS: Inverse Rendering of Glossy Objects With 3D Gaussian Splatting
abstract
Reconstructing objects from posed images is a crucial and complex task in computer graphics and computer vision. While NeRF-based neural reconstruction methods have exhibited impressive reconstruction ability, they tend to be time-comsuming. Recent strategies have adopted 3D Gaussian Splatting (3D-GS) for inverse rendering, which have led to quick and effective outcomes. However, these techniques generally have difficulty in producing believable geometries and materials for glossy objects, a challenge that stems from the inherent ambiguities of inverse rendering. To address this, we introduce GlossyGS, an innovative 3D-GS-based inverse rendering framework that aims to precisely reconstruct the geometry and materials of glossy objects by integrating material priors. The key idea is the use of micro-facet geometry segmentation prior, which helps to reduce the intrinsic ambiguities and improve the decomposition of geometries and materials. Additionally, we introduce a normal map prefiltering strategy to more accurately simulate the normal distribution of reflective surfaces. These strategies are integrated into a hybrid geometry and material representation that employs both explicit and implicit methods to depict glossy objects. We demonstrate through quantitative analysis and qualitative visualization that the proposed method is effective to reconstruct high-fidelity geometries and materials of glossy objects, and performs favorably against State-of-the-Arts.
Shuichang Lai, Letian Huang, Jie Guo 0001, Bowen Pan, Xiaoxiao Long, Jiangjing Lyu, Chengfei Lv, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.9
2025 Deep Point Cloud Edge Reconstruction via Surface Patch Segmentation
abstract
Parametric edge reconstruction for point cloud data is a fundamental problem in computer graphics. Existing methods first classify points as either edge points (including corners) or non-edge points, and then fit parametric edges to the edge points. However, few points are exactly sampled on edges in practical scenarios, leading to significant fitting errors in the reconstructed edges. Prominent deep learning-based methods also primarily emphasize edge points, overlooking the potential of non-edge areas. Given that sparse and non-uniform edge points cannot provide adequate information, we address this challenge by leveraging neighboring segmented patches to supply additional cues. We introduce a novel two-stage framework that reconstructs edges precisely and completely via surface patch segmentation. First, we propose PCER-Net, a Point Cloud Edge Reconstruction Network that segments surface patches, detects edge points, and predicts normals simultaneously. Second, a joint optimization module is designed to reconstruct a complete and precise 3D wireframe by fully utilizing the predicted results of the network. Concretely, the segmented patches enable accurate fitting of parametric edges, even when sparse points are not precisely distributed along the model's edges. Corners can also be naturally detected from the segmented patches. Benefiting from fitted edges and detected corners, a complete and precise 3D wireframe model with topology connections can be reconstructed by geometric optimization. Finally, we present a versatile patch-edge dataset, including CAD and everyday models (furniture), to generalize our method. Extensive experiments and comparisons against previous methods demonstrate our effectiveness and superiority. We will release the code and dataset to facilitate future research.
Yuanqi Li, Hongshen Wang, Jingcheng Huang, Jianwei Guo 0003, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.9
2025 Parameterize Structure With Differentiable Template for 3D Shape Generation
abstract
Structural representation is crucial for reconstructing and generating editable 3D shapes with part semantics. Recent 3D shape generation works employ complicated networks and structure definitions relying on hierarchical annotations and pay less attention to the details inside parts. In this paper, we propose the method that parameterizes the shared structure in the same category using a differentiable template and corresponding fixed-length parameters. Specific parameters are fed into the template to calculate cuboids that indicate a concrete shape. We utilize the boundaries of three-view renderings of each cuboid to further describe the inside details. Shapes are represented with the parameters and three-view details inside cuboids, from which the SDF can be calculated to recover the object. Benefiting from our fixed-length parameters and three-view details, our networks for reconstruction and generation are simple and effective to learn the latent space. Our method can reconstruct or generate diverse shapes with complicated details, and interpolate them smoothly. Extensive evaluations demonstrate the superiority of our method on reconstruction from point cloud, generation, and interpolation.
Changfeng Ma, Pengxiao Guo, Shuangyu Yang, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.7
2025 Make PBR Materials Tileable With Latent Diffusion Inpainting
abstract
Physically-based-rendering (PBR) materials are crucial in modern rendering pipelines, and many studies have focused on acquiring these materials from reality or images. However, existing methods may result in non-tileable results, since the realistic inputs usually have seams. Compared to non-tileable materials, tileable PBR materials have more universal application scenarios. To address this issue, we introduce MaTi, a novel pipeline that converts non-tileable PBR materials into tileable ones with minimal distortion. MaTi rearranges material patches to align boundaries at the center of the image, and then uses a diffusion model to inpaint the seams. We use scaled gamma correction to reduce the occurrence of collapse when processing special material maps. The color correction and triangular blending are adopt to preserve the original material information. Additionally, we design a division and blending strategy to efficiently handle high resolution materials. Our experiments demonstrate that MaTi can seamlessly modify PBR materials while preserving the original information, outperforming existing synthesis methods.
Xiaoyu Zhan, Jianxin Yang, Jun Wang 0039, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.6
2024 FASSET: Frame Supersampling and Extrapolation Using Implicit Neural Representations of Rendering Contents
Haoyu Qin, Jie Guo 0001, Wenyang Bai, Yanwen Guo 0001
CVM (1)6
2024 Practical Measurements of Translucent Materials with Inter-Pixel Translucency Prior
abstract
Material appearance is a key component of photorealism, with a pronounced impact on human perception. Although there are many prior works targeting at measuring opaque materials using light-weight setups (e.g., consumer-level cameras), little attention is paid on acquiring the optical properties of translucent materials which are also quite common in nature. In this paper, we present a practical method for acquiring scattering properties of translucent materials, based solely on ordinary images captured with unknown lighting and camera parameters. The key to our method is an inter-pixel translucency prior which states that image pixels of a given homogeneous translucent material typically form curves (dubbed translucent curves) in the RGB space, of which the shapes are determined by the parameters of the material. We leverage this prior in a specially-designed convolutional neural network comprising multiple encoders, a translucency-aware feature fusion module and a cascaded decoder. We demonstrate, through both visual comparisons and quantitative evaluations, that high accuracy can be achieved on a wide range of real-world translucent materials.
Zhenyu Chen 0001, Jie Guo 0001, Shuichang Lai, Ruoyu Fu, Mengxun Kong, Chen Wang 0149, Hongyu Sun 0001, Zhebin Zhang, Chen Li 0062, Yanwen Guo 0001
CVPR10
2024 LiDAR-Net: A Real-Scanned 3D Point Cloud Dataset for Indoor Scenes
abstract
In this paper, we present LiDAR-Net, a new real-scanned indoor point cloud dataset, containing nearly 3.6 billion precisely point-level annotated points, covering an expansive area of 30,000m2. It encompasses three prevalent daily environments, including learning scenes, working scenes, and living scenes. LiDAR-Net is characterized by its non-uniform point distribution, e.g., scanning holes and scanning lines. Additionally, it meticulously records and an-notates scanning anomalies, including reflection noise and ghost. These anomalies stem from specular reflections on glass or metal, as well as distortions due to moving persons. LiDAR-Net's realistic representation of non-uniform distribution and anomalies significantly enhances the training of deep learning models, leading to improved generalization in practical applications. We thoroughly evaluate the performance of state-of-the-art algorithms on LiDAR-Net and provide a detailed analysis of the results. Crucially, our research identifies several fundamental challenges in understanding indoor point clouds, contributing essential insights to future explorations in this field. Our dataset can be found online: http://lidar-net.njumeta.com.
Yanwen Guo 0001, Yuanqi Li, Dayong Ren, Xiaohong Zhang 0009, Liang Pu, Changfeng Ma, Xiaoyu Zhan, Jie Guo 0001, Mingqiang Wei, Yan Zhang 0057, Piaopiao Yu, Shuangyu Yang, Donghao Ji, Huisheng Ye
CVPR1
2024 FINER: Flexible Spectral-Bias Tuning in Implicit NEural Representation by Variableperiodic Activation Functions
abstract
Implicit Neural Representation (INR), which utilizes a neural network to map coordinate inputs to corresponding attributes, is causing a revolution in the field of signal processing. However, current INR techniques suffer from a re-stricted capability to tune their supported frequency set, re-sulting in imperfect performance when representing complex signals with multiple frequencies. We have identified that this frequency-related problem can be greatly alleviated by introducing variableperiodic activation functions, for which we propose FINER. By initializing the bias of the neural network within different ranges, sub-functions with various frequencies in the variableperiodic function are selected for activation. Consequently, the supported frequency set of FINER can be flexibly tuned, leading to improved performance in signal representation. We demon-strate the capabilities of FINER in the contexts of2D image fitting, 3D signed distance field representation, and 5D neural radiance fields optimization, and we show that it outper-forms existing INRs.
Zhen Liu 0031, Hao Zhu 0005, Qi Zhang 0029, Jingde Fu, Weibing Deng, Zhan Ma 0001, Yanwen Guo 0001, Xun Cao
CVPR7
2024 Semantic Human Mesh Reconstruction with Textures
abstract
The field of 3D detailed human mesh reconstruction has made significant progress in recent years. However, current methods still face challenges when used in industrial applications due to unstable results, low-quality meshes, and a lack of UV unwrapping and skinning weights. In this paper, we present SHERT, a novel pipeline that can reconstruct semantic human meshes with textures and high-precision details. SHERT applies semantic- and normal-based sampling between the detailed surface (e.g. mesh and SDF) and the corresponding SMPL-X model to obtain a partially sampled semantic mesh and then generates the complete semantic mesh by our specifically designed self-supervised completion and refinement networks. Using the complete semantic mesh as a basis, we employ a texture diffusion model to create human textures that are driven by both images and texts. Our reconstructed meshes have stable UV unwrapping, high-quality triangle meshes, and consistent semantic information. The given SMPL-X model provides semantic information and shape priors, allowing SHERT to perform well even with incorrect and incomplete inputs. The semantic information also makes it easy to substitute and animate different body parts such as the face, body, and hands. Quantitative and qualitative experiments demonstrate that SHERT is capable of producing high-fidelity and robust semantic meshes that outperform state-of-the-art methods.
Xiaoyu Zhan, Jianxin Yang, Yuanqi Li, Jie Guo 0001, Yanwen Guo 0001, Wenping Wang 0001
CVPR5
2024 Prompt3D: Random Prompt Assisted Weakly-Supervised 3D Object Detection
abstract
The prohibitive cost of annotations for fully supervised 3D indoor object detection limits its practicality. In this work, we propose Random Prompt Assisted Weakly-supervised 3D Object Detection, termed as Prompt3D, a weakly-supervised approach that leverages position-level labels to overcome this challenge. Explicitly, our method focuses on enhancing labeling using synthetic scenes crafted from 3D shapes generated via random prompts. First, a Synthetic Scene Generation (SSG) module is introduced to assemble synthetic scenes with a curated collection of 3D shapes, created via random prompts for each category. These scenes are enriched with automatically generated point-level annotations, providing a robust supervisory frame-work for training the detection algorithm. To enhance the transfer of knowledge from virtual to real datasets, we then introduce a Prototypical Proposal Feature Alignment (PPFA) module. This module effectively alleviates the domain gap by directly minimizing the distance between feature prototypes of the same class proposals across two domains. Compared with sota BR, our method improves by 5.4% and 8.7% on mAP with VoteNet and GroupFree3D serving as detectors respectively, demonstrating the effectiveness of our proposed method. Code is available at: https://github.com/huishengye/prompt3d.
Xiaohong Zhang 0009, Huisheng Ye, Qinyu Tang, Yuanqi Li, Yanwen Guo 0001, Jie Guo 0001
CVPR6
2024 On the Error Analysis of 3D Gaussian Splatting and an Optimal Projection Strategy
Letian Huang, Jiayang Bai, Jie Guo 0001, Yuanqi Li, Yanwen Guo 0001
ECCV (17)5
2024 GLPanoDepth: Global-to-Local Panoramic Depth Estimation
abstract
Depth estimation is a fundamental task in many vision applications. With the popularity of omnidirectional cameras, it becomes a new trend to tackle this problem in the spherical space. In this paper, we propose a learning-based method for predicting dense depth values of a scene from a monocular omnidirectional image. An omnidirectional image has a full field-of-view, providing much more complete descriptions of the scene than perspective images. However, fully-convolutional networks that most current solutions rely on fail to capture rich global contexts from the panorama. To address this issue and also the distortion of equirectangular projection in the panorama, we propose Cubemap Vision Transformers (CViT), a new transformer-based architecture that can model long-range dependencies and extract distortion-free global features from the panorama. We show that cubemap vision transformers have a global receptive field at every stage and can provide globally coherent predictions for spherical signals. As a general architecture, it removes any restriction that has been imposed on the panorama in many other monocular panoramic depth estimation methods. To preserve important local features, we further design a convolution-based branch in our pipeline (dubbed GLPanoDepth) and fuse global features from cubemap vision transformers at multiple scales. This global-to-local strategy allows us to fully exploit useful global and local features in the panorama, achieving state-of-the-art performance in panoramic depth estimation.
Jiayang Bai, Haoyu Qin, Shuichang Lai, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Image Process.5
2024 Specular Polynomials
abstract
Finding valid light paths that involve specular vertices in Monte Carlo rendering requires solving many non-linear, transcendental equations in high-dimensional space. Existing approaches heavily rely on Newton iterations in path space, which are limited to obtaining at most a single solution each time and easily diverge when initialized with improper seeds. We propose specular polynomials , a Newton iteration-free methodology for finding a complete set of admissible specular paths connecting two arbitrary endpoints in a scene. The core is a reformulation of specular constraints into polynomial systems, which makes it possible to reduce the task to a univariate root-finding problem. We first derive bivariate systems utilizing rational coordinate mapping between the coordinates of consecutive vertices. Subsequently, we adopt the hidden variable resultant method for variable elimination, converting the problem into finding zeros of the determinant of univariate matrix polynomials. This can be effectively solved through Laplacian expansion for one bounce and a bisection solver for more bounces. Our solution is generic, completely deterministic, accurate for the case of one bounce, and GPU-friendly. We develop efficient CPU and GPU implementations and apply them to challenging glints and caustic rendering. Experiments on various scenarios demonstrate the superiority of specular polynomial-based solutions compared to Newton iteration-based counterparts. Our implementation is available at https://github.com/mollnn/spoly.
Zhimin Fan 0001, Jie Guo 0001, Zhenyu Chen 0001, Pengpei Hong, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.9
2024 Conditional Mixture Path Guiding for Differentiable Rendering
abstract
The efficiency of inverse optimization in physically based differentiable rendering heavily depends on the variance of Monte Carlo estimation. Despite recent advancements emphasizing the necessity of tailored differential sampling strategies, the general approaches remain unexplored. In this paper, we investigate the interplay between local sampling decisions and the estimation of light path derivatives. Considering that modern differentiable rendering algorithms share the same path for estimating differential radiance and ordinary radiance, we demonstrate that conventional guiding approaches, conditioned solely on the last vertex, cannot attain this density. Instead, a mixture of different sampling distributions is required, where the weights are conditioned on all the previously sampled vertices in the path. To embody our theory, we implement a conditional mixture path guiding that explicitly computes optimal weights on the fly. Furthermore, we show how to perform positivization to eliminate sign variance and extend to scenes with millions of parameters. To the best of our knowledge, this is the first generic framework for applying path guiding to differentiable rendering. Extensive experiments demonstrate that our method achieves nearly one order of magnitude improvements over state-of-the-art methods in terms of variance reduction in gradient estimation and errors of inverse optimization. The implementation of our proposed method is available at https://github.com/mollnn/conditional-mixture.
Zhimin Fan 0001, Mufan Guo, Ruoyu Fu, Yanwen Guo 0001, Jie Guo 0001
ACM Trans. Graph.5
2024 Collaborative Completion and Segmentation for Partial Point Clouds With Outliers
abstract
Outliers will inevitably creep into the captured point cloud during 3D scanning, degrading cutting-edge models on various geometric tasks heavily. This paper looks at an intriguing question that whether point cloud completion and segmentation can promote each other to defeat outliers. To answer it, we propose a collaborative completion and segmentation network, termed CS-Net, for partial point clouds with outliers. Unlike most of existing methods, CS-Net does not need any clean (or say outlier-free) point cloud as input or any outlier removal operation. CS-Net is a new learning paradigm that makes completion and segmentation networks work collaboratively. With a cascaded architecture, our method refines the prediction progressively. Specifically, after the segmentation network, a cleaner point cloud is fed into the completion network. We design a novel completion network which harnesses the labels obtained by segmentation together with farthest point sampling to purify the point cloud and leverages KNN-grouping for better generation. Benefited from segmentation, the completion module can utilize the filtered point cloud which is cleaner for completion. Meanwhile, the segmentation module is able to distinguish outliers from target objects more accurately with the help of the clean and complete shape inferred by completion. Besides the designed collaborative mechanism of CS-Net, we establish a benchmark dataset of partial point clouds with outliers. Extensive experiments show clear improvements of our CS-Net over its competitors, in terms of outlier robustness and completion accuracy.
Changfeng Ma, Yang Yang 0092, Jie Guo 0001, Mingqiang Wei, Chong-Jun Wang, Yanwen Guo 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.6
2024 GeoSegNet: point cloud semantic segmentation via geometric encoder-decoder modeling
Chen Chen 0161, Yisen Wang 0003, Honghua Chen, Xuefeng Yan 0001, Dayong Ren, Yanwen Guo 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei
Vis. Comput.6
2024 MFFNet: multimodal feature fusion network for point cloud semantic segmentation
Dayong Ren, Zhengyi Wu, Jie Guo 0001, Mingqiang Wei, Yanwen Guo 0001
Vis. Comput.6
2023 Symmetric Shape-Preserving Autoencoder for Unsupervised Real Scene Point Cloud Completion
abstract
Unsupervised completion of real scene objects is of vital importance but still remains extremely challenging in preserving input shapes, predicting accurate results, and adapting to multi-category data. To solve these problems, we propose in this paper an Unsupervised Symmetric Shape-Preserving Autoencoding Network, termed USSPA, to predict complete point clouds of objects from real scenes. One of our main observations is that many natural and manmade objects exhibit significant symmetries. To accommodate this, we devise a symmetry learning module to learn from those objects and to preserve structural symmetries. Starting from an initial coarse predictor, our autoencoder refines the complete shape with a carefully designed upsampling refinement module. Besides the discriminative process on the latent space, the discriminators of our USSPA also take predicted point clouds as direct guidance, enabling more detailed shape prediction. Clearly different from previous methods which train each category separately, our USSPA can be adapted to the training of multi-category data in one pass through a classifier-guided discriminator, with consistent performance on single category. For more accurate evaluation, we contribute to the community a real scene dataset with paired CAD models as ground truth. Extensive experiments and comparisons demonstrate our superiority and generalization and show that our method achieves state-of-the-art performance on unsupervised completion of real scene objects.
Changfeng Ma, Pengxiao Guo, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001
CVPR6
2023 UHDNeRF: Ultra-High-Definition Neural Radiance Fields
abstract
We propose UHDNeRF, a new framework for novel view synthesis on the challenging ultra-high-resolution (e.g., 4K) real-world scenes. Previous NeRF methods are not specifically designed for rendering on extremely high resolutions, leading to burry results with notable detail-losing problems even though trained on 4K images. This is mainly due to the mismatch between the high-resolution inputs and the low-dimensional volumetric representation. To address this issue, we introduce an adaptive implicit-explicit scene representation with which an explicit sparse point cloud is used to boost the performance of an implicit volume on modeling subtle details. Specifically, we reconstruct the complex real-world scene with a frequency separation strategy that the implicit volume learns to represent the low-frequency properties of the whole scene, and the sparse point cloud is used for reproducing high-frequency details. To better explore the information embedded in the point cloud, we extract a global structure feature and a local point-wise feature from the point cloud for each sample located in the high-frequency regions. Furthermore, a patch-based sampling strategy is introduced to reduce the computational cost. The high-fidelity rendering results demonstrate the superiority of our method for retaining high-frequency details at 4K ultra-high-resolution scenarios against state-of-the-art NeRF-based solutions.
Quewei Li, Feichao Li, Jie Guo 0001, Yanwen Guo 0001
ICCV4
2023 Convolutional Self-attention Guided Graph Neural Network for Few-Shot Action Recognition
Jie Guo 0001, Yanwen Guo 0001
ICIC (2)3
2023 Few-Shot Action Recognition with A Transductive Maximum Margin Classifier
abstract
Few-shot action recognition aims to train a classifier that can generalize well when just a small number of labeled videos per class are given. We introduce a transductive maximum margin classifier for few-shot action recognition, which leverages the unlabeled query videos to improve the recognition performance in the test task. The basic idea of the classical maximum margin classifier is to search for a classifier with the largest geometric margin so that training data can be correctly classified. Due to the insufficient number of labeled videos in the support set, it is challenging to find such a classifier with good generalization ability. We observe that exploring the geometric relationship between the separating hyperplane of the classifier and the feature vectors of the query videos can bring improvements to the classifier. In order to improve data utilization efficiency in the few-shot setting, the class prototypes are also treated as examples, which participate in the iterative training process of the model. Experimental results on two action recognition datasets including Kinetics and Something-Something V2 show that our method achieves state-of-the-art performance.
Jie Guo 0001, Yanwen Guo 0001
IJCNN3
2023 SVBRDF Reconstruction by Transferring Lighting Knowledge
abstract
Abstract The problem of reconstructing spatially‐varying BRDFs from RGB images has been studied for decades. Researchers found themselves in a dilemma: opting for either higher quality with the inconvenience of camera and light calibration, or greater convenience at the expense of compromised quality without complex setups. We address this challenge by introducing a two‐branch network to learn the lighting effects in images. The two branches, referred to as Light‐known and Light‐aware, diverge in their need for light information. The Light‐aware branch is guided by the Light‐known branch to acquire the knowledge of discerning light effects and surface reflectance properties, but without the reliance of light positions. Both branches are trained using the synthetic dataset, but during testing on real‐world cases without calibration, only the Light‐aware branch is activated. To facilitate a more effective utilization of various light conditions, we employ gated recurrent units (GRUs) to fuse the features extracted from different images. The two modules mutually benefit when multiple inputs are provided. We present our reconstructed results on both synthetic and real‐world examples, demonstrating high quality while maintaining a lightweight characteristic in comparison to previous methods.
Pengfei Zhu 0001, Shuichang Lai, Mufan Chen, Jie Guo 0001, Yanwen Guo 0001
Comput. Graph. Forum6
2023 Deep graph learning for spatially-varying indoor lighting prediction
Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Piaopiao Yu, Yan Zhang 0057, Yanwen Guo 0001
Sci. China Inf. Sci.9
2023 Reflectance edge guided networks for detail-preserving intrinsic image decomposition
Quewei Li, Jie Guo 0001, Zhengyi Wu, Yanwen Guo 0001
Sci. China Inf. Sci.5
2023 Elastic temporal alignment for few-shot action recognition
abstract
Abstract Few‐shot action recognition aims to learn a classification model with good generalisation ability when trained with only a few labelled videos. However, it is difficult to learn discriminative feature representations for videos in such a setting. The Elastic Temporal Alignment (ETA) for few‐shot action recognition is proposed. First, a convolutional neural network is employed to extract feature representations of video frames sparsely sampled from videos. In order to obtain the similarity of two videos, a temporal alignment estimation function is utilised to estimate the matching score between each pair of frames from the two videos through an elastic alignment mechanism. The analysis shows that when we judge whether two frames from respective videos are matched, multiple adjacent frames in the videos should be considered, so as to embody the temporal information. Thus, before feeding per‐frame feature vectors of videos into the temporal alignment estimation function, a temporal message passing function is leveraged to propagate the information of per‐frame features in the temporal domain. The method has been evaluated on four action recognition datasets, including Kinetics, Something‐Something V2, HMDB51, and UCF101. The experimental results verify the effectiveness of ETA and show its superiority over state‐of‐the‐art methods.
Chunlei Xu, Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
IET Comput. Vis.5
2023 Improving Open Set Domain Adaptation Using Image-to-Image Translation and Instance-Weighted Adversarial Learning
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
J. Comput. Sci. Technol.4
2023 Adaptively feature matching via joint transformational-spatial clustering
Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Shaohua Wan 0001
Multim. Syst.4
2023 Learning Graph Convolutional Networks for Multi-Label Recognition and Applications
abstract
The task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important information, we propose graph convolutional networks (GCNs) based models for multi-label image recognition, where directed graphs are constructed over classes and information is propagated between classes to learn inter-dependent class-level representations. Following this idea, we design two particular models that approach multi-label classification from different views. In our first model, the prior knowledge about the class dependencies is integrated into classifier learning. Specifically, we propose Classifier Learning GCN (C-GCN) to map class-level semantic representations (e.g., word embeddings) into classifiers that maintain the inter-class topology. In our second model, we decompose the visual representation of an image into a set of label-aware features and propose prediction learning GCN (P-GCN) to encode such features into inter-dependent image-level prediction scores. Furthermore, we also present an effective correlation matrix construction approach to capture inter-class relationships and consequently guide information propagation among classes. Empirical results on generic multi-label image recognition demonstrate that both of the proposed models can obviously outperform other existing state-of-the-arts. Moreover, the proposed methods also show advantages in some other multi-label classification related applications.
Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 AGConv: Adaptive Graph Convolution on 3D Point Clouds
abstract
Convolution on 3D point clouds is widely researched yet far from perfect in geometric deep learning. The traditional wisdom of convolution characterises feature correspondences indistinguishably among 3D points, arising an intrinsic limitation of poor distinctive feature learning. In this article, we propose Adaptive Graph Convolution (AGConv) for wide applications of point cloud analysis. AGConv generates adaptive kernels for points according to their dynamically learned features. Compared with the solution of using fixed/isotropic kernels, AGConv improves the flexibility of point cloud convolutions, effectively and precisely capturing the diverse relations between points from different semantic parts. Unlike the popular attentional weight schemes, AGConv implements the adaptiveness inside the convolution operation instead of simply assigning different weights to the neighboring points. Extensive evaluations clearly show that our method outperforms state-of-the-arts of point cloud classification and segmentation on various benchmark datasets. Meanwhile, AGConv can flexibly serve more point cloud analysis approaches to boost their performance. To validate its flexibility and effectiveness, we explore AGConv-based paradigms of completion, denoising, upsampling, registration and circle extraction, which are comparable or even superior to their competitors.
Mingqiang Wei, Zeyong Wei, Huajian Si, Zhilei Chen, Zhe Zhu, Jingbo Qiu, Xuefeng Yan 0001, Yanwen Guo 0001, Jun Wang 0039, Harry Qin
IEEE Trans. Pattern Anal. Mach. Intell.10
2023 Ultra-High Resolution SVBRDF Recovery from a Single Image
abstract
Existing convolutional neural networks have achieved great success in recovering Spatially Varying Bidirectional Surface Reflectance Distribution Function (SVBRDF) maps from a single image. However, they mainly focus on handling low-resolution (e.g., 256 × 256) inputs. Ultra-High Resolution (UHR) material maps are notoriously difficult to acquire by existing networks because (1) finite computational resources set bounds for input receptive fields and output resolutions, and (2) convolutional layers operate locally and lack the ability to capture long-range structural dependencies in UHR images. We propose an implicit neural reflectance model and a divide-and-conquer solution to address these two challenges simultaneously. We first crop a UHR image into low-resolution patches, each of which are processed by a local feature extractor to extract important details. To fully exploit long-range spatial dependency and ensure global coherency, we incorporate a global feature extractor and several coordinate-aware feature assembly modules into our pipeline. The global feature extractor contains several lightweight material vision transformers that have a global receptive field at each scale and have the ability to infer long-term relationships in the material. After decoding globally coherent feature maps assembled by coordinate-aware feature assembly modules, the proposed end-to-end method is able to generate UHR SVBRDF maps from a single image with fine spatial details and consistent global structures.
Jie Guo 0001, Shuichang Lai, Qinghao Tu, Chengzhi Tao, Changqing Zou, Yanwen Guo 0001
ACM Trans. Graph.6
2023 Manifold Path Guiding for Importance Sampling Specular Chains
abstract
Complex visual effects such as caustics are often produced by light paths containing multiple consecutive specular vertices (dubbed specular chains) , which pose a challenge to unbiased estimation in Monte Carlo rendering. In this work, we study the light transport behavior within a sub-path that is comprised of a specular chain and two non-specular separators. We show that the specular manifolds formed by all the sub-paths could be exploited to provide coherence among sub-paths. By reconstructing continuous energy distributions from historical and coherent sub-paths, seed chains can be generated in the context of importance sampling and converge to admissible chains through manifold walks. We verify that importance sampling the seed chain in the continuous space reaches the goal of importance sampling the discrete admissible specular chain. Based on these observations and theoretical analyses, a progressive pipeline, manifold path guiding , is designed and implemented to importance sample challenging paths featuring long specular chains. To our best knowledge, this is the first general framework for importance sampling discrete specular chains in regular Monte Carlo rendering. Extensive experiments demonstrate that our method outperforms state-of-the-art unbiased solutions with up to 40 × variance reduction, especially in typical scenes containing long specular chains and complex visibility.
Zhimin Fan 0001, Pengpei Hong, Jie Guo 0001, Changqing Zou, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.5
2023 MetaLayer: A Meta-Learned BSDF Model for Layered Materials
abstract
Reproducing the appearance of arbitrary layered materials has long been a critical challenge in computer graphics, with regard to the demanding requirements of both physical accuracy and low computation cost. Recent studies have demonstrated promising results by learning-based representations that implicitly encode the appearance of complex (layered) materials by neural networks. However, existing generally-learned models often struggle between strong representation ability and high runtime performance, and also lack physical parameters for material editing. To address these concerns, we introduce MetaLayer , a new methodology leveraging meta-learning for modeling and rendering layered materials. MetaLayer contains two networks: a BSDFNet that compactly encodes layered materials into implicit neural representations, and a MetaNet that establishes the mapping between the physical parameters of each material and the weights of its corresponding implicit neural representation. A new positional encoding method and a well-designed training strategy are employed to improve the performance and quality of the neural model. As a new learning-based representation, the proposed MetaLayer model provides both fast responses to material editing and high-quality results for a wide range of layered materials, outperforming existing layered BSDF models.
Jie Guo 0001, Zeru Li, Xueyan He, Beibei Wang 0002, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.6
2023 Local-to-Global Panorama Inpainting for Locale-Aware Indoor Lighting Prediction
abstract
Predicting panoramic indoor lighting from a single perspective image is a fundamental but highly ill-posed problem in computer vision and graphics. To achieve locale-aware and robust prediction, this problem can be decomposed into three sub-tasks: depth-based image warping, panorama inpainting and high-dynamic-range (HDR) reconstruction, among which the success of panorama inpainting plays a key role. Recent methods mostly rely on convolutional neural networks (CNNs) to fill the missing contents in the warped panorama. However, they usually achieve suboptimal performance since the missing contents occupy a very large portion in the panoramic space while CNNs are plagued by limited receptive fields. The spatially-varying distortion in the spherical signals further increases the difficulty for conventional CNNs. To address these issues, we propose a local-to-global strategy for large-scale panorama inpainting. In our method, a depth-guided local inpainting is first applied on the warped panorama to fill small but dense holes. Then, a transformer-based network, dubbed PanoTransformer, is designed to hallucinate reasonable global structures in the large holes. To avoid distortion, we further employ cubemap projection in our design of PanoTransformer. The high-quality panorama recovered at any locale helps us to capture spatially-varying indoor illumination with physically-plausible global structures and fine details.
Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Yan Zhang 0057, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.7
2023 GeoDualCNN: Geometry-Supporting Dual Convolutional Neural Network for Noisy Point Clouds
abstract
We propose a geometry-supporting dual convolutional neural network (GeoDualCNN) for both point cloud normal estimation and denoising. GeoDualCNN fuses the geometry domain knowledge that the underlying surface of a noisy point cloud is piecewisely smooth with the fact that a point normal is properly defined only when local surface smoothness is guaranteed. Centered around this insight, we define the homogeneous neighborhood (HoNe) which stays clear of surface discontinuities, and associate each HoNe with a point whose geometry and normal orientation is mostly consistent with that of HoNe. Thus, we not only obtain initial estimates of the point normals by performing PCA on HoNes, but also for the first time optimize these initial point normals by learning the mapping from two proposed geometric descriptors to the ground-truth point normals. GeoDualCNN consists of two parallel branches that remove noise using the first geometric descriptor (a homogeneous height map, which encodes the point-position information), while preserving surface features using the second geometric descriptor (a homogeneous normal map, which encodes the point-normal information). Such geometry-supporting network architectures enable our model to leverage previous geometry expertise and to benefit from training data. Experiments with noisy point clouds show that GeoDualCNN outperforms the state-of-the-art methods in terms of both noise-robustness and feature preservation.
Mingqiang Wei, Honghua Chen, Yingkui Zhang, Haoran Xie 0001, Yanwen Guo 0001, Jun Wang 0039
IEEE Trans. Vis. Comput. Graph.5
2023 ShadowMover: Automatically Projecting Real Shadows onto Virtual Object
abstract
Inserting 3D virtual objects into real-world images has many applications in photo editing and augmented reality. One key issue to ensure the reality of the composite whole scene is to generate consistent shadows between virtual and real objects. However, it is challenging to synthesize visually realistic shadows for virtual and real objects without any explicit geometric information of the real scene or manual intervention, especially for the shadows on the virtual objects projected by real objects. In view of this challenge, we present, to our knowledge, the first end-to-end solution to fully automatically project real shadows onto virtual objects for outdoor scenes. In our method, we introduce the Shifted Shadow Map, a new shadow representation that encodes the binary mask of shifted real shadows after inserting virtual objects in an image. Based on the shifted shadow map, we propose a CNN-based shadow generation model named ShadowMover which first predicts the shifted shadow map for an input image and then automatically generates plausible shadows on any inserted virtual object. A large-scale dataset is constructed to train the model. Our ShadowMover is robust to various scene configurations without relying on any geometric information of the real scene and is free of manual intervention. Extensive experiments validate the effectiveness of our method.
Piaopiao Yu, Jie Guo 0001, Zhenyu Chen 0001, Chen Wang 0149, Yan Zhang 0057, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.7
2022 Unsupervised Point Cloud Completion and Segmentation by Generative Adversarial Autoencoding Network
abstract
Most existing point cloud completion methods assume the input partial point cloud is clean, which is not practical in practice, and are Most existing point cloud completion methods assume the input partial point cloud is clean, which is not the case in practice, and are generally based on supervised learning. In this paper, we present an unsupervised generative adversarial autoencoding network, named UGAAN, which completes the partial point cloud contaminated by surroundings from real scenes and cutouts the object simultaneously, only using artificial CAD models as assistance. The generator of UGAAN learns to predict the complete point clouds on real data from both the discriminator and the autoencoding process of artificial data. The latent codes from generator are also fed to discriminator which makes encoder only extract object features rather than noises. We also devise a refiner for generating better complete cloud with a segmentation module to separate the object from background. We train our UGAAN with one real scene dataset and evaluate it with the other two. Extensive experiments and visualization demonstrate our superiority, generalization and robustness. Comparisons against the previous method show that our method achieves the state-of-the-art performance on unsupervised point cloud completion and segmentation on real data.
Changfeng Ma, Yang Yang 0092, Jie Guo 0001, Chong-Jun Wang, Yanwen Guo 0001
NeurIPS6
2022 NeuLighting: Neural Lighting for Free Viewpoint Outdoor Scene Relighting with Unconstrained Photo Collections
abstract
We propose NeuLighting, a new framework for free viewpoint outdoor scene relighting from a sparse set of unconstrained in-the-wild photo collections. Our framework represents all the scene components as continuous functions parameterized by MLPs that take a 3D location and the lighting condition as input and output reflectance and necessary outdoor illumination properties. Unlike object-level relighting methods which often leverage training images with controllable and consistent indoor illumination, we concentrate on the more challenging outdoor situation where all the images are captured under arbitrary unknown illumination. The key to our method includes a neural lighting representation that compresses the per-image illumination into a disentangled latent vector, and a new free viewpoint relighting scheme that is robust to arbitrary lighting variations across images. The lighting representation is compressive to explain a wide range of illumination and can be easily fed into the query-based NeuLighting framework, enabling efficient shading effect evaluation under any kind of novel illumination. Furthermore, to produce high-quality cast shadows, we estimate the sun visibility map to indicate the shadow regions according to the scene geometry and the sun direction. Thanks to the flexible and explainable neural lighting representation, our system supports outdoor relighting with many different illumination sources, including natural images, environment maps, and time-lapse videos. The high-fidelity renderings under novel views and illumination prove the superiority of our method against state-of-the-art relighting solutions.
Quewei Li, Jie Guo 0001, Feichao Li, Yanwen Guo 0001
SIGGRAPH Asia5
2022 UTOPIC: Uncertainty-aware Overlap Prediction Network for Partial Point Cloud Registration
abstract
Abstract High‐confidence overlap prediction and accurate correspondences are critical for cutting‐edge models to align paired point clouds in a partial‐to‐partial manner. However, there inherently exists uncertainty between the overlapping and non‐overlapping regions, which has always been neglected and significantly affects the registration performance. Beyond the current wisdom, we propose a novel uncertainty‐aware overlap prediction network, dubbed UTOPIC, to tackle the ambiguous overlap prediction problem; to our knowledge, this is the first to explicitly introduce overlap uncertainty to point cloud registration. Moreover, we induce the feature extractor to implicitly perceive the shape knowledge through a completion decoder, and present a geometric relation embedding for Transformer to obtain transformation‐invariant geometry‐aware feature representations. With the merits of more reliable overlap scores and more precise dense correspondences, UTOPIC can achieve stable and accurate registration results, even for the inputs with limited overlapping areas. Extensive quantitative and qualitative experiments on synthetic and real benchmarks demonstrate the superiority of our approach over state‐of‐the‐art methods.
Zhilei Chen, Honghua Chen, Lina Gong, Xuefeng Yan 0001, Jun Wang 0039, Yanwen Guo 0001, Harry Qin, Mingqiang Wei
Comput. Graph. Forum6
2022 Real-time Deep Radiance Reconstruction from Imperfect Caches
abstract
Abstract Real‐time global illumination is a highly desirable yet challenging task in computer graphics. Existing works well solving this problem are mostly based on some kind of precomputed data (caches), while the final results depend significantly on the quality of the caches. In this paper, we propose a learning‐based pipeline that can reproduce a wide range of complex light transport phenomena, including high‐frequency glossy interreflection, at any viewpoint in real time (> 90 frames per‐second), using information from imperfect caches stored at the barycentre of every triangle in a 3D scene. These caches are generated at a precomputation stage by a physically‐based offline renderer at a low sampling rate (e.g., 32 samples per‐pixel) and a low image resolution (e.g., 64×16). At runtime, a deep radiance reconstruction method based on a dedicated neural network is then involved to reconstruct a high‐quality radiance map of full global illumination at any viewpoint from these imperfect caches, without introducing noise and aliasing artifacts. To further improve the reconstruction accuracy, a new feature fusion strategy is designed in the network to better exploit useful contents from cheap G‐buffers generated at runtime. The proposed framework ensures high‐quality rendering of images for moderate‐sized scenes with full global illumination effects, at the cost of reasonable precomputation time. We demonstrate the effectiveness and efficiency of the proposed pipeline by comparing it with alternative strategies, including real‐time path tracing and precomputed radiance transfer.
Tao Huang 0026, Yadong Song, Jie Guo 0001, Chengzhi Tao, Zijing Zong, Xihao Fu, Hongshan Li, Yanwen Guo 0001
Comput. Graph. Forum8
2022 GlassNet: Label Decoupling-based Three-stream Neural Network for Robust Image Glass Detection
abstract
Abstract Most of the existing object detection methods generate poor glass detection results, due to the fact that the transparent glass shares the same appearance with arbitrary objects behind it in an image. Different from traditional deep learning‐based wisdoms that simply use the object boundary as an auxiliary supervision, we exploit label decoupling to decompose the original labelled ground‐truth (GT) map into an interior‐diffusion map and a boundary‐diffusion map. The GT map in collaboration with the two newly generated maps breaks the imbalanced distribution of the object boundary, leading to improved glass detection quality. We have three key contributions to solve the transparent glass detection problem: (1) We propose a three‐stream neural network (call GlassNet for short) to fully absorb beneficial features in the three maps. (2) We design a multi‐scale interactive dilation module to explore a wider range of contextual information. (3) We develop an attention‐based boundary‐aware feature Mosaic module to integrate multi‐modal information. Extensive experiments on the benchmark dataset exhibit clear improvements of our method over SOTAs, in terms of both the overall glass detection accuracy and boundary clearness.
Ding Shi, Xuefeng Yan 0001, Dong Liang 0008, Mingqiang Wei, Xin Yang 0011, Yanwen Guo 0001, Haoran Xie 0001
Comput. Graph. Forum7
2022 Point attention network for point cloud semantic segmentation
Dayong Ren, Zhengyi Wu, Piaopiao Yu, Jie Guo 0001, Mingqiang Wei, Yanwen Guo 0001
Sci. China Inf. Sci.7
2022 Rendering discrete participating media using geometrical optics approximation
abstract
We consider the scattering of light in participating media composed of sparsely and randomly distributed discrete particles. The particle size is expected to range from the scale of the wavelength to several orders of magnitude greater, resulting in an appearance with distinct graininess as opposed to the smooth appearance of continuous media. One fundamental issue in the physically-based synthesis of such appearance is to determine the necessary optical properties in every local region. Since these properties vary spatially, we resort to geometrical optics approximation (GOA), a highly efficient alternative to rigorous Lorenz—Mie theory, to quantitatively represent the scattering of a single particle. This enables us to quickly compute bulk optical properties for any particle size distribution. We then use a practical Monte Carlo rendering solution to solve energy transfer in the discrete participating media. Our proposed framework is the first to simulate a wide range of discrete participating media with different levels of graininess, converging to the continuous media case as the particle concentration increases.
Jie Guo 0001, Bingyang Hu, Yuanqi Li, Yanwen Guo 0001, Lingqi Yan 0001
Comput. Vis. Media5
2022 Affinity Fusion Graph-Based Framework for Natural Image Segmentation
abstract
This paper proposes an affinity fusion graph framework to effectively connect different graphs with highly discriminating power and nonlinearity for natural image segmentation. The proposed framework combines adjacency-graphs and kernel spectral clustering based graphs (KSC-graphs) according to a new definition named affinity nodes of multi-scale superpixels. These affinity nodes are selected based on a better affiliation of superpixels, namely subspace-preserving representation which is generated by sparse subspace clustering based on subspace pursuit. Then a KSC-graph is built via a novel kernel spectral clustering to explore the nonlinear relationships among these affinity nodes. Moreover, an adjacency-graph at each scale is constructed, which is further used to update the proposed KSC-graph at affinity nodes. The fusion graph is built across different scales, and it is partitioned to obtain final segmentation result. Experimental results on the Berkeley segmentation dataset and Microsoft Research Cambridge dataset show the superiority of our framework in comparison with the state-of-the-art methods. The code is available athttps://github.com/Yangzhangcst/AF-graph.
Yang Zhang 0053, Moyun Liu, Jingwu He, Yanwen Guo 0001
IEEE Trans. Multim.5
2022 Efficient Light Probes for Real-Time Global Illumination
abstract
Reproducing physically-based global illumination (GI) effects has been a long-standing demand for many real-time graphical applications. In pursuit of this goal, many recent engines resort to some form of light probes baked in a precomputation stage. Unfortunately, the GI effects stemming from the precomputed probes are rather limited due to the constraints in the probe storage, representation or query. In this paper, we propose a new method for probe-based GI rendering which can generate a wide range of GI effects, including glossy reflection with multiple bounces, in complex scenes. The key contributions behind our work include a gradient-based search algorithm and a neural image reconstruction method. The search algorithm is designed to reproject the probes' contents to any query viewpoint, without introducing parallax errors, and converges fast to the optimal solution. The neural image reconstruction method, based on a dedicated neural network and several G-buffers, tries to recover high-quality images from low-quality inputs due to limited resolution or (potential) low sampling rate of the probes. This neural method makes the generation of light probes efficient. Moreover, a temporal reprojection strategy and a temporal loss are employed to improve temporal stability for animation sequences. The whole pipeline runs in realtime (>30 frames per second) even for high-resolution (1920×1080) outputs, thanks to the fast convergence rate of the gradient-based search algorithm and a light-weight design of the neural network. Extensive experiments on multiple complex scenes have been conducted to show the superiority of our method over the state-of-the-arts.
Jie Guo 0001, Zijing Zong, Yadong Song, Xihao Fu, Chengzhi Tao, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.6
2022 Video Vectorization via Bipartite Diffusion Curves Propagation and Optimization
abstract
We propose a new video vectorization approach for converting videos in the raster format to vector representation with the benefits of resolution independence and compact storage. Through classifying extracted curves in each video frame into salient ones and non-salient ones, we introduce a novel bipartite diffusion curves (BDCs) representation in order to preserve both important image features such as sharp boundaries and regions with smooth color variation. This bipartite representation allows us to propagate non-salient curves across frames such that the propagation, in conjunction with geometry optimization and color optimization of salient curves, ensures the preservation of fine details within each frame and across different frames, and meanwhile, achieves good spatial-temporal coherence. Thorough experiments on a variety of videos show that our method is capable of converting videos to the vector representation with low reconstruction errors, low computational cost, and fine details, demonstrating our superior performance over the state of the art. We also show that, when used for video upsampling, our method produces results comparable to video super-resolution.
Yuanqi Li, Chuan Wang 0001, Jie Guo 0001, Jue Wang 0001, Yanwen Guo 0001, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.7
2021 Tensor Composition Net for Visual Relationship Prediction
Yuting Qiang, Yongxin Yang, Yanwen Guo 0001, Timothy M. Hospedales
BMVC4
2021 GLAVNet: Global-Local Audio-Visual Cues for Fine-Grained Material Recognition
abstract
In this paper, we aim to recognize materials with combined use of auditory and visual perception. To this end, we construct a new dataset named GLAudio that consists of both the geometry of the object being struck and the sound captured from either modal sound synthesis (for virtual objects) or real measurements (for real objects). Besides global geometries, our dataset also takes local geometries around different hitpoints into consideration. This local information is less explored in existing datasets. We demonstrate that local geometry has a greater impact on the sound than the global geometry and offers more cues in material recognition. To extract features from different modalities and perform proper fusion, we propose a new deep neural network GLAVNet that comprises multiple branches and a well-designed fusion module. Once trained on GLAudio, our GLAVNet provides state-of-the-art performance on material identification and supports fine-grained material categorization.
Fengmin Shi, Jie Guo 0001, Yanwen Guo 0001
CVPR6
2021 Hierarchical Disentangled Representation Learning for Outdoor Illumination Estimation and Editing
abstract
Data-driven sky models have gained much attention in outdoor illumination prediction recently, showing superior performance against analytical models. However, naively compressing an outdoor panorama into a low-dimensional latent vector, as existing models have done, causes two major problems. One is the mutual interference between the HDR intensity of the sun and the complex textures of the surrounding sky, and the other is the lack of fine-grained control over independent lighting factors due to the entangled representation. To address these issues, we propose a hierarchical disentangled sky model (HDSky) for outdoor illumination prediction. With this model, any outdoor panorama can be hierarchically disentangled into several factors based on three well-designed autoencoders. The first autoencoder compresses each sunny panorama into a sky vector and a sun vector with some constraints. The second autoencoder and the third autoencoder further disentangle the sun intensity and the sky intensity from the sun vector and the sky vector with several customized loss functions respectively. Moreover, a unified framework is designed to predict all-weather sky information from a single outdoor image. Through extensive experiments, we demonstrate that the proposed model significantly improves the accuracy of outdoor illumination prediction. It also allows users to intuitively edit the predicted panorama (e.g., changing the position of the sun while preserving others), without sacrificing physical plausibility.
Piaopiao Yu, Jie Guo 0001, Hongwei Che, Yanwen Guo 0001
ICCV7
2021 Dual attention autoencoder for all-weather outdoor lighting estimation
Piaopiao Yu, Jie Guo 0001, Longhai Wu, Yanwen Guo 0001
Sci. China Inf. Sci.7
2021 A Unified Light Framework for Real-Time Fault Detection of Freight Train Images
abstract
Real-time fault detection for freight trains plays a vital role in guaranteeing the security and optimal operation of railway transportation under stringent resource requirements. Despite the promising results for deep-learning-based approaches, the performance of these fault detectors on freight train images is far from satisfactory in both accuracy and efficiency. This article proposes a unified light framework to improve detection accuracy while supporting a real-time operation with a low-resource requirement. We first design a novel lightweight backbone (real-time fault detection network-RFDNet) to improve the accuracy and reduce computational cost. Then, we propose a multiregion proposal network using multiscale feature maps generated from the RFDNet to improve the detection performance. Finally, we present multilevel position-sensitive score maps and region of interest pooling to further improve accuracy with few redundant computations. Extensive experimental results on public benchmark datasets suggest that our RFDNet can significantly improve the performance of the baseline network with higher accuracy and efficiency. Experiments on six fault datasets show that our method is capable of real-time detection at over 38 frames/s and achieves competitive accuracy and lower computation than the state-of-the-art detectors.
Yang Zhang 0053, Moyun Liu, Yang Yang 0092, Yanwen Guo 0001
IEEE Trans. Ind. Informatics4
2021 HCE: Hierarchical Context Embedding for Region-Based Object Detection
abstract
State-of-the-art two-stage object detectors apply a classifier to a sparse set of object proposals, relying on region-wise features extracted by RoIPool or RoIAlign as inputs. The region-wise features, in spite of aligning well with the proposal locations, may still lack the crucial context information which is necessary for filtering out noisy background detections, as well as recognizing objects possessing no distinctive appearances. To address this issue, we present a simple but effective Hierarchical Context Embedding (HCE) framework, which can be applied as a plug-and-play component, to facilitate the classification ability of a series of region-based detectors by mining contextual cues. Specifically, to advance the recognition of context-dependent object categories, we propose an image-level categorical embedding module which leverages the holistic image-level context to learn object-level concepts. Then, novel RoI features are generated by exploiting hierarchically embedded context information beneath both whole images and interested regions, which are also complementary to conventional RoI features. Moreover, to make full use of our hierarchical contextual RoI features, we propose the early-and-late fusion strategies (i.e., feature fusion and confidence fusion), which can be combined to boost the classification accuracy of region-based detectors. Comprehensive experiments demonstrate that our HCE framework is flexible and generalizable, leading to significant and consistent improvements upon various region-based detectors, including FPN, Cascade R-CNN, Mask R-CNN and PA-FPN. With simple modification, our HCE framework can be conveniently adapted to fit the structure of one-stage detectors, and achieve improved performance for SSD, RetinaNet and EfficientDet.
Xin Jin 0023, Borui Zhao, Xiaoqin Zhang 0002, Yanwen Guo 0001
IEEE Trans. Image Process.5
2021 Image Stitching Based on Semantic Planar Region Consensus
abstract
Image stitching for two images without a global transformation between them is notoriously difficult. In this paper, noticing the importance of semantic planar structures under perspective geometry, we propose a new image stitching method which stitches images by allowing for the alignment of a set of matched dominant semantic planar regions. Clearly different from previous methods resorting to plane segmentation, the key to our approach is to utilize rich semantic information directly from RGB images to extract semantic planar image regions with a deep Convolutional Neural Network (CNN). We specifically design a module implementing our newly proposed clustering loss to make full use of existing semantic segmentation networks to accommodate region segmentation. To train the network, a dataset for semantic planar region segmentation is constructed. With the prior of semantic planar region, a set of local transformation models can be obtained by constraining matched regions, enabling more precise alignment in the overlapping area. We also use this prior to estimate a transformation field over the whole image. The final mosaic is obtained by mesh-based optimization which maintains high alignment accuracy and relaxes similarity transformation at the same time. Extensive experiments with both qualitative and quantitative comparisons show that our method can deal with different situations and outperforms the state-of-the-arts on challenging scenes.
Aocheng Li, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Image Process.3
2021 Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided Network
abstract
In this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application.
Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu
IEEE Trans. Image Process.5
2021 Disentangling, Embedding and Ranking Label Cues for Multi-Label Image Recognition
abstract
Multi-label image recognition is a fundamental but challenging computer vision and multimedia task. Great progress has been achieved by exploiting label correlations among these multiple labels associated with a single image, which is the most crucial issue for multi-label image recognition. In this paper, to explicitly model label correlations, we propose a unified deep learning framework to Disentangle, Embed and Rank (DER) the corresponding label cues. Specifically, we first obtain class-aware disentangled maps (CADMs) by reforming deep activations in accordance with the class-specific recognition weights. Then, after transforming CADMs into the corresponding label vectors, we propose an embedding operation from a metric learning perspective to pull the relevant label vectors together and push irrelevant label vectors away. Furthermore, a ranking operation is employed, which aims to accurately and robustly measure the similarity/dissimilarity of these label vectors. Our model can be trained in an end-to-end manner with only image-level supervision, during which the proposed embedding and ranking operations can contribute to the CADMs learning through back-propagation. In addition, the obtained CADMs are aggregated and further used as an essential feature stream for the final multi-label classification. We conduct extensive experiments on three commonly used multi-label benchmark datasets. Quantitative results show that our model can significantly and consistently outperform previous competitive methods. Moreover, qualitative analysis of our DER proposal also reveals the effectiveness of our proposed model.
Quan Cui, Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001
IEEE Trans. Multim.5
2021 Highlight-aware two-stream network for single-image SVBRDF acquisition
abstract
This paper addresses the task of estimating spatially-varying reflectance (i.e., SVBRDF) from a single, casually captured image. Central to our method is a highlight-aware (HA) convolution operation and a two-stream neural network equipped with proper training losses. Our HA convolution, as a novel variant of standard (ST) convolution, directly modulates convolution kernels under the guidance of automatically learned masks representing potentially overexposed highlight regions. It helps to reduce the impact of strong specular highlights on diffuse components and at the same time, hallucinates plausible contents in saturated regions. Considering that variation of saturated pixels also contains important cues for inferring surface bumpiness and specular components, we design a two-stream network to extract features from two different branches stacked by HA convolutions and ST convolutions, respectively. These two groups of features are further fused in an attention-based manner to facilitate feature selection of each SVBRDF map. The whole network is trained end to end with a new perceptual adversarial loss which is particularly useful for enhancing the texture details. Such a design also allows the recovered material maps to be disentangled. We demonstrate through quantitative analysis and qualitative visualization that the proposed method is effective to recover clear SVBRDFs from a single casually captured image, and performs favorably against state-of-the-arts. Since we impose very few constraints on the capture process, even a non-expert user can create high-quality SVBRDFs that cater to many graphical applications.
Jie Guo 0001, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.6
2021 Volumetric appearance stylization with stylizing kernel prediction network
abstract
This paper aims to efficiently construct the volume of heterogeneous single-scattering albedo for a given medium that would lead to desired color appearance. We achieve this goal by formulating it as a volumetric style transfer problem in which an input 3D density volume is stylized using color features extracted from a reference 2D image. Unlike existing algorithms that require cumbersome iterative optimizations, our method leverages a feed-forward deep neural network with multiple well-designed modules. At the core of our network is a stylizing kernel predictor (SKP) that extracts multi-scale feature maps from a 2D style image and predicts a handful of stylizing kernels as a highly non-linear combination of the feature maps. Each group of stylizing kernels represents a specific style. A volume autoencoder (VolAE) is designed and jointly learned with the SKP to transform a density volume to an albedo volume based on these stylizing kernels. Since the autoencoder does not encode any style information, it can generate different albedo volumes with a wide range of appearance once training is completed. Additionally, a hybrid multi-scale loss function is used to learn plausible color features and guarantee temporal coherence for time-evolving volumes. Through comprehensive experiments, we validate the effectiveness of our method and show its superiority by comparing against state-of-the-arts. We show that with our method a novice user can easily create a diverse set of realistic translucent effects for 3D models (either static or dynamic), neglecting any cumbersome process of parameter tuning.
Jie Guo 0001, Zijing Zong, Jingwu He, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.6
2021 ExtraNet: real-time extrapolated rendering for low-latency temporal supersampling
abstract
Both the frame rate and the latency are crucial to the performance of realtime rendering applications such as video games. Spatial supersampling methods, such as the Deep Learning SuperSampling (DLSS), have been proven successful at decreasing the rendering time of each frame by rendering at a lower resolution. But temporal supersampling methods that directly aim at producing more frames on the fly are still not practically available. This is mainly due to both its own computational cost and the latency introduced by interpolating frames from the future. In this paper, we present ExtraNet, an efficient neural network that predicts accurate shading results on an extrapolated frame, to minimize both the performance overhead and the latency. With the help of the rendered auxiliary geometry buffers of the extrapolated frame, and the temporally reliable motion vectors, we train our ExtraNet to perform two tasks simultaneously: irradiance in-painting for regions that cannot find historical correspondences, and accurate ghosting-free shading prediction for regions where temporal information is available. We present a robust hole-marking strategy to automate the classification of these tasks, as well as the data generation from a series of high-quality production-ready scenes. Finally, we use lightweight gated convolutions to enable fast inference. As a result, our ExtraNet is able to produce plausibly extrapolated frames without easily noticeable artifacts, delivering a 1.5× to near 2× increase in frame rates with minimized latency in practice.
Jie Guo 0001, Xihao Fu, Liqiang Lin, Hengjun Ma, Yanwen Guo 0001, Shiqiu Liu, Lingqi Yan 0001
ACM Trans. Graph.5
2020 Hybrid Models for Open Set Recognition
Hongjie Zhang 0002, Jie Guo 0001, Yanwen Guo 0001
ECCV (3)4
2020 Hierarchical Context Embedding for Region-Based Object Detection
Xin Jin 0023, Borui Zhao, Xiu-Shen Wei, Yanwen Guo 0001
ECCV (21)5
2020 Deep Surface Normal Estimation on the 2-Sphere with Confidence Guided Semantic Attention
Quewei Li, Jie Guo 0001, Qinyu Tang, Wenxiu Sun, Jin Zeng 0004, Yanwen Guo 0001
ECCV (24)7
2020 Two-Stage Depth Video Recovery with Spatiotemporal Coherence
abstract
This paper proposes a practical two-stage method to enhance the depth channel of low-quality RGB-D videos. At the first stage of the proposed method, we select several key frames from an input RGB-D video and recover them with a new MRF-regularized low-rank matrix completion method. At the second stage, we introduce a novel Laplacian smoothing algorithm to smoothly expand the recovered key depth frames to other in-between frames. With associated weights extracted from color images and a depth confidence constraint, we are able to recover each in-between frame faithfully and guarantee long-range temporal consistency. Experiments show that our method can generate high-quality and spatiotemporally coherent RGB-D videos for a wide range of scene configurations and achieve the state-of-the-art performance. Several applications further validate the effectiveness of the proposed method.
Quewei Li, Jie Guo 0001, Qinyu Tang, Yanwen Guo 0001, Jinghui Qian
ICME4
2020 DeepBRDF: A Deep Representation for Manipulating Measured BRDF
abstract
Abstract Effective compression of densely sampled BRDF measurements is critical for many graphical or vision applications. In this paper, we present DeepBRDF, a deep‐learning‐based representation that can significantly reduce the dimensionality of measured BRDFs while enjoying high quality of recovery. We consider each measured BRDF as a sequence of image slices and design a deep autoencoder with a masked L 2 loss to discover a nonlinear low‐dimensional latent space of the high‐dimensional input data. Thorough experiments verify that the proposed method clearly outperforms PCA‐based strategies in BRDF data compression and is more robust. We demonstrate the effectiveness of DeepBRDF with two applications. For BRDF editing, we can easily create a new BRDF by navigating on the low‐dimensional manifold of DeepBRDF, guaranteeing smooth transitions and high physical plausibility. For BRDF recovery, we design another deep neural network to automatically generate the full BRDF data from a single input image. Aided by our DeepBRDF learned from real‐world materials, a wide range of reflectance behaviors can be recovered with high accuracy.
Bingyang Hu, Jie Guo 0001, Yanwen Guo 0001
Comput. Graph. Forum5
2020 Normal-Based Bas-Relief Modelling via Near-Lighting Photometric Stereo
abstract
Abstract We present a near‐lighting photometric stereo (NL‐PS) system to produce digital bas‐reliefs from a physical object (set) directly. Unlike both the 2D image and 3D model‐based modelling methods that require complicated interactions and transformations, the technique using NL‐PS is easy to use with cost‐effective hardware, providing users with a trade‐off between abstract and representation when creating bas‐reliefs. Our algorithm consists of two steps: normal map acquisition and constrained 3D reconstruction. First, we introduce a lighting model, named the quasi‐point lighting model (QPLM), and provide a two‐step calibration solution in our NL‐PS system to generate a dense normal map. Second, we filter the normal map into a detail layer and a structure layer, and formulate detail‐ or structure‐preserving bas‐relief modelling as a constrained surface reconstruction problem of solving a sparse linear system. The main contribution is a WYSIWYG (i.e. what you see is what you get) way of building new solvers that produces multi‐style bas‐reliefs with their geometric structures and/or details preserved. The performance of our approach is experimentally validated via comparisons with the state‐of‐the‐art methods.
Mingqiang Wei, Zhan Song, Ying Nie 0006, Jianhuang Wu, Zhongping Ji, Yanwen Guo 0001, Haoran Xie 0001, Jun Wang 0039, Fu Lee Wang
Comput. Graph. Forum6
2020 On large appearance change in visual tracking
Yun Liang 0003, Meihua Wang, Yanwen Guo 0001, Wei-Shi Zheng 0001
Neural Comput. Appl.3
2020 BRDF Analysis with Directional Statistics and its Applications
abstract
Data-driven BRDF models using real material measurements have become increasingly prevalent due to the development of novel gonioreflectometers, but efficient use of these models in many graphical applications remains challenging due to the few functionalities the raw data could provide. To ameliorate this issue, we propose to analyze BRDFs using directional statistics for better handling and exploring measured materials, especially isotropic materials, with efficient computation and compact storage. We conduct a thorough statistical analysis on both analytical BRDF models and measured materials from the MERL database. We show that different aspects of visual appearance can be characterized by different spherical moments, from which several descriptive measures can be derived to further facilitate their usage. We demonstrate how these measures are best leveraged in some graphical applications including gamut mapping using a new BRDF similarity measure, BRDF or SVBRDF reconstruction based on material clustering, and importance sampling for measured materials based on fast extracted GGX distributions. We finally show the potential of our approach in the categorization of surface reflectance types which is common for traditional photon mapping.
Jie Guo 0001, Yanwen Guo 0001, Jingui Pan, Wenzhou Lu
IEEE Trans. Vis. Comput. Graph.2
2020 Data-Driven Indoor Scene Modeling from a Single Color Image with Iterative Object Segmentation and Model Retrieval
abstract
We propose a new method for modeling the indoor scene from a single color image. With our system, the user only needs to drag a few semantic bounding boxes surrounding the objects of interest. Our system then automatically finds the most similar 3D models from the ShapeNet model repository and aligns them with the corresponding objects of interest. To achieve this, each 3D model is represented as a group of view-dependent representations generated from a set of synthesized views. We iteratively conduct object segmentation and 3D model retrieval, based on the observation that good segmentation of the objects of interest can significantly improve the accuracy of model retrieval and make it robust to cluttered background and occlusions, and in turn, the retrieved 3D models can be used to assist with object segmentation. Segmentation of all objects of interest is achieved simultaneously under a unified multi-labeling framework which fully utilizes the correspondences between the objects of interest and retrieved model images. Besides, we propose a new method to estimate the scene layout of the input image with the segmentation masks, which helps compose the resulting scene and further improves the modeling result remarkably. We verify the effectiveness of our approach through experimenting with a variety of indoor images and comparing against the relevant methods.
Mingming Liu 0004, Jun Wang 0039, Jie Guo 0001, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.6
2019 Multi-Label Image Recognition With Graph Convolutional Networks
abstract
The task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important dependencies, we propose a multi-label classification model based on Graph Convolutional Network (GCN). The model builds a directed graph over the object labels, where each node (label) is represented by word embeddings of a label, and GCN is learned to map this label graph into a set of inter-dependent object classifiers. These classifiers are applied to the image descriptors extracted by another sub-net, enabling the whole network to be end-to-end trainable. Furthermore, we propose a novel re-weighted scheme to create an effective label correlation matrix to guide information propagation among the nodes in GCN. Experiments on two multi-label image recognition datasets show that our approach obviously outperforms other existing state-of-the-art methods. In addition, visualization analyses reveal that the classifiers learned by our model maintain meaningful semantic topology.
Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001
CVPR4
2019 A Data-Driven Framework for Appearance Editing of Measured Materials
abstract
In this paper, we present an efficient framework for editing the appearance of measured materials. First, we separate the material data into diffuse and specular parts, which is important for measured BRDFs since editing without distinguishing the two parts makes the results less predictive. Additionally, the albedo of each part is calculated for later color editing. Next, we use a novel method to cluster the materials by their gloss levels and select the representative material of every level. With this at hand, the roughness, i.e., shape of the specular highlights can be adjusted among different levels. Finally, we reconstruct the edited materials which can be directly used in both online and offline rendering. We test the framework on the MERL dataset to validate the effectiveness of our method.
Jie Guo 0001, Bingyang Hu, Yanwen Guo 0001, Jingui Pan
ICME4
2019 Multi-Label Image Recognition with Joint Class-Aware Map Disentangling and Label Correlation Embedding
abstract
Multi-label image recognition is a fundamental but challenging computer vision task. Great progress has been achieved by exploring the label correlation among these multiple labels which is the most crucial issue for multi-label recognition. In this paper, we propose a unified deep learning framework to jointly disentangle class-specific maps corresponding to discriminative category-wise information and then evaluate the label co-occurrence of these maps. Specifically, after obtaining the general deep image features and conducting multi-label classification, we employ the classification weights to reform the feature maps into class-aware disentangled maps (CADMs). Then, based on CADMs, we first transfer them into label vectors and then formulate the label correlation dependency from an embedding perspective. The whole model is driven by both the classification loss and the label correlation embedding loss, which is end-to-end trainable with only image-level supervisions. Extensive quantitative results of two benchmark multi-label image datasets show our model consistently outperforms other competing methods by a large margin. Meanwhile, qualitative analyses also demonstrate our model can effectively capture relatively pure class-aware maps and model label correlation dependency as well.
Xiu-Shen Wei, Xin Jin 0023, Yanwen Guo 0001
ICME4
2019 PANet: A Context Based Predicate Association Network for Scene Graph Generation
abstract
Scene graph generation is widely studied in recent years, which tries to understand the interactions of different objects as a whole. The earlier researches only recognize a few relationships or model contexts among different relationships, neglecting the associations of predicates for each object pair. In this paper, we propose a two-stage framework named predicate association network (PANet) to properly extract contexts and model predicate association. In the first stage, instance-level and scene-level context are extracted for object classification and further used for predicate classification in the next stage. With a recurrent neural network, alignment technique and attention mechanism are combined to collect the associations of predicates in the second stage. The experiments on the Visual Genome dataset show that our method is effective and outperforms the state-of-the-art methods.
Yunian Chen, Yang Zhang 0053, Yanwen Guo 0001
ICME4
2019 Temporal Segment Convolutional Kernel Networks for Sequence Modeling of Videos
abstract
Sequence modeling is crucial for video action recognition. In this paper, we propose temporal segment convolutional kernel networks (TS-CKN), where we take advantage of convolutional neural networks to facilitate the extraction of appearance features, while time sequence is modeled with deep kernel networks. We employ the kernel methods to capture time-varying information of videos and propose a training method for kernel map approximation by matrix backpropagation. This leads to the model named deep kernel networks which can be easily integrated with existing deep learning models such as Resnet. Our approach also samples several video clips sparsely in the video and unifies class predictions from all clips. More importantly, all parameters of our model can be learned by stochastic optimization in an end-to-end manner. We evaluate our method on two standard action recognition datasets including HMDB-51 and UCF-101, achieving the state-of-the-art results.
Yanwen Guo 0001, Zhicheng Yan 0001, Jie Guo 0001
ICME2
2019 Improving Open Set Domain Adaptation Using Image-to-Image Translation
abstract
The open set domain adaptation problem was rarely studied and its existing solutions are mostly based on learning a joint latent space which may encounter issues when the domains differ significantly from each other. This work is driven by the question whether or not it is beneficial to operate the source images to another image domain as close to the target as possible. We propose to address the open set domain adaptation problem by aligning sample at both feature space and pixel space. Our approach, called Open Set Translation and Adaptation Network (Ostan), consists of two main components: translation and adaptation. The translation model is a cycle-consistent generative adversarial network, which translates any source sample to the "style" of a target domain. The adaptation network is built upon OpenBP, an open set domain adaptation framework, and trained using both (labeled) translated source images and (unlabeled) target images. The proposed Ostan model significantly outperforms the state-of-the-art open set domain adaptation methods on multiple public datasets. Our experiment also demonstrates that an image-to-image translation component can further improve the decision boundaries for both known and unknown classes.
Hongjie Zhang 0002, Yang Zhang 0053, Yanwen Guo 0001
ICME6
2019 An Adaptive Affinity Graph with Subspace Pursuit for Natural Image Segmentation
abstract
Graph-based segmentation methods have become a major trend in computer vision. Due to the advantages of assimilating different graphs, a multi-scale fusion graph have a better performance than a single graph with single-scale. However, it is not reliable to determine a principle of graph combination. In this paper, we propose an adaptive affinity graph with subspace pursuit (AASP-graph) for natural image segmentation. The input image is first over-segmented into superpixels at different scales. An improved affinity propagation clustering method is proposed to select global nodes of these superpixels adaptively. Then, a L0-graph at each scale is obtained by a sparse representation of global nodes based on subspace pursuit. The adjacency-graph is finally built upon all superpixels of each scale and updated by the L0-graph. Experimental results on the Berkeley segmentation database show the effectiveness of the proposed AASP-graph in comparison with state-of-the-art approaches.
Yang Zhang 0053, Yanwen Guo 0001, Jingwu He
ICME3
2019 Deep Spherical Gaussian Illumination Estimation for Indoor Scene
abstract
In this paper, we propose a learning-based method to estimate high dynamic range (HDR) indoor illumination from only a single low dynamic range (LDR) photograph of limited field-of-view. Considering the extreme complexity of indoor illumination that is virtually impossible to reconstruct perfectly, we choose to encode the environmental illumination in Spherical Gaussian (SG) functions with fixed centering directions and bandwidth and only allow the weights vary. An end-to-end convolutional neural network (CNN) is designed and trained to build the complex relationship between a photograph and its illumination represented by SG functions. Moreover, we employ a masked L2 loss instead of naive L2 loss to avoid the loss of high frequency information, and propose a glossy loss to improve the rendering quality. Our experiments demonstrate that the proposed approach outperforms the state-of-the-arts both qualitatively and quantitatively.
Jie Guo 0001, Xiufen Cui, Yanwen Guo 0001, Piaopiao Yu
MMAsia5
2019 Structure-guided shape-preserving mesh texture smoothing via joint low-rank matrix recovery
Honghua Chen, Oussama Remil, Haoran Xie 0001, Harry Qin, Yanwen Guo 0001, Mingqiang Wei, Jun Wang 0039
Comput. Aided Des.6
2019 Intrinsic shape matching via tensor-based optimization
Oussama Remil, Qian Xie 0001, Qiaoyun Wu, Yanwen Guo 0001, Jun Wang 0039
Comput. Aided Des.4
2019 Label transfer between images and 3D shapes via local correspondence encoding
Jie Guo 0001, Huikun Liu, Mingming Liu 0004, Yang Liu 0014, Yanwen Guo 0001
Comput. Aided Geom. Des.7
2019 3D model retrieval and pose estimation for indoor images by simulating scene context
Mingming Liu 0004, Jie Guo 0001, Yanwen Guo 0001
Graph. Model.3
2019 A unified two-parallel-branch deep neural network for joint gland contour and segmentation learning
Linbo Wang 0001, Hui Zhen, Xianyong Fang, Shaohua Wan 0001, Weiping Ding 0001, Yanwen Guo 0001
Future Gener. Comput. Syst.6
2019 Learning to Generate Posters of Scientific Papers by Probabilistic Graphical Models
Yuting Qiang, Yanwei Fu 0001, Yanwen Guo 0001, Zhi-Hua Zhou, Leonid Sigal
J. Comput. Sci. Technol.4
2019 Fractional gaussian fields for modeling and rendering of spatially-correlated media
abstract
Transmission of radiation through spatially-correlated media has demonstrated deviations from the classical exponential law of the corresponding uncorrelated media. In this paper, we propose a general, physically-based method for modeling such correlated media with non-exponential decay of transmittance. We describe spatial correlations by introducing the Fractional Gaussian Field (FGF), a powerful mathematical tool that has proven useful in many areas but remains under-explored in graphics. With the FGF, we study the effects of correlations in a unified manner, by modeling both high-frequency, noise-like fluctuations and k -th order fractional Brownian motion (fBm) with a stochastic continuity property. As a result, we are able to reproduce a wide variety of appearances stemming from different types of spatial correlations. Compared to previous work, our method is the first that addresses both short-range and long-range correlations using physically-based fluctuation models. We show that our method can simulate different extents of randomness in spatially-correlated media, resulting in a smooth transition in a range of appearances from exponential falloff to complete transparency. We further demonstrate how our method can be integrated into an energy-conserving RTE framework with a well-designed importance sampling scheme and validate its ability compared to the classical transport theory and previous work.
Jie Guo 0001, Bingyang Hu, Lingqi Yan 0001, Yanwen Guo 0001
ACM Trans. Graph.5
2019 GradNet: unsupervised deep screened poisson reconstruction for gradient-domain rendering
abstract
Monte Carlo (MC) methods for light transport simulation are flexible and general but typically suffer from high variance and slow convergence. Gradientdomain rendering alleviates this problem by additionally generating image gradients and reformulating rendering as a screened Poisson image reconstruction problem. To improve the quality and performance of the reconstruction, we propose a novel and practical deep learning based approach in this paper. The core of our approach is a multi-branch auto-encoder, termed GradNet, which end-to-end learns a mapping from a noisy input image and its corresponding image gradients to a high-quality image with low variance. Once trained, our network is fast to evaluate and does not require manual parameter tweaking. Due to the difficulty in preparing ground-truth images for training, we design and train our network in a completely unsupervised manner by learning directly from the input data. This is the first solution incorporating unsupervised deep learning into the gradient-domain rendering framework. The loss function is defined as an energy function including a data fidelity term and a gradient fidelity term. To further reduce the noise of the reconstructed image, the loss function is reinforced by adding a regularizer constructed from selected rendering-specific features. We demonstrate that our method improves the reconstruction quality for a diverse set of scenes, and reconstructing a high-resolution image takes far less than one second on a recent GPU.
Jie Guo 0001, Quewei Li, Yuting Qiang, Bingyang Hu, Yanwen Guo 0001, Lingqi Yan 0001
ACM Trans. Graph.6
2019 Viewpoint Assessment and Recommendation for Photographing Architectures
abstract
This paper studies the problem of how to assess the quality of photographing viewpoints and how to choose good viewpoints for taking photographs of architectures. We achieve this by learning from photographs of world famous landmarks that are available on the Internet and their viewpoint quality ranked by online user annotation. Unlike previous efforts devoted to photo quality assessment which mainly rely on 2D image features, we show in this paper combining 2D image features extracted from images with 3D geometric features computed on the 3D models can result in more reliable evaluation of viewpoint quality. Specifically, we collect a set of photographs for each of 15 world famous architectures as well as their 3D models from the Internet. Viewpoint recovery for images is carried out through an image-model registration process, after which a newly proposed viewpoint clustering strategy is exploited to validate users' viewpoint preferences when photographing landmarks. Finally, we extract a number of 2D and 3D features for each image based on multiple visual and geometric cues and perform viewpoint recommendation by learning from both 2D and 3D features using a specifically designed SVM-2K multi-view learner, achieving superior performance over using solely 2D or 3D features. We show the effectiveness of the proposed approach through extensive experiments. The experiments also demonstrate that our system can be used to recommend viewpoints for rendering textured 3D models of buildings for the use of architectural design, in addition to viewpoint evaluation of photographs and recommendation of viewpoints for photographing architectures in practice.
Jingwu He, Linbo Wang 0001, Wenzhe Zhou, Hongjie Zhang 0002, Xiufen Cui, Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.6
2018 A Unified Framework for Fault Detection of Freight Train Images Under Complex Environment
abstract
This paper proposes a novel unified framework for fault detection of the freight train images based on convolutional neural network (CNN) under complex environment. Firstly, the multi region proposal networks (MRPN) with a set of prior bounding boxes are introduced to achieve high quality fault proposal generation. And then, we apply a linear non-maximum suppression method to retain the most suitable anchor while removing redundant boxes. Finally, a powerful multi-level region-of-interest (ROI) pooling is proposed for proposal classification and accurate detection. The experimental results indicate that the proposed method can achieve high performance on four typical fault benchmarks, substantially outperforming the state-of-the-art methods.
Yang Zhang 0053, Yanwen Guo 0001, Guodong Sun 0002
ICIP4
2018 Image-Based 3D Model Retrieval for Indoor Scenes by Simulating Scene Context
abstract
We propose a single image-based 3D model retrieval method for indoor scenes. By simulating the scene context of the input image, our method is able to handle several challenging scenarios featuring cluttered backgrounds and severe occlusions. To use our system, the user only needs to drag a few semantic bounding boxes for the query objects. The proposed approach then retrieves the most similar 3D models from the ShapeNet model repository, and aligns them with the corresponding objects automatically. This requires that the 3D models are represented by calibrated view-dependent visual elements learned from the rendered views. With the estimated occlusion relationships, the rendered model images are stacked at the corresponding locations to simulate the scene context. By conducting matching between these synthesized scenes and the input image, the most similar 3D models under the approximate poses are retrieved. Moreover, we show that the retrieving time can be significantly reduced based on a novel greedy algorithm. Experimental results demonstrate the effectiveness of our proposed method.
Mingming Liu 0004, Jingwu He, Jie Guo 0001, Yanwen Guo 0001
ICIP5
2018 Modeling indoor scenes with repetitions from 3D raw point data
Jun Wang 0039, Qiaoyun Wu, Oussama Remil, Yanwen Guo 0001, Mingqiang Wei
Comput. Aided Des.5
2018 A Physically-based Appearance Model for Special Effect Pigments
abstract
Abstract An appearance model for materials adhered with massive collections of special effect pigments has to take both high‐frequency spatial details (e.g., glints) and wave‐optical effects (e.g., iridescence) due to thin‐film interference into account. However, either phenomenon is challenging to characterize and simulate in a physically accurate way. Capturing these fascinating effects in a unified framework is even harder as the normal distribution function and the reflectance term are highly correlated and cannot be treated separately. In this paper, we propose a multi‐scale BRDF model for reproducing the main visual effects generated by the discrete assembly of special effect pigments, enabling a smooth transition from fine‐scale surface details to large‐scale iridescent patterns. We demonstrate that the wavelength‐dependent reflectance inside the pixel's footprint follows a Gaussian distribution according to the central limit theorem, and is closely related to the distribution of the thin‐film's thickness. We efficiently determine the mean and the variance of this Gaussian distribution for each pixel whose closed‐form expressions can be derived by assuming that the thin‐film's thickness is uniformly distributed. To validate its effectiveness, the proposed model is compared against some previous methods and photographs of actual materials. Furthermore, since our method does not require any scene‐dependent precomputation, the distribution of thickness is allowed to be spatially‐varying.
Jie Guo 0001, Yanwen Guo 0001, Jingui Pan
Comput. Graph. Forum3
2018 A retroreflective BRDF model based on prismatic sheeting and microfacet theory
Jie Guo 0001, Yanwen Guo 0001, Jingui Pan
Graph. Model.2
2018 Object Detection and Tracking Under Occlusion for Object-Level RGB-D Video Segmentation
abstract
RGB-D video segmentation is important for many applications, including scene understanding, object tracking, and robotic grasping. However, to segment RGB-D frames over a long video sequence into globally consistent segmentation is still a challenging problem. Current methods often lose pixel correspondences between frames under occlusion and, thus, fail to generate consistent and continuous segmentation results. To address this problem, we propose a novel spatiotemporal RGB-D video segmentation framework that automatically segments and tracks objects with continuity and consistency over time. Our approach first produces consistent segments in some keyframes by region clustering, and then propagates the segmentation result to a whole video sequence via a mask propagation scheme in bilateral space. Instead of exploiting local optical, flow information to establish correspondences between adjacent frames, we leverage scale-invariant feature transform (SIFT) flow and bilateral representation to solve inconsistency under occlusion. Moreover, our method automatically extracts multiple objects of interest and tracks them without any user input hint. A variety of experiments demonstrates effectiveness and robustness of our proposed method.
Qian Xie 0001, Oussama Remil, Yanwen Guo 0001, Meng Wang 0001, Mingqiang Wei, Jun Wang 0039
IEEE Trans. Multim.3
2018 Correlation-Preserving Photo Collage
abstract
A new method is presented for producing photo collages that preserve content correlation of photos. We use deep learning techniques to find correlation among given photos to facilitate their embedding on the canvas, and develop an efficient combinatorial optimization technique to make correlated photos stay close to each other. To make efficient use of canvas space, our method first extracts salient regions of photos and packs only these salient regions. We allow the salient regions to have arbitrary shapes, therefore yielding informative, yet more compact collages than by other similar collage methods based on salient regions. We present extensive experimental results, user study results, and comparisons against the state-of-the-art methods to show the superiority of our method.
Lingjie Liu, Hongjie Zhang 0002, Guangmei Jing, Yanwen Guo 0001, Zhonggui Chen, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2018 A Data-Driven Approach for Furniture and Indoor Scene Colorization
abstract
We present a data-driven approach that colorizes 3D furniture models and indoor scenes by leveraging indoor images on the internet. Our approach is able to colorize the furniture automatically according to an example image. The core is to learn image-guided mesh segmentation to segment the model into different parts according to the image object. Given an indoor scene, the system supports colorization-by-example, and has the ability to recommend the colorization scheme that is consistent with a user-desired color theme. The latter is realized by formulating the problem as a Markov random field model that imposes user input as an additional constraint. Our system is able to imitate the colorization results for those scenes containing the same type of objects, but with spatially varied patterns. We contribute to the community a hierarchically organized image-model database with correspondences between each image and the corresponding model at the part-level. Our experiments and a user study show that our system produces perceptually convincing results comparable to those generated by interior designers.
Yanwen Guo 0001
IEEE Trans. Vis. Comput. Graph.2
2017 Denoising of Hyperspectral Images Using Nonconvex Low Rank Matrix Approximation
abstract
Hyperspectral image (HSI) denoising is challenging not only because of the difficulty in preserving both spectral and spatial structures simultaneously, but also due to the requirement of removing various noises, which are often mixed together. In this paper, we present a nonconvex low rank matrix approximation (NonLRMA) model and the corresponding HSI denoising method by reformulating the approximation problem using nonconvex regularizer instead of the traditional nuclear norm, resulting in a tighter approximation of the original sparsity-regularised rank function. NonLRMA aims to decompose the degraded HSI, represented in the form of a matrix, into a low rank component and a sparse term with a more robust and less biased formulation. In addition, we develop an iterative algorithm based on the augmented Lagrangian multipliers method and derive the closed-form solution of the resulting subproblems benefiting from the special property of the nonconvex surrogate function. We prove that our iterative optimization converges easily. Extensive experiments on both simulated and real HSIs indicate that our approach can not only suppress noise in both severely and slightly noised bands but also preserve large-scale image structures and small-scale details well. Comparisons against state-of-the-art LRMA-based HSI denoising approaches show our superior performance.
Yongyong Chen, Yanwen Guo 0001, Yongli Wang 0004, Chong Peng 0001, Guoping He
IEEE Trans. Geosci. Remote. Sens.2
2017 Cosegmentation for Object-Based Building Change Detection From High-Resolution Remotely Sensed Images
abstract
This paper presents a cosegmentation-based method for building change detection from multitemporal high-resolution (HR) remotely sensed images, providing a new solution to object-based change detection (OBCD). First, the magnitude of a difference image is calculated to represent the change feature. Next, cosegmentation is performed via graph-based energy minimization by combining the change feature with image features at each phase, directly resulting in foreground as multitemporal changed objects and background as unchanged area. Finally, the spatial correspondence between changed objects is established through overlay analysis. Cosegmentation provides a separate and associated, rather than a separate and independent, multitemporal image segmentation method for OBCD, which has two advantages: 1) both the image and change features are used to produce foreground segments as changed objects, which can take full advantage of multitemporal information and produce two spatially corresponded change detection maps by the association of the change feature, having the ability to reveal the thematic, geometric, and numeric changes of objects and 2) the background in the cosegmentation result represents the unchanged area, which naturally avoids the problem of matching inconsistent unchanged objects caused by the separate and independent multitemporal segmentation strategy. Experimental results on five HR datasets verify the effectiveness of the proposed method and the comparisons with the state-of-the-art OBCD methods further show its superiority.
Pengfeng Xiao, Xueliang Zhang 0002, Xuezhi Feng, Yanwen Guo 0001
IEEE Trans. Geosci. Remote. Sens.5
2017 Video Vectorization via Tetrahedral Remeshing
abstract
We present a video vectorization method that generates a video in vector representation from an input video in raster representation. A vector-based video representation offers the benefits of vector graphics, such as compactness and scalability. The vector video we generate is represented by a simplified tetrahedral control mesh over the spatial-temporal video volume, with color attributes defined at the mesh vertices. We present novel techniques for simplification and subdivision of a tetrahedral mesh to achieve high simplification ratio while preserving features and ensuring color fidelity. From an input raster video, our method is capable of generating a compact video in vector representation that allows a faithful reconstruction with low reconstruction errors.
Chuan Wang 0001, Yanwen Guo 0001, Wenping Wang 0001
IEEE Trans. Image Process.3
2017 Rendering Thin Transparent Layers with Extended Normal Distribution Functions
abstract
Realistic Rendering of thin transparent layers bounded by rough surfaces involves substantial expense of computation time to account for multiple internal reflections. Resorting to Monte Carlo rendering for such material is usually impractical since recursive importance sampling is inevitable. To reduce the burden of sampling for simulating subsurface scattering and hence improve rendering performance, we adapt the microfacet model to the material with a single thin layer by introducing the extended normal distribution function (ENDF), a new representation of this model, to express visually perceived roughness due to multiple bounces of reflections and refractions. With such a representation, both surface reflection and subsurface scattering can be treated in the same microfacet framework, and the sampling process can be reduced to only once for each bounce of scattering. We derive analytical expressions of the ENDF for several cases using joint spherical warping. We also show how to choose proper shadowing-masking and Fresnel terms to make the proposed bidirectional scattering distribution function (BSDF) model energy-conserving. Experiments demonstrate that our model can be easily incorporated into a Monte Carlo path tracer with little extra computational and storage overhead, enabling some real-time applications.
Jie Guo 0001, Jinghui Qian, Yanwen Guo 0001, Jingui Pan
IEEE Trans. Vis. Comput. Graph.3
2017 Indoor scene modeling from a single image using normal inference and edge features
Mingming Liu 0004, Yanwen Guo 0001, Jun Wang 0039
Vis. Comput.2
2016 Learning to Generate Posters of Scientific Papers
abstract
Researchers often summarize their work in the form of posters. Posters provide a coherent and efficient way to convey core ideas from scientific papers. Generating a good scientific poster, however, is a complex and time consuming cognitive task, since such posters need to be readable, informative, and visually aesthetic. In this paper, for the first time, we study the challenging problem of learning to generate posters from scientific papers. To this end, a data-driven framework, that utilizes graphical models, is proposed. Specifically, given content to display, the key elements of a good poster, including panel layout and attributes of each panel, are learned and inferred from data. Then, given inferred layout and attributes, composition of graphical elements within each panel is synthesized. To learn and validate our model, we collect and make public a Poster-Paper dataset, which consists of scientific papers and corresponding posters with exhaustively labelled panels and attributes. Qualitative and quantitative results indicate the effectiveness of our approach.
Yuting Qiang, Yanwei Fu 0001, Yanwen Guo 0001, Zhi-Hua Zhou, Leonid Sigal
AAAI3
2016 Normal Guided Data-Driven Semantic Modeling from a Single Indoor Image
abstract
We present in this paper an interactive approach for semantically modeling indoor environments given only a single indoor image as input, without requiring access to the scene or using any additional measurements like the RGBD cameras. Our key insight is that, although depth estimation from a single image is notoriously difficult, we can conveniently obtain a relatively accurate normal map, which essentially conveys a great deal of scene geometry. This enables us to model each object in a data-driven manner by representing the object as a normal-based graph and retrieving a similar model from the database by graph matching. We hypothesize a set of sparse surface orientations for the image, and further refine them in an intuitive and straightforward manner. With a small amount of simple user interaction, our approach is able to generate a plausible model of the scene. To verify the effectiveness of our proposed method, we show the modeling results on a variety of indoor images.
Mingming Liu 0004, Yanwen Guo 0001, Jun Wang 0039
CW2
2016 Multi-view Metric Learning for Multi-view Video Summarization
abstract
Traditional methods on video summarization are designed to generate summaries for single-view video records, and thus they cannot fully exploit the mutual information in multi-view video records. In this paper, we present a multiview metric learning framework for multi-view video summarization. It combines the advantages of maximum margin clustering with the disagreement minimization criterion. The learning framework thus has the ability to find a metric that best separates the input data, and meanwhile to force the learned metric to maintain underlying intrinsic structure of data points, for example geometric information. Facilitated by such a framework, a systematic solution to the multi-view video summarization problem is developed from the viewpoint of metric learning. The effectiveness of the proposed method is demonstrated by experiments.
Linbo Wang 0001, Xianyong Fang, Yanwen Guo 0001, Yanwei Fu 0001
CW3
2016 3D panorama reconstruction based on sitemap joining
abstract
We present a new approach for constructing the 3D panorama for an indoor environment by joining together aligned submaps. For each submap, the trajectory of the moving camera is estimated based on the Kanade-Lucas-Tomasi (KLT) features. Our method can update the feature set status by adding new features and removing expiring ones adaptively to accommodate scene changes. The accuracy of the estimated poses is further improved through sparse bundle adjustment. Furthermore, we utilize a linear optimization framework to align all submaps to obtain a consistently extended 3D panorama and to refine the visual odometry at the same time. We evaluated our approach on publicly available benchmark datasets. The experiments demonstrate that the proposed method achieves low translational drift and is robust even when the camera moves very fast.
Huixuan Wang, Yanwen Guo 0001, Minh N. Do, Caiming Zhang 0001, Changhe Tu
ICASSP2
2015 A 3D shape descriptor based on spectral analysis of medial axis
Shuiqing He, Yi-King Choi, Yanwen Guo 0001, Xiaohu Guo, Wenping Wang 0001
Comput. Aided Geom. Des.3
2015 Common Visual Pattern Discovery via Nonlinear Mean Shift Clustering
abstract
Discovering common visual patterns (CVPs) from two images is a challenging task due to the geometric and photometric deformations as well as noises and clutters. The problem is generally boiled down to recovering correspondences of local invariant features, and the conventionally addressed by graph-based quadratic optimization approaches, which often suffer from high computational cost. In this paper, we propose an efficient approach by viewing the problem from a novel perspective. In particular, we consider each CVP as a common object in two images with a group of coherently deformed local regions. A geometric space with matrix Lie group structure is constructed by stacking up transformations estimated from initially appearance-matched local interest region pairs. This is followed by a mean shift clustering stage to group together those close transformations in the space. Joining regions associated with transformations of the same group together within each input image forms two large regions sharing similar geometric configuration, which naturally leads to a CVP. To account for the non-Euclidean nature of the matrix Lie group, mean shift vectors are derived in the corresponding Lie algebra vector space with a newly provided effective distance measure. Extensive experiments on single and multiple common object discovery tasks as well as near-duplicate image retrieval verify the robustness and efficiency of the proposed approach.
Linbo Wang 0001, Yanwen Guo 0001, Minh N. Do
IEEE Trans. Image Process.3
2015 Content-Aware Video2Comics With Manga-Style Layout
abstract
We introduce in this paper a new approach that conveniently converts conversational videos into comics with manga-style layout. With our approach, the manga-style layout of a comic page is achieved in a content-driven manner, and the main components, including panels and word balloons, that constitute a visually pleasing comic page are intelligently organized . Our approach extracts key frames on speakers by using a speaker detection technique such that word balloons can be placed near the corresponding speakers. We qualitatively measure the information contained in a comic page. With the initial layout automatically determined, the final comic page is obtained by maximizing such a measure and optimizing the parameters relating to the optimal display of comics. An efficient Markov chain Monte Carlo sampling algorithm is designed for the optimization. Our user study demonstrates that users much prefer our manga-style comics to purely Western style comics. Extensive experiments and comparisons against previous work also verify the effectiveness of our approach.
Guangmei Jing, Yongtao Hu 0001, Yanwen Guo 0001, Yizhou Yu, Wenping Wang 0001
IEEE Trans. Multim.3
2014 A consistent pixel-wise blur measure for partially blurred images
abstract
Despite numerous efforts on blur measurement of partially blurred images, there still lacks an effective blur measure that is both pixel-wise and locally sharp consistent. The paper proposes a novel method with two contributions to overcome this limitation: 1) A new pixel-based blur metric, Multi-resolution Singular Value (MSV), which leverages the average singular value of high frequency bands to measure the blur of each pixel, and 2) a locally continuous strategy, maximum-likelihood estimation (MLE) based refinement, that ensures local continuity by imposing the local sharp consistency on pixel blur in a local correcting process. Experimental results show that our method is effective to smoothly measure the partially blurred images without local discontinuity.
Xianyong Fang, Yanwen Guo 0001, Christian Jacquemin, Jian Zhou 0006, Shanchun Huang
ICIP3
2014 Efficient view manipulation for cuboid-structured images
Yanwen Guo 0001, Guiping Zhang, Zili Lan, Wenping Wang 0001
Comput. Graph.1
2014 Confidence-driven image co-matting
Linbo Wang 0001, Tianchen Xia, Yanwen Guo 0001, Ligang Liu 0001, Jue Wang 0001
Comput. Graph.3
2014 Spectral Analysis on Medial Axis of 2D Shapes
abstract
Abstract Shape analysis finds many important applications in shape understanding, matching and retrieval. Among the various shape analysis methods, spectral shape analysis aims to study the spectrum of the Laplace–Beltrami operator of some well‐designed shape‐dependent equations and obtain a spectral shape descriptor that can in turn be used for shape analysis purposes. The success of such approaches depends greatly on the discriminating power of a shape descriptor. On the other hand, the medial axis of a shape is widely known for its complete shape representation. It is sensitive to small perturbation of the boundary of a shape which often poses difficulty in its effective use for shape analysis. In this paper, we propose a new spectral shape descriptor, called the medial axis spectrum for 2D shapes, which directly applies spectral analysis to the medial axes of the shapes. We extend the Laplace–Beltrami operator onto the medial axis, and take the solution to an extended Laplacian eigenvalue problem defined on the axis as the medial axis spectrum. The medial axis spectrum is robust in the presence of shape boundary noise, and is invariant under rigid transformations, uniform scaling and isometry of the medial axis. We demonstrate these benefits of such a medial axis spectrum representation through extensive experiments. The medial axis spectrum is further used for 2D shape retrieval, and its superiority over previous work is shown by comparison.
Shuiqing He, Yi-King Choi, Yanwen Guo 0001, Wenping Wang 0001
Comput. Graph. Forum3
2014 Object tracking using learned feature manifolds
Yanwen Guo 0001, Weitao Luo, Mingming Liu 0004
Comput. Vis. Image Underst.1
2014 Video Object Co-Segmentation via Subspace Clustering and Quadratic Pseudo-Boolean Optimization in an MRF Framework
abstract
Multiple videos may share a common foreground object, for instance a family member in home videos, or a leading role in various clips of a movie or TV series. In this paper, we present a novel method for co-segmenting the common foreground object from a group of video sequences. The issue was seldom touched on in the literature. Starting from over-segmentation of each video into Temporal Superpixels (TSPs), we first propose a new subspace clustering algorithm which segments the videos into consistent spatio-temporal regions with multiple classes, such that the common foreground has consistent labels across different videos. The subspace clustering algorithm exploits the fact that across different videos the common foreground shares similar appearance features, while motions can be used to better differentiate regions within each video, making accurate extraction of object boundaries easier. We further formulate video object co-segmentation as a Markov Random Field (MRF) model which imposes the constraint of foreground model automatically computed or specified with little user effort. The Quadratic Pseudo-Boolean Optimization (QPBO) is used to generate the results. Experiments show that this video co-segmentation framework can achieve good quality foreground extraction results without user interaction for those videos with unrelated background, and with only moderate user interaction for those videos with similar background. Comparisons with previous work also show the superiority of our approach.
Chuan Wang 0001, Yanwen Guo 0001, Linbo Wang 0001, Wenping Wang 0001
IEEE Trans. Multim.2
2014 Content-Aware Photo Collage Using Circle Packing
abstract
In this paper, we present a novel approach for automatically creating the photo collage that assembles the interest regions of a given group of images naturally. Previous methods on photo collage are generally built upon a well-defined optimization framework, which computes all the geometric parameters and layer indices for input photos on the given canvas by optimizing a unified objective function. The complex nonlinear form of optimization function limits their scalability and efficiency. From the geometric point of view, we recast the generation of collage as a region partition problem such that each image is displayed in its corresponding region partitioned from the canvas. The core of this is an efficient power-diagram-based circle packing algorithm that arranges a series of circles assigned to input photos compactly in the given canvas. To favor important photos, the circles are associated with image importances determined by an image ranking process. A heuristic search process is developed to ensure that salient information of each photo is displayed in the polygonal area resulting from circle packing. With our new formulation, each factor influencing the state of a photo is optimized in an independent stage, and computation of the optimal states for neighboring photos are completely decoupled. This improves the scalability of collage results and ensures their diversity. We also devise a saliency-based image fusion scheme to generate seamless compositive collage. Our approach can generate the collages on nonrectangular canvases and supports interactive collage that allows the user to refine collage results according to his/her personal preferences. We conduct extensive experiments and show the superiority of our algorithm by comparing against previous methods.
Zongqiao Yu, Lin Lu 0001, Yanwen Guo 0001, Rongfei Fan, Mingming Liu 0004, Wenping Wang 0001
IEEE Trans. Vis. Comput. Graph.3
2013 Exposing Blur Kernel from Retouch Image
abstract
The blurring in image comes either from the acquisition noise, or from image editing operation. The produced adverse noise during acquisition need to be eliminated, and the blurring generated by editing should be known in digital forensics, so the blur kernel recovery is significant in community of image processing and computer graphics. In the log-fourier domain, the images before and after blurring bears the isometry, therein an approach of gaussian blur kernel recovery based on Riemannian geodesic is proposed, which evaluates the blurring-invariant quantity between the original image and blurred one, and recovers the blur kernel from blur image. The presented method is testified in images convolved by gaussian blur, median blur, box blur and multiple gaussian blur, and the kernel could be robustly recovered from gaussian blur image. Moreover, the discussed method is capable of authenticating the retouch image generated from blurring.
Zhenlong Du, Xiaoli Li 0005, Yanwen Guo 0001
CAD/Graphics3
2013 Fuzzy quantization based bit transform for low bit-resolution motion estimation
Chuanming Song 0001, Yanwen Guo 0001, Xiang-Hai Wang 0001
Signal Process. Image Commun.2
2012 Improving Photo Composition Elegantly: Considering Image Similarity During Composition Optimization
abstract
Abstract Optimization of images with bad compositions has attracted increasing attention in recent years. Previous methods however seldomly consider image similarity when improving composition aesthetics. This may lead to significant content changes or bring large distortions, resulting in an unpleasant user experience. In this paper, we present a new algorithm for improving image composition aesthetics, while retaining faithful, as much as possible, to the original image content. Our method computes an improved image using a unified model of composition aesthetics and image similarity. The term of composition aesthetics obeys the rule of thirds and aims to enhance image composition. The similarity term in contrast penalizes image difference and distortion caused by composition adjustment. We use an edge‐based measure of structure similarity which nearly coincides with human visual perception to compare the optimized image with the original one. We describe an effective scheme to generate the optimized image with the objective model. Our algorithm is able to produce the recomposed images with minimal visual distortions in an elegant and user controllable manner. We show the superiority of our algorithm by comparing our results with those by previous methods.
Yanwen Guo 0001, Mingming Liu 0004, T. T. Gu, Wenping Wang 0001
Comput. Graph. Forum1
2011 Color-Mood-Aware Clothing Re-texturing
abstract
In this paper, we present a novel color-mood-aware technique to re-texture clothing in a photograph. An efficient classification algorithm is developed to classify clothing textures using color mood scheme. To re-texture the clothing, our approach first computes the gradient maps for the cloth region to be replaced and then calculates the texture distortion coordinates on the projected cloth region according to the gradient maps. After the user selects a target clothing texture from the classified clothing texture database, the lighting and shading effects on the original photograph is transferred using the HSV color space. Experimental results show that the proposed approach successfully re-textures the clothes in photographs while preserving the geometry and lighting features.
Jianbing Shen, Hanqiu Sun, Xiaoyang Mao, Yanwen Guo 0001, Xiaogang Jin 0001
CAD/Graphics4
2011 Multi-keyframe abstraction from videos
abstract
This paper presents a method for abstracting multi-keyframe from video datasets. Existing video abstraction methods focused on simple view videos, and the results will be unacceptable if applied to overlapping views directly due to limitations like unavoidable redundancy and complicated inner correlations. We propose a correlation map to naturally model the correlations with various attributes among multi-keyframe, keyframe importance and weighted correlations are then computed to construct the map. The weighted correlations, unlike the unweighted ones, not only model probabilistic relationship among keyframes but also address the temporal and visual similarity. We facilitate the abstraction process via SVM classification and keyframes reduction using rough set. The multi-keyframe correlation map, which serially assembles event-centered keyframes in temporal order, is presented for displaying the abstraction, which shows the correlations and improves the browsability of video datasets.
Ping Li 0016, Yanwen Guo 0001, Hanqiu Sun
ICIP2
2011 Exploiting feature correspondence constraints for image recognition
abstract
Image recognition is one of the fundamental problems in multimedia analysis. Typically in the training database, there will be more than one image for each object, however most existing bag-of-features based approaches treat them independently and completely ignore the feature correspondence relationship among them. As a result, features corresponding to the same physical point may be clustered into different clusters, which finally leads to inaccurate image representations for recognition. To tackle the problem, we present a supervised codebook construction algorithm exploiting the feature correspondence constraints in feature clustering. Features in different images of the same object are first matched, then ho-mography between images are computed to remove outliers as well as recover the feature correspondences that are not correctly matched. Features belonging to the same physical point are enforced to be in the same cluster. We show via experiments that codebook constructed using this approach can improve the recognition performance.
Linbo Wang 0001, Yanwen Guo 0001, Suk Hwan Lim, Nelson L. Chang
ICIP3
2011 Content-sensitive collection snapping
abstract
Interactive segmentation methods have greatly simplified the task of object cutout from an image. However, segmenting a large number of images in a collection is still a tedious task. In this paper, we present a content-sensitive group segmentation method that iteratively segments the images in a collection and incrementally refines the results. With our method, a user only provides a small number of strokes to segment and refine a few sample images. For each of the rest of images, our method finds relevant sample images and applies the corresponding appearance models to guide the segmentation. To improve the segmentation results using the user strokes on a few images with unsatisfactory segmentation results, our method calculates the relevance map that measures the probability that a stroke can be appropriately applied at each pixel/region of an image, and applies it accordingly. Our experiments show that our method can effectively segment an image collection with a wide variety of image content and significantly reduce user input.
Yanwei Fu 0001, Yanwen Guo 0001
ICME2
2010 Discriminative Nonorthogonal Binary Subspace Tracking
Yanwen Guo 0001
ECCV (3)3
2010 Multi-View Video Summarization
abstract
Previous video summarization studies focused on monocular videos, and the results would not be good if they were applied to multi-view videos directly, due to problems such as the redundancy in multiple views. In this paper, we present a method for summarizing multi-view videos. We construct a spatio-temporal shot graph and formulate the summarization problem as a graph labeling task. The spatio-temporal shot graph is derived from a hypergraph, which encodes the correlations with different attributes among multi-view video shots in hyperedges. We then partition the shot graph and identify clusters of event-centered shots with similar contents via random walks. The summarization result is generated through solving a multi-objective optimization problem based on shot importance evaluated using a Gaussian entropy fusion scheme. Different summarization objectives, such as minimum summary length and maximum information coverage, can be accomplished in the framework. Moreover, multi-level summarization can be achieved easily by configuring the optimization parameters. We also propose the multi-view storyboard and event board for presenting multi-view summaries. The storyboard naturally reflects correlations among multi-view summarized shots that describe the same important event. The event-board serially assembles event-centered multi-view shots in temporal order. Single video summary which facilitates quick browsing of the summarized multi-view video can be easily generated based on the event board representation.
Yanwei Fu 0001, Yanwen Guo 0001, Yanshu Zhu, Feng Liu 0015, Chuanming Song 0001, Zhi-Hua Zhou
IEEE Trans. Multim.2
2009 Binary Alpha-Plane Assisted Fast Motion Estimation of Video Objects in Wavelet Domain
abstract
Summary form only given. Shift-variance and computational complexity are bottleneck of existing wavelet-based motion estimation (ME). Moreover, to the best of our knowledge, few works have been reported on wavelet-domain ME of video objects (VOs). In this paper, we present an efficient wavelet-domain approach to ME of arbitrarily shaped VOs.
Chuanming Song 0001, Xiang-Hai Wang 0001, Yanwen Guo 0001, Fuyan Zhang
DCC3
2009 Pores-Preserving Face Cleaning Based on Improved Empirical Mode Decomposition
Yanli Liu 0006, Xiao-Gang Xu, Yanwen Guo 0001, Xin Duan, Qunsheng Peng 0001
J. Comput. Sci. Technol.3
2009 Image Retargeting Using Mesh Parametrization
abstract
Image retargeting aims to adapt images to displays of small sizes and different aspect ratios. Effective retargeting requires emphasizing the important content while retaining surrounding context with minimal visual distortion. In this paper, we present such an effective image retargeting method using saliency-based mesh parametrization. Our method first constructs a mesh image representation that is consistent with the underlying image structures. Such a mesh representation enables easy preservation of image structures during retargeting since it captures underlying image structures. Based on this mesh representation, we formulate the problem of retargeting an image to a desired size as a constrained image mesh parametrization problem that aims at finding a homomorphous target mesh with desired size. Specifically, to emphasize salient objects and minimize visual distortion, we associate image saliency into the image mesh and regard image structure as constraints for mesh parametrization. Through a stretch-based mesh parametrization process we obtain the homomorphous target mesh, which is then used to render the target image by texture mapping. The effectiveness of our algorithm is demonstrated by experiments.
Yanwen Guo 0001, Feng Liu 0015, Zhi-Hua Zhou, Michael Gleicher
IEEE Trans. Multim.1
2008 Constrained sampling for image retargeting
abstract
In this paper, we present a new approach for retargeting large images to mobile devices with small screens. As the core of image retargeting, information fidelity is adequately considered in terms of reservations of salient regions, edge integrity, and image layout. By taking these aspects as constraints, image retargeting is formulated as a constrained sampling task. Each pixel in image is first represented with a vector encoding the constraints. Then, pixels with the same vector values combine to form blocks, and the original image is thus converted into a graph representation. Thereafter, the sampling ratio of each block is determined with a balanced minimum cost flow algorithm. Final result is generated by an interpolated sampling scheme and direct scaling. Experiments demonstrate the effectiveness of the proposed approach.
Tongwei Ren, Yanwen Guo 0001, Gangshan Wu, Fuyan Zhang
ICME2
2008 A Robust and Fast Non-Local Means Algorithm for Image Denoising
Yanli Liu 0006, Yanwen Guo 0001, Qunsheng Peng 0001
J. Comput. Sci. Technol.4
2008 Mesh-Guided Optimized Retexturing for Image and Video
abstract
This paper presents an approach of replacing textures of specified regions in the input image and video using stretch-based mesh optimization.. The retexturing results have the similar distortion and shading effects conforming to the underlying geometry and lighting conditions. For replacing textures in single image,two important steps are developed: the stretch-based mesh parametrization incorporating the recovered normal information is deduced to imitate perspective distortion of the region of interest; the Poisson-based refinement process is exploited to account for texture distortion at fine scale. The luminance of the input image is preserved through color transfer in YCbCr color space. Our approach is independent of the replaced textures. Once the input image is processed, any new texture can be applied to efficiently generate the retexturing results. For video retexturing, we propose key-frame-based texture replacement extended and generalized from the image retexturing. Our approach repeatedly propagates the replacement result of key frame to the rest of the frames. We develop the local motion optimization scheme to deal with the inaccuracies and errors of robust optical flow when tracking moving objects. Visibility shifting and texture drifting are effectively alleviated using graphcut segmentation algorithm and the global optimization to smooth trajectories of the tracked points over temporal domain. Our experimental results showed that the proposed approach can generate visually pleasing results for both image and video.
Yanwen Guo 0001, Hanqiu Sun, Qunsheng Peng 0001, Zhongding Jiang
IEEE Trans. Vis. Comput. Graph.1
2007 A Robust and Fast Non-local Means Algorithm for Image Denoising
abstract
In the paper, we propose a robust and fast image denoising method. The approach integrates both Non-Local means algorithm and Laplacian Pyramid. Given an image to be denoised, we first decompose it into Laplacian pyramid. Exploiting the redundancy property of Laplacian pyramid, we then perform non-local means on every level image of Laplacian pyramid. Essentially, we use the similarity of image features in Laplacian pyramid to act as weight to denoise image. Since the features extracted in Laplacian pyramid are localized in spatial position and scale, they are much more able to describe image, and computing the similarity between them is more reasonable and more robust. Also, based on the efficient Summed Square Image (SSI) scheme and Fast Fourier Transform (FFT), we present an accelerating algorithm to break the bottleneck of non-local means algorithm - similarity computation of compare windows. After speedup, our algorithm is fifty times faster than original non-local means algorithm. Experiments demonstrated the effectiveness of our algorithm.
Yanli Liu 0006, Chen Xi, Yanwen Guo 0001, Qunsheng Peng 0001
CAD/Graphics4
2007 Learning-based 3D face detection using geometric context
abstract
Abstract In computer graphics community, face model is one of the most useful entities. The automatic detection of 3D face model has special significance to computer graphics, vision, and human‐computer interaction. However, few methods have been dedicated to this task. This paper proposes a machine learning approach for fully automatic 3D face detection. To exploit the facial features, we introducegeometric context, a novel shape descriptor which can compactly encode the distribution of local geometry and can be evaluated efficiently by using a new volume encoding form, namedintegral volume. Geometric contexts over 3D face offer the rich and discriminative representation of facial shapes and hence are quite suitable to classification. We adopt an AdaBoost learning algorithm to select the most effective geometric context‐based classifiers and to combine them into a strong classifier. Given an arbitrary 3D model, our method first identifies the symmetric parts as candidates with a new reflective symmetry detection algorithm. Then uses the learned classifier to judge whether the face part exists. Experiments are performed on a large set of 3D face and non‐face models and the results demonstrate high performance of our method. Copyright © 2007 John Wiley & Sons, Ltd.
Yanwen Guo 0001, Fuyan Zhang, Hanqiu Sun, Qunsheng Peng 0001
Comput. Animat. Virtual Worlds1
2007 Image completion based on views of large displacement
Yanwen Guo 0001, Liang Pan, Qunsheng Peng 0001, Fuyan Zhang
Vis. Comput.2
2006 Automatic Foreground Extraction of Head Shoulder Images
Yiting Ying, Yanwen Guo 0001, Qunsheng Peng 0001
Computer Graphics International3
2006 Fast Non-Local Algorithm for Image Denoising
abstract
For the non-local denoising approach presented by Buades et al., remarkable denoising results are obtained at high expense of computational cost. In this paper, a new algorithm that reduces the computational cost for calculating the similarity of neighborhood windows is proposed. We first introduce an approximate measure about the similarity of neighborhood windows, then we use an efficient summed square image (SSI) scheme and fast Fourier transform (FFT) to accelerate the calculation of this measure. Our algorithm is about fifty times faster than the original non-local algorithm both theoretically and experimentally, yet produces comparable results in terms of mean-squared error (MSE) and perceptual image quality.
Yanwen Guo 0001, Yiting Ying, Yanli Liu 0006, Qunsheng Peng 0001
ICIP2
2005 A New Constrained Texture Mapping Method
Yanwen Guo 0001, Xiufen Cui, Qunsheng Peng 0001
ICEC1
2005 A novel constrained texture mapping method based on harmonic map
Yanwen Guo 0001, Hanqiu Sun, Xiufen Cui, Qunsheng Peng 0001
Comput. Graph.1
2005 Image and video retexturing
abstract
Abstract We propose a novel image/video retexturing approach that preserves the original shading effects without knowing the underlying surface and lighting conditions. For static images, we first introduce the Poisson equation‐based algorithm to simulate the texture distortion on the projected interest region of the underlying surface, while preserving the shading effect of the original image. We further work on videos by retexturing the key frame as static image and then propagating the results onto the other frames. In video retexturing, we have introduced the mesh based optimization for object tracking to avoid texture drifting, and the graph cut algorithm to effectively deal with visibility shift between frames. The graph cut algorithm is applied on a trimap along the boundary of the object to extract the textured part inside the trimap. The proposed approach is developed in image/video retexturing at nearly interactive rate, and our experimental results have showed the satisfactory performance of our approach. Copyright © 2005 John Wiley & Sons, Ltd.
Yanwen Guo 0001, Zhongyi Xie, Hanqiu Sun, Qunsheng Peng 0001
Comput. Animat. Virtual Worlds1