VLDB 2026 Research / reviewers in the wild / expert
Qi Zhang 0029
dblp:52/323-29
· DBLP profile ↗
55ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0001-9611-6697ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 45 · 7 first-author · 37 since 2021Artificial intelligence and machine learning · 36 · 7 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RHINO: regularizing the hash-based implicit neural representation
Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Sci. China Inf. Sci. | 3 |
| 2026 | Multi-Reference Frames-Based Occupancy Information Difference and Correlation Entropy Coding for MPEG G-PCCabstractEfficient compression of massive 3D point clouds remains challenging under limited storage and bandwidth. The Moving Picture Experts Group (MPEG) has developed the geometry-based point cloud compression (G-PCC) standard and is currently at the FDIS stage, referred to as Enhanced G-PCC. This paper proposes a novel inter prediction attribute coding based on the Region Adaptive Hierarchical Transform (RAHT) and entropy coding framework to enhance the coding efficiency. 1)Multi-Reference Frames: A low-delay prediction architecture is introduced to improve temporal correlation utilization, which has not yet been explored in G-PCC inter RAHT attribute compression, with a lightweight motion-based switch to disable multi-reference prediction for large motion. 2)Occupancy Information Difference: The inter eligibility based on occupancy information difference scheme is employed to select valid inter reference blocks, while hybrid prediction scheme based on occupancy information difference is further introduced to enhance prediction accuracy. 3)Correlation Entropy Coding: A new entropy coder is developed to exploit the intrinsic correlation among color components. By jointly leveraging these techniques, the proposed method efficiently exploits spatial and temporal regularities of dynamic point clouds, achieving significant performance gains. Experimental results show that it outperforms the state-of-the-art Enhanced G-PCC reference software (TMC13-v31), with average coding gains of 3.17% for reflectance on the Cat3 dataset and 2.81%, 16.92%, and 12.21% for the Luma, Cb, and Cr components on the Cat2 dataset under MPEG common test conditions. The method,Occupancy Information Difference, has already been adopted into the MPEG Enhanced G-PCC standard. Qi Zhang 0029, Ge Li 0002, Hongyu Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | RelightFlow: An Inversion-Free Video Relighting Model via Dual-Trajectory Diffusion EditingabstractVideo relighting is a fundamental task with wide-ranging applications in contemporary visual computing, including film production, immersive virtual reality, augmented reality, and interactive digital worlds. Its objective is to generate temporally stable lighting effects while preserving the structural integrity, visual appearance, and intrinsic physical properties of objects in the source video. Recent studies combined image relighting models with video diffusion models, achieving notable progress in training-free video relighting. However, these methods rely on mapping noisy latents back to the pixel space during the relighting process, leading to degraded fidelity, consistency, and stability. In this work, we propose RelightFlow, a training-free video relighting framework built upon a flow-matching-based video DiT model, which requires neither inversion nor additional training, obtaining more feasible and robust relighting effects. Specifically, we first design a detail-preserving relighting module coupled with an interpolation trajectory, which injects stable lighting cues in the early generation stages while maintaining fine spatial details. Next, we develop a temporally consistent relighting module to form a flow-editing trajectory, leveraging the velocity fields predicted by the video DiT model to enhance temporal coherence. Finally, we introduce a dynamic fusion strategy that adaptively integrates these two trajectories to balance relighting intensity and temporal stability. Extensive experiments demonstrate that RelightFlow achieves high-quality video relighting with stable relighting intensity, superior fidelity, and temporal consistency. The code are provided in https://github.com/Yukun66/RelightFlow. Qi Zhang 0029, Longguang Wang, Yulan Guo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Lossless Dynamic Point Cloud Geometry Compression via Rate-Distortion Optimized Motion EstimationabstractDynamic point clouds are valid representations of three-dimensional moving entities in diverse application scenarios. The comprehensive redundancy in uncompressed point clouds necessitates efficient compression methods. Motion estimation (ME) plays a crucial role in eliminating the temporal redundancy of point cloud sequences. However, existing ME methods suffer from inaccurate geometry compensation distortion measures and imbalanced rate-distortion modeling, significantly impacting the coding performance. To address these challenges, we propose a rate-distortion (R-D) optimized ME scheme for dynamic point cloud geometry compression. First, a point cloud is hierarchically decomposed into a set of macroblocks and blocks via the octree, and variable block-size ME is introduced to capture complex local motions in point clouds. Second, we propose a block-matching motion search method that integrates a new geometry compensation distortion measure, combining voxel-wise Hamming distance and point-wise Euclidean distance as a joint criterion to improve the accuracy of the ME. To mitigate complexity, local motion features are extracted to efficiently accelerate the geometry distortion measure. Third, a rate-distortion model with an adaptive Lagrange multiplier is designed to facilitate better selection of inter predictors in the motion decision stage, thereby achieving efficient coding performance. Experimental results demonstrate the improvements of our proposed scheme over rival platforms in coding performance and computational efficiency. Additional results also validate the effectiveness of key modules within the proposed R-D optimized ME framework. Qi Zhang 0029, Yiting Shao, Lixuan Meng, Shan Liu 0001, Ge Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Zero-Pose-Prior NeRF: Recursive Radiance Field Reconstruction From Unposed and Unordered ImagesabstractThe dependence of neural radiance fields (NeRF) on accurate camera poses has emerged as a critical obstacle to their widespread real-world applications. While recent advances have demonstrated the potential for simultaneously addressing camera registration and scene reconstruction, these methods inherently rely on reasonable initialization derived from pose or scene priors and struggle with complex scenes involving large camera motions, particularly in unordered 360-degree scenes. In this work, we propose Zero-Pose-Prior NeRF to recover radiance fields from unposed and unordered image collections without any prior knowledge. Our key insight is to decompose this complex problem into smaller sub-problems, wherein the sub-problems' camera poses are initially estimated to provide self-bootstrapping priors for the global pose estimation, followed by a recursive registration and reconstruction. To achieve this, we first perform scene partitioning to establish a hierarchical structure that describes registration order from local to global. Thereafter, we devise a conditionally-decoupled positional encoding for NeRFs, which serves as the basic model for camera pose estimation and scene representation. Following this, we develop a recursive registration to recursively estimate the poses of local scenes and register them into a unified global pose space, ultimately enabling the reconstruction of the entire scene. Experiments on real-world scenes show that our approach outperforms the state-of-the-art pose-free methods in terms of accurate camera poses and robust radiance field reconstruction, resulting in high-fidelity view synthesis. Xinxin Liu 0020, Qi Zhang 0029, Xue Wang 0006, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Image Process. | 2 |
| 2025 | Mani-GS: Gaussian Splatting Manipulation with Triangular MeshabstractNeural 3D representations, such as Neural Radiation Fields (NeRF), excel at producing photorealistic rendering results but lack the flexibility for manipulation and editing which is crucial for content creation. However, manipulating NeRF is not highly controllable and requires a long training and inference time. With the emergence of 3D Gaussian Splatting (3DGS), extremely high-fidelity novel view synthesis can be achieved using an explicit point-based 3D representation with much faster training and rendering speed. However, there is still a lack of effective means to manipulate 3DGS freely while maintaining rendering quality. In this work, we aim to tackle the challenge of achieving manipulable photo-realistic rendering. We propose to utilize a triangular mesh to manipulate 3DGS directly with self-adaptation. This approach reduces the need to design various algorithms for different types of 3DGS manipulation. By utilizing a triangle shape-aware Gaussian binding and adapting method, we can achieve 3DGS manipulation and preserve high-fidelity rendering. In addition, our method is also effective with inaccurate meshes extracted from 3DGS. Experiments demonstrate our method’s effectiveness and superiority over baseline approaches. Xiangjun Gao, Xiaoyu Li 0002, Yiyu Zhuang, Qi Zhang 0029, Wenbo Hu 0002, Chaopeng Zhang, Yao Yao 0008, Ying Shan, Long Quan |
CVPR | 4 |
| 2025 | Mitigating Ambiguities in 3D Classification with Gaussian Splattingabstract3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective objects. To address these issues, we propose Gaussian Splatting (GS) point cloud-based 3D classification. We find that the scale and rotation coefficients in the GS point cloud help characterize surface types. Specifically, wire-like surfaces consist of multiple slender Gaussian ellipsoids, while flat surfaces are composed of a few flat Gaussian ellipsoids. Additionally, the opacity in the GS point cloud represents the transparency characteristics of objects. As a result, ambiguities in point cloud-based 3D classification can be mitigated utilizing GS point cloud as input. To verify the effectiveness of GS point cloud input, we construct the first real-world GS point cloud dataset in the community, which includes 20 categories with 200 objects in each category. Experiments not only validate the superiority of GS point cloud input, especially in distinguishing ambiguous objects, but also demonstrate the generalization ability across different classification methods. Our project page: https://ruiqi-nju.github.io/MACGS. Hao Zhu 0004, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 4 |
| 2025 | Rate-Distortion Optimized Motion Estimation for Dynamic Point Cloud Geometry CompressionabstractDynamic point clouds serve as crucial representations of three-dimensional moving entities across diverse applications. The substantial amount of data in point clouds necessitates the development of efficient compression techniques. Motion estimation (ME) plays a crucial role in eliminating the temporal redundancy of point cloud sequences. However, prevailing ME methods suffer from the inaccurate geometry distortion measure and the imbalanced rate-distortion modeling, significantly impacting the coding performance. To address these challenges, we propose a rate-distortion (R-D) optimized ME scheme for dynamic point cloud geometry compression. Qi Zhang 0029, Yiting Shao, Lixuan Meng, Hailong Jiao, Shan Liu 0001, Ge Li 0002 |
DCC | 1 |
| 2025 | CoMPaSS: Enhancing Spatial Understanding in Text-to-Image Diffusion ModelsabstractText-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning spatial relationships in existing datasets, and 2) the inability of current text encoders to accurately interpret the spatial semantics of input descriptions. We propose CoMPaSS, a versatile framework that enhances spatial understanding in T2I models. It first addresses data ambiguity with the Spatial Constraints-Oriented Pairing (SCOP) data engine, which curates spatially-accurate training data via principled constraints. To leverage these priors, CoMPaSS also introduces the Token ENcoding ORdering (TENOR) module, which preserves crucial token ordering information lost by text encoders, thereby reinforcing the prompt's linguistic structure. Extensive experiments on four popular T2I models (UNet and MMDiT-based) show CoMPaSS sets a new state of the art on key spatial benchmarks, with substantial relative gains on VISOR (+98%), T2I-CompBench Spatial (+67%), and GenEval Position (+131%). Code is available at https://github.com/blurgyy/CoMPaSS. Gaoyang Zhang, Bingtao Fu, Qingnan Fan, Qi Zhang 0029, Runxing Liu, Huaqi Zhang, Xinguo Liu |
ICCV | 4 |
| 2025 | Overfitted Point Cloud Attribute Codec Using Sparse Hierarchical Implicit Neural RepresentationsabstractCompressing attributes of 3D point clouds remains challenging due to their inherent sparsity and irregular distribution. To address this, we propose an efficient framework based on sparse hierarchical Implicit Neural Representations (INRs). Specifically, we introduce a novel vertex-based INR framework, which integrates interpolation to enable accurate and compact implicit representations of point cloud attributes. To effectively capture the varying importance of latent features, we design an adaptive quantization scheme. Furthermore, we develop efficient level-wise entropy models to exploit dependencies within and across hierarchical levels. Finally, point cloud attributes are reconstructed from concatenated multi-resolution latent representations via a sparse convolution-based reconstruction module. Experimental results demonstrate that our approach significantly outperforms previous INR-based methods, achieving superior performance compared to the latest G-PCC (TMC13v28) standard and state-of-the-art learning-based methods. Qi Zhang 0029, Shan Liu 0001, Ge Li 0002 |
ACM Multimedia | 3 |
| 2025 | UV Gaussians: Joint learning of mesh deformation and Gaussian textures for human avatar modeling
Yujiao Jiang, Qingmin Liao, Xiaoyu Li 0002, Qi Zhang 0029, Chaopeng Zhang, Zongqing Lu 0001, Ying Shan |
Knowl. Based Syst. | 5 |
| 2025 | GUS-IR: Gaussian Splatting With Unified Shading for Inverse RenderingabstractRecovering the intrinsic physical attributes of a scene from images, generally termed as the inverse rendering problem, has been a central and challenging task in computer vision and computer graphics. In this paper, we present GUS-IR, a novel framework designed to address the inverse rendering problem for complicated scenes featuring rough and glossy surfaces. This paper starts by analyzing and comparing two prominent shading techniques popularly used for inverse rendering, forward shading and deferred shading, effectiveness in handling complex materials. More importantly, we propose a unified shading solution that combines the advantages of both techniques for better decomposition. In addition, we analyze the normal modeling in 3D Gaussian Splatting (3DGS) and utilize the shortest axis as normal for each particle in GUS-IR, along with a depth-related regularization, resulting in improved geometric representation and better shape reconstruction. Furthermore, we enhance the probe-based baking scheme proposed by GS-IR to achieve more accurate ambient occlusion modeling to better handle indirect illumination. Extensive experiments have demonstrated the superior performance of GUS-IR in achieving precise intrinsic decomposition and geometric representation, supporting many downstream tasks (such as relighting, retouching) in computer vision, graphics, and extended reality. Zhihao Liang 0002, Hongdong Li, Kui Jia, Kailing Guo, Qi Zhang 0029 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | IR-Pro: Baking Probes to Model Indirect Illumination for Inverse Rendering of ScenesabstractModeling indirect illumination to handle global illumination and decompose materials from multi-view images is challenging, especially in complex scenes with self-occlusion. While recent implicit neural representations show promise in inverse rendering, they struggle with efficient and effective modeling of indirect illumination. Besides, real-time global illumination techniques (e.g. Indirect Lighting Cache) have been successful in gaming. Inspired by this, we present a novel three-stageInverseRenderer withProbes (IR-Pro), which efficiently caches occlusion to handle indirect illumination. Experiments demonstrate the superiority of IR-Pro over existing methods in the inverse rendering of complex scenes. Furthermore, we successfully integrate the results into digital content creation software and showcase their effectiveness in applications, like relighting, simulation, and editing. Zhihao Liang 0002, Qi Zhang 0029, Yirui Guan, Kui Jia |
IEEE Signal Process. Lett. | 2 |
| 2025 | HumanRef-GS: Image-to-3D Human Generation With Reference-Guided Diffusion and 3D Gaussian SplattingabstractGenerating a 3D human model from a single reference image is a challenging task as it involves inferring textures and geometries in unseen views while maintaining consistency with the reference image. Existing methods that rely on 3D generative models are limited by the availability of 3D training data. Optimization-based approaches that distill text-to-image diffusion models into 3D models often struggle to preserve the intricate texture details of the reference image, resulting in inconsistent appearances across different views. In this paper, we propose HumanRef-GS, a novel method for single image-to-3D clothed human generation based on 3D Gaussian Splatting (3DGS). To ensure the generated 3D model is both photorealistic and consistent with the input image, HumanRef-GS employs a unique technique called reference-guided score distillation sampling (Ref-SDS). This method effectively incorporates image guidance into the generation process, enhancing the quality of the results. Additionally, we introduce region-aware attention to Ref-SDS, which ensures accurate correspondence between different body regions. To mitigate the impact of view dependence in 3DGS and enhance the view-consistency of the generated results, we substitute the anisotropic Gaussians in the vanilla representation with isotropic Gaussians. By utilizing the 3D Gaussian representation, our method significantly enhances the generation efficiency and rendering speed of 3D clothed human models. This improvement allows for faster and more efficient generation of high-quality results. Experimental results demonstrate that HumanRef-GS surpasses state-of-the-art methods in generating 3D clothed humans with fine geometry, photorealistic textures, and view-consistent appearances. We are committed to making our code and model available upon acceptance for further research and exploration. Jingbo Zhang 0002, Xiaoyu Li 0002, Hongliang Zhong, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Jing Liao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF: Radiance Fields Reconstruction for Two-Hand-Held ObjectsabstractOur work aims to reconstruct the appearance and geometry of the two-hand-held object from a sequence of color images. In contrast to traditional single-hand-held manipulation, two-hand-holding allows more flexible interaction, thereby providing back views of the object, which is particularly convenient for reconstruction but generates complex view-dependent occlusions. The recent development of neural rendering provides new potential for hand-held object reconstruction. In this paper, we propose a novel neural representation-based framework to recover radiance fields of the two-hand-held object, named ${\rm{H}}_{2}{\rm{O}}$H2O-NeRF. We first design an object-centric semantic module based on the geometric signed distance function cues to predict 3D object-centric regions and develop the view-dependent visible module based on the image-related cues to label 2D occluded regions. We then combine them to obtain a 2D visible mask that adaptively guides ray sampling on the object for optimization. We also provide a newly collected ${\rm{H}}_{2}{\rm{O}}$H2O dataset to validate the proposed method. Experiments show that our method achieves superior performance on reconstruction completeness and view-consistency synthesis compared to the state-of-the-art methods. Xinxin Liu 0020, Qi Zhang 0029, Xin Huang 0021, Guoqing Zhou 0003, Qing Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | A Pre-convolved Representation for Plug-and-Play Neural Illumination FieldsabstractRecent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light modeling without a solid basis, can lead to an undesirable decomposition between lighting and material. To address this, we propose a fully differentiable framework named Neural Illumination Fields (NeIF) that uses radiance fields as a lighting model to handle complex lighting in a physically based way. Together with integral lobe encoding for roughness-adaptive specular lobe and leveraging the pre-convolved background for accurate decomposition, the proposed method represents a significant step towards integrating physically based rendering into the NeRF representation. The experiments demonstrate the superior performance of novel-view rendering compared to previous works, and the capability to re-render objects under arbitrary NeRF-style environments opens up exciting possibilities for bridging the gap between virtual and real-world scenes. Yiyu Zhuang, Qi Zhang 0029, Xuan Wang 0009, Hao Zhu 0004, Xiaoyu Li 0002, Ying Shan, Xun Cao |
AAAI | 2 |
| 2024 | ConTex-Human: Free-View Rendering of Human from a Single Image with Texture-Consistent SynthesisabstractIn this work, we propose a method to address the chal-lenge of rendering a 3D human from a single image in a free-view manner. Some existing approaches could achieve this by using generalizable pixel-aligned implicit fields to reconstruct a textured mesh of a human or by employing a 2D diffusion model as guidance with the Score Distillation Sampling (SDS) method, to lift the 2D image into 3D space. However, a generalizable implicit field often results in an over-smooth texture field, while the SDS method tends to lead to a texture-inconsistent novel view with the input image. In this paper, we introduce a texture-consistent back view synthesis module that could transfer the reference im-age content to the back view through depth and text-guided attention injection. Moreover, to alleviate the color distortion that occurs in the side region, we propose a visibility-aware patch consistency regularization for texture mapping and refinement combined with the synthesized back view texture. With the above techniques, we can achieve high-fidelity and texture-consistent human rendering from a single image. Experiments conducted on both real and synthetic data demonstrate the effectiveness of our method and show that our approach outperforms previous baseline methods. Xiangjun Gao, Xiaoyu Li 0002, Chaopeng Zhang, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Long Quan |
CVPR | 4 |
| 2024 | HumanNorm: Learning Normal Diffusion Model for High-quality and Realistic 3D Human GenerationabstractRecent text-to-3D methods employing diffusion models have made significant advancements in 3D human generation. However, these approaches face challenges due to the limitations of text-to-image diffusion models, which lack an understanding of 3D structures. Consequently, these methods struggle to achieve high-quality human generation, resulting in smooth geometry and cartoon-like appearances. In this paper, we propose HumanNorm, a novel approach for high-quality and realistic 3D human generation. The main idea is to enhance the model's 2D perception of 3D geometry by learning a normal-adapted diffusion model and a normal-aligned diffusion model. The normal-adapted diffusion model can generate high-fidelity normal maps corresponding to user prompts with view-dependent and body-aware text. The normal-aligned diffusion model learns to generate color images aligned with the normal maps, thereby transforming physical geometry details into realistic appearance. Leveraging the proposed normal diffusion model, we devise a progressive geometry generation strategy and a multi-step Score Distillation Sampling (SDS) loss to enhance the performance of 3D human generation. Comprehensive experiments substantiate HumanNorm's ability to generate 3D humans with intricate geometry and realistic appearances. HumanNorm outperforms existing text-to-3D methods in both geometry and texture quality. The project page of HumanNorm is https://humannorm.github.io/. Xin Huang 0021, Ruizhi Shao, Qi Zhang 0029, Hongwen Zhang 0001, Yebin Liu, Qing Wang 0006 |
CVPR | 3 |
| 2024 | GS-IR: 3D Gaussian Splatting for Inverse RenderingabstractWe propose GS-IR, a novel inverse rendering approach based on 3D Gaussian Splatting (3DGS) that leverages forward mapping volume rendering to achieve photorealistic novel view synthesis and relighting results. Unlike previous works that use implicit neural representations and volume rendering (e.g. NeRF), which suffer from low expressive power and high computational complexity, we extend 3DGS, a top-performance representation for novel view synthesis, to estimate scene geometry, surface material, and environment illumination from multi-view images captured under unknown lighting conditions. There are two main problems when introducing 3DGS to inverse rendering: 1) 3DGS does not support producing plausible normal natively; 2) forward mapping (e.g. rasterization and splatting) cannot trace the occlusion like backward mapping (e.g. ray tracing). To address these challenges, our GS-IR proposes an efficient optimization scheme incorporating a depth-derivation-based regularization for normal estimation and a baking-based occlusion to model indirect lighting. The flexible and expressive 3DGS representation allows us to achieve fast and compact geometry reconstruction, photore-alistic novel view synthesis, and effective physically-based rendering. We demonstrate the superiority of our method over baseline methods through qualitative and quantitative evaluations of various challenging scenes. The source code is available at https://github.com/lzhnb/GS-IR. Zhihao Liang 0002, Qi Zhang 0029, Ying Shan, Kui Jia |
CVPR | 2 |
| 2024 | FINER: Flexible Spectral-Bias Tuning in Implicit NEural Representation by Variableperiodic Activation FunctionsabstractImplicit Neural Representation (INR), which utilizes a neural network to map coordinate inputs to corresponding attributes, is causing a revolution in the field of signal processing. However, current INR techniques suffer from a re-stricted capability to tune their supported frequency set, re-sulting in imperfect performance when representing complex signals with multiple frequencies. We have identified that this frequency-related problem can be greatly alleviated by introducing variableperiodic activation functions, for which we propose FINER. By initializing the bias of the neural network within different ranges, sub-functions with various frequencies in the variableperiodic function are selected for activation. Consequently, the supported frequency set of FINER can be flexibly tuned, leading to improved performance in signal representation. We demon-strate the capabilities of FINER in the contexts of2D image fitting, 3D signed distance field representation, and 5D neural radiance fields optimization, and we show that it outper-forms existing INRs. Zhen Liu 0031, Hao Zhu 0005, Qi Zhang 0029, Jingde Fu, Weibing Deng, Zhan Ma 0001, Yanwen Guo 0001, Xun Cao |
CVPR | 3 |
| 2024 | HumanRef: Single Image to 3D Human Generation via Reference-Guided DiffusionabstractGenerating a 3D human model from a single reference image is challenging because it requires inferring textures and geometries in invisible views while maintaining consistency with the reference image. Previous methods utilizing 3D generative models are limited by the availability of 3D training data. Optimization-based methods that lift text-to-image diffusion models to 3D generation often fail to preserve the texture details of the reference image, resulting in inconsistent appearances in different views. In this paper, we propose HumanRef, a 3D human generation framework from a single-view input. To ensure the generated 3D model is photorealistic and consistent with the input image, HumanRef introduces a novel method called reference-guided score distillation sampling (Ref-SDS), which effectively incorporates image guidance into the generation process. Furthermore, we introduce region-aware attention to Ref-SDS, ensuring accurate correspondence between different body regions. Experimental results demonstrate that HumanRef outper-forms state-of-the-art methods in generating 3D clothed humans with fine geometry, photorealistic textures, and view-consistent appearances. Code and model are available at https./reckcrtrhang.github.io/HumanRef.github.io/. Jingbo Zhang 0002, Xiaoyu Li 0002, Qi Zhang 0029, Yan-Pei Cao 0001, Ying Shan, Jing Liao 0001 |
CVPR | 3 |
| 2024 | Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360$^\circ $
Yuxiao He, Yiyu Zhuang, Yao Yao 0008, Siyu Zhu 0001, Xiaoyu Li 0002, Qi Zhang 0029, Xun Cao, Hao Zhu 0004 |
ECCV (56) | 7 |
| 2024 | Analytic-Splatting: Anti-Aliased 3D Gaussian Splatting via Analytic Integration
Zhihao Liang 0002, Qi Zhang 0029, Wenbo Hu 0002, Lei Zhu 0016, Kui Jia |
ECCV (17) | 2 |
| 2024 | Neural Poisson Solver: A Universal and Continuous Framework for Natural Signal Blending
Delong Wu, Hao Zhu 0005, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
ECCV (80) | 3 |
| 2024 | Physically Plausible Color Correction for Neural Radiance Fields
Qi Zhang 0029, Hongdong Li |
ECCV (44) | 1 |
| 2024 | LTM-NeRF: Embedding 3D Local Tone Mapping in HDR Neural Radiance FieldabstractRecent advances in Neural Radiance Fields (NeRF) have provided a new geometric primitive for novel view synthesis. High Dynamic Range NeRF (HDR NeRF) can render novel views with a higher dynamic range. However, effectively displaying the scene contents of HDR NeRF on diverse devices with limited dynamic range poses a significant challenge. To address this, we present LTM-NeRF, a method designed to recover HDR NeRF and support 3D local tone mapping. LTM-NeRF allows for the synthesis of HDR views, tone-mapped views, and LDR views under different exposure settings, using only the multi-view multi-exposure LDR inputs for supervision. Specifically, we propose a differentiable Camera Response Function (CRF) module for HDR NeRF reconstruction, globally mapping the scene's HDR radiance to LDR pixels. Moreover, we introduce a Neural Exposure Field (NeEF) to represent the spatially varying exposure time of an HDR NeRF to achieve 3D local tone mapping, for compatibility with various displays. Comprehensive experiments demonstrate that our method can not only synthesize HDR views and exposure-varying LDR views accurately but also render locally tone-mapped views naturally. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the expressive power of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be greatly solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Furthermore, the expressive power of the DINER is determined by the width of the hash-table. Different width corresponds to different geometrical elements in the attribute space, e.g., 1D curve, 2D curved-plane and 3D curved-volume when the width is set as 1, 2 and 3, respectively. More covered areas of the geometrical elements result in stronger expressive power. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP versus SIREN) and various tasks (image/video representation, phase retrieval, refractive index recovery, and neural radiance field optimization) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Hao Zhu 0005, Shaowen Xie, Zhen Liu 0031, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Local-to-Global Registration for Bundle-Adjusting Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have achieved photorealistic novel views synthesis; however, the requirement of accurate camera poses limits its application. Despite analysis-by-synthesis extensions for jointly learning neural3D representations and registering camera frames exist, they are susceptible to suboptimal solutions if poorly initialized. We propose L2G-NeRF, a Local-to-Global registration method for bundle-adjusting Neural Radiance Fields: first, a pixel-wise flexible alignment, followed by a framewise constrained parametric alignment. Pixel-wise local alignment is learned in an unsupervised way via a deep network which optimizes photometric reconstruction errors. framewise global alignment is performed using differentiable parameter estimation solvers on the pixel-wise correspondences to find a global transformation. Experiments on synthetic and real-world data show that our method outperforms the current state-of-the-art in terms of high-fidelity reconstruction and resolving large camera pose misalignment. Our module is an easy-to-use plugin that can be applied to NeRF variants and other neural field applications. The Code and supplementary materials are available at https://rover-xingyu.github.io/L2G-NeRF/. Xuan Wang 0009, Qi Zhang 0029, Yu Guo 0006, Ying Shan, Fei Wang 0008 |
CVPR | 4 |
| 2023 | UV Volumes for Real-time Rendering of Editable Free-view Human PerformanceabstractNeural volume rendering enables photo-realistic renderings of a human performer in free-view, a critical task in immersive VR/AR applications. But the practice is severely limited by high computational costs in the rendering process. To solve this problem, we propose the UV Volumes, a new approach that can render an editable free-view video of a human performer in real-time. It separates the high-frequency (i.e., non-smooth) human appearance from the 3D volume, and encodes them into 2D neural texture stacks (NTS). The smooth UV volumes allow much smaller and shallower neural networks to obtain densities and texture coordinates in 3D while capturing detailed appearance in 2D NTS. For editability, the mapping between the parameterized human model and the smooth texture coordinates allows us a better generalization on novel poses and shapes. Furthermore, the use of NTS enables interesting applications, e.g., retexturing. Extensive experiments on CMU Panoptic, ZJU Mocap, and H36M datasets show that our model can render$960\times 540$images in 30FPS on average with comparable photo-realism to state-of-the-art methods. The project and supplementary materials are available at https://fanegg.github.io/UV-Volumes. Xuan Wang 0009, Qi Zhang 0029, Xiaoyu Li 0002, Yu Guo 0006, Jue Wang 0001, Fei Wang 0008 |
CVPR | 4 |
| 2023 | Inverting the Imaging Process by Learning an Implicit Camera ModelabstractRepresenting visual signals with implicit coordinate-based neural networks, as an effective replacement of the traditional discrete signal representation, has gained considerable popularity in computer vision and graphics. In contrast to existing implicit neural representations which focus on modelling the scene only, this paper proposes a novel implicit camera model which represents the physical imaging process of a camera as a deep neural network. We demonstrate the power of this new implicit camera model on two inverse imaging tasks: i) generating all-in-focus photos, and ii) HDR imaging. Specifically, we devise an implicit blur generator and an implicit tone mapper to model the aperture and exposure of the camera's imaging process, respectively. Our implicit camera model is jointly learned together with implicit scene models under multi-focus stack and multi-exposure bracket supervision. We have demonstrated the effectiveness of our new model on a large number of test images and videos, producing accurate and visually appealing all-in-focus and high dynamic range images. In principle, our new implicit neural camera model has the potential to benefit a wide array of other inverse imaging tasks. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 2 |
| 2023 | Local Implicit Ray Function for Generalizable Radiance Field RepresentationabstractWe propose LIRF (Local Implicit Ray Function), a generalizable neural rendering approach for novel view rendering. Current generalizable neural radiance fields (NeRF) methods sample a scene with a single ray per pixel and may therefore render blurred or aliased views when the input views and rendered views capture scene content with different resolutions. To solve this problem, we propose LIRF to aggregate the information from conical frustums to construct a ray. Given 3D positions within conical frustums, LIRF takes 3D coordinates and the features of conical frustums as inputs and predicts a local volumetric radiance field. Since the coordinates are continuous, LIRF renders high-quality novel views at a continuously-valued scale via volume rendering. Besides, we predict the visible weights for each input view via transformer-based feature matching to improve the performance in occluded areas. Experimental results on real-world scenes validate that our method outperforms state-of-the-art methods on novel view rendering of unseen scenes at arbitrary scales. Xin Huang 0021, Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 2 |
| 2023 | Fine-Grained Face Swapping Via Regional GAN InversionabstractWe present a novel paradigm for high-fidelity face swapping that faithfully preserves the desired subtle geometry and texture details. We rethink face swapping from the perspective of fine-grained face editing, i.e., “editing for swapping” (E4S), and propose a framework that is based on the explicit disentanglement of the shape and texture of facial components. Following the E4S principle, our framework enables both global and local swapping of facial features, as well as controlling the amount of partial swapping specified by the user. Furthermore, the E4S paradigm is in-herently capable of handling facial occlusions by means of facial masks. At the core of our system lies a novel Regional GAN Inversion (RGI) method, which allows the explicit disentanglement of shape and texture. It also allows face swapping to be performed in the latent space of Style-GAN. Specifically, we design a multi-scale mask-guided encoder to project the texture of each facial component into regional style codes. We also design a mask-guided injection module to manipulate the feature maps with the style codes. Based on the disentanglement, face swapping is re-formulated as a simplified problem of style and mask swapping. Extensive experiments and comparisons with current state-of-the-art methods demonstrate the superiority of our approach in preserving texture and shape details, as well as working with high resolution images. The project page is https://e4s2022.github.io Zhian Liu, Maomao Li, Yong Zhang 0034, Cairong Wang, Qi Zhang 0029, Jue Wang 0001, Yongwei Nie |
CVPR | 5 |
| 2023 | DINER: Disorder-Invariant Implicit Neural RepresentationabstractImplicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that such a frequency-related problem could be largely solved by re-arranging the coordinates of the input signal, for which we propose the disorder-invariant implicit neural representation (DINER) by augmenting a hash-table to a traditional INR backbone. Given discrete signals sharing the same histogram of attributes and different arrangement orders, the hash-table could project the coordinates into the same distribution for which the mapped signal can be better modeled using the subsequent INR network, leading to significantly alleviated spectral bias. Experiments not only reveal the generalization of the DINER for different INR backbones (MLP vs. SIREN) and various tasks (image/video representation, phase retrieval, and refractive index recovery) but also show the superiority over the state-of-the-art algorithms both in quality and speed. Project page: https://ezio77.github.io/DINER-website/ Shaowen Xie, Hao Zhu 0004, Zhen Liu 0031, Qi Zhang 0029, Xun Cao, Zhan Ma 0001 |
CVPR | 4 |
| 2023 | Wide-Angle Rectification via Content-Aware Conformal MappingabstractDespite the proliferation of ultra wide-angle lenses on smartphone cameras, such lenses often come with severe image distortion (e.g. curved linear structure, unnaturally skewed faces). Most existing rectification methods adopt a global warping transformation to undistort the input wideangle image, yet their performances are not entirely satisfactory, leaving many unwanted residue distortions uncorrected or at the sacrifice of the intended wide FoV (field- of-view). This paper proposes a new method to tackle these challenges. Specifically, we derive a locally-adaptive polardomain conformal mapping to rectify a wide-angle image. Parameters of the mapping are found automatically by analyzing image contents via deep neural networks. Experiments on a large number of photos have confirmed the superior performance of the proposed method compared with all available previous methods. Qi Zhang 0029, Hongdong Li, Qing Wang 0006 |
CVPR | 1 |
| 2023 | Anti-Aliased Neural Implicit Surfaces with Encoding Level of DetailabstractWe present LoD-NeuS, an efficient neural representation for high-frequency geometry detail recovery and anti-aliased novel view rendering. Drawing inspiration from voxel-based representations with the level of detail (LoD), we introduce a multi-scale tri-plane-based scene representation that is capable of capturing the LoD of the signed distance function (SDF) and the space radiance. Our representation aggregates space features from a multi-convolved featurization within a conical frustum along a ray and optimizes the LoD feature volume through differentiable rendering. Additionally, we propose an error-guided sampling strategy to guide the growth of the SDF during the optimization. Both qualitative and quantitative evaluations demonstrate that our method achieves superior surface reconstruction and photorealistic view synthesis compared to state-of-the-art approaches. Yiyu Zhuang, Qi Zhang 0029, Hao Zhu 0004, Yao Yao 0008, Xiaoyu Li 0002, Yan-Pei Cao 0001, Ying Shan, Xun Cao |
SIGGRAPH Asia | 2 |
| 2023 | Pyramid NeRF: Frequency Guided Fast Radiance Field Optimization
Junyu Zhu, Hao Zhu 0004, Qi Zhang 0029, Zhan Ma 0001, Xun Cao |
Int. J. Comput. Vis. | 3 |
| 2023 | Nonrigid Registration-Based Progressive Motion Compensation for Point Cloud Geometry CompressionabstractThere is a critical requirement for efficiently compressing point cloud geometries representing three-dimensional (3D) moving objects in various applications. The Moving Picture Experts Group 3D Graphics coding group (MPEG 3DG) set up an inter-exploration model for geometry-based point cloud compression (G-PCC interEM). However, the block-matching motion compensation scheme with a translational motion model has limited ability to handle dense point clouds with complex local motions. To overcome this problem, we propose a progressive non-rigid motion compensation framework for point cloud geometry compression, where the point cloud registration technique is introduced and tailored with our designed rate-distortion cost. In the coarse-grained stage, a point cloud is represented as deformable point patches, and the patch-wise non-rigid motion estimation task is formulated as a registration-based optimization problem that can be efficiently solved by the majorization-minimization method. In the fine-grained stage, we propose a block-based motion refinement to enhance the estimated motion field in the local region, followed by a multi-hypothesis motion compensation scheme enabling smooth reference reconstruction with patch-wise deformation and block-wise refined motions. Experiments demonstrate our proposed scheme outperforms several competitive platforms in terms of both coding performance and compensation quality. Compared with G-PCC interEM, our proposed framework achieves significant bitrate savings, i.e., 4.71% (32 frames) and 4.22% (200 frames), for point cloud lossless geometry compression. Yiting Shao, Ge Li 0002, Qi Zhang 0029, Wei Gao 0003, Shan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Hallucinated Neural Radiance Fields in the WildabstractNeural Radiance Fields (NeRF) has recently gained popularity for its impressive novel view synthesis ability. This paper studies the problem of hallucinated NeRF: i.e., recovering a realistic NeRF at a different time of day from a group of tourism images. Existing solutions adopt NeRF with a controllable appearance embedding to render novel views under various conditions, but they cannot render view-consistent images with an unseen appearance. To solve this problem, we present an end-to-end framework for constructing a hallucinated NeRF, dubbed as Ha-NeRF. Specifically, we propose an appearance hallucination module to handle time-varying appearances and transfer them to novel views. Considering the complex occlusions of tourism images, we introduce an anti-occlusion module to decompose the static subjects for visibility accurately. Experimental results on synthetic data and real tourism photo collections demonstrate that our method can hallucinate the desired appearances and render occlusion-free images from different views. The project and supplementary materials are available at https://rover-xingyu.github.io/Ha-NeRF/. Qi Zhang 0029, Xiaoyu Li 0002, Xuan Wang 0009, Jue Wang 0001 |
CVPR | 2 |
| 2022 | HDR-NeRF: High Dynamic Range Neural Radiance FieldsabstractWe present High Dynamic Range Neural Radiance Fields (HDR-NeRF) to recover an HDR radiance field from a set of low dynamic range (LDR) views with different exposures. Using the HDR-NeRF, we are able to generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the simplified physical imaging process, which dictates that the radiance of a scene point transforms to a pixel value in the LDR image with two implicit functions: a radiance field and a tone mapper. The radiance field encodes the scene radiance (values vary from 0 to$+\infty$), which outputs the density and radiance of a ray by giving corresponding ray origin and ray direction. The tone mapper models the mapping process that a ray hitting on the camera sensor becomes a pixel value. The color of the ray is predicted by feeding the radiance and the corresponding exposure time into the tone mapper. We use the classic volume rendering technique to project the output radiance, colors and densities into HDR and LDR images, while only the input LDR images are used as the supervision. We collect a new forward-facing HDR dataset to evaluate the proposed method. Experimental results on synthetic and real-world scenes validate that our method can not only accurately control the exposures of synthesized views but also render views with a high dynamic range. Xin Huang 0021, Qi Zhang 0029, Hongdong Li, Xuan Wang 0009, Qing Wang 0006 |
CVPR | 2 |
| 2022 | Deblur-NeRF: Neural Radiance Fields from Blurry ImagesabstractNeural Radiance Field (NeRF) has gained considerable attention recently for 3D scene reconstruction and novel view synthesis due to its remarkable synthesis quality. However, image blurriness caused by defocus or motion, which often occurs when capturing scenes in the wild, significantly degrades its reconstruction quality. To address this problem, We propose Deblur-NeRF, the first method that can recover a sharp NeRF from blurry input. We adopt an analysis-by-synthesis approach that reconstructs blurry views by simulating the blurring process, thus making NeRF robust to blurry inputs. The core of this simulation is a novel Deformable Sparse Kernel (DSK) module that models spatially-varying blur kernels by deforming a canonical sparse kernel at each spatial location. The ray origin of each kernel point is Jointly optimized, inspired by the physical blurring process. This module is parameterized as an MLP that has the ability to be generalized to various blur types. Jointly optimizing the NeRF and the DSK module allows us to restore a sharp NeRF. We demonstrate that our method can be used on both camera motion blur and defocus blur: the two most common types of blur in real scenes. Evaluation results on both synthetic and real-world data show that our method outperforms several baselines. The synthetic and real datasets along with the source code is publicly available at https://limacv.github.io/deblurNeRF/. Xiaoyu Li 0002, Jing Liao 0001, Qi Zhang 0029, Xuan Wang 0009, Jue Wang 0001, Pedro V. Sander |
CVPR | 4 |
| 2022 | FENeRF: Face Editing in Neural Radiance FieldsabstractPrevious portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome these limitations, we propose FENeRF, a 3D-aware generator that can produce view-consistent and locally-editable portrait images. Our method uses two decoupled latent codes to generate corresponding facial semantics and texture in a spatial-aligned 3D volume with shared geometry. Benefiting from such underlying 3D representation, FENeRF can Jointly render the boundary-aligned image and semantic mask and use the semantic mask to edit the 3D volume via GAN inversion. We further show such 3D representation can be learned from widely available monocular image and semantic mask pairs. Moreover, we reveal that Joint learning semantics and texture helps to generate finer geometry. Our experiments demonstrate that FENeRF outperforms state-of-the-art methods in various face editing tasks. Code is available at https://github.com/MrTornado24/FENeRF. Jingxiang Sun, Xuan Wang 0009, Yong Zhang 0034, Xiaoyu Li 0002, Qi Zhang 0029, Yebin Liu, Jue Wang 0001 |
CVPR | 5 |
| 2022 | Neural Color Operators for Sequential Image Retouching
Yili Wang 0003, Xin Li 0106, Kun Xu 0003, Dongliang He, Qi Zhang 0029, Fu Li 0003, Errui Ding |
ECCV (19) | 5 |
| 2022 | Ray-Space Epipolar Geometry for Light Field CamerasabstractLight field essentially represents rays in space. The epipolar geometry between two light fields is an important relationship that captures ray-ray correspondences and relative configuration of two views. Unfortunately, so far little work has been done in deriving a formal epipolar geometry model that is specifically tailored for light field cameras. This is primarily due to the high-dimensional nature of the ray sampling process with a light field camera. This paper fills in this gap by developing a novel ray-space epipolar geometry which intrinsically encapsulates the complete projective relationship between two light fields, while the generalized epipolar geometry which describes relationship of normalized light fields is the specialization of the proposed model to calibrated cameras. With Plücker parameterization, we propose the ray-space projection model involving a 6×6 ray-space intrinsic matrix for ray sampling of light field camera. Ray-space fundamental matrix and its properties are then derived to constrain ray-ray correspondences for general and special motions. Finally, based on ray-space epipolar geometry, we present two novel algorithms, one for fundamental matrix estimation, and the other for calibration. Experiments on synthetic and real data have validated the effectiveness of ray-space epipolar geometry in solving 3D computer vision tasks with light field cameras. Qi Zhang 0029, Qing Wang 0006, Hongdong Li, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Neural Parameterization for Dynamic Human Head EditingabstractImplicit radiance functions emerged as a powerful scene representation for reconstructing and rendering photo-realistic views of a 3D scene. These representations, however, suffer from poor editability. On the other hand, explicit representations such as polygonal meshes allow easy editing but are not as suitable for reconstructing accurate details in dynamic human heads, such as fine facial features, hair, teeth, and eyes. In this work, we present Neural Parameterization (NeP), a hybrid representation that provides the advantages of both implicit and explicit methods. NeP is capable of photo-realistic rendering while allowing fine-grained editing of the scene geometry and appearance. We first disentangle the geometry and appearance by parameterizing the 3D geometry into 2D texture space. We enable geometric editability by introducing an explicit linear deformation blending layer. The deformation is controlled by a set of sparse key points, which can be explicitly and intuitively displaced to edit the geometry. For appearance, we develop a hybrid 2D texture consisting of an explicit texture map for easy editing and implicit view and time-dependent residuals to model temporal and view variations. We compare our method to several reconstruction and editing baselines. The results show that the NeP achieves almost the same level of rendering accuracy while maintaining high editability. Xiaoyu Li 0002, Jing Liao 0001, Xuan Wang 0009, Qi Zhang 0029, Jue Wang 0001, Pedro V. Sander |
ACM Trans. Graph. | 5 |
| 2021 | 3D Scene Reconstruction with an Un-calibrated Light Field Camera
Qi Zhang 0029, Hongdong Li, Xue Wang 0006, Qing Wang 0006 |
Int. J. Comput. Vis. | 1 |
| 2021 | Region-based depth feature descriptor for saliency detection on light field
Xue Wang 0006, Yingying Dong, Qi Zhang 0029, Qing Wang 0006 |
Multim. Tools Appl. | 3 |
| 2020 | Fast Recolor Prediction Scheme in Point Cloud Attribute CompressionabstractDue to the emerging requirement of point cloud applications, efficient point cloud compression methods are in high demand for compact point cloud representation in limited bandwidth transmission. The compression standard GPCC (Geometry-based Point Cloud Compression) is led by the MPEG (Moving Picture Expert Group) in respond to industrial requirements. KNN (K-Nearest Neighbors) search based prediction method is adopted for point cloud attribute compression in current G-PCC, which only exploits Euclidean distance-based geometric relationship without fully consideration of underlying geometric distribution. In this paper, we propose a novel prediction scheme based on fast recolor technique for attribute lossless and near-lossless compression. Our method has been implemented upon G-PCC reference software of the latest version. Experimental results show that our method can take advantage of the correlation between the attributes of neighbors, which leads to better rate-distortion (R-D) performance than G-PCC anchor on point cloud dataset with negligible encode and decode time increase under the common test conditions. Ge Li 0002, Qi Zhang 0029, Yiting Shao, Jing Wang 0115, Shan Liu 0001 |
VCIP | 3 |
| 2020 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D images has been widely used in many computer vision tasks. Previous algorithms model the color, position, or higher spectral information for segmenting a 2D image. However, limited to the Gaussian imaging principle in a traditional camera, where each pixel is formed by summing lots of light rays from different angles, there is not a thorough segmentation solution to eliminate the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in light space, and propose light field superpixel (LFSP) to eliminate the ambiguity. The LFSP is first defined mathematically and then two evaluation metrics, named LFSP self-similarity and effective label ratio, are proposed to evaluate the refocus-invariant and full-sliced properties of segmentation. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the current state of the art in terms of traditional evaluation metrics. Additionally, the LFSP self-similarity evaluations under different light field refocus levels show the refocus-invariance of the proposed algorithm. The full-sliced property of the proposed LFSP algorithm is verified by comparing it with the classical supervoxel algorithms. Finally, an LFSP-based application is demonstrated to show the effectiveness of LFSP in light field editing. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006, Hongdong Li |
IEEE Trans. Image Process. | 2 |
| 2019 | Ray-Space Projection Model for Light Field CameraabstractLight field essentially represents the collection of rays in space. The rays captured by multiple light field cameras form subsets of full rays in 3D space and can be transformed to each other. However, most previous approaches model the projection from an arbitrary point in 3D space to corresponding pixel on the sensor. There are few models on describing the ray sampling and transformation among multiple light field cameras. In the paper, we propose a novel ray-space projection model to transform sets of rays captured by multiple light field cameras in term of the Plucker coordinates. We first derive a 6×6 ray-space intrinsic matrix based on multi-projection-center (MPC) model. A homogeneous ray-space projection matrix and a fundamental matrix are then proposed to establish ray-ray correspondences among multiple light fields. Finally, based on the ray-space projection matrix, a novel camera calibration method is proposed to verify the proposed model. A linear constraint and a ray-ray cost function are established for linear initial solution and non-linear optimization respectively. Experimental results on both synthetic and real light field data have verified the effectiveness and robustness of the proposed model. Qi Zhang 0029, Jinbo Ling, Qing Wang 0006, Jingyi Yu 0001 |
CVPR | 1 |
| 2019 | A Generic Multi-Projection-Center Model and Calibration Method for Light Field CamerasabstractLight field cameras can capture both spatial and angular information of light rays, enabling 3D reconstruction by a single exposure. The geometry of 3D reconstruction is affected by intrinsic parameters of a light field camera significantly. In the paper, we propose a multi-projection-center (MPC) model with 6 intrinsic parameters to characterize light field cameras based on traditional two-parallel-plane (TPP) representation. The MPC model can generally parameterize light field in different imaging formations, including conventional and focused light field cameras. By the constraints of 4D ray and 3D geometry, a 3D projective transformation is deduced to describe the relationship between geometric structure and the MPC coordinates. Based on the MPC model and projective transformation, we propose a calibration algorithm to verify our light field camera model. Our calibration method includes a close-form solution and a non-linear optimization by minimizing re-projection errors. Experimental results on both simulated and real scene data have verified the performance of our algorithm. Qi Zhang 0029, Chunping Zhang, Jinbo Ling, Qing Wang 0006, Jingyi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Common Self-polar Triangle of Concentric Conics for Light Field Camera Calibration
Qi Zhang 0029, Qing Wang 0006 |
ACCV (6) | 1 |
| 2018 | Hybrid Point Cloud Attribute Compression Using Slice-based Layered Structure and Block-based Intra PredictionabstractPoint cloud compression is a key enabler for the emerging applications of immersive visual communication, autonomous driving and smart cities, etc. In this paper, we propose a hybrid point cloud attribute compression scheme built on an original layered data structure. First, a slice partition scheme and a geometry-adaptive k-dimensional tree (k-d tree) method are devised to generate layer structures. Second, we introduce an efficient block-based intra prediction scheme containing to exploit spatial correlations among adjacent points. Third, an adaptive transform scheme based on Graph Fourier Transform (GFT) is Lagrangian optimized to achieve better transform efficiency. The Lagrange multiplier is off-line derived based on the statistics of attribute coding. Last but not least, multiple scan modes are dedicated to improve coding efficiency for entropy coding. Experimental results demonstrate that our method performs better than the state-of-the-art region-adaptive hierarchical transform (RAHT) system, and on average a 37.21% BD-rate gain is achieved. Comparing with the test model for category 1 (TMC1) anchors, which were recently published by MPEG-3DG group on 121st MPEG meeting, a 8.81% BD-rate gain is obtained. Yiting Shao, Qi Zhang 0029, Ge Li 0002, Zhu Li 0001, Li Li 0040 |
ACM Multimedia | 2 |
| 2018 | Point Clouds Attribute Compression Using Data-Adaptive Intra predictionabstractIn recent years, 3D sensing and capture technologies have made constant progress, leading to point clouds with higher resolution and fidelity. Since most applications demand compact storage and fast transmission, the issue of how to compress point clouds efficiently becomes an intractable problem. While previous GFT-based solutions use the transform tool to decorrelate attributes directly, ignoring the overall attribute's data spatial redundancy, Graph Fourier Transform (GFT) has shown good performance on point cloud attribute compression. So, motivated by coding tools in traditional image and video coding, we propose a block-based data-adaptive intra prediction tool before graph transform processing to further reduce the redundancy. We adopt uniform quantizing and context-based arithmetic coding to get the final bitstream. Experimental results on different datasets demonstrate that our method improves the compression efficiency of other GFT-based schemes and has much better BD-BR performance than the state-of-the-art Region-Adaptive Hierarchical Transform (RAHT) approach on most specified point cloud contents. Qi Zhang 0029, Yiting Shao, Ge Li 0002 |
VCIP | 1 |
| 2017 | 4D Light Field Superpixel and SegmentationabstractSuperpixel segmentation of 2D image has been widely used in many computer vision tasks. However, limited to the Gaussian imaging principle, there is not a thorough segmentation solution to the ambiguity in defocus and occlusion boundary areas. In this paper, we consider the essential element of image pixel, i.e., rays in the light space and propose light field superpixel (LFSP) segmentation to eliminate the ambiguity. The LFSP is first defined mathematically and then a refocus-invariant metric named LFSP self-similarity is proposed to evaluate the segmentation performance. By building a clique system containing 80 neighbors in light field, a robust refocus-invariant LFSP segmentation algorithm is developed. Experimental results on both synthetic and real light field datasets demonstrate the advantages over the state-of-the-arts in terms of traditional evaluation metrics. Additionally the LFSP self-similarity evaluation under different light field refocus levels shows the refocus-invariance of the proposed algorithm. Hao Zhu 0005, Qi Zhang 0029, Qing Wang 0006 |
CVPR | 2 |
| 2017 | Extending the FOV from disparity and color consistencies in multiview light fieldsabstractLight field, which is captured by a plenoptic camera, is always limited in its narrow field of view (FOV) by the physical size of the aperture. To break through the restriction, we propose to extend the FOV using multiview light fields. A series of light fields are acquired by translating the camera at isometric spatial positions. In contrast to previous methods, our algorithm is the first that achieves light field registration and rendering based on epipolar plane image (EPI) properties, including disparity and color consistencies. Furthermore, the aliasing caused by the under-sampling in the angular space is eliminated by synthesizing novel views in the EPI space. Experimental results on the real scene data have demonstrated the effectiveness of our algorithm. Zhao Ren, Qi Zhang 0029, Hao Zhu 0005, Qing Wang 0006 |
ICIP | 2 |