Qi Zhang 0071

dblp:52/323-71 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-0548-6121ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations
abstract
The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical in real-world scenarios. In this paper, we propose a novel learning from Unposed views under Varied illuminations Relightable 3D Gaussian Splatting (dubbed UV-RGS), to address this challenge by jointly optimizing camera poses, 3DGS representations, surface materials, and environment illuminations (i.e., unknown and varied lighting conditions in training) using only unposed views under varied lightings. Firstly, UV-RGS presents a viewpoint dividing strategy to group inputs into constituent units, enabling each unit can perform similar poses and illuminations. Next, for each unit, to get the constituent model, UV-RGS establishes an incrementally pose learning module to estimate coarse camera parameters, which also enjoy a proxy-view refinement to alleviate the sparse view learning. Additionally, for all constituent unit models, we introduce a holistic model learning strategy that integrates progressive unit aggregation component and the 3DGS coupled with camera poses joint optimization, which realizes the scene high-fidelity perception by the physical-based rendering. Extensive experiments on both real-world and synthetic challenging datasets demonstrate the effectiveness of UV-RGS, achieving the state-of-the-art performance for scene inverse rendering by learning 3DGS from only unposed views under varied illuminations.
Wei Feng 0005, Chi Huang, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048
AAAI3
2026 Imaging-sensitive defect detection for high-surface-quality products
Nan Li 0048, Qian Zhang 0051, Qi Zhang 0071, Wei Feng 0005
Expert Syst. Appl.3
2025 Generative Hard Example Augmentation for Semantic Point Cloud Segmentation
abstract
The recent progress in semantic point cloud segmentation is attributed to deep networks, which require a large amount of point cloud data for training. However, how to collect substantial point-wise annotations of the point clouds at affordable cost for the end-to-end network training still needs to be solved. In this paper, we propose Generative Hard Example Augmentation (GHEA) to achieve novel examples of point clouds, which enrich the data for training the segmentation network. Firstly, GHEA employs the generative network to embed the discrepancy between the point clouds into the latent space. From the latent space, we sample multiple discrepancies for reshaping a point cloud to various examples, contributing to the richness of the training data. Secondly, GHEA mixes the reshaped point clouds by respecting their segmentation errors. This mixup allows the reshaped point clouds, which are difficult to segment, to join as the challenging example for network training. We evaluate the effectiveness of GHEA, which helps the popular segmentation networks to improve the performances.
Qi Zhang 0071, Jibin Peng, Wei Feng 0005, Di Lin 0002
CVPR1
2025 SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views Under Unconstrained Illuminations
Qi Zhang 0071, Chi Huang, Qian Zhang 0051, Nan Li 0048, Wei Feng 0005
ICCV1
2025 2D Gaussian Splatting for Outdoor Scene Decomposition and Relighting
abstract
Gaussian splatting techniques have recently revolutionized outdoor scene decomposition and relighting through multi-view images. However, achieving high rendering quality still requires a fixed lighting condition among all input views, which is costly or even impractical to capture in outdoor scenes. In this paper, we propose outdoor scene decomposition and relighting with 2D Gaussian splatting (OSDR-GS), a novel inverse rendering strategy under outdoor changing and unknown lighting conditions. Firstly, we present a lighting-based group learning framework that categorizes input images into multiple lighting groups, to learn the separate lighting from each group individually. Secondly, OSDR-GS introduces a fine-grained outdoor lighting component to represent sun-light and sky-light, respectively, which are also adjusted via the correlative exposure factors adaptively. Finally, we construct a visibility-driven shadow module to characterize the nuanced interplay of light and occlusion realistically, for eliminating the uncertainty of dark pixels on lighting-based group learning. Extensive experiments on multiple challenging outdoor datasets validate the effectiveness of OSDR-GS, which achieves the state-of-the-art performance in changing lighting scene inverse rendering.
Wei Feng 0005, Kangrui Ye, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048
IJCAI3
2025 TriGS: Tri-consistency 3D Gaussian Splatting from Sparse and Unposed Views
abstract
Recent advances in 3D scene representation, particularly 3D Gaussian Splatting (3DGS), have demonstrated remarkable photorealistic rendering capabilities. However, the heavy reliance on dense and precisely calibrated camera configurations limits effectiveness in sparse view and unposed scenarios. In this paper, we present Tri-consistency 3D Gaussian Splatting (dubbed TriGS), a novel framework that jointly optimizes 3DGS parameters and camera poses only from sparse and unposed images via triple consistency supervisions coupled with the adaptive regularization strategy. We first estimate coarse camera poses by exploiting 3DGS's anisotropic properties through iterative relative pose optimization. Building upon this foundation, we introduce cross-view consistency enforcement through synchronized photometric color, geometric structure, and deep feature, effectively resolving rendering ambiguities with auxiliary supervisions. A unified rendering paradigm is also proposed to jointly refine Gaussian primitives and camera poses by transforming positions, covariances, and spherical harmonics. To combat overfitting inherent in joint optimization, we devise an adaptive regularization mechanism that strategically samples hard viewpoints based on baseline distances and training dynamics, enforcing projection consistency through deep feature priors. Extensive experiments on multiple challenging real-world datasets validate the effectiveness of TriGS, which achieves satisfactory results to set a new state-of-the-art without the reliance on external pose priors only under sparse and unposed view inputs.
Chi Huang, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048, Yipu Gong, Wei Feng 0005
ACM Multimedia2
2024 Learning Geometry Consistent Neural Radiance Fields from Sparse and Unposed Views
abstract
The latest progress in novel view synthesis can be attributed to the Neural Radiance Field (NeRF), which requires densely sampled images with precise camera poses. However, collecting dense input images for a NeRF with accurate camera poses is highly expensive in many real-world scenarios. In this paper, we propose to learn Geometry Consistent Neural Radiance Field (GC-NeRF), to tackle this challenge by jointly optimizing a NeRF and its corresponding camera poses with sparse (as low as 2) and unposed views. First, the proposed GC-NeRF establishes image-level geometric consistencies, by producing photometric constraints from inter- and intra-views to update the NeRF and the camera poses in a fine-grained manner. Then, we adopt geometry projection with camera extrinsic parameters to further provide region-level consistency supervisions, which constructs pseudo-pixel labels to capture critical matching correlations. Moreover, we present an adaptive high-frequency mapping function to augment the geometry and texture information of the 3D scene. Extensive experiments on multiple challenging real-world datasets validate the effectiveness of the proposed GC-NeRF, which sets a new state-of-the-art for effectively learning NeRF with sparse and unposed views.
Qi Zhang 0071, Chi Huang, Qian Zhang 0051, Nan Li 0048, Wei Feng 0005
ACM Multimedia1
2023 GCRec: Graph-Augmented Capsule Network for Next-Item Recommendation
abstract
Next-item recommendation has been a hot research, which aims at predicting the next action by modeling users' behavior sequences. While previous efforts toward this task have been made in capturing complex item transition patterns, we argue that they still suffer from three limitations: 1) they have difficulty in explicitly capturing the impact of inherent order of item transition patterns; 2) only a simple and crude embedding is insufficient to yield satisfactory long-term users' representations from limited training sequences; and 3) they are incapable of dynamically integrating long-term and short-term user interest modeling. In this work, we propose a novel solution named graph-augmented capsule network (GCRec), which exploits sequential user behaviors in a more fine-grained manner. Specifically, we employ a linear graph convolution module to learn informative long-term representations of users. Furthermore, we devise a user-specific capsule module and a position-aware gating module, which are sensitive to the relative sequential order of the recently interacted items, to capture sequential patterns at union-level and point-level. To aggregate the long-term and short-term user interests as a representative vector, we design a dual-gating mechanism, which could decide the contribution ratio of each module given different contextual information. Through extensive experiments on four benchmarks, we validate the rationality and effectiveness of GCRec on the next-item recommendation task.
Bin Wu 0019, Xiangnan He 0001, Qi Zhang 0071, Meng Wang 0001, Yangdong Ye
IEEE Trans. Neural Networks Learn. Syst.3
2022 Gating augmented capsule network for sequential recommendation
Qi Zhang 0071, Bin Wu 0019, Zhongchuan Sun, Yangdong Ye
Knowl. Based Syst.1