Qing Li 0029

dblp:181/2689-29 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-7816-9733ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Realistic and Stable 3D Gaussian Relighting: Semantic-Guided Stable Material Estimation Through Object-Oriented Differentiable Gaussian Path Tracing
abstract
The Gaussian relighting is critical for appearance editing and physical-based rendering, which rely on the decomposition and estimation of light and materials. However, existing methods treat both the environmental lighting and material features as unknowns and estimate them jointly, which results in inaccurate material decomposition and estimation due to the inherent variation of environmental illumination across datasets. In this study, we introduce the realistic and stable 3D Gaussian relighting (RS3DGR) that is enabled by the semantic-guided stable material estimation through object-oriented differentiable Gaussian path tracing. First, the target object is segmented from environment via semantic 3D Gaussian segmentation and the semantic label is designed as the roughness regularization term in optimization. Then, the object is illuminated by the environmental light sampled from the radiance field of environmental Gaussians or other types of radiance caches. The object-oriented path tracing rendering is differentiable and efficient. Under the alpha blending of primary rays penetrating through object Gaussians, the secondary rays are uniformly emitted from those primary-hit Gaussians to directly sample the environmental radiance field to model the reflection on object surface. As a result, the influence of unstable environmental lighting is removed, and the semantic prior successfully promotes the convergence on material estimation. Experimental results demonstrate significant improvements in stability and accuracy for material reconstruction on both synthetic and self-constructed datasets.
Xinzhu Sang, Zhidong Chen, Hongkun Cao, Luyu Ji, Qing Li 0029
IEEE Trans. Circuits Syst. Video Technol.6
2026 PSAvatar: A Point-Based Shape Model for Real-Time Head Avatar Animation With 3D Gaussian Splatting
abstract
Despite much progress, achieving real-time highfidelity head avatar animation is still difficult and existing methods have to trade-off between speed and quality. 3DMM based methods often fail to model non-facial structures such as eyeglasses and hairstyles, while neural implicit models suffer from deformation inflexibility and rendering inefficiency. Although 3D Gaussian has been demonstrated to possess promising capability for geometry representation and radiance field reconstruction, applying 3D Gaussian in head avatar creation remains a major challenge since it is difficult for 3D Gaussian to model the head shape variations caused by changing poses and expressions. In this paper, we introduce PSAvatar, a novel framework for animatable head avatar creation that utilizes discrete geometric primitive to create a parametric shape model and employs 3D Gaussian for fine detail representation and high fidelity rendering. The parametric shape model is a Point-based Shape Model (PSM) which uses points instead of meshes for 3D representation to achieve enhanced representation flexibility. Specifically, PSM first converts the FLAME mesh to points by sampling on the surfaces as well as off the meshes to enable the reconstruction of not only surface-like structures but also complex geometries such as eyeglasses and hairstyles. By aligning these points with the head shape in an analysis-by-synthesis manner, the PSM makes it possible to utilize 3D Gaussian for fine detail representation and appearance modeling, thus enabling the creation of high-fidelity avatars. We show that PSAvatar can reconstruct high-fidelity head avatars of varieties of subjects and the avatars can be animated in real-time.
Zhenyu Bao, Qing Li 0029, Guoping Qiu, Kanglin Liu
IEEE Trans. Vis. Comput. Graph.3
2025 SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputs
abstract
3D Gaussian Splatting-based indoor open-world free-view synthesis approaches have shown significant performance with dense input images. However, they exhibit poor performance when confronted with sparse inputs, primarily due to the sparse distribution of Gaussian points and insufficient view supervision. To relieve these challenges, we propose SPC-GS, leveraging Scene-layout-based Gaussian Initialization (SGI) and Semantic-Prompt Consistency (SPC) Regularization for open-world free view synthesis with sparse inputs. Specifically, SGI provides a dense, scene-layout-based Gaussian distribution by utilizing view-changed images generated from the video generation model and view-constraint Gaussian points densification. Additionally, SPC mitigates limited view supervision by employing semantic-prompt-based consistency constraints developed by SAM2. This approach leverages available semantics from training views, serving as instructive prompts, to optimize visually overlapping regions in novel views with 2D and 3D consistency constraints. Extensive experiments demonstrate the superior performance of SPC-GS across Replica and ScanNet benchmarks. Notably, our SPC-GS achieves a 3.06 dB gain in PSNR for reconstruction quality and a 7.3% improvement in mIoU for open-world semantic segmentation. Project website at: https://gbliao.github.io/SPC-GS.github.io.
Guibiao Liao, Qing Li 0029, Zhenyu Bao, Guoping Qiu, Kanglin Liu
CVPR2
2025 A Unified Inverse-Tone-Mapped HDR Video Quality Assessment Method across Two HDR Formats
abstract
High Dynamic Range Video Quality Assessment (HDR VQA) plays a pivotal role in Inverse Tone Mapping (ITM) research. Existing HDR VQA datasets and models mainly focus on a single HDR format, leading to poor generalization and limited application scope. To address this problem, this paper proposes a reference-free CONTrastive ITM-VQA (CONT-ITM-VQA) model via format transformation-based data augmentation and contrastive learning. Specifically, the format transformation-based data augmentation improves the model generalization, via applying Opto-electronic Transfer Function (OETF) transformations between the HDR formats; while the contrastive learning-based quality-related feature alignment aligns the quality features from different HDR formats of the same video to obtain more effective quality representations. It is worth noting that our method can be extended to other HDR-related quality assessment, not limited to ITM-HDR VQA. Experimental results demonstrate that our model closely mimics subjective judgments.
Leidong Fan, Xiongkuo Min, Qing Li 0029, Anjie Wang
ICME3
2025 Inverse-Tone-Mapped HDR Video Quality Assessment for Broadcast Television: A Comprehensive Dataset and SDR-Referenced Method
abstract
Inverse-Tone-Mapped High Dynamic Range Video Quality Assessment (ITM-HDR VQA) plays a pivotal role in evaluating the visual quality of ITM-enhanced HDR videos. The research community tackles this issue from the dataset and method perspectives. However, current ITM-HDR VQA datasets exhibit three key limitations: narrow scene diversity, partial HDR format representation, and inadequate distortion coverage; existing methods face challenges in HDR and SDR domain discrepancy and insufficient ITM-induced quality feature extraction. To bridge these gaps, we introduce a comprehensive ITM HDR Video Quality Assessment dataset tailored to Broadcast Television (BT-ITM-VQA), along with a novel SDR-Referenced Bidirectional Quality Interaction (SDR-R-BQI) method. The BT-ITM-VQA dataset features rich broadcast scenes, multiple HDR-format support of Hybrid Log-Gamma (HLG) and Perceptual Quantizer (PQ), and real-world distortions induced by super-resolution and deinterlacing, providing a systematic foundation for ITM-HDR VQA model development and validation. The SDR-R-BQI method effectively mitigates HDR and SDR discrepancies through luminance dynamic range alignment and color gamut alignment, and then extracts ITM-induced quality alterations by bidirectional, cross-quality-based computation in a unified feature space. Extensive validation on four datasets demonstrates the effectiveness of our newly constructed dataset and proposed method.
Leidong Fan, Qian Zhang 0096, Qing Li 0029
ACM Multimedia3
2025 LoopSparseGS: Loop-Based Sparse-View Friendly Gaussian Splatting
abstract
Despite the photorealistic novel view synthesis (NVS) performance achieved by the original 3D Gaussian splatting (3DGS), its rendering quality significantly degrades with sparse input views. This performance drop is mainly caused by the limited number of initial points generated from the sparse input, lacking reliable geometric supervision during the training process, and inadequate regularization of the oversized Gaussian ellipsoids. To handle these issues, we propose the LoopSparseGS, a loop-based 3DGS framework for the sparse novel view synthesis task. In specific, we propose a loop-based Progressive Gaussian Initialization (PGI) strategy that could iteratively densify the initialized point cloud using the rendered pseudo images during the training process. Then, the sparse and reliable depth from the Structure from Motion, and the window-based dense monocular depth are leveraged to provide precise geometric supervision via the proposed Depth-alignment Regularization (DAR). Additionally, we introduce a novel Sparse-friendly Sampling (SFS) strategy to handle oversized Gaussian ellipsoids leading to large pixel errors. Comprehensive experiments on four datasets demonstrate that LoopSparseGS outperforms existing state-of-the-art methods for sparse-input novel view synthesis, across indoor, outdoor, and object-level scenes with various image resolutions. Code is available at: https://github.com/pcl3dv/LoopSparseGS.
Zhenyu Bao, Guibiao Liao, Kaichen Zhou, Kanglin Liu, Qing Li 0029, Guoping Qiu
IEEE Trans. Image Process.5
2025 CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding
abstract
Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus. Existing methods typically attach high-dimensional CLIP semantic embeddings to 3D Gaussians and leverage view-inconsistent 2D CLIP semantics as Gaussian supervision, resulting in efficiency bottlenecks and deficient 3D semantic consistency. To address these challenges, we present CLIP-GS, efficiently achieving a coherent semantic understanding of 3D indoor scenes via the proposed Semantic Attribute Compactness (SAC) and 3D Coherent Regularization (3DCR). SAC approach exploits the naturally unified semantics within objects to learn compact, yet effective, semantic Gaussian representations, enabling highly efficient rendering (>100 FPS). 3DCR enforces semantic consistency in 2D and 3D domains: In 2D, 3DCR utilizes refined view-consistent semantic outcomes derived from 3DGS to establish cross-view coherence constraints; in 3D, 3DCR encourages features similar among 3D Gaussian primitives associated with the same object, leading to more precise and coherent segmentation results. Extensive experimental results demonstrate that our method remarkably suppresses existing state-of-the-art approaches, achieving mIoU improvements of 21.20% and 13.05% on ScanNet and Replica datasets, respectively, while maintaining real-time rendering speed. Furthermore, our approach exhibits superior performance even with sparse input data, substantiating its robustness.
Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Qing Li 0029, Kanglin Liu
ACM Trans. Multim. Comput. Commun. Appl.5
2024 3D Reconstruction and Novel View Synthesis of Indoor Environments Based on a Dual Neural Radiance Field
Zhenyu Bao, Guibiao Liao, Kanglin Liu, Qing Li 0029, Guoping Qiu
ACM Multimedia5
2024 OV-NeRF: Open-Vocabulary Neural Radiance Fields With Vision and Language Foundation Models for 3D Semantic Understanding
abstract
The development of Neural Radiance Fields (NeRFs) has provided a potent representation for encapsulating the geometric and appearance characteristics of 3D scenes. Enhancing the capabilities of NeRFs in open-vocabulary 3D semantic perception tasks has been a recent focus. However, current methods that extract semantics directly from Contrastive Language-Image Pretraining (CLIP) for semantic field learning encounter difficulties due to noisy and view-inconsistent semantics provided by CLIP. To tackle these limitations, we propose OV-NeRF, which exploits the potential of pre-trained vision and language foundation models to enhance semantic field learning through proposed single-view and cross-view strategies. First, from the single-view perspective, we introduce Region Semantic Ranking (RSR) regularization by leveraging 2D mask proposals derived from Segment Anything (SAM) to rectify the noisy semantics of each training view, facilitating accurate semantic field learning. Second, from the cross-view perspective, we propose a Cross-view Self-enhancement (CSE) strategy to address the challenge raised by view-inconsistent semantics. Rather than invariably utilizing the 2D inconsistent semantics from CLIP, CSE leverages the 3D consistent semantics generated from the well-trained semantic field itself for semantic field training, aiming to reduce ambiguity and enhance overall semantic consistency across different views. Extensive experiments validate our OV-NeRF outperforms current state-of-the-art methods, achieving a significant improvement of 20.31% and 18.42% in mIoU metric on Replica and ScanNet, respectively. Furthermore, our approach exhibits consistent superior results across various CLIP configurations, further verifying its robustness. Codes are available at:https://github.com/pcl3dv/OV-NeRF.
Guibiao Liao, Kaichen Zhou, Zhenyu Bao, Kanglin Liu, Qing Li 0029
IEEE Trans. Circuits Syst. Video Technol.5
2024 DarkLoc+: Thermal Image-Based Indoor Localization for Dark Environments With Relative Geometry Constraints
abstract
Thermal images capture temperature information of the environments instead of texture, making it well suitable for obtaining position in dark environments. Many methods have been proposed to handle RGB images, while thermal image-based localization methods are not well studied. To address it, we propose DarkLoc+, a thermal image-based indoor localization method based on the attention model and relative constraints between images under a learning-based localization framework. To be specific, we utilize self-attention to extract reprehensive features from thermal images and exploit relative constraints to enforce the convolutional neural networks (CNNs) to predict global poses. Relative pose loss(RelLoss)and relative regression loss are designed to work with global poses to constrain the network in feature and pose space simultaneously. We evaluate the proposed method on the public thermal images indoor dataset and our own dataset. The experimental results demonstrate that our method can obtain accurate position information.
Baoding Zhou, Yufeng Xiao, Qing Li 0029, Bing Wang 0013, Longmin Pan, Dejin Zhang, Jiasong Zhu, Qingquan Li 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 DarkLoc: Attention-based Indoor Localization Method for Dark Environments Using Thermal Images
abstract
Image-based localization is an essential component for many applications such as autonomous driving, virtual reality. Many researchers focus on developing methods for daytime via RGB images. Few research study the methods for night condition. The main reason is that the RGB-based image localization methods fail to work in dark scenes due to the low illumination. Thermal images capture temperature information instead of texture, making it well suitable for dark environments. However, thermal image-based localization methods are not well studied. To address it, we propose an attention-based localization method for night condition (DarkLoc) using thermal images. The proposed method introduce the attention mechanism in a deep learning-based framework to extract key information from low quality thermal images. The attention model can enforce the whole network focus on geometry meaningful feature in thermal images and thus improve the localization accuracy. We perform extensive experiment on the thermal image dataset. The results show that the attention model can enforce the whole network to learn geometry meaningful feature from thermal images and effectively locate the thermal image in real-time.
Baoding Zhou, Longming Pan, Qing Li 0029, Gang Liu 0028, Aiwu Xiong, Qingquan Li 0001
IPIN3
2021 Relative geometry-aware siamese neural network for 6DOF camera relocalization
Qing Li 0029, Jiasong Zhu, Rui Cao 0001, Ke Sun 0006, Jonathan M. Garibaldi, Qingquan Li 0001, Guoping Qiu
Neurocomputing1
2020 A contextual conditional random field network for monocular depth estimation
Qing Li 0029, Rui Cao 0001, Wenming Tang, Guoping Qiu
Image Vis. Comput.2
2019 Learning Spatial-Aware Cross-View Embeddings for Ground-to-Aerial Geolocalization
Rui Cao 0001, Jiasong Zhu, Qing Li 0029, Qian Zhang 0018, Qingquan Li 0001, Guoping Qiu
ICIG (1)3