VLDB 2026 Research / reviewers in the wild / expert
Guibiao Liao
dblp:276/3125
· DBLP profile ↗
14ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-5714-1926ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Monocular Depth via Cascaded Iterative Refinement in Visual-Echo ScenesabstractIn recent years, integrating multimodal information, particularly visual and echo data, has shown great promise for improving depth estimation performance. While existing works demonstrate that combining binaural echo features with image attributes can enhance depth estimation, they often use rudimentary feature alignment and fusion methods, failing to fully exploit the complementary nature of cross-modal information and limiting integration effectiveness. To address these challenges, this paper introduces an innovative multimodal fusion framework. First, the framework incorporates a combination of multi-scale self-attention and cross-attention mechanisms, establishing correlations between features and facilitating cohesive interactions between the visual and echo domains. Furthermore, we propose an incremental feature updating mechanism based on Convolutional Gated Recurrent Units (ConvGRU), which implements cascaded iterative optimization, integrating contextual features with the multi-scale fused features from both echo and image modalities. In each iteration, the framework preserves contextual information from previous steps while employing a multi-level loss function to guide result updates. This approach effectively captures spatial structural information and progressively enhances depth estimation accuracy. Comprehensive experimental evaluations on the Replica, Matterport3D and BatVision (BV1) datasets validate the effectiveness of the proposed method. Comparative analyses with state-of-the-art monocular plus echo methods underscore the superior performance achievable through this novel framework. Anjie Wang, Zhijun Fang 0001, Leidong Fan, Guibiao Liao, Siwei Ma 0001, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputsabstract3D Gaussian Splatting-based indoor open-world free-view synthesis approaches have shown significant performance with dense input images. However, they exhibit poor performance when confronted with sparse inputs, primarily due to the sparse distribution of Gaussian points and insufficient view supervision. To relieve these challenges, we propose SPC-GS, leveraging Scene-layout-based Gaussian Initialization (SGI) and Semantic-Prompt Consistency (SPC) Regularization for open-world free view synthesis with sparse inputs. Specifically, SGI provides a dense, scene-layout-based Gaussian distribution by utilizing view-changed images generated from the video generation model and view-constraint Gaussian points densification. Additionally, SPC mitigates limited view supervision by employing semantic-prompt-based consistency constraints developed by SAM2. This approach leverages available semantics from training views, serving as instructive prompts, to optimize visually overlapping regions in novel views with 2D and 3D consistency constraints. Extensive experiments demonstrate the superior performance of SPC-GS across Replica and ScanNet benchmarks. Notably, our SPC-GS achieves a 3.06 dB gain in PSNR for reconstruction quality and a 7.3% improvement in mIoU for open-world semantic segmentation. Project website at: https://gbliao.github.io/SPC-GS.github.io. Guibiao Liao, Qing Li 0029, Zhenyu Bao, Guoping Qiu, Kanglin Liu |
CVPR | 1 |
| 2025 | Toward Realistic Co-Speech Motion via Cross-Modal Spatial-Temporal Attention and Hand Memory ModuleabstractCo-speech video generation focuses on improving the authenticity of virtual characters by aligning their gestures and facial expressions with spoken audio. Despite recent advancements, existing methods often struggle with speech-gesture misalignment and unnatural hand motions. To address the issues, we propose a novel audio-driven gesture generation framework. This framework integrates a hierarchical diffusion model with multimodal feature disentanglement and dynamic fusion strategies. The core of our approach is the Cross-modal Spatial-Temporal Attention mechanism (CSTA), which ensures high-fidelity synchronization between audio and human motion while capturing fine-grained dynamics of hand and facial. By effectively disentangling different motion modalities, CSTA enhances the alignment between body part movements and the audio signal, leading to more natural and coherent video synthesis. Furthermore, to improve the physical plausibility and diversity of generated gestures, we introduce a Hand Memory Module (HMM). This module leverages a Vector Quantization-Variational Autoencoder (VQ-VAE) to learn a discrete gesture prior. By embedding these learned priors during the generation process, our method not only enhances temporal consistency but also preserves intricate details, mitigating common issues in prior work such as motion blur and detail loss. Experiments on the PATS and BEAT2 datasets demonstrate that CSTA surpasses existing methods in generating co-speech videos with more synchronized and natural hand motions, achieving state-of-the-art performance in both qualitative and quantitative evaluations. Project Page: CSTA-HMM Dingwei Liu, Guibiao Liao, Xiuhua Jiang, Jiangbo Xu |
ECAI | 3 |
| 2025 | Sparse-view 3D Open-vocabulary Gaussian Splatting via Collaborative Contrastive Learningabstract3D Gaussian Splatting-based Open-vocabulary 3D segmentation has shown impressive performance with dense input images. However, existing methods exhibit poor results when confronted with sparse inputs, primarily due to limited overlap among input views and insufficient view supervision provided. To tackle these challenges, we propose SpContrast, a novel framework that creates additional semantic constraints to enhance sparse-view 3D open-vocabulary segmentation. First, we introduce Collaborative Contrastive Learning (CCL), which creates instructive multi-view semantic constraints by collaboratively mining semantic interactions between training and online-rendered novel views. Motivated by the principle that semantically consistent features should converge and divergent ones separate, CCL establishes cross-view contrastive constraints to enhance semantic coherence. Second, to alleviate the adverse impact of false negative samples caused by semantic inconsistencies within the same object, we present Region-aware Negative Sampling (RNS). RNS rectifies these false negative samples, and treats them as hard samples during our contrastive optimization, leading to improved object completeness and more accurate segmentation. Extensive experiments on challenging sparse-input datasets, including Replica and ScanNet, demonstrate the superiority of SpContrast, achieving 7.6% and 8.3% mIoU improvements for 3D open-vocabulary segmentation. Guibiao Liao, Anjie Wang, Mingxuan Chen, Zhijun Fang 0001 |
ICME | 1 |
| 2025 | LoopSparseGS: Loop-Based Sparse-View Friendly Gaussian SplattingabstractDespite the photorealistic novel view synthesis (NVS) performance achieved by the original 3D Gaussian splatting (3DGS), its rendering quality significantly degrades with sparse input views. This performance drop is mainly caused by the limited number of initial points generated from the sparse input, lacking reliable geometric supervision during the training process, and inadequate regularization of the oversized Gaussian ellipsoids. To handle these issues, we propose the LoopSparseGS, a loop-based 3DGS framework for the sparse novel view synthesis task. In specific, we propose a loop-based Progressive Gaussian Initialization (PGI) strategy that could iteratively densify the initialized point cloud using the rendered pseudo images during the training process. Then, the sparse and reliable depth from the Structure from Motion, and the window-based dense monocular depth are leveraged to provide precise geometric supervision via the proposed Depth-alignment Regularization (DAR). Additionally, we introduce a novel Sparse-friendly Sampling (SFS) strategy to handle oversized Gaussian ellipsoids leading to large pixel errors. Comprehensive experiments on four datasets demonstrate that LoopSparseGS outperforms existing state-of-the-art methods for sparse-input novel view synthesis, across indoor, outdoor, and object-level scenes with various image resolutions. Code is available at: https://github.com/pcl3dv/LoopSparseGS. Zhenyu Bao, Guibiao Liao, Kaichen Zhou, Kanglin Liu, Qing Li 0029, Guoping Qiu |
IEEE Trans. Image Process. | 2 |
| 2025 | CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic UnderstandingabstractExploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus. Existing methods typically attach high-dimensional CLIP semantic embeddings to 3D Gaussians and leverage view-inconsistent 2D CLIP semantics as Gaussian supervision, resulting in efficiency bottlenecks and deficient 3D semantic consistency. To address these challenges, we present CLIP-GS, efficiently achieving a coherent semantic understanding of 3D indoor scenes via the proposed Semantic Attribute Compactness (SAC) and 3D Coherent Regularization (3DCR). SAC approach exploits the naturally unified semantics within objects to learn compact, yet effective, semantic Gaussian representations, enabling highly efficient rendering (>100 FPS). 3DCR enforces semantic consistency in 2D and 3D domains: In 2D, 3DCR utilizes refined view-consistent semantic outcomes derived from 3DGS to establish cross-view coherence constraints; in 3D, 3DCR encourages features similar among 3D Gaussian primitives associated with the same object, leading to more precise and coherent segmentation results. Extensive experimental results demonstrate that our method remarkably suppresses existing state-of-the-art approaches, achieving mIoU improvements of 21.20% and 13.05% on ScanNet and Replica datasets, respectively, while maintaining real-time rendering speed. Furthermore, our approach exhibits superior performance even with sparse input data, substantiating its robustness. Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Qing Li 0029, Kanglin Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | VLM2Scene: Self-Supervised Image-Text-LiDAR Learning with Foundation Models for Autonomous Driving Scene UnderstandingabstractVision and language foundation models (VLMs) have showcased impressive capabilities in 2D scene understanding. However, their latent potential in elevating the understanding of 3D autonomous driving scenes remains untapped. In this paper, we propose VLM2Scene, which exploits the potential of VLMs to enhance 3D self-supervised representation learning through our proposed image-text-LiDAR contrastive learning strategy. Specifically, in the realm of autonomous driving scenes, the inherent sparsity of LiDAR point clouds poses a notable challenge for point-level contrastive learning methods. This method often grapples with limitations tied to a restricted receptive field and the presence of noisy points. To tackle this challenge, our approach emphasizes region-level learning, leveraging regional masks without semantics derived from the vision foundation model. This approach capitalizes on valuable contextual information to enhance the learning of point cloud representations. First, we introduce Region Caption Prompts to generate fine-grained language descriptions for the corresponding regions, utilizing the language foundation model. These region prompts then facilitate the establishment of positive and negative text-point pairs within the contrastive loss framework. Second, we propose a Region Semantic Concordance Regularization, which involves a semantic-filtered region learning and a region semantic assignment strategy. The former aims to filter the false negative samples based on the semantic distance, and the latter mitigates potential inaccuracies in pixel semantics, thereby enhancing overall semantic consistency. Extensive experiments on representative autonomous driving datasets demonstrate that our self-supervised method significantly outperforms other counterparts. Codes are available at https://github.com/gbliao/VLM2Scene. Guibiao Liao, Jiankun Li, Xiaoqing Ye |
AAAI | 1 |
| 2024 | 3D Reconstruction and Novel View Synthesis of Indoor Environments Based on a Dual Neural Radiance Field
Zhenyu Bao, Guibiao Liao, Kanglin Liu, Qing Li 0029, Guoping Qiu |
ACM Multimedia | 2 |
| 2024 | OV-NeRF: Open-Vocabulary Neural Radiance Fields With Vision and Language Foundation Models for 3D Semantic UnderstandingabstractThe development of Neural Radiance Fields (NeRFs) has provided a potent representation for encapsulating the geometric and appearance characteristics of 3D scenes. Enhancing the capabilities of NeRFs in open-vocabulary 3D semantic perception tasks has been a recent focus. However, current methods that extract semantics directly from Contrastive Language-Image Pretraining (CLIP) for semantic field learning encounter difficulties due to noisy and view-inconsistent semantics provided by CLIP. To tackle these limitations, we propose OV-NeRF, which exploits the potential of pre-trained vision and language foundation models to enhance semantic field learning through proposed single-view and cross-view strategies. First, from the single-view perspective, we introduce Region Semantic Ranking (RSR) regularization by leveraging 2D mask proposals derived from Segment Anything (SAM) to rectify the noisy semantics of each training view, facilitating accurate semantic field learning. Second, from the cross-view perspective, we propose a Cross-view Self-enhancement (CSE) strategy to address the challenge raised by view-inconsistent semantics. Rather than invariably utilizing the 2D inconsistent semantics from CLIP, CSE leverages the 3D consistent semantics generated from the well-trained semantic field itself for semantic field training, aiming to reduce ambiguity and enhance overall semantic consistency across different views. Extensive experiments validate our OV-NeRF outperforms current state-of-the-art methods, achieving a significant improvement of 20.31% and 18.42% in mIoU metric on Replica and ScanNet, respectively. Furthermore, our approach exhibits consistent superior results across various CLIP configurations, further verifying its robustness. Codes are available at:https://github.com/pcl3dv/OV-NeRF. Guibiao Liao, Kaichen Zhou, Zhenyu Bao, Kanglin Liu, Qing Li 0029 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Rethinking Feature Mining for Light Field Salient Object DetectionabstractLight field salient object detection (LF SOD) has recently received increasing attention. However, most current works typically rely on an individual focal stack backbone for feature extraction. This manner ignores the characteristic of blurred saliency-related regions and contour within focal slices, resulting in insufficient or even inaccurate saliency responses. Aiming at addressing this issue, we rethink the feature mining (i.e., exploration) within focal slices and focus on exploiting informative focal slice features and fully leveraging contour information for accurate LF SOD. First, we observe that the geometric relation between different regions within the focal slices is conducive to useful saliency feature mining if utilized properly. In light of this, we propose an implicit graph learning (IGL) approach. The IGL constructs graph structures to propagate informative geometric relations within the focal slices and all-focus features, and promotes crucial and discriminative focal stack feature mining via graph feature distillation. Second, unlike previous works that rarely utilize contour information, we propose a reciprocal refinement fusion (RRF) strategy. This strategy encourages saliency features and object contour cues to effectively complement each other. Furthermore, a contour hint injection mechanism is introduced to refine the feature expressions. Extensive experiments showcase the superiority of our approach over previous state-of-the-art models with an efficient real-time inference speed. Codes are available at https://github.com/gbliao/IRNet and https://openi.pcl.ac.cn/OpenVision/IRNet . Guibiao Liao, Wei Gao 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | TDRNet: Transformer-Based Dual-Branch Restoration Network for Geometry Based Point Cloud Compression ArtifactsabstractWith the development of 3D point cloud applications, compression plays an important role in lightweight transmission. However, there are barely related works for recovering the compressed point clouds. In this paper, we investigate the geometry-based point cloud compression (G-PCC) artifact removal problem, which is reflected in position offsets and quantity reductions. To address these issues, we propose a novel Transformer-based Dual-branch Restoration Network (TDRNet). First, we design a Transformer Feature Extractor (TFE), which aims to accurately model structural features for position correction and handle inputs of different sizes. Second, a Dual-Branch Restoration (DBR) module is proposed to deeply exploit global shape and local geometry information to restore the quantity reduction distortion. In this way, the proposed TFE and DBR can work cooperatively for the overall compression artifact removal to reconstruct dense point cloud with better completeness and fine-grained details. Experiments show that our proposed TDRNet achieves state-of-the-art results, and our model is expected to provide a baseline for future point cloud compression distortion restoration studies. Xiaoyu Zhang 0002, Guibiao Liao, Wei Gao 0003, Ge Li 0002 |
ICME | 2 |
| 2022 | Unified Information Fusion Network for Multi-Modal RGB-D and RGB-T Salient Object DetectionabstractThe use of complementary information, namely depth or thermal information, has shown its benefits to salient object detection (SOD) during recent years. However, the RGB-D or RGB-T SOD problems are currently only solved independently, and most of them directly extract and fuse raw features from backbones. Such methods can be easily restricted by low-quality modality data and redundant cross-modal features. In this work, a unified end-to-end framework is designed to simultaneously analyze RGB-D and RGB-T SOD tasks. Specifically, to effectively tackle multi-modal features, we propose a novel multi-stage and multi-scale fusion network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the visual color stage doctrine in the human visual system (HVS), the proposed CMFM aims to explore important feature representations in feature response stage, and integrate them into cross-modal features in adversarial combination stage. Moreover, the proposed BMD learns the combination of multi-level cross-modal fused features to capture both local and global information of salient objects, and can further boost the multi-modal SOD performance. The proposed unified cross-modality feature analysis framework based on two-stage and multi-scale information fusion can be used for diverse multi-modal SOD tasks. Comprehensive experiments ($\sim 92\text{K}$image-pairs) demonstrate that the proposed method consistently outperforms the other 21 state-of-the-art methods on nine benchmark datasets. This validates that our proposed method can work well on diverse multi-modal SOD tasks with good generalization and robustness, and provides a good multi-modal SOD benchmark. Wei Gao 0003, Guibiao Liao, Siwei Ma 0001, Ge Li 0002, Yongsheng Liang 0001, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Cross-Collaborative Fusion-Encoder Network for Robust RGB-Thermal Salient Object DetectionabstractWith the prevalence of thermal cameras, RGB-T multi-modal data have become more available for salient object detection (SOD) in complex scenes. Most RGB-T SOD works first individually extract RGB and thermal features from two separate encoders and directly integrate them, which pay less attention to the issue of defective modalities. However, such an indiscriminate feature extraction strategy may produce contaminated features and thus lead to poor SOD performance. To address this issue, we propose a novel CCFENet for a perspective to perform robust and accurate multi-modal expression encoding. First, we propose an essential cross-collaboration enhancement strategy (CCE), which concentrates on facilitating the interactions across the encoders and encouraging different modalities to complement each other during encoding. Such a cross-collaborative-encoder paradigm induces our network to collaboratively suppress the negative feature responses of defective modality data and effectively exploit modality-informative features. Moreover, as the network goes deeper, we embed several CCEs into the encoder, further enabling more representative and robust feature generation. Second, benefiting from the proposed robust encoding paradigm, a simple yet effective cross-scale cross-modal decoder (CCD) is designed to aggregate multi-level complementary multi-modal features, and thus encourages efficient and accurate RGB-T SOD. Extensive experiments reveal that our CCFENet outperforms the state-of-the-art models on three RGB-T datasets with a fast inference speed of 62 FPS. In addition, the advantages of our approach in complex scenarios (e.g., bad weather, motion blur, etc.) and RGB-D SOD further verify its robustness and generality. The source code will be publicly available via our project page:https://git.openi.org.cn/OpenVision/CCFENet. Guibiao Liao, Wei Gao 0003, Ge Li 0002, Junle Wang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionabstractMost existing RGB-D salient object detection (SOD) methods directly extract and fuse raw features from RGB and depth backbones. Such methods can be easily restricted by low-quality depth maps and redundant cross-modal features. To effectively capture multi-scale cross-modal fusion features, this paper proposes a novel Multi-stage and Multi-Scale Fusion Network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the mechanism of visual color stage doctrine in human visual system, the proposed CMFM aims to explore the useful and important feature representations in feature response stage, and effectively integrate them into available cross-modal fusion features in adversarial combination stage. Moreover, the proposed BMD learns the combination of cross-modal fusion features from multiple levels to capture both local and global information of salient objects and further reasonably boost the performance of the proposed method. Comprehensive experiments demonstrate that the proposed method can achieve consistently superior performance over the other 14 state-of-the-art methods on six popular RGB-D datasets when evaluated by 8 different metrics. Guibiao Liao, Wei Gao 0003, Qiuping Jiang, Ronggang Wang, Ge Li 0002 |
ACM Multimedia | 1 |