Boliang Guan

dblp:276/6973 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-5221-8018ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Visual Boundary-Guided Pseudo-Labeling for Weakly Supervised 3D Point Cloud Segmentation in Indoor Environments
abstract
Accurate segmentation of 3D point clouds in indoor scenes remains a challenging task, often hindered by the labor-intensive nature of data annotation. While weakly supervised learning approaches have shown promise in leveraging partial annotations, they frequently struggle with imbalanced performance between foreground and background elements due to the complex structures and proximity of objects in indoor environments. To address this issue, we propose a novel foreground-aware label enhancement method utilizing visual boundary priors. Our approach projects 3D point clouds onto 2D planes and applies 2D image segmentation to generate pseudo-labels for foreground objects. These labels are subsequently back-projected into 3D space and used to train an initial segmentation model. We further refine this process by incorporating prior knowledge from projected images to filter the predicted labels, followed by model retraining. We introduce this technique as the Foreground Boundary Prior (FBP), a versatile, plug-and-play module designed to enhance various weakly supervised point cloud segmentation methods. We demonstrate the efficacy of our approach on the widely-used 2D-3D-Semantic dataset, employing both random-sample and bounding-box based weak labeling strategies. Our experimental results show significant improvements in segmentation performance across different architectural backbones, highlighting the method's effectiveness and portability.
Zhuo Su 0001, Yudi Tan, Boliang Guan, Fan Zhou 0001
IEEE Trans. Vis. Comput. Graph.4
2023 OaIF: Occlusion-Aware Implicit Function for Clothed Human Re-construction
abstract
Abstract Clothed human re‐construction from a monocular image is challenging due to occlusion, depth‐ambiguity and variations of body poses. Recently, shape representation based on an implicit function, compared to explicit representation such as mesh and voxel, is more capable with complex topology of clothed human. This is mainly achieved by using pixel‐aligned features, facilitating implicit function to capture local details. But such methods utilize an identical feature map for all sampled points to get local features, making their models occlusion‐agnostic in the encoding stage. The decoder, as implicit function, only maps features and does not take occlusion into account explicitly. Thus, these methods fail to generalize well in poses with severe self‐occlusion. To address this, we present OaIF to encode local features conditioned in visibility of SMPL vertices. OaIF projects SMPL vertices onto image plane to obtain image features masked by visibility. Vertices features integrated with geometry information of mesh are then feed into a GAT network to encode jointly. We query hybrid features and occlusion factors for points through cross attention and learn occupancy fields for clothed human. The experiments demonstrate that OaIF achieves more robust and accurate re‐construction than the state of the art on both public datasets and wild images.
Yudi Tan, Boliang Guan, Fan Zhou 0001, Zhuo Su 0001
Comput. Graph. Forum2
2022 Multistage Spatio-Temporal Networks for Robust Sketch Recognition
abstract
Sketch recognition relies on two types of information, namely, spatial contexts like the local structures in images and temporal contexts like the orders of strokes. Existing methods usually adopt convolutional neural networks (CNNs) to model spatial contexts, and recurrent neural networks (RNNs) for temporal contexts. However, most of them combine spatial and temporal features with late fusion or single-stage transformation, which is prone to losing the informative details in sketches. To tackle this problem, we propose a novel framework that aims at the multi-stage interactions and refinements of spatial and temporal features. Specifically, given a sketch represented by a stroke array, we first generate a temporal-enriched image (TEI), which is a pseudo-color image retaining the temporal order of strokes, to overcome the difficulty of CNNs in leveraging temporal information. We then construct a dual-branch network, in which a CNN branch and a RNN branch are adopted to process the stroke array and the TEI respectively. In the early stages of our network, considering the limited ability of RNNs in capturing spatial structures, we utilize multiple enhancement modules to enhance the stroke features with the TEI features. While in the last stage of our network, we propose a spatio-temporal enhancement module that refines stroke features and TEI features in a joint feature space. Furthermore, a bidirectional temporal-compatible unit that adaptively merges features in opposite temporal orders, is proposed to help RNNs tackle abrupt strokes. Comprehensive experimental results on QuickDraw and TU-Berlin demonstrate that the proposed method is a robust and efficient solution for sketch recognition.
Xudong Jiang 0001, Boliang Guan, Ruomei Wang 0001, Nadia Magnenat-Thalmann
IEEE Trans. Image Process.3
2021 Efficient Sketch Recognition Via Compact Spatial Embedding Graph Neural Networks
abstract
Sketches are descriptive, high-level visual media in many systems and applications. However, current methods for sketch recognition are mainly based on large neural networks, which have millions of parameters and are too cumbersome to be deployed on edge devices. Besides, convolutional neural networks are not optimal, since many areas in sketches are blank and without any information. Hence, this paper aims at designing an efficient network that maintains the state-of-the-art accuracy. Our solution is a novel graph neural network that utilizes densely connected grouped convolutions on the graph representation of sketches. It allows us to extract and aggregate spatio-temporal features efficiently. Moreover, a compact spatial embedding module is introduced to explore spatial contexts, and consequently facilitates recognition. In this way, our network is small (about 10% parameters of ResNet- 18) and efficient (inference speed in 800 ~ 2,000 FPS), meanwhile has the competitive accuracy on the QuickDraw dataset.
Xudong Jiang 0001, Boliang Guan, Nadia Magnenat-Thalmann
ICME3
2021 LGCPNet : Local-global combined point-based network for shape segmentation
Boliang Guan, Fan Zhou 0001, Shujin Lin, Ruomei Wang 0001
Comput. Graph.1
2021 Joint Feature Optimization and Fusion for Compressed Action Recognition
abstract
Recent methods including CoViAR and DMC-Net provide a new paradigm for action recognition since they are directly targeted at compressed videos (e.g., MPEG4 files). It avoids the cumbersome decoding procedure of traditional methods, and leverages the pre-encoded motion vectors and residuals in compressed videos to complete recognition efficiently. However, motion vectors and residuals are noisy, sparse and highly correlated information, which cannot be effectively exploited by plain and separated networks. To tackle these issues, we propose a joint feature optimization and fusion framework that better utilizes motion vectors and residuals in the following three aspects. (i) We model the feature optimization problem as a reconstruction process that represents features by a set of bases, and propose a joint feature optimization module that extracts bases in the both modalities. (ii) A low-rank non-local attention module, which combines the non-local operation with the low-rank constraint, is proposed to tackle the noise and sparsity problem during the feature reconstruction process. (iii) A lightweight feature fusion module and a self-adaptive knowledge distillation method are introduced, which use motion vectors and residuals to generate predictions similar to those from networks with optical flows. With these proposed components embedded in a baseline network, the proposed network not only achieves the state-of-the-art performance on HMDB-51 and UCF-101, but also maintains its advantage in computational complexity.
Xudong Jiang 0001, Boliang Guan, Raymond Rui Ming Tan, Ruomei Wang 0001, Nadia Magnenat-Thalmann
IEEE Trans. Image Process.3
2020 Voxel-based quadrilateral mesh generation from point cloud
Boliang Guan, Shujin Lin, Ruomei Wang 0001, Fan Zhou 0001, Yongchuan Zheng
Multim. Tools Appl.1