VLDB 2026 Research / reviewers in the wild / expert
Guo-Rong Cai
dblp:30/6749 · also Guorong Cai
· DBLP profile ↗
38ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0001-8091-271XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 10 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiDAR-DHMT: LiDAR-Adaptive Dual Hierarchical Mask Transformer for Robust Freespace Detection and Semantic SegmentationabstractInaccurate freespace detection remains a significant challenge to the safety of autonomous driving. However, we observe that current multisource fusion approaches rely on converting LiDAR point clouds into depth maps, often lose crucial 3D geometric cues. This compromises the spatial consistency of predictions, especially in complex urban scenes. To address this limitation, we propose LiDAR-DHMT (LiDAR-Adaptive Dual-branch Hierarchical Mask Transformer), a novel framework designed for spatial-consistent freespace detection and semantic segmentation. Our key innovation lies in the introduction of a 3D Relative Position Bias module, which effectively captures LiDAR’s inherent spatial priors. This is coupled with a Dynamic Bias Attention mechanism that adaptively incorporates the 3D positional cues into the Transformer’s attention computation, enhancing spatial coherence. Additionally, we employ a Mask Interaction module and a global-local fusion strategy to jointly model contextual semantics and fine-grained structural details. Extensive experiments conducted on the KITTI Road, KITTI-360, Cityscapes datasets demonstrate that LiDAR-DHMT consistently outperforms existing state-of-the-art methods, achieving a competitive 97.59% F1 score in freespace detection and 69.45% and 84.4% mIoU in semantic segmentation. Our findings suggest that LiDAR-DHMT offers a practical solution for deploying robust freespace perception in complex urban driving environments. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Huan Chen 0025, Meiliu Wu, Guo-Rong Cai, Jinhe Su |
WACV | 7 |
| 2026 | An efficient single-branch network for semantic scene completion of building spacesabstractSemantic scene completion (SSC) aims to reconstruct a complete 3D scene from a sparse point cloud while simultaneously predicting semantic labels. However, existing SSC methods still struggle with two major types of completion errors. First, when the geometric completion branch lacks an immediate understanding of semantic transitions at object boundaries, it tends to prioritize surface continuity or smoothness, leading to erroneous overfilling. Second, in severely occluded regions, insufficient understanding of scene context and object semantics often leads to implausible or incomplete reconstructions. To address these challenges, we propose Efficient-SSC , an efficient framework composed of three complementary modules. First, we introduce the Structural Semantic Joint ( SSJ ) module, which implicitly fuses structural cues and semantic features through a Structural Semantic Attention mechanism, thereby preserving geometric consistency and reducing superfluous filling in regions with clear boundaries. Second, the Hierarchical Region Structure Extraction ( HRSE ) module captures scene layout and structural relationships through multi-scale graph convolution to provide semantic priors for scene completion, thereby improving completion integrity. Finally, we develop the Awareness Offset Estimation ( AOE ) module, which jointly refines geometric structure and semantic information through a heterogeneous densification strategy, enabling predictive refinement and improving the fidelity of fine-grained scene details. Our network adopts a single-branch architecture for semantic scene completion. The proposed approach is computationally efficient and achieves competitive performance on widely used benchmark datasets, including SSC-PC and NYUCAD-PC. For example, on the SSC-PC dataset, compared with previous single-modality methods, it improves mIoU and mAcc by 2.54% and 1.82% , respectively, while reducing the parameter count by 28% . Duxin Zhu, Jinhe Su, Zhaohong Huang, Ting Han 0001, Jiasheng Su, Guo-Rong Cai |
Neurocomputing | 6 |
| 2025 | EdgeDiff: Edge-aware Diffusion Network for Building Reconstruction from Point CloudsabstractBuilding reconstruction is a challenging problem at the intersection of computer vision, photogrammetry and computer graphics. 3D wireframe presents a compelling representation for building modeling through its compact structure. Existing wireframe reconstruction methods employing vertex detection and edge regression have achieved promising results. In this paper, we develop an Edge-aware Diffusion network, dubbed EdgeDiff. As a novel paradigm for wireframe reconstruction, the EdgeDiff generates wireframe models from noise using a conditional diffusion model. During the training process, the ground truth wireframes firstly are formulated as a set of parameterized edges and then diffused into a random noise distribution. EdgeDiff learns both the noise reversal process and the network structure simultaneously. During inference, EdgeDiff iteratively refines the generated edge distribution using the denoising diffusion implicit model, enabling flexible single- or multi-step denoising and dynamic adaptation to buildings of varying complexity. Additionally, given the unique structure of wireframes, we introduce an edge attention module to extract point-wise attention from point features, using it as auxiliary information to facilitate learning of edge cues and guide the network toward improved edge awareness. Extensive experiments on the real-world Building3D dataset demonstrate that our approach achieves state-of-the-art performance. Yujun Liu 0005, Ruisheng Wang 0001, Shangfeng Huang, Guo-Rong Cai |
CVPR | 4 |
| 2025 | Edge First: Edge-Guided Geometry for Superior 3D Roof Wireframe ReconstructionabstractRoof wireframe reconstruction has shown great success in 3D building reconstruction due to its lightweight nature and straightforward representation. However, previous methods consider all roof points, which result in edge redundancy and omissions. In this paper, we propose a novel and streamlined Edge-guided Geometric wireframe reconstruction framework, named EDGE. We find that points distributed along the roof edges make a significant contribution to the precise geometric structure of wireframe. Therefore, we design an edge point extractor (EPE) to capture the spatial relationship between points and edges, filtering out internal plane points. Moreover, we discover that the previous edge detectors rely solely on corner points, leading to error accumulation. To address this, we present the Hybrid Edge Detector (HED) feeding corner points with edge contextual features, which not only enhances edge completeness but also mitigates edge redundancy. Comprehensive experiments demonstrate EDGE outperforms existing wireframe reconstruction methods with 0.86 Corner F1-score and 0.71 Edge F1-score on Building3D dataset, striking the significant improvement of accuracy between corner and edge. Notably, our EDGE achieves a significant improvement of over 11% in Edge Recall, demonstrating the effectiveness and robustness of the proposed method. Qiaoqiao Hao, Ting Han 0001, Yujun Liu 0005, Shangfeng Huang, Duxin Zhu, Jinhe Su, Yun-Dong Wu, Guo-Rong Cai |
ICASSP | 8 |
| 2025 | Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic Segmentation
Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Meiliu Wu, Guo-Rong Cai, Jinhe Su |
ICCV | 6 |
| 2025 | Depth Matters: Exploring Deep Interactions of RGB-D for Semantic Segmentation in Traffic ScenesabstractRGB-D has gradually become a crucial data source for understanding complex scenes in assisted driving. However, existing studies have paid insufficient attention to the intrinsic spatial properties of depth maps. This oversight significantly impacts the attention representation, leading to prediction errors caused by attention shift issues. To this end, we propose a novel learnable Depth interaction Pyramid Transformer (DiPFormer) to explore the effectiveness of depth. Firstly, we introduce Depth Spatial-Aware Optimization (Depth SAO) as offset to represent real-world spatial relationships. Secondly, the similarity in the feature space of RGB-D is learned by Depth Linear Cross-Attention (Depth LCA) to clarify spatial differences at the pixel level. Finally, an MLP Decoder is utilized to effectively fuse multi-scale features for meeting real-time requirements. Comprehensive experiments demonstrate that the proposed DiPFormer significantly addresses the issue of attention misalignment in both road detection (+7.5%) and semantic segmentation (+4.9% / +1.5%) tasks. DiPFormer achieves state-of-the-art performance on the KITTI (97.57% F-score on KITTI road and 68.74% mIoU on KITTI-360) and Cityscapes (83.4% mIoU) datasets. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Weiquan Liu, Jinhe Su, Zongyue Wang, Guo-Rong Cai |
IROS | 7 |
| 2025 | Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic SegmentationabstractOpen-Vocabulary semantic segmentation (OVSS) and domain generalization in semantic segmentation (DGSS) highlight a subtle complementarity that motivates Open-Vocabulary Domain-Generalized Semantic Segmentation (OV-DGSS). OV-DGSS aims to generate pixel-level masks for unseen categories while maintaining robustness across unseen domains, a critical capability for real-world scenarios such as autonomous driving in adverse conditions. We introduce Vireo, a novel single-stage framework for OV-DGSS that unifies the strengths of OVSS and DGSS for the first time. Vireo builds upon the frozen Visual Foundation Models (VFMs) and incorporates scene geometry via Depth VFMs to extract domain-invariant structural features. To bridge the gap between visual and textual modalities under domain shift, we propose three key components: (1) GeoText Query, which align geometric features with language cues and progressively refine VFM encoder representations; (2) Coarse Mask Prior Embedding (CMPE) for enhancing gradient flow for faster convergence and stronger textual influence; and (3) the Domain-Open-Vocabulary Vector Embedding Head (DOV-VEH), which fuses refined structural and semantic features for robust prediction. Comprehensive evaluation on these components demonstrates the effectiveness of our designs. Our proposed Vireo achieves the state-of-the-art performance and surpasses existing methods by a large margin in both domain generalization and open-vocabulary recognition, offering a unified and scalable solution for robust visual understanding in diverse and dynamic environments. Code is available at https://github.com/SY-Ch/Vireo. Siyu Chen 0004, Ting Han 0001, Chengzheng Fu, Changshe Zhang, Chaolei Wang, Jinhe Su, Guo-Rong Cai, Meiliu Wu |
NeurIPS | 7 |
| 2025 | MTCloud: Multi-type convolutional linkage network for point cloud instance segmentation
Jing Du 0007, Guo-Rong Cai, Zongyue Wang, Jinhe Su, Min Huang 0004, John S. Zelek, José Marcato Junior, Jonathan Li 0001 |
Expert Syst. Appl. | 2 |
| 2025 | A robust few-shot classifier with image as set of pointsabstractAbstract In recent years, many few‐shot classification methods have been proposed. However, only a few of them have explored robust classification, which is an important aspect of human visual intelligence. Humans can effortlessly recognise visual patterns, including lines, circles, and even characters, from image data that has been corrupted or degraded. In this paper, the authors investigate a robust classification method that extends the classical paradigm of robust geometric model fitting. The method views an image as a set of points in a low‐dimensional space and analyses each image through low‐dimensional geometric model fitting. In contrast, the majority of other methods, such as deep learning methods, treat an image as a single point in a high‐dimensional space. The authors evaluate the performance of the method using a noisy Omniglot dataset. The experimental results demonstrate that the proposed method is significantly more robust than other methods. The source code and data for this paper are available at https://github.com/pengsuhua/PMF_OMNIGLOT . Suhua Peng, Zongliang Zhang, Xingwang Huang, Zongyue Wang, Shubing Su, Guo-Rong Cai |
IET Comput. Vis. | 6 |
| 2025 | LVP: Leverage Virtual Points in Multimodal Early Fusion for 3-D Object DetectionabstractDue to the sparsity and occlusion of point clouds, pure point cloud detection has limited effectiveness in detecting such samples. Researchers have been actively exploring the fusion of multimodal data, attempting to address the bottleneck issue based on LiDAR. In particular, virtual points, generated through depth completion from front-view RGB image, offer the potential for better integration with point clouds. Nevertheless, recent approaches fuse these two modalities in the region of interest (RoI), which limits the fusion effectiveness due to the inaccurate RoI region issue in the point cloud’s branch, especially in hard samples. To overcome it and unleash the potential of virtual points, while combining late fusion, we present leverage virtual point (LVP), a high-performance 3-D object detector which LVPs in early fusion to enhance the quality of RoI generation. LVP consists of three early fusion modules: virtual points painting (VPP), virtual points auxiliary (VPA), and virtual points completion (VPC) to achieve point-level fusion and global-level fusion. The integration of these modules effectively improves occlusion handling and improves the detection of distant small objects. In the KITTI benchmark, LVP achieves 85.45% 3-D mAP. As for large dataset nuScenes, we could improve the detection accuracy of large objects by compensating for errors in depth estimation. Without whistles and bells, these results establish LVP as an impressive solution for a 3-D outdoor object detection algorithm. Yidong Chen 0006, Guo-Rong Cai, Ziying Song, Zhaoliang Liu, Binghui Zeng, Jonathan Li 0001, Zongyue Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | CSFNet: Cross-Modal Semantic Focus Network for Semantic Segmentation of Large-Scale Point CloudsabstractSemantic segmentation of large-scale point clouds is an indispensable component of outdoor scene perception, providing essential 3-D semantic insights for applications in scene reconstruction, urban planning, autonomous driving, and more. However, the discriminative capability of point clouds features declines with increasing distance from the sensor, causing current methods to usually perform poorly in segmenting distant objects. To overcome this challenge and improve the differentiation between classes with similar geometric features, we propose the cross-modal semantic focus network (CSFNet). Firstly, we design a multiscale feature dynamic fusion (MDF) module to leverage multiscale image features, thereby enriching the feature representation of point clouds with additional images color and texture information. Then, in order to extract the distinguishing features of distant and different categories of objects more efficiently, we propose a semantic focus module (SFM) that employs a multiclass contrastive learning strategy to enhance feature discrimination. Finally, we introduce cross-modal knowledge distillation (KD) to augment the model’s comprehension of point clouds. Extensive experiments conducted on the SemanticKITTI and nuScenes datasets demonstrate the effectiveness of our method. Notably, our method achieves superior segmentation accuracy across multiple classes at various distances compared to current methods. Yang Luo 0002, Ting Han 0001, Yujun Liu 0005, Jinhe Su, Yiping Chen 0002, Yun-Dong Wu, Guo-Rong Cai |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | HSPFormer: Hierarchical Spatial Perception Transformer for Semantic SegmentationabstractSemantic perception in driving scenarios plays a crucial role in intelligent transportation systems. However, existing Transformer-based semantic segmentation methods often do not fully exploit their potential in understanding driving scene dynamically. These methods typically lack spatial reasoning, failing to effectively correlate image pixels with their spatial positions, leading to attention drift. To address this issue, we propose a novel architecture, the Hierarchical Spatial Perception Transformer (HSPFormer), which integrates monocular depth estimation and semantic segmentation into a unified framework for the first time. We introduce the Spatial Depth Perception Auxiliary Network (SDPNet), a framework for multiscale feature extraction and multilayer depth map prediction to establish hierarchical spatial coherence. Additionally, we design the Hierarchical Pyramid Transformer Network (HPTNet), which uses depth estimation as learnable position embeddings to form spatially correlated semantic representations and generate global contextual information. Experiments on benchmark datasets such as KITTI-360, Cityscapes, and NYU Depth V2, demonstrate that HSPFormer outperforms several state-of-the-art networks, and achieves promising performance with 66.82% top-1 mIoU on KITTI-360, 83.8% mIoU on Cityscapes, and 57.7% mIoU on NYU Depth V2, respectively. The code will be made publicly available athttps://github.com/SY-Ch/HSPFormer. Siyu Chen 0004, Ting Han 0001, Changshe Zhang, Jinhe Su, Ruisheng Wang 0001, Yiping Chen 0002, Zongyue Wang, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | Giving loss a personal course: Universal loss reweighting to improve stereo matching via uncertainty guidance
Yujun Liu 0005, Xiangchen Zhang, Qiaoqiao Hao, Yang Luo 0002, Jinhe Su, Guo-Rong Cai |
Image Vis. Comput. | 6 |
| 2024 | GeoRGS: Geometric Regularization for Real-Time Novel View Synthesis From Sparse InputsabstractWhen the number of available training views is limited, NeRF and 3DGS will soon overfit the optimization and learn the wrong scene geometry. For this challenge, a common solution is to provide depth prior as supervision to correct scene geometry. In this work, we present Geometric Regularized 3D Gaussian Splatting (GeoRGS), a priors-independent method for improving novel view synthesis from sparse inputs. We analyze the problems of the density control strategy in 3DGS with sparse inputs, and find that correcting the erroneous Gaussian growth trend at the beginning of training is effective in mitigating overfitting. Based on this analysis, we propose two geometric regularization methods that do not require prior information. One is based on selecting seed patches of 3D Gaussian from the scene, which guides growth to form correct scene geometry, while the other focuses on regularizing depth similarity between object surfaces and edges. GeoRGS achieves state-of-the-art performance in novel view synthesis from sparse input on LLFF, Blender, RealEstate10K and MipNeRF360 datasets, while also demonstrating significantly faster training speeds and rendering efficiency compared to other baselines. Zhaoliang Liu, Jinhe Su, Guo-Rong Cai, Yidong Chen 0006, Binghui Zeng, Zongyue Wang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Epurate-Net: Efficient Progressive Uncertainty Refinement Analysis for Traffic Environment Urban Road DetectionabstractHigh-performance and real-time road detection plays an essential role in Advanced Driver Assistance Systems (ADAS) of intelligent transportation. However, existing approaches still suffer from ambiguous road contour in traffic environment because deep learning methods lack explicit constraints on road boundaries with similar textures and structures. To address the unsatisfactory boundaries, we propose an efficient architecture for urban road detection to refine road edges adaptively. First, we design a lightweight symmetrical data-fusion network to merge spatial responses into visual features. Second, we construct cross-layer attention transformation to aggregate non-local contextual information. Moreover, a progressive uncertainty analysis module eliminates indistinct road and obstacle edges. Finally, we introduce upgrade uncertainty loss and improved deep supervision to constrain margin error for multi-scale predictions. Results of experiments using three famous datasets confirm the superiority of our method (F1-measure of 96.91% in KITTI, 98.86% in Cityscapes, and 95.18% in R2D, processing speed of 0.02s) over previous approaches. We demonstrate that, to ensure the safety of autonomous driving, the Epurate-Net adaptively refines road contour to reach exquisite road margins. The source code will be available soon. Ting Han 0001, Siyu Chen 0004, Chuanmu Li, Zongyue Wang, Jinhe Su, Min Huang 0004, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Guard-Net: Lightweight Stereo Matching Network via Global and Uncertainty-Aware Refinement for Autonomous DrivingabstractStereo matching is a prominent research area in autonomous driving and computer vision. Despite significant progress made by learning-based methods, accurately predicting disparities in hazardous regions, which is crucial for ensuring safe vehicle operation, remains challenging. The limitations of methods based on Convolutional Neural Networks (CNNs) are most noticeable in textureless regions and repetitive patterns, leading to unreliable predictions. Furthermore, calculating disparities for boundaries and thin structures, where the disparity jump phenomenon is prominent remains difficult. To address these issues, we propose a lightweight stereo matching architecture that focuses on obtaining real-time and high-precision disparity maps in hazardous areas. We exploit an efficient global enhanced path to provide global representations in ill-posed regions, where CNN-based approaches often struggle. Second, our model integrates local and global features to generate more reliable cost volume. Finally, our innovative uncertainty-aware module refines disparity, making full use of high-frequency detailed information and uncertainty attention, effectively preserving complex structures. Comprehensive experimental studies on SceneFlow demonstrate our method outperforms state-of-the-art methods, achieving an End-Point Error (EPE) of 0.47 with only 3.60M parameters. The effectiveness of our method speed-accuracy trade-off is further confirmed by competitive results obtained from the KITTI 2012 and KITTI 2015 experiments. Code is available at: https://github.com/YJLCV/Guard-Net. Yujun Liu 0005, Xiangchen Zhang, Yang Luo 0002, Qiaoqiao Hao, Jinhe Su, Guo-Rong Cai |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | AAEE-Net: Attention-guided aggregation and error-aware enhancement network for accurate and efficient stereo matchingabstractAbstract Stereo matching is a fundamental and long‐standing task in computer vision. Although learning‐based stereo matching algorithms have made remarkable progress, two major challenges still persist. Firstly, existing cost aggregation methods that use stacked three‐dimensional convolutions are complex, leading to heavy computation and memory costs. Secondly these methods continue to struggle with establishing reliable matches in weakly matchable such as that edges and thin structures. To overcome these limitations, we propose an accurate and efficient network called Attention‐guided Aggregation and Error‐aware Enhancement Network (AAEE‐Net). Our approach involves designing an Attention‐guided Aggregation Mechanism (AAM) based on simple image features. This mechanism uses attention weights generated from image features to guide cost aggregation with a more efficient and effective strategy. Additionally, we propose an Error‐aware Enhancement Module (EEM) that refines the raw disparity by combining high‐frequency information from the original image and warp error between the left and right views. EEM enables the network to learn error correction capabilities that produce excellent subtle details and sharp edges. The experimental results on the SceneFlow and KITTI benchmark datasets demonstrate that AAEE‐Net achieves state‐of‐the‐art performance with low inference time. The qualitative results show that AAEE‐Net significantly improves predictions, especially for thin structures. Yujun Liu 0005, Xiangchen Zhang, Jinhe Su, Guo-Rong Cai |
Concurr. Comput. Pract. Exp. | 4 |
| 2023 | 3-D HANet: A Flexible 3-D Heatmap Auxiliary Network for Object Detectionabstract3-D object detection is a vital part of outdoor scene perception. Learning the complete size and accurate positioning of objects from an incomplete point cloud spatial structure is essential to 3-D object detection. We propose a novel flexible 3-D heatmap auxiliary network (3-D HANet) for object detection. To obtain complete structure and location information from an incomplete point cloud structure, we propose a 3-D heatmap to reflect object information. Also, we design a plug-and-play auxiliary network based on 3-D heatmap, which improves the accuracy of the entire detection network without extra computation in the inference stage. We validate the 3-D HANet on the basis of three classic 3-D object detection networks: PointPillars, sparsely embedded convolutional detection (SECOND), and structure aware single-stage 3-D object detection from point cloud (SASSD). Experimental results show that our auxiliary network augments the feature extraction ability of the backbone network, which is manifested in that the predicted boxes and the ground-truth boxes are more suitable in size and more aligned in direction. Furthermore, we conducted verification experiments on the state-of-the-art (SOTA) detector, CasA, and made a further improvement on the official ranking of the KITTI dataset. Qiming Xia, Yidong Chen 0006, Guo-Rong Cai, Guikun Chen, Daoshun Xie, Jinhe Su, Zongyue Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | MSSA-Net: A novel multi-scale feature fusion and global self-attention network for lesion segmentationabstractAbstract In medical image segmentation tasks, it is typical to adopt convolutional neural networks with a serial encoder‐decoder structure. However, mainstream networks cannot simultaneously achieve sufficient extraction of global features and the fusion of multi‐scale information, which may lead to unpromising results for the segmentation of pathological images. Therefore, this article proposed a novel multi‐scale feature fusion and global self‐attention network (MSSA‐Net) for medical image segmentation. Specifically, we designed a parallel double‐encoder network with a multi‐scale feature fusion encoder (MS‐Encoder) and a self‐attention encoder (SA‐Encoder). The SA‐Encoder introduces the transformer's global self‐attention mechanism to extract global features, and the MS‐Encoder adopts atrous spatial pyramid pooling (ASPP) to realize multi‐scale fusion. We have evaluated the proposed MSSA‐Net using three medical segmentation datasets, covering various imaging modalities such as colonoscopy and magnetic resonance imaging. Experiments on the CVC‐ClinicDC, the 2015 MICCAI subchallenge on automatic polyp detection dataset, and anatomical tracings of lesions after stroke (ATLAS) show that our MSSA‐Net outperforms mainstream methods such as DoubleU‐Net and TransUNet. Moreover, MSSA‐Net can predict more accurate segmentation masks, especially in the case of ATLAS, which has challenging images such as multiple shadow areas and discrete lesions. Zhaohong Huang, Xiangchen Zhang, Guo-Rong Cai |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Remote Sensing Image Registration Based on Local Affine Constraint With Circle DescriptorabstractMany methods have been developed to improve the performance of image registration. In this letter, we introduce a novel method based on a local affine constraint for remote sensing image registration, which can be widely used in image processing and pattern recognition. Our algorithm has three components. First, we exploit the scale invariant feature transform (SIFT) method to extract feature points and calculate the gradient magnitude to establish feature descriptors with a circular instead of square neighborhood. Second, an initial matching is implemented by the nearest neighbor distance ratio (NNDR) and the fast sample consensus (FSC) algorithm. Finally, fine registration is established using more correct matches obtained by the local affine transformation circular region search algorithm. Experimental results show that the proposed method achieves subpixel accuracy. In addition, both the correct matching rate and registration demonstrate the effectiveness and efficiency of our method. Chengcai Leng, Guo-Rong Cai, Zhao Pei, Naigong Yu, Anup Basu |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | SSA3D: Semantic Segmentation Assisted One-Stage Three-Dimensional Vehicle Object DetectionabstractOne-stage 3D object detection using mobile light detection and ranging (LiDAR) has developed rapidly in recent years. Specifically, one-stage methods have attracted attention because of their high efficiency and light weight compared with two-stage methods. Inspired by this, we present the semantic segmentation assisted one-stage three-dimensional vehicle object detection (SSA3D), a network for the rapid detection of objects that keeps the advantages of the semantic segmentation module in the two-stage methods without increasing redundant computational load. First, we modified the sampling of the farthest point to improve the quality of the sampling points. This helps to reduce sampling outlier points and bad points that are difficult to perceive in the spatial structure information surrounding the point. Second, a neighbor attention group module is devoted to selectively add extra weight to neighbor points because of the different importance of neighbor points for the corresponding sampling point. Correctly increasing the weight is helpful to obtain richer spatial structure information. Finally, a delicate box generation module is included as a voted center point layer based on the generalized Hoff vote method and an anchor-free regression. We used the feature aggregation module as the backbone and the feature propagation module as the auxiliary network to achieve efficiency. At the same time, the auxiliary network retains the ability to extract point-wise features from the state-of-the-art semantic segmentation network. In experiments, we evaluated and tested the SSA3D on a common KITTI dataset and achieved improved performance in the class of car accuracy. Shangfeng Huang, Guo-Rong Cai, Zongyue Wang, Qiming Xia, Ruisheng Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Convertible Sparse Convolution for Point Cloud Instace SegmentationabstractInstance segmentation based on 3D point cloud is a key step in scene understanding. It is widely used in indoor robot navigation, outdoor autonomous driving, and other fields. But research in this area is still in its infancy. Instance segmentation not only needs to predict the semantic label of each point but also the instance label of each point. Therefore, semantic segmentation can be considered the basis of instance segmentation to some extent. Based on this motivation, we designed a voxel-based branch based on convertible sparse convolution and residual optimization modules. We design a point-based branch so that the network can maintain high-resolution representation. Then the two branches are combined to optimize the semantic segmentation results. Breadth-first search (BFS) performs well in indoor point clouds and is simple to operate. Therefore, we use this clustering operation to group the points of the same instance to obtain the instance segmentation result. The proposed method was tested on the indoor dataset Scan-Net v2 and achieved relatively good instance segmentation precision. Jing Du 0007, Guo-Rong Cai, Zongyue Wang, Jinhe Su, Yun-Dong Wu |
IGARSS | 2 |
| 2021 | Multi-Scale Cascade Guided Object Detection in Aerial ImagesabstractObject detection in aerial images has received increasing attention during the last few years. Scale variation is one of the main challenges in large scene aerial images. Existing object detection pipelines usually detect objects of different scale objects at multiple scale layers. However, the conventional detection approaches with multi-scale density predictions could cause duplicate detections of the same object. In this paper, we proposed a multi-scale cascade guided detection framework (MCGNet) to address these issues by guiding the different scales in detector focus on different scale objects. In particular, we proposed a multi-scale cascade module to predict the different scale objects with an explicit constraint in the loss function. Experiments on benchmark DOTA show promising performance of MCGNet compared with other detectors. Code will be released at https://github.com/jason-su/MCGNET. Jiajia Liao, Yingchao Piao, Guo-Rong Cai, Yun-Dong Wu, Jinhe Su |
IGARSS | 3 |
| 2021 | A Multiple Encoders Network for Stroke Lesion Segmentation
Xiangchen Zhang, Yujun Liu 0005, Jiajia Liao, Guo-Rong Cai, Jinhe Su, Yehua Song |
PRCV (3) | 5 |
| 2021 | Deep 3D caricature face generation with identity and structure consistency
Songzhi Su, Juncong Lin, Guo-Rong Cai, Li Sun 0005 |
Neurocomputing | 4 |
| 2020 | Multi-layer Pointpillars: Multi-layer Feature Abstraction for Object Detection from Point Cloud
Shangfeng Huang, Qiming Xia, Yanhao Lin, Haiyan Lian, Zongyue Wang, Guo-Rong Cai, Jinhe Su |
PRCV (1) | 6 |
| 2019 | Semi-supervised Deep Neural Networks for Object Detection in Video Surveillance Systems
Jinshan Chen, Yujun Liu 0005, Kaiming Ding, Songxin Cai, Jinhe Su, Zongyue Wang, Guo-Rong Cai |
PRCV (1) | 8 |
| 2019 | Multi-scale Convolutional Neural Network Based on 3D Context Fusion for Lesion Detection
Zebiao Wu, Jinshan Chen, Zongyue Wang, Jinhe Su, Guo-Rong Cai |
PRCV (1) | 5 |
| 2019 | Cover patches: A general feature extraction strategy for spoofing detectionabstractSummary Face anti‐spoofing has attracted many attentions in security applications, such as mobile payment and entrance guard. Until now, face anti‐spoofing technique is still a challenging task. Mainstream image‐based spoofing algorithms usually use global motion or texture information to distinguish whether an input face is live or fake. However, the performance of these methods are sensitive in light changes, or images acquired from different sensors. The main reason is that spoofed face image always has slight different texture in local areas, such as landmark or salient region of face. To this end, this paper proposes a novel multi‐patches feature extraction strategy to detect spoofing. First, a set of patches with specific combination scheme is selected to cover the face image. Second, features such as hand‐crafted Gray Level Co‐occurrence Matrix (GLCM), Local Binary Patterns (LBP), or deep features are extracted from these patches. Third, all features are combined as the global descriptor of the face image, then fed into an SVM classifier to verify the anti‐spoofing detection. Experimental results show that the proposed strategy can effectively enhance the performance, concerning with the accuracy of spoofed face detection in four widely used anti‐spoofing databases. Guo-Rong Cai, Songzhi Su, Chengcai Leng, Jipeng Wu, Yun-Dong Wu, Shaozi Li |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Combining 2D and 3D features to improve road detection based on stereo camerasabstractRoad detection is a fundamental component of autonomous driving systems since it provides validspace and candidate regions of objects for driving decision. The core of roaddetection methods is extracting effective and discriminative features. Sincetwo‐dimensional (2D) and 3D features are complementary, the authors propose arobust multi‐feature combination and optimisation framework for stereo imagepairs, called Feature++. First, several 2D and 3D features such as Gabor andplane are, respectively, extracted after the generation of 2D super‐pixel and a3D depth image from stereo matching. Second, the combined features are fed intoa three‐layer shallow neural network classifier to decide whether a super‐pixelis road region or not. Finally, the classified results are further refined usingfully connected conditional random field (CRF), taking the content informationinto consideration. We extensively evaluate the performance of four 2D features,four 3D features, and their combinations. Experiments conducted on the KITTIROAD benchmark show that (i) the combinations of 2D and 3D features greatlyimprove the road detection performance and (ii) using CRF as a refinement stepis necessary. Overall, their proposed ‘Feature + +’ method outperforms mostmanually designed features, and is comparable with state‐of‐the‐art methods thatare based on deep learning methods. Guo-Rong Cai, Songzhi Su, Wenli He, Yun-Dong Wu, Shaozi Li |
IET Comput. Vis. | 1 |
| 2018 | Discriminative parts learning for 3D human action recognition
Min Huang 0004, Guo-Rong Cai, Hongbo Zhang 0002, Sheng Yu 0007, Dong-Ying Gong, Donglin Cao, Shaozi Li, Songzhi Su |
Neurocomputing | 2 |
| 2018 | Multifeature Selection for 3D Human Action RecognitionabstractIn mainstream approaches for 3D human action recognition, depth and skeleton features are combined to improve recognition accuracy. However, this strategy results in high feature dimensions and low discrimination due to redundant feature vectors. To solve this drawback, a multi-feature selection approach for 3D human action recognition is proposed in this paper. First, three novel single-modal features are proposed to describe depth appearance, depth motion, and skeleton motion. Second, a classification entropy of random forest is used to evaluate the discrimination of the depth appearance based features. Finally, one of the three features is selected to recognize the sample according to the discrimination evaluation. Experimental results show that the proposed multi-feature selection approach significantly outperforms other approaches based on single-modal feature and feature fusion. Min Huang 0004, Songzhi Su, Hongbo Zhang 0002, Guo-Rong Cai, Dong-Ying Gong, Donglin Cao, Shaozi Li |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Feature++: Cross dimension feature fusion for road detectionabstractRoad detection is a key component of Advanced Driving Assistance Systems, which provides valid space and candidate regions of objects for vehicles. Mainstream road detection methods have focused on extracting discriminative features. In this paper, we propose a robust feature fusion framework, called “Feature++”, which is combined with superpixel feature and 3D feature extracted from stereo images. Then a neural network classifier is been trained to decide whether a superpixel is road region or not. Finally, the classified results are further refined by conditional random field. Experiments conducted on the KITTI ROAD benchmark show that the proposed “Feature++” method outperforms most manually designed features, and are comparable with state-of-the-art methods that based on deep learning architecture. Wenli He, Guo-Rong Cai, Zhun Zhong, Songzhi Su |
ICASSP | 2 |
| 2017 | Meta-action descriptor for action recognition in RGBD videoabstractAction recognition is one of the hottest research topics in computer vision. Recent methods represent actions based on global or local video features. These approaches, however, lack semantic structure and may not provide a deep insight into the essence of an action. In this work, the authors argue that semantic clues, such as joint positions and part‐level motion clustering, help verify actions. To this end, a meta‐action descriptor for action recognition in RGBD video is proposed in this study. Specifically, two discrimination‐based strategies – dynamic and discriminative part clustering – are introduced to improve accuracy. Experiments conducted on the MSR Action 3D dataset show that the proposed method significantly outperforms the methods without joint position semantic. Min Huang 0004, Songzhi Su, Guo-Rong Cai, Hongbo Zhang 0002, Donglin Cao, Shaozi Li |
IET Comput. Vis. | 3 |
| 2017 | Stratified pooling based deep convolutional neural networks for human action recognition
Sheng Yu 0007, Songzhi Su, Guo-Rong Cai, Shaozi Li |
Multim. Tools Appl. | 4 |
| 2017 | Adaptive total-variation for non-negative matrix factorization on manifold
Chengcai Leng, Guo-Rong Cai, Dongdong Yu, Zongyue Wang |
Pattern Recognit. Lett. | 2 |
| 2015 | Novel Graph Cuts Method for Multi-Frame Super-ResolutionabstractIn this letter, we propose a new graph cuts multi-frame super resolution method. The method is carried out in 3 steps. First, we project each high-resolution pixel p onto the low-resolution images and select low-resolution pixels which fall within the zone of influence of p. Second, we weigh the contribution of the low-resolution pixels via a soft switching function and add them to construct a virtual low resolution pixel. The high resolution image is then recovered after minimizing a Maximum a posteriori Markov Random Field (MAP-MRF) energy function. This is done by approximating our energy function to make it graph representable and minimize it with a graph cuts α-expansion algorithm. Experimental results show that our approach outperforms state-of-the-art methods. Dongxiao Zhang, Pierre-Marc Jodoin, Cuihua Li, Yun-Dong Wu, Guo-Rong Cai |
IEEE Signal Process. Lett. | 5 |
| 2013 | Perspective-SIFT: An efficient tool for low-altitude remote sensing image registration
Guo-Rong Cai, Pierre-Marc Jodoin, Shaozi Li, Yun-Dong Wu, Songzhi Su, Zhenkun Huang |
Signal Process. | 1 |