Suqin Bai

dblp:169/5367 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A robotic grasping method of box-shaped objects based on Dual-Stream You Only Look Once framework
Jiale Cui, Suqin Bai
Eng. Appl. Artif. Intell.4
2025 SYN-PBOX: A large-scale benchmark dataset of box-shaped objects for scene understanding of bin picking
Jiale Cui, Caisheng Liu, Suqin Bai
Neurocomputing4
2025 SRDA: Self-regulating distribution alignment based on prompt learning for unsupervised domain adaptation
Yun Cui, Suqin Bai
Neurocomputing5
2025 Psg-6d: prior-free implicit category-level 6D pose estimation with SO(3)-equivariant network and point cloud global enhancement
Hanqi Jiang, Yongjie Gao, Suqin Bai
Multim. Syst.5
2025 Extended Receptive Field UDA Semantic Segmentation Based on Spatial Alignment and Knowledge Distillation
abstract
In recent years, unsupervised domain adaptation (UDA) has significantly advanced, addressing the issue of requiring large amounts of labeled data in deep learning. Some UDA strategies have effectively alleviated domain shift, but they are still struggling to tackle challenges such as the model’s inability to continuously extract features, the difficulty to achieve better segmentation boundaries in the target domain, and tend to neglect previously acquired knowledge. To address these issues, we propose an extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation (ERF). Firstly, based on the idea of combining serial and parallel, we design a novel large continuous receptive field decoder (largeCF) to extract large continuous receptive field features. This approach alleviates the bias of model in feature extraction between objects of different sizes and simultaneously reduces model complexity. Secondly, we propose an edge consistency strategy that aligns edge features and matches the spatial arrangement between predicted and ground truth labels, improving edges segmentation accuracy of the target domain. Finally, we employ a knowledge distillation module to achieve an optimized student-teacher framework, where the teacher effectively guides the student to retain previously learned information, resulting in more accurate segmentation of the target domain. Experimental results demonstrated the effectiveness of the proposed approach, which achieved mIoU of 76.5% and 68.8% on UDA benchmark tasks GTA$\rightarrow $CityScapes and SYNTHIA$\rightarrow $CityScapes, respectively. The code is available at:https://github.com/fz-ss/ERF. Note to Practitioners—This paper focuses on the challenges of semantic segmentation in autonomous driving, particularly facing the issue of extensive manual annotation required for dense semantic labels. We propose a novel extended receptive field UDA semantic segmentation based on spatial alignment and knowledge distillation. The article begins by outlining the initial implementation process of UDA, laying the foundation for subsequent in-depth discussions. Subsequently, we conduct a theoretical analysis of the Large Continuous Decoder, Boundary Consistency Strategy, and Knowledge Distillation Scheme, which constitute the core components of our method. Finally, experiments on two UDA benchmark tasks demonstrate the feasibility of our approach. However, the performance gap between UDA and supervised semantic segmentation still exists. In future research, we will focus on reducing the feature gap between different domains. Additionally, we will strive to fully use image features and distilled features to make greater progress, thereby driving advancements in semantic segmentation for autonomous driving.
Yunna Song, Caisheng Liu, Suqin Bai, Xin Shu 0001, Yunhan Sun
IEEE Trans Autom. Sci. Eng.4
2025 Ellipsoid-SLAM: enhancing dynamic scene understanding through ellipsoidal object representation and trajectory tracking
Haowei Zhu, Suqin Bai, Shucheng Huang
Vis. Comput.2
2025 IOFusion: instance segmentation and optical-flow guided 3D reconstruction in dynamic scenes
Haowei Zhu, Suqin Bai, Chenggen Wang, Yunhan Sun, Shucheng Huang
Vis. Comput.2
2024 Transformer framework for depth-assisted UDA semantic segmentation
Yunna Song, Danping Zou, Caisheng Liu, Suqin Bai, Yunhan Sun
Eng. Appl. Artif. Intell.5
2024 Clear-Plenoxels: Floaters free radiance fields without neural networks
Weichen Yang, Suqin Bai, Zhen Ou, Yunhan Sun
Knowl. Based Syst.3
2024 EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation
abstract
Estimating the 6-D poses of objects from RGB-D images holds great potential for several applications. However, given that the 6-D pose estimation accuracy is significantly affected by occlusion and noise between the objects in an image, this paper proposes a novel 6-D pose estimation method based on Efficient feature extraction and Point-pair feature matching. Specifically, we develop the Efficient channel attention Convolutional Neural Network (ECNN) and SO(3)-Encoder modules to extract 2-D features from the RGB image and SO(3)-equivariant features from the depth image, respectively. These features are fused in the DenseFusion module to obtain 3-D features in the camera space. Meanwhile, we exploit CAD model priors to obtain 3-D features in the model space through the model feature encoder, and then we globally regress the 3-D features in the camera and model space. According to these features, we generate oriented point clouds in each space, and then conduct point-pair feature matching to obtain pose information. Finally, we perform direct pose regression on the 3-D features in the camera and model space, and then resulting point-pair feature matching pose information is combined with the direct point-wise pose regression information to enhance pose prediction accuracy. Experimental results on three widely used benchmarking datasets demonstrate that our method achieves state-of-the-art performance, particularly for severe occluded scenes.
Danping Zou, Xin Shu 0001, Suqin Bai, Haowei Zhu, Yunhan Sun
IEEE Trans. Multim.5
2021 A self-supervised method of single-image depth estimation by feeding forward information using max-pooling layers
Yunhan Sun, Suqin Bai, Zhengxing Sun, Zhaohui Tian
Vis. Comput.3
2020 Single View Depth Estimation via Dense Convolution Network with Self-supervision
Yunhan Sun, Suqin Bai, Zhengxing Sun
MMM (2)3
2018 3D reconstruction framework via combining one 3D scanner and multiple stereo trackers
Zhengxing Sun, Suqin Bai
Vis. Comput.3