Bohan Ren

dblp:276/6083 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-4664-267XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RoadsideSplat: Robust 3D Gaussian Reconstruction from Monocular Roadside Surveillance
abstract
Reconstructing dynamic roads from roadside traffic surveillance cameras is crucial for smart cities and digital twin applications. While the latest monocular depth estimation methods demonstrate strong performance, they exhibit instability in roadside scenarios. Existing reconstruction approaches for autonomous driving scenes predominantly adopt vehicle-mounted perspectives, accumulating vehicle point clouds from per-frame depth maps using 3D bounding boxes. These point clouds are used to initialize the center positions and colors of 3D Gaussians to improve reconstruction performance. However, the compressed depth discrepancy between vehicles and road surfaces in roadside views leads to model confusion between vehicle and background depth estimations. To address these challenges, we propose a robust reconstruction framework based on a single fixed RGB traffic camera. Differing from conventional frame-wise depth prediction followed by 3D box-based accumulation, our method processes masked vehicle fore-ground sequences through existing models, directly predicting complete vehicle point clouds via local feature matching and global alignment while iteratively refining 3D boxes to enhance reconstruction quality. Leveraging the explicit nature of 3D Gaussians for scene editing, we introduce simple yet effective road constraints to mitigate penetration artifacts during scene manipulation. Extensive evaluations on the TUMTraf-V2X and RCooper datasets under monocular roadside settings validate the effectiveness of our approach.
Zhaoxiang Liang, Wenjun Guo, Bohan Ren, Yi Yang 0009
IROS3
2025 Automated 3D-GS Registration and Fusion via Skeleton Alignment and Gaussian-Adaptive Features
abstract
In recent years, 3D Gaussian Splatting (3D-GS)based scene representation demonstrates significant potential in real-time rendering and training efficiency. However, most existing methods primarily focus on single-map reconstruction, while the registration and fusion of multiple 3D-GS submaps remain underexplored. Existing methods typically rely on manual intervention to select a reference sub-map as a template and use point cloud matching for registration. Moreover, hard-threshold filtering of 3D-GS primitives often degrades rendering quality after fusion. In this paper, we present a novel approach for automated 3D-GS sub-map alignment and fusion, eliminating the need for manual intervention while enhancing registration accuracy and fusion quality. First, we extract geometric skeletons across multiple scenes and leverage ellipsoid-aware convolution to capture 3D-GS attributes, facilitating robust scene registration. Second, we introduce a multi-factor Gaussian fusion strategy to mitigate the scene element loss caused by rigid thresholding. Experiments on the ScanNet-GSReg and our Coord datasets demonstrate the effectiveness of the proposed method in registration and fusion. For registration, it achieves a 41.9% reduction in RRE on complex scenes, ensuring more precise pose estimation. For fusion, it improves PSNR by 10.11 dB, highlighting superior structural preservation. These results confirm its ability to enhance scene alignment and reconstruction fidelity, ensuring more consistent and accurate 3D scene representation for robotic perception and autonomous navigation.
Shiyang Liu, Dianyi Yang, Yu Gao 0040, Bohan Ren, Yi Yang 0009, Mengyin Fu
IROS4
2025 OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding
abstract
Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise 3D object-level understanding given open-ended queries. In this paper, we present OpenGS-Fusion, an innovative open-vocabulary dense mapping framework that improves semantic modeling and refines object-level understanding. OpenGS-Fusion combines 3D Gaussian representation with a Truncated Signed Distance Field to facilitate lossless fusion of semantic features on-the-fly. Furthermore, we introduce a novel multimodal language-guided approach named MLLM-Assisted Adaptive Thresholding, which refines the segmentation of 3D objects by adaptively adjusting similarity thresholds, achieving an improvement 17% in 3D mIoU compared to the fixed threshold strategy. Extensive experiments demonstrate that our method outperforms existing methods in 3D object understanding and scene reconstruction quality, as well as showcasing its effectiveness in language-guided scene interaction. The code is available at https://young-bit.github.io/opengs-fusion.github.io/.
Dianyi Yang, Xihan Wang, Yu Gao 0040, Shiyang Liu, Bohan Ren, Yufeng Yue, Yi Yang 0009
IROS5
2025 Transferring Prior Thermal Knowledge for Snowy Urban Scene Semantic Segmentation
abstract
RGB-thermal (RGB-T) semantic segmentation enables intelligent vehicles to understand environments while operating in urban scenes. However, the research encounters two main challenges: 1) scarcity of training samples under snowy conditions and 2) challenge in applying the model in practice. To address the first challenge, we proposed a publicly accessible RGB-T semantic segmentation dataset in snowy urban scenes (SUS dataset). The SUS dataset comprises 1035 pairs of precisely registered RGB-T images, and provides pixel-level semantic annotations for five categories for all images. To tackle the second challenge, we introduced MCNet-S*, a novel semantic segmentation model that leverages knowledge distillation (KD). The KD structure consists of an RGB-T teacher model, named MCNet-T, and an RGB student model, named MCNet-S. Within MCNet-T, we proposed a cross-modal dual association (CDA) module to enhance utilization of RGB-T information in snowy urban scenes. Within MCNet-S, a depth-wise separable pyramid (DSP) module was proposed to improve the efficiency of RGB information utilization and align the feature dimensions with those of MCNet-T. Between MCNet-S and MCNet-T, memory-based contrastive learning distillation (MCLD) was proposed to transfer the prior thermal knowledge, improving the segmentation accuracy of MCNet-S and obtaining optimized MCNet-S*. Extensive experiments on the SUS and MFNet datasets show that the proposed models outperform state-of-the-art models. The SUS dataset and codes are available at https://github.com/xiaodonguo/SUS_dataset.
Tong Liu 0009, Yefeng Mou, Bohan Ren, Wujie Zhou
IEEE Trans. Intell. Transp. Syst.5
2025 CQformer: Learning Dynamics Across Slices in Medical Image Segmentation
abstract
Prevalent studies on deep learning-based 3D medical image segmentation capture the continuous variation across 2D slices mainly via convolution, Transformer, inter-slice interaction, and time series models. In this work, via modeling this variation by an ordinary differential equation (ODE), we propose a cross instance query-guided Transformer architecture (CQformer) that leverages features from preceding slices to improve the segmentation performance of subsequent slices. Its key components include a cross-attention mechanism in an ODE formulation, which bridges the features of contiguous 2D slices of the 3D volumetric data. In addition, a regression head is employed to shorten the gap between the bottleneck and the prediction layer. Extensive experiments on 7 datasets with various modalities (CT, MRI) and tasks (organ, tissue, and lesion) demonstrate that CQformer outperforms previous state-of-the-art segmentation algorithms on 6 datasets by 0.44%-2.45%, and achieves the second highest performance of 88.30% on the BTCV dataset. The code is available at https://github.com/qbmizsj/CQformer.
Xiang Chen 0031, Bohan Ren, Haibo Yang 0002, Xiao-Yong Zhang, Yuan Zhou 0004
IEEE Trans. Medical Imaging5
2023 A-GCL: Adversarial graph contrastive learning for fMRI analysis to diagnose neurodevelopmental disorders
Xiang Chen 0031, Bohan Ren, Haibo Yang 0002, Xi Jiang 0001, Dinggang Shen, Yuan Zhou 0004, Xiao-Yong Zhang
Medical Image Anal.4
2023 TW-Net: Transformer Weighted Network for Neonatal Brain MRI Segmentation
abstract
Accurate neonatal brain MRI segmentation is valuable for investigating brain growth patterns and tracking the progression of neurodevelopmental disorders. However, it is a challenging task to use intensity-based methods to segment neonatal brain structures because of small contrast differences between brain regions caused by the inherent myelination process. Although convolutional neural networks offer the potential to segment brain structures in an intensity-independent manner, they suffer from lack of in-plane long-range dependency which is essential for the segmentation. To solve this problem, we propose a novel Transformer-Weighted network (TW-Net) to incorporate in-plane long-range dependency information. TW-Net employs a conventional encoder-decoder architecture with a Transformer module in the middle. The Transformer module uses a rotate-and-flip layer to better calculate the similarity between two patches in a slice to leverage similar patterns of geometrical and texture features within brain structures. In addition, a deep supervision module and squeeze-and-excitation blocks are introduced to incorporate boundary information of brain structures. Compared with state-of-the-art deep learning algorithms, TW-Net outperforms these methods for multiple-label tasks in 2D and 2.5D configurations on two independent public datasets, demonstrating that TW-Net is a promising method for neonatal brain MRI segmentation.
Bohan Ren, Haibo Yang 0002, Xiaoyang Han, Xiang Chen 0031, Yuan Zhou 0004, Dinggang Shen, Xiao-Yong Zhang
IEEE J. Biomed. Health Informatics2
2022 3D Global Fourier Network for Alzheimer's Disease Diagnosis Using Structural MRI
Xiang Chen 0031, Bohan Ren, Haibo Yang 0002, Xiao-Yong Zhang, Yuan Zhou 0004
MICCAI (1)3