EDBT 2026 Demo / reviewers in the wild / expert
Bohuan Xue
dblp:251/9055
· DBLP profile ↗
10ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-0332-3967ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MDMU:Multimodal Dynamic Mamba UNet for Multimodal sentiment analysisabstractIn multimodal sentiment analysis, linguistic, visual, and audio sequence data are utilized to assess users’ emotional intensity. However, due to discrepancies in sampling rates across different modalities, sequence alignment poses a significant challenge. While cross-attention-based methods can effectively address this issue, the pairwise attention computation across three modalities incurs substantial computational overhead. To address these challenges, we propose the Multimodal Dynamic Mamba UNet (MDMU) framework, which represents the first integration of an UNet-like structure with Mamba for multimodal sequences modeling. The UNet architecture is employed to capture temporal interactions across modalities, while the Mamba module is utilized for semantic feature modeling, ensuring computational efficiency with near-linear complexity. Additionally, we introduce the Multimodal Momentum Contrast (MMC) method, which also eschews pairwise interactions between the three modalities in favor of a unified approach. MMC facilitates fine-grained fusion between modalities by dynamically constructing a large set of hard negative samples to enhance intermodal interactions. Experimental results demonstrate that even non-strictly aligned temporal interactions benefit the model, offering a novel perspective for multimodal sequence modeling. Our code is available at https://github.com/SCNU-RISLAB/MDMU. Jiazheng Huang, Bohuan Xue, Wenhao Shao |
ICME | 3 |
| 2025 | From Satellite to Street: Semantic and Depth Information for Enhanced Geo-LocalizationabstractAccurate positioning is essential for autonomous driving, but localization using 2D maps is challenging due to the domain gap between perspective view and 2D map. While GNSS accuracy is often limited by atmospheric effects, multipath, and signal blockages. We propose a novel positioning method that combines perspective view images with satellite images retrieved based on rough GNSS positions to achieve precise three-degree-of-freedom (3-DoF) pose estimation. Our method leverages the Swin Transformer for satellite image processing and semantic completion for monocular image analysis. By extracting depth and semantic information from monocular images, we convert these to overhead projections, effectively bridging the gap between different viewpoints. This cross-view transformation allows for precise alignment of features from monocular images onto semantically enriched satellite images. Additionally, we integrate a robust global position estimator using the semantic information from satellite images to further enhance accuracy and robustness. The experimental results demonstrate that our method excels in various complex scenarios; we successfully improved the positioning accuracy within 1 m to 80.67% and the heading in 1° to 33.78%. However, longitudinal localization remains more challenging, with higher errors than lateral positioning. Yilong Zhu, Jianhao Jiao, Hexiang Wei, Jin Wu 0002, Bohuan Xue, Shaojie Shen |
IROS | 5 |
| 2025 | Three-Filters-to-Normal+: Revisiting Discontinuity Discrimination in Depth-to-Normal TranslationabstractThis article introduces three-filters-to-normal$+$(3F2N$+$), an extension of our previous work three-filters-to-normal (3F2N), with a specific focus on incorporating discontinuity discrimination capability into surface normal estimators (SNEs). 3F2N$+$achieves this capability by utilizing a novel discontinuity discrimination module (DDM), which combines depth curvature minimization and correlation coefficient maximization through conditional random fields (CRFs). To evaluate the robustness of SNEs on noisy data, we create a large-scale synthetic surface normal (SSN) dataset containing 20 scenarios (ten indoor scenarios and ten outdoor scenarios with and without random Gaussian noise added to depth images). Extensive experiments demonstrate that 3F2N$+$achieves greater performance than all other geometry-based surface normal estimators, with average angular errors of 7.85$^\circ$, 8.95$^\circ$, 9.25$^\circ$, and 11.98$^\circ$on the clean-indoor, clean-outdoor, noisy-indoor, and noisy-outdoor datasets, respectively. We conduct three additional experiments to demonstrate the effectiveness of incorporating our proposed 3F2N$+$into downstream robot perception tasks, including freespace detection, 6D object pose estimation, and point cloud completion. Our source code and datasets are publicly available at https://mias.group/3F2Nplus.Note to Practitioners—The primary motivation behind this work arises from the need to develop a high-performing surface normal estimator for practical robotics and computer vision applications. While geometry-based surface normal estimators have been widely used in these domains, the existing solutions focus merely on discontinuity discrimination. To tackle this problem, this article introduces a plug-and-play module that leverages both depth curvature and correlation coefficient to quantify discontinuity levels, thereby optimizing surface normal estimation, particularly near or on discontinuous regions. Moreover, this article also introduces a large-scale public dataset with random noise added to depth images, providing a more realistic and robust platform for algorithm evaluation within this research community. Extensive experimental results demonstrate that our method outperforms other state-of-the-art algorithms. Jingwei Yang 0002, Bohuan Xue, Deming Wang, Rui Fan 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Masked PaCONet: Self-Supervised Part-Aware Implicit Shape Reconstruction Scalability, Flexibility, Multi-scale and Semantic ConsistencyabstractLocalized neural implicit representation methods have recently been proven effective for shape reconstruction. However, while some recent neural implicit representation-based approaches have investigated part awareness, there is still room for improvement in leveraging the rich geometry information contained in parts, which is crucial for accurate reconstruction. This study aims to enhance the accuracy of shape reconstruction by incorporatingpart awareness. This principle faces a fundamental technical challenge: manually defining parts across various categories is ambiguous and expensive. To address it, we propose a new self-supervised learning paradigm that automatically discovers meaningful parts. Our proposed paradigm has several prominent advantages as compared with the prior arts: (1) It allows masked part modeling thatscaleswell with available data; (2) It is aflexibleformulation that allows a variable number of parts; (3) It allows the fusion ofmulti-scale(global-level and part-level) features at an arbitrarily given coordinate; (4) Thesemantic consistencyof learned parts leads to transferable features. Extensive experiments validate our approach, named Masked PaCONet, showcasing its superiority in qualitative and quantitative results on public benchmarks, even under challenging settings. Codes and models will be released. Tianyu Liu 0008, Hao Zhao 0002, Bohuan Xue, Guyue Zhou, Ming Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | An Image Acquisition Scheme for Visual Odometry based on Image Bracketing and Online Attribute ControlabstractVisual odometry (VO) system is challenged by complex illumination environments. Image quality and its consistency in the time domain directly determine feature detection and tracking performance, which further affect the robustness and accuracy of the entire system. In this paper, an image acquisition scheme with image bracketing patterns is proposed. Images with different exposure levels are continuously captured to sufficiently explore the scene under varying illumination. An attribute control method is designed to adjust image exposures within the brackets online. Gaussian process regression fits the relationship between image quality metric and exposure via image synthesis technique. The optimal exposures for the next bracket are obtained directly without attempts to ensure a quick response. Experiments show our acquisition system’s effectiveness and performance improvement for VO tasks in complex illumination scenes. Jinhao He, Bohuan Xue, Jin Wu 0002, Pengyu Yin, Jianhao Jiao, Ming Liu 0001 |
ICRA | 3 |
| 2024 | PGO-IPM: Enhance IPM Accuracy with Pose-guided Optimization for Low-cost High-definition Angular Marking Map GenerationabstractHigh-definition angular marking maps (HDAM maps) are vital in large-scale environments with variable appearances. In these scenarios, unmanned ground vehicles (UGVs) can use angular markings for localization because they are easy to identify and informative for localization. However, creating such a marking map relies heavily on manual measurement and annotation, which is time-consuming and laborious. Although Inverse Perspective Mapping (IPM) offers a low-cost and automated alternative, its accuracy is compromised by vehicle motion and the arduous pre-calibration of the IPM matrix. To fill these gaps, we propose a pose-guided optimization framework for IPM. This framework enables the automated generation of HDAM maps, while concurrently refining the preliminary IPM matrix. We deployed the proposed method in two different automated ports, and the method yielded HDAM maps with near-centimeter precision. Moreover, the refined IPM matrix matched the accuracy of manual calibrations. The supplementary materials and videos are available at http://liuhongji.site/PGO-IPM/. Hongji Liu, Linwei Zheng, Xiaoyang Yan, Zhenhua Xu 0003, Bohuan Xue, Yang Yu 0028, Ming Liu 0001 |
IV | 5 |
| 2024 | A lightweight and continuous dimensional emotion analysis system of facial expression recognition under complex background
Jiewen Feng, Jinbo Huang, Qiuchi Xiang, Bohuan Xue |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Robust Embedded Autonomous Driving Positioning System Fusing LiDAR and Inertial SensorsabstractAutonomous driving emphasizes precise multi-sensor fusion positioning on limit resource embedded systems. LiDAR-centered sensor fusion system serves as a mainstream navigation system due to its insensitivity to illumination and viewpoint change. However, these types of systems suffer from handling large-scale sequential LiDAR data using limited resources on board, leading LiDAR-centralized sensor fusion unpractical. As a result, hand-crafted features such as plane and edge are leveraged in majority mainstream positioning methods to alleviate this unsatisfaction, triggering a new cornerstone in LiDAR Inertial sensor fusion. However, such super light weight feature extraction, although it achieves real-time constraint in LiDAR-centered sensor fusion, encounters severe vulnerability under high speed rotational or translational perturbation. In this paper, we propose a sparse tensor based LiDAR Inertial fusion method for autonomous driving embedded system. Leveraging the power of sparse tensor, the global geometrical feature is fetched so that the point cloud sparsity defect is alleviated. Inertial sensor is deployed to conquer the time-consuming step caused by the coarse level point-wise inlier matching. We construct our experiments on both representative dataset benchmarks and realistic scenes. The evaluation results show the robustness and accuracy of our proposed solution compared to classical methods. Zhijian He, Bohuan Xue, Xiangcheng Hu, Zhaoyan Shen, Xiangyue Zeng, Ming Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2023 | D2NT: A High-Performing Depth-to-Normal TranslatorabstractSurface normal holds significant importance in visual environmental perception, serving as a source of rich geometric information. However, the state-of-the-art (SoTA) surface normal estimators (SNEs) generally suffer from an unsatisfactory trade-off between efficiency and accuracy. To resolve this dilemma, this paper first presents a superfast depth-to-normal translator (D2NT), which can directly translate depth images into surface normal maps without calculating 3D coordinates. We then propose a discontinuity-aware gradient (DAG) filter, which adaptively generates gradient convolution kernels to improve depth gradient estimation. Finally, we propose a surface normal refinement module that can easily be integrated into any depth-to-normal SNEs, substantially improving the surface normal estimation accuracy. Our proposed algorithm demonstrates the best accuracy among all other existing real-time SNEs and achieves the SoTA trade-off between efficiency and accuracy. Bohuan Xue, Ming Liu 0001, Rui Fan 0001 |
ICRA | 2 |
| 2023 | Completely Rational $\text{SO}(n)$ OrthonormalizationabstractThe rotation orthonormalization on the special orthogonal group$\text{SO}(n)$, also known as the high dimensional nearest rotation problem, has been revisited. A new generalized simple iterative formula has been proposed that solves this problem in a completely rational manner. Rational operations allow for efficient implementation on various platforms and also significantly simplify the synthesis of large-scale circuitization. The developed scheme is also capable of designing efficient fundamental rational algorithms, for example, quaternion normalization, which outperforms long-exisiting solvers. Furthermore, an$\text{SO}(n)$neural network has been developed for further learning purpose on the rotation group. Simulation results verify the effectiveness of the proposed scheme and show the superiority against existing representatives. Applications show that the proposed orthonormalizer is of potential in robotic pose estimation problems, e.g., hand-eye calibration. Jin Wu 0002, Soheil Sarabandi, Jianhao Jiao, Huaiyang Huang, Bohuan Xue, Ruoyu Geng, Lujia Wang 0001, Ming Liu 0001 |
ICRA | 5 |