Bang Du

dblp:232/8241 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DynaGSLAM: Real-Time Gaussian-Splatting SLAM for Online Rendering, Tracking, Motion Predictions of Moving Objects in Dynamic Scenes
abstract
Simultaneous Localization and Mapping (SLAM) is one of the most important environment-perception and navigation algorithms for computer vision, robotics, and autonomous cars/drones. Hence, high quality and fast mapping becomes a fundamental problem. With the advent of 3D Gaussian Splatting (3DGS) as an explicit representation with excellent rendering quality and speed, state-of-the-art (SOTA) works introduce GS to SLAM. Compared to classical pointcloud-SLAM, GS-SLAM generates photometric information by learning from input camera views and synthesizing unseen views with high-quality textures. However, these GS-SLAM fail when moving objects occupy the scene that violates the static assumption of bundle adjustment. The failed updates of moving GS affects the static GS and contaminates the full map over the video sequence. Although some efforts have been made by concurrent works to consider moving objects for GS-SLAM, they simply detect and remove the moving regions from GS rendering ("anti" dynamic GS-SLAM), where only the static background could benefit from GS. To this end, we propose the first real-time GS-SLAM, "DynaGSLAM", that achieves high-quality online GS rendering, tracking, motion predictions of moving objects in dynamic scenes while jointly estimating accurate ego motion. Our DynaGSLAM outperforms SOTA static & "Anti" dynamic GS-SLAM on three dynamic real datasets, while keeping speed and memory efficiency in practice. https://blarklee.github.io/dynagslam/
Runfa Blark Li, Mahdi Shaghaghi, Keito Suzuki, Xinshuang Liu, Varun Moparthi, Bang Du, Walker Curtis, Martin Renschler, Ki Myung Brian Lee, Nikolay Atanasov 0001, Truong Q. Nguyen
WACV6
2026 Few-Shot Medical Image Segmentation With Hierarchical Hypercorrelation Vision State-Space Model
Zheng Cao 0005, Bang Du, Danqing Hu, Wei Zhou 0063, Shangde Gao, Rong Tan, Jun Xu 0005
IEEE Trans. Comput. Soc. Syst.2
2025 Open-Vocabulary Semantic Part Segmentation of 3D Human
abstract
3D part segmentation is still an open problem in the field of 3D vision and AR/VR. Due to limited 3D labeled data, traditional supervised segmentation methods fall short in generalizing to unseen shapes and categories. Recently, the advancement in vision-language models' zero-shot abilities has brought a surge in open-world 3D segmentation methods. While these methods show promising results for$3 D$scenes or objects, they do not generalize well to 3D humans. In this paper, we present the first open-vocabulary segmentation method capable of handling 3D human. Our framework can segment the human category into desired fine-grained parts based on the textual prompt. We design a simple segmentation pipeline, leveraging SAM to generate multi-view proposals in 2D and proposing a novel Human-CLIP model to create unified embeddings for visual and textual inputs. Compared with existing pre-trained CLIP models, the HumanCLIP model yields more accurate embeddings for human-centric contents. We also design a simple-yet-effective MaskFusion module, which classifies and fuses multi-view features into 3D semantic masks without complex voting and grouping mechanisms. The design of decoupling mask proposals and text input also significantly boosts the efficiency of per-prompt inference. Experimental results on various 3D human datasets show that our method outperforms current state-of-the-art open-vocabulary 3D segmentation methods by a large margin. In addition, we showthat our method can be directly applied to various 3D representations including meshes, point clouds, and 3D Gaussian Splatting.
Keito Suzuki, Bang Du, Girish Krishnan, Kunyao Chen, Runfa Blark Li, Truong Q. Nguyen
3DV2
2025 Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac Challenge
abstract
Rib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis.
Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni
IEEE Trans. Medical Imaging14
2024 Select-Sliced Wasserstein Distance for Point Cloud Learning
abstract
Quantifying the discrepancy between point sets is a critical component for point cloud learning tasks. The mainstream point cloud learning tasks utilize Chamfer distance and Earth-Mover’s distance. Chamfer distance is computationally efficient but may not fully capture differences between sets of points. Earth-Mover’s distance, while precise, is computationally expensive and can be impractical to use with high-definition data. Several variants of Sliced Wasserstein distances (SW) are introduced to reduce the computation cost, but bring new problems to the situation: The vanilla SW treats sampled slices equally, resulting in redundant projections; Distributional Sliced Wasserstein distance requires gradient-based optimization, offsetting its benefits. To overcome this limitation and leverage the advantages of Sliced Wasserstein distance over EMD, we propose a novel metric, Select-Sliced Wasserstein distance. This new distance analyzes drawn samples of slices and quantifies their informativeness for each point in a single shot, which eliminates unnecessary projections as well as costly optimizations, but perpetuates the performance. Extensive experiments on various point cloud learning tasks to demonstrate the efficiency and effectiveness of the proposed distance metric. Our code is available at https://github.com/VideoProcessingLab/SSW_Distance
Bang Du, Kunyao Chen, Truong Q. Nguyen
3DV1
2024 Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
abstract
Weize Liu, Guocong Li, Kai Zhang, Bang Du, Qiyuan Chen, Xuming Hu, Hongxia Xu, Jintai Chen, Jian Wu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Weize Liu, Guocong Li, Kai Zhang 0053, Bang Du, Qiyuan Chen 0003, Xuming Hu, Jintai Chen, Jian Wu 0001
NAACL-HLT4
2024 New techniques to identify the tissue of origin for cancer of unknown primary in the era of precision medicine: progress and challenges
abstract
Despite a standardized diagnostic examination, cancer of unknown primary (CUP) is a rare metastatic malignancy with an unidentified tissue of origin (TOO). Patients diagnosed with CUP are typically treated with empiric chemotherapy, although their prognosis is worse than those with metastatic cancer of a known origin. TOO identification of CUP has been employed in precision medicine, and subsequent site-specific therapy is clinically helpful. For example, molecular profiling, including genomic profiling, gene expression profiling, epigenetics and proteins, has facilitated TOO identification. Moreover, machine learning has improved identification accuracy, and non-invasive methods, such as liquid biopsy and image omics, are gaining momentum. However, the heterogeneity in prediction accuracy, sample requirements and technical fundamentals among the various techniques is noteworthy. Accordingly, we systematically reviewed the development and limitations of novel TOO identification methods, compared their pros and cons and assessed their potential clinical usefulness. Our study may help patients shift from empirical to customized care and improve their prognoses.
Wenyuan Ma, Yiran Chen 0021, Bang Du, Mingyu Wan, Xiaolu Ma, Lili Lin, Xinhui Su, Xuanwen Bao, Yifei Shen 0001, Nong Xu, Jian Ruan, Haiping Jiang, Yongfeng Ding
Briefings Bioinform.6
2024 Polygonal Approximation Learning for Convex Object Segmentation in Biomedical Images With Bounding Box Supervision
abstract
As a common and critical medical image analysis task, deep learning based biomedical image segmentation is hindered by the dependence on costly fine-grained annotations. To alleviate this data dependence, in this article, a novel approach, called Polygonal Approximation Learning (PAL), is proposed for convex object instance segmentation with only bounding-box supervision. The key idea behind PAL is that the detection model for convex objects already contains the necessary information for segmenting them since their convex hulls, which can be generated approximately by the intersection of bounding boxes, are equivalent to the masks representing the objects. To extract the essential information from the detection model, a repeated detection approach is employed on biomedical images where various rotation angles are applied and a dice loss with the projection of the rotated detection results is utilized as a supervised signal in training our segmentation model. In biomedical imaging tasks involving convex objects, such as nuclei instance segmentation, PAL outperforms the known models (e.g., BoxInst) that rely solely on box supervision. Furthermore, PAL achieves comparable performance with mask-supervised models including Mask R-CNN and Cascade Mask R-CNN. Interestingly, PAL also demonstrates remarkable performance on non-convex object instance segmentation tasks, for example, surgical instrument and organ instance segmentation.
Jintai Chen, Kai Zhang 0053, Jiahuan Yan, Bang Du, Danny Ziyi Chen, Honghao Gao, Jian Wu 0001
IEEE J. Biomed. Health Informatics7
2023 Efficient Registration for Human Surfaces via Isometric Regularization on Embedded Deformation
abstract
3D registration is a fundamental step to obtain the correspondences between surfaces. Traditional mesh alignment methods tackle this problem through non-rigid deformation, mostly accomplished by applying ICP-based (Iterative Closest Point) optimization. The embedded deformation method is proposed for the purpose of acceleration, which enables various real-time applications. However, it regularizes on an underlying simplified structure, which could be problematic for intricate cases when the simplified graph doesn't fully represent the surface attributes. Moreover, without elaborate parameter-tuning, deformation usually performs suboptimally, leading to slow convergence or a local minimum if all regions on the surface are assumed to share the same rigidity during the optimization. In this article, we propose a novel solution that decouples regularization from the underlying deformation model by explicitly managing the rigidity of vertex clusters. We further design an efficient two-step solution that alternates between isometric deformation and embedded deformation with cluster-based regularization. Our method can easily support region-adaptive regularization with cluster refinement and execute efficiently. Extensive experiments demonstrate the effectiveness of our approach for mesh alignment tasks even under large-scale deformation and imperfect data. Our method outperforms state-of-the-art methods both numerically and visually.
Kunyao Chen, Bang Du, Baichuan Wu, Truong Q. Nguyen
IEEE Trans. Vis. Comput. Graph.3
2021 In Defense of Scene Graphs for Image Captioning
abstract
The mainstream image captioning models rely on Convolutional Neural Network (CNN) image features to generate captions via recurrent models. Recently, image scene graphs have been used to augment captioning models so as to leverage their structural semantics, such as object entities, relationships and attributes. Several studies have noted that the naive use of scene graphs from a black-box scene graph generator harms image captioning performance and that scene graph-based captioning models have to incur the overhead of explicit use of image features to generate decent captions. Addressing these challenges, we propose SG2Caps, a framework that utilizes only the scene graph labels for competitive image captioning performance. The basic idea is to close the semantic gap between the two scene graphs - one derived from the input image and the other from its caption. In order to achieve this, we leverage the spatial location of objects and the Human-Object-Interaction (HOI) labels as an additional HOI graph. SG2Caps outperforms existing scene graph-only captioning models by a large margin, indicating scene graphs as a promising representation for image captioning. Direct utilization of scene graph labels avoids expensive graph convolutions over high-dimensional CNN features resulting in 49% fewer trainable parameters. Our code is available at: https://github.com/Kien085/SG2Caps
Kien Nguyen 0006, Subarna Tripathi, Bang Du, Tanaya Guha, Truong Q. Nguyen
ICCV3
2021 Mesh Completion with Virtual Scans
abstract
Meshes generated by range scanners are often incomplete and contain complex holes due to limited input coverage and occlusion. In this paper, we present an effective method to fill the gap regions on meshes by leveraging the templates inferred from the learning-based method. We first segment both source and template models into corresponding parts. Each part will be aligned with non-rigid deformation. We then modify the gap regions by “virtual” depth maps rendered using the aligned parts from newly selected viewpoints. Comparing with the template-based mesh completion approaches, our algorithm can generate natural appearances without any user interaction. Comparing with the state-of-the-art volumetric fusion methods, our approach supports selective blending, which only modifies the regions of interest and prevents bad template inference from impacting the source.
Kunyao Chen, Baichuan Wu, Bang Du, Truong Q. Nguyen
ICIP4