EDBT 2026 Demo / reviewers in the wild / expert
Shuo Gu
dblp:146/1376
· DBLP profile ↗
22ranked-venue papers
8as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 9 since 2021Systems, architecture and hardware · 9 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Trackingabstract3D multiple object tracking (MOT) plays a crucial role in autonomous driving perception. Recent end-to-end query-based trackers simultaneously detect and track objects, which have shown promising potential for the 3D MOT task. However, existing methods are still in the early stages of development and lack systematic improvements, failing to track objects in certain complex scenarios, like occlusions and the small size of target object’s situations. In this paper, we first summarize the current end-to-end 3D MOT framework by decomposing it into three constituent parts: query initialization, query propagation, and query matching. Then we propose corresponding improvements, which lead to a strong yet simple tracker: S2-Track. Specifically, for query initialization, we present 2D-Prompted Query Initialization, which leverages predicted 2D object and depth information to prompt an initial estimate of the object’s 3D location. For query propagation, we introduce an Uncertainty-aware Probabilistic Decoder to capture the uncertainty of complex environment in object prediction with probabilistic attention. For query matching, we propose a Hierarchical Query Denoising strategy to enhance training robustness and convergence. As a result, our S2-Track achieves state-of-the-art performance on nuScenes benchmark, i.e., 66.3% AMOTA on test split, surpassing the previous best end-to-end solution by a significant margin of 8.9% AMOTA. We achieve 1st place on the nuScenes tracking task leaderboard. Pengkun Hao, Kalok Ho, Shuo Gu, Zhihui Hao, Kun Zhan, Peng Jia 0007, Xianpeng Lang, Xiaodan Liang |
ICML | 6 |
| 2025 | Off-Road Freespace Detection with LiDAR-Camera Fusion and Self-Distillation
Shuo Gu |
ICRA | 1 |
| 2025 | PCMF2-Net: A Pyramid Cross-Modal Feature Fusion Network for Off-Road Freespace DetectionabstractFreespace detection plays an important role in autonomous driving. In recent years, deep learning based freespace detection methods have performed well in urban scenes. However, for off-road scenes, freespace detection poses significant challenges due to the complexity of the scenes and the lack of clear edges. The existing methods have not effectively fused LiDAR data and camera images. In this paper, we propose a Pyramid Cross-Modal Feature Fusion Network (PCMF2-Net) for off-road freespace detection. The dense depth maps are concatenated with RGB images and used as input along with surface normal maps. The dual branch CNN-Transformer encoder combines convolutional neural networks and transformers to extract local and global features from RGBD images and surface normal maps, respectively. Then, in the pyramid cross-modal feature fusion module, the multi-scale and multimodal encoder features are fused in a top-down manner. In addition, we also use an edge segmentation task and a two-step training strategy to further improve performance. Experiments on the off-road freespace detection dataset (ORFD) demonstrate that the proposed PCMF2-Net achieves a competitive result of 93.9% IoU at a speed of 23 Hz. Chunpeng Lu, Shuo Gu, Yigong Zhang, Hui Kong 0001 |
IROS | 3 |
| 2024 | SGNet: Salient Geometric Network for Point Cloud RegistrationabstractPoint Cloud Registration (PCR) is a critical and challenging task in computer vision and robotics. One of the primary difficulties in PCR is identifying salient and meaningful points that exhibit consistent semantic and geometric properties across different scans. Previous methods have encountered challenges with ambiguous matching due to the similarity among patch blocks throughout the entire point cloud and the lack of consideration for efficient global geometric consistency. To address these issues, we propose a new framework that includes several novel techniques. Firstly, we introduce a semantic-aware geometric encoder that combines object-level and patch-level semantic information. This encoder significantly improves registration recall by reducing ambiguity in patch-level superpoint matching. Additionally, we incorporate a prior knowledge approach that utilizes an intrinsic shape signature to identify salient points. This enables us to extract the most salient super points and meaningful dense points in the scene. Secondly, we introduce an innovative transformer that encodes High-Order (HO) geometric features. These features are crucial for identifying salient points within initial overlap regions while considering global high-order geometric consistency. We introduce an anchor node selection strategy to optimize this high-order transformer further. By encoding inter-frame triangle or polyhedron consistency features based on these anchor nodes, we can effectively learn high-order geometric features of salient super points. These high-order features are then propagated to dense points and utilized by a Sinkhorn matching module to identify critical correspondences for successful registration. The experiments conducted on the 3DMatch/3DLoMatch and KITTI datasets demonstrate the effectiveness of our method. Qianliang Wu, Yaqing Ding 0001, Lei Luo 0001, Haobo Jiang, Shuo Gu, Chuanwei Zhou, Jin Xie 0001, Jian Yang 0003 |
IROS | 5 |
| 2024 | A Practical Method to Detect Evaporation Ducts Based on BDS-3 Signals Received by a Single AntennaabstractEvaporation ducts have a major impact on the antennas that receive and transmit radar signals, and these ducts could support over-the-horizon radar detection at sea. Therefore, the detection of evaporation ducts has significant military value. Global navigation satellite systems (GNSSs) are potential tools for duct remote sensing because they can perform stealthy passive measurements. Conventional duct sensing studies have mostly used the global positioning system (GPS) coarse/acquisition (C/A) signals, which require the use of special antennas, e.g., dual antennas or high gain antennas. The new generation of GNSS signals has better autocorrelation properties that allow these signals to be used to detect evaporation ducts more effectively, but determining how to detect ducts using common GNSS equipment continues to present difficulties. In this letter, we describe a new method that uses the new BeiDou navigation satellite system (BDS-3) signals B1C and B2a to identify evaporation ducts. The proposed method has improved applicability because it only requires a single common antenna. A validation experiment was conducted on the sea by Weihai City, China, and raw GNSS intermediate frequency (IF) data were collected. A software-defined receiver (SDR) was used to perform post-processing. Our first results show that the proposed method can detect both direct and ducted signals using only one signal channel, and it thus represents a low-cost and effective technique for evaporation duct height (EDH) inversion. Shuo Gu, Fan Gao 0002, Runqi Liu, Quanchao He |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Semi-supervised vanishing point detection with contrastive learning
Shuo Gu, Yinbo Liu, Hui Kong 0001 |
Pattern Recognit. | 2 |
| 2023 | Center-Based Decoupled Point Cloud Registration for 6D Object Pose EstimationabstractIn this paper, we propose a novel center-based decoupled point cloud registration framework for robust 6D object pose estimation in real-world scenarios. Our method decouples the translation from the entire transformation by predicting the object center and estimating the rotation in a center-aware manner. This center offset-based translation estimation is correspondence-free, freeing us from the difficulty of constructing correspondences in challenging scenarios, thus improving robustness. To obtain reliable center predictions, we use a multi-view (bird’s eye view and front view) object shape description of the source-point features, with both views jointly voting for the object center. Additionally, we propose an effective shape embedding module to augment the source features, largely completing the missing shape information due to partial scanning, thus facilitating the center prediction. With the center-aligned source and model point clouds, the rotation predictor utilizes feature similarity to establish putative correspondences for SVD-based rotation estimation. In particular, we introduce a center-aware hybrid feature descriptor with a normal correction technique to extract discriminative, part-aware features for high-quality correspondence construction. Our experiments show that our method outperforms the state-of-the-art methods by a large margin on real-world datasets such as TUD-L, LINEMOD, and Occluded-LINEMOD. Code is available at https://github.com/JiangHB/CenterReg. Haobo Jiang, Zheng Dang, Shuo Gu, Jin Xie 0001, Mathieu Salzmann, Jian Yang 0003 |
ICCV | 3 |
| 2023 | Dual Fusion Network for Hyperspectral Semantic Segmentation
Shuo Gu, Jian Yang 0003 |
ICIG (2) | 2 |
| 2023 | LiDAR-SGMOS: Semantics-Guided Moving Object Segmentation with 3D LiDARabstractMost of the existing moving object segmentation (MOS) methods regard MOS as an independent task, in this paper, we associate the MOS task with semantic segmentation, and propose a semantics-guided network for moving object segmentation (LiDAR-SGMOS). We first transform the range image and semantic features of the past scan into the range view of current scan based on the relative pose between scans. The residual image is obtained by calculating the normalized absolute difference between the current and transformed range images. Then, we apply a Meta-Kernel based cross scan fusion (CSF) module to adaptively fuse the range images and semantic features of current scan, the residual image and transformed features. Finally, the fused features with rich motion and semantic information are processed to obtain reliable MOS results. We also introduce a residual image augmentation method to further improve the MOS performance. Our method outperforms most LiDAR-MOS methods with only two sequential LiDAR scans as inputs on the SemanticKITTI MOS dataset. Shuo Gu, Suling Yao, Jian Yang 0003, Cheng-Zhong Xu 0001, Hui Kong 0001 |
IROS | 1 |
| 2023 | Implicit Obstacle Map-driven Indoor Navigation Model for Robust Obstacle AvoidanceabstractRobust obstacle avoidance is one of the critical steps for successful goal-driven indoor navigation tasks. Due to the obstacle missing in the visual image and the possible missed detection issue, visual image-based obstacle avoidance techniques still suffer from unsatisfactory robustness. To mitigate it, in this paper, we propose a novel implicit obstacle map-driven indoor navigation framework for robust obstacle avoidance, where an implicit obstacle map is learned based on the historical trial-and-error experience rather than the visual image. In order to further improve the navigation efficiency, a non-local target memory aggregation module is designed to leverage a non-local network to model the intrinsic relationship between the target semantic and the target orientation clues during the navigation process so as to mine the most target-correlated object clues for the navigation decision. Extensive experimental results on AI2-Thor and RoboTHOR benchmarks verify the excellent obstacle avoidance and navigation efficiency of our proposed method.The core source code is available at https://github.com/xwaiyy123/object-navigation. Wei Xie 0019, Haobo Jiang, Shuo Gu, Jin Xie 0001 |
ACM Multimedia | 3 |
| 2023 | Mining electronic health records using artificial intelligence: Bibliometric and content analyses for current research status and product conversion
Yunfan He, Xianming Fan, Qinglian Wen, Dongxia Shen, Shuo Gu, Jianbo Lei |
J. Biomed. Informatics | 9 |
| 2022 | A Cylindrical Convolution Network for Dense Top-View Semantic Segmentation with LiDAR Point Clouds
Shuo Gu, Cheng-Zhong Xu 0001, Hui Kong 0001 |
ACCV (7) | 2 |
| 2021 | A Cascaded LiDAR-Camera Fusion Network for Road DetectionabstractMost of the existing road detection methods are either single-modal based, e.g., based on LiDAR or camera, or multi-modal based with LiDAR-camera fusion. The algorithms are designed for a specific data type, and cannot cope with input data changes. In addition, the LiDAR-camera based methods can only work in day time with enough light. In this paper, we develop a novel LiDAR-camera fusion strategy, which combines the LiDAR point clouds and the camera images in a cascaded way. The proposed network has two working modes, the single-modal mode with LiDAR point clouds only and the multimodal mode with both LiDAR and camera data, so it can be used in all day scenes. The whole network consists of three parts: 1) LiDAR segmentation module, which segments road points in the LiDAR’s imagery view. 2) Sparse-to-dense module, which upsamples the sparse LiDAR feature maps to dense road detection results. 3) LiDAR-camera fusion module, which fuses the dense LiDAR feature maps with the dense camera images to obtain accurate road estimations. Experiments on the KITTI-Road dataset show that the proposed cascaded LiDAR-camera fusion network can obtain very competitive road detection performance, with a MaxF value of 96.38%, and achieve the state-of-the-art in the single-modal mode among all LiDAR-only methods. Shuo Gu, Jian Yang 0003, Hui Kong 0001 |
ICRA | 1 |
| 2021 | Tumor IsomiR Encyclopedia (TIE): a pan-cancer database of miRNA isoformsabstractSUMMARY: MicroRNAs (miRNAs) are master regulators of gene expression in cancers. Their sequence variants or isoforms (isomiRs) are highly abundant and possess unique functions. Given their short sequence length and high heterogeneity, mapping isomiRs can be challenging; without adequate depth and data aggregation, low frequency events are often disregarded. To address these challenges, we present the Tumor IsomiR Encyclopedia (TIE): a dynamic database of isomiRs from over 10 000 adult and pediatric tumor samples in The Cancer Genome Atlas (TCGA) and The Therapeutically Applicable Research to Generate Effective Treatments (TARGET) projects. A key novelty of TIE is its ability to annotate heterogeneous isomiR sequences and aggregate the variants obtained across all datasets. Results can be browsed online or downloaded as spreadsheets. Here, we show analysis of isomiRs of miR-21 and miR-30a to demonstrate the utility of TIE. AVAILABILITY AND IMPLEMENTATION: TIE search engine and data are freely available to use at https://isomir.ccr.cancer.gov/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xavier Bofill-De Ros, Brian T. Luke, Robert Guthridge, Uma Mudunuri, Michael A. Loss, Shuo Gu |
Bioinform. | 6 |
| 2019 | Road Detection through CRF based LiDAR-Camera FusionabstractIn this paper, we propose a road detection method with LiDAR-camera fusion in a novel conditional random field (CRF) framework to exploit both range and color information. In the LiDAR based part, a fast height-difference based scanning strategy is applied in the 2D LiDAR range-image domain and a dense road detection result in camera image domain can be obtained through geometric upsampling given the LiDAR-camera calibration parameters. In the camera based part, a fully convolutional network is applied in the camera image domain. Finally, we fuse the dense and binary road detection results from both LiDAR and camera in a single CRF framework. Experiments show that using a single thread of CPU, the proposed LiDAR based part can operate at a frequency of over 250Hz with sparse output in range image and 40Hz with dense result in camera image for the 64-beam Velodyne scanner. Our CRF fusion method achieves very promising road detection performance on the KITTI-Road dataset. Shuo Gu, Yigong Zhang, Jinhui Tang 0001, Jian Yang 0003, Hui Kong 0001 |
ICRA | 1 |
| 2019 | Build your own hybrid thermal/EO camera for autonomous vehicleabstractIn this work, we propose a novel paradigm to design a hybrid thermal/EO (Electro-Optical or visible-light) camera, whose thermal and RGB frames are pixel-wisely aligned and temporally synchronized. Compared with the existing schemes, we innovate in three ways in order to make it more compact in dimension, and thus more practical and extendable for real-world applications. The first is a redesign of the structure layout of the thermal and EO cameras. The second is on obtaining a pixel-wise spatial registration of the thermal and RGB frames by a coarse mechanical adjustment and a fine alignment through a constant homography warping. The third innovation is on extending one single hybrid camera to a hybrid camera array, through which we can obtain wide-view spatially aligned thermal, RGB and disparity images simultaneously. The experimental results show that the average error of spatial-alignment of two image modalities can be less than one pixel. Yigong Zhang, Shuo Gu, Yubin Guo, Minghao Liu 0003, Zezhou Sun, Zhixing Hou, Ying Wang 0007, Jian Yang 0003, Jean Ponce, Hui Kong 0001 |
ICRA | 3 |
| 2019 | Two-View Fusion based Convolutional Neural Network for Urban Road DetectionabstractIn this paper, we propose a two-view fusion based convolutional neural network to estimate road areas in urban environments with LiDAR point clouds as input only. The proposed network takes two transformed LiDAR data representations, the LiDAR imageries and the camera-perspective maps, as inputs. It outputs pixel-wise road detection results in both the LiDAR's imagery view and the camera's perspective view simultaneously, in an end-to-end manner. To make better use of the data associations between two representations, we construct a novel mapping layer to transform features from the LiDAR's imagery view to the camera's perspective view in order to strengthen the road detection performance in the camera's perspective view. Experiments on the KITTI-Road dataset show that the proposed network can achieve the state-of-the-art performance among all LiDAR-only methods in real time. Shuo Gu, Yigong Zhang, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001 |
IROS | 1 |
| 2019 | QuagmiR: a cloud-based application for isomiR big data analyticsabstractSUMMARY: MicroRNAs (miRNAs) function as master regulators of gene expression. Recent studies demonstrate that miRNA isoforms (isomiRs) play a unique role in cancer development. Here, we present QuagmiR, the first cloud-based tool to analyze isomiRs from next generation sequencing data. Using a novel and flexible searching algorithm designed for the detection and annotation of heterogeneous isomiRs, it permits extensive customization of the query process and reference databases to meet the user 's diverse research needs. AVAILABILITY AND IMPLEMENTATION: QuagmiR is written in Python and can be obtained freely from GitHub (https://github.com/Gu-Lab-RBL-NCI/QuagmiR). QuagmiR can be run from the command line on local machines, as well as on high-performance servers. A web-accessible version of the tool has also been made available for use by academic researchers through the National Cancer Institute-funded Seven Bridges Cancer Genomics Cloud (https://cancergenomicscloud.org). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xavier Bofill-De Ros, Kevin Chen 0004, Susanna Chen, Nikola Tesic, Dusan Randjelovic, Nikola Skundric, Svetozar Nesic, Vojislav Varjacic, Elizabeth H. Williams, Raunaq Malhotra, Minjie Jiang, Shuo Gu |
Bioinform. | 12 |
| 2019 | Histograms of the Normalized Inverse Depth and Line Scanning for Urban Road DetectionabstractIn this paper, we propose to fuse the geometric information of a 3-D LiDAR and a monocular camera to detect the urban road region ahead of an autonomous vehicle. Our method takes advantage of both the high definition of 3-D LiDAR data and the continuity of road in image representation. First, we obtain an efficient representation of LiDAR data and an organized 2-D inverse depth map, by projecting the 3-D LiDAR points onto the camera's image plane. Through the new representation, we can acquire the intermediate representations of road scenes by extracting the vertical and horizontal histograms of the normalized inverse depth. The approximate road regions can be quickly estimated with both histogram-based schemes. To accurately find the road area, we propose a row and column scanning strategy in the approximate road region to refine the detected road area. We have carried out experiments on the public KITTI-Road benchmark, and have achieved one of the best performances among the LiDAR-based road detection methods without learning procedure. Shuo Gu, Yigong Zhang, Xia Yuan, Jian Yang 0003, Tao Wu 0001, Hui Kong 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2018 | Fusion of LiDAR and Camera by Scanning in LiDAR Imagery and Image-Guided Diffusion for Urban Road DetectionabstractThis paper proposes a new method for road detection based on a 3D LiDAR and a camera. First, the original LiDAR point cloud is re-organized in an ordered way to generate a LiDAR imagery. Then the flat region is extracted from the LiDAR imagery as the candidate road region. Next, a strategy of row- and column- scanning is given in the LiDAR imagery to detect a finer road region from the candidate region. To fuse the point cloud with image information, we transform the point cloud that corresponds to the above detected road region to the image space according to the calibration parameters between the LiDAR and camera. Then we give two image-guided diffusion schemes to conduct image segmentation of road area, respectively. Our experiments demonstrate that this training free approach detects the road region fast, accurately and robustly, and compares favorably with the state-of-the-art on the KITTI benchmark. Yigong Zhang, Shuo Gu, Jian Yang 0003, José M. Álvarez 0004, Hui Kong 0001 |
Intelligent Vehicles Symposium | 2 |
| 2014 | BiELL: A bisection ELLPACK-based storage format for optimizing SpMV on GPUs
Cong Zheng, Shuo Gu, Tongxiang Gu, Xing-Ping Liu |
J. Parallel Distributed Comput. | 2 |
| 2014 | Quantitatively Characterizing the Ligand Binding Mechanisms of Choline Binding Protein Using Markov State Model AnalysisabstractProtein-ligand recognition plays key roles in many biological processes. One of the most fascinating questions about protein-ligand recognition is to understand its underlying mechanism, which often results from a combination of induced fit and conformational selection. In this study, we have developed a three-pronged approach of Markov State Models, Molecular Dynamics simulations, and flux analysis to determine the contribution of each model. Using this approach, we have quantified the recognition mechanism of the choline binding protein (ChoX) to be ∼90% conformational selection dominant under experimental conditions. This is achieved by recovering all the necessary parameters for the flux analysis in combination with available experimental data. Our results also suggest that ChoX has several metastable conformational states, of which an apo-closed state is dominant, consistent with previous experimental findings. Our methodology holds great potential to be widely applied to understand recognition mechanisms underlining many fundamental biological processes. Shuo Gu, Daniel-Adriano Silva, Luming Meng, Alexander Yue, Xuhui Huang |
PLoS Comput. Biol. | 1 |