EDBT 2026 Demo / reviewers in the wild / expert
Siwen Quan
dblp:198/9333
· DBLP profile ↗
24ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-7579-937XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFF-Matcher: Robust cross-source registration with density-fused feature and bidirectional consensus matching
Zhenxuan Zeng, Xiyu Zhang 0001, Siwen Quan, Zhongwen Hu, Yu Zhu 0004, Jiaqi Yang 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | Rethinking the refinement stage of 3D object detection: A multi-task learning perspective with Mixture-of-Experts
Bingqian Wu, Pei An, Siwen Quan, Qiao Wu, Chu'ai Zhang, Jiaqi Yang 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Single Voter Spreading for Efficient Correspondence Grouping and 3D RegistrationabstractObtaining highly consistent correspondences between point clouds is crucial for computer vision tasks such as 3D registration and recognition. Due to nuisances such as limited overlap and noise, initial correspondences often contain a large number of outliers, imposing a great challenge to downstream tasks. In this paper, we present a novel single voter spreading (SVOS) method for efficient 3D correspondence grouping and 3D registration. Our core insight is to leverage low-order graph constraints only in a single voter spreading voting scheme to achieve comparable constrain-ability as complex constraints without searching them. First, a simple first-order graph is constructed for the initial correspondence set. Second, a two-stage voting method is proposed, including single voter voting and spread voters voting. Each voting stage involves both local and global voting via edge constraints only. This promises good selectivity while making the voting process time- and storage-efficient. Finally, top-scored correspondences are opted for robust transformation estimation. Experiments on U3M, 3DMatch/3DLoMatch, ETH, and KITTI-LC datasets verify that SVOS achieves new state-of-the-art correspondence grouping and registration performance, while being light-weight and robust to graph construction parameters. Siwen Quan, Zhao Zeng, Xiyu Zhang 0001, Jiaqi Yang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | A Hierarchical Prior Mining Approach for Non-Local Multi-View StereoabstractAs a fundamental problem in computer vision, multi-view stereo (MVS) aims at recovering the 3D geometry of the target from a set of 2D images. However, the reconstructed quality is significantly impacted by the presence of low-textured areas. In this paper, we propose a Hierarchical Prior Mining (HPM) framework for non-local multi-view stereo. Different from most existing works dedicated to focusing on local information and only using a single prior, HPM captures non-local structural cues and leverages multi-source priors for geometry recovery. Based on the framework, we first propose HPM-MVS, which obtains precise initial hypotheses through non-local operations, simultaneously constructing a better planar prior model in an HPM framework to further facilitate hypothesis generation. In addition, we futher propose HPM-MVS++, which excavates the structured region information of images and spatial geometric relationships of hypotheses as prior knowledge. Then, it incorporates them into probabilistic graphical models, ultimately deducing two novel multi-view matching costs. This significantly enhances the robustness to challenging situations and improves the completeness of the reconstruction. Experimental results on the ETH3D and Tanks & Temples have verified the superior performance and strong generalization capability of our approach. Jiaqi Yang 0002, Yanan He, Chunlin Ren, Qingshan Xu 0001, Siwen Quan, Xiyu Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | SPU-IMR: Self-supervised Arbitrary-scale Point Cloud Upsampling via Iterative Mask-recovery NetworkabstractPoint cloud upsampling aims to generate dense and uniformly distributed point sets from sparse point clouds. Existing point cloud upsampling methods typically approach the task as an interpolation problem. They achieve upsampling by performing local interpolation between point clouds or in the feature space, then regressing the interpolated points to appropriate positions. By contrast, our proposed method treats point cloud upsampling as a global shape completion problem. Specifically, our method first divides the point cloud into multiple patches. Then a masking operation is applied to remove some patches, leaving visible point cloud patches. Finally, our custom-designed neural network iterative completes the missing sections of the point cloud through the visible parts. During testing, by selecting different mask sequences, we can restore various complete patches. A sufficiently dense upsampled point cloud can be obtained by merging all the completed patches. We demonstrate the superior performance of our method through both quantitative and qualitative experiments, showing overall superiority against both existing self-supervised and supervised methods. Ziming Nie, Qiao Wu, Chenlei Lv, Siwen Quan, Zhaoshuai Qi, Muze Wang, Jiaqi Yang 0002 |
AAAI | 4 |
| 2025 | Pre-training meets iteration: Learning for robust 3D point cloud denoising
Siwen Quan, Hebin Zhao, Zhao Zeng, Ziming Nie, Jiaqi Yang 0002 |
Pattern Recognit. Lett. | 1 |
| 2025 | DRGAN: A Detail Recovery-Based Model for Optical Remote Sensing Images Super-ResolutionabstractThe need for high-resolution (HR) remote sensing images has grown significantly in recent years as a result of the rapid advancement of fine-sensing technologies. However, increasing sensor resolution usually requires a costly investment. To tackle this challenge, super-resolution (SR) methods for remote sensing images have emerged as a cost-effective alternative to enhance the quality and usability of existing low-resolution (LR) images. Although many current methods have achieved some reconstruction results, they often suffer from problems such as transition smoothing and artifacts. To solve these problems, we propose an SR reconstruction model for detail recovery based on generative adversarial networks (GANs), referred to as DRGAN. Specifically, unlike the traditional residual-in-residual dense block network (RRDBNet), we propose a novel dense residual network (OSRRDBNet). It uses dynamic convolution and self-attention mechanisms to recover the rich detailed information in the image more effectively. In addition, we employ an average pooling layer to enhance the ability to capture HR image features. By conducting experiments on three different remote sensing datasets, DRGAN shows remarkable reconstruction results and successfully recovers the rich detail information in the images. Yongchao Song, Jiping Bi, Siwen Quan, Xuan Wang 0021 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Lane Detection for Autonomous Driving: Comprehensive Reviews, Current Challenges, and Future PredictionsabstractLane detection is crucial for autonomous driving systems (ADS), utilizing sensors like cameras and LiDAR to identify lanes and understand vehicle position, direction, and lane shape. It provides data support for the control system to make informed driving decisions. In this survey, we review recent advancements in lane detection, focusing on both 2D techniques and emerging 3D methods. We begin with an overview of the significance of lane detection in ADS, followed by an analysis of the evolution of 2D techniques over the past decade, covering traditional and deep learning approaches. We also examine recent advancements in 3D lane detection. Additionally, we summarize evaluation metrics and popular datasets in the field. Finally, we discuss current challenges and future directions in lane detection, aiming to provide valuable insights for researchers and developers in this technology. Jiping Bi, Yongchao Song, Yahong Jiang, Xuan Wang 0021, Zhaowei Liu 0001, Siwen Quan, Weiqing Yan |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | FRCFNet: Feature Reassembly and Context Information Fusion Network for Road ExtractionabstractExisting road extraction methods based on very high resolution (VHR) satellite imagery suffer from insufficient multidimensional feature expression and difficulty capturing global context. We propose a grouping multidimensional feature reassembly (GMFR) module, performing channel, height, and width reassembly of multiscale features between network layers via gating to focus on valid information. Given the distinct geometric structure of roads, we propose a novel module, multidirectional context information fusion (MCIF), utilizing four strip convolutions to capture the long-distance context in various directions within VHR images. It aggregates global information through two pooling branches. Based on these, we designed a road extraction network, FRCFNet, with an encoder–decoder structure and skip connections. The proposed network efficiently fuses multiscale features while capturing global context from various directions and reducing complexity. Experimental results show that the proposed method achieves 68.97% and$80.23\%~F1$-score on CHN6-CUG and DeepGlobe datasets, respectively, outperforming other comparison methods. The code will be posted athttps://github.com/CHD-IPAC/FRCFNet. Haijuan Wang, Danni Xue, Moslema Chowdhuray Momi, Zhen Ye 0007, Siwen Quan |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | ELLK-Net: An Efficient Lightweight Large Kernel Network for SAR Ship DetectionabstractELLK-Net, an efficient, lightweight network with a large kernel, is proposed for synthetic aperture radar (SAR) ship detection. It addresses background variations, different ship scales, and noise interference challenges. ELLK-Net uses an anchor-free detector framework and sequentially decomposes large kernel convolutions to capture comprehensive global information and long-range dependencies. It adaptively selects convolution kernels on the basis of target characteristics, enhancing multiscale feature expression. A novel large kernel multiscale attention (LKMA) module is introduced to enhance interlayer feature fusion and semantic alignment, mitigating the impacts of overlapping ships and scattering noise. Structural reparameterization techniques optimize inference speed across devices without compromising accuracy. The experimental results on the SAR ship detection dataset (SSDD) and high-resolution SAR image dataset (HRSID) datasets demonstrate that ELLK-Net achieves impressive AP50 values of 95.6% and 90.6% for horizontal box detection and 89.7% and 79.7% for rotating box detection, respectively. The reparameterized detector exhibits a significant 48.7% FPS improvement on the Nvidia Jetson NX platform, indicating its suitability for edge computing deployment. The code is available athttps://github.com/CHD-IPAC/ELLK-Net. Moslema Chowdhuray Momi, Siwen Quan, Zhen Ye 0007 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and TrendsabstractA system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving. Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | ESC-Net: Alleviating Triple Sparsity on 3D LiDAR Point Clouds for Extreme Sparse Scene Completionabstract3D scene completion (SC) has made progress in the last three years. From the application of mobile robot system, SC should support the downstream task (i.e. mapping or perception), instead of only predicting the completed scenes. However, as the low-cost few-beam LiDAR is widely applied in mobile robot, gap between SC and downstream tasks is large. To generate the high quality completion result, the bottleneck lies in the triple sparsity of input, ground truth (GT) occupancy, and GT foreground. To deal with the triple sparsity, we present an extreme sparse scene completion network (ESC-Net). At first, input sparsity hides most of the spatial information of the scene. A feature completion (FC) decoder is designed to mine the spatial feature using feature-level completion. Then, GT occupancy sparsity hinders representation learning of the real scene with continuous surfaces. A multi-view multi-task attention (MMA) loss is presented to recover the high-quality object boundaries via correcting occupancy and semantic labels of regions from 3D and bird's eye view (BEV) spaces. After that, GT foreground sparsity is the imbalance of foreground and background GT labels. It causes the inaccuracy of local 3D object completion. A combination network (ESC-Net-D) is presented to recover 3D structural details of both foreground and background. Experiment is conducted on KITTI and SemanticPOSS datasets. It shows that ESC-Net has the performance higher than current methods not only on completion task, but also on the downstream tasks (i.e. 3D registration, 3D object detection). Hence, we believe that ESC-Net benefits to the community of mobile robot. Source code is released soon. Pei An, Siwen Quan, Junfeng Ding, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | VOID: 3D object recognition based on voxelization in invariant distance space
Jiaqi Yang 0002, Shichao Fan, Siwen Quan, Yanning Zhang 0001 |
Vis. Comput. | 4 |
| 2022 | Dual spin-image: A bi-directional spin-image variant using multi-scale radii for 3D local shape description
Daryl L. Bibissi, Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
Comput. Graph. | 3 |
| 2022 | Toward Efficient and Robust Metrics for RANSAC Hypotheses and 3D Rigid RegistrationabstractThis paper focuses on developing efficient and robust evaluation metrics for RANSAC hypotheses to achieve accurate 3D rigid registration. Estimating six-degree-of-freedom (6-DoF) pose from feature correspondences remains a popular approach to 3D rigid registration, where random sample consensus (RANSAC) is a well-known solution to this problem. However, existing metrics for RANSAC hypotheses are either time-consuming or sensitive to common nuisances, parameter variations, and different application scenarios, resulting in performance deterioration with respect to overall registration accuracy and speed. We alleviate this problem by first analyzing the contributions of inliers and outliers and then proposing several efficient and robust metrics with different designing motivations for RANSAC hypotheses. Comparative experiments on four standard datasets with different nuisances and application scenarios verify that our considered metrics can significantly improve the registration performance and are more robust than several state-of-the-art competitors, making them good gifts to practical applications. This work also draws an interesting conclusion, i.e., not all inliers are equal while all outliers should be equal, which may shed new light on this research problem. Jiaqi Yang 0002, Siwen Quan, Qian Zhang 0046, Yanning Zhang 0001, Zhiguo Cao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Correspondence Selection With Loose-Tight Geometric Voting for 3-D Point Cloud RegistrationabstractThis article presents a simple yet effective method for 3-D correspondence selection and point cloud registration. It first models the initial correspondence set as a graph with nodes representing correspondences and edges connecting geometrically compatible nodes. Such graphs offer either loose or tight geometric constraints for judging the correctness of correspondence, e.g., edges, loops, and cliques. Then, we render these constraints dynamic voters to judge the correctness of a node. More specifically, we develop a loose–tight geometric voting (LT-GV) method that employs both loose and tight geometric constraints in the graph to score 3-D feature correspondences. The motivation behind this is to strike a balanced performance in terms of precision and recall because loose and tight constraints are complementary to each other. Under the dynamic voting scheme with both loose and tight voters, consistent correspondences can be retrieved based on the voting score. Both feature-matching and 3-D point cloud registration experiments on datasets with different modalities, challenges, application scenarios, and comparisons with state-of-the-art methods (including deep learned methods) verify that our LT-GV is effective for correspondence selection, robust to a number of nuisances, and able to dramatically boost 3-D point cloud registration performance. Jiaqi Yang 0002, Siwen Quan, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | SAC-COT: Sample Consensus by Sampling Compatibility Triangles in Graphs for 3-D Point Cloud RegistrationabstractSix-degree-of-freedom (6-DOF) pose estimation from feature correspondences remains a popular and robust approach for 3-D registration. However, heavy outliers that existed in the initial correspondence set pose a great challenge to this problem. This article presents a simple yet effective estimator called SAmple Consensus by sampling COmpatibility Triangles in graphs (SAC-COT) for robust 6-DOF pose estimation and 3-D registration. The key novelty is a guided three-point sampling approach. It is based on a novel correspondence sample representation, i.e., COmpatibility Triangle (COT). We first model the correspondence set as a graph with nodes connecting compatible correspondences. Then, by ranking and sampling COTs formed by ternary loops, we show that correct hypotheses can be generated in early iteration stage. Finally, the hypothesis generated by the COT yielding to the maximum consensus is the output of SAC-COT. Extensive experiments on six data sets and extensive comparisons with the state-of-the-art estimators confirm that: 1) SAC-COT can achieve accurate registrations with a few iterations and 2) SAC-COT outperforms all competitors and is ultrarobust when confronted with Gaussian noise, data decimation, holes, clutter, partial overlap, varying scales of input correspondences, and data modality variation. Jiaqi Yang 0002, Siwen Quan, Zhaoshuai Qi, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | On shortened 3D local binary descriptors
Siwen Quan, Jie Ma 0003 |
Inf. Sci. | 1 |
| 2020 | Compatibility-Guided Sampling Consensus for 3-D Point Cloud RegistrationabstractThis article presents an efficient and robust estimator called compatibility-guided sampling consensus (CG-SAC) to achieve accurate 3-D point cloud registration. For correspondence-based registration methods, the random sample consensus (RANSAC) is served as a de facto solution for rigid transformation estimation from a number of feature correspondences. Unfortunately, RANSAC still suffers from two major limitations. First, it generates a hypothesis with at least three samples and desires a very large number of iterations to attain reasonable results, making it relatively time consuming. Second, the randomness during sampling can result in inaccurate results as it is highly potential to miss the optimal hypothesis. To solve these problems, we propose a compatibility-guided sampling strategy to eliminate randomness during sampling. In particular, only two correspondences are required by our method for hypothesis generation. We then rank correspondence pairs according to their compatibility scores because compatible correspondences are more likely to be correct and can yield more reasonable hypotheses. In addition, we propose a new geometric constraint named the distance between salient points (DSP) to measure the compatibility of two correspondences. Experiments on a set of real-world point cloud data with different application contexts and data modalities confirm the effectiveness of the proposed method. Comparison with several state-of-the-art estimators demonstrates the overall superiority of our CG-SAC estimator with regards to precision and time efficiency. Siwen Quan, Jiaqi Yang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Evaluating Local Geometric Feature Representations for 3D Rigid Data MatchingabstractLocal geometric descriptors act as an essential component for 3D rigid data matching. A rotational invariant local geometric descriptor usually consists of two components: local reference frame (LRF) and feature representation. However, existing evaluation efforts have mainly been paid on the LRF or the overall descriptor and the quantitative comparison of feature representations remains unexplored. This paper fills the gap by comprehensively evaluating nine state-of-the-art local geometric feature representations. In particular, our evaluation assesses feature representations based on ground-truth LRFs such that the ranking of tested methods is more convincing as compared with existing studies. The experiments are deployed on six standard datasets with various application scenarios (shape retrieval, point cloud registration, and object recognition) and data modalities (LiDAR, Kinect, and Space Time) as well as perturbations including Gaussian noise, shot noise, data decimation, clutter, occlusion, and limited overlap. The evaluated terms cover the major concerns for a feature representation, e.g., distinctiveness, robustness, compactness, and efficiency. The outcomes present interesting findings that may shed new light on this community and provide complementary perspectives to existing evaluations on the topic of local geometric feature description. A summary of evaluated methods regarding their peculiarities is finally presented to guide real-world applications and new descriptor crafting. Jiaqi Yang 0002, Siwen Quan, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Local voxelized structure for 3D binary feature representation and robust registration of point clouds from low-cost sensors
Siwen Quan, Jie Ma 0003, Fangyu Hu, Bin Fang 0007, Tao Ma 0004 |
Inf. Sci. | 1 |
| 2018 | Representing local shape geometry from multi-view silhouette perspective: A distinctive and robust binary 3D feature
Siwen Quan, Jie Ma 0003, Tao Ma 0004, Fangyu Hu, Bin Fang 0007 |
Signal Process. Image Commun. | 1 |
| 2017 | Local voxelized structure for 3D local shape description: A binary representationabstractThis paper proposes a novel binary descriptor named local voxelized structure (LoVS) for 3D local shape description. Unlike many previous local shape descriptors relying on geometric attributes such as curvature and normals, LoVS simply uses point spatial locations to encode the local shape structure represented by point clouds into bit string. Specifically, LoVS is computed on a local cubic volume around the keypoint. The orientation of the cubic is determined by a local reference frame (LRF) to achieve rotation invariance. Then, the cubic is uniformly split into a set of voxels. A voxel is attached with label 1 if there are points inside, otherwise, it produces a 0 bit. All these labels therefore integrates into the LoVS descriptor. We evaluate our method on three public datasets. On each dataset, the LoVS descriptor outperforms all other descriptors tested. Siwen Quan, Jie Ma 0003, Fangyu Hu, Bin Fang 0007, Tao Ma 0004 |
ICIP | 1 |
| 2016 | Fast motion deblurring using gyroscopes and strong edge predictionabstractThis paper presents a fast deblurring algorithm to remove camera motion blur from a single photograph using built-in gyroscopes and strong edge prediction. An inaccurate blur kernel or point spread function (PSF) usually leads to an unsatisfying restored result. Hence, we propose a robust three-phase method for accurate PSF estimation. In the first stage, we utilize the embedded gyroscopes to compute a coarse version of the PSF from the camera's angular velocity during an exposure. In order to reduce the execution time of the later PSF modification, we introduce a patch selection procedure in the second stage to choose a suitable region from the blurry image based on the size of the coarse PSF estimated in stage one. The third phase aims to modify the coarse PSF to obtain an accurate one by predicting strong edges from an estimated latent image. In our experiments, we compare the restoration performance of several state-of-the-art approaches including ours and find that the proposed method outperforms others qualitatively as well as quantitatively. In addition, our method is also compared with the multi-scale approach without gyroscope data and shows shorter processing time and comparable deblurring quality. To the best of our knowledge, this is the first work that combines the sensor-aided method with the image-based approach to estimate the blur kernel. Jiacai Zhao, Jie Ma 0003, Bin Fang 0007, Siwen Quan, Fangyu Hu |
ICPR | 4 |