EDBT 2026 Demo / reviewers in the wild / expert
Jie Ma 0003
dblp:62/5110-3
· DBLP profile ↗
51ranked-venue papers
0as first author
35since 2021 · last 2026
0000-0003-1996-5163ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 19 since 2021Artificial intelligence and machine learning · 22 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic QueriesabstractSemantic occupancy has emerged as a powerful representation in world models for its ability to capture rich spatial semantics. However, most existing occupancy world models rely on static and fixed embeddings or grids, which inherently limit the flexibility of perception. Moreover, their ``in-place classification" over grids exhibits a potential misalignment with the dynamic and continuous nature of real scenarios. In this paper, we propose SparseWorld, a novel 4D occupancy world model that is flexible, adaptive, and efficient, powered by sparse and dynamic queries. We propose a Range-Adaptive Perception module, in which learnable queries are modulated by the ego vehicle states and enriched with temporal-spatial associations to enable extended-range perception. To effectively capture the dynamics of the scene, we design a State-Conditioned Forecasting module, which replaces classification-based forecasting with regression-guided formulation, precisely aligning the dynamic queries with the continuity of the 4D environment. In addition, We specifically devise a Temporal-Aware Self-Scheduling training strategy to enable smooth and efficient training. Extensive experiments demonstrate that SparseWorld achieves state-of-the-art performance across perception, forecasting, and planning tasks. Comprehensive visualizations and ablation studies further validate the advantages of SparseWorld in terms of flexibility, adaptability, and efficiency. Chenxu Dang, Jason Bao, Pei An, Xinyue Tang, An Pan, Jie Ma 0003, Bingchuan Sun |
AAAI | 7 |
| 2026 | HGS-OCC: Real-time 3D occupancy prediction via hybrid depth optimization with multi-representation geometry-semantic integration
Zaipeng Duan, Xuzhong Hu, Pei An, Chenxu Dang, Jie Ma 0003 |
Pattern Recognit. | 6 |
| 2026 | Dual-domain homogeneous fusion with cross-modal mamba and progressive decoder for 3D object detection
Xuzhong Hu, Zaipeng Duan, Pei An, Ziwen Xu, Jie Ma 0003 |
Pattern Recognit. | 6 |
| 2025 | Co-Progression Knowledge Distillation with Knowledge Prototype for Industrial Anomaly DetectionabstractUnsupervised anomaly detection has emerged as a powerful technique for identifying abnormal patterns in images without relying on pre-labeled defective samples. Many unsupervised methods use pre-trained feature extractors from large datasets, with knowledge distillation between teacher and student models being a leading technique. However, due to the similar structures of teacher and student, these methods face challenges like excessive specialization and inadequate generalization, reducing detection performance. In this paper, we introduce a Co-Progression Knowledge Distillation (CPKD) framework, enabling bidirectional learning between teacher and student models. This innovative framework enables concurrent evolution of both models, fostering mutual improvement and enhanced adaptability. To maintain system stability and prevent overspecialization, we introduce a knowledge prototype as a regulatory mechanism for the teacher's learning process. Our method effectively addresses key challenges in anomaly detection, including insufficient learning and overadaptation, by striking a balance between acquiring new knowledge and preserving core competencies. We demonstrate significant improvements in detection accuracy, achieving SOTA performance on the MVTec dataset. Bokang Yang, Zhe Zhang 0037, Jie Ma 0003 |
AAAI | 3 |
| 2025 | FASTer: Focal token Acquiring-and-Scaling Transformer for Long-term 3D Objection DetectionabstractRecent top-performing temporal 3D detectors based on Lidars have increasingly adopted region-based paradigms. They first generate coarse proposals, followed by encoding and fusing regional features. However, indiscriminate sampling and fusion often overlook the varying contributions of individual points and lead to exponentially increased complexity as the number of input frames grows. Moreover, arbitrary result-level concatenation limits the global information extraction. In this paper, we propose a Focal Token Acquring-and-Scaling Transformer (FASTer), which dynamically selects focal tokens and condenses token sequences in an adaptive and lightweight manner. Empha-sizing the contribution of individual tokens, we propose a simple but effective Adaptive Scaling mechanism to capture geometric contexts while sifting out focal points. Adaptively storing and processing only focal points in historical frames dramatically reduces the overall complexity. Furthermore, a novel Grouped Hierarchical Fusion strategy is proposed, progressively performing sequence scaling and Intra-Group Fusion operations to facilitate the exchange of global spatial and temporal information. Experiments on the Waymo Open Dataset demonstrate that our FASTer significantly outperforms other state-of-the-art detectors in both performance and efficiency while also exhibiting improved flexibility and robustness. The code is available at https://github.com/MSunDYY/FASTer.git. Chenxu Dang, Zaipeng Duan, Pei An, Xuzhong Hu, Jie Ma 0003 |
CVPR | 6 |
| 2025 | SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy PredictionabstractMultimodal 3D occupancy prediction has garnered significant attention for its potential in autonomous driving. However, most existing approaches are single-modality: camera-based methods lack depth information, while LiDAR-based methods struggle with occlusions. Current lightweight methods primarily rely on the Lift-Splat-Shoot (LSS) pipeline, which suffers from inaccurate depth estimation and fails to fully exploit the geometric and semantic information of 3D LiDAR points. Therefore, we propose a novel multimodal occupancy prediction network called SDG-OCC, which incorporates a joint semantic and depth-guided view transformation coupled with a fusion-to-occupancy-driven active distillation. The enhanced view transformation constructs accurate depth distributions by integrating pixel semantics and co-point depth through diffusion and bilinear discretization. The fusion-to-occupancy-driven active distillation extracts rich semantic information from multimodal data and selectively transfers knowledge to image features based on LiDAR-identified regions. Finally, for optimal performance, we introduce SDG-Fusion, which uses fusion alone, and SDG-KL, which integrates both fusion and distillation for faster inference. Our method achieves state-of-the-art (SOTA) performance with real-time processing on the Occ3D-nuScenes dataset and shows comparable performance on the more challenging SurroundOcc-nuScenes dataset, demonstrating its effectiveness and robustness. The code will be released at https://github.com/DzpLab/SDGOCC. Zaipeng Duan, Chenxu Dang, Xuzhong Hu, Pei An, Junfeng Ding, Jie Zhan, YunBiao Xu, Jie Ma 0003 |
CVPR | 8 |
| 2025 | Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology RelationshipabstractImage-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan |
IJCAI | 6 |
| 2025 | Op-Occ-Net: Neural Ordinary Differential Equation Based 4D Occupancy Forecasting for Autonomous Vehicles
Junfeng Ding, Erxin Guo, Pei An, Jie Ma 0003, Ruixiang Zhao |
PRCV (11) | 4 |
| 2025 | Restore Nighttime Flare Distribution: Flare Removal via Light Source Preservation and Physical Prior Rendering
Shaojun Lin, Jie Ma 0003 |
PRCV (9) | 2 |
| 2025 | DSConv: Fine-Grained Dynamic Sequence Convolution for 3D Understanding
Pei An, Chenxu Dang, Zaipeng Duan, Jie Ma 0003 |
PRCV (10) | 5 |
| 2025 | Defect-LoRA: Controllable defect data augmentation based on low-rank adaptation for surface defect recognition under limited data
Zhe Zhang 0037, Jie Ma 0003 |
Expert Syst. Appl. | 3 |
| 2025 | Lidar-camera range-view fusion for 3D object detection in autonomous driving
Xuzhong Hu, Zaipeng Duan, Pei An, Jun Zhang 0062, Jie Ma 0003 |
Multim. Syst. | 5 |
| 2025 | IIIM-SAM: Zero-Shot Texture Anomaly Detection Without External PromptsabstractAnomaly detection is a crucial aspect in ensuring the reliability of industrial products in smart manufacturing. Despite employing unsupervised anomaly detection methods to mitigate the challenges associated with obtaining defect data, there remains the challenge of adapting anomaly detectors to drift within normal data distributions, especially when there are significant intra-class variations in normal samples. Consequently, zero-shot anomaly detection techniques, which do not require training data and thus avoid interference from training data, have emerged as a new research direction. Meanwhile, the zero-shot segmentation capability demonstrated by the Segment Anything Mode (SAM) visual foundation model has further fueled our research enthusiasm. In this paper, we focus on texture images and propose a no-external prompt SAM-based method for zero-shot texture anomaly detection, called Image Internal Information Mining SAM (IIIM-SAM). We utilize the Image Internal Information Mining Prompter to mine relationships between internal regions of the image, identifying specific background and potential defect points. The two types of prompt points are fed into the SAM model in different stages, imparting the SAM with incorporating category information segmentation capability. Unlike current methods, our zero-shot texture anomaly detection method requires no external prompts or expert knowledge, and also avoids the uncertainty introduced by text prompts necessary for CLIP-based approaches. Our experimental results demonstrate the superiority of our zero-shot texture anomaly detection method compared to other approaches. Specifically, in the texture subset of MVTecAD/KolektorSDD2, our IIIM-SAM achieves 99.2%/93.6% image-level AUROC and 98.6%/92.3% pixel-level AUROC. Zhe Zhang 0037, Jiahe Yue, Runchu Zhang, Jie Ma 0003 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Attention4Align: Align Multi-view Parts Via Part2Part Hierarchical Attention Map for Fine-Grained 3D Object Classification
Runchu Zhang, Jiahe Yue, Zhe Zhang 0037, Jie Ma 0003 |
ACCV (9) | 4 |
| 2024 | Low-Rank Completion Based Normal Guided Lidar Point Cloud Up-SamplingabstractCommercial inexpensive LiDAR sensor generally suffers low vertical resolutions, whose point cloud is sparse and may not be able to satisfy future metaverse applications. LiDAR point cloud up-sampling is a task to increase the vertical resolution while preserving the structural details. Scene representation is the central pillar of point cloud up-sampling. However, the sparsity of point cloud hinders the extraction of scene representation. In this paper, we find that low-rank representation can describe the primary scene structure approximately, and convert up-sampling as low-rank tensor completion problem. To decrease problem complexity, we leverage range view projection to convert the problem as low-rank depth completion, and present a low-rank normal guided up-sampling approach. It uses normal as guidance to smooth range depth. Extensive experiments show that our method outperforms current methods. In 2× up-sampling task, it achieves as low as 41cm of mean absolute error (MAE), which is 282% and 32% smaller than interpolation and traditional matrix completion methods, respectively. Hence, we believe the proposed method benefits to the field of metaverse. Pei An, You Yang 0002, Jie Ma 0003 |
ICASSP | 4 |
| 2024 | LiDAR-Camera Extrinsic Calibration with Hierachical and Iterative Feature MatchingabstractIn autonomous driving, the LiDAR-Camera system plays a crucial role in a vehicle’s perception of 3D environments. To effectively fuse information from both camera and LiDAR, extrinsic calibration is indispensable. Recently, some researchers have proposed deep learning-based methods that utilize convolutional networks to automatically extract features from LiDAR depth images and RGB images for calibration. However, these features do not sufficiently interact during feature matching, which limits the calibration accuracy. To this end, we introduce a novel extrinsic calibration network (HIFM-Net) in this paper. It establishes a comprehensive connection between camera and LiDAR features by calculating a globally-aware map-to-map cost volume and hierachical point-to-map cost volumes. The former is used to regress large extrinsic offsets. The latter is employed to iteratively fine-tune extrinsic parameters, while the rigidity of LiDAR points is considered in each iteration to enhance regression robustness. Extensive experiments on the KITTI-odometry dataset demonstrate the superior performance of our HIFMNet compared to other state-of-the-art learning-based methods. Xuzhong Hu, Zaipeng Duan, Junfeng Ding, Zhe Zhang 0037, Xiao Huang 0008, Jie Ma 0003 |
ICRA | 6 |
| 2024 | Towards Visibility Estimation and Noise-Distribution-Based Defogging for LiDAR in Autonomous DrivingabstractPoint clouds play a crucial role in robots and intelligent vehicles. Noise caused by fog droplets seriously degrades the quality of point clouds. Previous researches have shown that the extent of degradation is correlated with visibility. The fog attenuation coefficient is associated with visibility. In light of this background, this paper proposes a noise-distribution-based defogging method for point clouds. Our approach hinges on the estimation of the fog attenuation coefficient, facilitated by road-based prior knowledge. Subsequently, our method integrates the fog-induced noise distribution inferred from the LiDAR imaging model with the spatially non-uniform distribution of point clouds caused by LiDAR structure. The fused results are input to a statistical filter based on the relative sparsity of noise to achieve defogging. This paper is one of the early works focusing on point cloud defogging. Its core insight lies in the estimation of the attenuation coefficient and the employment of fog-induced noise distribution for defogging. Experiments demonstrate that our method can accurately mitigate the impact of fog and meanwhile enhance the performance of 3D object detection network. Jie Zhan, Yucong Duan, Junfeng Ding, Xuzhong Hu, Xiao Huang 0008, Jie Ma 0003 |
ICRA | 6 |
| 2024 | How SAM helps Unsupervised Video Object Segmentation?abstractAs an emerging vision foundation model, Segment Anything Model (SAM) has been successfully applied to Video Object Segmentation (VOS). However, previous methods rely on first frame’s mask or manual interaction, which belong to semi-supervised or interactive video object segmentation. The potential of SAM in Unsupervised Video Object Segmentation (UVOS) remains to be explored. In this work, we propose a two-stage training framework to explore how SAM helps UVOS. In Stage 1, utilizing SAM powerful feature extraction ability, we only train a lightweight decoder with feature aggregation module. This stage employs relaxed flow reconstruction as an unsupervised proxy task for object discovery. In Stage 2, based on SAM promptability, we design a refinement process that involves rough correction based on in-clip temporal consistency and heuristic refinement guided by Prompt-closure, to generate refined results from the coarse mask produced by Stage 1. Finally, update the segmentation head using refined results along with corresponding confidence. The proposed method demonstrates competitive performance on three common UVOS datasets (DAVIS2016, SegTrackv2, FBMS59) with higher accuracy, faster convergence, lower training cost, validating the effectiveness of our approach. Jiahe Yue, Runchu Zhang, Zhe Zhang 0037, Ruixiang Zhao, Wu Lv, Jie Ma 0003 |
IJCNN | 6 |
| 2024 | OL-Reg: Registration of Image and Sparse LiDAR Point Cloud With Object-Level Dense CorrespondencesabstractImage and point cloud registration (2D-3D registration) is an essential prerequisite for multi-modal feature fusion. However, due to the significant feature difference of point cloud and image, it is challenging to establish 2D-3D correspondences. Targeting for the background of autonomous driving, we propose 2D-3D registration method with object-level correspondence (OL-Reg) in this paper. Object-level correspondence consists of object bounding box and object contour in 2D image and 3D space. The first step is to match 2D-3D objects. Due to sensor pose and field of view (FoV) difference, object shape and occlusion is different in image and point cloud, causing the difficulty of object matching. To solve this issue, we represent object as 3D bounding box, and design 2D-3D object matching with 3D box projection (Box-Proj) constraint. It aligns object 3D bounding box in image and point cloud. After that, the next step is to build 2D-3D correspondence from the matched objects. To extract correspondence from object with irregular shape, we notice the distance constraint of object surface and rays back-projected from object contour, and present projection based iterative closest point (Proj-ICP). Towards the stability of Proj-ICP, object-level regularization term is designed. Experiment is conducted in KITTI object and odometry dataset. With the pre-trained 3D object detector, results suggest that OL-Reg has the better performance than current approaches in tasks of re-localization and extrinsic calibration. And source code will be released soon1. Pei An, Xuzhong Hu, Junfeng Ding, Jun Zhang 0062, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Combined Anomaly Aware Weakly Supervised Lightweight Model for Surface Defect InspectionabstractSurface defect inspection plays a vital role in the industrial production process. Many detection methods based on deep learning have been gradually applied because of their better generalization performance. However, achieving accurate annotations for training deep learning models remains a challenge due to the difficult definition of defect boundaries and the high cost of manual annotation work. Meanwhile, the detection performance of the current deep-learning methods still cannot meet the needs of industrial applications. To address these issues, this article proposes a combined anomaly aware weakly supervised lightweight model that only requires image-level labels for training and outputs defect localization. In the framework, we first design a lightweight backbone to obtain feature maps. Then, we propose a novel weakly supervised localization (WSL) method to obtain anomaly responses and use them as prior knowledge of the downstream network. Finally, the final defect detection result is obtained through the work of the designed downstream fine inspection network. In addition, we will employ multiple supervisions throughout the framework for full data use. The results of the evaluation on four real-world defect datasets demonstrate that the proposed method is superior and more generalized than state-of-the-art WSL methods and defect detection methods on average precision. Zhe Zhang 0037, Zhenqiao Shang, Xin Wang 0201, Jie Ma 0003 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and TrendsabstractA system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving. Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | SP-Det: Leveraging Saliency Prediction for Voxel-Based 3D Object Detection in Sparse Point CloudabstractVoxel is one of the common structural representation of 3D point cloud. Due to the sparsity of point cloud generated by light detection and ranging (LiDAR), there is the extreme imbalance in the foreground and background voxels. It decreases the accuracy of 3D object detection, has the negative effect on intelligent driving safety. To overcome this problem, we present a saliency prediction based 3D object detector SP-Det in this article. Although foreground voxels have the sufficient feature of object, it is difficult to localize the foreground region from voxel space with the larger background region. We design an auxiliary learning task, saliency prediction (SP). It benefits 3D detector in identifying the foreground region. SP task uses label diffusion to alleviate the label imbalance. It reduces the learning difficulty of saliency in voxel and bird's eye view (BEV) spaces. After that, to strengthen feature interaction from the sparse foreground region, we design saliency fusion (SF) module to fuse the learning result in SP task. It utilizes voxel and BEV saliency maps as progressive attention to resist the redundant feature from background region. To aggregate more foreground feature inside 3D and BEV region of interest (RoI), we design hybrid grid maps based RoI pooling (Hybrid-RoI pooling). Experiments are conducted in STF dataset. The adverse weather enlarges the sparsity of LiDAR point cloud, increasing the difficulty of object detection. SP-Det identifies and leverages foreground region, and achieves the performance better than the current methods. Hence, we believe that SP-Det benefits to LiDAR based 3D scene understanding in the adverse weather. Pei An, Yucong Duan, Yuliang Huang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | ESC-Net: Alleviating Triple Sparsity on 3D LiDAR Point Clouds for Extreme Sparse Scene Completionabstract3D scene completion (SC) has made progress in the last three years. From the application of mobile robot system, SC should support the downstream task (i.e. mapping or perception), instead of only predicting the completed scenes. However, as the low-cost few-beam LiDAR is widely applied in mobile robot, gap between SC and downstream tasks is large. To generate the high quality completion result, the bottleneck lies in the triple sparsity of input, ground truth (GT) occupancy, and GT foreground. To deal with the triple sparsity, we present an extreme sparse scene completion network (ESC-Net). At first, input sparsity hides most of the spatial information of the scene. A feature completion (FC) decoder is designed to mine the spatial feature using feature-level completion. Then, GT occupancy sparsity hinders representation learning of the real scene with continuous surfaces. A multi-view multi-task attention (MMA) loss is presented to recover the high-quality object boundaries via correcting occupancy and semantic labels of regions from 3D and bird's eye view (BEV) spaces. After that, GT foreground sparsity is the imbalance of foreground and background GT labels. It causes the inaccuracy of local 3D object completion. A combination network (ESC-Net-D) is presented to recover 3D structural details of both foreground and background. Experiment is conducted on KITTI and SemanticPOSS datasets. It shows that ESC-Net has the performance higher than current methods not only on completion task, but also on the downstream tasks (i.e. 3D registration, 3D object detection). Hence, we believe that ESC-Net benefits to the community of mobile robot. Source code is released soon. Pei An, Siwen Quan, Junfeng Ding, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Context-Aware Data Augmentation for LIDAR 3d Object DetectionabstractFor 3D LIDAR object detection, data augmentation is an important module to make full use of precious annotated data. As a widely used data augmentation method, GT-aug effectively improves detection performance by inserting sampled groundtruths into LIDAR frames. However, they are often placed in unreasonable areas, leading to the loss of the semantic information between targets and backgrounds during training. To address this problem, we propose a context-aware data augmentation method (CA-aug), which ensures the proper placement of inserted objects by a simple strategy and produces realistic augmented scenes. CA-aug is lightweight and compatible with other augmentation methods. Experiments conducted on KITTI benckmark show that compared with the GT-aug and the similar method in LIDAR-aug (SOTA), it brings higher accuracy to the existing models especially for the detection of cyclists and perdestrians. We also present an in-depth study of augmentation strategies for the range-view-based (RV-based) models and demonstrate that CA-aug can fully exploit the potential of RV-based networks, boosting the moderate mAP of our test model by 8%. Xuzhong Hu, Zaipeng Duan, Xiao Huang 0008, Ziwen Xu, Delie Ming, Jie Ma 0003 |
ICIP | 6 |
| 2023 | Transformer-Based Cross-Modal Information Fusion Network for Semantic Segmentation
Zaipeng Duan, Xiao Huang 0008, Jie Ma 0003 |
Neural Process. Lett. | 3 |
| 2023 | ProUDA: Progressive unsupervised data augmentation for semi-Supervised 3D object detection on point cloud
Pei An, Junxiong Liang, Tao Ma 0004, Yanfei Chen, Liheng Wang, Jie Ma 0003 |
Pattern Recognit. Lett. | 6 |
| 2023 | RS-Aug: Improve 3D Object Detection on LiDAR With Realistic Simulator Based Data AugmentationabstractLight detection and ranging (LiDAR) is an essential sensor for three dimensional (3D) object detection via generating 3D point cloud of the surroundings, and it has been widely used in the various visual applications, especially autonomous driving. However, limited numbers of labeled LiDAR datasets brutally restrain the development of 3D object detector, and this situation breeds an urgent demand on data augmentation in this field. By far, most of the traditional methods reuse the labeled samples, while those unlabeled are hastily untaken. Motivated by this, we propose aRealisticSimulator based data augmentation (RS-Aug). It aims to construct augmented real scenes to enrich the diversity of training dataset. To train 3D object detector in a supervised learning way, the first step of RS-Aug is auto-annotation. Time-continuous LiDAR frames are used to construct the dense scene, which is beneficial to annotation and the subsequent rendering augmentation. However, 3D points with incorrect semantic labels are naturally gathered during multi-view reconstruction, causing the negative effect on auto-annotation. We propose an algorithm of cluster guided$k$-nearest neighbor (c-$k$NN). It emphasizes on de-nosing semantic labels of clustered points using distance and intensity constraints. Then, the next step of RS-Aug is rendering augmentation on the real scene. To enhance the rendering quality using collision and distance constraints with the less computation complexity, we propose a scheme of heuristic search (HS) based object insertion. It estimates the proper position of the inserted object from 2D bird’s eye view (BEV). Experiments demonstrate the de-noising accuracy of c-$k$NN, rendering quality of HS based object insertion, and improvement of RS-Aug on object detection. Pei An, Junxiong Liang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Distribution Aware VoteNet for 3D Object DetectionabstractOcclusion is common in the actual 3D scenes, causing the boundary ambiguity of the targeted object. This uncertainty brings difficulty for labeling and learning. Current 3D detectors predict the bounding box directly, regarding it as Dirac delta distribution. However, it does not fully consider such ambiguity. To deal with it, distribution learning is used to efficiently represent the boundary ambiguity. In this paper, we revise the common regression method by predicting the distribution of the 3D box and then present a distribution-aware regression (DAR) module for box refinement and localization quality estimation. It contains scale adaptive (SA) encoder and joint localization quality estimator (JLQE). With the adaptive receptive field, SA encoder refines discriminative features for precise distribution learning. JLQE provides a reliable location score by further leveraging the distribution statistics, correlating with the localization quality of the targeted object. Combining DAR module and the baseline VoteNet, we propose a novel 3D detector called DAVNet. Extensive experiments on both ScanNet V2 and SUN RGB-D datasets demonstrate that the proposed DAVNet achieves significant improvement and outperforms state-of-the-art 3D detectors. Junxiong Liang, Pei An, Jie Ma 0003 |
AAAI | 3 |
| 2022 | Deep structural information fusion for 3D object detection on LiDAR-camera system
Pei An, Junxiong Liang, Bin Fang 0007, Jie Ma 0003 |
Comput. Vis. Image Underst. | 5 |
| 2022 | Lambertian Model-Based Normal Guided Depth Completion for LiDAR-Camera SystemabstractDepth completion is an essential task for the dense scene reconstruction on light detection and ranging (LiDAR)-camera system. Learning-based method achieves precise depth completion results on specific data sets. However, for the general outdoor scenes with insufficient labeled data sets, an efficient nonlearning method is still required. In this letter, from the geometrical constraint between depth and normal, a novel nonlearning normal guided depth completion method is proposed. For the objects in the outdoor scene, local brightness normal (LBN) constraint is derived from the Lambertian model. It is used to recover dense normal from RGB image and sparse normal. After that, we present a pipeline for depth completion with the guidance of dense normal. Extensive experiments on the KITTI depth completion data set demonstrate that our method achieves smaller root mean squared error (RMSE) than current nonlearning methods. Pei An, Wenxing Fu, Yingshuo Gao, Jie Ma 0003, Jun Zhang 0062, Bin Fang 0007 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | NCFT: Automatic Matching of Multimodal Image Based on Nonlinear Consistent Feature TransformabstractAutomatic matching of multimodal images remains a critical challenging task in many remote sensing and computer vision applications. Due to significant nonlinear radiation distortions (NRDs) between multimodal images, it is difficult for many traditional feature matching methods which are sensitive to NRD to achieve satisfactory matching performance. To cope with this problem, this letter proposes a novel feature matching method named nonlinear consistent feature transform (NCFT) that is robust to large NRD. There are three main contributions to NCFT. First, we propose a new consistent feature map instead of image intensity for feature point detection and description, the consistent feature map encodes the structure information and provides a rich and robust feature. Second, we propose a mean-residual maximum index map (MR-MIM) for feature description and the MR-MIM is constructed from the Log-Gabor convolution sequence on the consistent feature map. Finally, the structure descriptors are built according to the MR-MIM, and multimodal image matching is achieved by computing the correspondence. The extensive experimental results demonstrate that NCFT can effectively overcome the problem of NRD, NCFT outperforms other state-of-the-art methods and improves the matching accuracy and robustness on different multimodal image datasets. Yucong Duan, Bin Fang 0007, Pei An, Jie Ma 0003 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2021 | Feature Interactive Representation for Point Cloud RegistrationabstractPoint cloud registration is the process of using the common structures in two point clouds to splice them together. To find out these common structures and make these structures match more accurately, interacting information of the source and target point clouds is essential. However, limited attention has been paid to explicitly model such feature interaction. To this end, we propose a Feature Interactive Representation learning Network (FIRE-Net), which can explore feature interaction among the source and target point clouds from different levels. Specifically, we first introduce a Combined Feature Encoder (CFE) based on feature interaction intra point cloud. The CFE extracts interactive features intra each point cloud and combines them to enhance the ability of the network to describe the local geometric structure. Then, we propose a feature interaction mechanism inter point clouds which includes a Local Interaction Unit (LIU) and a Global Interaction Unit (GIU). The former is used to interact information between point pairs across two point clouds, thus the point features in one point cloud and its similar point features in another point cloud can be aware of each other. The latter is applied to change the per-point features depending on the global cross information of two point clouds, thus one point cloud has the global perception of another. Extensive experiments on partially overlapping point cloud registration show that our method achieves state-of-the-art performance. Bingli Wu, Jie Ma 0003, Gaojie Chen 0003, Pei An |
ICCV | 2 |
| 2021 | Attention-Based Local Region Aggregation Network For Hierarchical Point Cloud LearningabstractIn point cloud processing, efficient feature aggregation of local regions is essential for hierarchical representations of 3D shapes. However, most prior works directly aggregate local features by symmetric functions, which will result in the desertion of vital shape information. In this paper, we propose a novel aggregation module based on 3D attention mechanism, named Local Region Attention Aggregation, which can capture the local shape implied in irregular points by emphasizing informative features and suppressing unnecessary ones. Specifically, for each local region, our module first encodes point relations, and then sequentially infers two attention maps generated by 3D geometric relations and fusion of multi-scale features respectively. After that, we multiply encoded point features with the attention maps to refine features adaptively. Consequently, by aggregating the higher-quality local features, shape awareness can be enhanced. Extensive experiments on various challenging benchmarks verify our method achieves state-of-the-art performance. Gaojie Chen 0003, Jie Ma 0003, Bingli Wu |
ICIP | 3 |
| 2021 | Straight Sampling Network for Point Cloud LearningabstractSampling operation is a bottleneck of the hierarchical point cloud learning. Existing learnable sampling methods generate a “soft” virtual subset in the training phase, thus distorting the original underlying shape and losing 3D geometric information. In this paper, we propose a novel end-to-end discrete sampling method, named Straight Sampling, to output a “hard” authentic subset with the assistance of Straight Through Estimator. Equipped with Straight Sampling, a hierarchical architecture is developed to learn an effective representation. By grouping and pooling the sampled points in 3D Euclidean space, the network benefits from semantic features as well as 3D geometric information to achieve state-of-the-art performance. Gaojie Chen 0003, Jie Ma 0003, Pei An |
ICIP | 3 |
| 2021 | Multispectral Remote Sensing Image Matching via Image Transfer by Regularized Conditional Generative Adversarial Networks and Local FeatureabstractMultispectral image matching is at the base for many remote sensing and computer vision applications. Due to the different imaging principles and spectra, there are significant nonlinear variations in intensity, texture, and style in multispectral images. This makes it difficult for many classic methods designed for the images of the same spectrum to achieve satisfactory matching performance. To cope with this problem, this letter proposes a new method based on image transfer and local feature for multispectral image matching. First, we propose a new regularized conditional generative adversarial network (GAN) for image transfer to preprocess the multispectral images. This step eliminates the differences in grayscale, texture, and style between the multispectral images. Then, we use a classic local feature to match the generated and original images. We evaluate our method on two commonly used data sets and compare with several state-of-the-art methods. The experiments show that our method performs well by significantly improving the matching accuracy and robustness, and slightly increasing the runtime. Tao Ma 0004, Jie Ma 0003, Jun Zhang 0062, Wenxing Fu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | On shortened 3D local binary descriptors
Siwen Quan, Jie Ma 0003 |
Inf. Sci. | 2 |
| 2018 | Deep joint rain and haze removal from a single imageabstractRain removal from a single image is a challenge which has been studied for a long time. In this paper, a novel convolutional neural network based on wavelet and dark channel is proposed. On one hand, we think that rain streaks correspond to high frequency component of the image. Therefore, haar wavelet transform is a good choice to separate the rain streaks and background to some extent. More specifically, the LL subband of a rain image is more inclined to express the background information, while HL, LH subband tend to represent the rain streaks and the edges respectively. On the other hand, the accumulation of rain streaks from long distance makes the rain image look like haze veil. We extract dark channel of rain image as a feature map in network. By increasing this mapping between the dark channel of input and output images, we achieve haze removal in an indirect way. All of the parameters are optimized by back-propagation. Experiments on both synthetic and realworld datasets reveal that our method outperforms other state-of-the-art methods from a qualitative and quantitative perspective. Zihan Yue, Jie Ma 0003 |
ICPR | 5 |
| 2018 | Local voxelized structure for 3D binary feature representation and robust registration of point clouds from low-cost sensors
Siwen Quan, Jie Ma 0003, Fangyu Hu, Bin Fang 0007, Tao Ma 0004 |
Inf. Sci. | 2 |
| 2018 | Digital video stabilization based on multilayer gray projection
Fangyu Hu, Jie Ma 0003, Huajun Du |
Signal Process. Image Commun. | 2 |
| 2018 | Representing local shape geometry from multi-view silhouette perspective: A distinctive and robust binary 3D feature
Siwen Quan, Jie Ma 0003, Tao Ma 0004, Fangyu Hu, Bin Fang 0007 |
Signal Process. Image Commun. | 2 |
| 2017 | Local voxelized structure for 3D local shape description: A binary representationabstractThis paper proposes a novel binary descriptor named local voxelized structure (LoVS) for 3D local shape description. Unlike many previous local shape descriptors relying on geometric attributes such as curvature and normals, LoVS simply uses point spatial locations to encode the local shape structure represented by point clouds into bit string. Specifically, LoVS is computed on a local cubic volume around the keypoint. The orientation of the cubic is determined by a local reference frame (LRF) to achieve rotation invariance. Then, the cubic is uniformly split into a set of voxels. A voxel is attached with label 1 if there are points inside, otherwise, it produces a 0 bit. All these labels therefore integrates into the LoVS descriptor. We evaluate our method on three public datasets. On each dataset, the LoVS descriptor outperforms all other descriptors tested. Siwen Quan, Jie Ma 0003, Fangyu Hu, Bin Fang 0007, Tao Ma 0004 |
ICIP | 2 |
| 2016 | Fast motion deblurring using gyroscopes and strong edge predictionabstractThis paper presents a fast deblurring algorithm to remove camera motion blur from a single photograph using built-in gyroscopes and strong edge prediction. An inaccurate blur kernel or point spread function (PSF) usually leads to an unsatisfying restored result. Hence, we propose a robust three-phase method for accurate PSF estimation. In the first stage, we utilize the embedded gyroscopes to compute a coarse version of the PSF from the camera's angular velocity during an exposure. In order to reduce the execution time of the later PSF modification, we introduce a patch selection procedure in the second stage to choose a suitable region from the blurry image based on the size of the coarse PSF estimated in stage one. The third phase aims to modify the coarse PSF to obtain an accurate one by predicting strong edges from an estimated latent image. In our experiments, we compare the restoration performance of several state-of-the-art approaches including ours and find that the proposed method outperforms others qualitatively as well as quantitatively. In addition, our method is also compared with the multi-scale approach without gyroscope data and shows shorter processing time and comparable deblurring quality. To the best of our knowledge, this is the first work that combines the sensor-aided method with the image-based approach to estimate the blur kernel. Jiacai Zhao, Jie Ma 0003, Bin Fang 0007, Siwen Quan, Fangyu Hu |
ICPR | 2 |
| 2015 | Salient object detection via contrast information and object vision organization cues
Shengxiang Qi, Jin-Gang Yu, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian |
Neurocomputing | 3 |
| 2015 | Unsupervised Ship Detection Based on Saliency and S-HOG Descriptor From Optical Satellite ImagesabstractWith the development of high-resolution imagery, ship detection in optical satellite images has attracted a lot of research interest because of the broad applications in fishery management, vessel salvage, etc. Major challenges for this task include cloud, wave, and wake clutters, and even the variability of ship sizes. In this letter, we propose an unsupervised ship detection method toward overcoming these existing issues. Visual saliency, which focuses on highlighting salient signals from scenes, is applied to extract candidate regions followed by a homogeneous filter presented to confirm suspected ship targets with complete profiles. Then, a novel descriptor, ship histogram of oriented gradient, which characterizes the gradient symmetry of ship sides, is provided to discriminate real ships. Experimental results on numerous panchromatic satellite images demonstrate the good performance of our method compared to state-of-the-art methods. Shengxiang Qi, Jie Ma 0003, Yansheng Li 0001, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Accurate Aerial Object Localization Using Gravity and Gravity Gradient AnomalyabstractAutonomous underwater vehicles (AUVs) have been widely used in diverse contexts, especially military affairs. The smooth operation of the AUV requires accurate localization of surrounding objects, especially the aerial objects. In this letter, a novel and practical method is presented for aerial object localization by using gravity and gravity gradient anomaly. Different from the state-of-the-art methods, such as GPS, radar, and laser, the proposed method runs in a passive manner and achieves AUV invisibility without energy emission. Compared with the object localization methods based on gravity and gravity gradient inversion, the proposed method is more practical as no large area gravity and gravity gradient measurements are needed to estimate the object mass. Experimental results demonstrate that the proposed method performs better than the existing methods. Zu Yan, Jie Ma 0003, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Visual saliency detection using feature activity weighted decorrelation cuesabstractIn this paper, a novel model based on feature activity weighted decorrelation cues is proposed for visual saliency detection in natural images. It consists of two parts: the feature decorrelation and feature information-activity. For the first part, Laplacian sparse coding and low-rank decomposition are used to extract decorrelated features from the scenes. For the second part, Incremental Coding Length is applied to measure the information-activity contained in features, which is then employed to weight the decorrelated features. Finally, visual saliency is estimated through a max pooling strategy. Experimental results on a publicly available benchmark demonstrate the effectiveness of our proposed model with good performance against the state-of-the-art methods. Shengxiang Qi, Jin-Gang Yu, Ji Zhao 0001, Jie Ma 0003, Jinwen Tian |
ICIP | 4 |
| 2014 | Modeling Local Gravity Anomaly Self-Adaption Quotient Reference Maps for Underwater Autonomous NavigationabstractGravity navigation, with its independent, passive, concealment and all weather, has become one of the best options of aided inertial navigation system (INS). Precise local gravity or gravity anomaly reference maps will greatly improve the accuracy of autonomous underwater vehicles (AUVs) navigation. Due to the lack of measured gravity data, the previous methods generally used digital elevation model (DEM) to model gravity anomaly reference maps, however, which neglected the impact of terrain density in homogeneity. In this paper, a novel and practical method is proposed for modeling a reference map which takes a full consideration of the terrain density difference. Experimental results show that the proposed method performs better than the existing methods. Zu Yan, Jie Ma 0003, Jinwen Tian |
ICTAI | 2 |
| 2014 | A Gravity Gradient Differential Ratio Method for Underwater Object DetectionabstractThe smooth operation of autonomous underwater vehicles (AUVs) relies heavily on the accurate detection of surrounding objects. Toward this end, this letter presents a novel method for underwater object detection based on the gravity gradient differential and the gravity gradient differential ratio caused by the relative motion between the AUV and the object. Unlike the existing techniques, the proposed method works in a passive manner and achieves AUV invisibility without energy emission. In addition, for the proposed method, no gravity map or gravity gradient map is required, which improves its practicality. Experimental results demonstrate that the proposed method performs better than the existing methods. Zu Yan, Jie Ma 0003, Jinwen Tian, Hai Liu 0004, Jingang Yu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2013 | A Robust Directional Saliency-Based Method for Infrared Small-Target Detection Under Various Complex BackgroundsabstractInfrared small-target detection plays an important role in image processing for infrared remote sensing. In this letter, different from traditional algorithms, we formulate this problem as salient region detection, which is inspired by the fact that a small target can often attract attention of human eyes in infrared images. This visual effect arises from the discrepancy that a small target resembles isotropic Gaussian-like shape due to the optics point spread function of the thermal imaging system at a long distance, whereas background clutters are generally local orientational. Based on this observation, a new robust directional saliency-based method is proposed incorporating with visual attention theory for infrared small-target detection. Experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods for real infrared images with various typical complex backgrounds. Shengxiang Qi, Jie Ma 0003, Chao Tao 0001, Changcai Yang, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2011 | A robust method for vector field learning with application to mismatch removingabstractWe propose a method for vector field learning with outliers, called vector field consensus (VFC). It could distinguish inliers from outliers and learn a vector field fitting for the inliers simultaneously. A prior is taken to force the smoothness of the field, which is based on the Tiknonov regularization in vector-valued reproducing kernel Hilbert space. Under a Bayesian framework, we associate each sample with a latent variable which indicates whether it is an inlier, and then formulate the problem as maximum a posteriori problem and use Expectation Maximization algorithm to solve it. The proposed method possesses two characteristics: 1) robust to outliers, and being able to tolerate 90% outliers and even more, 2) computationally efficient. As an application, we apply VFC to solve the problem of mismatch removing. The results demonstrate that our method outperforms many state-of-the-art methods, and it is very robust. Ji Zhao 0001, Jiayi Ma 0001, Jinwen Tian, Jie Ma 0003, Dazhi Zhang |
CVPR | 4 |
| 2010 | Underwater Object Detection Based on Gravity GradientabstractA novel method of underwater object detection based on gravity gradient is presented, which can be used on autonomous underwater vehicles (AUVs) to detect abnormal objects underwater. Gravity gradient anomalies of partial area, which are caused by the object, can be measured by a gravity gradiometer on an AUV. Then, anomalies can be inversed with a gravity gradient inversion algorithm, so the mass and barycenter of an object can be estimated. Simulation results show that approximate information of an object can be provided by the proposed method. Xin Tian 0006, Jie Ma 0003, Jinwen Tian |
IEEE Geosci. Remote. Sens. Lett. | 3 |