EDBT 2026 Demo / reviewers in the wild / expert
Lijun Zhao 0003
dblp:06/5162-3
· DBLP profile ↗
40ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0002-9108-8276ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 11 since 2021Computer networks · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Spatio-Temporal Alignment of End-to-End 3D PerceptionabstractSpatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model with a unified explicit physical model (constant velocity, etc.). These approaches prefer semantic features for implicit alignment, challenging the importance of explicit motion modeling in the traditional perception paradigm. However, variations in motion states and object features across categories and frames render this alignment suboptimal. To address this, we propose HAT, a spatio-temporal alignment module that allows each object to adaptively decode the optimal alignment proposal from multiple hypotheses without direct supervision. Specifically, HAT first utilizes multiple explicit motion models to generate spatial anchors and motion-aware feature proposals for historical instances. It then performs multi-hypothesis decoding by incorporating semantic and motion cues embedded in cached object queries, ultimately providing the optimal alignment proposal for the target frame. On nuScenes, HAT consistently improves 3D temporal detectors and trackers across diverse baselines. It achieves state-of-the-art tracking results with 46.0% AMOTA on the test set when paired with the DETR3D detector. In an object-centric E2E AD method, HAT enhances perception accuracy (+1.3% mAP, +3.1% AMOTA) and reduces the collision rate by 32%. When semantics are corrupted (nuScenes-C), the enhancement of motion modeling by HAT enables more robust perception and planning in the E2E AD. Peidong Li, Dedong Liu, Jiajia Fu, Dixiao Cui, Lijun Zhao 0003, Lining Sun |
AAAI | 9 |
| 2026 | GFMLLM: Enhance multi-modal large language model for global and fine-grained visual spatial perception
Zhendong Fan, Jinyang Gao, Hongbo Gao 0008, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 8 |
| 2026 | MTNet: A Mixed Transformer Network for High-Quality 3-D Object Detectionabstract3D object detection from point clouds represents a formidable challenge, necessitating the accurate identification and localization of objects within a 3D space. Recent advancements have showcased the efficacy of point-based detectors, leveraging local aggregators to encode intricate structural details of the point cloud. However, a notable limitation resides in their treatment of each point and object proposal in isolation, devoid of considering the interrelationships among them, thus impeding the overall detection performance. In this work, we argue that the integration of contextual information is paramount, particularly in the realm of indoor 3D object detection. Indoor environments are inherently characterized by robust contextual constraints, providing a rich tapestry for enhanced scene comprehension. In this way, we introduce MTNet, a mixed transformer network for high-quality indoor 3D object detection. Technically, we develop a mixed transformer (MixFormer) block that is purpose-built to intricately model the synergistic interplay between local structural information and global contextual features of 3D point clouds. In contrast to the classical transformer, our proposed MixFormer incorporates a local feature aggregator engineered to capture local geometric information while leveraging a KNN(K-Nearest Neighbors)-based attention mechanism to aggregate global contextual information. Furthermore, we suggest a feature compensator to adaptively fuse the strengths of local and global features, further bolstering detection performance. The culmination of these proposed components results in our MTNet framework, a hierarchical, versatile pipeline that consistently outperforms existing works across a multitude of benchmarks. In addition, we affirm the potential of proposed MixFormer and FC as generic modules that are capable of augmenting performance across a spectrum of 3D downstream point cloud tasks. Ruqi Liu, Shuaiyan Liu, Linqi Yang, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
IEEE Internet Things J. | 7 |
| 2026 | OFVL-MS++: Once for visual localization across multiple scenes via a two-stage framework
Chunsheng Yang, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 8 |
| 2026 | An MoE-Driven Unified Image Restoration Framework for Adverse Weather ConditionsabstractAdverse natural weather conditions frequently cause substantial performance degradation in outdoor vision systems, underscoring the critical importance of research on image restoration techniques. Employing a unified set of network parameters to restore degraded images across diverse weather conditions has emerged as a key research direction in the field of image restoration. In this work, we propose MUIRF, a Mixture-of-Experts (MoE)-driven unified image restoration framework for multiple adverse weather conditions. Specifically, our technical contribution includes a novel channel-level parameter sharing strategy guided by a shallow-feature-based MoE (CPSM). This fine-grained parameter sharing strategy adaptively selects convolution weight channels for cross-task sharing based on the input image, enabling the network to accurately capture weather-general features, while the remaining channels encode weather-specific features corresponding to each weather condition. CPSM facilitates precise channel selection, thereby enhancing the robustness and accuracy of MUIRF during joint training across diverse image restoration tasks under varying weather conditions. Additionally, gradient conflicts inevitably arise in shared parameters due to the divergent optimization objectives across tasks. To address this challenge, we propose a meta-vector-guided gradient homogenization (MVGH) algorithm that mitigates inter-task gradient conflicts and improves image restoration quality. Comprehensive experimental evaluations demonstrate that our proposed network outperforms most state-of-the-art approaches, validating its superior performance and effectiveness. Hongbo Gao 0008, Ruqi Liu, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | COME: A Collaborative Optimization Framework With Low-Rank MoE for Indoor 3D Object DetectionabstractIndoor 3D object detection serves as a fundamental task in computer vision and robotics. Existing research predominantly focuses on training domain-specific optimal models for individual datasets, yet it overlooks the potential value of capturing universal geometric attributes that can substantially enhance object detection performance across diverse domains. To resolve this gap, we propose COME, a novel and effective collaborative optimization framework designed to seamlessly integrate these universal attributes while preserving the domain-specific characteristics of each dataset domain. COME is built on VoteNet and incorporates a Cross-Domain Expert Parameter Sharing Strategy (CEPSS) that draws inspiration from the Mixture of Experts (MoE) framework. Its core innovation resides in the dual-expert design of CEPSS: domain-shared experts capture universal geometric relationships across datasets, whereas domain-specific experts encode unique features for individual datasets. This separation enables the model to focus on learning both generic and domain-specialized visual cues, without mutual interference. In addition, to dynamically adapt to different domains, we design a lightweight gating network that automatically selects relevant experts, eliminating irrelevant feature interference and enhancing model specialization. Compared to standard parameter-sharing architectures, this design significantly reduces gradient conflicts during multi-domain training. We further optimize computational efficiency by implementing low-rank structures for domain-shared and domain-specific experts, thus striking a better balance between memory overhead and detection performance. Experiments show that COME achieves state-of-the-art results across benchmarks, with acceptable parameter growth, and outperforms existing multi-domain detection methods. Hongbo Gao 0008, Zimeng Tong, Fuyuan Qiu, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 6 |
| 2025 | DVDS: A deep visual dynamic slam system
Tao Xie 0010, Qihao Sun, Tao Sun 0024, Jinhang Zhang, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
Expert Syst. Appl. | 6 |
| 2025 | GCPNet: Gradient-aware channel pruning network with bilateral coupled sampling strategyabstractIn the realm of deep neural network optimization, network slimming has emerged as a key technique for reducing model size without significantly compromising performance. Traditional pruning algorithms use neural architecture search (NAS) to identify networks with adjustable widths, focusing on extracting representative subnets across varied pruning ratios. However, a significant challenge lies in ensuring that the pruned network maintains high performance (accuracy) while substantially reducing the model size and computational cost. To tackle this challenge, we introduce GCPNet, a gradient-aware channel pruning network designed for efficient neural network slimming. Specifically, GCPNet incorporates a bilateral coupled sampling strategy (BCSS) to sample the smallest, largest, and several middle-sized models and perform forward and backward passes in each training iteration. The gradients of these models are then fused to update the overarching supernet. In addition, we develop a gradient-aware homogenization technique (GHT) to mitigate gradient conflicts between the supernet and the sampled models due to differing gradient directions , this accelerates the convergence of the supernet and ensures comprehensive training for GCPNet. The trained supernet serves as a reliable performance indicator, with the performance of architectures ranked by our supernet exhibiting a high correlation with true performance. Extensive experiments validate that GCPNet outperforms current advanced channel pruning approaches on the ImageNet dataset while employing comparable FLOPs and parameters. For example, under the constraints of 100 Flops and 50 Flops, our pruned MobilenetV2 achieved 68.7% and 63.5% Top-1 accuracy on the ImageNet dataset, outperforming the most advanced BCNet by 0.7% and 0.8% respectively. Chuqing Cao, Fangjun Zheng, Tao Sun 0024, Lijun Zhao 0003 |
Expert Syst. Appl. | 5 |
| 2025 | Adaptive-LIO: Enhancing Robustness and Precision Through Environmental Adaptation in LiDAR Inertial OdometryabstractThe emerging Internet of Things (IoT) applications, such as driverless cars, have a growing demand for high-precision positioning and navigation. Nowadays, LiDAR inertial odometry (LIO) becomes increasingly prevalent in robotics and autonomous driving. However, many current SLAM systems lack sufficient adaptability to various scenarios. Challenges include decreased point cloud accuracy with longer frame intervals under the constant velocity assumption, coupling of erroneous IMU information when IMU saturation occurs, and decreased localization accuracy due to the use of fixed-resolution maps during indoor-outdoor scene transitions. To address these issues, we propose a loosely coupled adaptive LIO named Adaptive-LIO, which incorporates adaptive segmentation to enhance mapping accuracy, adapts motion modality through IMU saturation and fault detection, and adjusts map resolution adaptively using multiresolution voxel maps based on the distance from the LiDAR center. Our proposed method has been tested in various challenging scenarios, demonstrating the effectiveness of the improvements we introduce. The code is open-source on GitHub: Adaptive-LIO. Chengwei Zhao 0003, Kun Hu 0016, Jie Xu 0066, Lijun Zhao 0003, Baiwen Han, Kaidi Wu, Maoshan Tian, Shenghai Yuan 0001 |
IEEE Internet Things J. | 4 |
| 2025 | Poly-TF: A polymeric transformer framework for multiple visual tasks at once
Xuan Fan, Tao Sun 0024, Jinghan Gao, Tao Xie 0010, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 7 |
| 2025 | MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation
Zimeng Tong, Youran Du, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 8 |
| 2025 | Ada-Matcher: A deep detector-based local feature matcher with adaptive weight sharing
Fangjun Zheng, Chuqing Cao, Tao Sun 0024, Jinhang Zhang, Lijun Zhao 0003 |
Knowl. Based Syst. | 6 |
| 2025 | FSPDD: A double-branch attention guided network for few-shot PCB defect detectionabstractAbstract During the production of printed circuit board (PCB), there will be defects due to inappropriate operations, which will affect the use of electronic products. Majority defect detection methods cost a large number of annotated samples to train detection models. However, PCB defect samples are difficult to collect. Moreover, existing few-shot object detection methods tend to extracting low-level features from support and query images via the shared backbone such as ResNet-50. However, it is not sufficient to obtain fine-grained prior guidance. To address the above issues, we propose a few-shot PCB defect detection model with double-branch attention. Specifically, the joint attention enhancement (JAE) module is proposed to fully mine effective information of query PCB images in multiple dimensions to enhance the representation of latent defects. Then, the multi-scale guidance (MSG) module is proposed to integrate prior knowledge within support PCB images into vectors to reweight query PCB images. Experiments on the PCB defect dataset demonstrate that AP of FSPDD outperforms state-of-the-art methods under different shot settings (k=1,2,3,5,10,30) and our proposed FSPDD has a good generalization ability, in which AP reachs 0.273 when $$k=30$$ k = 30 and is 5.28% higher than SOTA methods. Kehao Shi, Zhenyi Xu, Yang Cao 0010, Lijun Zhao 0003, Yu Kang 0001 |
Multim. Tools Appl. | 4 |
| 2025 | CTFS: A consolidated transformer framework for instance and semantic segmentation tasks
Fuyuan Qiu, Hongbo Gao 0008, Tao Xie 0010, Chuqing Cao, Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028 |
Neural Networks | 7 |
| 2025 | AAPMatcher: Adaptive attention pruning matcher for accurate local feature matching
Xuan Fan, Shuaiyan Liu, Lijun Zhao 0003, Ruifeng Li 0001 |
Neural Networks | 4 |
| 2025 | VD-Matcher: A Very Deep Local Feature Matcher With Weight Recycling and Keypoint DetectionabstractEstablishing local feature matches between image pairs serves as a fundamental component of plentiful vision tasks, such as visual localization and Structure from Motion (SfM). Recently, detector-free techniques equipped with the transformer have exhibited exceptional performance. Theoretically, optimizing the transformer architecture and stacking more transformer blocks could emphasize crucial features and filter out extraneous information by progressively narrowing the effective perception regions of the network within images, thereby enhancing the matching performance. Nevertheless, this paradigm results in a linear escalation of model size with respect to the number of blocks. In this study, we introduce VD-Matcher to address this issue. A principal innovation of VD-Matcher is the utilization of a weight recycling technique (WRT) that enables partial weights to be reutilized across successive transformer blocks, along with specific transformations designed to sufficiently enhance feature representations. This approach enables VD-Matcher to construct a deep transformer architecture for accurate local feature matching while maintaining a manageable parameter size. Furthermore, we propose a lightweight multi-scale keypoint detection module that captures representative keypoints to replace all keypoints for compact global information aggregation intra-/inter- images, which reduces the computational overhead induced by excessively deep transformer layers while alleviating redundant information propagation to a certain extent. Extensive experiments verify that VD-Matcher exceeds state-of-the-art algorithms on multiple benchmarks while maintaining less parameters. The source code is available at https://github.com/mooncake199809/VD-Matcher. Qihao Sun, Tao Xie 0010, Hongbo Gao 0008, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | AGHL: Anchor-Guided Point Cloud Registration Network With Hybrid Local Feature PerceptionabstractPoint cloud registration, which estimates a rigid transformation matrix between two point clouds, is a fundamental process in numerous applications. While existing detector-free techniques present exceptional performance, they overlook the extraction of hybrid local features that capture correlations between points and their neighbours, thereby limiting the quality of point cloud recognition. Moreover, these approaches typically treat point clouds as sequential data and employ the transformer to integrate global context from all points, which inevitably introduces interference from irrelevant regions, hence affecting the registration accuracy. In this work, we propose a novel detector-free approach AGHL to address these challenges. For the first issue, AGHL introduces a hybrid local feature perception module that designs two parallel branches to concurrently extract low-level and high-level local features, which effectively encode the correlations between each point and its neighborhood points in both Euclidean space and high-dimensional feature space. For the second issue, AGHL develops an anchor-guided cross attention that adheres to the local geometric consistency to constrain the network's attention on reliable anchors, thereby effectively suppressing interference from irrelevant regions. Benefiting from these techniques, AGHL achieves impressive point cloud registration accuracy across all synthetic, indoor, and outdoor datasets. Furthermore, we build an experimental platform and conduct a real-world robot localization experiment, with results showing the strong generalization ability of AGHL. Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Chuqing Cao |
IEEE Trans. Image Process. | 6 |
| 2025 | SOFW: A Synergistic Optimization Framework for Indoor 3D Object DetectionabstractIn this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar |
IEEE Trans. Multim. | 8 |
| 2025 | Centra-Net: A Centralized Network for Visual Localization Spanning Multiple ScenesabstractWe present Centra-Net, a centralized network that concurrently optimizes visual localization over numerous scenes under heterogeneous dataset domains. Centra-Net exemplifies storage efficiency by amalgamating multiple models with task-shared parameters into a singular cohesive structure. Technically, we develop abasic feature extraction unit (BFEU)with two parallel branches: one dedicated to local feature extraction and the other adept at adaptively generating a task-specific attention mask for feature calibration, thus bolstering its feature extraction capability across diverse scenes. Based on the BFEU, we introduce afilter-wise sharing mechanism (FSM)that adaptively determines parameter sharing within the unit, thus facilitating fine-grained parameter allocation. The key insight of FSM resides in reconceptualizing the parameter sharing of the unit as a learnable paradigm, enabling the determination of shared parameters to be made post-training. Finally, we suggest acomplexity-prioritized gradient algorithm (CPGA)that capitalizes on task complexity to attain a harmonious learning space for various tasks, thus safeguarding optimal performances across all tasks. Through rigorous experiments on numerous benchmarks, Centra-Net demonstrates a notable edge over existing state-of-the-art works while operating with a significantly reduced parameter footprint. Ke Wang 0028, Tao Xie 0010, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Multim. | 8 |
| 2025 | HVLF: A Holistic Visual Localization Framework Across Diverse ScenesabstractRecently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods. Fuyuan Qiu, Dedong Liu, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | ViT-MVT: A Unified Vision Transformer Network for Multiple Vision TasksabstractIn this work, we seek to learn multiple mainstream vision tasks concurrently using a unified network, which is storage-efficient as numerous networks with task-shared parameters can be implanted into a single consolidated network. Our framework, vision transformer (ViT)-MVT, built on a plain and nonhierarchical ViT, incorporates numerous visual tasks into a modest supernet and optimizes them jointly across various dataset domains. For the design of ViT-MVT, we augment the ViT with a multihead self-attention (MHSE) to offer complementary cues in the channel and spatial dimension, as well as a local perception unit (LPU) and locality feed-forward network (locality FFN) for information exchange in the local region, thus endowing ViT-MVT with the ability to effectively optimize multiple tasks. Besides, we construct a search space comprising potential architectures with a broad spectrum of model sizes to offer various optimum candidates for diverse tasks. After that, we design a layer-adaptive sharing technique that automatically determines whether each layer of the transformer block is shared or not for all tasks, enabling ViT-MVT to obtain task-shared parameters for a reduction of storage and task-specific parameters to learn task-related features such that boosting performance. Finally, we introduce a joint-task evolutionary search algorithm to discover an optimal backbone for all tasks under total model size constraint, which challenges the conventional wisdom that visual tasks are typically supplied with backbone networks developed for image classification. Extensive experiments reveal that ViT-MVT delivers exceptional performances for multiple visual tasks over state-of-the-art methods while necessitating considerably fewer total storage costs. We further demonstrate that once ViT-MVT has been trained, ViT-MVT is capable of incremental learning when generalized to new tasks while retaining identical performances for trained tasks. The code is available at https://github.com/XT-1997/vitmvt. Tao Xie 0010, Ruifeng Li 0001, Shouren Mao, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | I2EKF-LO: A Dual-Iteration Extended Kalman Filter Based LiDAR OdometryabstractLiDAR odometry is a pivotal technology in the fields of autonomous driving and autonomous mobile robotics. However, most of the current works focus on nonlinear optimization methods, and still existing many challenges in using the traditional Iterative Extended Kalman Filter (IEKF) framework to tackle the problem: IEKF only iterates over the observation equation, relying on a rough estimate of the initial state, which is insufficient to fully eliminate motion distortion in the input point cloud; the system process noise is difficult to be determined during state estimation of the complex motions; and the varying motion models across different sensor carriers. To address these issues, we propose the Dual-Iteration Extended Kalman Filter (I2EKF) and the LiDAR odometry based on I2EKF (I2EKF-LO). This approach not only iterates over the observation equation but also leverages state updates to iteratively mitigate motion distortion in LiDAR point clouds. Moreover, it dynamically adjusts process noise based on the confidence level of prior predictions during state estimation and establishes motion models for different sensor carriers to achieve accurate and efficient state estimation. Comprehensive experiments demonstrate that I2EKF-LO achieves outstanding levels of accuracy and computational efficiency in the realm of LiDAR odometry. Additionally, to foster community development, our code is open-sourced.1 Wenlu Yu, Jie Xu 0066, Chengwei Zhao 0003, Lijun Zhao 0003, Thien-Minh Nguyen, Shenghai Yuan 0001, Mingming Bai, Lihua Xie 0001 |
IROS | 4 |
| 2024 | FMAP: Learning robust and accurate local feature matching with anchor points
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 6 |
| 2024 | ALNet: An adaptive channel attention network with local discrepancy perception for accurate indoor visual localization
Hongbo Gao 0008, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Mengyuan Wu |
Expert Syst. Appl. | 5 |
| 2024 | APM: Adaptive parameter multiplexing for class incremental learning
Jinghan Gao, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
Expert Syst. Appl. | 5 |
| 2024 | CorMatcher: A corners-guided graph neural network for local feature matching
Hainan Luo, Tao Xie 0010, Chuqing Cao, Lijun Zhao 0003 |
Expert Syst. Appl. | 6 |
| 2024 | DeepMatcher: A deep transformer-based network for robust and accurate local feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 5 |
| 2024 | ActionMixer: Temporal action detection with Optimal Action Segment Assignment and mixers
Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
Expert Syst. Appl. | 3 |
| 2024 | M-DIVO: Multiple ToF RGB-D Cameras-Enhanced Depth-Inertial-Visual OdometryabstractTime-of-Flight (ToF) RGB-D cameras provide a wealth of information for SLAM systems. However, the limited field of view (FOV) of a single ToF RGB-D camera and the small range of its depth measurement module make it prone to degeneracy when relying solely on visual or depth information for SLAM, a problem typical of unimodal SLAM algorithms. To address this issue, this article presents M-DIVO: an IEKF-based odometry that fuses visual, depth (similar to LiDAR), and inertial modules from multiple ToF RGB-D cameras. It comprises two direct method subsystems: 1) the depth–inertial odometry (DIO) subsystem, which constructs point-to-plane constraints from multiple depth modules and 2) the visual–inertial odometry (VIO) subsystem, which optimizes pose using photometric error constructed by multiple cameras. Additionally, to manage the significant computational load from processing multiple sensors and multimodal information, we introduce a multimodal redundancy scheduling mechanism (MRSM): prioritizing the DIO subsystem with the VIO subsystem as auxiliary, executing the VIO subsystem only when degeneracy occurs in the DIO subsystem. We also propose a “External First, Internal Last” strategy for calibrating multiple external and internal sensors. Experiments demonstrate that compared to unimodal SLAM, our method achieves higher robustness and precision, as well as satisfactory real-time performance. The proposed calibration strategy is demonstrated to be more accurate than the traditional inertial measurement unit-centric approach. The code is open source. Jie Xu 0066, Wenlu Yu, Shenghai Yuan 0001, Lijun Zhao 0003, Ruifeng Li 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 5 |
| 2024 | CO-Net++: A Cohesive Network for Multiple Point Cloud Tasks at Once With Two-Stage Feature RectificationabstractWe present CO-Net++, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains with a two-stage feature rectification strategy. The core of CO-Net++ lies in optimizing task-shared parameters to capture universal features across various tasks while discerning task-specific parameters tailored to encapsulate the unique characteristics of each task. Specifically, CO-Net++ develops a two-stage feature rectification strategy (TFRS) that distinctly separates the optimization processes for task-shared and task-specific parameters. At the first stage, TFRS configures all parameters in backbone as task-shared, which encourages CO-Net++ to thoroughly assimilate universal attributes pertinent to all tasks. In addition, TFRS introduces a sign-based gradient surgery to facilitate the optimization of task-shared parameters, thus alleviating conflicting gradients induced by various dataset domains. In the second stage, TFRS freezes task-shared parameters and flexibly integrates task-specific parameters into the network for encoding specific characteristics of each dataset domain. CO-Net++ prominently mitigates conflicting optimization caused by parameter entanglement, ensuring the sufficient identification of universal and specific features. Extensive experiments reveal that CO-Net++ realizes exceptional performances on both 3D object detection and 3D semantic segmentation tasks. Moreover, CO-Net++ delivers an impressive incremental learning capability and prevents catastrophic amnesia when generalizing to new point cloud tasks. Tao Xie 0010, Qihao Sun, Chuqing Cao, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | OAMatcher: An overlapping areas-based network with label credibility for robust and accurate feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 6 |
| 2023 | OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesabstractIn this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net. Tao Xie 0010, Siyi Lu, Ke Wang 0028, Jinghan Gao, Dedong Liu, Jie Xu 0066, Lijun Zhao 0003, Ruifeng Li 0001 |
ICCV | 9 |
| 2023 | CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive NetworkabstractWe present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task. Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001 |
ICCV | 9 |
| 2023 | Poly-MOT: A Polyhedral Framework For 3D Multi-Object Trackingabstract3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association and state estimation for all objects. With large-scale modern datasets and real scenes, there are a variety of object categories that commonly exhibit distinctive geometric properties and motion patterns. In this way, such distinctions would enable various object categories to behave differently under the same standard, resulting in erroneous matches between trajectories and detections, and jeopardizing the reliability of downstream tasks (navigation, etc.). Towards this end, we propose Poly-MOT, an efficient 3D MOT method based on the Tracking-By-Detection framework that enables the tracker to choose the most appropriate tracking criteria for each object category. Specifically, Poly-MOT leverages different motion models for various object categories to characterize distinct types of motion accurately. We also introduce the constraint of the rigid structure of objects into a specific motion model to accurately describe the highly nonlinear motion of the object. Additionally, we introduce a two-stage data association strategy to ensure that objects can find the optimal similarity metric from three custom metrics for their categories and reduce missing matches. On the NuScenes dataset, our proposed method achieves state-of-the-art performance with 75.4% AMOTA. The code is available at https://github.com/lixiaoyu20001P0Iy-MOT. Tao Xie 0010, Dedong Liu, Jinghan Gao, Lijun Zhao 0003, Ke Wang 0028 |
IROS | 7 |
| 2023 | Point-NAS: A Novel Neural Architecture Search Framework for Point Cloud AnalysisabstractRecently, point-based networks have exhibited extraordinary potential for 3D point cloud processing. However, owing to the meticulous design of both parameters and hyperparameters inside the network, constructing a promising network for each point cloud task can be an expensive endeavor. In this work, we develop a novel one-shot search framework called Point-NAS to automatically determine optimum architectures for various point cloud tasks. Specifically, we design an elastic feature extraction (EFE) module that serves as a basic unit for architecture search, which expands seamlessly alongside both the width and depth of the network for efficient feature extraction. Based on the EFE module, we devise a searching space, which is encoded into a supernet to provide a wide number of latent network structures for a particular point cloud task. To fully optimize the weights of the supernet, we propose a weight coupling sandwich rule that samples the largest, smallest, and multiple medium models at each iteration and fuses their gradients to update the supernet. Furthermore, we present a united gradient adjustment algorithm that mitigates gradient conflict induced by distinct gradient directions of sampled models and supernet, thus expediting the convergence of the supernet and assuring that it can be comprehensively trained. Pursuant to the provided techniques, the trained supernet enables a multitude of subnets to be incredibly well-optimized. Finally, we conduct an evolutionary search for the supernet under resource constraints to find promising architectures for different tasks. Experimentally, the searched Point-NAS with weights inherited from the supernet realizes outstanding results across a variety of benchmarks. i.e., 94.2% and 88.9% overall accuracy under ModelNet40 and ScanObjectNN, 68.6% mIoU under S3DIS, 63.6% and 69.3% [email protected] under SUN RGB-D and ScanNet V2 datasets. Tao Xie 0010, Linqi Yang, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 7 |
| 2022 | Adaptive Update Tracking Algorithm for Fast Motion ObjectabstractAiming at the problem of object loss and tracking precision decrease due to the fast movement of an object, this paper proposes an adaptive updating algorithm for a fast moving object in video based on the KCF algorithm. The target motion information is obtained by simple motion estimation based on the motion law of the object, and an object tracking strategy is proposed for the fast motion scene based on the adaptive search area and template updating combined with the relevant response. OTB100 dataset was selected for experiments. Experimental results show that the proposed algorithm has improved tracking precision and success rate compared with the baseline algorithm in the scene of fast object movement, which can effectively solve the problem of tracking failure caused by fast object movement and achieve stable object tracking. Haozheng Qian, Mingxing Fang, Jinhua She, Lijun Zhao 0003, Youwu Du |
IECON | 4 |
| 2022 | Fast Detection of Multi-Direction Remote Sensing Ship Object Based on Scale Space PyramidabstractShips in remote sensing images are usually arranged in arbitrary direction, small in size, and densely arranged. As a result, existing object detection algorithms cannot detect ships quickly and accurately. In order to solve the above problems, a lightweight object detection network for fast detection of ships is proposed. The network is composed of backbone network, four-scale fusion network and rotation branch. First, a lightweight network unit S-LeanNet is designed and used to build a low-computing and accurate backbone network. Then, a four-scale feature fusion module is designed to generate a four-scale feature pyramid, which contains more features such as ship shape and texture, and at the same time is conducive to the detection of small ships. Finally, a novel rotation branch module is designed, using balance L1 loss function and R-NMS for post-processing, to realize the precise positioning and regression of the rotating bounding box in one step. Experimental results show that the detection precision of our method in the DOT A remote sensing data set is compared with the latest SCRDet detection method, the precision is increased by 1.1%, and the operating speed is increased by 8 times, which can meet the fast detection requirements of ships. Ziying Song, Li Wang 0092, Caiyan Jia, Jiangfeng Bi, Haiyue Wei, Yongchao Xia, Lijun Zhao 0003 |
MSN | 9 |
| 2021 | Multi-obstacle path planning and optimization for mobile robot
Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028, Xichun Gui |
Expert Syst. Appl. | 3 |
| 2018 | Feature-Based and Convolutional Neural Network Fusion Method for Visual RelocalizationabstractRelocalization is one of the necessary modules for mobile robots in long-term autonomous movement in an environment. Currently, visual relocalization algorithms mainly include feature-based methods and CNN-based (Convolutional Neural Network) methods. Feature-based methods can achieve high localization accuracy in feature-rich scenes, but the error is quite large or it even fails in cases with motion blur, texture-less scene and changing view angle. CNN-based methods usually have better robustness but poor localization accuracy. For this reason, a visual relocalization algorithm that combines the advantages of the two methods is proposed in this paper. The BoVW (Bag of Visual Words) model is used to search for the most similar image in the training dataset. PnP (Perspective n Points) and RANSAC (Random Sample Consensus) are employed to estimate an initial pose. Then the number of inliers is utilized as a criterion whether the feature-based method or the CNN-based method is to be leveraged. Compared with a previous CNN-based method, PoseNet, the average position error is reduced by 45.6% and the average orientation error is reduced by 67.4% on Microsoft's 7-Scenes datasets, which verifies the effectiveness of the proposed algorithm. Li Wang 0092, Ruifeng Li 0001, Seah Hock Soon, Chee Kwang Quah, Lijun Zhao 0003 |
ICARCV | 6 |
| 2018 | Robot teaching by teleoperation based on visual interaction and extreme learning machine
Chenguang Yang 0001, Junpei Zhong, Ning Wang 0009, Lijun Zhao 0003 |
Neurocomputing | 5 |