EDBT 2026 Demo / reviewers in the wild / expert
Ruifeng Li 0001
dblp:38/6776-1
· DBLP profile ↗
51ranked-venue papers
0as first author
45since 2021 · last 2026
0000-0002-1383-7745ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFMLLM: Enhance multi-modal large language model for global and fine-grained visual spatial perception
Zhendong Fan, Jinyang Gao, Hongbo Gao 0008, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 7 |
| 2026 | MTNet: A Mixed Transformer Network for High-Quality 3-D Object Detectionabstract3D object detection from point clouds represents a formidable challenge, necessitating the accurate identification and localization of objects within a 3D space. Recent advancements have showcased the efficacy of point-based detectors, leveraging local aggregators to encode intricate structural details of the point cloud. However, a notable limitation resides in their treatment of each point and object proposal in isolation, devoid of considering the interrelationships among them, thus impeding the overall detection performance. In this work, we argue that the integration of contextual information is paramount, particularly in the realm of indoor 3D object detection. Indoor environments are inherently characterized by robust contextual constraints, providing a rich tapestry for enhanced scene comprehension. In this way, we introduce MTNet, a mixed transformer network for high-quality indoor 3D object detection. Technically, we develop a mixed transformer (MixFormer) block that is purpose-built to intricately model the synergistic interplay between local structural information and global contextual features of 3D point clouds. In contrast to the classical transformer, our proposed MixFormer incorporates a local feature aggregator engineered to capture local geometric information while leveraging a KNN(K-Nearest Neighbors)-based attention mechanism to aggregate global contextual information. Furthermore, we suggest a feature compensator to adaptively fuse the strengths of local and global features, further bolstering detection performance. The culmination of these proposed components results in our MTNet framework, a hierarchical, versatile pipeline that consistently outperforms existing works across a multitude of benchmarks. In addition, we affirm the potential of proposed MixFormer and FC as generic modules that are capable of augmenting performance across a spectrum of 3D downstream point cloud tasks. Ruqi Liu, Shuaiyan Liu, Linqi Yang, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
IEEE Internet Things J. | 8 |
| 2026 | OFVL-MS++: Once for visual localization across multiple scenes via a two-stage framework
Chunsheng Yang, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 7 |
| 2026 | An MoE-Driven Unified Image Restoration Framework for Adverse Weather ConditionsabstractAdverse natural weather conditions frequently cause substantial performance degradation in outdoor vision systems, underscoring the critical importance of research on image restoration techniques. Employing a unified set of network parameters to restore degraded images across diverse weather conditions has emerged as a key research direction in the field of image restoration. In this work, we propose MUIRF, a Mixture-of-Experts (MoE)-driven unified image restoration framework for multiple adverse weather conditions. Specifically, our technical contribution includes a novel channel-level parameter sharing strategy guided by a shallow-feature-based MoE (CPSM). This fine-grained parameter sharing strategy adaptively selects convolution weight channels for cross-task sharing based on the input image, enabling the network to accurately capture weather-general features, while the remaining channels encode weather-specific features corresponding to each weather condition. CPSM facilitates precise channel selection, thereby enhancing the robustness and accuracy of MUIRF during joint training across diverse image restoration tasks under varying weather conditions. Additionally, gradient conflicts inevitably arise in shared parameters due to the divergent optimization objectives across tasks. To address this challenge, we propose a meta-vector-guided gradient homogenization (MVGH) algorithm that mitigates inter-task gradient conflicts and improves image restoration quality. Comprehensive experimental evaluations demonstrate that our proposed network outperforms most state-of-the-art approaches, validating its superior performance and effectiveness. Hongbo Gao 0008, Ruqi Liu, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | COME: A Collaborative Optimization Framework With Low-Rank MoE for Indoor 3D Object DetectionabstractIndoor 3D object detection serves as a fundamental task in computer vision and robotics. Existing research predominantly focuses on training domain-specific optimal models for individual datasets, yet it overlooks the potential value of capturing universal geometric attributes that can substantially enhance object detection performance across diverse domains. To resolve this gap, we propose COME, a novel and effective collaborative optimization framework designed to seamlessly integrate these universal attributes while preserving the domain-specific characteristics of each dataset domain. COME is built on VoteNet and incorporates a Cross-Domain Expert Parameter Sharing Strategy (CEPSS) that draws inspiration from the Mixture of Experts (MoE) framework. Its core innovation resides in the dual-expert design of CEPSS: domain-shared experts capture universal geometric relationships across datasets, whereas domain-specific experts encode unique features for individual datasets. This separation enables the model to focus on learning both generic and domain-specialized visual cues, without mutual interference. In addition, to dynamically adapt to different domains, we design a lightweight gating network that automatically selects relevant experts, eliminating irrelevant feature interference and enhancing model specialization. Compared to standard parameter-sharing architectures, this design significantly reduces gradient conflicts during multi-domain training. We further optimize computational efficiency by implementing low-rank structures for domain-shared and domain-specific experts, thus striking a better balance between memory overhead and detection performance. Experiments show that COME achieves state-of-the-art results across benchmarks, with acceptable parameter growth, and outperforms existing multi-domain detection methods. Hongbo Gao 0008, Zimeng Tong, Fuyuan Qiu, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 5 |
| 2025 | R-SIEL: a physics-informed learning algorithm for discovering dynamics of serial manipulators
Mohamed Omar, Ruifeng Li 0001, Ke Wang 0028, Ahmed Asker |
Expert Syst. Appl. | 2 |
| 2025 | DVDS: A deep visual dynamic slam system
Tao Xie 0010, Qihao Sun, Tao Sun 0024, Jinhang Zhang, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
Expert Syst. Appl. | 8 |
| 2025 | Poly-TF: A polymeric transformer framework for multiple visual tasks at once
Xuan Fan, Tao Sun 0024, Jinghan Gao, Tao Xie 0010, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 8 |
| 2025 | MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation
Zimeng Tong, Youran Du, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 9 |
| 2025 | CTFS: A consolidated transformer framework for instance and semantic segmentation tasks
Fuyuan Qiu, Hongbo Gao 0008, Tao Xie 0010, Chuqing Cao, Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028 |
Neural Networks | 6 |
| 2025 | AAPMatcher: Adaptive attention pruning matcher for accurate local feature matching
Xuan Fan, Shuaiyan Liu, Lijun Zhao 0003, Ruifeng Li 0001 |
Neural Networks | 5 |
| 2025 | NL-WCS: A Novel Data-Driven Algorithm for Extracting the Dynamics of Serial Robots Considering Non-Linear FrictionabstractRecently, developed data-driven SINDy-based techniques can identify the dynamic model of serial robots without simplifying assumptions nor pre-knowledge of all kinematics and geometric details. However, these techniques cannot handle non-linear friction models, which significantly affects the precision of the dynamic model identification. This study proposes a novel data-driven approach for dynamic model identification considering the non-linear friction model along with SINDy concept. This approach is termed as non-linear-weighted-constrained SINDy (NL-WCS). The SINDy concept is extended to accommodate any non-linear friction model by efficiently incorporating the Levenberg-Marquardt (LM) algorithm. Weighted L1 regularization is combined with the physics and the data constraints, to promote sparsity in the recovery of the dynamic equations. Moreover, this combination makes the approach robust against the regression matrix’s ill-conditionality and noise. LM is integrated with the robot’s SIMULINK model to get initial values for the non-linear friction empirical parameters. NL-WCS is experimentally evaluated by utilizing three distinct trajectories to validate the extracted dynamic model of 6-DOF UR10 and 7-DOF KUKA robots. In addition, five non-linear friction models are compared. NL-WCS outperforms all the previous SINDy-data-driven methods since it reduced the RMSE significantly up to 60.65%. NL-WCS also demonstrated robustness when tested against different levels of noise. Note to Practitioners—High-performance serial robot model-based controllers need precise dynamic model identification. The classical analytical methods for modeling serial robots rely on deriving the dynamic equations with simplified assumptions and then identifying the inertial parameters. Specifically, they assume that all the robot’s geometric details are known. Furthermore, there are uncertainties in geometric parameter values due to manufacturing/assembly errors. Data-driven-based SINDy methods can derive the dynamic model without pre-knowledge of the geometric parameters of the robot. However, it can’t incorporate non-linear friction models. Thus, this paper proposed a novel data-driven technique to extract the dynamic model of any serial manipulator by coupling SINDy approach with the nonlinear friction models which is the realistic case. The practitioners can benefit from a technique like that, in such a way of applying NL-WCS to different types of industrial robots, particularly, those that have missing manufacturer data sets or sheets. Simple and reliable dynamic models can be derived only by running the robot with various trajectories. Also, this helps eliminate any uncertainties in the dynamic model discovery/building process. Furthermore, Incorporating the non-linear friction model in the derivation process enhances the dynamic modeling accuracy as it effectively takes into account all friction characteristics. The proposed approach can be converted into a software package that the practitioners can use to derive the dynamics without getting their heads around the overwhelming robot’s kinematic details. Mohamed Omar, Ke Wang 0028, Ruifeng Li 0001, Ossama B. Abouelatta, Tao Xie 0010, Mohamed Gouda Alkalla |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | VD-Matcher: A Very Deep Local Feature Matcher With Weight Recycling and Keypoint DetectionabstractEstablishing local feature matches between image pairs serves as a fundamental component of plentiful vision tasks, such as visual localization and Structure from Motion (SfM). Recently, detector-free techniques equipped with the transformer have exhibited exceptional performance. Theoretically, optimizing the transformer architecture and stacking more transformer blocks could emphasize crucial features and filter out extraneous information by progressively narrowing the effective perception regions of the network within images, thereby enhancing the matching performance. Nevertheless, this paradigm results in a linear escalation of model size with respect to the number of blocks. In this study, we introduce VD-Matcher to address this issue. A principal innovation of VD-Matcher is the utilization of a weight recycling technique (WRT) that enables partial weights to be reutilized across successive transformer blocks, along with specific transformations designed to sufficiently enhance feature representations. This approach enables VD-Matcher to construct a deep transformer architecture for accurate local feature matching while maintaining a manageable parameter size. Furthermore, we propose a lightweight multi-scale keypoint detection module that captures representative keypoints to replace all keypoints for compact global information aggregation intra-/inter- images, which reduces the computational overhead induced by excessively deep transformer layers while alleviating redundant information propagation to a certain extent. Extensive experiments verify that VD-Matcher exceeds state-of-the-art algorithms on multiple benchmarks while maintaining less parameters. The source code is available at https://github.com/mooncake199809/VD-Matcher. Qihao Sun, Tao Xie 0010, Hongbo Gao 0008, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | AGHL: Anchor-Guided Point Cloud Registration Network With Hybrid Local Feature PerceptionabstractPoint cloud registration, which estimates a rigid transformation matrix between two point clouds, is a fundamental process in numerous applications. While existing detector-free techniques present exceptional performance, they overlook the extraction of hybrid local features that capture correlations between points and their neighbours, thereby limiting the quality of point cloud recognition. Moreover, these approaches typically treat point clouds as sequential data and employ the transformer to integrate global context from all points, which inevitably introduces interference from irrelevant regions, hence affecting the registration accuracy. In this work, we propose a novel detector-free approach AGHL to address these challenges. For the first issue, AGHL introduces a hybrid local feature perception module that designs two parallel branches to concurrently extract low-level and high-level local features, which effectively encode the correlations between each point and its neighborhood points in both Euclidean space and high-dimensional feature space. For the second issue, AGHL develops an anchor-guided cross attention that adheres to the local geometric consistency to constrain the network's attention on reliable anchors, thereby effectively suppressing interference from irrelevant regions. Benefiting from these techniques, AGHL achieves impressive point cloud registration accuracy across all synthetic, indoor, and outdoor datasets. Furthermore, we build an experimental platform and conduct a real-world robot localization experiment, with results showing the strong generalization ability of AGHL. Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Chuqing Cao |
IEEE Trans. Image Process. | 5 |
| 2025 | SOFW: A Synergistic Optimization Framework for Indoor 3D Object DetectionabstractIn this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar |
IEEE Trans. Multim. | 7 |
| 2025 | Centra-Net: A Centralized Network for Visual Localization Spanning Multiple ScenesabstractWe present Centra-Net, a centralized network that concurrently optimizes visual localization over numerous scenes under heterogeneous dataset domains. Centra-Net exemplifies storage efficiency by amalgamating multiple models with task-shared parameters into a singular cohesive structure. Technically, we develop abasic feature extraction unit (BFEU)with two parallel branches: one dedicated to local feature extraction and the other adept at adaptively generating a task-specific attention mask for feature calibration, thus bolstering its feature extraction capability across diverse scenes. Based on the BFEU, we introduce afilter-wise sharing mechanism (FSM)that adaptively determines parameter sharing within the unit, thus facilitating fine-grained parameter allocation. The key insight of FSM resides in reconceptualizing the parameter sharing of the unit as a learnable paradigm, enabling the determination of shared parameters to be made post-training. Finally, we suggest acomplexity-prioritized gradient algorithm (CPGA)that capitalizes on task complexity to attain a harmonious learning space for various tasks, thus safeguarding optimal performances across all tasks. Through rigorous experiments on numerous benchmarks, Centra-Net demonstrates a notable edge over existing state-of-the-art works while operating with a significantly reduced parameter footprint. Ke Wang 0028, Tao Xie 0010, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Multim. | 6 |
| 2025 | COFP: A Collaborative Optimization Framework With Polyhedral Feature Extraction for Multi-Weather Image RestorationabstractImage restoration in adverse weather conditions is a critical research focus in computer vision and autonomous driving. In this work, we introduce COFP, a collaborative optimization framework designed to simultaneously enhance the performance of image de-raining, de-snowing, and de-hazing tasks across diverse datasets. The core of COFP lies in its adaptive optimization of weather-shared and weather-specific parameters, enabling the extraction of polyhedral features that effectively integrate both weather-shared and weather-specific attributes, thus substantially boosting multi-weather performance of the network. Technically, we design a polyhedral feature extraction module (PFEM) to facilitate the acquisition of weather-shared attribute and weather-specific attribute. In PFEM, we first introduce an element-adaptive sharing strategy (ESS) that dynamically activates either weather-shared or weather-specific parameters for each element of the weight matrix based on learnable scores, thereby adaptively determining which parameters are shared or not. Secondly, we develop a feature extraction enhancement strategy (FEES), which extends two pathways in PFEM comprised of standard convolutional layers to enhance the capture of polyhedral features, further promoting the model performance. Furthermore, we propose a gradient balancing algorithm (GBA) that mitigates the unequal competition among tasks for shared parameters during network optimization by adaptively adjusting the direction of task gradients, effectively addressing the negative transfer issue induced by domain variations in multi-task learning. Experimental results demonstrate that COFP delivers state-of-the-art performances across various adverse weather image restoration benchmarks. Tao Xie 0010, Ruifeng Li 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | HVLF: A Holistic Visual Localization Framework Across Diverse ScenesabstractRecently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods. Fuyuan Qiu, Dedong Liu, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | ViT-MVT: A Unified Vision Transformer Network for Multiple Vision TasksabstractIn this work, we seek to learn multiple mainstream vision tasks concurrently using a unified network, which is storage-efficient as numerous networks with task-shared parameters can be implanted into a single consolidated network. Our framework, vision transformer (ViT)-MVT, built on a plain and nonhierarchical ViT, incorporates numerous visual tasks into a modest supernet and optimizes them jointly across various dataset domains. For the design of ViT-MVT, we augment the ViT with a multihead self-attention (MHSE) to offer complementary cues in the channel and spatial dimension, as well as a local perception unit (LPU) and locality feed-forward network (locality FFN) for information exchange in the local region, thus endowing ViT-MVT with the ability to effectively optimize multiple tasks. Besides, we construct a search space comprising potential architectures with a broad spectrum of model sizes to offer various optimum candidates for diverse tasks. After that, we design a layer-adaptive sharing technique that automatically determines whether each layer of the transformer block is shared or not for all tasks, enabling ViT-MVT to obtain task-shared parameters for a reduction of storage and task-specific parameters to learn task-related features such that boosting performance. Finally, we introduce a joint-task evolutionary search algorithm to discover an optimal backbone for all tasks under total model size constraint, which challenges the conventional wisdom that visual tasks are typically supplied with backbone networks developed for image classification. Extensive experiments reveal that ViT-MVT delivers exceptional performances for multiple visual tasks over state-of-the-art methods while necessitating considerably fewer total storage costs. We further demonstrate that once ViT-MVT has been trained, ViT-MVT is capable of incremental learning when generalized to new tasks while retaining identical performances for trained tasks. The code is available at https://github.com/XT-1997/vitmvt. Tao Xie 0010, Ruifeng Li 0001, Shouren Mao, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | RSG-Search Plus: An Advanced Traffic Scene Retrieval Methods based on Road Scene GraphabstractCurrently, with the rapid growth of training datasets for autonomous driving systems, we are faced with a challenge: how to efficiently retrieve specific traffic scenes from massive amount of scene in multiple datasets. This challenge primarily stems from the heterogeneity of existing datasets, meaning these datasets contain different types of data, follow different data formats, and use different sensors for data collection. To address this issue, we present RSG-Search Plus, a universal traffic scene searching method based on Road Scene Graph and Large Language Models (LLMs). Our approach first transform datasets into scene graphs to exclude irrelevant details, then efficiently retrieving specific configurations among thousands of traffic scenes by matching isomorphic sub-graphs between input graph and road scene graph. Experimental results demonstrate that our graph searching method can accurately match the scenes described by input condition. Additionally, this method is easily adaptable to different datasets, significantly simplifying the scene search process. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Simon Thompson 0002, Kazuya Takeda |
IV | 3 |
| 2024 | FMAP: Learning robust and accurate local feature matching with anchor points
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 5 |
| 2024 | ALNet: An adaptive channel attention network with local discrepancy perception for accurate indoor visual localization
Hongbo Gao 0008, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Mengyuan Wu |
Expert Syst. Appl. | 4 |
| 2024 | APM: Adaptive parameter multiplexing for class incremental learning
Jinghan Gao, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
Expert Syst. Appl. | 3 |
| 2024 | DeepMatcher: A deep transformer-based network for robust and accurate local feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 4 |
| 2024 | ActionMixer: Temporal action detection with Optimal Action Segment Assignment and mixers
Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
Expert Syst. Appl. | 5 |
| 2024 | M-DIVO: Multiple ToF RGB-D Cameras-Enhanced Depth-Inertial-Visual OdometryabstractTime-of-Flight (ToF) RGB-D cameras provide a wealth of information for SLAM systems. However, the limited field of view (FOV) of a single ToF RGB-D camera and the small range of its depth measurement module make it prone to degeneracy when relying solely on visual or depth information for SLAM, a problem typical of unimodal SLAM algorithms. To address this issue, this article presents M-DIVO: an IEKF-based odometry that fuses visual, depth (similar to LiDAR), and inertial modules from multiple ToF RGB-D cameras. It comprises two direct method subsystems: 1) the depth–inertial odometry (DIO) subsystem, which constructs point-to-plane constraints from multiple depth modules and 2) the visual–inertial odometry (VIO) subsystem, which optimizes pose using photometric error constructed by multiple cameras. Additionally, to manage the significant computational load from processing multiple sensors and multimodal information, we introduce a multimodal redundancy scheduling mechanism (MRSM): prioritizing the DIO subsystem with the VIO subsystem as auxiliary, executing the VIO subsystem only when degeneracy occurs in the DIO subsystem. We also propose a “External First, Internal Last” strategy for calibrating multiple external and internal sensors. Experiments demonstrate that compared to unimodal SLAM, our method achieves higher robustness and precision, as well as satisfactory real-time performance. The proposed calibration strategy is demonstrated to be more accurate than the traditional inertial measurement unit-centric approach. The code is open source. Jie Xu 0066, Wenlu Yu, Shenghai Yuan 0001, Lijun Zhao 0003, Ruifeng Li 0001, Lihua Xie 0001 |
IEEE Internet Things J. | 6 |
| 2024 | CO-Net++: A Cohesive Network for Multiple Point Cloud Tasks at Once With Two-Stage Feature RectificationabstractWe present CO-Net++, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains with a two-stage feature rectification strategy. The core of CO-Net++ lies in optimizing task-shared parameters to capture universal features across various tasks while discerning task-specific parameters tailored to encapsulate the unique characteristics of each task. Specifically, CO-Net++ develops a two-stage feature rectification strategy (TFRS) that distinctly separates the optimization processes for task-shared and task-specific parameters. At the first stage, TFRS configures all parameters in backbone as task-shared, which encourages CO-Net++ to thoroughly assimilate universal attributes pertinent to all tasks. In addition, TFRS introduces a sign-based gradient surgery to facilitate the optimization of task-shared parameters, thus alleviating conflicting gradients induced by various dataset domains. In the second stage, TFRS freezes task-shared parameters and flexibly integrates task-specific parameters into the network for encoding specific characteristics of each dataset domain. CO-Net++ prominently mitigates conflicting optimization caused by parameter entanglement, ensuring the sufficient identification of universal and specific features. Extensive experiments reveal that CO-Net++ realizes exceptional performances on both 3D object detection and 3D semantic segmentation tasks. Moreover, CO-Net++ delivers an impressive incremental learning capability and prevents catastrophic amnesia when generalizing to new point cloud tasks. Tao Xie 0010, Qihao Sun, Chuqing Cao, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | OAMatcher: An overlapping areas-based network with label credibility for robust and accurate feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 5 |
| 2024 | A Convex Optimization Method to Time-Optimal Trajectory Planning With Jerk Constraint for Industrial Robotic ManipulatorsabstractThis paper proposes a convex time-optimal trajectory planning method for industrial robotic manipulators with jerk constraints. To achieve smooth and efficient trajectories, the square of the pseudo velocity profile is constructed using a cubic uniform B-spline, and a linear relationship is defined with the control points of the B-spline to preserve convexity in the pseudo states. Bi-linear and non-convex jerk constraints are introduced in the optimization problem, and a convex restriction method is applied to achieve convexity. The proposed method is evaluated through three case studies: two contour following tasks and a pick-and-place task. Comparative optimization results demonstrate that the proposed method achieves time optimality and trajectory smoothness simultaneously in the reformulated and jerk-restricted optimization problem. The proposed method provides a practical approach to address the non-convexity of jerk constraints in trajectory optimization for industrial robotic manipulators.Note to Practitioners—This work was motivated by the challenge of generating smooth and time-optimal trajectories for industrial manipulators while considering jerk and dynamic constraints. Existing approaches typically use multi-objective methods that treat jerk constraints as soft constraints. In contrast, this work proposes a B-spline-based convex optimization method that treats jerk constraints as hard constraints. Three case studies are conducted, involving butterfly-type, door-type and ‘OPTEC’-type paths, with the UR5 cooperative robot. The numerical and experimental results demonstrate that the proposed method sacrifices some time optimality to achieve trajectory smoothness and computational efficiency due to the convex restriction reformulation of the jerk constraints. Moreover, the proposed method is applicable to general industrial manipulators. Supplementary video is available at https://youtu.be/gVz5IuPKD-o. Guanggui Cheng, Minxiu Kong, Ruifeng Li 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object DetectionabstractIn this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance. Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082 |
IEEE Trans. Multim. | 4 |
| 2023 | Poly-PC: A Polyhedral Network for Multiple Point Cloud Tasks at OnceabstractIn this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by multiple dataset domains, etc.) of multi-task learning on point cloud. Specifically, we propose a residual set abstraction (Res-SA) layer for efficient and effective scaling in both width and depth of the network, hence accommodating the needs of various tasks. We develop a weight-entanglement- based one-shot NAS technique to find optimal architectures for all tasks. Moreover, such technique entangles the weights of multiple tasks in each layer to offer task-shared parameters for efficient storage deployment while providing ancillary task-specific parameters for learning task-related features. Finally, to facilitate the training of Poly-PC, we introduce a task-prioritization-based gradient balance algorithm that leverages task prioritization to reconcile conflicting gradients, ensuring high performance for all tasks. Benefiting from the suggested techniques, models optimized by Poly-PC collectively for all tasks keep fewer total FLOPs and parameters and outperform previous methods. We also demonstrate that Poly-PC allows incremental learning and evades catastrophic forgetting when tuned to a new task. Tao Xie 0010, Shiguang Wang, Ke Wang 0028, Linqi Yang, Xingcheng Zhang, Ruifeng Li 0001, Jian Cheng 0003 |
CVPR | 8 |
| 2023 | OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesabstractIn this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net. Tao Xie 0010, Siyi Lu, Ke Wang 0028, Jinghan Gao, Dedong Liu, Jie Xu 0066, Lijun Zhao 0003, Ruifeng Li 0001 |
ICCV | 10 |
| 2023 | CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive NetworkabstractWe present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task. Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001 |
ICCV | 11 |
| 2023 | A Real-Time Hardware-Accelerator-Aided MJ-EKF SLAM Algorithm for Large-Scale Map Based on Heterogeneous Multi-Core SoCabstractAs a classic SLAM algorithm, LIDAR-based EKF SLAM still remains a problem. An open issue is high computational complexity, which leading a very limited map size to sufficient real-time requirements. To address this problem, in this paper, we propose an MJ-EKF SLAM system that is fully implemented on a Heterogeneous Multi-Core SoC(HMS). To limit the EKF SLAM computation complexity, the Map Joining algorithm participates in the framework. The most computation cost step in EKF SLAM and Map Joining algorithms is implemented on an MJ-EKF dedicated accelerator. Meanwhile, several hardware optimization methods are proposed to save logic resources, avoid redundant computation and reduce data communication between on-chip memory and off-chip main memory. Field experiments demonstrate the HMS-based MJ-EKF algorithm achieved high real-time performance (over 30Hz) in a large-scale map with 500 landmarks. Zhendong Fan, Ke Wang 0028, Minjie Bao, Ruifeng Li 0001 |
IECON | 4 |
| 2023 | GCA-Net: A Global Context Aggregation Network for Effective Optical FlowabstractOptical flow seeks to estimate the per-pixel 2D motion between two frames by identifying corresponding pixels. Current flow estimators typically involve per-pixel feature extraction, multi-scale 4D correlation volume construction, and iterative flow field updates through a Conv-GRU module. However, the locality of convolutional features in these methods renders the calculated correlations vulnerable to different noises. In addition, the Conv-GRU module of these methods is only executed with a convolution layer, which is incapable of exploiting context clues from larger window sizes even inside the image being queried itself, thus raising it more difficult for the model to process images with challenging regions, e.g., textureless areas. To this end, we introduce GCA-Net, a global context aggregation network for credible yet effective optical flow estimation. More specifically, we propose a highly efficient multi-scale transformer (MSFormer) layer which enables the per-pixel feature to aggregate long-range information from other features, hence building more accurate 4D correlation volumes. Besides, we develop an attention-enhanced Conv-GRU block (AttGRU) that can incorporate information alongside a larger context window even within itself empowering our network to estimate optical flow in even the most challenging regions. Experimentally, we demonstrate that GCA-Net outperforms previous state-of-the-art methods by large margins on Sintel (Final) and KITTI 2015 (background and foreground) benchmarks. Tao Xie 0010, Jinghan Gao, Ke Wang 0028, Ruifeng Li 0001 |
IECON | 4 |
| 2023 | RSG-Search: Semantic Traffic Scene Retrieval Using Graph-Based Scene RepresentationabstractBrowsing specific traffic scene in large-scale dataset is an increasing demand from researchers, self-driving community and insurance companies. It is easy to search scenes with specific tags such as "rain", "snow", or "on highway". However, searching specific scene configurations, like "two vehicles waiting for a person crossing the road", is still an open problem. In this paper, we provide RSG-search, a scene-graph based traffic scene retrieval method, based on our previous research on traffic scene-graph generation. By previously translating open datasets to scene graphs, we can ignore irrelevant details, and efficiently search specific scene configuration among thousands of traffic scenes. Experiment results shows that our graph searching method is able to retrieve results for a given query with high accuracy. Our method simplifies the task of scene retrieval, opening opportunities for new applications. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 3 |
| 2023 | Actions as points: a simple and efficient detector for skeleton-based temporal action detection
Ke Wang 0028, Ruifeng Li 0001 |
Mach. Vis. Appl. | 3 |
| 2023 | Cascading spatio-temporal attention network for real-time action detection
Ke Wang 0028, Ruifeng Li 0001, Petra Perner |
Mach. Vis. Appl. | 3 |
| 2023 | Point-NAS: A Novel Neural Architecture Search Framework for Point Cloud AnalysisabstractRecently, point-based networks have exhibited extraordinary potential for 3D point cloud processing. However, owing to the meticulous design of both parameters and hyperparameters inside the network, constructing a promising network for each point cloud task can be an expensive endeavor. In this work, we develop a novel one-shot search framework called Point-NAS to automatically determine optimum architectures for various point cloud tasks. Specifically, we design an elastic feature extraction (EFE) module that serves as a basic unit for architecture search, which expands seamlessly alongside both the width and depth of the network for efficient feature extraction. Based on the EFE module, we devise a searching space, which is encoded into a supernet to provide a wide number of latent network structures for a particular point cloud task. To fully optimize the weights of the supernet, we propose a weight coupling sandwich rule that samples the largest, smallest, and multiple medium models at each iteration and fuses their gradients to update the supernet. Furthermore, we present a united gradient adjustment algorithm that mitigates gradient conflict induced by distinct gradient directions of sampled models and supernet, thus expediting the convergence of the supernet and assuring that it can be comprehensively trained. Pursuant to the provided techniques, the trained supernet enables a multitude of subnets to be incredibly well-optimized. Finally, we conduct an evolutionary search for the supernet under resource constraints to find promising architectures for different tasks. Experimentally, the searched Point-NAS with weights inherited from the supernet realizes outstanding results across a variety of benchmarks. i.e., 94.2% and 88.9% overall accuracy under ModelNet40 and ScanObjectNN, 68.6% mIoU under S3DIS, 63.6% and 69.3% [email protected] under SUN RGB-D and ScanNet V2 datasets. Tao Xie 0010, Linqi Yang, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 6 |
| 2023 | Real-Time Point Cloud Object Detection via Voxel-Point Geometry AbstractionabstractRecent advances in 3D object detection typically learn voxel-based or point-based representations on point clouds. Point-based methods preserve precise point positions but incur high computational load, whereas voxel-based methods rasterize unordered points into voxel grids efficiently but give rise to an accuracy bottleneck. To take advantage of voxel- and point-based representations, we develop an effective and efficient 3D object detector via a novel voxel-point geometry abstraction scheme. Our motivation is to use coarse voxel representation to accelerate proposal generation while using precise point representation to facilitate proposal refinement. For voxel representation learning, we propose a context enrichment module with a novel 3D sparse interpolation layer to augment raw points with multi-scale context. We further develop a point-based RoI pooling module with explicit position augmentation for proposal refinement. Extensive experiments on the widely used KITTI Dataset and the latest Waymo Open Dataset show that the proposed algorithm outperforms state-of-the-art point-voxel-based methods while running at 24 FPS on the TITAN XP GPU. Guangsheng Shi, Ke Wang 0028, Ruifeng Li 0001, Chao Ma 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Real-to-Synthetic: Generating Simulator Friendly Traffic Scenes from Graph RepresentationabstractReproducing real-world traffic scenes in the simulator is fundamental to training self-driving systems. Creating a simulation scenario is a complex task, generally done manually: the ego-vehicle and other entities are placed and their trajectories defined, trying to recreate some situation found in real traffic. To reduce the manual burden, here we propose the Real-to-Synthetic toolset. This toolset provides synthetic traffic scene in openDrive format, which can be directly simulated in many simulators such as SUMO or CARLA. Also, we provide a scene generator which generates near-realistic scene from minimum user effort. To maintain the similarity between real-world scene and generated one, here we introduce the concept “Road Scene Graph”(RSG). In this graph, nodes represent entities while edges stand for pairwise relationships. These relationships could be maintained in the scene generation process while the actor is generated according to the distribution sampled from real-world data. Experiments proved that by using “Road Scene Graph”, our scene generator proposes a much more convenient way to conFigure traffic scenes rather than manually defining every actor’s initial status and trajectories. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 3 |
| 2022 | A novel fast combine-and-conquer object detector based on only one-level feature map
Ke Wang 0028, Ruifeng Li 0001, Zhonghao Qin, Petra Perner |
Comput. Vis. Image Underst. | 3 |
| 2022 | Welding splash and arc noise reduction imaging model based on computationally efficient pairwise response serving welding process library
Zhonghao Qin, Ke Wang 0028, Ruifeng Li 0001 |
Mach. Vis. Appl. | 3 |
| 2021 | RSG-Net: Towards Rich Sematic Relationship Prediction for Intelligent Vehicle in Complex EnvironmentsabstractBehavioral and semantic relationships play a vital role on intelligent self-driving vehicles and ADAS systems. Different from other research focused on trajectory, position, and bounding boxes, relationship data provides a human understandable description of the object's behavior, and it could describe an object's past and future status in an amazingly brief way. Therefore it is a fundamental method for tasks such as risk detection, environment understanding, and decision making. In this paper, we propose RSG-Net (Road Scene Graph Net): a graph convolutional network designed to predict potential semantic relationships from object proposals, and to produce a graph-structured result, called “Road Scene Graph”. The experimental results indicate that this network, trained on Road Scene Graph dataset, could efficiently predict potential semantic relationships among objects around the ego-vehicle. Yafu Tian, Alexander Carballo, Ruifeng Li 0001, Kazuya Takeda |
IV | 3 |
| 2021 | Multi-obstacle path planning and optimization for mobile robot
Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028, Xichun Gui |
Expert Syst. Appl. | 2 |
| 2018 | Feature-Based and Convolutional Neural Network Fusion Method for Visual RelocalizationabstractRelocalization is one of the necessary modules for mobile robots in long-term autonomous movement in an environment. Currently, visual relocalization algorithms mainly include feature-based methods and CNN-based (Convolutional Neural Network) methods. Feature-based methods can achieve high localization accuracy in feature-rich scenes, but the error is quite large or it even fails in cases with motion blur, texture-less scene and changing view angle. CNN-based methods usually have better robustness but poor localization accuracy. For this reason, a visual relocalization algorithm that combines the advantages of the two methods is proposed in this paper. The BoVW (Bag of Visual Words) model is used to search for the most similar image in the training dataset. PnP (Perspective n Points) and RANSAC (Random Sample Consensus) are employed to estimate an initial pose. Then the number of inliers is utilized as a criterion whether the feature-based method or the CNN-based method is to be leveraged. Compared with a previous CNN-based method, PoseNet, the average position error is reduced by 45.6% and the average orientation error is reduced by 67.4% on Microsoft's 7-Scenes datasets, which verifies the effectiveness of the proposed algorithm. Li Wang 0092, Ruifeng Li 0001, Seah Hock Soon, Chee Kwang Quah, Lijun Zhao 0003 |
ICARCV | 2 |
| 2018 | Interface Design of a Physical Human-Robot Interaction System for Human Impedance Adaptive Skill TransferabstractIt has been established that the transfer of human adaptive impedance is of great significance for physical human-robot interaction (pHRI). By processing the electromyography (EMG) signals collected from human muscles, the limb impedance could be extracted and transferred to robots. The existing impedance transfer interfaces rely only on visual feedback and, thus, may be insufficient for skill transfer in a sophisticated environment. In this paper, physical haptic feedback mechanism is introduced to result in muscle activity that would generate EMG signals in a natural manner, in order to achieve intuitive human impedance transfer through a designed coupling interface. Relevant processing methods are integrated into the system, including the spectral collaborative representation-based classifications method used for hand motion recognition; fast smooth envelop and dimensionality reduction algorithm for arm endpoint stiffness estimation. The tutor's arm endpoint motion trajectory is directly transferred to the robot by the designed coupling module without the restriction of hands. Haptic feedback is provided to the human tutor according to skill learning performance to enhance the teaching experience. The interface has been experimentally tested by a plugging-in task and a cutting task. Compared with the existing interfaces, the developed one has shown a better performance. Chenguang Yang 0001, Chao Zeng 0002, Peidong Liang, Zhijun Li 0001, Ruifeng Li 0001, Chun-Yi Su |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2017 | Power difference template for action recognition
Ruifeng Li 0001, Yajun Fang |
Mach. Vis. Appl. | 2 |
| 2017 | Three-stream CNNs for action recognition
Lianzheng Ge, Ruifeng Li 0001, Yajun Fang |
Pattern Recognit. Lett. | 3 |
| 2016 | Gradient-layer feature transform for action detection and recognition
Ruifeng Li 0001, Yajun Fang |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Unsupervised feature selection based on spectral regression from manifold learning for facial expression recognitionabstractIn this study, an unsupervised feature selection method is proposed for facial feature recognition (FER) in the absence of class labels. The contribution is the descriptive feature components selector spectral regression representative coefficient scores based on graph manifold learning from high‐dimensional feature space. The spectral regression analysis and L1‐regularised least square are then used to compute the importance of features in the original space, so that less representative features with lower coefficient scores will be removed without prior distribution assumption. To verify the performance of the authors’ method, some classifiers are used to classify facial expressions on three benchmark facial expression databases. The recognition results indicate the availability and effectiveness of the proposed method for FER. Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001 |
IET Comput. Vis. | 3 |