Ke Wang 0028

dblp:181/2613-28 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0002-5615-0847ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MTNet: A Mixed Transformer Network for High-Quality 3-D Object Detection
abstract
3D object detection from point clouds represents a formidable challenge, necessitating the accurate identification and localization of objects within a 3D space. Recent advancements have showcased the efficacy of point-based detectors, leveraging local aggregators to encode intricate structural details of the point cloud. However, a notable limitation resides in their treatment of each point and object proposal in isolation, devoid of considering the interrelationships among them, thus impeding the overall detection performance. In this work, we argue that the integration of contextual information is paramount, particularly in the realm of indoor 3D object detection. Indoor environments are inherently characterized by robust contextual constraints, providing a rich tapestry for enhanced scene comprehension. In this way, we introduce MTNet, a mixed transformer network for high-quality indoor 3D object detection. Technically, we develop a mixed transformer (MixFormer) block that is purpose-built to intricately model the synergistic interplay between local structural information and global contextual features of 3D point clouds. In contrast to the classical transformer, our proposed MixFormer incorporates a local feature aggregator engineered to capture local geometric information while leveraging a KNN(K-Nearest Neighbors)-based attention mechanism to aggregate global contextual information. Furthermore, we suggest a feature compensator to adaptively fuse the strengths of local and global features, further bolstering detection performance. The culmination of these proposed components results in our MTNet framework, a hierarchical, versatile pipeline that consistently outperforms existing works across a multitude of benchmarks. In addition, we affirm the potential of proposed MixFormer and FC as generic modules that are capable of augmenting performance across a spectrum of 3D downstream point cloud tasks.
Ruqi Liu, Shuaiyan Liu, Linqi Yang, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001
IEEE Internet Things J.6
2026 OFVL-MS++: Once for visual localization across multiple scenes via a two-stage framework
Chunsheng Yang, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
Pattern Recognit.6
2026 An MoE-Driven Unified Image Restoration Framework for Adverse Weather Conditions
abstract
Adverse natural weather conditions frequently cause substantial performance degradation in outdoor vision systems, underscoring the critical importance of research on image restoration techniques. Employing a unified set of network parameters to restore degraded images across diverse weather conditions has emerged as a key research direction in the field of image restoration. In this work, we propose MUIRF, a Mixture-of-Experts (MoE)-driven unified image restoration framework for multiple adverse weather conditions. Specifically, our technical contribution includes a novel channel-level parameter sharing strategy guided by a shallow-feature-based MoE (CPSM). This fine-grained parameter sharing strategy adaptively selects convolution weight channels for cross-task sharing based on the input image, enabling the network to accurately capture weather-general features, while the remaining channels encode weather-specific features corresponding to each weather condition. CPSM facilitates precise channel selection, thereby enhancing the robustness and accuracy of MUIRF during joint training across diverse image restoration tasks under varying weather conditions. Additionally, gradient conflicts inevitably arise in shared parameters due to the divergent optimization objectives across tasks. To address this challenge, we propose a meta-vector-guided gradient homogenization (MVGH) algorithm that mitigates inter-task gradient conflicts and improves image restoration quality. Comprehensive experimental evaluations demonstrate that our proposed network outperforms most state-of-the-art approaches, validating its superior performance and effectiveness.
Hongbo Gao 0008, Ruqi Liu, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003
IEEE Trans. Circuits Syst. Video Technol.7
2025 SplatPose: Geometry-Aware 6-DoF Pose Estimation from Single RGB Image via 3D Gaussian Splatting
abstract
6-DoF pose estimation is a fundamental task in computer vision with wide-ranging applications in augmented reality and robotics. Existing single RGB-based methods often compromise accuracy due to their reliance on initial pose estimates and susceptibility to rotational ambiguity, while approaches requiring depth sensors or multi-view setups incur significant deployment costs. To address these limitations, we introduce SplatPose, a novel framework that synergizes 3D Gaussian Splatting (3DGS) with a dual-branch neural architecture to achieve high-precision pose estimation using only a single RGB image. Central to our approach is the Dual-Attention Ray Scoring Network (DARS-Net), which innovatively decouples positional and angular alignment through geometry-domain attention mechanisms, explicitly modeling directional dependencies to mitigate rotational ambiguity. Additionally, a coarse-to-fine optimization pipeline progressively refines pose estimates by aligning dense 2D features between query images and 3DGS-synthesized views, effectively correcting feature misalignment and depth errors from sparse ray sampling. Experiments on three benchmark datasets demonstrate that SplatPose achieves state-of-the-art 6-DoF pose estimation accuracy in single RGB settings, rivaling approaches that depend on depth or multi-view images.
Linqi Yang, Xiongwei Zhao, Qihao Sun, Ke Wang 0028
IROS4
2025 R-SIEL: a physics-informed learning algorithm for discovering dynamics of serial manipulators
Mohamed Omar, Ruifeng Li 0001, Ke Wang 0028, Ahmed Asker
Expert Syst. Appl.3
2025 DVDS: A deep visual dynamic slam system
Tao Xie 0010, Qihao Sun, Tao Sun 0024, Jinhang Zhang, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001
Expert Syst. Appl.7
2025 MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation
Zimeng Tong, Youran Du, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001
Knowl. Based Syst.7
2025 CTFS: A consolidated transformer framework for instance and semantic segmentation tasks
Fuyuan Qiu, Hongbo Gao 0008, Tao Xie 0010, Chuqing Cao, Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028
Neural Networks8
2025 NL-WCS: A Novel Data-Driven Algorithm for Extracting the Dynamics of Serial Robots Considering Non-Linear Friction
abstract
Recently, developed data-driven SINDy-based techniques can identify the dynamic model of serial robots without simplifying assumptions nor pre-knowledge of all kinematics and geometric details. However, these techniques cannot handle non-linear friction models, which significantly affects the precision of the dynamic model identification. This study proposes a novel data-driven approach for dynamic model identification considering the non-linear friction model along with SINDy concept. This approach is termed as non-linear-weighted-constrained SINDy (NL-WCS). The SINDy concept is extended to accommodate any non-linear friction model by efficiently incorporating the Levenberg-Marquardt (LM) algorithm. Weighted L1 regularization is combined with the physics and the data constraints, to promote sparsity in the recovery of the dynamic equations. Moreover, this combination makes the approach robust against the regression matrix’s ill-conditionality and noise. LM is integrated with the robot’s SIMULINK model to get initial values for the non-linear friction empirical parameters. NL-WCS is experimentally evaluated by utilizing three distinct trajectories to validate the extracted dynamic model of 6-DOF UR10 and 7-DOF KUKA robots. In addition, five non-linear friction models are compared. NL-WCS outperforms all the previous SINDy-data-driven methods since it reduced the RMSE significantly up to 60.65%. NL-WCS also demonstrated robustness when tested against different levels of noise. Note to Practitioners—High-performance serial robot model-based controllers need precise dynamic model identification. The classical analytical methods for modeling serial robots rely on deriving the dynamic equations with simplified assumptions and then identifying the inertial parameters. Specifically, they assume that all the robot’s geometric details are known. Furthermore, there are uncertainties in geometric parameter values due to manufacturing/assembly errors. Data-driven-based SINDy methods can derive the dynamic model without pre-knowledge of the geometric parameters of the robot. However, it can’t incorporate non-linear friction models. Thus, this paper proposed a novel data-driven technique to extract the dynamic model of any serial manipulator by coupling SINDy approach with the nonlinear friction models which is the realistic case. The practitioners can benefit from a technique like that, in such a way of applying NL-WCS to different types of industrial robots, particularly, those that have missing manufacturer data sets or sheets. Simple and reliable dynamic models can be derived only by running the robot with various trajectories. Also, this helps eliminate any uncertainties in the dynamic model discovery/building process. Furthermore, Incorporating the non-linear friction model in the derivation process enhances the dynamic modeling accuracy as it effectively takes into account all friction characteristics. The proposed approach can be converted into a software package that the practitioners can use to derive the dynamics without getting their heads around the overwhelming robot’s kinematic details.
Mohamed Omar, Ke Wang 0028, Ruifeng Li 0001, Ossama B. Abouelatta, Tao Xie 0010, Mohamed Gouda Alkalla
IEEE Trans Autom. Sci. Eng.2
2025 An Energy-Efficient, High-Frame-Rate, and Reconfigurable EKF-SLAM Processor With Full Acceleration for Autonomous Mobile Robots
abstract
In many intelligent edge applications involving Autonomous Mobile Robots (AMRs), efficient and real-time localization and mapping is a fundamental issue. Extended Kalman Filter Simultaneous Localization and Mapping (EKFSLAM) algorithm is a classic and successful solution to realize localization and mapping, while it is computationally intensive and poses a challenge for real-time tasks in small and micro robots. To address this issue, this work proposes an energy-efficient, highframe-rate, and reconfigurable EKF-SLAM processor. Firstly, a heterogeneous dual-core architecture is proposed to enable full acceleration of both matrix operations and nonlinear calculations in EKF-SLAM at the hardware architecture level. Secondly, a Reconfigurable Matrix Accelerator (RMA) and Reconfigurable Nonlinear Accelerator (RNA) are proposed to maximize data reuse and support diverse nonlinear functions at the data flow level. Thirdly, a data property-aware strategy is proposed at the data property level, which exploits matrix symmetry, sparsity, and dependency to reduce storage significantly and eliminate redundant computations. FPGA validation results show that the proposed design can achieve a frame rate of 774 fps and an energy efficiency of 0.66 mJ/frame, when performing mapping processes involving 60 landmarks at 100 MHz.
Bingqiang Liu, Yequan Zhao, Minjie Bao, Zhendong Fan, Dingcheng Jiang, Zixuan Shen, Yulong Tan, Zaisheng He, Dengke Xu, Ke Wang 0028, Chao Wang 0096, Lining Sun
IEEE Trans. Circuits Syst. I Regul. Pap.11
2025 AGHL: Anchor-Guided Point Cloud Registration Network With Hybrid Local Feature Perception
abstract
Point cloud registration, which estimates a rigid transformation matrix between two point clouds, is a fundamental process in numerous applications. While existing detector-free techniques present exceptional performance, they overlook the extraction of hybrid local features that capture correlations between points and their neighbours, thereby limiting the quality of point cloud recognition. Moreover, these approaches typically treat point clouds as sequential data and employ the transformer to integrate global context from all points, which inevitably introduces interference from irrelevant regions, hence affecting the registration accuracy. In this work, we propose a novel detector-free approach AGHL to address these challenges. For the first issue, AGHL introduces a hybrid local feature perception module that designs two parallel branches to concurrently extract low-level and high-level local features, which effectively encode the correlations between each point and its neighborhood points in both Euclidean space and high-dimensional feature space. For the second issue, AGHL develops an anchor-guided cross attention that adheres to the local geometric consistency to constrain the network's attention on reliable anchors, thereby effectively suppressing interference from irrelevant regions. Benefiting from these techniques, AGHL achieves impressive point cloud registration accuracy across all synthetic, indoor, and outdoor datasets. Furthermore, we build an experimental platform and conduct a real-world robot localization experiment, with results showing the strong generalization ability of AGHL.
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Chuqing Cao
IEEE Trans. Image Process.4
2025 SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection
abstract
In this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW
Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar
IEEE Trans. Multim.4
2025 Centra-Net: A Centralized Network for Visual Localization Spanning Multiple Scenes
abstract
We present Centra-Net, a centralized network that concurrently optimizes visual localization over numerous scenes under heterogeneous dataset domains. Centra-Net exemplifies storage efficiency by amalgamating multiple models with task-shared parameters into a singular cohesive structure. Technically, we develop abasic feature extraction unit (BFEU)with two parallel branches: one dedicated to local feature extraction and the other adept at adaptively generating a task-specific attention mask for feature calibration, thus bolstering its feature extraction capability across diverse scenes. Based on the BFEU, we introduce afilter-wise sharing mechanism (FSM)that adaptively determines parameter sharing within the unit, thus facilitating fine-grained parameter allocation. The key insight of FSM resides in reconceptualizing the parameter sharing of the unit as a learnable paradigm, enabling the determination of shared parameters to be made post-training. Finally, we suggest acomplexity-prioritized gradient algorithm (CPGA)that capitalizes on task complexity to attain a harmonious learning space for various tasks, thus safeguarding optimal performances across all tasks. Through rigorous experiments on numerous benchmarks, Centra-Net demonstrates a notable edge over existing state-of-the-art works while operating with a significantly reduced parameter footprint.
Ke Wang 0028, Tao Xie 0010, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003
IEEE Trans. Multim.3
2025 HVLF: A Holistic Visual Localization Framework Across Diverse Scenes
abstract
Recently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods.
Fuyuan Qiu, Dedong Liu, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
IEEE Trans. Neural Networks Learn. Syst.6
2025 ViT-MVT: A Unified Vision Transformer Network for Multiple Vision Tasks
abstract
In this work, we seek to learn multiple mainstream vision tasks concurrently using a unified network, which is storage-efficient as numerous networks with task-shared parameters can be implanted into a single consolidated network. Our framework, vision transformer (ViT)-MVT, built on a plain and nonhierarchical ViT, incorporates numerous visual tasks into a modest supernet and optimizes them jointly across various dataset domains. For the design of ViT-MVT, we augment the ViT with a multihead self-attention (MHSE) to offer complementary cues in the channel and spatial dimension, as well as a local perception unit (LPU) and locality feed-forward network (locality FFN) for information exchange in the local region, thus endowing ViT-MVT with the ability to effectively optimize multiple tasks. Besides, we construct a search space comprising potential architectures with a broad spectrum of model sizes to offer various optimum candidates for diverse tasks. After that, we design a layer-adaptive sharing technique that automatically determines whether each layer of the transformer block is shared or not for all tasks, enabling ViT-MVT to obtain task-shared parameters for a reduction of storage and task-specific parameters to learn task-related features such that boosting performance. Finally, we introduce a joint-task evolutionary search algorithm to discover an optimal backbone for all tasks under total model size constraint, which challenges the conventional wisdom that visual tasks are typically supplied with backbone networks developed for image classification. Extensive experiments reveal that ViT-MVT delivers exceptional performances for multiple visual tasks over state-of-the-art methods while necessitating considerably fewer total storage costs. We further demonstrate that once ViT-MVT has been trained, ViT-MVT is capable of incremental learning when generalized to new tasks while retaining identical performances for trained tasks. The code is available at https://github.com/XT-1997/vitmvt.
Tao Xie 0010, Ruifeng Li 0001, Shouren Mao, Ke Wang 0028, Lijun Zhao 0003
IEEE Trans. Neural Networks Learn. Syst.6
2024 Live Demonstration: A Reconfigurable, Energy-efficient and High-frame-rate EKF-SLAM Accelerator Based SoC Design for Autonomous Mobile Robot Applications
abstract
This demonstration shows a Extend Kalman Filter-Simultaneous Localization And Mapping (EKF-SLAM) accelerator based System On Chip (SoC) design for Autonomous Mobile Robots (AMR). The AMR platform consists of a multi-sensor system with a wheel encoder and LiDAR, and a ZYNQ-7000 FPGA based SoC featuring an EKF-SLAM hardware accelerator. This AMR system achieves real-time SLAM with significant energy efficient improvement against the state-of-the-art designs.
Dingcheng Jiang, Bingqiang Liu, Ao Hu, Yequan Zhao, Minjie Bao, Zhendong Fan, Zixuan Shen, Ke Wang 0028, Chao Wang 0096
ISCAS9
2024 FMAP: Learning robust and accurate local feature matching with anchor points
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
Expert Syst. Appl.3
2024 ALNet: An adaptive channel attention network with local discrepancy perception for accurate indoor visual localization
Hongbo Gao 0008, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Mengyuan Wu
Expert Syst. Appl.3
2024 APM: Adaptive parameter multiplexing for class incremental learning
Jinghan Gao, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003
Expert Syst. Appl.4
2024 DeepMatcher: A deep transformer-based network for robust and accurate local feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
Expert Syst. Appl.3
2024 ActionMixer: Temporal action detection with Optimal Action Segment Assignment and mixers
Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001
Expert Syst. Appl.2
2024 SMTCNN - A global spatio-temporal texture convolutional neural network for 3D dynamic texture recognition
Liangliang Wang 0007, Lei Zhou 0008, Peidong Liang, Ke Wang 0028, Lianzheng Ge
Image Vis. Comput.4
2024 CO-Net++: A Cohesive Network for Multiple Point Cloud Tasks at Once With Two-Stage Feature Rectification
abstract
We present CO-Net++, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains with a two-stage feature rectification strategy. The core of CO-Net++ lies in optimizing task-shared parameters to capture universal features across various tasks while discerning task-specific parameters tailored to encapsulate the unique characteristics of each task. Specifically, CO-Net++ develops a two-stage feature rectification strategy (TFRS) that distinctly separates the optimization processes for task-shared and task-specific parameters. At the first stage, TFRS configures all parameters in backbone as task-shared, which encourages CO-Net++ to thoroughly assimilate universal attributes pertinent to all tasks. In addition, TFRS introduces a sign-based gradient surgery to facilitate the optimization of task-shared parameters, thus alleviating conflicting gradients induced by various dataset domains. In the second stage, TFRS freezes task-shared parameters and flexibly integrates task-specific parameters into the network for encoding specific characteristics of each dataset domain. CO-Net++ prominently mitigates conflicting optimization caused by parameter entanglement, ensuring the sufficient identification of universal and specific features. Extensive experiments reveal that CO-Net++ realizes exceptional performances on both 3D object detection and 3D semantic segmentation tasks. Moreover, CO-Net++ delivers an impressive incremental learning capability and prevents catastrophic amnesia when generalizing to new point cloud tasks.
Tao Xie 0010, Qihao Sun, Chuqing Cao, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2024 OAMatcher: An overlapping areas-based network with label credibility for robust and accurate feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
Pattern Recognit.3
2024 FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object Detection
abstract
In this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance.
Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082
IEEE Trans. Multim.3
2023 Poly-PC: A Polyhedral Network for Multiple Point Cloud Tasks at Once
abstract
In this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by multiple dataset domains, etc.) of multi-task learning on point cloud. Specifically, we propose a residual set abstraction (Res-SA) layer for efficient and effective scaling in both width and depth of the network, hence accommodating the needs of various tasks. We develop a weight-entanglement- based one-shot NAS technique to find optimal architectures for all tasks. Moreover, such technique entangles the weights of multiple tasks in each layer to offer task-shared parameters for efficient storage deployment while providing ancillary task-specific parameters for learning task-related features. Finally, to facilitate the training of Poly-PC, we introduce a task-prioritization-based gradient balance algorithm that leverages task prioritization to reconcile conflicting gradients, ensuring high performance for all tasks. Benefiting from the suggested techniques, models optimized by Poly-PC collectively for all tasks keep fewer total FLOPs and parameters and outperform previous methods. We also demonstrate that Poly-PC allows incremental learning and evades catastrophic forgetting when tuned to a new task.
Tao Xie 0010, Shiguang Wang, Ke Wang 0028, Linqi Yang, Xingcheng Zhang, Ruifeng Li 0001, Jian Cheng 0003
CVPR3
2023 OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes
abstract
In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net.
Tao Xie 0010, Siyi Lu, Ke Wang 0028, Jinghan Gao, Dedong Liu, Jie Xu 0066, Lijun Zhao 0003, Ruifeng Li 0001
ICCV4
2023 CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive Network
abstract
We present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task.
Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001
ICCV2
2023 A Real-Time Hardware-Accelerator-Aided MJ-EKF SLAM Algorithm for Large-Scale Map Based on Heterogeneous Multi-Core SoC
abstract
As a classic SLAM algorithm, LIDAR-based EKF SLAM still remains a problem. An open issue is high computational complexity, which leading a very limited map size to sufficient real-time requirements. To address this problem, in this paper, we propose an MJ-EKF SLAM system that is fully implemented on a Heterogeneous Multi-Core SoC(HMS). To limit the EKF SLAM computation complexity, the Map Joining algorithm participates in the framework. The most computation cost step in EKF SLAM and Map Joining algorithms is implemented on an MJ-EKF dedicated accelerator. Meanwhile, several hardware optimization methods are proposed to save logic resources, avoid redundant computation and reduce data communication between on-chip memory and off-chip main memory. Field experiments demonstrate the HMS-based MJ-EKF algorithm achieved high real-time performance (over 30Hz) in a large-scale map with 500 landmarks.
Zhendong Fan, Ke Wang 0028, Minjie Bao, Ruifeng Li 0001
IECON2
2023 GCA-Net: A Global Context Aggregation Network for Effective Optical Flow
abstract
Optical flow seeks to estimate the per-pixel 2D motion between two frames by identifying corresponding pixels. Current flow estimators typically involve per-pixel feature extraction, multi-scale 4D correlation volume construction, and iterative flow field updates through a Conv-GRU module. However, the locality of convolutional features in these methods renders the calculated correlations vulnerable to different noises. In addition, the Conv-GRU module of these methods is only executed with a convolution layer, which is incapable of exploiting context clues from larger window sizes even inside the image being queried itself, thus raising it more difficult for the model to process images with challenging regions, e.g., textureless areas. To this end, we introduce GCA-Net, a global context aggregation network for credible yet effective optical flow estimation. More specifically, we propose a highly efficient multi-scale transformer (MSFormer) layer which enables the per-pixel feature to aggregate long-range information from other features, hence building more accurate 4D correlation volumes. Besides, we develop an attention-enhanced Conv-GRU block (AttGRU) that can incorporate information alongside a larger context window even within itself empowering our network to estimate optical flow in even the most challenging regions. Experimentally, we demonstrate that GCA-Net outperforms previous state-of-the-art methods by large margins on Sintel (Final) and KITTI 2015 (background and foreground) benchmarks.
Tao Xie 0010, Jinghan Gao, Ke Wang 0028, Ruifeng Li 0001
IECON3
2023 Poly-MOT: A Polyhedral Framework For 3D Multi-Object Tracking
abstract
3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association and state estimation for all objects. With large-scale modern datasets and real scenes, there are a variety of object categories that commonly exhibit distinctive geometric properties and motion patterns. In this way, such distinctions would enable various object categories to behave differently under the same standard, resulting in erroneous matches between trajectories and detections, and jeopardizing the reliability of downstream tasks (navigation, etc.). Towards this end, we propose Poly-MOT, an efficient 3D MOT method based on the Tracking-By-Detection framework that enables the tracker to choose the most appropriate tracking criteria for each object category. Specifically, Poly-MOT leverages different motion models for various object categories to characterize distinct types of motion accurately. We also introduce the constraint of the rigid structure of objects into a specific motion model to accurately describe the highly nonlinear motion of the object. Additionally, we introduce a two-stage data association strategy to ensure that objects can find the optimal similarity metric from three custom metrics for their categories and reduce missing matches. On the NuScenes dataset, our proposed method achieves state-of-the-art performance with 75.4% AMOTA. The code is available at https://github.com/lixiaoyu20001P0Iy-MOT.
Tao Xie 0010, Dedong Liu, Jinghan Gao, Lijun Zhao 0003, Ke Wang 0028
IROS8
2023 Actions as points: a simple and efficient detector for skeleton-based temporal action detection
Ke Wang 0028, Ruifeng Li 0001
Mach. Vis. Appl.2
2023 Cascading spatio-temporal attention network for real-time action detection
Ke Wang 0028, Ruifeng Li 0001, Petra Perner
Mach. Vis. Appl.2
2023 Point-NAS: A Novel Neural Architecture Search Framework for Point Cloud Analysis
abstract
Recently, point-based networks have exhibited extraordinary potential for 3D point cloud processing. However, owing to the meticulous design of both parameters and hyperparameters inside the network, constructing a promising network for each point cloud task can be an expensive endeavor. In this work, we develop a novel one-shot search framework called Point-NAS to automatically determine optimum architectures for various point cloud tasks. Specifically, we design an elastic feature extraction (EFE) module that serves as a basic unit for architecture search, which expands seamlessly alongside both the width and depth of the network for efficient feature extraction. Based on the EFE module, we devise a searching space, which is encoded into a supernet to provide a wide number of latent network structures for a particular point cloud task. To fully optimize the weights of the supernet, we propose a weight coupling sandwich rule that samples the largest, smallest, and multiple medium models at each iteration and fuses their gradients to update the supernet. Furthermore, we present a united gradient adjustment algorithm that mitigates gradient conflict induced by distinct gradient directions of sampled models and supernet, thus expediting the convergence of the supernet and assuring that it can be comprehensively trained. Pursuant to the provided techniques, the trained supernet enables a multitude of subnets to be incredibly well-optimized. Finally, we conduct an evolutionary search for the supernet under resource constraints to find promising architectures for different tasks. Experimentally, the searched Point-NAS with weights inherited from the supernet realizes outstanding results across a variety of benchmarks. i.e., 94.2% and 88.9% overall accuracy under ModelNet40 and ScanObjectNN, 68.6% mIoU under S3DIS, 63.6% and 69.3% [email protected] under SUN RGB-D and ScanNet V2 datasets.
Tao Xie 0010, Linqi Yang, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
IEEE Trans. Image Process.4
2023 Real-Time Point Cloud Object Detection via Voxel-Point Geometry Abstraction
abstract
Recent advances in 3D object detection typically learn voxel-based or point-based representations on point clouds. Point-based methods preserve precise point positions but incur high computational load, whereas voxel-based methods rasterize unordered points into voxel grids efficiently but give rise to an accuracy bottleneck. To take advantage of voxel- and point-based representations, we develop an effective and efficient 3D object detector via a novel voxel-point geometry abstraction scheme. Our motivation is to use coarse voxel representation to accelerate proposal generation while using precise point representation to facilitate proposal refinement. For voxel representation learning, we propose a context enrichment module with a novel 3D sparse interpolation layer to augment raw points with multi-scale context. We further develop a point-based RoI pooling module with explicit position augmentation for proposal refinement. Extensive experiments on the widely used KITTI Dataset and the latest Waymo Open Dataset show that the proposed algorithm outperforms state-of-the-art point-voxel-based methods while running at 24 FPS on the TITAN XP GPU.
Guangsheng Shi, Ke Wang 0028, Ruifeng Li 0001, Chao Ma 0004
IEEE Trans. Intell. Transp. Syst.2
2022 A novel fast combine-and-conquer object detector based on only one-level feature map
Ke Wang 0028, Ruifeng Li 0001, Zhonghao Qin, Petra Perner
Comput. Vis. Image Underst.2
2022 Welding splash and arc noise reduction imaging model based on computationally efficient pairwise response serving welding process library
Zhonghao Qin, Ke Wang 0028, Ruifeng Li 0001
Mach. Vis. Appl.2
2021 Multi-obstacle path planning and optimization for mobile robot
Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028, Xichun Gui
Expert Syst. Appl.4
2015 Unsupervised feature selection based on spectral regression from manifold learning for facial expression recognition
abstract
In this study, an unsupervised feature selection method is proposed for facial feature recognition (FER) in the absence of class labels. The contribution is the descriptive feature components selector spectral regression representative coefficient scores based on graph manifold learning from high‐dimensional feature space. The spectral regression analysis and L1‐regularised least square are then used to compute the importance of features in the original space, so that less representative features with lower coefficient scores will be removed without prior distribution assumption. To verify the performance of the authors’ method, some classifiers are used to classify facial expressions on three benchmark facial expression databases. The recognition results indicate the availability and effectiveness of the proposed method for FER.
Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001
IET Comput. Vis.2