EDBT 2026 Demo / reviewers in the wild / expert
Tao Xie 0010
dblp:181/2328-10
· DBLP profile ↗
35ranked-venue papers
10as first author
35since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 14 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFMLLM: Enhance multi-modal large language model for global and fine-grained visual spatial perception
Zhendong Fan, Jinyang Gao, Hongbo Gao 0008, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 6 |
| 2026 | V 2 -Fusion: Virtual voxel enhanced 4D radar-image feature fusion for 3D object detection
Li Wang 0092, Xinyu Zhang 0001, Yuxuan Fan, Tao Xie 0010, Lei Yang 0060, Bin Xu 0003 |
Expert Syst. Appl. | 6 |
| 2026 | MTNet: A Mixed Transformer Network for High-Quality 3-D Object Detectionabstract3D object detection from point clouds represents a formidable challenge, necessitating the accurate identification and localization of objects within a 3D space. Recent advancements have showcased the efficacy of point-based detectors, leveraging local aggregators to encode intricate structural details of the point cloud. However, a notable limitation resides in their treatment of each point and object proposal in isolation, devoid of considering the interrelationships among them, thus impeding the overall detection performance. In this work, we argue that the integration of contextual information is paramount, particularly in the realm of indoor 3D object detection. Indoor environments are inherently characterized by robust contextual constraints, providing a rich tapestry for enhanced scene comprehension. In this way, we introduce MTNet, a mixed transformer network for high-quality indoor 3D object detection. Technically, we develop a mixed transformer (MixFormer) block that is purpose-built to intricately model the synergistic interplay between local structural information and global contextual features of 3D point clouds. In contrast to the classical transformer, our proposed MixFormer incorporates a local feature aggregator engineered to capture local geometric information while leveraging a KNN(K-Nearest Neighbors)-based attention mechanism to aggregate global contextual information. Furthermore, we suggest a feature compensator to adaptively fuse the strengths of local and global features, further bolstering detection performance. The culmination of these proposed components results in our MTNet framework, a hierarchical, versatile pipeline that consistently outperforms existing works across a multitude of benchmarks. In addition, we affirm the potential of proposed MixFormer and FC as generic modules that are capable of augmenting performance across a spectrum of 3D downstream point cloud tasks. Ruqi Liu, Shuaiyan Liu, Linqi Yang, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
IEEE Internet Things J. | 5 |
| 2026 | OFVL-MS++: Once for visual localization across multiple scenes via a two-stage framework
Chunsheng Yang, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 5 |
| 2026 | An MoE-Driven Unified Image Restoration Framework for Adverse Weather ConditionsabstractAdverse natural weather conditions frequently cause substantial performance degradation in outdoor vision systems, underscoring the critical importance of research on image restoration techniques. Employing a unified set of network parameters to restore degraded images across diverse weather conditions has emerged as a key research direction in the field of image restoration. In this work, we propose MUIRF, a Mixture-of-Experts (MoE)-driven unified image restoration framework for multiple adverse weather conditions. Specifically, our technical contribution includes a novel channel-level parameter sharing strategy guided by a shallow-feature-based MoE (CPSM). This fine-grained parameter sharing strategy adaptively selects convolution weight channels for cross-task sharing based on the input image, enabling the network to accurately capture weather-general features, while the remaining channels encode weather-specific features corresponding to each weather condition. CPSM facilitates precise channel selection, thereby enhancing the robustness and accuracy of MUIRF during joint training across diverse image restoration tasks under varying weather conditions. Additionally, gradient conflicts inevitably arise in shared parameters due to the divergent optimization objectives across tasks. To address this challenge, we propose a meta-vector-guided gradient homogenization (MVGH) algorithm that mitigates inter-task gradient conflicts and improves image restoration quality. Comprehensive experimental evaluations demonstrate that our proposed network outperforms most state-of-the-art approaches, validating its superior performance and effectiveness. Hongbo Gao 0008, Ruqi Liu, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | COME: A Collaborative Optimization Framework With Low-Rank MoE for Indoor 3D Object DetectionabstractIndoor 3D object detection serves as a fundamental task in computer vision and robotics. Existing research predominantly focuses on training domain-specific optimal models for individual datasets, yet it overlooks the potential value of capturing universal geometric attributes that can substantially enhance object detection performance across diverse domains. To resolve this gap, we propose COME, a novel and effective collaborative optimization framework designed to seamlessly integrate these universal attributes while preserving the domain-specific characteristics of each dataset domain. COME is built on VoteNet and incorporates a Cross-Domain Expert Parameter Sharing Strategy (CEPSS) that draws inspiration from the Mixture of Experts (MoE) framework. Its core innovation resides in the dual-expert design of CEPSS: domain-shared experts capture universal geometric relationships across datasets, whereas domain-specific experts encode unique features for individual datasets. This separation enables the model to focus on learning both generic and domain-specialized visual cues, without mutual interference. In addition, to dynamically adapt to different domains, we design a lightweight gating network that automatically selects relevant experts, eliminating irrelevant feature interference and enhancing model specialization. Compared to standard parameter-sharing architectures, this design significantly reduces gradient conflicts during multi-domain training. We further optimize computational efficiency by implementing low-rank structures for domain-shared and domain-specific experts, thus striking a better balance between memory overhead and detection performance. Experiments show that COME achieves state-of-the-art results across benchmarks, with acceptable parameter growth, and outperforms existing multi-domain detection methods. Hongbo Gao 0008, Zimeng Tong, Fuyuan Qiu, Tao Xie 0010, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 4 |
| 2025 | DVDS: A deep visual dynamic slam system
Tao Xie 0010, Qihao Sun, Tao Sun 0024, Jinhang Zhang, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
Expert Syst. Appl. | 1 |
| 2025 | Poly-TF: A polymeric transformer framework for multiple visual tasks at once
Xuan Fan, Tao Sun 0024, Jinghan Gao, Tao Xie 0010, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 5 |
| 2025 | MTF-Net: A mediator transformer-based fusion network with MOE for 6D object pose estimation
Zimeng Tong, Youran Du, Tao Xie 0010, Ke Wang 0028, Lijun Zhao 0003, Ruifeng Li 0001 |
Knowl. Based Syst. | 6 |
| 2025 | CTFS: A consolidated transformer framework for instance and semantic segmentation tasks
Fuyuan Qiu, Hongbo Gao 0008, Tao Xie 0010, Chuqing Cao, Ruifeng Li 0001, Lijun Zhao 0003, Ke Wang 0028 |
Neural Networks | 4 |
| 2025 | NL-WCS: A Novel Data-Driven Algorithm for Extracting the Dynamics of Serial Robots Considering Non-Linear FrictionabstractRecently, developed data-driven SINDy-based techniques can identify the dynamic model of serial robots without simplifying assumptions nor pre-knowledge of all kinematics and geometric details. However, these techniques cannot handle non-linear friction models, which significantly affects the precision of the dynamic model identification. This study proposes a novel data-driven approach for dynamic model identification considering the non-linear friction model along with SINDy concept. This approach is termed as non-linear-weighted-constrained SINDy (NL-WCS). The SINDy concept is extended to accommodate any non-linear friction model by efficiently incorporating the Levenberg-Marquardt (LM) algorithm. Weighted L1 regularization is combined with the physics and the data constraints, to promote sparsity in the recovery of the dynamic equations. Moreover, this combination makes the approach robust against the regression matrix’s ill-conditionality and noise. LM is integrated with the robot’s SIMULINK model to get initial values for the non-linear friction empirical parameters. NL-WCS is experimentally evaluated by utilizing three distinct trajectories to validate the extracted dynamic model of 6-DOF UR10 and 7-DOF KUKA robots. In addition, five non-linear friction models are compared. NL-WCS outperforms all the previous SINDy-data-driven methods since it reduced the RMSE significantly up to 60.65%. NL-WCS also demonstrated robustness when tested against different levels of noise. Note to Practitioners—High-performance serial robot model-based controllers need precise dynamic model identification. The classical analytical methods for modeling serial robots rely on deriving the dynamic equations with simplified assumptions and then identifying the inertial parameters. Specifically, they assume that all the robot’s geometric details are known. Furthermore, there are uncertainties in geometric parameter values due to manufacturing/assembly errors. Data-driven-based SINDy methods can derive the dynamic model without pre-knowledge of the geometric parameters of the robot. However, it can’t incorporate non-linear friction models. Thus, this paper proposed a novel data-driven technique to extract the dynamic model of any serial manipulator by coupling SINDy approach with the nonlinear friction models which is the realistic case. The practitioners can benefit from a technique like that, in such a way of applying NL-WCS to different types of industrial robots, particularly, those that have missing manufacturer data sets or sheets. Simple and reliable dynamic models can be derived only by running the robot with various trajectories. Also, this helps eliminate any uncertainties in the dynamic model discovery/building process. Furthermore, Incorporating the non-linear friction model in the derivation process enhances the dynamic modeling accuracy as it effectively takes into account all friction characteristics. The proposed approach can be converted into a software package that the practitioners can use to derive the dynamics without getting their heads around the overwhelming robot’s kinematic details. Mohamed Omar, Ke Wang 0028, Ruifeng Li 0001, Ossama B. Abouelatta, Tao Xie 0010, Mohamed Gouda Alkalla |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | FMRT: Learning Accurate Feature Matching With Reconciliatory TransformerabstractLocal Feature Matching, a pivotal component of numerous computer vision tasks (e.g., structure from motion and visual localization), has been effectively addressed by Transformer-based methods. Nevertheless, these methods solely incorporate long-range context information among keypoints with a fixed receptive field, which constrains the network from appropriately reconciling the importance of features with diverse receptive fields to realize complete image perception, hence limiting feature matching accuracy. In addition, these methods employ a conventional handcrafted encoding approach to incorporate positional information of keypoints into visual descriptors, which limits the capability of networks to extract effective positional encoding message. In this study, we propose FMRT, a novel detector-free method that reconciles local features with diverse receptive fields adaptively and utilizes parallel networks to realize reliable positional encoding. Specifically, FMRT proposes a dedicated reconciliatory transformer (RecFormer) that contains a global perception attention layer to identify visual descriptors with different receptive fields and integrate global context information under various scales, a perception weight layer to measure the importance of various receptive fields adaptively, and a local perception feed-forward network to extract deep aggregated multi-scale local feature representation. Moreover, we introduce a novel axis-wise position encoder (AWPE) that views positional encoding as two keypoints encoding tasks along the row and column dimensions, decouples the x- and y-coordinates of keypoints into two independent 1D vectors, and designs two parallel network branches to explicitly encodes geometric correlations among keypoints, hence realizing reliable positional encoding. Extensive experiments indicate that FMRT yields impressive performance on multiple tasks, including relative pose estimation, visual localization, homography estimation, and image matching. Besides, we integrate FMRT into a localization framework and conduct a visual localization experiment in a real scene, which further demonstrate the superiority of FMRT. Note to Practitioners—This paper presents a novel approach to enhancing the performance of local feature matching in computer vision tasks. Traditional methods often rely on fixed receptive fields for integrating context among keypoints, which can limit the perception of the complete image and, consequently, the precision of feature matching. Our work introduces a Reconciliatory Transformer that not only addresses these limitations by effectively reconciling the importance of features across varying receptive fields but also improves the integration of positional information into visual descriptors. The techniques developed here can be adapted to a wide range of systems, e.g., image matching for computer vision and visual localization for autonomous driving, offering practitioners a tool to significantly improve the fidelity of feature matching, which is foundational for accurate interaction with the surrounding environment. Li Wang 0092, Xinyu Zhang 0001, Tao Xie 0010, Lei Yang 0060, Wenhao Yu 0006, Yang Shen 0005, Bin Xu 0003, Jun Li 0082 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | VD-Matcher: A Very Deep Local Feature Matcher With Weight Recycling and Keypoint DetectionabstractEstablishing local feature matches between image pairs serves as a fundamental component of plentiful vision tasks, such as visual localization and Structure from Motion (SfM). Recently, detector-free techniques equipped with the transformer have exhibited exceptional performance. Theoretically, optimizing the transformer architecture and stacking more transformer blocks could emphasize crucial features and filter out extraneous information by progressively narrowing the effective perception regions of the network within images, thereby enhancing the matching performance. Nevertheless, this paradigm results in a linear escalation of model size with respect to the number of blocks. In this study, we introduce VD-Matcher to address this issue. A principal innovation of VD-Matcher is the utilization of a weight recycling technique (WRT) that enables partial weights to be reutilized across successive transformer blocks, along with specific transformations designed to sufficiently enhance feature representations. This approach enables VD-Matcher to construct a deep transformer architecture for accurate local feature matching while maintaining a manageable parameter size. Furthermore, we propose a lightweight multi-scale keypoint detection module that captures representative keypoints to replace all keypoints for compact global information aggregation intra-/inter- images, which reduces the computational overhead induced by excessively deep transformer layers while alleviating redundant information propagation to a certain extent. Extensive experiments verify that VD-Matcher exceeds state-of-the-art algorithms on multiple benchmarks while maintaining less parameters. The source code is available at https://github.com/mooncake199809/VD-Matcher. Qihao Sun, Tao Xie 0010, Hongbo Gao 0008, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | AGHL: Anchor-Guided Point Cloud Registration Network With Hybrid Local Feature PerceptionabstractPoint cloud registration, which estimates a rigid transformation matrix between two point clouds, is a fundamental process in numerous applications. While existing detector-free techniques present exceptional performance, they overlook the extraction of hybrid local features that capture correlations between points and their neighbours, thereby limiting the quality of point cloud recognition. Moreover, these approaches typically treat point clouds as sequential data and employ the transformer to integrate global context from all points, which inevitably introduces interference from irrelevant regions, hence affecting the registration accuracy. In this work, we propose a novel detector-free approach AGHL to address these challenges. For the first issue, AGHL introduces a hybrid local feature perception module that designs two parallel branches to concurrently extract low-level and high-level local features, which effectively encode the correlations between each point and its neighborhood points in both Euclidean space and high-dimensional feature space. For the second issue, AGHL develops an anchor-guided cross attention that adheres to the local geometric consistency to constrain the network's attention on reliable anchors, thereby effectively suppressing interference from irrelevant regions. Benefiting from these techniques, AGHL achieves impressive point cloud registration accuracy across all synthetic, indoor, and outdoor datasets. Furthermore, we build an experimental platform and conduct a real-world robot localization experiment, with results showing the strong generalization ability of AGHL. Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003, Chuqing Cao |
IEEE Trans. Image Process. | 2 |
| 2025 | SOFW: A Synergistic Optimization Framework for Indoor 3D Object DetectionabstractIn this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar |
IEEE Trans. Multim. | 3 |
| 2025 | Centra-Net: A Centralized Network for Visual Localization Spanning Multiple ScenesabstractWe present Centra-Net, a centralized network that concurrently optimizes visual localization over numerous scenes under heterogeneous dataset domains. Centra-Net exemplifies storage efficiency by amalgamating multiple models with task-shared parameters into a singular cohesive structure. Technically, we develop abasic feature extraction unit (BFEU)with two parallel branches: one dedicated to local feature extraction and the other adept at adaptively generating a task-specific attention mask for feature calibration, thus bolstering its feature extraction capability across diverse scenes. Based on the BFEU, we introduce afilter-wise sharing mechanism (FSM)that adaptively determines parameter sharing within the unit, thus facilitating fine-grained parameter allocation. The key insight of FSM resides in reconceptualizing the parameter sharing of the unit as a learnable paradigm, enabling the determination of shared parameters to be made post-training. Finally, we suggest acomplexity-prioritized gradient algorithm (CPGA)that capitalizes on task complexity to attain a harmonious learning space for various tasks, thus safeguarding optimal performances across all tasks. Through rigorous experiments on numerous benchmarks, Centra-Net demonstrates a notable edge over existing state-of-the-art works while operating with a significantly reduced parameter footprint. Ke Wang 0028, Tao Xie 0010, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | COFP: A Collaborative Optimization Framework With Polyhedral Feature Extraction for Multi-Weather Image RestorationabstractImage restoration in adverse weather conditions is a critical research focus in computer vision and autonomous driving. In this work, we introduce COFP, a collaborative optimization framework designed to simultaneously enhance the performance of image de-raining, de-snowing, and de-hazing tasks across diverse datasets. The core of COFP lies in its adaptive optimization of weather-shared and weather-specific parameters, enabling the extraction of polyhedral features that effectively integrate both weather-shared and weather-specific attributes, thus substantially boosting multi-weather performance of the network. Technically, we design a polyhedral feature extraction module (PFEM) to facilitate the acquisition of weather-shared attribute and weather-specific attribute. In PFEM, we first introduce an element-adaptive sharing strategy (ESS) that dynamically activates either weather-shared or weather-specific parameters for each element of the weight matrix based on learnable scores, thereby adaptively determining which parameters are shared or not. Secondly, we develop a feature extraction enhancement strategy (FEES), which extends two pathways in PFEM comprised of standard convolutional layers to enhance the capture of polyhedral features, further promoting the model performance. Furthermore, we propose a gradient balancing algorithm (GBA) that mitigates the unequal competition among tasks for shared parameters during network optimization by adaptively adjusting the direction of task gradients, effectively addressing the negative transfer issue induced by domain variations in multi-task learning. Experimental results demonstrate that COFP delivers state-of-the-art performances across various adverse weather image restoration benchmarks. Tao Xie 0010, Ruifeng Li 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | HVLF: A Holistic Visual Localization Framework Across Diverse ScenesabstractRecently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods. Fuyuan Qiu, Dedong Liu, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | ViT-MVT: A Unified Vision Transformer Network for Multiple Vision TasksabstractIn this work, we seek to learn multiple mainstream vision tasks concurrently using a unified network, which is storage-efficient as numerous networks with task-shared parameters can be implanted into a single consolidated network. Our framework, vision transformer (ViT)-MVT, built on a plain and nonhierarchical ViT, incorporates numerous visual tasks into a modest supernet and optimizes them jointly across various dataset domains. For the design of ViT-MVT, we augment the ViT with a multihead self-attention (MHSE) to offer complementary cues in the channel and spatial dimension, as well as a local perception unit (LPU) and locality feed-forward network (locality FFN) for information exchange in the local region, thus endowing ViT-MVT with the ability to effectively optimize multiple tasks. Besides, we construct a search space comprising potential architectures with a broad spectrum of model sizes to offer various optimum candidates for diverse tasks. After that, we design a layer-adaptive sharing technique that automatically determines whether each layer of the transformer block is shared or not for all tasks, enabling ViT-MVT to obtain task-shared parameters for a reduction of storage and task-specific parameters to learn task-related features such that boosting performance. Finally, we introduce a joint-task evolutionary search algorithm to discover an optimal backbone for all tasks under total model size constraint, which challenges the conventional wisdom that visual tasks are typically supplied with backbone networks developed for image classification. Extensive experiments reveal that ViT-MVT delivers exceptional performances for multiple visual tasks over state-of-the-art methods while necessitating considerably fewer total storage costs. We further demonstrate that once ViT-MVT has been trained, ViT-MVT is capable of incremental learning when generalized to new tasks while retaining identical performances for trained tasks. The code is available at https://github.com/XT-1997/vitmvt. Tao Xie 0010, Ruifeng Li 0001, Shouren Mao, Ke Wang 0028, Lijun Zhao 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | FMAP: Learning robust and accurate local feature matching with anchor points
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 2 |
| 2024 | APM: Adaptive parameter multiplexing for class incremental learning
Jinghan Gao, Tao Xie 0010, Ruifeng Li 0001, Ke Wang 0028, Lijun Zhao 0003 |
Expert Syst. Appl. | 2 |
| 2024 | CorMatcher: A corners-guided graph neural network for local feature matching
Hainan Luo, Tao Xie 0010, Chuqing Cao, Lijun Zhao 0003 |
Expert Syst. Appl. | 2 |
| 2024 | DeepMatcher: A deep transformer-based network for robust and accurate local feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Expert Syst. Appl. | 1 |
| 2024 | PSE-Net: Channel pruning for Convolutional Neural Networks with parallel-subnets estimator
Shiguang Wang, Tao Xie 0010, Haijun Liu 0001, Xingcheng Zhang, Jian Cheng 0003 |
Neural Networks | 2 |
| 2024 | CO-Net++: A Cohesive Network for Multiple Point Cloud Tasks at Once With Two-Stage Feature RectificationabstractWe present CO-Net++, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains with a two-stage feature rectification strategy. The core of CO-Net++ lies in optimizing task-shared parameters to capture universal features across various tasks while discerning task-specific parameters tailored to encapsulate the unique characteristics of each task. Specifically, CO-Net++ develops a two-stage feature rectification strategy (TFRS) that distinctly separates the optimization processes for task-shared and task-specific parameters. At the first stage, TFRS configures all parameters in backbone as task-shared, which encourages CO-Net++ to thoroughly assimilate universal attributes pertinent to all tasks. In addition, TFRS introduces a sign-based gradient surgery to facilitate the optimization of task-shared parameters, thus alleviating conflicting gradients induced by various dataset domains. In the second stage, TFRS freezes task-shared parameters and flexibly integrates task-specific parameters into the network for encoding specific characteristics of each dataset domain. CO-Net++ prominently mitigates conflicting optimization caused by parameter entanglement, ensuring the sufficient identification of universal and specific features. Extensive experiments reveal that CO-Net++ realizes exceptional performances on both 3D object detection and 3D semantic segmentation tasks. Moreover, CO-Net++ delivers an impressive incremental learning capability and prevents catastrophic amnesia when generalizing to new point cloud tasks. Tao Xie 0010, Qihao Sun, Chuqing Cao, Lijun Zhao 0003, Ke Wang 0028, Ruifeng Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | OAMatcher: An overlapping areas-based network with label credibility for robust and accurate feature matching
Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
Pattern Recognit. | 2 |
| 2024 | Auto-Points: Automatic Learning for Point Cloud Analysis With Neural Architecture SearchabstractPure point-based neural networks have recently shown tremendous promise for point cloud tasks, including 3D object classification, 3D object part segmentation, 3D semantic segmentation, and 3D object detection. Nevertheless, it is a laborious process to construct a network for each task due to the artificial parameters and hyperparameters involved, e.g., the depths and widths of the network and the number of sampled points at each stage. In this work, we propose Auto-Points, a novel one-shot search framework that automatically seeks the optimal architecture configuration for point cloud tasks. Technically, we introduce a set abstraction mixer (SAM) layer that is capable of scaling up flexibly along the depth and width of the network. Each SAM layer consists of numerous child candidates, which simplifies architecture search and enables us to discover the optimum design for each point cloud task pursuant to resource constraint from an enormous search space. To fully optimize the child candidates, we develop a weight-entwinement neural architecture search (NAS) technique that entwines the weights of different candidates in the same layer during supernet training such that all candidates can be extremely optimized. Benefiting from the proposed techniques, the trained supernet allows the searched subnets to be exceptionally well-optimized without further retraining or finetuning. In particular, the searched models deliver superior performances on multiple extensively employed benchmarks, 93.9% overall accuracy (OA) on ModelNet40, 89.1% OA on ScanObjectNN, 87.1% instance average IoU on ShapeNetPart, 69.1% mIoU on S3DIS, 70.4% [email protected] on ScanNet V2, and 64.4% [email protected] on SUN RGB-D. Li Wang 0092, Tao Xie 0010, Xinyu Zhang 0001, Linqi Yang, Yilong Ren, Haiyang Yu 0002, Jun Li 0082, Huaping Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | FARP-Net: Local-Global Feature Aggregation and Relation-Aware Proposals for 3D Object DetectionabstractIn this work, we introduce FARP-Net, an adaptive local-global feature aggregation and relation-aware proposal network for high-quality 3D object detection from pure point clouds. Our key insight is that learning adaptive local-global feature aggregation from an irregular yet sparse point cloud and generating superb proposals are both pivotal for detection. Technically, we propose a novel local-global feature aggregation layer (LGFAL) that fully exploits the complementary correlation between local features and global features, and fuses their strengths adaptively via an attention-based fusion module. Furthermore, we incorporate a lightweight feature affine module (LFAM) into LGFAL to map the local features into a normal distribution, thus acquiring fine-grained features of each local region in a weight-sharing manner. During object proposal generation, we propose a weighted relation-aware proposal module (WRPM) that uses an objectness-aware formalism to weigh the relation importance among object candidates for a clear and principal context, thereby facilitating the generation of high-quality proposals. The WRPM challenges the traditional practice of extracting contextual information among all object candidates, which is inefficient as object candidates are always noisy and redundant. Experimentally, FARP-Net delivers superior performance on two widely used benchmarks with fewer parameters, 64.0% [email protected] on the SUN RGB-D dataset and 70.9% [email protected] on the ScanNet V2 dataset. We further validate that the proposed LGFAL and WRPM can be integrated into both indoor and outdoor detectors to boost performance. Tao Xie 0010, Li Wang 0092, Ke Wang 0028, Ruifeng Li 0001, Xinyu Zhang 0001, Linqi Yang, Huaping Liu 0001, Jun Li 0082 |
IEEE Trans. Multim. | 1 |
| 2023 | MDL-NAS: A Joint Multi-domain Learning Framework for Vision TransformerabstractIn this work, we introduce MDL-NAS, a unified frame-work that integrates multiple vision tasks into a manageable supernet and optimizes these tasks collectively under diverse dataset domains. MDL-NAS is storage-efficient since multiple models with a majority of shared parameters can be deposited into a single one. Technically, MDL-NAS constructs a coarse-to-fine search space, where the coarse search space offers various optimal architectures for different tasks while the fine search space provides fine-grained parameter sharing to tackle the inherent obstacles of multi-domain learning. In the fine search space, we suggest two parameter sharing policies, i.e., sequential sharing policy and mask sharing policy. Compared with previous works, such two sharing policies allow for the partial sharing and non-sharing of parameters at each layer of the network, hence attaining real fine-grained parameter sharing. Finally, we present a joint-subnet search algorithm that finds the optimal architecture and sharing parameters for each task within total resource constraints, challenging the traditional practice that downstream vision tasks are typically equipped with backbone networks designed for image classification. Experimentally, we demonstrate that MDL-NAS families fitted with non-hierarchical or hierarchical transformers deliver competitive performance for all tasks compared with state-of-the-art methods while maintaining efficient storage deployment and computation. We also demonstrate that MDL-NAS allows incremental learning and evades catastrophic forgetting when generalizing to a new task. Shiguang Wang, Tao Xie 0010, Jian Cheng 0003, Xingcheng Zhang, Haijun Liu 0001 |
CVPR | 2 |
| 2023 | Poly-PC: A Polyhedral Network for Multiple Point Cloud Tasks at OnceabstractIn this work, we show that it is feasible to perform multiple tasks concurrently on point cloud with a straightforward yet effective multi-task network. Our framework, Poly-PC, tackles the inherent obstacles (e.g., different model architectures caused by task bias and conflicting gradients caused by multiple dataset domains, etc.) of multi-task learning on point cloud. Specifically, we propose a residual set abstraction (Res-SA) layer for efficient and effective scaling in both width and depth of the network, hence accommodating the needs of various tasks. We develop a weight-entanglement- based one-shot NAS technique to find optimal architectures for all tasks. Moreover, such technique entangles the weights of multiple tasks in each layer to offer task-shared parameters for efficient storage deployment while providing ancillary task-specific parameters for learning task-related features. Finally, to facilitate the training of Poly-PC, we introduce a task-prioritization-based gradient balance algorithm that leverages task prioritization to reconcile conflicting gradients, ensuring high performance for all tasks. Benefiting from the suggested techniques, models optimized by Poly-PC collectively for all tasks keep fewer total FLOPs and parameters and outperform previous methods. We also demonstrate that Poly-PC allows incremental learning and evades catastrophic forgetting when tuned to a new task. Tao Xie 0010, Shiguang Wang, Ke Wang 0028, Linqi Yang, Xingcheng Zhang, Ruifeng Li 0001, Jian Cheng 0003 |
CVPR | 1 |
| 2023 | OFVL-MS: Once for Visual Localization across Multiple Indoor ScenesabstractIn this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net. Tao Xie 0010, Siyi Lu, Ke Wang 0028, Jinghan Gao, Dedong Liu, Jie Xu 0066, Lijun Zhao 0003, Ruifeng Li 0001 |
ICCV | 1 |
| 2023 | CO-Net: Learning Multiple Point Cloud Tasks at Once with A Cohesive NetworkabstractWe present CO-Net, a cohesive framework that optimizes multiple point cloud tasks collectively across heterogeneous dataset domains. CO-Net maintains the characteristics of high storage efficiency since models with the preponderance of shared parameters can be assembled into a single model. Specifically, we leverage residual MLP (Res-MLP) block for effective feature extraction and scale it gracefully along the depth and width of the network to meet the demands of different tasks. Based on the block, we propose a novel nested layer-wise processing policy, which identifies the optimal architecture for each task while provides partial sharing parameters and partial non-sharing parameters inside each layer of the block. Such policy tackles the inherent challenges of multi-task learning on point cloud, e.g., diverse model topologies resulting from task skew and conflicting gradients induced by heterogeneous dataset domains. Finally, we propose a sign-based gradient surgery to promote the training of CO-Net, thereby emphasizing the usage of task-shared parameters and guaranteeing that each task can be thoroughly optimized. Experimental results reveal that models optimized by CO-Net jointly for all point cloud tasks maintain much fewer computation cost and overall storage cost yet outpace prior methods by a significant margin. We also demonstrate that CO-Net allows incremental learning and prevents catastrophic amnesia when adapting to a new point cloud task. Tao Xie 0010, Ke Wang 0028, Siyi Lu, Jie Xu 0066, Li Wang 0092, Lijun Zhao 0003, Xinyu Zhang 0001, Ruifeng Li 0001 |
ICCV | 1 |
| 2023 | GCA-Net: A Global Context Aggregation Network for Effective Optical FlowabstractOptical flow seeks to estimate the per-pixel 2D motion between two frames by identifying corresponding pixels. Current flow estimators typically involve per-pixel feature extraction, multi-scale 4D correlation volume construction, and iterative flow field updates through a Conv-GRU module. However, the locality of convolutional features in these methods renders the calculated correlations vulnerable to different noises. In addition, the Conv-GRU module of these methods is only executed with a convolution layer, which is incapable of exploiting context clues from larger window sizes even inside the image being queried itself, thus raising it more difficult for the model to process images with challenging regions, e.g., textureless areas. To this end, we introduce GCA-Net, a global context aggregation network for credible yet effective optical flow estimation. More specifically, we propose a highly efficient multi-scale transformer (MSFormer) layer which enables the per-pixel feature to aggregate long-range information from other features, hence building more accurate 4D correlation volumes. Besides, we develop an attention-enhanced Conv-GRU block (AttGRU) that can incorporate information alongside a larger context window even within itself empowering our network to estimate optical flow in even the most challenging regions. Experimentally, we demonstrate that GCA-Net outperforms previous state-of-the-art methods by large margins on Sintel (Final) and KITTI 2015 (background and foreground) benchmarks. Tao Xie 0010, Jinghan Gao, Ke Wang 0028, Ruifeng Li 0001 |
IECON | 1 |
| 2023 | Poly-MOT: A Polyhedral Framework For 3D Multi-Object Trackingabstract3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association and state estimation for all objects. With large-scale modern datasets and real scenes, there are a variety of object categories that commonly exhibit distinctive geometric properties and motion patterns. In this way, such distinctions would enable various object categories to behave differently under the same standard, resulting in erroneous matches between trajectories and detections, and jeopardizing the reliability of downstream tasks (navigation, etc.). Towards this end, we propose Poly-MOT, an efficient 3D MOT method based on the Tracking-By-Detection framework that enables the tracker to choose the most appropriate tracking criteria for each object category. Specifically, Poly-MOT leverages different motion models for various object categories to characterize distinct types of motion accurately. We also introduce the constraint of the rigid structure of objects into a specific motion model to accurately describe the highly nonlinear motion of the object. Additionally, we introduce a two-stage data association strategy to ensure that objects can find the optimal similarity metric from three custom metrics for their categories and reduce missing matches. On the NuScenes dataset, our proposed method achieves state-of-the-art performance with 75.4% AMOTA. The code is available at https://github.com/lixiaoyu20001P0Iy-MOT. Tao Xie 0010, Dedong Liu, Jinghan Gao, Lijun Zhao 0003, Ke Wang 0028 |
IROS | 2 |
| 2023 | Point-NAS: A Novel Neural Architecture Search Framework for Point Cloud AnalysisabstractRecently, point-based networks have exhibited extraordinary potential for 3D point cloud processing. However, owing to the meticulous design of both parameters and hyperparameters inside the network, constructing a promising network for each point cloud task can be an expensive endeavor. In this work, we develop a novel one-shot search framework called Point-NAS to automatically determine optimum architectures for various point cloud tasks. Specifically, we design an elastic feature extraction (EFE) module that serves as a basic unit for architecture search, which expands seamlessly alongside both the width and depth of the network for efficient feature extraction. Based on the EFE module, we devise a searching space, which is encoded into a supernet to provide a wide number of latent network structures for a particular point cloud task. To fully optimize the weights of the supernet, we propose a weight coupling sandwich rule that samples the largest, smallest, and multiple medium models at each iteration and fuses their gradients to update the supernet. Furthermore, we present a united gradient adjustment algorithm that mitigates gradient conflict induced by distinct gradient directions of sampled models and supernet, thus expediting the convergence of the supernet and assuring that it can be comprehensively trained. Pursuant to the provided techniques, the trained supernet enables a multitude of subnets to be incredibly well-optimized. Finally, we conduct an evolutionary search for the supernet under resource constraints to find promising architectures for different tasks. Experimentally, the searched Point-NAS with weights inherited from the supernet realizes outstanding results across a variety of benchmarks. i.e., 94.2% and 88.9% overall accuracy under ModelNet40 and ScanObjectNN, 68.6% mIoU under S3DIS, 63.6% and 69.3% [email protected] under SUN RGB-D and ScanNet V2 datasets. Tao Xie 0010, Linqi Yang, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003 |
IEEE Trans. Image Process. | 1 |