Dedong Liu

dblp:55/7578 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0009-4100-9590ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 57% Video understanding and tracking · 18% Autonomous driving · 12%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d object detection
1.922026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection · IEEE Trans. Multim. 2025
Computer vision › Video understanding and tracking › object tracking
3d object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Robotics › Autonomous driving › perception
3d perception
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › Video understanding and tracking
multi-object tracking
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision
spatio-temporal alignment
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision › 3d object detection
temporal 3d detection
1.012026
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception · AAAI 2026
Computer vision › 3D vision › 3d object detection
indoor 3d object detection
0.912025
SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection · IEEE Trans. Multim. 2025
Machine learning › Transfer learning and domain adaptation › cross-domain learning
multi-domain learning
0.912025
SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection · IEEE Trans. Multim. 2025
Computer vision › 3D vision
camera geometry and pose estimation
0.712023
OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes · ICCV 2023
Machine learning › Learning paradigms
multi-task learning
0.712023
OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes · ICCV 2023
Computer vision › 3D vision
visual localization
0.712023
OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes · ICCV 2023
Computer vision › 3D vision
point cloud analysis
0.312025
SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection · IEEE Trans. Multim. 2025

Methods — techniques the papers use, named apart from their topics

multi-hypothesis decoding · 1.0motion model · 1.0attention mechanism · 1.0set abstraction · 0.9elementwise parameter sharing · 0.9sparse penalty · 0.7layer-adaptive sharing · 0.7gradient normalization · 0.7
YearPublicationVenuePosition
2026 Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
abstract
Spatio-temporal alignment is crucial for temporal modeling of end-to-end (E2E) perception in autonomous driving (AD), providing valuable structural and textural prior information. Existing methods typically rely on the attention mechanism to align objects across frames, simplifying the motion model with a unified explicit physical model (constant velocity, etc.). These approaches prefer semantic features for implicit alignment, challenging the importance of explicit motion modeling in the traditional perception paradigm. However, variations in motion states and object features across categories and frames render this alignment suboptimal. To address this, we propose HAT, a spatio-temporal alignment module that allows each object to adaptively decode the optimal alignment proposal from multiple hypotheses without direct supervision. Specifically, HAT first utilizes multiple explicit motion models to generate spatial anchors and motion-aware feature proposals for historical instances. It then performs multi-hypothesis decoding by incorporating semantic and motion cues embedded in cached object queries, ultimately providing the optimal alignment proposal for the target frame. On nuScenes, HAT consistently improves 3D temporal detectors and trackers across diverse baselines. It achieves state-of-the-art tracking results with 46.0% AMOTA on the test set when paired with the DETR3D detector. In an object-centric E2E AD method, HAT enhances perception accuracy (+1.3% mAP, +3.1% AMOTA) and reduces the collision rate by 32%. When semantics are corrupted (nuScenes-C), the enhancement of motion modeling by HAT enables more robust perception and planning in the E2E AD.
Peidong Li, Dedong Liu, Jiajia Fu, Dixiao Cui, Lijun Zhao 0003, Lining Sun
AAAI5
2025 SOFW: A Synergistic Optimization Framework for Indoor 3D Object Detection
abstract
In this work, we observe that indoor 3D object detection across varied scene domains encompasses both universal attributes and specific features. Based on this insight, we propose SOFW, a synergistic optimization framework that investigates the feasibility of optimizing 3D object detection tasks concurrently spanning several dataset domains. The core of SOFW is identifying domain-shared parameters to encode universal scene attributes, while employing domain-specific parameters to delve into the particularities of each scene domain. Technically, we introduce a set abstraction alteration strategy (SAAS) that embeds learnable domain-specific features into set abstraction layers, thus empowering the network with a refined comprehension for each scene domain. Besides, we develop an elementwise sharing strategy (ESS) to facilitate fine-grained adaptive discernment between domain-shared and domain-specific parameters for network layers. Benefited from the proposed techniques, SOFW crafts feature representations for each scene domain by learning domain-specific parameters, whilst encoding generic attributes and contextual interdependencies via domain-shared parameters. Built upon the classical detection framework VoteNet without any complicated modules, SOFW delivers impressive performances under multiple benchmarks with much fewer total storage footprint. Additionally, we demonstrate that the proposed ESS is a universal strategy and applying it to a voxels-based approach TR3D can realize cutting-edge detection accuracy on all S3DIS, ScanNet, and SUN RGB-D datasets. The source code is available at https://github.com/mooncake199809/SOFW
Tao Xie 0010, Ke Wang 0028, Dedong Liu, Zhendong Fan, Ruifeng Li 0001, Lijun Zhao 0003, Mohamed Omar
IEEE Trans. Multim.5
2025 HVLF: A Holistic Visual Localization Framework Across Diverse Scenes
abstract
Recently, integrating the multitask learning (MTL) paradigm into scene coordinate regression (SCoRe) techniques has achieved significant success in visual localization tasks. However, the feature extraction ability of existing frameworks is inherently constrained by the rigid weight activation strategy, which prevents each layer from concurrently capturing scene-universal features across diverse scenes and scene-particular attributes unique to each individual scene. In addition, the straightforward network architecture further exacerbates the issue of insufficient feature representation. To address these limitations, we introduce HVLF, a holistic framework that ensures flexible identification of both scene-universal and scene-particular attributes while integrating various attention mechanisms to enhance feature representation effectively. Technically, for the first issue, HVLF proposes a soft weight activation strategy (SWAS) equipped with polyhedral convolution to concurrently optimize scene-shared and scene-specific weights within each layer, which facilitates sufficient discernment of both scene-universal features and scene-particular attributes, thereby boosting the network's capability for comprehensive scene perception. For the second issue, HVLF introduces a mixed attention perception module (MAPM) that incorporates channelwise, spatialwise, and elementwise attention mechanisms to perform multilevel feature fusion, hence extracting discriminative features to regress precise scene coordinates. Extensive experiments on indoor and outdoor datasets prove that HVLF realizes impressive localization performance. In addition, experiments conducted on 3-D object detection and feature matching tasks prove that the two proposed techniques are universal and can be seamlessly inserted into other methods.
Fuyuan Qiu, Dedong Liu, Tao Xie 0010, Ke Wang 0028, Ruifeng Li 0001, Lijun Zhao 0003
IEEE Trans. Neural Networks Learn. Syst.4
2023 OFVL-MS: Once for Visual Localization across Multiple Indoor Scenes
abstract
In this work, we seek to predict camera poses across scenes with a multi-task learning manner, where we view the localization of each scene as a new task. We propose OFVL-MS, a unified framework that dispenses with the traditional practice of training a model for each individual scene and relieves gradient conflict induced by optimizing multiple scenes collectively, enabling efficient storage yet precise visual localization for all scenes. Technically, in the forward pass of OFVL-MS, we design a layer-adaptive sharing policy with a learnable score for each layer to automatically determine whether the layer is shared or not. Such sharing policy empowers us to acquire task-shared parameters for a reduction of storage cost and task-specific parameters for learning scene-related features to alleviate gradient conflict. In the backward pass of OFVL-MS, we introduce a gradient normalization algorithm that homogenizes the gradient magnitude of the task-shared parameters so that all tasks converge at the same pace. Furthermore, a sparse penalty loss is applied on the learnable scores to facilitate parameter sharing for all tasks without performance degradation. We conduct comprehensive experiments on multiple benchmarks and our new released indoor dataset LIVL, showing that OFVL-MS families significantly outperform the state-of-the-arts with fewer parameters. We also verify that OFVL-MS can generalize to a new scene with much few parameters while gaining superior localization performance. The dataset and evaluation code is available at https://github.com/mooncake199809/UFVL-Net.
Tao Xie 0010, Siyi Lu, Ke Wang 0028, Jinghan Gao, Dedong Liu, Jie Xu 0066, Lijun Zhao 0003, Ruifeng Li 0001
ICCV7
2023 Poly-MOT: A Polyhedral Framework For 3D Multi-Object Tracking
abstract
3D Multi-object tracking (MOT) empowers mobile robots to accomplish well-informed motion planning and navigation tasks by providing motion trajectories of surrounding objects. However, existing 3D MOT methods typically employ a single similarity metric and physical model to perform data association and state estimation for all objects. With large-scale modern datasets and real scenes, there are a variety of object categories that commonly exhibit distinctive geometric properties and motion patterns. In this way, such distinctions would enable various object categories to behave differently under the same standard, resulting in erroneous matches between trajectories and detections, and jeopardizing the reliability of downstream tasks (navigation, etc.). Towards this end, we propose Poly-MOT, an efficient 3D MOT method based on the Tracking-By-Detection framework that enables the tracker to choose the most appropriate tracking criteria for each object category. Specifically, Poly-MOT leverages different motion models for various object categories to characterize distinct types of motion accurately. We also introduce the constraint of the rigid structure of objects into a specific motion model to accurately describe the highly nonlinear motion of the object. Additionally, we introduce a two-stage data association strategy to ensure that objects can find the optimal similarity metric from three custom metrics for their categories and reduce missing matches. On the NuScenes dataset, our proposed method achieves state-of-the-art performance with 75.4% AMOTA. The code is available at https://github.com/lixiaoyu20001P0Iy-MOT.
Tao Xie 0010, Dedong Liu, Jinghan Gao, Lijun Zhao 0003, Ke Wang 0028
IROS3
2009 Division-based rainfall-runoff simulations with BP neural networks and Xinanjiang model
Qin Ju, Zhongbo Yu, Zhenchun Hao, Gengxin Ou, Dedong Liu
Neurocomputing6