VLDB 2026 Research / reviewers in the wild / expert
Dongdong Li 0004
dblp:14/5457-4
· DBLP profile ↗
15ranked-venue papers
8as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 67% Video understanding and tracking · 33% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Image and video processing › image restoration
degradation-aware enhancement |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Image and video processing
image fusion |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Image and video processing
image restoration |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Image and video processing › image fusion
multi-modal image fusion |
1.0 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Computer vision › Video understanding and tracking
object tracking |
0.9 | 1 | 2025 | Exploring Efficient and Effective Sequence Learning for Visual Object Tracking · IJCAI 2025 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.9 | 1 | 2025 | Exploring Efficient and Effective Sequence Learning for Visual Object Tracking · IJCAI 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | Exploring Efficient and Effective Sequence Learning for Visual Object Tracking · IJCAI 2025 |
Image and video processing
image enhancement |
0.3 | 1 | 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image Fusion · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
prompt-based encoding · 1.0cross-modal enhancement · 1.0transformer · 0.9early exit · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dream-IF: Dynamic Relative EnhAnceMent for Image FusionabstractImage fusion aims to integrate comprehensive information from images acquired through multiple sources. However, images captured by diverse sensors often encounter various degradations that can negatively affect fusion quality. Traditional fusion methods generally treat image enhancement and fusion as separate processes, overlooking the inherent correlation between them; notably, the dominant regions in one modality of a fused image often indicate areas where the other modality might benefit from enhancement. Inspired by this observation, we introduce the concept of dominant regions for image enhancement and present a Dynamic Relative EnhAnceMent framework for Image Fusion (Dream-IF). This framework quantifies the relative dominance of each modality across different layers and leverages this information to facilitate reciprocal cross-modal enhancement. By integrating the relative dominance derived from image fusion, our approach supports not only image restoration but also a broader range of image enhancement applications. Furthermore, we employ prompt-based encoding to capture degradation-specific details, which dynamically steer the restoration process and promote coordinated enhancement in both multi-modal image fusion and image enhancement scenarios. Extensive experimental results demonstrate that Dream-IF consistently outperforms its counterparts. Xingxin Xu, Bing Cao 0002, Dongdong Li 0004, Qinghua Hu, Pengfei Zhu 0001 |
AAAI | 3 |
| 2025 | Exploring Efficient and Effective Sequence Learning for Visual Object TrackingabstractSequence learning based tracking frameworks are popular in the tracking community. In practice, its auto-regressive sequence generation manner leads to inferior performance and high latency compared with latest advanced trackers. In this paper, to mitigate this issue, we propose an efficient and effective sequence-to-sequence tracking framework named FastSeqTrack. FastSeqTrack differs from previous sequence learning based trackers in terms of token initialization and sequence generation manner. Four tracking tokens are appended to patch embeddings and generated in the encoder as initial guesses for the bounding box sequence, which improves the tracking accuracy compared with randomly initialized tokens. Tracking tokens are then parallelly fed into the decoder in a one-pass manner and greatly boost the forward inference speed compared with the auto-regressive manner. Inspired by the early-exit mechanism, we inject internal classifiers after each decoder layer to early terminate forward inference when the softmax confidence is sufficiently reliable. In easy tracking frames, early exits avoid network overthinking and unnecessary computation. Extensive experiments on multiple benchmarks demonstrate that FastSeqTrack runs over 100 fps and showcases superior performance against state-of-the-art trackers. Codes and models are available at https://github.com/vision4drones/FastSeqTrack. Dongdong Li 0004, Zhinan Gao, Yangliu Kuai |
IJCAI | 1 |
| 2025 | Exploring a Hierarchical Cross-Attention Transformer for High-Speed Tracking
Xin Chen 0032, Ben Kang, Jiawen Zhu 0003, Dongdong Li 0004, Chunjuan Bo, Dong Wang 0004 |
Comput. Vis. Media | 4 |
| 2025 | Visible-Infrared Image Alignment for UAVs: Benchmark and New BaselineabstractWith the extensive use of multisensors in uncrewed aerial vehicles (UAVs), multimodality information processing has become the research focus. In academic research pertaining to object detection and tracking tasks in UAVs, researchers often align visible-infrared image pairs as a preprocessing step. However, in actual tasks, the dual-modality image pair acquired by UAVs is unaligned, which significantly limits the application of downstream tasks. At present, there are no publicly available multimodality image alignment datasets for UAVs. In this article, we present a large-scale benchmark for the dual-modality image alignment task in UAVs, including 81000 training image pairs and 15000 testing image pairs. Meanwhile, we propose a transformer-based dual-modality image alignment network as the baseline for this benchmark. First, the algorithm extracts multiscale features for image representation to address unaligned image pairs with varying resolutions. Second, a transformer-based alignment network is proposed to improve the fusion of features from heterogeneous modalities. Finally, deformable attention is adopted to alleviate the problem of memory explosion. Numerous experiments on this dual-modality image alignment benchmark are conducted to demonstrate the effectiveness of our algorithm. Source codes are available athttps://github.com/gaozhinanjiu/UAVmatch. Zhinan Gao, Dongdong Li 0004, Yangliu Kuai, GongJian Wen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Vision-Based Anti-UAV Detection and TrackingabstractUnmanned aerial vehicles (UAV) have been widely used in various fields, and their invasion of security and privacy has aroused social concern. Several detection and tracking systems for UAVs have been introduced in recent years, but most of them are based on radio frequency, radar, and other media. We assume that the field of computer vision is mature enough to detect and track invading UAVs. Thus we propose a visible light mode dataset called Dalian University of Technology Anti-UAV dataset, DUT Anti-UAV for short. It contains a detection dataset with a total of 10,000 images and a tracking dataset with 20 videos that include short-term and long-term sequences. All frames and images are manually annotated precisely. We use this dataset to train several existing detection algorithms and evaluate the algorithms’ performance. Several tracking methods are also tested on our tracking dataset. Furthermore, we propose a clear and simple tracking algorithm combined with detection that inherits the detector’s high precision. Extensive experiments show that the tracking performance is improved considerably after fusing detection, thus providing a new attempt at UAV tracking using our dataset. The datasets and results are publicly available at:https://github.com/wangdongdut/DUT-Anti-UAV. Jie Zhao 0014, Jingshu Zhang, Dongdong Li 0004, Dong Wang 0004 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Robust visual tracking with channel attention and focal loss
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Lingxiao Zhu, Fatih Porikli |
Neurocomputing | 1 |
| 2020 | When Correlation Filters Meet Siamese Networks for Real-Time Complementary TrackingabstractDiscriminative correlation filter (DCF)-based trackers have recently exhibited high efficiency and impressive robustness to challenging factors, such as illumination change and partial occlusion. However, in cases with fast motion and full occlusion, these trackers drift off soon and can hardly re-detect the target from the restricted search region due to the boundary effect. On the contrary, recent work using a fully convolutional Siamese network (Siamfc) locates the exemplar image within a large search image but suffers from coarse location and distractors. In this paper, we propose a real-time complementary tracker (RCT) by integrating DCF and Siamfc into a two-stage tracking framework where DCF and Siamfc share mutual advantages and complement each other. In the first stage of this framework, RCT locates the target coarsely but robustly with Siamfc. In the second stage, the derived coarse location is refined by DCF for higher accuracy. For efficiency reasons, Siamfc in the first stage is activated occasionally based on the tracking status inferred from the correlation response map of DCF in the second stage. Comprehensive experiments are performed on three popular benchmark datasets: OTB2013, OTB2015, and VOT2016. On OTB2013, RCT runs with over 40 f/s and achieves an absolute gain of 4.8% and 5.2% in mean overlap precision compared with two base trackers (Staple and Siamfc). On VOT2016, RCT makes a good balance between performance and efficiency, ranking fifth in EAO and first in EFO compared with the top five trackers. Dongdong Li 0004, Fatih Porikli, GongJian Wen, Yangliu Kuai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Masked and dynamic Siamese network for robust visual tracking
Yangliu Kuai, GongJian Wen, Dongdong Li 0004 |
Inf. Sci. | 3 |
| 2019 | Learning target-aware correlation filters for visual tracking
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Beyond feature integration: a coarse-to-fine framework for cascade correlation tracking
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
Mach. Vis. Appl. | 1 |
| 2018 | A Rigorous Solution for Closed-form Correlation Filter TrackingabstractRecently, Discriminative Correlation Filters (DCF) have achieved enormous popularity in the tracking community due to high efficiency and fair robustness. With the circular structure, DCF transform computationally consuming spatial correlation into efficient element-wise operation in the Fourier domain. In this paper, we argue that this element-wise solution can be derived only in the case of single-channel features. In terms of tracking with multi-channel features, this element-wise solution trains each feature dimension independently and fails to learn a joint correlation filter. To tackle this problem, we propose a rigorous solution to closed-form correlation filter tracking. This rigorous solution can be computed pixel by pixel from a small linear equation system. Experimental results demonstrate that our rigorous pixel-wise solution achieves better tracking performance than the baseline element-wise solution. Dongdong Li 0004, GongJian Wen, Yangliu Kuai |
ICPR | 1 |
| 2018 | LCO: Lightweight Convolution Operators for fast tracking
Dongdong Li 0004, GongJian Wen, Yangliu Kuai, BingWei Hui |
Image Vis. Comput. | 1 |
| 2018 | Learning adaptively windowed correlation filters for robust tracking
Yangliu Kuai, GongJian Wen, Dongdong Li 0004 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | When correlation filters meet fully-convolutional Siamese networks for distractor-aware tracking
Yangliu Kuai, GongJian Wen, Dongdong Li 0004 |
Signal Process. Image Commun. | 3 |
| 2018 | End-to-End Feature Integration for Correlation Filter Tracking With Channel AttentionabstractRecently, the performance advancement of discriminative correlation filter (DCF) based trackers is predominantly driven by the use of deep convolutional features. As convolutional features from multiple layers capture different target information, existing works integrate hierarchical convolutional features to enhance target representation. However, these works separate feature integration from DCF learning and hardly benefit from end-to-end training. In this letter, we incorporates feature integration and DCF learning in a unified convolutional neural network. This network reformulates feature integration as a differential module that concatenates features from the shallow and deep layers. A channel attention mechanism is introduced to adaptively impose channel-wise weight on the integrated features. Experimental results on OTB100 and UAV123 demonstrate that our method achieves significant performance improvement while running in real-time. Dongdong Li 0004, GongJian Wen, Yangliu Kuai, Fatih Porikli |
IEEE Signal Process. Lett. | 1 |