EDBT 2026 Demo / reviewers in the wild / expert
Ding Ma 0001
dblp:30/6718-1
· DBLP profile ↗
19ranked-venue papers
13as first author
13since 2021 · last 2026
0000-0002-0683-9523ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cluster-filter-based pseudo-label refinement for source-free domain adaptation fundus image segmentation
Yanqin Zhang, Ding Ma 0001, Xiangqian Wu 0002 |
Pattern Recognit. | 2 |
| 2026 | EyeKey: Self-Supervised Keypoint Detection and Description Network Based on Local Feature Saliency for Retinal Image Global RegistrationabstractRetinal image registration (RIR) plays an important role in the diagnosis and long-term monitoring of retinal diseases. Retinal image global registration (RIGR) is usually the first step of RIR. Traditional methods often struggle to achieve robust keypoint detection and description when faced with high-resolution, fine-textured retinal images. Deep learning-based methods for this task have not been widely developed. Therefore, we propose a keypoint detection and description network based on local feature saliency, EyeKey, for RIGR. EyeKey uses the "Detect While Describing (DWD)" design. Specifically, two proposed UDPAM++ modules are embedded into the feature description network to enhance its feature description capability. Concurrently, these modules detect distinctive keypoints based on local feature saliency, combined with a Mapping Module featuring only three learnable parameters. Moreover, we achieve self-supervised feature description network training on high-resolution, fine-textured retinal images through the Random Local Hardest Example Mining strategy. Additionally, we realize robust unsupervised keypoint detection network training based on the High Matching Probability Defines Keypoints strategy and the proposed Cumulative Salient Keypoint Expansion, which, together with the DWD design, mutually reinforce the training of the keypoint detection and description network. Finally, combined with the feature-based RIGR pipeline, our method achieves outstanding performance while maintaining excellent inference speed on monomodal and multimodal RIGR evaluation datasets. Yanchao Liang, Ding Ma 0001, Xiangqian Wu 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | Self-Chained Dynamic Context Perception to Tracking by Natural Language SpecificationabstractVision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned with the aid of pre-trained image-text representations. Although such representations provide strong static semantic cues, they are often less effective in explicitly modeling target-state changes described by action-related phrases in natural language queries. As a result, dynamic linguistic cues, such as verbs and motion-related descriptions, may be insufficiently emphasized during cross-modal matching. To address this issue, we propose Self-Chained Dynamic Context Perception (SeDCP), a self-chained framework for explicit dynamic query modulation and language-guided visual refinement in TNL. Specifically, SeDCP consists of two coupled chains. First, the Forward Chain performs visual-evidence-guided dynamic query modulation by injecting trajectory-aware spatiotemporal cues into the language representation, thereby enhancing phrases that describe target-state changes. Second, the Backward Chain uses the dynamically enhanced query representation to refine visual spatiotemporal features, strengthening the alignment between language cues and target-state evolution. In addition, we introduce sequence-level matching rather than isolated pairwise matching to better exploit temporal dynamics, and design a Global-Local enhanced video Transformer to capture both long-range contextual dependencies and fine-grained target details. Extensive experiments on seven standard TNL benchmarks and an additional unseen $\mathrm {LaSOT}_{\mathrm {ext}}$ benchmark demonstrate that SeDCP consistently outperforms state-of-the-art methods and generalizes well to unseen categories and video characteristics. Ding Ma 0001, Zexu Zhang, Xiangqian Wu 0002 |
IEEE Trans. Image Process. | 1 |
| 2025 | Source-free domain adaptation framework based on confidence constrained mean teacher for fundus image segmentation
Yanqin Zhang, Ding Ma 0001, Xiangqian Wu 0002 |
Neurocomputing | 2 |
| 2023 | Tracking by Natural Language Specification with Long Short-term Context DecouplingabstractThe main challenge of Tracking by Natural Language Specification (TNL) is to predict the movement of the target object by giving two heterogeneous information, e.g., one is the static description of the main characteristics of a video contained in the textual query, i.e., long-term context; the other one is an image patch containing the object and its surroundings cropped from the current frame, i.e., the search area. Currently, most methods still struggle with the rationality of using those two information and simply fusing the two. However, the linguistic information contained in the textual query and the visual representation stored in the search area may sometimes be inconsistent, in which case the direct fusion of the two may lead to conflicts. To address this problem, we propose DecoupleTNL, introducing a video clip containing short-term context information into the framework of TNL and exploring a proper way to reduce the impact when visual representation is inconsistent with linguistic information. Concretely, we design two jointly optimized tasks, i.e., short-term context-matching and long-term context-perceiving. The context-matching task aims to gather the dynamic short-term context information in a period, while the context-perceiving task tends to extract the static long-term context information. After that, we design a long short-term modulation module to integrate both context information for accurate tracking. Extensive experiments have been conducted on three tracking benchmark datasets to demonstrate the superiority of DecoupleTNL. Ding Ma 0001, Xiangqian Wu 0002 |
ICCV | 1 |
| 2023 | Enhancing Feature Representation for Anomaly Detection via Local-and-Global Temporal Relations and a Multi-stage Memory
Ding Ma 0001, Xiangqian Wu 0002 |
PRCV (6) | 2 |
| 2023 | Feature Refinement from Multiple Perspectives for High Performance Salient Object Detection
Congao Wang, Ding Ma 0001, Xiangqian Wu 0002 |
PRCV (12) | 3 |
| 2023 | Capsule-Based Regression Tracking via Background InpaintingabstractBackground cues play an accompanying role in most regression trackers, where they directly learn a mapping from dense sampling to soft label by giving a search area. In essence, the trackers need to identify a large amount of background information (i.e., other objects and distractor objects) under the circumstance of extreme target-background data imbalance. Therefore, we believe that it is more worth performing regression tracking depending on the informative background cues and using target cues as supplementary. To do this, we propose a capsule-based approach, referred to as CapsuleBI, which performs regression tracking based on a background inpainting network and a target-aware network. The background inpainting network explores the background representations by restoring the region of the target with all available scenes, and a target-aware network captures the target representations by focusing on the target itself only. To explore the subjects/distractors in the whole scene, we propose a global-guided feature construction module, which helps enhance the local features with global information. Both the background and target are encoded in capsules, which can model the relationships between objects or object parts in the background scene. Apart from this, the target-aware network assists the background inpainting network with a novel background-target routing algorithm that guides the background and target capsules to estimate the target location with multi-video relationships information precisely. Extensive experimental results show that the proposed tracker achieves favorably against state-of-the-art methods. Ding Ma 0001, Xiangqian Wu 0002 |
IEEE Trans. Image Process. | 1 |
| 2022 | QuadTreeCapsule: QuadTree Capsules for Deep Regression TrackingabstractBenefit from the capability of capturing part-to-whole relationships, Capsule Network has been successful in many vision tasks. However, their high computational complexity poses a significant obstacle to applying them to visual tracking, requiring fast inference. In this paper, we introduce the idea of QuadTree Capsules, which explores the property of part-to-whole relationships endowed by the Capsule Network by significantly reducing the computational complexity. We build capsule pyramids and select meaningful relationships in a coarse-to-fine manner, dubbed as QuadTreeCapsule. Specifically, the top K capsules with the highest activation values are selected, and routing is only calculated within the relevant regions corresponding to these top K capsules with a novel symmetric guided routing algorithm. Additionally, considering the importance of temporal relationships, a multi-spectral pose matrix attention mechanism is developed for more accurate spatio-temporal capsule assignments between two sets of capsules. Moreover, during online inference, we shift part of the spatio-temporal capsules long the temporal dimension, facilitating information exchanged among neighboring frames. Extensive experimentation has proved the effectiveness of our methodology, which achieves state-of-the-art results compared with other tracking methods on eight widely-used benchmarks. Our tracker runs at approximately 43 fps on GPU. Ding Ma 0001, Xiangqian Wu 0002 |
ACM Multimedia | 1 |
| 2021 | CapsuleRRT: Relationships-Aware Regression Tracking via CapsulesabstractRegression tracking has gained more and more attention thanks to its easy-to-implement characteristics, while existing regression trackers rarely consider the relationships between the object parts and the complete object. This would ultimately result in drift from the target object when missing some parts of the target object. Recently, Capsule Network (CapsNet) has shown promising results for image classification benefits from its part-object relationships mechanism, while CapsNet is known for its high computational demand even when carrying out simple tasks. Therefore, a primitive adaptation of CapsNet to regression tracking does not make sense, since this will seriously affect speed of a tracker. To solve these problems, we first explore the spatial-temporal relationships endowed by the CapsNet for regression tracking. The entire regression framework, dubbed CapsuleRRT, consists of three parts. One is S-Caps, which captures the spatial relationships between the parts and the object. Meanwhile, a T-Caps module is designed to exploit the temporal relationships within the target. The response of the target is obtained by STCaps Learning. Further, a prior-guided capsule routing algorithm is proposed to generate more accurate capsule assignments for subsequent frames. Apart from this, the heavy computation burden in CapsNet is addressed with a knowledge distillation pose matrix compression strategy that exploits more tight and discriminative representation with few samples. Extensive experimental results show that CapsuleRRT performs favorably against state-of-the-art methods in terms of accuracy and speed. Ding Ma 0001, Xiangqian Wu 0002 |
CVPR | 1 |
| 2021 | Capsule-based Object Tracking with Natural Language SpecificationabstractTracking with Natural-Language Specification (TNL) is a joint topic of understanding the vision and natural language with a wide range of applications. In previous works, the communication between two heterogeneous features of vision and language is mainly through a simple dynamic convolution. However, the performance of prior works is capped by the difficulty of linguistic variation of natural language in modeling the dynamically changing target and its surroundings. In the meanwhile, natural language and vision are firstly fused and then utilized for tracking, which is hard to model the query-focused context. Query-focused should pay more attention to context modeling to promote the correlation between these two features. To address these issues, we propose a capsule-based network, referred to as CapsuleTNL, which performs regression tracking with natural language query. In the beginning, the visual and textual input is encoded with capsules, which can not only establish the relationship between entities but also the relationship between the parts of the entity itself. Then, we devise two interaction routing modules, which consist of visual-textual routing module to reduce the linguistic variation of input query and textual-visual routing module to precisely incorporate query-based visual cues simultaneously. To validate the potential of the proposed network for visual object tracking, we evaluate our method on two large tracking benchmarks. The experimental evaluation demonstrates the effectiveness of our capsule-based network. Ding Ma 0001, Xiangqian Wu 0002 |
ACM Multimedia | 1 |
| 2021 | Conditioners for Adaptive Regression Tracking
Ding Ma 0001, Xiangqian Wu 0002 |
PRCV (1) | 1 |
| 2021 | Distillation-Based Multi-exit Fully Convolutional Network for Visual Tracking
Ding Ma 0001, Xiangqian Wu 0002 |
PRCV (1) | 1 |
| 2020 | FurcaNeXt: End-to-End Monaural Speech Separation with Dynamic Gated Dilated Temporal Convolutional Networks
Liwen Zhang 0001, Ziqiang Shi, Jiqing Han 0001, Anyan Shi, Ding Ma 0001 |
MMM (1) | 5 |
| 2019 | High Speed Recurrent Regression Network for Visual TrackingabstractFor some recently released trackers, the spatial-temporal information of the target are processed separately, which is time consuming and inefficient for locating the target in the sequential data-videos. To solve this problem, we present a recurrent regression framework(RRNet), which leverages spatial and temporal information coherence on feature level simultaneously. The RRNet is composed of a regression network and a long short term memory network(LSTM). The regression network is learned on static image level for focusing on the spatial information of the target, and the whole framework is fine-tuned on videos by fixing the parameters of the regression network, which improves the per-frame regression by aggregation of recurrent prior. Especially, there is no need to online training for adapting the unseen targets. And the experimental results show that the proposed RRNet gets better performance than the compared trackers with a high speed (45 fps). Ding Ma 0001, Xiangqian Wu 0002 |
ICME | 1 |
| 2018 | Multi-Scale Recurrent Tracking via Pyramid Recurrent Network and Optical Flow
Ding Ma 0001, Wei Bu, Xiangqian Wu 0002 |
BMVC | 1 |
| 2018 | Dual-SVM tracker via Multiple Support Instance and LEVER StrategyabstractVisual tracking can be modeled as a binary classification problem, and the classic classifiersupport vector machine (SVM) based methods have been demonstrated encouraging performance in recent object tracking benchmarks. However, the performance of SVM is too sensitive to noisy training data during online update. In this paper, we propose an efficient dual-SVM based tracker to improve classification performance for visual tracking. The tracker proposed consists of two models: the holistic model and the part model. To learn the holistic model, the support instances are derived from the RMI-SVM trained in a deep feature space. As for the part model to highlight local structure of the target, a linear SVM is learned to further encode local details of the target, selecting candidate instances from the support instances by the confidence as input. To fuse the holistic model and the part model, we design a simple but efficient decision strategy (LEVER) to enforce the dual-SVM to focus on the target. The proposed LEVER is updated incrementally to capture changes of the appearance of the target. Extensive experimental results show that the proposed tracker performs favorably against state-of-the-art methods. Ding Ma 0001, Wei Bu, Xiangqian Wu 0002 |
ICPR | 1 |
| 2018 | Learning Collaborative Model for Visual TrackingabstractThis paper proposes a robust visual tracking method by designing a collaborative model. The collaborative model employs a two-stage tracker and a HOG-based detector, which exploits both holistic and local information of the target. The two-stage tracker learns a linear classifier from the patches of original images and the HOG-based detector trains a linear discriminant analysis classifier with the object exemplar. Finally, a result decision making strategy is developed by considering both the original template and the appearance variations, making the tracker and the detector collaborate with each other. The proposed method has been evaluated on OTB-50, OTB-100 and Temple-Color datasets, and results demonstrate that the proposed method is able to effectively address the challenging cases such as scale variation and out-of-view and gets better performance than the state-of-the-art trackers. Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002 |
ICPR | 1 |
| 2018 | Segmentation-Guided Tracking with Prior Map DecisionabstractFor visual tracking, the target object is represented by an appearance model and the location of the target is estimated in each frame. Numerous tracking algorithms model the appearance of the target with a confidence score and rarely take into account the semantic information of the target. In this paper, we propose an efficient tracking algorithm that models the appearance of the target based on semantic segmentation. The overall architecture consists of two parts: the segmentation part and the tracking part. In the segmentation part, an attention model is employed, providing spatial highlights of the candidate region of the target. In the tracking part, the tracker is constructed by an online updated convolutional neural networks to identify the target in subsequent frames, taking advantage of the segmentation information of the target from the segmentation part. To enhance the performance of this architecture, we design an incremental updated prior map taking both the segmentation signal and the tracking signal into consideration. Extensive experiments on two benchmarks including OTB-50, OTB-100, and Temple-Color, show that the proposed method outperforms other trackers. Ding Ma 0001, Wei Bu, Yuehua Cui, Xiangqian Wu 0002 |
ICPR | 1 |