EDBT 2026 Demo / reviewers in the wild / expert
Xingping Dong
dblp:156/3020
· DBLP profile ↗
36ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0003-1613-9288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 22 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware DiffusionabstractRecent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually struggle to produce high-fidelity 3D models, frequently yielding excessively blurred textures. We attribute this issue to the insufficient consideration of cross-view consistency during the diffusion process, resulting in significant disparities between different views and ultimately leading to blurred 3D representations. In this paper, we address this issue by comprehensively exploiting multi-view priors in both the conditioning and diffusion procedures to produce consistent, detail-rich portraits. From the conditioning standpoint, we propose a Hybrid Priors Diffusion model, which explicitly and implicitly incorporates multi-view priors as conditions to enhance the status consistency of the generated multi-view portraits. From the diffusion perspective, considering the significant impact of the diffusion noise distribution on detailed texture generation, we propose a Multi-View Noise Resampling Strategy integrated within the optimization process leveraging cross-view priors to enhance representation consistency. Extensive experiments show that our method produces 3D portraits with accurate geometry and rich details from a single image. Wencheng Han, Xingping Dong, Jianbing Shen |
AAAI | 3 |
| 2026 | Language Interprets Vision: Adaptive Encoding and Decoding for Referring Image SegmentationabstractReferring image segmentation aims to segment the referent with natural linguistic expressions. Due to the distinct modality properties of the image and language, it is challenging to effectively align token embeddings with visual regions. Different from existing methods of coordinate linguistics for the specific visual region, we propose a novel referring image segmentation paradigm, language interprets vision (LIV), which densely fine-grained aligns the visual and linguistic modalities, and fuse the multi-modal biases effectively. LIV resorts to re-encoding visual features on compositional dimensions of, which interprets vision through linguistic expression and makes cross-modality alignment denser. More specifically, we innovatively consider the adjacency of visual regions on the channel level to promote channel semantic consistency and propagate fine-grained semantics in the whole segmentation procedure. In addition, we also theoretically analyze that LIV effectively enriches the representation space and makes the comprehensive modality-fused biases more generalized, which boosts the precision of mask prediction. Extensive experimental results on three benchmarks validate that our proposed framework significantly outperforms other methods by a remarkable margin. Qi A, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
Comput. Vis. Media | 3 |
| 2026 | Condition-Guided Diffusion for Multi-Modal Pedestrian Trajectory Prediction Incorporating Intention and Interaction PriorsabstractPedestrian behavior exhibits inherent multi-modality, necessitating predictions that balance accuracy and diversity to adapt effectively to various complex scenarios. However, conventional noise addition in diffusion models is often aimless and unguided, leading to redundant noise reduction steps and the generation of uncontrollable samples. To address these issues, we propose a Prior Condition-Guided Diffusion Model (CGD-TraP) for multi-modal pedestrian trajectory prediction. Instead of directly adding Gaussian noise to trajectories at each timestep during the forward process, our approach leverages internal intention and external interaction to guide noise estimation. Specifically, we design two specialized modules to extract and aggregate intention and interaction features. These features are then adaptively fused through a spatial-temporal fusion based on selective state space, which estimates a controllable noisy trajectory distribution. By optimizing the noise addition process in a more controlled and efficient manner, our method ensures that the denoising process is effectively guided, resulting in predictions that are both accurate and diverse. Extensive experiments on the ETH-UCY, SDD, and NBA datasets demonstrate that CGD-TraP surpasses state-of-the-art diffusion-based and other generative methods, achieving superior efficiency, accuracy, and diversity. Yanghong Liu, Xingping Dong, Yutian Lin, Mang Ye, Kaihao Zhang, Bo Du 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy PredictionabstractWe present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and fusion strategies. It systematically examines the entire VisionOcc pipeline, identifying three fundamental yet previously overlooked temporal cues: scene-level consistency, motion calibration, and geometric complementation. These cues capture diverse facets of temporal evolution and make distinct contributions across various modules in the VisionOcc framework. To effectively fuse temporal signals across heterogeneous representations, we propose a novel fusion strategy by reinterpreting the formulation of vanilla RNNs. This reinterpretation leverages gradient descent on features to unify the integration of diverse temporal information, seamlessly embedding the proposed temporal cues into the network. Extensive experiments on nuScenes demonstrate that GDFusion significantly outperforms established baselines, achieving 2.2%–4.7% mIoU improvement and reducing memory consumption by 30%–72%. Codes are available at https: //github.com/cdb342/GDFusion. Dubing Chen, Xingping Dong, Xianfei Li, Wenlong Liao, Jianbing Shen |
CVPR | 4 |
| 2025 | Semantic-Aware Pseudo-Labeling for Unsupervised Meta-LearningabstractIn unsupervised meta-learning, the clustering-based pseudo-labeling approach is an attractive framework, since it is model-agnostic, allowing it to synergize with supervised algorithms to learn from unlabeled data. However, the pseudo-labels suffer from clustering noise and semantic chaos problems, further impacting the effectiveness of meta-learning. In this paper, we analyze and optimize the pseudo-labeling process, including encoding and clustering, aiming to generate semantic-like pseudo-labels to narrow the gap between unsupervised and supervised meta-learning. First, during the encoding, we observe that the embedding space of existing methods lacks clustering-friendly properties, which is the primary reason for clustering noise. To address this issue, we minimize the inter-to-intra-class similarity ratio to generate clustering-friendly embedding features and validate our approach through comprehensive experiments. Then, during the clustering, we find that the semantic quality of pseudo-labels is not adequately controlled, resulting in semantic chaos of pseudo-labels. We propose a semantic-stability index to measure the semantic quality of pseudo-labels quantitatively. Based on this index, we propose the Semantic-aware Pseudo-label Reassignment mechanism to generate semantic-like pseudo-labels for all samples. Our approach is model-agnostic and can easily be integrated into existing supervised methods. To demonstrate its generalization ability, we integrate it into two representative algorithms: MAML and EP. The results on three main few-shot benchmarks clearly show that the proposed method achieves significant improvement compared to state-of-the-art models. Notably, our approach also outperforms the corresponding supervised method in three tasks. Tianran Ouyang, Xingping Dong, Mang Ye, Bo Du 0001, Ling Shao 0001, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | A Deep Learning-Based Precipitation Nowcasting Model Fusing GNSS-PWV and Radar Echo ObservationsabstractNowcasting plays a critical role in disaster warning systems, and recent advancements in deep learning have shown great potential in improving the accuracy and timeliness of such predictions. This study proposes a novel deep learning-based model for precipitation nowcasting, which integrates global navigation satellite system (GNSS)-derived precipitable water vapor (PWV) data with radar observations. The model introduces two key innovations: multi-source data fusion and time-dimension attention mechanism. These advancements enhance the model’s capability to accurately forecast precipitation events, particularly under challenging conditions with high rainfall intensity. In comparative experiments conducted using radar and GNSS data from Hong Kong, the model, incorporating both data fusion and the attention mechanism, demonstrated the best overall performance, with critical success index (CSI) scores increasing by 26% and Heidke skill score (HSS) scores by 23% at the 30 mm/h threshold. Moreover, it effectively simulates rainfall regions and their changing trends, demonstrating the complementary value of GNSS PWV data to radar observations. Yidong Lou, Xingping Dong, Xiaohong Zhang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Perception Assisted Transformer for Unsupervised Object Re-IdentificationabstractUnsupervised object re-identification (Re-ID) aims to learn discriminative features without identity annotations. Existing mainstream methods are usually developed based on convolutional neural networks for feature extraction and pseudo-label estimation. However, convolutional neural networks suffer from limitations in capturing dispersed long-range dependencies and integrating global information. In comparison, vision transformers demonstrate superior robustness in complex environments, leveraging their versatile modeling capabilities to process diverse data structures with greater precision. In this paper, we delve into the potential of vision transformers in unsupervised Re-ID, proposing a Transformer-based perception-assisted framework (PAT). Considering Re-ID is a typical fine-grained task, existing unsupervised Re-ID methods relying on pseudo-labels generated by clustering algorithms provide only category-level discriminative supervision, with limited attention to local details. Therefore, we propose a novel target-aware mask alignment (TMA) strategy that provides additional supervision signals by leveraging low-level visual cues. Specifically, we employ pseudo-labels to guide the fine-grained alignment of features with local pixel information from critical discriminative regions. This method establishes a mutual learning mechanism via a shared Transformer, effectively balancing discriminative learning and detailed understanding. Furthermore, we propose a perceptual fusion feature augmentation (PFA) method to optimize instance-level discriminative learning. The proposed method is evaluated on multiple Re-ID datasets, demonstrating superior performance and robustness in comparison to state-of-the-art techniques. Notably, without annotations, our method achieves better results than many supervised counterparts. The code will be released. Shuoyi Chen, Mang Ye, Xingping Dong, Bo Du 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | High-Fidelity and High-Efficiency Talking Portrait Synthesis With Detail-Aware Neural Radiance FieldsabstractIn this paper, we propose a novel rendering framework based on neural radiance fields (NeRF) named HH-NeRF that can generate high-resolution audio-driven talking portrait videos with high fidelity and fast rendering. Specifically, our framework includes a detail-aware NeRF module and an efficient conditional super-resolution module. First, a detail-aware NeRF is proposed to efficiently generate a high-fidelity low-resolution talking head, by using the encoded volume density estimation and audio-eye-aware color calculation. This module can capture natural eye blinks and high-frequency details, and maintain a similar rendering time as previous fast methods. Secondly, we present an efficient conditional super-resolution module on the dynamic scene to directly generate the high-resolution portrait with our low-resolution head. Incorporated with the prior information, such as depth map and audio features, our new proposed efficient conditional super resolution module can adopt a lightweight network to efficiently generate realistic and distinct high-resolution videos. Extensive experiments demonstrate that our method can generate more distinct and fidelity talking portraits on high resolution (900 × 900) videos compared to state-of-the-art methods. Muyu Wang, Sanyuan Zhao, Xingping Dong, Jianbing Shen |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | DifTraj: Diffusion Inspired by Intrinsic Intention and Extrinsic Interaction for Multi-Modal Trajectory Prediction
Yanghong Liu, Xingping Dong, Yutian Lin, Mang Ye |
IJCAI | 2 |
| 2024 | Asymmetric Convolution: An Efficient and Generalized Method to Fuse Feature Maps in Multiple Vision TasksabstractFusing features from different sources is a critical aspect of many computer vision tasks. Existing approaches can be roughly categorized as parameter-free or learnable operations. However, parameter-free modules are limited in their ability to benefit from offline learning, leading to poor performance in some challenging situations. Learnable fusing methods are often space-consuming and time-consuming, particularly when fusing features with different shapes. To address these shortcomings, we conducted an in-depth analysis of the limitations associated with both fusion methods. Based on our findings, we propose a generalized module named Asymmetric Convolution Module (ACM). This module can learn to encode effective priors during offline training and efficiently fuse feature maps with different shapes in specific tasks. Specifically, we propose a mathematically equivalent method for replacing costly convolutions on concatenated features. This method can be widely applied to fuse feature maps across different shapes. Furthermore, distinguished from parameter-free operations that can only fuse two features of the same type, our ACM is general, flexible, and can fuse multiple features of different types. To demonstrate the generality and efficiency of ACM, we integrate it into several state-of-the-art models on three representative vision tasks. Extensive experimental results on three tasks and several datasets demonstrate that our new module can bring significant improvements and noteworthy efficiency. Wencheng Han, Xingping Dong, David Crandall, Cheng-Zhong Xu 0001, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Pseudo-Labeling Based Practical Semi-Supervised Meta-Training for Few-Shot LearningabstractMost existing few-shot learning (FSL) methods require a large amount of labeled data in meta-training, which is a major limit. To reduce the requirement of labels, a semi-supervised meta-training (SSMT) setting has been proposed for FSL, which includes only a few labeled samples and numbers of unlabeled samples in base classes. However, existing methods under this setting require class-aware sample selection from the unlabeled set, which violates the assumption of unlabeled set. In this paper, we propose a practical semi-supervised meta-training setting with truly unlabeled data to facilitate the applications of FSL in realistic scenarios. To better utilize both the labeled and truly unlabeled data, we propose a simple and effective meta-training framework, called pseudo-labeling based meta-learning (PLML). Firstly, we train a classifier via common semi-supervised learning (SSL) and use it to obtain the pseudo-labels of unlabeled data. Then we build few-shot tasks from labeled and pseudo-labeled data and design a novel finetuning method with feature smoothing and noise suppression to better learn the FSL model from noise labels. Surprisingly, through extensive experiments across two FSL datasets, we find that this simple meta-training framework effectively prevents the performance degradation of various FSL models under limited labeled data, and also significantly outperforms the representative SSMT models. Besides, benefiting from meta-training, our method also improves several representative SSL algorithms as well. We provide the training code and usage examples at https://github.com/ouyangtianran/PLML. Xingping Dong, Tianran Ouyang, Shengcai Liao, Bo Du 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Referring Multi-Object TrackingabstractExisting referring understanding tasks tend to involve the detection of a single text-referred object. In this paper, we propose a new and general referring understanding task, termed referring multi-object tracking (RMOT). Its core idea is to employ a language expression as a semantic cue to guide the prediction of multi-object tracking. To the best of our knowledge, it is the first work to achieve an arbitrary number of referent object predictions in videos. To push forward RMOT, we construct one benchmark with scalable expressions based on KITTI, named Refer-KITTI. Specifically, it provides 18 videos with 818 expressions, and each expression in a video is annotated with an average of 10.7 objects. Further, we develop a transformer-based architecture TransRMOT to tackle the new task in an online manner, which achieves impressive detection performance and out-performs other counterparts. The Refer-KITTI dataset and the code are released at https://referringmot.github.io. Dongming Wu 0005, Wencheng Han, Tiancai Wang, Xingping Dong, Xiangyu Zhang 0005, Jianbing Shen |
CVPR | 4 |
| 2023 | Adaptive Siamese Tracking With a Compact Latent NetworkabstractIn this article, we provide an intuitive viewing to simplify the Siamese-based trackers by converting the tracking task to a classification. Under this viewing, we perform an in-depth analysis for them through visual simulations and real tracking examples, and find that the failure cases in some challenging situations can be regarded as the issue of missing decisive samples in offline training. Since the samples in the initial (first) frame contain rich sequence-specific information, we can regard them as the decisive samples to represent the whole sequence. To quickly adapt the base model to new scenes, a compact latent network is presented via fully using these decisive samples. Specifically, we present a statistics-based compact latent feature for fast adjustment by efficiently extracting the sequence-specific information. Furthermore, a new diverse sample mining strategy is designed for training to further improve the discrimination ability of the proposed compact latent network. Finally, a conditional updating strategy is proposed to efficiently update the basic models to handle scene variation during the tracking phase. To evaluate the generalization ability and effectiveness and of our method, we apply it to adjust three classical Siamese-based trackers, namely SiamRPN++, SiamFC, and SiamBAN. Extensive experimental results on six recent datasets demonstrate that all three adjusted trackers obtain the superior performance in terms of the accuracy, while having high running speed. Xingping Dong, Jianbing Shen, Fatih Porikli, Jiebo Luo 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Multi-Level Representation Learning with Semantic Alignment for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-based spatial granularity. The limitation of visual representation is prone to causing vision-language mismatching and producing poor segmentation results. To address this, we propose a novel multi-level representation learning approach, which explores the inherent structure of the video content to provide a set of discriminative visual embedding, enabling more effective vision-language semantic alignment. Specifically, we embed different visual cues in terms of visual granularity, including multi-frame long-temporal information at video level, intra-frame spatial semantics at frame level, and enhanced object-aware feature prior at object level. With the powerful multi-level visual embedding and carefully-designed dynamic alignment, our model can generate a robust representation for accurate video object segmentation. Extensive experiments on Refer-DAVIS17and Refer-YouTube-VOS demonstrate that our model achieves superior performance both in segmentation accuracy and inference speed. Dongming Wu 0005, Xingping Dong, Ling Shao 0001, Jianbing Shen |
CVPR | 2 |
| 2022 | Learning Disentanglement with Decoupled Labels for Vision-Language Navigation
Wenhao Cheng, Xingping Dong, Salman Khan 0001, Jianbing Shen |
ECCV (36) | 2 |
| 2022 | Rethinking Clustering-Based Pseudo-Labeling for Unsupervised Meta-Learning
Xingping Dong, Jianbing Shen, Ling Shao 0001 |
ECCV (20) | 1 |
| 2022 | Distilled Siamese Networks for Visual TrackingabstractIn recent years, Siamese network based trackers have significantly advanced the state-of-the-art in real-time tracking. Despite their success, Siamese trackers tend to suffer from high memory costs, which restrict their applicability to mobile devices with tight memory budgets. To address this issue, we propose a distilled Siamese tracking framework to learn small, fast and accurate trackers (students), which capture critical knowledge from large Siamese trackers (teachers) by a teacher-students knowledge distillation model. This model is intuitively inspired by the one teacher versus multiple students learning method typically employed in schools. In particular, our model contains a single teacher-student distillation module and a student-student knowledge sharing mechanism. The former is designed using a tracking-specific distillation strategy to transfer knowledge from a teacher to students. The latter is utilized for mutual learning between students to enable in-depth knowledge understanding. Extensive empirical evaluations on several popular Siamese trackers demonstrate the generality and effectiveness of our framework. Moreover, the results on five tracking benchmarks show that the proposed distilled trackers achieve compression rates of up to 18× and frame-rates of 265 FPS, while obtaining comparable tracking accuracy compared to base models. Jianbing Shen, Yuanpei Liu, Xingping Dong, Xiankai Lu, Fahad Shahbaz Khan, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Learning To Fuse Asymmetric Feature Maps in Siamese TrackersabstractRecently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCorr has several limitations within Siamese-based tracking: it can easily be fooled by distractors, has fewer activated channels and provides weak discrimination of object boundaries. Further, DW-XCorr is a handcrafted parameter-free module and cannot fully benefit from offline learning on large-scale data.We propose a learnable module, called the asymmetric convolution (ACM), which learns to better capture the se-mantic correlation information in offline training on large-scale data. Different from DW-XCorr and its predecessor (XCorr), which regard a single feature map as the convolution kernel, our ACM decomposes the convolution operation on a concatenated feature map into two mathematically equivalent operations, thereby avoiding the need for the feature maps to be of the same size (width and height) during concatenation. Our ACM can incorporate useful prior information, such as bounding-box size, with standard visual features. Furthermore, ACM can easily be integrated into existing Siamese trackers based on DW-XCorr or XCorr. To demonstrate its generalization ability, we integrate ACM into three representative trackers: SiamFC, SiamRPN++ and SiamBAN. Our experiments reveal the benefits of the proposed ACM, which outperforms existing methods on six tracking benchmarks. On the LaSOT test set, our ACM-based tracker obtains a significant improvement of 5.8% in terms of success (AUC), over the baseline. Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao 0001, Jianbing Shen |
CVPR | 2 |
| 2021 | Dynamical Hyperparameter Optimization via Deep Reinforcement Learning in TrackingabstractHyperparameters are numerical pre-sets whose values are assigned prior to the commencement of a learning process. Selecting appropriate hyperparameters is often critical for achieving satisfactory performance in many vision problems, such as deep learning-based visual object tracking. However, it is often difficult to determine their optimal values, especially if they are specific to each video input. Most hyperparameter optimization algorithms tend to search a generic range and are imposed blindly on all sequences. In this paper, we propose a novel dynamical hyperparameter optimization method that adaptively optimizes hyperparameters for a given sequence using an action-prediction network leveraged on continuous deep Q-learning. Since the observation space for object tracking is significantly more complex than those in traditional control problems, existing continuous deep Q-learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic strategy to handle high dimensional state space, while also accelerating the convergence behavior. The proposed algorithm is applied to improve two representative trackers, a Siamese-based one and a correlation-filter-based one, to evaluate its generalizability. Their superior performances on several popular benchmarks are clearly demonstrated. Our source code is available at https://github.com/shenjianbing/dqltracking. Xingping Dong, Jianbing Shen, Wenguan Wang, Ling Shao 0001, Haibin Ling, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | CLNet: A Compact Latent Network for Fast Adjusting Siamese Trackers
Xingping Dong, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
ECCV (20) | 1 |
| 2020 | Inferring Salient Objects from Human FixationsabstractPrevious research in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this work, we propose to employ the former model type to identify salient objects. We build a novel Attentive Saliency Network (ASNet)11.Available at: https://github.com/wenguanwang/ASNet. that learns to detect salient objects from fixations. The fixation map, derived at the upper network layers, mimics human visual attention mechanisms and captures a high-level understanding of the scene from a global view. Salient object detection is then viewed as fine-grained object-level saliency segmentation and is progressively optimized with the guidance of the fixation map in a top-down manner. ASNet is based on a hierarchy of convLSTMs that offers an efficient recurrent mechanism to sequentially refine the saliency features over multiple steps. Several loss functions, derived from existing saliency evaluation metrics, are incorporated to further boost the performance. Extensive experiments on several challenging datasets show that our ASNet outperforms existing methods and is capable of generating accurate segmentation maps with the help of the computed fixation prior. Our work offers a deeper insight into the mechanisms of attention and narrows the gap between salient object detection and fixation prediction. Wenguan Wang, Jianbing Shen, Xingping Dong, Ali Borji, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Adaptive Nonlocal Random Walks for Image Superpixel SegmentationabstractIn this paper, we propose a novel superpixel segmentation method using an adaptive nonlocal random walk (ANRW) algorithm. There are three main steps in our image superpixel segmentation algorithm. Our method is based on the random walk model, in which the seed points are produced to generate the initial superpixels by a gradient-based method in the first step. In the second step, the ANRW is proposed to get the initial superpixels by adjusting the NRW to obtain a better image and superpixel segmentation. In the last step, these small superpixels are merged to get the final regular and compact superpixels. The experimental results demonstrate that our method achieves a better superpixel performance than the state-of-the-art methods. Our source code will be available at: http://github.com/shenjianbing/ANRW. Jianbing Shen, Junbo Yin, Xingping Dong, Hanqiu Sun, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Visual Object Tracking by Hierarchical Attention Siamese NetworkabstractVisual tracking addresses the problem of localizing an arbitrary target in video according to the annotated bounding box. In this article, we present a novel tracking method by introducing the attention mechanism into the Siamese network to increase its matching discrimination. We propose a new way to compute attention weights to improve matching performance by a sub-Siamese network [Attention Net (A-Net)], which locates attentive parts for solving the searching problem. In addition, features in higher layers can preserve more semantic information while features in lower layers preserve more location information. Thus, in order to solve the tracking failure cases by the higher layer features, we fully utilize location and semantic information by multilevel features and propose a new way to fuse multiscale response maps from each layer to obtain a more accurate position estimation of the object. We further propose a hierarchical attention Siamese network by combining the attention weights and multilayer integration for tracking. Our method is implemented with a pretrained network which can outperform most well-trained Siamese trackers even without any fine-tuning and online updating. The comparison results with the state-of-the-art methods on popular tracking benchmarks show that our method achieves better performance. Our source code and results will be available at https://github.com/shenjianbing/HASN. Jianbing Shen, Xingping Dong, Ling Shao 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Reducing Estimation Bias via Triplet-Average Deep Deterministic Policy GradientabstractThe overestimation caused by function approximation is a well-known property in Q-learning algorithms, especially in single-critic models, which leads to poor performance in practical tasks. However, the opposite property, underestimation, which often occurs in Q-learning methods with double critics, has been largely left untouched. In this article, we investigate the underestimation phenomenon in the recent twin delay deep deterministic actor-critic algorithm and theoretically demonstrate its existence. We also observe that this underestimation bias does indeed hurt performance in various experiments. Considering the opposite properties of single-critic and double-critic methods, we propose a novel triplet-average deep deterministic policy gradient algorithm that takes the weighted action value of three target critics to reduce the estimation bias. Given the connection between estimation bias and approximation error, we suggest averaging previous target values to reduce per-update error and further improve performance. Extensive empirical results over various continuous control tasks in OpenAI gym show that our approach outperforms the state-of-the-art methods. Dongming Wu 0005, Xingping Dong, Jianbing Shen, Steven C. H. Hoi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Quadruplet Network With One-Shot Learning for Fast Visual Object TrackingabstractIn the same vein of discriminative one-shot learning, Siamese networks allow recognizing an object from a single exemplar with the same class label. However, they do not take advantage of the underlying structure of the data and the relationship among the multitude of samples as they only rely on the pairs of instances for training. In this paper, we propose a new quadruplet deep network to examine the potential connections among the training instances, aiming to achieve a more powerful representation. We design a shared network with four branches that receive a multi-tuple of instances as inputs and are connected by a novel loss function consisting of pair loss and triplet loss. According to the similarity metric, we select the most similar and the most dissimilar instances as the positive and negative inputs of triplet loss from each multi-tuple. We show that this scheme improves the training performance. Furthermore, we introduce a new weight layer to automatically select suitable combination weights, which will avoid the conflict between triplet and pair loss leading to worse performance. We evaluate our quadruplet framework by model-free tracking-by-detection of objects from a single initial exemplar in several visual object tracking benchmarks. Our extensive experimental analysis demonstrates that our tracker achieves superior performance with a real-time processing speed of 78 frames/s. Our source code is available. Xingping Dong, Jianbing Shen, Dongming Wu 0005, Kan Guo, Xiaogang Jin 0001, Fatih Porikli |
IEEE Trans. Image Process. | 1 |
| 2019 | Submodular Function Optimization for Motion Clustering and Image SegmentationabstractIn this paper, we propose a framework of maximizing quadratic submodular energy with a knapsack constraint approximately, to solve certain computer vision problems. The proposed submodular maximization problem can be viewed as a generalization of the classic 0/1 knapsack problem. Importantly, maximization of our knapsack constrained submodular energy function can be solved via dynamic programing. We further introduce a range-reduction step prior to dynamic programing as a two-stage procedure for more efficient maximization. In order to demonstrate the effectiveness of the proposed energy function and its maximization algorithm, we apply it to two representative computer vision tasks: image segmentation and motion trajectory clustering. Experimental results of image segmentation demonstrate that our method outperforms the classic segmentation algorithms of graph cuts and random walks. Moreover, our framework achieves better performance than state-of-the-art methods on the motion trajectory clustering task. Jianbing Shen, Xingping Dong, Jianteng Peng, Xiaogang Jin 0001, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Hyperparameter Optimization for Tracking With Continuous Deep Q-LearningabstractHyperparameters are numerical presets whose values are assigned prior to the commencement of the learning process. Selecting appropriate hyperparameters is critical for the accuracy of tracking algorithms, yet it is difficult to determine their optimal values, in particular, adaptive ones for each specific video sequence. Most hyperparameter optimization algorithms depend on searching a generic range and they are imposed blindly on all sequences. Here, we propose a novel hyperparameter optimization method that can find optimal hyperparameters for a given sequence using an action-prediction network leveraged on Continuous Deep Q-Learning. Since the common state-spaces for object tracking tasks are significantly more complex than the ones in traditional control problems, existing Continuous Deep Q-Learning algorithms cannot be directly applied. To overcome this challenge, we introduce an efficient heuristic to accelerate the convergence behavior. We evaluate our method on several tracking benchmarks and demonstrate its superior performance1. Xingping Dong, Jianbing Shen, Wenguan Wang, Yu Liu 0074, Ling Shao 0001, Fatih Porikli |
CVPR | 1 |
| 2018 | Salient Object Detection Driven by Fixation PredictionabstractResearch in visual saliency has been focused on two major types of models namely fixation prediction and salient object detection. The relationship between the two, however, has been less explored. In this paper, we propose to employ the former model type to identify and segment salient objects in scenes. We build a novel neural network called Attentive Saliency Network (ASNet) that learns to detect salient objects from fixation maps. The fixation map, derived at the upper network layers, captures a high-level understanding of the scene. Salient object detection is then viewed as fine-grained object-level saliency segmentation and is progressively optimized with the guidance of the fixation map in a top-down manner. ASNet is based on a hierarchy of convolutional LSTMs (convLSTMs) that offers an efficient recurrent mechanism for sequential refinement of the segmentation map. Several loss functions are introduced for boosting the performance of the ASNet. Extensive experimental evaluation shows that our proposed ASNet is capable of generating accurate segmentation maps with the help of the computed fixation map. Our work offers a deeper insight into the mechanisms of attention and narrows the gap between salient object detection and fixation prediction. Wenguan Wang, Jianbing Shen, Xingping Dong, Ali Borji |
CVPR | 3 |
| 2018 | Triplet Loss in Siamese Network for Object Tracking
Xingping Dong, Jianbing Shen |
ECCV (13) | 1 |
| 2018 | Fast Online Tracking With Detection RefinementabstractMost of the existing multiple object tracking (MOT) methods employ the tracking-by-detection framework. Among them, the min-cost network flow optimization techniques become the most popular and standard ones. In these methods, the graph structure models the MOT problem and finds the optimal flow in a connected graph of detections to encode the accurate track trajectories. However, the existing network flow is not suitable for directly online tracking, where the tracking results depend too much on the initial detections. To solve these problems, we present a fast online MOT algorithm by introducing the minimum output sum of squared error filter. The proposed method can adaptively refine the tracking targets according to the proposed rules of correcting the detection mistakes. Furthermore, we introduce an alternative targets hypotheses to reduce the dependence on detections and adaptively refine the object detection boxes. The experimental results on the MOT 2015 benchmark demonstrate that our method achieves comparable or even better results than previous approaches. Jianbing Shen, Dajiang Yu, Leyao Deng, Xingping Dong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Hierarchical Superpixel-to-Pixel Dense MatchingabstractIn this paper, we propose a novel matching method to establish dense correspondences automatically between two images in a hierarchical superpixel-to-pixel manner. Our method first estimates dense superpixel pairings between the two images in the coarse-grained level to overcome large patch displacements and then utilizes superpixel level pairings to drive the matchings in the pixel level to obtain fine texture details. In order to compensate for the influence of color and illumination variations, we apply a regularization technique to rectify images by a color transfer function. Experimental validation on benchmark data sets demonstrates that our approach achieves better visual quality outperforming the state-of-the-art dense matching algorithms. Xingping Dong, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Higher Order Energies for Image SegmentationabstractA novel energy minimization method for general higher order binary energy functions is proposed in this paper. We first relax a discrete higher order function to a continuous one, and use the Taylor expansion to obtain an approximate lower order function, which is optimized by the quadratic pseudo-Boolean optimization or other discrete optimizers. The minimum solution of this lower order function is then used as a new local point, where we expand the original higher order energy function again. Our algorithm does not restrict to any specific form of the higher order binary function or bring in extra auxiliary variables. For concreteness, we show an application of segmentation with the appearance entropy, which is efficiently solved by our method. Experimental results demonstrate that our method outperforms the state-of-the-art methods. Jianbing Shen, Jianteng Peng, Xingping Dong, Ling Shao 0001, Fatih Porikli |
IEEE Trans. Image Process. | 3 |
| 2017 | Occlusion-Aware Real-Time Object TrackingabstractThe online learning methods are popular for visual tracking because of their robust performance for most video sequences. However, the drifting problem caused by noisy updates is still a challenge for most highly adaptive online classifiers. In visual tracking, target object appearance variation, such as deformation and long-term occlusion, easily causes noisy updates. To overcome this problem, a new real-time occlusion-aware visual tracking algorithm is introduced. First, we learn a novel two-stage classifier with circulant structure with kernel, named integrated circulant structure kernels (ICSK). The first stage is applied for transition estimation and the second is used for scale estimation. The circulant structure makes our algorithm realize fast learning and detection. Then, the ICSK is used to detect the target without occlusion and build a classifier pool to save these classifiers with noisy updates. When the target is in heavy occlusion or after long-term occlusion, we redetect it using an optimal classifier selected from the classifier-pool according to an entropy minimization criterion. Extensive experimental results on the full benchmark demonstrate our real-time algorithm achieves better performance than state-of-the-art methods. Xingping Dong, Jianbing Shen, Dajiang Yu, Wenguan Wang, Jianhong Liu |
IEEE Trans. Multim. | 1 |
| 2016 | Video Supervoxels Using Partially Absorbing Random WalksabstractSupervoxels have been widely used as a preprocessing step to exploit object boundaries to improve the performance of video processing tasks. However, most of the traditional supervoxel algorithms do not perform well in regions with complex textures or weak boundaries. These methods may generate supervoxels with overlapping boundaries. In this paper, we present the novel video supervoxel generation algorithm using partially absorbing random walks to get more accurate supervoxels in these regions. Our spatial-temporal framework is introduced by making full use of the appearance and motion cues, which effectively exploits the temporal consistency in video sequence. Moreover, we build a novel Laplacian optimization structure using two adjacent frames to make our approach more efficient. Experimental results demonstrated that our method achieved better performance than the state-of-the-art supervoxel algorithms. Yuling Liang, Jianbing Shen, Xingping Dong, Hanqiu Sun, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Sub-Markov Random Walk for Image SegmentationabstractA novel sub-Markov random walk (subRW) algorithm with label prior is proposed for seeded image segmentation, which can be interpreted as a traditional random walker on a graph with added auxiliary nodes. Under this explanation, we unify the proposed subRW and other popular random walk (RW) algorithms. This unifying view will make it possible for transferring intrinsic findings between different RW algorithms, and offer new ideas for designing novel RW algorithms by adding or changing auxiliary nodes. To verify the second benefit, we design a new subRW algorithm with label prior to solve the segmentation problem of objects with thin and elongated parts. The experimental results on both synthetic and natural images with twigs demonstrate that the proposed subRW method outperforms previous RW algorithms for seeded image segmentation. Xingping Dong, Jianbing Shen, Ling Shao 0001, Luc Van Gool |
IEEE Trans. Image Process. | 1 |
| 2015 | Interactive Cosegmentation Using Global and Local Energy OptimizationabstractWe propose a novel interactive cosegmentation method using global and local energy optimization. The global energy includes two terms: 1) the global scribbled energy and 2) the interimage energy. The first one utilizes the user scribbles to build the Gaussian mixture model and improve the cosegmentation performance. The second one is a global constraint, which attempts to match the histograms of common objects. To minimize the local energy, we apply the spline regression to learn the smoothness in a local neighborhood. This energy optimization can be converted into a constrained quadratic programming problem. To reduce the computational complexity, we propose an iterative optimization algorithm to decompose this optimization problem into several subproblems. The experimental results show that our method outperforms the state-of-the-art unsupervised cosegmentation and interactive cosegmentation methods on the iCoseg and MSRC benchmark data sets. Xingping Dong, Jianbing Shen, Ling Shao 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 1 |