VLDB 2026 Research / reviewers in the wild / expert
Tianpeng Liu
dblp:169/0649
· DBLP profile ↗
30ranked-venue papers
7as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpinTrack: Cyclic token permutation for transformer-based tracking
Tianpeng Liu, Lezhi Lian |
Neurocomputing | 1 |
| 2026 | Self-Supervised SAR Despeckling by Integrating Denoiser Prior With Attributed Scattering CenterabstractDeep learning has become the mainstream approach for SAR image despeckling. However, the lack of speckle-free SAR images poses a significant challenge for supervised deep learning methods. To overcome this limitation, we propose a self-supervised SAR despeckling method that leverages a denoiser prior and attributed scattering centers to enhance the training process. Specifically, we use the output of an external denoiser as a pseudo-label for despeckling, while spatially correlated speckle noise in SAR images is decorrelated through random downsampling. The network is then updated by optimizing the similarity between its output and the pseudo-label. Additionally, an attributed scattering center map is introduced to help the network recognize strong scatterers and better preserve image details. Experiments on both synthetic and real SAR datasets demonstrate that our method outperforms existing despeckling approaches. Specifically, our method improves ENL by 16% over the recent self-supervised MERLIN framework on the TerraSAR scene. Meanwhile, our method achieves MOR and VOR values that are closer to the ideal value of 1.0. The code of our work is made freely available at https://github.com/Duan-NUDT/SARDIDP_ASC. Xiandong Duan, Huangxing Lin, Tianpeng Liu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2026 | Graph-CFRNet: contextual fusion refinement for multi-agent trajectory prediction in autonomous driving
Tianpeng Liu, Huilong Ai |
Multim. Syst. | 2 |
| 2026 | ATRNet-STAR: A Large Dataset and Benchmark Toward Remote Sensing Object Recognition in the WildabstractThe absence of publicly available, large-scale, high-quality datasets for Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) has significantly hindered the application of rapidly advancing deep learning techniques, which hold huge potential to unlock new capabilities in this field. This is primarily because collecting large volumes of diverse target samples from SAR images is prohibitively expensive, largely due to privacy concerns, the characteristics of microwave radar imagery perception, and the need for specialized expertise in data annotation. Throughout the history of SAR ATR research, there have been only a number of small datasets, mainly including targets like ships, airplanes, buildings, etc. There is only one vehicle dataset MSTAR collected in the 1990 s, which has been a valuable source for SAR ATR. To fill this gap, this paper introduces a large-scale, new dataset named ATRNet-STAR with 40 different vehicle categories collected under various realistic imaging conditions and scenes. It marks a substantial advancement in dataset scale and diversity, comprising over 190,000 well-annotated samples-$10\times$ larger than its predecessor, the famous MSTAR. Building such a large dataset is a challenging task, and the data collection scheme will be detailed. Secondly, we illustrate the value of ATRNet-STAR via extensively evaluating the performance of 15 representative methods with 7 different experimental settings on challenging classification and detection benchmarks derived from the dataset. Finally, based on our extensive experiments, we identify valuable insights for SAR ATR and discuss potential future research directions in this field. We hope that the scale, diversity, and benchmark of ATRNet-STAR can significantly facilitate the advancement of SAR ATR. Yongxiang Liu, Li Liu 0002, Jie Zhou 0031, Bowen Peng, Xuying Xiong, Wei Yang 0046, Tianpeng Liu, Zhen Liu 0004, Xiang Li 0014 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | Step-Wise Distribution-Aligned Style Prompt Tuning for Source-Free Cross-Domain Few-Shot LearningabstractExisting cross-domain few-shot learning (CDFSL) methods, which develop training strategies in the source domain to enhance model transferability, face challenges when applied to large-scale pre-trained models (LMs), as their source domains and training strategies are not accessible. Besides, fine-tuning LMs specifically for CDFSL requires substantial computational resources, which limits their practicality. Therefore, this paper investigates the source-free CDFSL (SF-CDFSL) problem to solve the few-shot learning (FSL) task in target domain using only a pre-trained model and a few target samples, without requiring source data or training strategies. However, the inaccessibility of source data prevents explicitly reducing the domain gaps between the source and target. To tackle this challenge, this paper proposes a novel approach, Step-wise Distribution-aligned Style Prompt Tuning (StepSPT), to implicitly narrow the domain gaps from the perspective of prediction distribution optimization. StepSPT initially proposes a style prompt that adjusts the target samples to mirror the expected distribution. Furthermore, StepSPT tunes the style prompt and classifier by exploring a dual-phase optimization process (external and internal processes). In the external process, a step-wise distribution alignment strategy is introduced to tune the proposed style prompt by factorizing the prediction distribution optimization problem into the multi-step distribution alignment problem. In the internal process, the classifier is updated via standard cross-entropy loss. Evaluation on 5 datasets illustrates the superiority of StepSPT over existing prompt tuning-based methods and state-of-the-art methods (SOTAs). Furthermore, ablation studies and performance analyzes highlight the efficacy of StepSPT. Huali Xu, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Shuzhou Sun, Ming-Ming Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Priori-assisted soft actor-critic based interrupted sampling repeater jamming method
Jiaqi Tang 0014, Tianpeng Liu, Weidong Jiang, Dewang Wang, Zhongguo Wu |
Signal Process. | 2 |
| 2026 | Reinforcement Learning-Based Multi-Target Detection Method for MIMO Radar Assisted by Strong Target LimitationabstractUnder the background of co-located MIMO radar, the existing reinforcement learning (RL)-based multi-target detection methods generally perform poorly on weak targets. In our previous work, we have proposed a beam optimization scheme with strong target limitation and gave a solution approach based on multi-rank beamformer, to achieve focusing more radar transmit power on weak targets. In this letter, we further propose a solution approach based on inner convex approximation, which can achieve a higher power gain due to its improved freedom. In addition, we also design an approach for choosing the focused angle cells of radar by fusing the statistical prior information from previous time. Summarizing the above improvements, we propose a RL-based multi-target detection method for MIMO radar assisted by strong target limitation. The experiments show that our method owns better performance on weak targets than its competitors while maintaining the excellent performance on strong targets. Xijie Wu, Tianpeng Liu, Yongxiang Liu, Li Liu 0002 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Policy Generalization Enhancement for UAV Active Object Detection via Divide-and-Conquer Sharpness-Aware Gradient MatchingabstractTarget detection in aerial images captured by unmanned aerial vehicles has long been hampered by occlusion. Active Object Detection (AOD) aims to fundamentally address this issue from the active vision perspective, typically realized through the Deep Reinforcement Learning (DRL) paradigm. However, the active observation policy often suffers from low generalization ability, thus limiting its practical application. In this paper, we propose Divide-and-Conquer Sharpness-Aware Gradient Matching (DC-SAGM), a novel sharpness-based Domain Generalization (DG) method, to effectively enhance the generalization capacity of the agent’s policy. Specifically, we train the agent to learn the active observation policy using the conventional DRL approach. Sharpness-Aware Gradient Matching (SAGM) is employed during training, improving the model’s generalization performance by minimizing the sharpness metric of the loss landscape. Nevertheless, the imperfect state representation and classifier preference in the AOD problem lead to fierce gradient conflicts, deteriorating the effectiveness of SAGM. We address this incompatibility by using a divide-and-conquer strategy and exclude gradient conflicts via the majority-rule gradient surgery operation. Extensive experimental results on the UEVAVD dataset validate DC-SAGM’s superiority in helping the agent’s policy achieve better generalization compared to extensive policy learning approaches. Xinhua Jiang, Tianpeng Liu, Li Liu 0002, Zhenghui Gong, Yongxiang Liu, Xiang Li 0014 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesabstractUnmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality dataset. However, the existing dataset struggles to fully capture real-world complexity for limited imaging conditions. To this end, we introduce a high-diversity dataset ATR-UMOD covering varying scenarios, spanning altitudes from 80m to 300m, angles from 0° to 75°, and all-day, all-year time variations in rich weather and illumination conditions. Moreover, each RGB-IR image pair is annotated with 6 condition attributes, offering valuable high-level contextual information. To meet the challenge raised by such diverse conditions, we propose a novel prompt-guided condition-aware dynamic fusion (PCDF) to adaptively reassign multimodal contributions by leveraging annotated condition cues. By encoding imaging conditions as text prompts, PCDF effectively models the relationship between conditions and multimodal contributions through a task-specific soft-gating transformation. A prompt-guided condition-decoupling module further ensures the availability in practice without condition annotations. Experiments on ATR-UMOD dataset reveal the effectiveness of PCDF. Chen Chen 0152, Kangcheng Bin, Jiahao Qi, Tianpeng Liu, Zhen Liu 0004, Yongxiang Liu, Ping Zhong 0001 |
ICCV | 6 |
| 2025 | TS-BiT: Two-Stage Binary Transformer for ORSI Salient Object DetectionabstractVision transformers (ViTs) have demonstrated superior performance in various remote sensing tasks, such as optical remote sensing image salient object detection (ORSI-SOD). However, the high resolution of remote sensing images and the substantial computational costs pose significant challenges for deploying existing methods on resource-constrained devices. Model binarization significantly reduces computational costs and storage requirements by constraining weights and activations to 1-bit representations, which has been widely explored in convolutional neural networks (CNNs). However, directly applying binary methods to ViTs poses challenges since quantization errors hinder the ability to capture the similarity between tokens, resulting in significant performance degradation in detecting salient objects in complex ORSI scenarios. To address this issue, we propose two-stage binary transformer (TS-BiT) for the ORSI-SOD task to preserve information on salient objects under 1-bit representation. Specifically, we design a two-stage central-aware softmax binarization (TCSB) strategy to reduce quantization errors arising from substantial discrepancies in the long-tail distribution of multihead attention. Furthermore, we develop a scalable hyperbolic tangent function to approximate the gradients of the Sign function within each binarization group, substantially mitigating quantization errors during the binarization of softmax attention. Extensive experiments demonstrate that our method outperforms existing binary ViT approaches on ORSSD, EORSSD, and ORSI-4199 datasets. Tianpeng Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph GenerationabstractExisting two-stage Scene Graph Generation (SGG) frameworks typically incorporate a detector to extract relationship features and a classifier to categorize these relationships; therefore, the training paradigm follows a causal chain structure, where the detector's inputs determine the classifier's inputs, which in turn influence the final predictions. However, such a causal chain structure can yield spurious correlations between the detector's inputs and the final predictions, i.e., the prediction of a certain relationship may be influenced by other relationships. This influence can induce at least two observable biases: tail relationships are predicted as head ones, and foreground relationships are predicted as background ones; notably, the latter bias is seldom discussed in the literature. To address this issue, we propose reconstructing the causal chain structure into a reverse causal structure, wherein the classifier's inputs are treated as the confounder, and both the detector's inputs and the final predictions are viewed as causal variables. Specifically, we term the reconstructed causal paradigm as the Reverse causal Framework for SGG (RcSGG). RcSGG initially employs the proposed Active Reverse Estimation (ARE) to intervene on the confounder to estimate the reverse causality, i.e., the causality from final predictions to the classifier's inputs. Then, the Maximum Information Sampling (MIS) is suggested to enhance the reverse causality estimation further by considering the relationship information. Theoretically, RcSGG can mitigate the spurious correlations inherent in the SGG framework, subsequently eliminating the induced biases. Comprehensive experiments on popular benchmarks and diverse SGG frameworks show the state-of-the-art mean recall rate. Shuzhou Sun, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Ming-Ming Cheng, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | HEART: Historically Information Embedding and Subspace Re-Weighting Transformer-Based TrackingabstractTransformers-based trackers offer significant potential for integrating semantic interdependence between template and search features in tracking tasks. Transformers possess inherent capabilities for processing long sequences and extracting correlations within them. Several researchers have explored the feasibility of incorporating Transformers to model continuously changing search areas in tracking tasks. However, their approach has substantially increased the computational cost of an already resource-intensive Transformer. Additionally, existing Transformers-based trackers rely solely on mechanically employing multi-head attention to obtain representations in different subspaces, without any inherent bias. To address these challenges, we propose HEART (Historical Information Embedding And Subspace Re-weighting Tracker). Our method embeds historical information into the queries in a lightweight and Markovian manner to extract discriminative attention maps for robust tracking. Furthermore, we develop a multi-head attention distribution mechanism to retrieve the most promising subspace weights for tracking tasks. HEART has demonstrated its effectiveness on five datasets, including OTB-100, LaSOT, UAV123, TrackingNet, and GOT-10k. Tianpeng Liu, Jing Li 0055, Amin Beheshti, Jia Wu 0001, Beihang Song, Lezhi Lian |
IEEE Trans. Big Data | 1 |
| 2025 | Observations Temporal Permutation-Based Self-Supervised Reinforcement Learning for UAV Active Object DetectionabstractIn passive ground target detection using Unmanned Aerial Vehicles (UAVs), some detrimental factors like occlusion significantly impact target detection performance. Active Object Detection offers an effective way to address it, which usually uses Deep Reinforcement Learning (DRL) to plan UAV’s viewpoint for favorable observations. However, existing DRL-based AOD methods often suffer from low sample efficiency and poor generalization due to inadequate state representation learned by the policy network. Inspired by human scene understanding where their spatial representation of the scene remains consistent despite different observation orders, we design a self-supervised state representation learning method based on Observations Temporal Permutation (OTP) to improve the state representation of the agent’s policy network. We require the policy network to output consistent action value estimates for observation sequences with the same content but different temporal orders. Besides, we use the state representation to predict the target orientation variations in the observation sequence, which further regularizes and facilitates the state representation learning process. Finally, we design multiple experiments based on the UEVAVD dataset to compare the proposed method with existing self-supervised state representation learning methods for the AOD task. The experimental results demonstrate that the OTP method can help the agent’s policy network learn a better state representation, thus achieving higher policy learning sample efficiency and stronger policy generalization. Xinhua Jiang, Tianpeng Liu, Li Liu 0002, Zhen Liu 0004, Yongxiang Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Heterogeneous Binary Pixel Difference Networks for Remote Sensing Object DetectionabstractRecent research in remote sensing object detection (RSOD) has significantly advanced the development of vision foundation models. However, deploying these models on resource-constrained edge devices is challenging due to their high computational demands. Binarized detectors utilize binary neural networks (BNNs) to achieve extreme compression by quantizing weights and activations to +1 or −1, which have been extensively studied for generic object detection tasks. In remote sensing images, the objects of interest typically exhibit weak responses, and the images often contain numerous unique local areas. Feature binarization in these images can lead to substantial loss of object contrast and scale prior information, which exacerbates performance issues, particularly for small objects, resulting in significant performance degradation. To address these challenges, we propose a novel binarized detector for RSOD named the heterogeneous binary pixel difference network (HBiPiDiNet). Initially, we developed a binary pixel difference convolution (BiPDC) that integrates local binary patterns (LBPs) to capture local contrast information with traditional binary convolution, thereby enhancing the representation of small objects. Subsequently, we constructed heterogeneous kernel fusion convolution blocks (HKFCB) based on BiPDC and standard binary convolution. The HKFCB comprises multiple BiPDCs at different scales, effectively representing BiPDC under multiscale LBP and multiscale binary convolutions. Extensive experiments demonstrate that our proposed method significantly enhances the performance of state-of-the-art binary detection methods across three remote sensing datasets: AI-TOD, VisDrone2019, and DIOR. We have released our code and models athttps://github.com/yuhua666/HBiPiDiNet/tree/main. Jialei Zhan, Liang Bai 0003, Tianpeng Liu, Fan Shi 0003, Yongxiang Liu, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Advancing Segment Anything Model for Efficient Salient Object Detection in Remote Sensing ImagesabstractSalient object detection in optical remote sensing images (ORSI-SOD) often relies on leveraging pre-trained knowledge from natural images to achieve high accuracy with limited training data. Traditional methods typically employ vision backbones (e.g., Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs)) pre-trained on ImageNet to extract features from ORSI scenes. However, these backbones exhibit limited generalization across diverse scenarios compared to recent vision foundation models. To this end, we propose ORSI-SAM, a novel ORSI-SOD framework based on the Segment Anything Model (SAM), leveraging its superior generalization capabilities to achieve an exceptional efficiency-accuracy trade-off. Specifically, ORSI-SAM adopts lightweight SAM as the backbone, effectively reducing parameter size and computational overhead to enable efficient deployment on satellite devices while retaining the rich knowledge learned from large-scale natural image datasets. To mitigate the impact of unavailable prompts in ORSI-SOD on the prediction capability of the SAM decoder, we introduce a Hierarchical Interaction Prompt Generator (HIPG), which aggregates hierarchical features and generates mask prompts tailored for salient objects to guide the decoder in producing high-quality saliency maps. Furthermore, to address the recognition challenges caused by the inherent characteristics of ORSIs, we propose a Semantic-Aware Refinement Decoder (SARD). SARD integrates structural details from low-level features to enrich fine-grained object information while leveraging high-level features to suppress redundant interference in shallow layers, thereby improving the detailed information in the predicted saliency map. ORSI-SAM is the first work to explore the accuracy-efficiency trade-offs for ORSI-SOD based on SAM architecture. Extensive experiments on benchmark datasets show that ORSI-SAM achieves superior performance compared to recent state-of-the-art methods with 12.2M parameters and 8.9G FLOPs. Li Liu 0002, Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Matti Pietikäinen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Boosting Convolutional Neural Networks With Middle Spectrum Grouped ConvolutionabstractThis article proposes a novel module called middle spectrum grouped convolution (MSGC) for efficient deep convolutional neural networks (DCNNs) with the mechanism of grouped convolution. It explores the broad "middle spectrum" area between channel pruning and conventional grouped convolution. Compared with channel pruning, MSGC can retain most of the information from the input feature maps due to the group mechanism; compared with grouped convolution, MSGC benefits from the learnability, the core of channel pruning, for constructing its group topology, leading to better channel division. The middle spectrum area is unfolded along four dimensions: groupwise, layerwise, samplewise, and attentionwise, making it possible to reveal more powerful and interpretable structures. As a result, the proposed module acts as a booster that can reduce the computational cost of the host backbones for general image recognition with even improved predictive accuracy. For example, in the experiments on the ImageNet dataset for image classification, MSGC can reduce the multiply-accumulates (MACs) of ResNet-18 and ResNet-50 by half but still increase the Top-1 accuracy by more than 1%. With a 35% reduction of MACs, MSGC can also increase the Top-1 accuracy of the MobileNetV2 backbone. Results on the MS COCO dataset for object detection show similar observations. Our code and trained models are available at https://github.com/hellozhuo/msgc. Zhuo Su 0002, Tianpeng Liu, Zhen Liu 0004, Shuanghui Zhang, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Unsupervised Pan-Sharpening via Mutually Guided Detail RestorationabstractPan-sharpening is a task that aims to super-resolve the low-resolution multispectral (LRMS) image with the guidance of a corresponding high-resolution panchromatic (PAN) image. The key challenge in pan-sharpening is to accurately modeling the relationship between the MS and PAN images. While supervised deep learning methods are commonly employed to address this task, the unavailability of ground-truth severely limits their effectiveness. In this paper, we propose a mutually guided detail restoration method for unsupervised pan-sharpening. Specifically, we treat pan-sharpening as a blind image deblurring task, in which the blur kernel can be estimated by a CNN. Constrained by the blur kernel, the pan-sharpened image retains spectral information consistent with the LRMS image. Once the pan-sharpened image is obtained, the PAN image is blurred using a pre-defined blur operator. The pan-sharpened image, in turn, is used to guide the detail restoration of the blurred PAN image. By leveraging the mutual guidance between MS and PAN images, the pan-sharpening network can implicitly learn the spatial relationship between the two modalities. Extensive experiments show that the proposed method significantly outperforms existing unsupervised pan-sharpening methods. Huangxing Lin, Xinghao Ding, Tianpeng Liu, Yongxiang Liu |
AAAI | 4 |
| 2024 | An Integrated Network for SA-ISAR Image Processing With Adaptive Denoising and Super-Resolution ModulesabstractThis letter focuses on developing an effective and generalizable deep learning approach for inverse synthetic aperture radar (ISAR) image super-resolution (SR). Since the ISAR imaging process is typically carried out under sparse aperture (SA) conditions, imaging results may exhibit striped noise caused by echoes missing, making it challenging to apply conventional SR methods directly. In view of this, we present a blind SR (BSR) method specifically designed for ISAR images with striped noise. The proposed method employs an integrated network that includes an adaptive denoising module and a SR module (AD-SRNet). Experimental results on both synthetic and real ISAR samples demonstrate the superior performance and strong generalization capability of our approach. Mingyao Chen, Jingyuan Xia, Tianpeng Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | An Autonomous Feature Detection Method of Slow Small Targets on Sea Surface Based on Contextual BanditabstractUnder the background of slow small target detection on sea surface, it is a serious problem that the detection performance of existing feature detection methods decreases when the number of coherent pulses is less. In this letter, we first model the slow small target detection on sea surface as a contextual bandit problem. On this basis, we propose an autonomous feature detection method by modifying the classical feature detection process. The method can autonomously choose the detectors with better performance under current sea scene from the constructed 18 feature detectors, and obtain fine detection performance by fusing their detection results. The performance superiority and the real-time capability of proposed method are verified by the experiments on 7 CSIR datasets. Xijie Wu, Tianpeng Liu, Yongxiang Liu, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | Sparsity-Based Adaptive Beamforming for Coherent Signals With Polarized Sensor ArraysabstractA sparsity-based adaptive beamforming (ABF) method is introduced to effectively process coherent signals with polarized sensor arrays (PSA). This method exploits the spatial sparsity of observed signals by transforming it into row-sparsity within a waveform-polarization composite matrix through data reorganization. This row-sparsity is subsequently cast as an$\ell _{2,1}$norm minimization problem, characterized by a gridless and compact mathematical expression with a Hermitian Toeplitz matrix. Then, a matrix factorization-based gradient descent (GD) algorithm is introduced to effectively resolve this optimization problem. The experimental evaluations demonstrate that the GD algorithm significantly outperforms the MOSEK solver in terms of computational efficiency. Further comparative analysis demonstrates that the proposed method outperforms the existing techniques, especially in contexts of low signal-to-noise ratio (SNR), with a moderate increase in computational runtime. Tianpeng Liu, Junpeng Shi, Zhen Liu 0004, Yongxiang Liu |
IEEE Signal Process. Lett. | 1 |
| 2024 | Tracking With Saliency Region TransformerabstractTransformers show a great impact on visual tracking thanks to their powerful representation learning capabilities. As the capacity of the model grows, the speed of the tracker tends to decrease gradually. Our work focuses on dealing with massively redundant information in tracking sequences with the Saliency Region Tracker (SRTrack). SRTrack is a heuristic two-stage tracker consisting of a lightweight tracking stage and a saliency stage. The former can handle simple tracking sequences while the latter is designed to perform delicate tracking on challenging frames with more discriminative features. However, the two-stage design leads to feature extrapolation, creating inconsistencies between training and inference features. In order to mitigate this problem, we develop an attention scaling factor that guarantees model robustness while yielding a slight performance gain. Our SRTrack achieves a state-of-the-art 0.699 AUC running at 61 FPS on LaSOT. Several experiments on large benchmarks demonstrate the high efficiency and accuracy of SRTrack. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005, Lezhi Lian |
IEEE Trans. Image Process. | 1 |
| 2023 | Variational Information Bottleneck for Cross Domain Object DetectionabstractCross domain object detection leverages a labeled source domain to learn an object detector which performs well in a novel unlabeled target domain. Most existing works mainly align the distribution utilizing the entire image knowledge ignoring the obstacles of task-uncorrelated information to alleviate the domain discrepancy. To tackle this issue, we propose a novel module called Variational Instance Disentanglement (VID) based on information theory which aims to decouple the information of task-correlated while filtering out the task-uncorrelated factors at the instance level. Notably, the proposed VID can be used as a plug-and-play module without bringing extra network parameter cost. We equip it with adversarial network and self-training network forming Variational Instance Disentanglement Adversarial Network (VIDAN) and Variational Instance Disentanglement Self-training Network (VIDSN), respectively. Extensive experiments on multiple widely-used scenarios show that the proposed method improves the performance of the popular frameworks and outperforms state-of-the-art methods. Jiangming Chen, Wanxia Deng, Tianpeng Liu, Yingmei Wei, Li Liu 0002 |
ICME | 4 |
| 2023 | Visual tracking with dumbbell selection network
Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Yafu Xiao, Yan Hong 0005 |
Neurocomputing | 1 |
| 2023 | Redundancy-Reduced Sparsity-Based Adaptive Beamforming for Polarization-Sensitive ArraysabstractA sparse reconstruction approach for adaptive beamforming (ABF) with polarization-sensitive arrays (PSA) is introduced in this letter. It first represents the spatial sparsity of incoming signals as the row sparsity of a power-scaled polarization matrix, which arises from the matrization of the redundancy-reduced covariance vector. Then the row sparsity issue is relaxed to an$\ell _{2,1}$norm minimization form and solved in a gridless way via a compact formulation, where a dimension reduction method is introduced to reduce the problem size. Compared to existing techniques, the proposed method processes the polarization information holistically and derives each signal parameter in the continuous domain. Simulation results substantiate the advantages of the proposed method over competing methods. Tianpeng Liu, Junpeng Shi, Li Liu 0002, Yongxiang Liu |
IEEE Signal Process. Lett. | 2 |
| 2023 | SRDF: Single-Stage Rotate Object Detector via Dense Prediction and False Positive SuppressionabstractOriented object detection has made astonishing progress. However, existing methods neglect to address the issue of false positives caused by the background or nearby clutter objects. Meanwhile, class imbalance and boundary overflow issues caused by the predicting rotation angles may affect the accuracy of rotated bounding box predictions. To address the above issues, we propose a Single-stage Rotate object detector via Dense prediction and False positive suppression (SRDF). Specifically, we design an Instance-level False Positive Suppression Module (IFPSM), IFPSM acquires the weight information of target and non-target regions by supervised learning of spatial feature encoding, and applies these weight values to the deep feature map, thereby attenuating the response signals of non-target regions within the deep feature map. Compared to commonly used attention mechanisms, this approach more accurately suppresses false positive regions. Then, we introduce a hybrid classification and regression method to represent the object orientation, the proposed mothed divide the angle into two segments for prediction, reducing the number of categories and narrowing the range of regression. This alleviates the issue of class imbalance caused by treating one degree as a single category in classification prediction, as well as the problem of boundary overflow caused by directly regressing the angle. In addition, we transform the traditional post-processing steps based on matching and searching to a two-dimensional probability distribution mathematical model, which accurately and quickly extracts the bounding boxes from dense prediction results. Extensive experiments on Remote Sensing, Synthetic Aperture Radar, and Scene Text benchmarks demonstrate the superiority of the proposed SRDF method over state-of-the-art rotated object detection methods. Our codes are available at https://github.com/TomZandJerryZ/SRDF. Beihang Song, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005, Tianpeng Liu |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Tracking With Mutual Attention NetworkabstractVisual tracking is a visual task that tracks a specific target by only giving its first frame location and size. To punish the low-quality but high-scoring tracking results, researchers resorted to foreground reinforcement learning to suppress the scores of positive samples near edges. However, for training with negative samples, all backgrounds are equally labeled as false. In this way, the interdependence and difference between the foreground and the background are not considered. We interpret the underlying reason for drifts as the imbalance between the embedding of background and foreground information. Specifically, some catastrophic tracking results and common tracking errors should not be treated equally but should strengthen the implicit connection between the foreground and background. In this paper, we propose a Mutual Attention (MA) module to strengthen the interdependence between positive and negative samples. It can aggregate the rich contextual interdependence between the target template and the search area, thereby providing an implicit way to update the target template accordingly. As for the difference, we design a background training enhancement (BTE) mechanism to distinguish negative samples with varying degrees of error, that is, to down-weight outrageous and absurd tracking results to improve the robustness of the tracker. The results on a large number of benchmarks indicate the validity of our results, such as OTB-100, VOT-2018, VOT-2019, and LaSOT. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Beihang Song |
IEEE Trans. Multim. | 1 |
| 2020 | Large-Scale Point Cloud Contour Extraction via 3-D-Guided Multiconditional Residual Generative Adversarial NetworkabstractAs one of the most important features for human perception, contours are widely applied in graphics and mapping applications. However, it is considerably challenging to extract contours from large-scale point clouds due to the irregular distribution of point clouds. In this letter, we propose a 3-D-guided multiconditional residual generative adversarial network (3-D-GMRGAN), the first deep-learning framework to generate contours for large-scale outdoor point clouds. To make the network handle huge amounts of points, we operate contours in the parametric space rather than raw point space, associated with a parametric chamfer distance. Then, to gather contour features from potential positions and avoid the huge solution space, we propose a guided residual generative adversarial framework, by utilizing a simple feature-based method to get the “over extraction” potential contour distribution. Experiments demonstrate that the proposed method is able to generate contours efficiently for large-scale point clouds, with fewer outliers and pseudo contours compared with state-of-the-art approaches. Yang Zhang 0036, Zhen Liu 0004, Tianpeng Liu, Xiang Li 0014 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | A survey of event analysis and mining from social multimedia
Tianpeng Liu |
Multim. Tools Appl. | 1 |
| 2019 | Sparse Subspace Clustering for Evolving Data StreamsabstractThe data streams arising in many applications can be modeled as a union of low-dimensional subspaces known as multi-subspace data streams (MSDSs). Clustering MSDSs according to their underlying low-dimensional subspaces is a challenging problem which has not been resolved satisfactorily by existing data stream clustering (DSC) algorithms. In this paper, we propose a sparse-based DSC algorithm, which we refer to as dynamic sparse subspace clustering (D-SSC). This algorithm recovers the low-dimensional subspaces (structures) of high-dimensional data streams and finds an explicit assignment of points to subspaces in an online manner. Moreover, as an online algorithm, D-SSC is able to cope with the time-varying structure of MSDSs. The effectiveness of D-SSC is evaluated using numerical experiments. Jinping Sui, Zhen Liu 0004, Li Liu 0002, Alexander Jung 0001, Tianpeng Liu, Xiang Li 0014 |
ICASSP | 5 |
| 2019 | Social multi-modal event analysis via knowledge-based weighted topic model
Feng Xue 0002, Xueliang Liu, Tianpeng Liu, Qiang Lu 0002 |
J. Vis. Commun. Image Represent. | 4 |