Shu Tian

dblp:120/8868 · DBLP profile ↗
← Back
28ranked-venue papers
10as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 5 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Joint Feature and Kernel Fusion for Improved Depth-Aware Panoptic Segmentation
abstract
Depth-aware Panoptic Segmentation, which combines panoptic segmentation and monocular depth estimation, is a challenging task that requires a comprehensive understanding of both scene geometry and object semantics. Recent multi-task learning approaches have leveraged dynamic kernel methods to tackle these tasks simultaneously. However, these methods often treat feature extraction and kernel generation for each task in isolation, failing to fully exploit the rich interdependencies between depth and semantic information. To address this, we propose a novel framework with Cross-Task Feature Fusion and Kernel Fusion mechanisms to enhance Depth-aware Panoptic Segmentation. Our approach enables deeper integration of features and kernels, promoting more effective information exchange and mutual reinforcement between tasks. Experiments show that our method brings significant improvements, demonstrating the potential of a more integrated multi-task learning strategy for Depth-aware Panoptic Segmentation.
Shu Tian, Xin Zhao 0012, Xu-Cheng Yin
ICASSP2
2025 CTIDRNet: Cross-Temporal Interaction With Difference Refinement Network for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RSCD) has achieved creditable success in recent years. However, the challenge of identifying changed objects with shape details persists in RSCD. In this letter, we proposed a cross-temporal interaction with difference refinement network (CTIDRNet) to solve interference-caused fake change and incomplete irregular change shape in RSCD tasks. Specifically, by combining cross-attention and self-attention to steer the temporal feature interaction of each input, we design a temporal feature attention (TFA) module to excavate the potential relation of change areas and suppress the unchanged object interference. Afterward, a deformable convolution is used to design a difference feature refinement (DFR) architecture to capture temporal difference information at diverse feature levels. At last, we proposed a multiscale-guided fusion (MGF) module to fuse pyramid features, thereby dealing with scaling changes. Experimental results on three datasets show that CTIDRNet can extract irregularly changed areas effectively, and the evaluation result outperforms other SOTA methods, with an improvement of 1.79%–19.82%, 2.9%–11.07%, and 0.97%–8.91% in terms of F1 for CDD, SYSU, and LEVIR datasets, respectively. The demo code of this work is publicly available athttps://github.com/lucyjiong/CTIDR.
Kangning Du, Xian Sun 0001, Lin Cao 0003, Shu Tian
IEEE Geosci. Remote. Sens. Lett.5
2025 Decoupling and Interaction: task coordination in single-stage object detection
Jia-Wei Ma, Shu Tian, Haixia Man, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin
Multim. Tools Appl.2
2025 MSSI-Net: Multiscale Semantic-Guided Synergistic Interaction Network for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RSCD) has become an essential tool in observing and analyzing geographical information. However, existing deep learning approaches dependent solely on visual modalities may encounter challenges in discerning subtle variations amidst noise interference. To overcome these issues, we propose a multiscale semantic-guided synergistic interaction network (MSSI-Net), which utilizes the advanced multimodal semantic representations for enhancing the capacity to perceive hierarchical changes. Specifically, we first devise a multiscale interaction module (MIM) which leverages multiscale attention mechanism to guide the interaction between the coarse and fine stages of different visual features. The fine-grained visual features subsequently complement the semantic features through scale weight reassignment to enhance the discriminative capability of vision-language features. Furthermore, driven by the semantic-guided synergistic interaction mechanism, our developed cross-modal feature fusion module (CFFM) exploits both homogeneous and heterogeneous features among modalities. This ensures that the generated vision-language features are semantically representative. Finally, we formulate a manifold differential perception head (MDPH) to optimize the detection of changes by efficiently fusing diverse differential feature representations, achieving comprehensive performance enhancement. Extensive experiments conducted on four benchmark datasets (LEVIR-CD, CDD, SYSU-CD and WHU-CD) indicate that the designed MSSI-Net achieves state-of-the-art performance compared to existing methods.
Shu Tian, Jiyuan Shen, Lin Cao 0003, Lihong Kang, Xian Sun 0001, Xiangwei Xing, Chunzhuo Fan, Kangning Du, Chong Fu 0001, Ye Zhang 0008
IEEE Trans. Geosci. Remote. Sens.1
2024 Attention Decoupling for Query-Based Object Detection
abstract
Benefiting from attention mechanisms, query-based detectors have a strong model capacity. They predict classification and regression by utilizing their shared queries and features in the decoder. Inter-task biases cause multi-directional gradients that disturb each other to limit model optimization. In this work, we introduce an attention decoupling (AD) for query-based detectors to explicitly align multi-task features. Specifically, AD consists of a Dense-to-Sparse Query Generator (DSQG) and a Split Cross-Attention (SCA), enabling query and feature decoupling respectively in decoding phase. Then, we propose a task consistency loss (TCL) which integrates a novel task alignment metric to classification loss to further improve task consistency across multiple decoding stages. Thus, AD effectively mitigates query-based detectors’ task misalignment problem and inspires subsequent multi-task paradigms. Moreover, extensive experiments on COCO dataset demonstrate that the proposed AD can enhance a variety of representative detectors. Remarkably, AD-DINO achieves the state-of-the-art performance.
Jia-Wei Ma, Haixia Man, Shu Tian, Jingyan Qin, Xu-Cheng Yin
ICASSP4
2024 Dynamic receptive field adaptation for scene text recognition
Shu Tian, Kang-Xi Zhu, Haibo Qin
Pattern Recognit. Lett.1
2024 Sample Weighting with Hierarchical Equalization Loss for Dense Object Detection
abstract
Label assignment (LA) is one of the essential phases in the object detection paradigm and aims to classify samples as foreground or background. Current LA strategies generally discriminate samples by explicit thresholds and then calculate weighted losses based on their significances. However, existing methods mostly neglect to consider the importance of samples comprehensively due to the uneven distribution of objects and the limitations of detector structures. In this paper, we propose a hierarchical equalization loss (HEL) by reconsidering the underlying factors affecting sample weights. First, we mitigate sample imbalance at three progressive levels. (1) Task level. We propose task-reconciled weights (TRW) to overcome the effects caused by inter-task inconsistencies (i.e., the inherent differences of classification and localization). (2) Instance level. We propose instance-aware normalization (IAN) for reconstructing the distribution of sample weights within an instance to suppress environmental noise. (3) Pyramid level. We propose hierarchical modulation (HM) to alleviate the unbalanced distribution of multi-scale objects on feature pyramids. Then, we stack the above three mechanisms and formulate the effective weighted loss. Moreover, we propose a staggered candidate bag construction (SCBC) mechanism to further improve the robustness of our method. Without adding any extra overhead, HEL can improve the performance of representative detectors by an impressive margin. Equipped with HEL, a single “ResNet-50+FPN+Head” detector can achieve a performance of 41.9 AP on COCO under 1× schedule, outperforming other existing LA methods. Extensive experiments conducted on multiple backbones and datasets demonstrate the effectiveness of our method.
Jia-Wei Ma, Lei Chen 0069, Shu Tian, Song-Lu Chen, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Multim.4
2022 A Relation-Augmented Embedded Graph Attention Network for Remote Sensing Object Detection
abstract
Multiclass geospatial object detection in high spatial resolution remote sensing imagery (HSRI) is still a challenging task. The main reason is that the objects in HRSI are location-variable and semantic-confusable, which results in the difficulties in differentiating the complicated spatial patterns and deriving the implicitly semantic labels among different categories of objects. In this article, we propose a relation-augmented embedded graph attention network (EGAT), which enables the full exploitation of the underlying spatial and semantic relations among objects for improving the detection performance. Specifically, we first construct two sets of spatial and semantic graphs of objects–objects for object relations modeling. Second, a Siamese architecture-based embedding spatial and semantic graph attention network is designed for relations reasoning, which is implemented by introducing the long short-term memory (LSTM) mechanism into the EGAT, for learning the relations among different categories of intraobjects and interobjects. Driven by the spatial and semantic LSTM, the EGAT-LSTM can adaptively focus on the critical information of reason graphs for spatial–semantic correlation discrimination in the embedding non-Euclidean feature space. By this way, the EGAT-LSTM can effectively capture the global and local spatial–semantic relationships of objects–objects, and then produce relations-augmented features for improving the performance of object detection. We conduct comprehensive experiments on three public datasets for multiclass geospatial object detection. Our method achieves state-of-the-art performance, which demonstrates the superiority and effectiveness of the proposed method.
Shu Tian, Lihong Kang, Xiangwei Xing, Chunzhuo Fan, Ye Zhang 0008
IEEE Trans. Geosci. Remote. Sens.1
2022 Depth-Guided Progressive Network for Object Detection
abstract
Multi-scale object detection in natural scenes is still challenging. To enhance the multi-scale perception capability, some algorithms combine the lower-level and higher-level information via multi-scale feature fusion strategies. However, the inherent spatial properties among instances and relations between foreground and background are ignored. In addition, the human-defined “center-based” regression quality evaluation strategy, predicting a high-to-low score based on a linear relationship with the distance to the center of ground-truth box, is not robust to scale-variant objects. In this work, we propose a Depth-Guided Progressive Network (DGPNet) for multi-scale object detection. Specifically, besides the prediction of classification and localization, the depth is estimated and used to guide the image features in a weighted manner to obtain a better spatial representation. Therefore, depth estimation and 2D object detection are simultaneously learned via a unified network, where the depth features are merged as auxiliary information into the detection branch to enhance the discrimination among multi-scale objects. Moreover, to overcome the difficulty of empirically fitting the localization quality function, high-quality predicted boxes on scale-variant objects are more adaptively obtained by an IoU-aware progressive sampling strategy. We divide the sampling process into two stages, i.e., “statistical-aware” and “IoU-aware”. The former selects thresholds for positive samples based on statistical characteristics of multi-scale instances, and the latter further selects high-quality samples by IoU on the basis of the former. Therefore, the final ranking scores better reflect the quality of localization. Experiments verify that our method outperforms state-of-the-art methods on the KINS and Cityscapes dataset.
Jia-Wei Ma, Song-Lu Chen, Feng Chen 0040, Shu Tian, Jingyan Qin, Xu-Cheng Yin
IEEE Trans. Intell. Transp. Syst.5
2021 Online Scene Text Tracking with Spatial-Temporal Relation
Yan Xiu, Shu Tian, Xu-Cheng Yin
ICIG (3)3
2021 A Complete Building Extraction Framework for Airborne Laser Scanning Point Cloud
abstract
In this paper we proposed a complete building extraction framework (CBEF) for airborne laser scanning point clouds. By using 3D instance segmentation to extract rough buildings, we proposed a post-processing method to optimize the extract results, and used a point cloud completion network to repair the incomplete building instances. Experimental results show that this proposed framework can better extract building instances from airborne laser scanning point cloud, and can repair incomplete building point clouds those lost facades.
Chunhui Zhao 0003, Hemin Lin, Nan Su 0001, Shu Tian
IGARSS5
2021 End-to-end trainable network for degraded license plate detection via vehicle-plate relation mining
Song-Lu Chen, Shu Tian, Jia-Wei Ma, Qi Liu 0041, Feng Chen 0040, Xu-Cheng Yin
Neurocomputing2
2021 Siamese Graph Embedding Network for Object Detection in Remote Sensing Images
abstract
Multiclass geospatial object detection is a vital fundamental task for many remote sensing applications. However, it still faces several challenges in very high-resolution (VHR) images in remote sensing, such as the ambiguity of object appearance and the complexity of spatial distribution. In this letter, we propose a novel Siamese graph embedding network (SGEN) that leverages the spatial and semantic information to jointly extract the high-level feature representation for object detection. The main purpose of our SGEN is to learn an embedding discriminative feature space that strengthens the interclass compactness while alleviating the intraclass separability. Specifically, we first design a novel contrastive loss in terms of spatial dependence and semantic correspondence for graph similarity metric learning (ML). Then, the SGEN architecture is adopted for spatial and semantic similarity learning by training the novel contrastive loss function. The SGEN model contains two-stream graph convolutional networks (GCNs) for ML, which is helpful to capture the discriminative features. At last, these extracted features with high spatial and semantic discrimination are served to improve the performance of object detection. The comprehensive evaluations on a combined data set consisting of two public object detection data sets demonstrate the effectiveness of the proposed method.
Shu Tian, Lihong Kang, Xiangwei Xing, Zhou Li 0002, Chunzhuo Fan, Ye Zhang 0008
IEEE Geosci. Remote. Sens. Lett.1
2019 Combined Correlation Filters with Siamese Region Proposal Network for Visual Tracking
Shugang Cui, Shu Tian, Xu-Cheng Yin
ICONIP (2)2
2018 Building Reconstrucion Using Three-Dimensional Zernike Moments in Digital Surface Model
abstract
With the development of airborne Lidar, a large number of digital surface models(DSMs) are now available, which has greatly contributed to the rapid development of the applications for building reconstruction. Similar to two-dimensional Zernike moments(2D-ZMs), the three-dimensional Zernike moments(3D-ZMs) consist of two components, amplitude component and phase component, respectively. Benefiting from rotation invariance of the amplitude characteristic and orientation difference of the phase characteristic, the 3D-ZMs have excellent performance for building recognition. Image reconstruction using 2D-ZMs has been achieved in early research. However, few studies have referred to three-dimensional object reconstruction based on 3D-ZMs. This paper proposes a method using 3D-ZMs to approximately reconstruct the shape of a general 3D building. Firstly, we study the amplitude and phase characteristic of 3D-ZMs to distinguish buildings, and conversely, the 3D-ZMs are used to reconstruct 3D buildings. Experimental results illustrated that this proposed method could be used to approximately reconstruct the shape of a general 3D building from a small number of moments.
Ye Zhang 0008, Shu Tian
IGARSS3
2018 A Unified Framework for Tracking Based Text Detection and Recognition from Web Videos
abstract
Video text extraction plays an important role for multimedia understanding and retrieval. Most previous research efforts are conducted within individual frames. A few of recent methods, which pay attention to text tracking using multiple frames, however, do not effectively mine the relations among text detection, tracking and recognition. In this paper, we propose a generic Bayesian-based framework of Tracking based Text Detection And Recognition (T DAR) from web videos for embedded captions, which is composed of three major components, i.e., text tracking, tracking based text detection, and tracking based text recognition. In this unified framework, text tracking is first conducted by tracking-by-detection. Tracking trajectories are then revised and refined with detection or recognition results. Text detection or recognition is finally improved with multi-frame integration. Moreover, a challenging video text (embedded caption text) database (USTB-VidTEXT) is constructed and publicly available. A variety of experiments on this dataset verify that our proposed approach largely improves the performance of text detection and recognition from web videos.
Shu Tian, Xu-Cheng Yin, Ya Su, Hongwei Hao
IEEE Trans. Pattern Anal. Mach. Intell.1
2017 A Novel Deep Embedding Network for Building Shape Recognition
abstract
Building shape, as a key structured element, plays a significant role in various urban remote sensing applications. However, because of high complexity and intraclass variations between building structures, the capability of building shape description and recognition becomes limited or even impoverished. In this letter, a novel deep embedding network is proposed for building shape recognition, which combines the strength of the unsupervised feature learning of convolutional neural networks (CNNs) and a novel triplet loss. Specifically, we take advantage of the strong discriminative power of CNNs to learn an efficient building shape representation for shape recognition. With this deep embedding network, the high-dimensional image space can be mapped into a low-dimensional feature space, and the deep features can effectively reduce the intraclass variations while increasing the interclass variation between different building shape images. Afterward, the derived deep features are exploited for the process of building shape recognition. This method consists of two stages. In the first stage, for standard building shape image queries stored in the shape primitives library and the building shape data set, two sets of deep features are extracted with the deep embedding network. In the second stage, we formulate the shape recognition task into a feature matching problem and the final building shape recognition results can be achieved by set-to-set feature matching method. Experiments on the VHR-10 and UCML data sets demonstrate the effectiveness and precision of the proposed method.
Shu Tian, Ye Zhang 0008, Junping Zhang, Nan Su 0001
IEEE Geosci. Remote. Sens. Lett.1
2017 Tracking Based Multi-Orientation Scene Text Detection: A Unified Framework With Dynamic Programming
abstract
There are a variety of grand challenges for multi-orientation text detection in scene videos, where the typical issues include skew distortion, low contrast, and arbitrary motion. Most conventional video text detection methods using individual frames have limited performance. In this paper, we propose a novel tracking based multi-orientation scene text detection method using multiple frames within a unified framework via dynamic programming. First, a multi-information fusion-based multi-orientation text detection method in each frame is proposed to extensively locate possible character candidates and extract text regions with multiple channels and scales. Second, an optimal tracking trajectory is learned and linked globally over consecutive frames by dynamic programming to finally refine the detection results with all detection, recognition, and prediction information. Moreover, the effectiveness of our proposed system is evaluated with the state-of-the-art performances on several public data sets of multi-orientation scene text images and videos, including MSRA-TD500, USTB-SV1K, and ICDAR 2015 Scene Videos.
Xu-Cheng Yin, Wei-Yi Pei, Shu Tian, Ze-Yu Zuo, Chao Zhu 0003, Junchi Yan
IEEE Trans. Image Process.4
2016 Scene Text Detection in Video by Learning Locally and Globally
Shu Tian, Wei-Yi Pei, Ze-Yu Zuo, Xu-Cheng Yin
IJCAI1
2016 Multi-object tracking with inter-feedback between detection and tracking
Shu Tian, Fei Yuan 0003, Gui-Song Xia
Neurocomputing1
2016 Phase Analysis of Three-Dimensional Zernike Moment for Building Classification and Orientation in Digital Surface Model
abstract
In this letter, we proposed a phase analysis of the 3-D Zernike moment (3D-ZM), to estimate the orientation difference between buildings in a digital surface model (DSM). A 3-D analysis using the DSM is an important way for building reconstruction and many further remote sensing applications. By using the 3D-ZM, we could decompose the 3-D structure of an object in a complex domain. Benefiting from rotation invariance of the amplitude component, the 3D-ZM has excellent performance for object classification. However, the phase component of 3D-ZM is ignored in early research, by which orientation of different objects could be analyzed, and similar buildings (within one class) could be distinguished meticulously, whereas other traditional geometric features may fail to do so. Therefore, we studied the phase analysis of 3D-ZM and introduced a flow frame of orientation-difference estimation. Experimental results illustrated that our method could robustly find the orientation difference between similar buildings, and improvement on accuracy was achieved for building classification and orientation.
Ye Zhang 0008, Shu Tian, Fengjiao Gao
IEEE Geosci. Remote. Sens. Lett.3
2016 Text Detection, Tracking and Recognition in Video: A Comprehensive Survey
abstract
The intelligent analysis of video data is currently in wide demand because a video is a major source of sensory data in our lives. Text is a prominent and direct source of information in video, while the recent surveys of text detection and recognition in imagery focus mainly on text extraction from scene images. Here, this paper presents a comprehensive survey of text detection, tracking, and recognition in video with three major contributions. First, a generic framework is proposed for video text extraction that uniformly describes detection, tracking, recognition, and their relations and interactions. Second, within this framework, a variety of methods, systems, and evaluation protocols of video text extraction are summarized, compared, and analyzed. Existing text tracking techniques, tracking-based detection and recognition techniques are specifically highlighted. Third, related applications, prominent challenges, and future directions for video text extraction (especially from scene videos and web videos) are also thoroughly discussed.
Xu-Cheng Yin, Ze-Yu Zuo, Shu Tian, Cheng-Lin Liu 0001
IEEE Trans. Image Process.3
2015 Multi-strategy tracking based text detection in scene videos
abstract
Text detection and tracking in scene videos are important prerequisites for content-based video analysis and retrieval, wearable camera systems and mobile devices augmented reality translators. Here, we present a novel multi-strategy tracking based text detection approach in scene videos. In this approach, a state-of-the-art scene text detection module [1] is first used to detect text in each video frame. Then a multi-strategy text tracking technique is proposed, which uses tracking by detection, spatio-temporal context learning, and linear prediction to predict the candidate text location sequentially, and adaptively integrates and selects the best matching text block from the candidate blocks with a rule-based method. This multi-strategy tracking technique can combine the advantages of the three different tracking techniques and afterwards make remedies to the disadvantages of them. Experiments on a variety of scene videos show that our proposed approach is effective and robust to reduce false alarm and improve the accuracy of detection.
Ze-Yu Zuo, Shu Tian, Wei-Yi Pei, Xu-Cheng Yin
ICDAR2
2015 Building surface texture segmentation in urban remote sensing image using improved ORTSEG algorithm
abstract
Texture segmentation is a critical step in building-based analysis in urban remote sensing images to obtain more detail information for further applications. Most existing segmentation algorithms rely on region or edge information to segment, which failed to segment building surfaces with almost same texture and unclear edge between the surfaces. Therefore, in order to solve this challenging task, based on the ORTSEG algorithm in Michael T. McCann's paper, an improved ORTSEG algorithm is proposed, in which, a sparse non-negative matrix factorization (SNMF) is used for the optimization. The final segmentation results show the superiority of this improved ORTSEG algorithm.
Shupei Deng, Ye Zhang 0008, Shu Tian
IGARSS3
2015 A roof-contour guided multi-side interpolation method for building texture-mapping using remote sensing resource
abstract
In this paper, we proposed a novel texture-mapping method for buildings using remote sensing images and digital surface model. For generating better 3D map, boring manually or semi-automatic texture-mapping of buildings is always needed. However, only with remote sensing images and digital surface model, it is difficult to generate `sides' of buildings, and corresponding relation between triangular mesh and texture image is also hard to be found. Inspired by building extraction result, we found that roof-contour could guide interpolation of `sides' of building, and generate triangular mesh, and an interactive frame of texture mapping is introduced for processing the generated multi-sides and corresponding texture-images of buildings. Experiments show that texture-mapping of buildings using remote sensing resource is easily realized with our frame, and excellent results could be obtained.
Xi Chen 0004, Fengjiao Gao, Ye Zhang 0008, Yi Shen 0001, Nan Su 0001, Shu Tian
IGARSS7
2014 Orientation estimation of building using DSM and optical images based on Zernike moments
abstract
Digital Surface Models (DSMs) generated from airborne laser-scanning or stereo satellite images provide a very useful source of information for building orientation estimation. However, due to the high complexity of the building structure, it is hardly to get the details of the building boundary information. In this paper, a new orientation estimation method is proposed, which merge the edge information extracted from optical images into DSM, and take advantage of the ZM phase information in the orientation estimation process. The classical way of extracting Zernike features only takes into account the magnitude of the moments and loses the phase information. The novelty of our approach is to take advantage of the phase information of Zernike moments to capture these variabilities in a way that makes it robust in the context of rotation angle estimation.
Shu Tian, Ye Zhang 0008, Yanfeng Gu
IGARSS1
2013 Transfer learning based compressive tracking
abstract
Although existing online tracking algorithms can solve the problems of scene illumination changes, partial or full object occlusions, and pose variation, there are still two weaknesses, inadequacy of training data and drift problem. Considering these, Compressive Tracking algorithm (CT) [1] extracts features from compressed domain, and classified object and background via a naive Bayes classier with online update. To further solve the problems of drift and inadequacy of training data, we introduce transfer learning into CT to take full advantage of prior information and propose a self-traininglike transfer learning algorithm. It selects training samples from samples collection to update classifier by the conduction of the classifier constructed at first frame. Eventually we introduce self-training-like transfer learning algorithm into CT to construct a novel tracking algorithm called Transfer Learning based Compressive Tracking (TLCT). Experimental results on 17 publicly available challenging sequences have shown the effectiveness and robustness of our algorithm.
Shu Tian, Xu-Cheng Yin, Hongwei Hao
IJCNN1
2012 Pedestrian Analysis and Counting System with Videos
Zhi-Bin Wang, Hongwei Hao, Xu-Cheng Yin, Shu Tian
ICONIP (5)5