Tao Chen 0002

dblp:69/510-2 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0001-6447-6153ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Full-Space DOA Estimation for STAR-RIS via ADMM
abstract
This paper proposes an efficient gridless direction of arrival (DOA) estimation method for simultaneously transmitting and reflecting reconfigurable intelligent surface (STARRIS) assisted full-space localization. To handle the uncertain reflection–transmission parameter mapping induced by different metasurface constraints and hardware architectures, we adopt a generalized model and construct subspace-specific atomic sets, thereby casting DOA estimation as an atomic norm minimization (ANM) problem.We further derive the corresponding dual model and exploit the dual polynomial property to recover the DOAs via polynomial rooting. To enhance real-time capability, an efficient alternating direction method of multipliers (ADMM) solver is developed, ensuring stable and fast convergence to the optimum of the convex program. Furthermore, we investigate the impact of key parameters through numerical simulations and validate the effectiveness and robustness of the proposed methods.
Tao Chen 0002, Muran Guo
IEEE Signal Process. Lett.1
2026 LKAFormer: A Lightweight Kolmogorov-Arnold Transformer Model for Image Semantic Segmentation
abstract
Transformer-based semantic segmentation methods have demonstrated outstanding performance by leveraging global self-attention to effectively capture long-range dependence. However, there still exist two issues in existing works: (1) Most of them utilize the full-rank weight matrix to support the self-attention mechanism and feed-forward network in modelling long-range dependence between patches/pixels, resulting in a high computational cost during both training and inference. (2) Most of them ignore information interactions between high-level semantics and low-level structures during the image resolution recovery, which leads to the performance degradation in segmenting objects with complex boundaries. To tackle these challenges, a lightweight Kolmogorov-Arnold Transformer model (LKAFormer) is proposed for the image semantic segmentation, containing a two-stream lightweight Transformer encoder and a graph feature pyramid aggregation KAN-decoder. The former constructs a hierarchical feature cross-scale fusion pipeline to obtain sufficient semantics containing comprehensive multi-scale information via setting coarse-grained and fine-grained streams with different-size patches of images. In that pipeline, feature lightweight focusing modules model complex and long-range dependence across patches/pixels to refine image semantics with less computational costs by lightweight multi-head self-attention and lightweight feed-forward network designs. The latter leverages the learnable nonlinear transformation mechanism of the Kolmogorov-Arnold Transformer architecture to adaptively capture spatial structure dependence of distinct sub-regions of images. And then, it jointly performs the intra-scale graph fusion and cross-scale graph fusion during the image resolution recovery to enhance information interactions between high-level semantics and low-level structures, which achieves the robust boundary localization and texture refinement of segmentation objects. Finally, plentiful experiments are conducted on three challenging datasets, and the results show LKAFormer sets a new baseline in the image segmentation task in comparison with 11 methods.
Shoulin Yin, Liguo Wang 0001, Tao Chen 0002, Huafei Huang 0001, Jing Gao 0007, Jianing Zhang 0001, Meng Liu 0025, Peng Li 0027, Chengpei Xu
ACM Trans. Intell. Syst. Technol.3
2025 Radar Signal Intra-Pulse Modulation Recognition Based on Point Cloud Network
abstract
Aiming at the existing deep learning radar signal modulation recognition methods are mostly based on time-frequency image (TFI) and consequently result in networks with a large number of parameters due to the significant amount of redundant information contained in TFI, this paper proposes a radar signal intra-pulse modulation recognition method based on point cloud which removes redundant information. Radar signals of different modulation types are mapped into point cloud after Smoothed Pseudo Wigner-Ville Distribution (SPWVD) transformation. Then, PointNet++ is used to classify the point cloud data according to its modulation type and output its corresponding modulation type labels. Simulation results show that the proposed method can effectively recognize radar signals of typical modulation types, and show strong effectiveness and reliability at low signal-to-noise ratio (SNR). Besides, the lightweight characteristics of PointNet++ make the operation of the method more efficient.
Tao Chen 0002, Yingming Liu, Yihan Xiao, Boyi Yang
IEEE Signal Process. Lett.1
2024 Attention Enhancement With Parallel Groups for Remote Sensing Object Detection
abstract
Nowadays, remote sensing object detection has benefited a lot from the development of convolutional neural networks (CNNs). However, it is still a challenging task due to arbitrary orientation and dense distribution of objects in remote sensing images. To deal with these difficulties, we propose two effective attention mechanisms with parallel groups strategy to enhance feature representations in the detection head, named PGAE-head. Significantly, our designs can achieve competitive performance improvement by only introducing tiny parameters and computations in the model. Firstly, the features received by the PGAE-head are divided into multiple groups, which ensures the independence of each group during subsequent attention enhancement. Then, PGAE-head processes these sub-features with enhanced attention mechanisms based on spatial and channel dimensions in parallel to detect more accurate results. Experiments on DOTA and HRSC datasets show that the proposed PGAE-head achieves comparable performances with other state-of-the-art CNN-based models at minimal optimization costs, demonstrating its effectiveness.
Zhigang Yang 0003, Zehao Gao, Jiayue He, Tao Chen 0002, Wei Zhang 0098
ICIP5
2024 Object Detection in Remote Sensing Images With Parallel Feature Fusion and Cascade Global Attention Head
abstract
Convolutional neural networks (CNNs) have driven significant development in remote sensing (RS) object detection. To achieve concise and effective optimization, we propose a two-stage detector with parallel feature fusion strategy and cascade global attention mechanism for object detection in RS images, named PC-RCNN. We first design a feature pyramid network with two parallel branches (PB-FPN), corresponding to the top-down and bottom-up feature fusion pathways, respectively. Different optimization modules can be adopted in different pathways to avoid potential module incompatibility when connected in series. Such parallel feature fusion strategy can achieve both higher detection accuracy and higher computational efficiency compared with previous series fusion modes. Furthermore, we design a global attention block to enhance feature representations of regions and propose a cascade global attention head network (CGA-Head) for accurate category prediction and location estimation. Experiments on a challenging large-scale dataset, namely DOTA, show that the proposed PC-RCNN achieves a mean average precision (mAP) of 77.63%, which is comparable to other state-of-the-art CNN-based models.
Zhigang Yang 0003, Guiwei Wen, Xiangyu Xia, Wei Zhang 0098, Tao Chen 0002
IEEE Geosci. Remote. Sens. Lett.6
2024 S$^{{\text{3}}}$Seg: A Three-Stage Unsupervised Foreground and Background Segmentation Network
abstract
Generative adversarial networks (GANs) gather increasing attention in the field of unsupervised image segmentation. However, most GAN-based unsupervised segmentation methods cannot directly segment a specific individual image, and the segmentation performance can be further boosted. To address these issues, we propose a three-stage unsupervised foreground and background segmentation network (S$^{{\text{3}}}$Seg). In the first stage, we design low-to-high dimensional attention (LHA) to enhance the image understanding ability of the StyleGAN2 generator, which can capture more effective semantic information for the following segmentation task. In the second stage, we introduce an inversion network to form an encoder-generator structure, which can acquire semantic features of a given image. In the third stage, we devise a radial loss to explore the edge of the foreground from the center to the outside, which is beneficial for producing a high-quality mask. S$^{{\text{3}}}$Seg not only provides a solution to direct segmentation of a given image but also outperforms previous unsupervised methods on three public datasets.
Zhigang Yang 0003, Yahui Shen, Wei Zhang 0098, Tao Chen 0002
IEEE Signal Process. Lett.5
2024 Multi-component signal separation based on ALSAE
Tao Chen 0002, Boyi Yang
Wirel. Networks1
2023 Radar Pulse Stream Clustering Based on MaskRCNN Instance Segmentation Network
abstract
In this paper, a clustering algorithm based on the MaskRCNN instance segmentation network is proposed to address the radar pulse stream clustering problem. Compared with traditional algorithms, the suggested algorithm successfully avoids the problem of preset parameters and algorithm structure constraints by using instance segmentation networks as clustering algorithms, and provides an end-to-end solution for the radar signal sorting issue. A PDW lattice mapping method is proposed to adapt to the input form of the instance segmentation network. The instance segmentation network based on ResNet, FPN, RPN, and RolAlign structure detects and masks the individual from the same or different radar signals. The simulation results show that the suggested algorithm performs the sorting of radar pulse stream under the condition of serious overlapping of radar parameters.
Tao Chen 0002, Boyi Yang
IEEE Signal Process. Lett.1
2022 Object Detection in Remote Sensing Images With Balanced Rotational and Horizontal Bounding Boxes
abstract
Object detection in remote sensing images is a challenging task because these images usually contain a number of targets with arbitrary orientations. To annotate the arbitrary-oriented objects accurately, rotational bounding boxes are more effective than horizontal bounding boxes. However, rotational bounding box deformation often happens when objects are near horizontal, since regressing angles is a highly nonlinear task. Aiming at this problem, we propose a region-based convolution neural network with balanced rotational and horizontal bounding boxes (RH-RCNN) for arbitrary-oriented object detection. We first design a multi-layer-enhanced feature pyramid network (ML-FPN) to obtain powerful feature representations. Then we take both stability and accuracy into consideration, and devise an RH-head network to distinguish near horizontal objects from inclined ones. Angle prediction for near horizontal objects is prohibited to avoid large regression deviation. Rotational bounding boxes are only used to locate the obviously inclined objects. Finally, we evaluate the proposed RH-RCNN on the DOTA benchmark. Experimental results show that not only the objective detection accuracy in terms of mAP but also the visualized predicted results are greatly improved.
Zhigang Yang 0003, Junyu Kong, Binxi Zheng, Wei Zhang 0098, Tao Chen 0002
IEEE Geosci. Remote. Sens. Lett.6
2021 Generalized Sparse Polarization Array for DOA Estimation Using Compressive Measurements
abstract
The compressive array method, where a compression matrix is designed to reduce the dimension of the received signal vector, is an effective solution to obtain high estimation performance with low system complexity. While sparse arrays are often used to obtain higher degrees of freedom (DOFs), in this paper, an orthogonal dipole sparse array structure exploiting compressive measurements is proposed to estimate the direction of arrival (DOA) and polarization signal parameters jointly. Based on the proposed structure, we also propose an estimation algorithm using the compressed sensing (CS) method, where the DOAs are accurately estimated by the CS algorithm and the polarization parameters are obtained via the least‐square method exploiting the previously estimated DOAs. Furthermore, the performance of the estimation of DOA and polarization parameters is explicitly discussed through the Cramér‐Rao bound (CRB). The CRB expression for elevation angle and auxiliary polarization angle is derived to reveal the limit of estimation performance mathematically. The difference between the results given in this paper and the CRB results of other polarized reception structures is mainly due to the use of the compression matrix. Simulation results verify that, compared with the uncompressed structure, the proposed structure can achieve higher estimated performance with a given number of channels.
Tao Chen 0002, Weitong Wang, Muran Guo
Wirel. Commun. Mob. Comput.1
2020 Gridless Direction of Arrival Estimation Exploiting Sparse Linear Array
abstract
In this letter, a dual formulation of atomic norm minimization (ANM) approach is proposed by exploiting the vectorized covariance data of the signals received by the extended virtual array. Compared with the traditional ANM-based gridless direction of arrival (DOA) estimation algorithms for sparse linear array (SLA), the proposed algorithm not only realizes the parameter continuity and data completion, but also reduces the size of optimization model and the influence of noise on the estimation accuracy, which can significantly improve the performance of the algorithm. In addition, after establishing an ANM model based on complete vectorized covariance data, the corresponding dual problem is formulated to make the subsequent DOA recovery process no longer need the number of sources as a priori information. Numerial simulations demonstrate the superiority of the proposed algorithm over state of the art gridless DOA estimation algorithms.
Tao Chen 0002, Lin Shi 0004
IEEE Signal Process. Lett.1
2019 Soft Actor-Critic-Based Continuous Control Optimization for Moving Target Tracking
Tao Chen 0002, Xingxing Ma, Shixun You
ICIG (2)1
2018 Performance Analysis for Uniform Linear Arrays Exploiting Two Coprime Frequencies
abstract
Sparse arrays can achieve a higher number of degrees of freedom (DOFs) compared with uniform linear array (ULA) counterparts. To further reduce the number of physical sensors while keeping a high number of DOFs, a direction of arrival (DOA) estimation algorithm by exploiting coprime frequencies base on a sparse ULA is recently proposed. However, the performance of such approach is not properly analyzed. In this letter, we analyze the Cramér-Rao bound (CRB) as the lower bound of the DOA estimation performance. The difference between the results presented in this letter and the recent CRB results on sparse arrays lies primarily in the additional phases occurred when utilizing different frequencies. It is shown in this letter that the phases affect the covariance matrix of the received data vector and, as a result, change the number of resolvable sources and alter the achieved CRB. We first demonstrate the effect of the additional phases with an example of two closely spaced sources, and the CRB for a sparse ULA exploiting two coprime frequencies is then derived. Numerical simulations are provided to validate the analyses.
Muran Guo, Yimin Zhang 0001, Tao Chen 0002
IEEE Signal Process. Lett.3