Shu Tang

dblp:79/10856 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 A spatial-frequency hybrid restoration network for JPEG compressed image deblurring
Shu Tang, Xinbo Gao 0001, Shuli Yang, Jiaxu Leng, Zengdan Pan
Neural Networks1
2026 A multi-scale gate network for high-quality image deblurring
Shu Tang, Yufu Lin, Xinbo Gao 0001, Shuli Yang, Jiaxu Leng, Zengdan Pan
Pattern Recognit.1
2026 Efficient bidirectional fusion multi-gate network for lightweight single image super-resolution
Shuli Yang, Shu Tang, Xinbo Gao 0001, Jiaxu Leng, Xianzhong Xie
Pattern Recognit.2
2026 A Lightweight Frequency-Selection-Based Progressive Patch Transformer Network for Single Image Super-Resolution
abstract
Recently, lightweight networks for single image super-resolution (SISR) have surged due to the need of resource-constrained devices, where divide-and-conquer multi-route model exhibits impressive trade-off between performance and computational cost. However, most existing divide-and-conquer multi-route models face two key limitations: (1) possible suboptimal decoupling of image components (e.g. smooth regions, edges and texture details) due to spatial-domain-only processing, and (2) inability to model global dependencies explicitly and capture structural information, hindering further performance gains. To address these drawbacks, we propose a lightweight frequency-selection-based progressive patch Transformer network (FSPPTN) for higher-quality SISR reconstruction. Specifically, we first propose a frequency selection module, in which we develop a frequency enhancement branch (FEB) to dynamically decouple different image components by introducing the window-based Fast Fourier transform (WFFT) and a learnable weight matrix, and a spatial restoration branch (SRB) to recalibrate and fuse cross-granularity features by designing a multi-gate mechanism for reconstructing the component information screened out by the FEB at current level. Secondly, we propose a lightweight multi-branch gradient-guided inter-patch self-attention to explicitly capture global structural similarities by summarizing structural information of each patch into a lower-dimensional space using the statistical properties of first-order gradients, thereby achieving explicit global dependencies modeling and lightweight. Extensive experimental results demonstrate that, in the vast majority of cases, FSPPTN outperforms state-of-the-art lightweight SISR methods in terms of both performance and computational overhead, especially for ×3 and ×4 SR, e.g. FSPPTN outperforms MaIR-Small by even 0.14dB PSNR on Manga109 dataset for ×4 SR even with 48.3% fewer parameters and 63.5% lower FLOPs. The code is available at: https://github.com/yslyangshuli/FSPPTN-main.
Shuli Yang, Shu Tang, Xinbo Gao 0001, Xianzhong Xie, Jiaxu Leng
IEEE Trans. Circuits Syst. Video Technol.2
2025 Integrating BIM and AIoT for Smart Engineering Laboratory Monitoring and Management
abstract
Automatic indoor environmental quality (IEQ) monitoring plays a pivotal role in the management of green building operations. Traditional monitoring methods that integrate Building Information Modeling (BIM) and the Internet of Things (IoT) are unable to perform automatic detection. This study addresses the limitation by introducing a BIM-AIoT based ‘LabMonitor’ approach for real-time IEQ monitoring and prediction. To enhance the accuracy of detecting occupants comfort, a convolutional neural network (CNN) based YOLOv8 model is utilized. The effectiveness of the proposed approach is validated in diverse engineering laboratories within a university setting. Results demonstrated that the BIM-AIoT based LabMonitor approach achieves a high mean average precision (mAP) of 0.939 in real-time indoor comfort level detection tasks. This work provides a scalable and interoperable solution for smart engineering laboratory management, addressing key challenges in multimodal data fusion and automatic real-time monitoring.
Guofeng Qiang, Shu Tang, Jianli Hao, Luigi Di Sarno, Guangdong Wu, Mengtian Yin
IEEE Internet Things J.2
2024 M-NeuS: Volume rendering based surface reconstruction and material estimation
Shu Tang, Jiabin He, Shuli Yang, Hongxing Qin
Comput. Aided Geom. Des.1
2023 Assistance from the Ambient Intelligence: Cyber-physical​ system applications in smart buildings for cognitively declined occupants
Xinghua Gao, Saeid Alimoradi, Jianli Chen, Yuqing Hu 0002, Shu Tang
Eng. Appl. Artif. Intell.5
2022 A Stage-Mutual-Affine Network for Single Remote Sensing Image Super-Resolution
Shu Tang, Xianbo Xie, Shuli Yang, Wanling Zeng
PRCV (4)1
2022 BiDFNet: Bi-decoder and Feedback Network for Automatic Polyp Segmentation with Vision Transformers
Shu Tang, Junlin Qiu, Xianzhong Xie, Haiheng Ran, Guoli Zhang
PRCV (2)1
2021 Deep Convolutional-Neural-Network-Based Channel Attention for Single Image Dynamic Scene Blind Deblurring
abstract
The success of convolutional neural network (CNN) based single image dynamic scene blind deblurring (SIDSBD) methods mainly stems from the multi-scale/multi-patch model and the designs of the encoder-decoder architecture, and the residual block structure, which make different contributions to SIDSBD. In this paper, we further exploit the advantages of the multi-scale model, the encoder-decoder module, and the residual block structure, respectively, and propose a novel multi-scale channel attention network (MSCAN) for effective single image dynamic scene blind deblurring. Different from existing multi-scale models, in our proposed network, each scale consists of multiple levels, in which a novel spatial pyramid pooling channel attention (SPPCA) strategy is proposed to adaptively rescale the channel-wise features by using both the global and local feature statistics for more powerful network representation. Extensive experiments on both the synthetic benchmark datasets and the real blurred images show that our method can produce better deblurring results than the state-of-the-art SIDSBD methods in terms of both qualitative evaluation and quantitative metrics.
Shengdao Wan, Shu Tang, Xianzhong Xie, Bin Ma 0005, Lei Luo 0003
IEEE Trans. Circuits Syst. Video Technol.2
2020 Real-Time Tracking of Vehicles with Siamese Network and Backward Prediction
abstract
Tracking of vehicles is a key technique for Intelligent transportation system, which commonly follows tracking-by-detection strategy. Due to high appearance similarity among vehicles and heavy occlusion caused by busy traffic flow, a major challenge in such a tracking system is the limited performance of the underlying detector which may produce noisy detections. Consequently, Siamese network and backward prediction-based vehicle tracking approach is proposed. Siamese network based forward position prediction is designed to alleviate the interference of noisy detections, while backward prediction verification is performed to reduce the false positives arising with forward prediction. The final tracklets are obtained through weighted merging based on the detection confidence and forward prediction confidence. The experiment results demonstrate that the proposed method outperforms the state-of-the-art on the UA-DETRAC vehicle tracking dataset, as well as maintains real-time processing at an average tracking speed of 20.1fps, which can be used for real-time applications.
Ao Li 0007, Lei Luo 0003, Shu Tang
ICME3
2020 Bi-Labeled LDA: Inferring Interest Tags for Non-famous Users in Social Network
abstract
Abstract User tags in social network are valuable information for many applications such as Web search, recommender systems and online advertising. Thus, extracting high quality tags to capture user interest has attracted many researchers’ study in recent years. Most previous studies inferred users’ interest based on text posted in social network. In some cases, ordinary users usually only publish a small number of text posts and text information is not related to their interest very much. Compared with famous user, it is more challenging to find non-famous (ordinary) user’s interest. In this paper, we propose a probabilistic topic model,Bi-Labeled LDA,to automatically find interest tags for non-famous users in social network such as Twitter. Instead of extracting tags from text posts, tags of non-famous users are inferred from interest topics of famous users. With the proposed model, the formulation of social relationship between non-famous users and famous user is simulated and interest tags of famous users are exploited to supervise the training of the model and to make use of latent relation among famous users. Furthermore, the influence of popularity of famous user and popular tags are considered, and tags of non-famous users are ranked based on random walk model. Experiments were conducted on Twitter real datasets. Comparison with state-of-the-art methods shows that our method is more superior in terms of both ranking and quality of the tagging results.
Jun He 0008, Hongyan Liu 0002, Yiqing Zheng, Shu Tang, Xiaoyong Du 0001
Data Sci. Eng.4
2019 Joint Texture/Depth Power Allocation for 3-D Video SoftCast
abstract
Recently, a novel uncoded (pseudoanalog) scheme called SoftCast is proposed for wireless video transmission, which eliminates the cliff effect of the state-of-the-art source-channel coding based schemes and achieves linear quality transition within a wide range of channel signal-to-noise ratio. Therefore, SoftCast-like uncoded and hybrid transmission has become an attractive research issue for natural 2-D video. However, very few studies focus on the SoftCast-based wireless transmission of the 3-D video (3DV) currently. One critical issue of 3DV SoftCast is how to allocate the limited power budget of the transmitter to the texture videos and depth maps of the 3DV to achieve the optimal overall quality on the receiver side, including the transmission quality of the reference views and the synthesis quality of the virtual views. This paper attempts to solve the optimal joint power allocation problem in an efficient way. First, we formulate the target problem as a constrained power-distortion optimization (PDO) problem mathematically. Then, each part of the distortion is analyzed and formulated in a closed form. Finally, the PDO problem is mapped to an unconstrained convex optimization problem and solved by the Lagrangian multiplier method. Simulation results demonstrate that the performance of the proposed method is close to that of the full search method, which can provide the best performance theoretically. Nevertheless, the complexity of the proposed method is negligible compared with that of the full search method. In addition, as compared with the fixed ratio (e.g., 1:1) power allocation between texture and depth, the proposed method can achieve a PNSR gain up to 1.8 dB.
Lei Luo 0003, Taihai Yang, Ce Zhu, Zhi Jin 0002, Shu Tang
IEEE Trans. Multim.5
2018 Spatial-scale-regularized blur kernel estimation for blind image deblurring
Shu Tang, Xianzhong Xie, Lei Luo 0003, Peisong Liu
Signal Process. Image Commun.1
2016 A Novel Parameter Estimation Algorithm Based on GHMM for Vertical Handover
abstract
Great efforts have been driven to improve the optimal decision for network selection, which is significant for vertical handover to meet user's rapidly increased requirements. Unfortunately, the instability of the basic decisive parameters is usually ignored, resulting in a high misjudgement ratio. Therefore, this paper proposed a novel parameter estimation algorithm to bridge the observed values and the inherent characters for the decisive parameters via Gaussian Hidden Markov Model. After trained by given parameter sequences in off-line phase, the parameter pre-estimated algorithm can reveal the parameters' inherent statuses in the on-line phase. Furthermore, the pre- estimated results can be regarded as normalized inputs for network selections. The simulation results show that the proposed algorithm could reduce the misjudgement ratio of the network selection algorithms, especially exhibiting a good performance for the low speed mobile users.
Shu Tang, Lin Ma 0001, Yubin Xu
GLOBECOM1
2015 Extracting Interest Tags for Non-famous Users in Social Network
abstract
Inferring interests of users in social network is important for many applications such as personalized search, recommender systems and online advertising. Most previous studies inferred users' interests based on text posted in social network, which is usually not related to their interests. In this paper, we propose a modified topic model, Bi-Labeled LDA with a term weighting scheme, to extract interest tags for users in social network. The proposed model utilize only users' relationship information without requirement for text information, and incorporates supervision into traditional LDA. Specifically, we introduce method to extract tags for non-famous user through their relationship with famous users in Twitter, and study why a non-famous user follows famous users simultaneously. Comparison with state-of-the-art methods on real dataset shows that our method is far more superior in terms of precision and recall of the extracted tag set, and also more applicable for many personalized applications. Besides, we find that a reasonable term weighting scheme can actually improve the performance further.
Hongyan Liu 0002, Jun He 0008, Shu Tang, Xiaoyong Du 0001
CIKM4
2014 Temporal Consistency Based Method for Blind Video Deblurring
abstract
In this paper, we propose a video deblurring method for recovering effectively the blurred video caused by camera shake. First, we estimate the blur kernel for each blurred video frame by introducing a preprocessing strategy which employs the anisotropic diffusion and shock filter to get strong edges from each video frame. Second, for video deblurring, we propose a temporal cubic rhombic mask technique, which can not only suppress ringing artifacts, but also can enhance temporal consistency of the deblurred video. The proposed technique uses the relationship of the current frame and the frames before the current frame, and the frame behind the current frame for obtaining a high-quality deblurred video. Experiments demonstrate that proposed method outperforms state-of-the-art deblurring methods for the blurred video.
Weiguo Gong, Weihong Li 0001, Shu Tang
ICPR4
2014 Non-blind image deblurring method by local and nonlocal total variation models
Shu Tang, Weiguo Gong, Weihong Li 0001
Signal Process.1
2012 Total variation blind deconvolution employing split Bregman iteration
Weihong Li 0001, Quanli Li, Weiguo Gong, Shu Tang
J. Vis. Commun. Image Represent.4