Tinglong Tang

dblp:239/5324 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-6301-3592ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Few-shot video summarization via cross-video temporal invariance
Tinglong Tang, Fanyuan Wu, Shengyong Chen, Xu Cheng 0003
Neurocomputing1
2025 HTR-VT: Handwritten text recognition with vision transformer
Yuting Li 0001, Dexiong Chen, Tinglong Tang, Xi Shen 0001
Pattern Recognit.3
2025 Deformable Blur Sensing and Regression Analysis ReID Feature Fusion for Multitarget Multicamera Tracking Systems in Highway Scenarios
abstract
In highway scenarios, the rapid motion of vehicles can cause deformation and blur in camera footage, significantly affecting the accuracy of vehicle detection and re-identification (ReID) in multitarget multicamera tracking (MTMCT) systems. To address this issue, this article develops the deformable and blur sensing and regression analysis ReID feature fusion MTMCT system (DSRF). First, a deformable and blur sensing detection module (DFB) in DSRF is designed to overcome the limitations of cameras in capturing fast-moving objects, thereby accurately detecting vehicles moving at high speeds on highways. Then, a regression-based ReID feature fusion algorithm (RARF) in DSRF is proposed, which enhances ReID features by modeling the relationship between vehicle motion and its features, thereby better associating the detected vehicles in consecutive frames into trajectories and establishing intertrajectory relationships. Finally, extensive experiments are conducted on the highway surveillance traffic (HST) dataset developed by our team and the public dataset (CityFlow). Promising results are achieved, validating the effectiveness of our proposed method.
Sixian Chan 0001, Shenghao Ni, Jie Hu 0041, Tinglong Tang, Xiaolong Zhou 0001, Pengyi Hao
IEEE Trans. Comput. Soc. Syst.5
2025 Adaptive Target-Oriented Tracking
abstract
The current one-stream tracking pipelines are early relation modeling in feature extraction. However, insufficient discrimination may result in ambiguous relation modeling during early feature extraction. Moreover, the non-target information occupies most of the search image, rendering most relation modeling futile. To tackle the above issues, we propose tracking via learning adaptive target-oriented representation, named ATOTrack . We design an Untied positional encoding to mark the template token and the search region token separately, which reduces the confused relationship between the template and the search region. Besides, we introduce an Auto-Mask Learner to decouple the target and non-target information in the search region. Interestingly, the Auto-Mask Learner can self-learn and mask the ineffective information to interpret adaptive target-oriented representation. Extensive experiments demonstrate that ATOTrack is superior to existing methods, which achieves the state-of-the-art performance on six tracking benchmarks. In particular, ATOTrack establishes a new record on AViST with 57% AO. The code and models will be released as soon.
Sixian Chan 0001, Xianpeng Zeng, Zhoujian Wu, Yu Wang 0254, Xiaolong Zhou 0001, Tinglong Tang, Jie Hu 0041
ACM Trans. Intell. Syst. Technol.6
2025 Digital twin-enabled deep learning for real-time fire situation awareness
Tinglong Tang, Chunli Zhao, Xinqiong Liu, Shuifa Sun
Vis. Comput.1
2024 Digital Twin-Based Office Equipment Management and Personnel Detection System
abstract
In traditional office management, it is labor-intensive to perform real-time oversight on equipment and personnel. To address this challenge, this paper proposes a digital twin-based office management system. The system leverages the ESP8266 wireless module for device control and data collection, and employs the YOLOv5 deep learning model for real-time detection of employees’ working conditions. Additionally, a virtual office environment is constructed using the Unity engine. The system implemented herein enables real-time monitoring and analysis of office utilization, and assists managers in optimally allocating resources to enhance resource utilization efficiency by leveraging intelligent sensing and decision-making technologies. The system incurs low hardware and software costs, minimal data transmission latency, and rapid response times across its modules. Moreover, through data masking techniques, the system can protect the privacy of office personnel while enabling real-time monitoring.
Tinglong Tang, Shuifa Sun, Yirong Wu
CSCWD2
2024 Optical Flow Guided Pyramid Network for Video Salient Object Detection
abstract
Video Salient Object Detection (VSOD) is a significant pre-work for many vision applications. Different for Salient Object Detection (SOD), an effective VSOD model requires not only the spatial domain of origin image but also temporal domain. In this paper, we proposed an optical flow guided pyramid network (OFPN) for VSOD, which exploit the temporal optical flow (OF) to assist VSOD. Due to the fact that optical flow maps have slightly lower quality compared to depth maps, we designed two modules for seeking better improvement. To this end, we render optical flow maps from RGB images firstly. Then, an adaptive cross-modal attention module (ACA) is designed for multi-modal fusion. The high-level encoded features are aggregated into a shared decoder for primary prediction. Besides, the low-level features are separately sent into multi-scale context attention module (MCA) for multi-scale context fusion with the assist of the primary prediction level by level. Further, we exploit a multi-scale loss to take full advantage of the hierarchical details through image pyramid structure. Extensive experiments on five benchmark datasets demonstrate the superiority of our method against 12 state-of-the-art methods.
Tinglong Tang, Sheng Hua, Shuifa Sun, Yirong Wu, Chonghao Yue
CSCWD1
2024 Campus intelligent decision system based on digital twin
abstract
This paper presents a Unity engine-based digital twin intelligent decision-making system for campuses, aiming to improve campus management efficiency and student experience. The system combines shapefile information and tilt-shot fusion technology to achieve 3D reconstruction of campus buildings and environments, creating a digital twin model. For intelligent decision-making, we simulated a virtual energy environment and applied a reinforcement learning algorithm to address real-world energy decision challenges. Concurrently, IoT technology is used to monitor campus devices and resources, including energy utilization, security, and environmental quality. The integrated data is presented through a visual and interactive interface for real-time monitoring and management of campus resources by administrators and students. This system enhances resource utilization efficiency, sustainability, and security, showcasing the practical application of digital twin technology in modernizing campus management and improving the student experience in institutions.
Tinglong Tang, Yongjie Wu, Shuifa Sun, Yirong Wu
CSCWD1
2024 A label information fused medical image report generation framework
Shuifa Sun, Zhoujunsen Mei, Tinglong Tang, Zhanglin Su, Yirong Wu
Artif. Intell. Medicine4
2023 A Defect Detection Method Based on Parallel Multiple AutoEncoders
abstract
There is an urgent need for industrial manufacturing to fully integrate with emerging technologies to build enterprise core competitiveness. Currently, existing methods have difficulty meeting the high-precision and stability practical requirements with diversified industrial products. In this study, a PMAE (Parallel Multiple AutoEncoders, PMAE) model is proposed, which is designed with parallel multiple encoders based on the AutoEncoder framework. It uses the network of parallel multiple encoders with different encoder structures to obtain latent features that have rich and precise semantic information. A consistency objective function is proposed to make the PMAE network converge stably and rapidly, which allows the aggregated latent features to be simultaneously reconstructed and adaptively classified. Compared with the state-of-the-art methods on the NEU-CLS, Data-Crack, and DAGM2007 datasets, our method achieves the most stable performance and the highest accuracy in different defect detection tasks.
Yirong Wu, Shuifa Sun, Tinglong Tang
CSCWD4
2023 Construction Site Fence Recognition Method Based on Multi-Scale Attention Fusion ENet Segmentation Network (S)
abstract
In this paper, we propose a fence recognition method based on the ENet (Efficient neural Network) segmentation network to address the problems of traditional segmentation networks, which have poor performance in recognizing fences with a large range of scale variations and hollow structures.Firstly, a multi-scale attention fusion ENet segmentation network is designed, which is trained using the fence with obvious color features.Then, a morphological algorithm is used to process the predicted image to restore the fence segmentation results.The designed multi-scale attention fusion segmentation network performs better on fence datasets than traditional methods.In addition, the activation function Leaky_Relu6 further enhances the stability and generalization ability of the network.The experiments are conducted on 540 fence images from different construction sites, and the computed IoU is 90%.The processing speed is about 28 frames per second.The experimental results show that our proposed network outperforms traditional segmentation algorithms in fence recognition performance, and achieves robustness in different construction scenarios while meeting the requirements of both accuracy and speed.
Tinglong Tang, Yirong Wu, Tingwei Quan
SEKE2
2022 An Attention Mechanism-based Relation Network for Few-Shot Image Classification
abstract
The key to solving the few-shot image classification problem is learning image category information from a handful of image samples. Few-shot image classification can easily cause overfitting problems that impact classification performance due to insufficient labeled data. In this paper, an attention mechanism-based relation network model for few-shot image classification is proposed. Inspired by classic methods of relation networks for few-shot learning, the attention mechanism in a convolutional block attention mechanism (CBAM) for metric learning is utilized in the feature extraction network. Then, an update strategy selecting a validation set during the training is adopted to reduce the possibility of overfitting. Through comparison experiments on different datasets, the results demonstrate that our model has better accuracy than traditional methods.
Tinglong Tang, Yirong Wu
CSCWD1
2022 Feature Enhanced Graph Neural Network for Few-Shot Image Classification
abstract
There has been a rising trend of solving few-shot image classification problems utilizing graph neural networks (GNN) in recent years. However, most GNN-based approaches fail to fully exploit relationships among samples, resulting in poor image classification. To address this problem, we propose a feature-enhanced GNN model appropriate for few-shot image classification tasks. The suggested method first exploits an efficient convolutional block to generate accurate feature maps, enhancing the expressivity of the feature extraction module. Then, our technique employs flexible combinatorial distance metric functions to compute the exact relation score between samples and determine image relevance by minimizing the matching cost. Moreover, a multilayer perceptron based on the residual structure with an attention mechanism is developed to produce focused feature representations, allowing the model to obtain selective and relevant information among the samples. The proposed model is evaluated on a supervised few-shot image classification task utilizing four benchmark datasets, with the corresponding results demonstrating that our model achieves a higher accuracy performance than traditional few-shot image classification methods.
Yirong Wu, Tinglong Tang
CSCWD3
2022 Image Classification Based on Deep Graph Convolutional Networks
abstract
Due to their powerful modeling and reasoning capabilities, graph neural networks have not only achieved adequate performance in unstructured data but, in recent years, their research interests in Euclidean data such as images have also been on the rise. In this context, the most common task is image classification, whose method, based on graph neural networks, is roughly divided into two stages. The first is the graph construction stage, where the images are converted into graph structure data (composed of a node-edge-node form). The second is the graph classification stage, where the processed graph data is loaded into the graph classification network for graph classification. Subsequently, the images classification results are obtained. However, nearly two problems are faced in this setting. One is that the graph construction stage takes a long time due to the considerable computation that is required in order to convert images into graph structured data, the other problem is related to the graph classification used in the graph classification stage. The number of layers in the network tends to be limited, usually 4 layers or less. Hence, the graph’s classification accuracy is affected to some extent. Subsequently, this study proposes a deep graph neural image classification model based on gSLIC, combining the attention mechanism to conduct related experiments, in order to prove that while the speed of graph construction is greatly improved, the image classification accuracy exceeds that of most existing models based on graph neural networks for image classification.
Tinglong Tang, Xiaowang Chen, Yirong Wu, Shuifa Sun
DSAA1
2022 Memory Reconstruction Based Dual Encoders for Anomaly Detection
abstract
Anomaly detection technology relying on memory reconstruction leverages the difference in reconstruction errors between the normal and abnormal frames to achieve superior detection performance. However, there are still some challenges with this technology. First, the memory has insufficient representation capacity for features. Second, there is a contradiction between feature fusion and reconstruction. As feature fusion copies the abnormal patterns into the reconstructed frames, the abnormal frames are effectively reconstructed, reducing the detection performance. In response to these challenges, we use a memory update threshold to improve the representational power of memory. We also propose a dual-encoder anomaly detection model to restrict anomaly feature propagation. Experiment results demonstrate the effectiveness and robustness of our approach.
Yirong Wu, Qi Ren, Shuifa Sun, Tinglong Tang
SMC4
2019 Very large-scale data classification based on K-means clustering and multi-kernel SVM
Tinglong Tang, Shengyong Chen, Meng Zhao 0001, Wei Huang 0015, Jake Luo
Soft Comput.1