Jingzheng Tu

dblp:255/6071 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0003-0048-4669ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 CIVS: Communication-Aware Industrial Video Surveillance With Edge-End Collaboration
abstract
Multi-object tracking (MOT) is crucial for edge-enabled industrial video surveillance. Many video streams are delivered to the edge via limited, dynamic communication channels to be processed efficiently for accurate, on-time surveillance responses. However, time-varying communication conditions and constrained edge computational resources challenge surveillance accuracy and latency. This paper proposes an edge-based industrial video surveillance framework with bandwidth adaptation and edge-end collaboration that operates under limited communication conditions. Besides, a communication-aware industrial video surveillance with edge-end collaboration method (CIVS) is devised to improve surveillance accuracy and latency performance across different bandwidth environments. Furthermore, an NP-hard integer non-linear problem is formulated to maximize the accuracy while minimizing the latency of MOT by configuring the matches between an industrial video network and an edge computing network, which is solved by a branch and bound algorithm. Simulations show CIVS achieves 0.5892s latency and 68.9% multi-object tracking accuracy, outperforming other state-of-the-art.
Jingzheng Tu, Cailian Chen, Xin-Ping Guan
IEEE Internet Things J.1
2024 Cross-Modality Guided Multi-Scale Feature Representation Based Image Compression for Industrial Scene Understanding
abstract
Industrial surveillance videos capture subsequent video frames and deliver them to processing servers for video analytics. In such a case, image compression is a promising technology in industrial scene understanding due to its reduced transmission burden. Existing works focus on establishing macroscopic visual feature maps that contain compact image information. In this article, a cross-modality guided multi-scale feature representation based image compression method is proposed for understanding industrial scenes. Firstly, an image compressing framework with multi-grained representation (MGRIC) is proposed for industrial scene understanding, which builds image patch, structure, semantic, and signal levels feature maps. Secondly, an encoder-decoder architecture is designed for image compression with diverse granularities. Thirdly, an interaction policy between different levels is developed to ensure the subsequent level's decoding. Experiments on the CUB-200-2011 dataset demonstrate that the proposed MGRIC method achieves competitive performance, which verifies its effectiveness.
Jingzheng Tu, Cailian Chen, Shiru Zhou
INDIN1
2024 Reine: Reinspection Necessity-Based Video Collaborative Edge Caching in Smart Factory
abstract
In intelligent factories, multiple industrial cameras capture continuous videos and upload them to edge nodes for automatic preinspection. For reliability, video chunks with low preinspection accuracy must be delivered to quality inspectors for manual reinspection. Video edge caching is an urgent technology for fast and efficient manual reinspection. However, time-varying and constrained industrial network conditions cannot meet the increasing demand for video-oriented quality reinspection. This article proposes a reinspection necessity-based video collaborative edge caching method, called Reine, utilizing video superresolution (VSR) for industrial quality reinspection. First, a video edge collaborative caching framework is established for industrial quality reinspection. Second, a reinspection necessity metric is designed, and a video edge caching strategy based on VSR benefit is proposed for edge nodes. Third, an NP-hard integer nonlinear programming problem is formulated for collaborative video caching and adaptive bitrate decisions. Simulation validates Reine outperforms the state-of-the-art video edge caching methods.
Jingzheng Tu, Cailian Chen, Qimin Xu, Xin-Ping Guan
IEEE Trans. Ind. Informatics1
2023 Self-Supervised Implicit Glyph Attention for Text Recognition
abstract
The attention mechanism has become the de facto module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit attention based and supervised attention based, depended on how the attention is computed, i.e., implicit attention and supervised attention are learned from sequence-level text annotations and or character-level bounding box annotations, respectively. Implicit attention, as it may extract coarse or even incorrect spatial regions as character attention, is prone to suffering from an alignment-drifted issue. Supervised attention can alleviate the above issue, but it is character category-specific, which requires extra laborious character-level bounding box annotations and would be memory-intensive when handling languages with larger character categories. To address the aforementioned issues, we propose a novel attention mechanism for STR, self-supervised implicit glyph attention (SICA). SICA delineates the glyph structures of text images by jointly self-supervised text seg-mentation and implicit attention alignment, which serve as the supervision to improve attention correctness without extra character-level annotations. Experimental results demonstrate that SIGA performs consistently and significantly better than previous attention-based STR methods, in terms of both attention correctness and final recognition performance on publicly available context benchmarks and our contributed contextless benchmarks.
Tongkun Guan, Chaochen Gu, Jingzheng Tu, Xue Yang 0005, Yudi Zhao, Wei Shen 0002
CVPR3
2023 EdgeLeague: Camera Network Configuration With Dynamic Edge Grouping for Industrial Surveillance
abstract
Object detection is crucial for surveillance in edge-enabled Industrial Internet-of-Things. Massive high-dimensional video streams without considering priority differences connect to edges via narrow and time-varying uplink channels, which should be analyzed efficiently for accurate and fast surveillance responses. However, time-varying network environments and constrained edge resources degrade surveillance's accuracy and real-time performance. This article proposes EdgeLeague for multiple video streams with different quality of service, which maintains high surveillance performance under edge resource limitations and uplink bandwidth dynamics by edge collaboration and camera network configuration. The EdgeLeague scheme is formulated by an NP-hard integer nonlinear problem to dynamically configure camera network resolutions and detection models on cooperative edges. To accelerate configuration responses, the formulated problem is decomposed into edge league grouping, video-league matching, and video configuration, solved by low-complexity algorithms. Theoretical analysis is provided for optimal video-league matching. Simulations show EdgeLeague achieves 0.312 s latency and 86.3% surveillance accuracy.
Jingzheng Tu, Cailian Chen, Qimin Xu, Xin-Ping Guan
IEEE Trans. Ind. Informatics1
2023 PSTile: Perception-Sensitivity-Based 360$^\circ$ Tiled Video Streaming for Industrial Surveillance
abstract
360$^\circ$video becomes increasingly attractive in smart factories due to its immersive experience for industrial surveillance. However, transmitting this kind of video requires a fairly large demand on the bandwidth due to its high-resolution and panoramic view. Moreover, low-latency responses of the video to users' head movements are required. This leads to the tradeoff between high video quality and low-latency response under limited bandwidth in factories. This article proposes a perception-sensitivity (PS)-based 360$^\circ$tiled video streaming method called PSTile for industrial surveillance. Specifically, a PS tiling strategy is constructed for valid tile grouping based on the designed PS index. Then, a tile bitrate adaptation problem is formulated to allocate bitrates to valid tiles. It jointly optimizes surveillance accuracy, end-to-end latency, and quality-of-experience of users under accuracy, latency, and bandwidth constraints. Simulations demonstrate that PSTile achieves at least 32.6% lower end-to-end latency and 19.3% higher average video bitrate than grid-tiling and clus-tiling methods.
Jingzheng Tu, Cailian Chen, Ziwen Yang, Qimin Xu, Xin-Ping Guan
IEEE Trans. Ind. Informatics1
2022 Resource-Efficient Visual Multiobject Tracking on Embedded Device
abstract
Multiobject tracking (MOT) is a crucial technology for security surveillance, which is computationally intensive due to the requirement of processing a large number of video streams within low latency in practice. The input video streams of MOT are processed on a cloud computing center with abundant computational capability, posing heavy pressures on delivering video streams to the cloud. Recent advances in the Internet-of-Things (IoT) technology provide edge-computing-based solutions for video analytics at scale. However, the gap between MOT’s high computational capability demand and IoT devices’ resource-constrained nature remains significant. In this article, a resource-efficient MOT (REMOT) method is proposed for real-time surveillance on IoT embedded devices, including an affinity measurement based on an appearance model with angular triplet loss and a motion association that substitutes the time-consuming graph-based data association stage. Considering the tradeoff between latency and accuracy, we design an optimization strategy on the parallel processing of deep learning models’ layers to accelerate the inference speed with less accuracy loss. Besides, we employ a model compression strategy for model size reduction. Experiments on MOT16 and MOT17 benchmarks demonstrate that REMOT reduces 2.4$\times $latency compared with the original implementation and achieves a running speed of 81 frames per second (fps) on an embedded device with only a marginal accuracy loss (6%), which meets the requirements of real-time processing and low-latency response for surveillance.
Jingzheng Tu, Cailian Chen, Qimin Xu, Bo Yang 0006, Xin-Ping Guan
IEEE Internet Things J.1
2022 Learning-Based Scalable Scheduling and Routing Co-Design With Stream Similarity Partitioning for Time-Sensitive Networking
abstract
The deterministic and real-time communication is the indispensable requirement in Industrial Internet of Things (IIoT) application areas. Time-sensitive networking (TSN) is a promising technology for this kind of communication demands through designing proper scheduling and routing mechanisms. However, it is still challenging to design the mechanisms for large-scale instances due to high computational complexity. In order to guarantee schedulability and scalability, a learning-based scalable scheduling and routing co-design (LSSR) architecture is proposed in this article for TSN. A stream partition method combining classification and graph-based clustering is established to reduce interpartition conflicts to enhance schedulability based on the explored domain knowledge and the characterized stream data set for practical requirements. Integrated with the stream partition method, we construct the constraints of scheduling and routing co-design to guarantee the deterministic and real-time transmission. An iterative scheduling algorithm is proposed to reduce the computational complexity and thus, to enhance scalability. Simulations demonstrate the effectiveness and advantages of the proposed LSSR scheme.
Lei Xu 0043, Qimin Xu, Jingzheng Tu, Yanzhou Zhang, Cailian Chen, Xin-Ping Guan
IEEE Internet Things J.3
2022 DFR-ST: Discriminative feature representation with spatio-temporal cues for vehicle re-identification
Jingzheng Tu, Cailian Chen, Xiaolin Huang, Jianping He 0001, Xin-Ping Guan
Pattern Recognit.1
2022 Industrial Scene Text Detection With Refined Feature-Attentive Network
abstract
Detecting the marking characters of industrial metal parts remains challenging due to low visual contrast, uneven illumination, corroded surfaces, and cluttered background of metal part images. Affected by these factors, bounding boxes generated by most existing methods could not locate low-contrast text areas very well. In this paper, we propose a refined feature-attentive network (RFN) to solve the inaccurate localization problem. Specifically, we first design a parallel feature integration mechanism to construct an adaptive feature representation from multi-resolution features, which enhances the perception of multi-scale texts at each scale-specific level to generate a high-quality attention map. Then, an attentive proposal refinement module is developed by the attention map to rectify the location deviation of candidate boxes. Besides, a re-scoring mechanism is designed to select text boxes with the best rectified location. To promote the research towards industrial scene text detection, we contribute two industrial scene text datasets, including a total of 102156 images and 1948809 text instances with various character structures and metal parts. Extensive experiments on our dataset and four public datasets demonstrate that our proposed method achieves the state-of-the-art performance. Both code and dataset are available at:https://github.com/TongkunGuan/RFN.
Tongkun Guan, Chaochen Gu, Changsheng Lu, Jingzheng Tu, Kaijie Wu 0002, Xin-Ping Guan
IEEE Trans. Circuits Syst. Video Technol.4
2021 CANS: Communication Limited Camera Network Self-Configuration for Intelligent Industrial Surveillance
abstract
Realtime and intelligent video surveillance via camera networks involve computation-intensive vision detection tasks with massive video data, which is crucial for safety in the edge-enabled industrial Internet of Things (IIoT). Multiple video streams compete for limited communication resources on the link between edge devices and camera networks, resulting in considerable communication congestion. It postpones the completion time and degrades the accuracy of vision detection tasks. Thus, achieving high accuracy of vision detection tasks under the communication constraints and vision task deadline constraints is challenging. Previous works focus on single camera configuration to balance the tradeoff between accuracy and processing time of detection tasks by setting video quality parameters. In this paper, an adaptive camera network self-configuration method (CANS) of video surveillance is proposed to cope with multiple video streams of heterogeneous quality of service (QoS) demands for edge-enabled IIoT. Moreover, it adapts to video content and network dynamics. Specifically, the tradeoff between two key performance metrics, i.e., accuracy and latency, is formulated as an NP-hard optimization problem with latency constraints. A low-complexity algorithm is proposed to solve the optimization problem based on greedy searching. Simulation on real-world surveillance datasets demonstrates that the proposed CANS method achieves low end-to-end latency (13 ms on average) with high accuracy (92%) with network dynamics, which validates its effectiveness.
Jingzheng Tu, Qimin Xu, Cailian Chen
IECON1
2021 Diagnosis for IGBT Open-circuit Faults in Photovoltaic Inverters: A Compressed Sensing and CNN based Method
abstract
The inverter is the most vulnerable module of photovoltaic (PV) systems. The insulated gate bipolar transistor (IGBT) is the core part of inverters and the root source of PV inverter failures. How to effectively diagnose the IGBT faults is critical for reliability, high efficiency, and safety of PV systems. Recently, deep learning (DL) methods are widely used for fault detection and diagnosis. Different from traditional diagnosis methods, DL methods use deep neural networks which can automatically extract the useful representative features from raw data. However, DL methods require large amounts of data, which leads to the high cost of communication, storage, and computation. To tackle these issues, a data-driven fault detection and diagnosis method for IGBT open-circuit faults based on compressed sensing (CS) and convolutional neural networks (CNN) is proposed in this paper. CS is adopted to compress raw signals, and the optimal value of compression ratio (CR) is determined by considering the trade-off between classification accuracy and model training time. The overlap sampling method is adopted for data segmentation. Meanwhile, overlap sampling can also increase the number of training samples and improve the sample correlation. The compressed signals are segmented and reconstructed into two-dimensional feature maps for model training. Finally, compared with CNN of the same structure, the developed CS-CNN model can compress 85% of data without accuracy loss. The performance comparison with the state-of-the-art networks demonstrates that the test accuracy is 98.68% and the model training time is much shorter than other methods.
Bo Yang 0006, Qi Liu 0014, Jingzheng Tu, Cailian Chen
INDIN4
2020 Attention-Aware Answers of the Crowd
abstract
Crowdsourcing is a relatively economic and efficient solution to collect annotations from the crowd through online platforms. Answers collected from workers with different expertise may be noisy and unreliable, and the quality of annotated data needs to be further maintained. Various solutions have been attempted to obtain high-quality annotations. However, they all assume that workers' label quality is stable over time (always at the same level whenever they conduct the tasks). In practice, workers' attention level changes over time, and the ignorance of which can affect the reliability of the annotations. In this paper, we focus on a novel and realistic crowdsourcing scenario involving attention-aware annotations. We propose a new probabilistic model that takes into account workers' attention to estimate the label quality. Expectation propagation is adopted for efficient Bayesian inference of our model, and a generalized Expectation Maximization algorithm is derived to estimate both the ground truth of all tasks and the label-quality of each individual crowd worker with attention. In addition, the number of tasks best suited for a worker is estimated according to changes in attention. Experiments against related methods on three real-world and one semi-simulated datasets demonstrate that our method quantifies the relationship between workers' attention and label-quality on the given tasks, and improves the aggregated labels.
Jingzheng Tu, Guoxian Yu, Jun Wang 0035, Carlotta Domeniconi, Xiangliang Zhang 0001
SDM1