Jiaqi Wu 0012

dblp:214/8039-12 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-1587-163XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real-Time Underground Fire Detection on Coal Mine IoVT Systems: An Edge-Deployed Efficient YOLO-Architecture
abstract
Underground fires pose a significant threat to production safety in coal mines, and existing detection methods suffer from drawbacks such as poor adaptability to complex subterranean environments and excessive model parameters. To address the need for deploying object detection models on resource-constrained devices, this paper proposes a novel and efficient algorithm forUndergroundFireYOLOdetection, named UF-YOLO. The core innovation of this method is threefold: first, the StarNet module is introduced into the backbone to significantly reduce model parameters and computational complexity without sacrificing accuracy; second, the Cross-scale Context Fusion Module (CCFM) is integrated into the neck to enhance the model’s detection capability for fires of various scales, particularly small targets; and finally, Partial Convolution (PConv) is integrated to extract spatial features more efficiently, further reducing redundant computations and memory access. On our self-built Mine Fire Image Dataset (MFID), compared to the baseline model YOLOv11m, UF-YOLO reduces parameters by 77.1%, increases inference speed by 60.6%. Experimental results on the public COCO val 2017 dataset demonstrate that the proposed method outperforms state-of-the-art (SOTA) models such as YOLOv12. The results confirm that UF-YOLO can be efficiently deployed on the edge-side of coal mine IoVT monitoring systems to performe accurate and real-time fire detection. This work provides a new intelligent paradigm for the real-time monitoring of underground fires.
Wei Yang 0063, Jiaqi Wu 0012, Zehua Wang 0001, Qi-Chong Tian, Tao Ye 0002, Wei Chen 0036, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.3
2026 Efficient Detection Framework Adaptation for Edge Computing: A Plug-and-Play Neural Network Toolbox Enabling Edge Deployment
abstract
Recently, edge computing has emerged as a prevailing paradigm in applying deep learning-based object detection models, offering a promising solution for time-sensitive tasks. However, existing edge object detection faces several challenges: 1) These methods struggle to balance detection precision and model lightweightness. 2) Existing generalized edge-deployment designs offer limited adaptability for object detection. 3) Current works lack real-world evaluation and validation. To address these challenges, we propose theEdgeDetectionToolbox(ED-TOOLBOX), which leverages generalizable plug-and-play components to enable edge-site adaptation of object detection models. Specifically, we propose a lightweightReparameterized Dynamic Convolutional Network(Rep-DConvNet) that employs a weighted multi-shape convolutional branch structure to enhance detection performance. Furthermore, ED-TOOLBOX includes aSparse Cross-Attention(SC-A) network that adopts a localized-mapping-assisted self-attention mechanism to facilitate a well-craftedJoint Modulein adaptively transferring features for further performance improvement. Moreover, we propose anEfficient Headfor the classification and location modules to achieve more efficient prediction. Additionally, in practical industrial scenarios, we identify that helmet detection-one of the most representative edge object detection tasks-overlooks band fastening, which introduces potential safety hazards. To address this, we build aHelmet Band Detection Dataset(HBDD) and apply an edge object detection model optimized by the ED-TOOLBOX to tackle this real-world task. Extensive experiments validate the effectiveness of components in ED-TOOLBOX. In visual surveillance simulations, ED-TOOLBOX-assisted edge detection models outperform sixstate-of-the-artmethods, enabling real-time and accurate detection. These results demonstrate that our approach offers a superior solution for edge object detection.
Jiaqi Wu 0012, Lixu Wang, Zehua Wang 0001, Wei Chen 0036, Fangyuan He, F. Richard Yu, Victor C. M. Leung
IEEE Trans. Mob. Comput.1
2025 FDPT: Federated Discrete Prompt Tuning for Black-Box Visual-Language Models
Jiaqi Wu 0012, Yuzhe Yang 0002, Lixu Wang, Zehua Wang 0001, Wei Chen 0036
ICCV1
2025 CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT Under Uneven Illumination
abstract
Clear monitoring images are crucial for the safe operation of coal mine Internet of Video Things (IoVT) systems. However, low illumination and uneven brightness in underground environments significantly degrade image quality, posing challenges for enhancement methods that often rely on difficult-to-obtain paired reference images. Additionally, there is a tradeoff between enhancement performance and computational efficiency on edge devices within IoVT systems.To address these issues, we propose a multimodal image enhancement method tailored for coal mine IoVT, utilizing an ISP operations within a differentiable CNN framework fusion architecture optimized for uneven illumination. This two-stage strategy combines global enhancement with detail optimization, effectively improving image quality, especially in poorly lit areas. A contrastive language-image pretraining (CLIP)-based multimodal iterative optimization allows for unsupervised training of the enhancement algorithm. By integrating traditional image signal processing (ISP) with convolutional neural networks (CNN), our approach reduces computational complexity while maintaining high performance, making it suitable for real-time deployment on edge devices. Experimental results demonstrate that our method effectively mitigates uneven brightness and enhances key image quality metrics, with preservation of original visual information (PSNR) improvements of 2.9%–4.9%, structural similarity (SSIM) by 4.3%–11.4%, and visual information fidelity (VIF) by 4.9%–17.8% compared to seven state-of-the-art algorithms. Simulated coal mine monitoring scenarios validate our method’s ability to balance performance and computational demands, facilitating real-time enhancement and supporting safer mining operations.
Shuai Wang 0039, Jiaqi Wu 0012, Wei Chen 0036, Tongzhu Jin, Miaomiao Xue, Zehua Wang 0001, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.3
2025 A self-supervised enhancement method for real world low-light images using Retinex and camera response function
Wei Yang 0063, Shuai Wang 0039, Jiaqi Wu 0012, Wei Chen 0036
Multim. Syst.3
2025 CLIP-AE: A Multi-Modal Unsupervised Images Enhancement Method Based on High-Order Adaptive Curve for Visual Disbalance Defects
abstract
For visual disbalance defects (VDDs) in low-light images, such as brightness unevenness and color imbalance, existing enhancement methods struggle to extract defect features from local regions and apply adaptive enhancement based on varying degrees of these defects. To address these challenges, we propose an unsupervised multi-modal enhancement method based on a high-order adaptive curve, named CLIP-AE. Specifically, we introduce a multi-modal recurrent optimization approach utilizing contrastive language-image pre-training (CLIP). This method iteratively optimizes variable embedded prompts and an Adaptive Enhancement Module (AEM) to establish dependencies between the prompts and detailed style features in the images, guiding the AEM to perform adaptive image enhancement. Additionally, we implement a progressive feature alignment strategy to enhance the model's ability to perceive style features and improve optimization efficiency by using multiple enhanced images with identical content features and incremental style features. In the AEM, the optimized Hyperparameters Generative Network (HGN) generates the optimal hyperparameters, which drive a High-Dimensional Nested Gamma correction (HDN-Gamma) to perform pixel-wise adaptive enhancement for VDDs. HDN-Gamma further maps pixel values using specific enhancement curves to avoid artifacts. Extensive experiments demonstrate that our method effectively improves visual disbalance defects and reduces artifacts. Compared to seven state-of-the-art algorithms, our method shows significant improvements (PSNR: 16.46%, 16.89%, and 15.14%; SSIM: 9.26%, 8.02%, and 9.85%; MUSIQ: 6.37%, 6.54%, and 7.45%) on the LOL, SICE, and MIT-Adobe FiveK datasets. Our approach offers a novel solution for applying multimedia technology in low-light image enhancement tasks.
Jiaqi Wu 0012, Mingshuo Hou, Zehua Wang 0001, Wei Chen 0036, F. Richard Yu, Victor C. M. Leung
IEEE Trans. Multim.1
2024 FedSAR for Heterogeneous Federated learning:A Client Selection Algorithm Based on SARSA
Dufeng Chen, Rui Jing, Jiaqi Wu 0012, Zehua Wang 0001, Fan Zhang 0057, Wei Chen 0036
ICIC (1)3
2024 A TransISP Based Image Enhancement Method for Visual Disbalance in Low-light Images
abstract
Abstract Existing image enhancement algorithms often fail to effectively address issues of visual disbalance, such as brightness unevenness and color distortion, in low‐light images. To overcome these challenges, we propose a TransISP‐based image enhancement method specifically designed for low‐light images. To mitigate color distortion, we design dual encoders based on decoupled representation learning, which enable complete decoupling of the reflection and illumination components, thereby preventing mutual interference during the image enhancement process. To address brightness unevenness, we introduce CNNformer, a hybrid model combining CNN and Transformer. This model efficiently captures local details and long‐distance dependencies between pixels, contributing to the enhancement of brightness features across various local regions. Additionally, we integrate traditional image signal processing algorithms to achieve efficient color correction and denoising of the reflection component. Furthermore, we employ a generative adversarial network (GAN) as the overarching framework to facilitate unsupervised learning. The experimental results show that, compared with six SOTA image enhancement algorithms, our method obtains significant improvement in evaluation indexes (e.g., on LOL, PSNR: 15.59%, SSIM: 9.77%, VIF: 9.65%), and it can improve visual disbalance defects in low‐light images captured from real‐world coal mine underground scenarios.
Jiaqi Wu 0012, Rui Jing, Wei Chen 0036, Zehua Wang 0001
Comput. Graph. Forum1
2024 PrFu-YOLO: A Lightweight Network Model for UAV-Assisted Real-Time Vehicle Detection Toward an IoT Underlayer
abstract
With the rapid development of Internet of Things (IoT) and UAV technology, for the whole IoT system of vehicle detection, the middle and high level of information transmission and server processing has made a breakthrough, so at the bottom of the real-time detection of the vehicle by the UAV is the key to the whole system. However, UAV vehicle detection faces the challenges of too many small targets in the image leading to low detection accuracy, limited hardware platform resources requiring control of model size, and real-time detection requiring high inference speed. Aiming at the above problems, we propose a lightweight model PrFu-YOLO based on YOLOv8 improvement, which achieves a good balance between the accuracy, inference speed, and model size. And it realizes real-time vehicle detection embedded in an UAV platform. To solve the problem of low vehicle detection accuracy, we design a new structure PrFuFPN based on adding a small target detection layer to achieve more advanced feature fusion. To address the limited resources of the platform and the problem of real-time vehicle detection, we add GhostConv to the structure and constantly try to adjust the parameters of the network. Finally, extensive experiments were conducted on the VisDrone2019 and CARPK data sets to fully evaluate the model. Compared to YOLOv8s on the VisDrone2019 test set, mAP50 was improved by 10.05%, mA95 by 14.54%, the number of parameters was reduced by 8.13%, and the model size was reduced by 12.5%, while the FPS of 67.13 fully met the needs of real-time detection.
Haishun Liu, Jiaqi Wu 0012, Wei Chen 0036, Ruihan Zheng, Zehua Wang 0001
IEEE Internet Things J.3
2024 Small Insulator Defects Detection Based on Multiscale Feature Interaction Transformer for UAV-Assisted Power IoVT
abstract
The power inspection is an important application of UAV-assisted power internet of video things (IoVT) for maintaining the safety of the power system. Due to the limitations of distance and angle, the resolution of the images captured by UAV is low, which seriously impacts the effects of small insulator defects detection. To address this problem, we propose a small-size defects detection method based on multi-scale feature interaction transformer for UAV-assisted Power IoVT. For the algorithm, we design a super-resolution reconstruction-assisted small object detection algorithm, the super-resolution module generates high-resolution images with the requirements of object detection function, which greatly improves the small object detection performance. Moreover, we design multi-scale feature interaction transformer network (MFITN), compared with the traditional non-local attention mechanism, the network structure can capture dependencies in multi-scales features, furthermore, the advantage assist the super-resolution module to generate more realistic image information to further improve small object detection. In addition, we propose a distributed model deployment strategy to deploy our high computational complexity algorithm in the edge side of the IoVT system, which can drive the overall algorithm to perform low-latency edge computation by relying only on the limited computing power devices. Experiments demonstrate that our method has better small object detection performance (mAP=81.3%, FPS=49.7), the super-resolution reconstruction is able to recover more realistic detail information, the distributed computing method can reduce the response latency by 33.4%-87.2%, which all contribute UAV-assisted Power IoVT system to realize accurate and fast power insulator defects detection.
Jiaqi Wu 0012, Rui Jing, Yishuo Bai, Wei Chen 0036, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.1
2024 A Lightweight Small Object Detection Method Based on Multilayer Coordination Federated Intelligence for Coal Mine IoVT
abstract
Video surveillance as an important function of internet of video things (IoVT) system has been widely used in coal mine monitoring for coal mine safety with excellent results, however, there are still many shortcomings: 1) Existing coal mine IoVT systems have limited detection accuracy for small-sized objects; 2) Coal mine video surveillance systems generally adopt centralized cloud computing, transmission of massive data causes high latency, which seriously affects the response speed of object detection function; 3) The concept drift caused by the data stream seriously affect the detection effect of the offline algorithm. To address the above issues, we propose a small object detection method based federated intelligence to assist coal mine IoVT for object detection. First, we design a lightweight neural network Rep-ShuffleNet to improve YOLOv8, the state-of-the-art YOLO algorithm, to maintain high detection accuracy while dramatically increasing the inference speed, and with the advantage of lightweight, it can be deployed to embedded devices for low-latency edge computing; Moreover, we design a federated learning-based MLC-FL algorithm for local algorithms’ automatic and efficient optimization by asynchronous communication and data interaction reduction strategy. The experimental results show that with the assistance of federated intelligence model optimization strategies, the lightweight YOLOv8 has excellent detection performance (mAP: 94.6%, APsmall: 86.7%, FPS: 21.6), thus to assist coal mine IoVT to realize accurate and real-time underground small object detection.
Jiaqi Wu 0012, Ruihan Zheng, Jiade Jiang, Wei Chen 0036, Zehua Wang 0001, F. Richard Yu, Victor C. M. Leung
IEEE Internet Things J.1