Minghua Zhu

dblp:73/6884 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-6000-5837ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 10 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 3 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 FQuant: Fast Quantization with Adaptive Resolution via the Clustering Algorithm
Linagwei Li, Jingfei Jiang, Jinwei Xu, Shunan Zhou, Minghua Zhu
ICIC (21)5
2025 NP4Q: Nice Point for Post-training Quantization of Object Detection Models
abstract
Post-Training Quantization (PTQ) methods have been widely applied in neural network model compression due to their ability to compress models without the need for retraining the model weights. Recently, various quantization methods for object detection models have been proposed. Unfortunately, object detection models are highly sensitive to quantization, particularly in low-bit-width scenarios, due to the high dynamic range of values. Addressing the long-tailed distribution of activations in object detection models, we propose NP4Q, a quantization framework that employs a piecewise non-uniform clustering quantization method. NP4Q is inspired by our findings that long-tailed activations can be characterized using the mean and standard deviation in a similar manner. Consequently, NP4Q utilizes a piecewise quantization strategy. Specifically, NP4Q first introduces a Nice Boundary Point (NBP) using the mean and standard deviation to partition the activations into head and tail segments. For the tail segment, a simple MinMax uniform quantization is sufficient, while for the head segment, the Meanshift Clustering Quantization (MCQ) is proposed to better handle the majority of the data. By leveraging NBP and MCQ, NP4Q can rapidly generate high-precision quantized models. Experiments demonstrate that NP4Q effectively segments the long tail and enhances the detection accuracy of quantized models, particularly in low-bit configurations. For instance, NP4Q pushes the accuracy of YOLOv5s to 52.2% and RetinaNet to 35.3% in 4-bit INT.
Shunan Zhou, Jingfei Jiang, Liangwei Li, Minghua Zhu
IJCNN5
2025 ACR: Adaptive Computation Reuse for Video Analytics in Collaborative Edge Computing
abstract
Video analytics typically requires substantial computational resources and energy. In edge computing application scenarios (e.g., smart cities), multiple users and devices may be spatially proximate, leading to offloaded tasks with high similarity and redundant computation. Computational results from previously executed tasks can be cached and reused for subsequent tasks based on similarity to enhance system efficiency. However, existing computation reuse methods often employ fixed similarity thresholds tailored to specific tasks, struggling to adapt to dynamically changing scenarios. This results in low reuse accuracy under low similarity thresholds and high latency under high similarity thresholds. Additionally, while edge servers possess larger storage capacities, they significantly increase real-time retrieval overhead. To address these issues, this paper proposes ACR, an adaptive computation reuse framework based on edge-cloud collaboration. ACR leverages the localized deployment of edge gateways to reduce cache query latency and introduces a deep Q-Network (DQN) algorithm with n-step temporal difference (TD) to adaptively adjust similarity thresholds. Our evaluation results demonstrate that, on typical urban surveillance datasets, ACR effectively balances overall system latency and reuse accuracy compared to other computation reuse methods.
Jiale Yu, Minghua Zhu
SMC2
2023 A Novel Hybrid Model Based on Deep Learning and Autoregressive for Air Quality Prediction
abstract
With the rapid development of urbanization and industrialization, air pollution problems are gradually aggravated. Efficient and accurate air quality prediction can provide a reference for air pollution control and air quality improvement. Predicting air quality by traditional methods is difficult as air quality is influenced by many complex factors, i.e., temporal and multivariate correlations. Meanwhile, under the constraints of data scale, neural network methods have difficulty capturing the long-term associations of input features. To address these challenges, we propose a novel hybrid model, namely Deep Learning and Autoregressive Network (DLARN). Specifically, our model consists of two parts: a deep learning part for analyzing the nonlinear features of air quality series, and an autoregressive module for capturing the linear trend of the series. The former combines CNN, BiLSTM and the temporal attention mechanism to learn the temporal and multivariate dependence patterns in air quality series, and the latter can address the problem of neural network methods that the output scale is insensitive to the input scale. Experiments on two real-world datasets show that our model outperforms other baselines and improves performance by 7.04% - 10.81% compared to the state-of-the-art baseline.
Minghua Zhu
IJCNN2
2023 RACCS: Real-Time Cloud-Edge Collaborative Anomaly Analysis System with Compressive Sensing Joint Optimization
abstract
In the context of the Internet of Things (IoT), ensuring public safety relies heavily on the real-time detection of anomalies in monitoring systems. Thus, video anomaly analysis has garnered significant attention. We propose a collaborative anomaly analysis system(RACCS) that leverages cloud-edge-end computing. The system receives real-time video packets generated by cameras through edge gateways and utilizes load balancing algorithms to determine the destination node, reducing system latency. Using the MQTT protocol of the IoT, the video packets are transmitted to edge nodes or the cloud for inference. To prevent highly confident video packets from being transmitted to the cloud for secondary inference, we deploy a lightweight model on the edge nodes for initial inference, thereby reducing bandwidth consumption and latency. To achieve this goal, we adopt specific quantization methods for both the input video packets and complex cloud model parameters. This converts the complex cloud model into a lightweight model, reducing model storage and improving inference speed. We also propose a joint optimization framework for compressive sensing and anomaly classification that meets both high compression rates for video packets and high accuracy for anomaly classification. We further conduct a series of experiments to determine the optimal configuration and compare our approach with other mainstream methods, demonstrating the superiority of our system.
Minghua Zhu, Dian Zhong
SMC2
2023 QueryEdge: Real-Time Muti-Video Query in Edge-Cloud Collaborative System
abstract
The real-time query of surveillance video plays a significant role in many fields such as public safety, smart city, and abnormality monitoring. However, with the exponential growth of surveillance video data, traditional cloud-based intelligent video processing faces significant challenges in terms of latency and bandwidth, while the pure edge computing approach is deficient in query accuracy due to its lack of computational power. Existing edge cloud collaboration approaches, such as SurveilEdge, focus on real-time target queries within a single video stream and do not show promising results during target queries across multiple video streams. For this reason, this paper proposes Query Edge, an edge-cloud collaborative real-time query system for multiple video streams. Specifically, we design a real-time query system based on an edge-cloud collaboration framework to achieve highly accurate and low-latency target query services in multiple video streams. In addition, we introduce a prioritization mechanism and a load-balancing strategy in the query task scheduling process to further improve query efficiency. The evaluation proves that QueryEdge has a significant improvement in query latency and bandwidth consumption compared with pure cloud computing, pure edge computing, and SurveilEdge.
Jihua Zhong, Yannian Niu, Minghua Zhu
SMC3
2023 MACEdge: Real-Time Video Analytics Based on Multi-Access Collaborative Edge Computing
abstract
Video analysis typically requires a significant amount of computing resources and energy. Traditional cloud-based video analysis relies on concentrating computing resources in the cloud, which puts a tremendous load on network bandwidth and introduces poor user experience due to latency. Edge computing offers a solution by offloading a portion of the video analysis tasks to the edge, significantly reducing the pressure on bandwidth and latency. However, the fixed video configurations (resolution, frame rate, etc.) uploaded to the edge may not be optimal for specific situations. Moreover, using a low video configuration makes it difficult for edge video analysis systems to recognize environmental changes, while using a high video configuration poses challenges to latency and bandwidth. Therefore, in this paper, we propose the MACEdge-an edge video analysis framework by leveraging the multi-access capability of the edge gateway to enhance edge perception by accessing heterogeneous sensors. We offer an algorithm based on Q Learning for dynamic adaptive adjustment of sensor threshold and video configurations. In experiments using real-world data, our collaborative solution improved accuracy and reduced high bandwidth costs compared to other video configuration adjustment benchmarks.
Dian Zhong, Minghua Zhu
SMC2
2022 Assembler: A Throughput-Efficient Module for Network Driver of Edge Computing Gateways
abstract
With the rise of IoT and edge computing, terminal data generated by end devices is also growing explosively. The limited data transmission links between edge data centers and the end devices is becoming a performance bottleneck of IoT edge-based architectures. This paper proposes Assembler, a module integrated into network drivers of the gateway to realize high-throughput forwarding in small-packet-intensive scenarios. Motivated by observations that transmitting large packets obtains far more data throughput than transmitting small ones, Assembler assembles small packets received by the driver into large ones before the driver transmits them. Meanwhile, an adaptive assembling size algorithm is introduced to balance data throughput and packet throughput in scenarios with abrupt changes in data traffic. Our evaluation shows that the network driver incorporated with Assembler can achieve 1.75x the data throughput in small-packet-intensive scenarios compared to the state-of-the-art work. Furthermore, the adaptive assembling size algorithm can help the driver adapt to scenarios with changing data traffic and keep steady throughput.
Yannian Niu, Minghua Zhu
APNOMS2
2022 A Feature Fusion Analysis Model of Heterogeneous Data Based on Tensor Decomposition
abstract
The vast volume of heterogeneous multi-source data generated by terminal devices is expanding exponentially as the Intelligent Internet of Things develops. The construction of an intelligent Internet of Things requires the extraction of crucial data features, which is from heterogeneous multi-source. However, traditional data feature extraction approaches are limited to single-source, vector space and are incapable of capturing high-dimensional nonlinear correlations of heterogeneous data. This paper innovatively proposes a feature fusion analysis model of heterogeneous data based on tensor decomposition. We optimize the tensor ring decomposition method to initially extract low-dimensional latent features from heterogeneous data effectively. We construct a tensor autoencoder network based on subspace mapping for core feature extraction of subtensor space data, and then use the Kronecker product operation to associate and fuse core data features from different subtensor spaces. After multiple fusions, the final fusion features are obtained. Experiments were carried out on the multi-modal dataset CUAVE. Experimental results show that, in terms of analysis accuracy, this model has an improvement of 18% compared with a single data source analysis model; compared with other heterogeneous fusion models, it can be increased by a maximum of 34%.
Xianyang Chu, Minghua Zhu, Hongyan Mao, Yunzhou Qiu
IJCNN2
2022 Task Offloading for Multi-Gateway-Assisted Mobile Edge Computing Based on Deep Reinforcement Learning
abstract
An effective task offloading strategy in the mobile edge computing enables terminals to migrate their tasks to the edge server, accelerating the execution of terminal tasks. However, most researches on task offloading are limited to the single edge server, while in practice, it is difficult for a single edge server to support joint offloading requests from multiple terminals. Edge gateways can be flexibly deployed around terminal devices to further reduce the computing load of edge servers. Therefore, we jointly study the task offloading problem in the multi-gateway-assisted mobile edge computing scenario. Constrained by discrete environmental variables, the offloading process jointly optimizes user scheduling, task offloading rate, and gateway resource allocation, with evaluation indexes defined by the average task delay and energy consumption. Aiming at minimizing the long-term cost of the whole system, we design a deep reinforcement learning algorithm with dynamically adjusted offloading strategies and allocated resources with only the partial state information. The simulation results demonstrate that the algorithm can obtain the optimal computation offloading policy in an uncontrollable dynamic environment. Compared with the other four benchmark algorithms, it has better system cost performance, and can quickly converge to the optimum. Meanwhile, In order to ensure the relative load balance on multiple gateways, we design a low-complexity balanced offloading strategy among multiple gateways and verify its performance.
Xianyang Chu, Minghua Zhu, Hongyan Mao, Yunzhou Qiu
SMC2
2022 Throughput-Efficient Communication Device Driver for IoT Gateways
abstract
As bridges for data exchange between end devices and cloud servers, gateways play an important role in IoT and cloud computing fields. However, with the dramatic increase of data generated by the end devices in recent years, the limited data forwarding capability of the gateway has become the bottleneck of the performance of the IoT cloud-based architectures. To efficiently improve the data throughput of links between end devices and cloud servers, this paper proposes a user-space driver for gateways that places the packet-forwarding behavior in the user space of the gateway OS and achieves high throughput by bypassing the kernel and simplifying the packet delivery path. In addition, to obtain more significant data throughput in small packet-intensive scenarios, we integrate packet assembling methods for the driver, which assembles received small packets into a large one before forwarding it. Our evaluation shows that the user-space driver has tens of times throughput compared to the kernel driver. Furthermore, our driver cooperated with the packet assembling method shows a significant advantage compared to the latest related work in small packet-intensive scenarios.
Yannian Niu, Minghua Zhu
SMC2
2022 Modeling and verifying NDN-based IoV using CSP
abstract
Abstract As a crucial component of intelligent transportation system, Internet of Vehicles (IoV) plays an important role in the smart and intelligent cities. However, current Internet architectures cannot guarantee efficient data delivery and adequate data security for IoV. Therefore, Named Data Networking (NDN), a leading architecture of Information‐Centric Networking (ICN), is introduced into IoV. Although problems about data distribution can be resolved effectively, the combination of NDN and IoV causes some new security issues. In this paper, we apply Communicating Sequential Processes (CSP) to formalize NDN‐based IoV. We mainly focus on its data access mechanism and model this mechanism in detail. By feeding the formalized model into the model checker Process Analysis Toolkit (PAT), we verify four vital properties, namely, deadlock freedom, data reliability, PIT deletion faking, and CS caching pollution. According to verification results, the model cannot ensure the security of data with the appearance of intruders. To solve these problems, we construct a blockchain‐based mechanism by creating a blockchain‐based distribution trusted platform on top of NDN‐based IoV. Through the analysis of the improved model, the blockchain‐based mechanism can truly guarantee the security of NDN‐based IoV.
Ningning Chen, Huibiao Zhu, Yuan Fei, Lili Xiao, Minghua Zhu
J. Softw. Evol. Process.6
2021 Accurate Indoor Localization Using Magnetic Sequence Fingerprints with Deep Learning
Xuedong Ding, Minghua Zhu, Bo Xiao 0004
ICA3PP (1)2
2021 An Efficient and Low-Cost FPGAs-Accelerated CNN-Based Edge Intelligent Garbage Classification System on ZYNQ
abstract
Garbage classification has achieved a high classification accuracy using convolutional neural network (CNN). However, due to the high complexity of CNNs, CPU or GPU-based solutions are too expensive for daily garbage classification. In this paper, we propose an efficient and low-cost field programmable logic gate arrays (FPGAs) accelerated CNN-based intelligent garbage classification system with two ARM+FPGA heterogeneous platforms. In this work, an image preprocessing module is deployed in the programmable logic of xc7z010clg400-1 to perform the real-time preprocessing of a garbage image. A pre-trained 18-layer CNN model is deployed in the programmable logic of xc7z020clg400-1 with the help of hardware computing module reusing to perform the garbage classification in real time. In the processing system of xc7z010clg400-1 and xc7z020clg400-1, a software to implement the scheduling of hardware is designed. According to the experiments, the system achieved a calculation speed of 21.709 times faster than the Intel(R) Core(TM) i7-7700HQ processor (CPU) and the power consumption is 1/17 of the CPU's. Its classification accuracy on the garbage dataset TrashNet is 93%, which is only 0.333% less than the pure software on the CPU. It is a cheaper solution and can also enhance the efficiency of daily garbage classification based on embedded devices. In the future work, the system we proposed will be a promising solution of the large consumption of time, energy and comnuting resource in previous solutions.
Minghua Zhu
IJCNN2
2021 A Novel Fault Diagnosis Method Based on Ensemble Feature Selection in The Industrial IoT Scenario
abstract
Fault diagnosis as a research hotspot in the field of prognostics and health management (PHM) has attracted the attention of academia and industry. Deep learning is widely used in academia due to its strong self-learning ability. However, as a "black-box" model, the features extracted by deep learning have poor interpretability which can’t be understood by IoT devices. In this paper, we propose a novel fault diagnosis method based on ensemble feature selection. We separately evaluate our method on the four-fault hydraulic components, which are derived from the real-world hydraulic time series data sets. The results show that compared with the deep learning method the proposed method can greatly reduce the complexity of the model without much reduction accuracy. Meanwhile, the balanced ensemble feature selection method we proposed is better than traditional ensemble methods such as union, intersection, and weighted linear aggregation. After testing, the method we proposed can be applied to the industrial IoT fault diagnosis.
Huadong Xu, Minghua Zhu, Bo Xiao 0004, Yunzhou Qiu
SMC2
2021 S-CNN-ESystem: An end-to-end embedded CNN inference system with low hardware cost and hardware-software time-balancing
Minghua Zhu, Xiaotong Chi, Huadong Xu
J. Syst. Archit.2
2020 Improved Model Structure with Cosine Margin OIM Loss for End-to-End Person Search
Minghua Zhu, Xuesong Cai, Jufeng Luo, Yunzhou Qiu
MMM (1)2
2020 Joint RFID and UWB Technologies in Intelligent Warehousing Management System
abstract
Joint application of radio-frequency identification (RFID) and ultrawideband (UWB) technologies in the intelligent warehousing management system is proposed. In this system, we regard forklift as infrastructure, and both the UWB mobile terminal (MT) and the RFID reader are mounted on the forklift. The RFID reader is used not only to read the information of goods but also to determine the goods' status of loading and unloading. The UWB MT is used to locate the forklift. The goods or pallets are labeled with RFID tags. Utilizing the integration of these two technologies, the dual goals of goods information and goods location perception are achieved. An M/N-K sliding window method is proposed to determine the loading and unloading of goods in this article. Our experiments reveal that this novel method can quickly, accurately, and efficiently determine the states of the goods on the forklift. For the indoor localization, an algorithm is proposed based on RSS residual weighting (RRW). Experiments show that RRW can mitigate the nonline-of-sight error substantially compared with the conventional Taylor algorithm and recent Convex approximation algorithm. Finally, a real working practice in a warehouse of a company is introduced, and it illustrates the feasibility of the system.
Minghua Zhu, Bo Xiao 0004, Xuguang Yang, Changlei Gong
IEEE Internet Things J.2
2019 Accurate Magnetic Object Localization Using Artificial Neural Network
abstract
Magnetic object localization technology is emerging to bring substantial benefits for human medical investigation, such as tracking wireless endoscopic devices or catheters with embedded permanent magnets. Towards accurate magnetic localization, we propose a novel method of fitting the magnetic field intensity and magnetic gradient tensor (MGT) with artificial neural network (ANN) for magnetic object localization, which permits accurate localization of the magnetic object in a certain range. In the simulation, we analyzed and compared the capability of the proposed method and a traditional method based on tensor module gradient. The results show that the average localization error of the proposed method is at least 8 times more accurate than the traditional method in the noisy environment. Besides, a magnetic localization system based on a sensor array was used to carry out a localization experiment and the result has shown the feasibility of the proposed method with the average localization error of 2.15 cm when the magnetic target distance was from 61 to 79 cm.
Shengzhi Chen, Minghua Zhu, Xuesong Cai, Bo Xiao 0004
MSN2
2018 A Hardware/Software Co-design Approach for Real-Time Binocular Stereo Vision Based on ZYNQ (Short Paper)
Yukun Pan, Minghua Zhu, Jufeng Luo, Yunzhou Qiu
CollaborateCom2
2018 RFID Based Motion Direction Estimation in Gate Systems
abstract
The RFID based school gate system is used to estimate students entering or leaving the school when they go through the school gate with RFID tags. In general, the accuracy of the estimation of RFID is sensitive to complex electromagnetic environment changing. For example, the estimation success rate is high in some circumstance but low when there is a car parking in front of directional readers. In this paper, a method is presented by which the direction of the readers would be automatically adjusted according to the received information of the tags. With this method, the accuracy of the estimation of RFID could maintain high even if the electromagnetic environment changes. Experiments also were carried out to evaluate the feasibility of the proposed method. The results showed that the method is highly suitable to keep a stable success rate of motion direction estimation in school gate systems.
Jie Wu 0018, Minghua Zhu, Bo Xiao 0004
CSCWD2
2018 Graph-Based Indoor Localization with the Fusion of PDR and RFID Technologies
Jie Wu 0018, Minghua Zhu, Bo Xiao 0004, Yunzhou Qiu
ICA3PP (3)2
2018 The Improved Fingerprint-Based Indoor Localization with RFID/PDR/MM Technologies
abstract
The fingerprint based indoor localization is becoming a dominant solution for its high applicability in complex indoor environment. However, the extensive site survey efforts on manpower and time have become a major bottleneck. Based on the crowdsourcing method, the paper puts forward a novel indoor localization with the fusion of RFID (radio frequency identification devices), PDR (pedestrian dead reckoning)and MM (magnetic matching)technologies. First, a zero-effort fingerprint automated construction and site survey update scheme is proposed with the dual-frequency RFID. Second, in order to solve the problem that step length would vary from person to person which results into positioning bias, the RSS technology and floor map is introduced to aid step length estimation. Third, the particle filter is adopted to fuse the advantages of three different technologies to conduct the accurate indoor positioning. In the experiment, the obtained fingerprint database is demonstrated to possess a comparable accuracy with the human-annotated database. Also, the experiment results show that the proposed method achieves the comparable positioning accuracy and the positioning accuracy is 40% higher than the PDR system or RFID positioning system alone.
Jie Wu 0018, Minghua Zhu, Bo Xiao 0004, Yunzhou Qiu
ICPADS2
2017 MAC-ILoc: Multiple Antennas Cooperation Based Indoor Localization Using Cylindrical Antenna Arrays
Jie Wu 0018, Minghua Zhu, Bo Xiao 0004
CollaborateCom2
2016 A deep tongue image features analysis model for medical application
abstract
With the improvement of people's living standards, there is no doubt that people are paying more and more attention to their health. However, shortage of medical resources is a critical global problem. As a result, an intelligent prognostics system has a great potential to play important roles in computer aided diagnosis. Numerous papers reported that tongue features have been closely related to a human's state. Among them, the majority of the existing tongue image analyses and classification methods are based on the low-level features, which may not provide a holistic view of the tongue. Inspired by a deep convolutional neural network (CNN), we propose a deep tongue image feature analysis system to extract unbiased features and reduce human labor for tongue diagnosis. With the unbalanced sample distribution, it is hard to form a balanced classification model based on feature representations obtained by existing low-level and high-level methods. Our proposed deep tongue image feature analysis model learns high-level features and provide more classification information during training time, which may result in higher accuracy when predicting testing samples. We tested the proposed system on a set of 267 gastritis patients, and a control group of 48 healthy volunteers (labeled according to Western medical practices). Test results show that the proposed deep tongue image feature analysis model can classify a given tongue image into healthy and diseased state with an average accuracy of 91.49%, which demonstrates the relationship between human body's state and its deep tongue image features.
Guitao Cao, Ye Duan, Minghua Zhu, Liping Tu, Jiatuo Xu, Dong Xu 0002
BIBM4