Zhuwei Qin

dblp:211/0080 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-5465-7740ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 EdgeEMG: On-Device Neural Network Training for Real-Time EMG Pattern Recognition
abstract
Surface electromyography (sEMG) is a non-invasive technique that records bioelectrical signals generated by muscle activity via electrodes placed on the skin. Its ability to capture a user's motor intent in real time has enabled a wide range of applications, including prosthetic control, rehabilitation robotics, and human-computer interaction. Recent advances in machine learning (ML), particularly deep learning (DL), have enabled automated processing of complex biosignals. While DL-based approaches for sEMG gesture recognition have shown strong performance on embedded systems, they typically rely on pretraining models on high-performance computing platforms (e.g., PCs or supercomputers) before deployment to low-power devices. This off-device pretraining limits portability and adaptability, as it requires prior collection and processing of sEMG data on nonportable hardware. In this paper, we present EdgeEMG, the first fully on-device training approach for sEMG gesture recognition using a deep neural network. Our approach is implemented using the AIfES library, developed by the Fraunhofer Institute, which provides efficient operations for feedforward neural networks and supports deployment on resource-constrained microcontrollers. We validate our system on the Sony Spresense MCU and benchmark it against a conventional Linear Discriminant Analysis (LDA) classifier. Experimental results demonstrate that our approach achieves an average real-time accuracy of 70% across different hand gestures. These findings highlight the feasibility of real-time, user-adaptive EMG decoding entirely on embedded hardware, without reliance on external compute resources.
Jiahua Tang, Philip Liang, Zhuwei Qin
BSN3
2022 Privacy-preserving federated learning for transportation mode prediction based on personal mobility data
abstract
Personal daily mobility trajectories/traces like Google Location Service integrates many valuable information from individuals and could benefit a lot of application scenarios, such as pandemic control and precaution, product recommendation, customized user profile analysis, traffic management in smart cities, etc. However, utilizing such personal mobility data faces many challenges since users’ private information, such as home/work addresses, can be unintentionally leaked. In this work, we build an FL system for transportation mode prediction based on personal mobility data. Utilizing FL-based training scheme, all user’s data are kept in local without uploading to central nodes, providing high privacy preserving capability. At the same time, we could train accurate DNN models that is close to the centralized training performance. The resulted transportation mode prediction system serves as a prototype on user’s traffic mode classification, which could potentially benefit the transportation data analysis and help make wise decisions to manage public transportation resources.
Fuxun Yu, Zhuwei Qin, Xiang Chen 0010
High Confid. Comput.3
2022 CaptorX: A Class-Adaptive Convolutional Neural Network Reconfiguration Framework
abstract
Nowadays, the evolution of deep learning and cloud service significantly promotes neural network-based mobile applications. Although intelligent and prolific, those applications still lack certain flexibility: for classification tasks, neural networks are generally trained with vast classification targets to cover various utilization contexts. However, only partial classes are practically inferred due to individual mobile user preference and application specificity, which causes unnecessary computation consumption. Thus, we proposedCaptorX—a class-adaptive convolutional neural network (CNN) reconfiguration framework to adaptively prune convolutional filters associated with unneeded classes.CaptorXcan reconfigure a pretrained full-class CNN model into class-specific lightweight models based on the visualization analysis of convolutional filters’ exclusive functionality for a single class. These lightweight models can be directly deployed to mobile devices without the retraining cost of traditional pruning-based reconfiguration. Furthermore, we can apply theCaptorXframework into a distributed collaboration setting. With dedicated local training regulation and collaborative aggregation schemes, the class-adaptive models on individual mobile devices can further contribute back to the central full-class model. Experiments on representative CNNs and image classification datasets show that,CaptorXcan reduce the CNN computation workload up to 50.22% and save 46.58% energy consumption for varied local devices, meanwhile improving accuracy for their targeted classes with better task focus. With our distributed collaboration paradigm,CaptorXalso provides further potential to enhance the central model accuracy, while reducing up to 37.58% communication cost compared to traditional distributed learning methods.
Zhuwei Qin, Fuxun Yu, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 Fed2: Feature-Aligned Federated Learning
abstract
Federated learning learns from scattered data by fusing collaborative models from local nodes. However, conventional coordinate-based model averaging by FedAvg ignored the random information encoded per parameter and may suffer from structural feature misalignment. In this work, we propose Fed2, a feature-aligned federated learning framework to resolve this issue by establishing a firm structure-feature alignment across the collaborative models. Fed2 is composed of two major designs: First, we design a feature-oriented model structure adaptation method to ensure explicit feature allocation in different neural network structures. Applying the structure adaptation to collaborative models, matchable structures with similar feature information can be initialized at the very early training stage. During the federated learning process, we then propose a feature paired averaging scheme to guarantee aligned feature distribution and maintain no feature fusion conflicts under either IID or non-IID scenarios. Eventually, Fed2 could effectively enhance the federated learning convergence performance under extensive homo- and heterogeneous settings, providing excellent convergence speed, accuracy, and computation/communication efficiency.
Fuxun Yu, Weishan Zhang, Zhuwei Qin, Di Wang 0003, Zhi Tian, Xiang Chen 0010
KDD3
2021 DiReCtX: Dynamic Resource-Aware CNN Reconfiguration Framework for Real-Time Mobile Applications
abstract
Although convolutional neural networks (CNNs) have been widely applied in various cognitive applications, they are still very computationally intensive for resource-constrained mobile systems. To reduce the resource consumption of CNN computation, many optimization works have been proposed for mobile CNN deployment. However, most works are merely targeting CNN model compression from the perspective of parameter size or model structure, ignoring different resource constraints in mobile systems with respect to memory, energy, and real-time requirement. Moreover, previous works take accuracy as their primary consideration, requiring a time-costing retraining process to compensate the inference accuracy loss after compression. To address these issues, we propose DiReCtX-a dynamic resource-aware CNN model reconfiguration framework. DiReCtX is based on a set of accurate CNN profiling models for different resource consumption and inference accuracy estimation. With manageable consumption/accuracy tradeoffs, DiReCtX can reconfigure a CNN model to meet distinct resource constraint types and levels with expected inference performance maintained. To further achieve fast model reconfiguration in real-time, improved CNN model pruning and its corresponding accuracy tuning strategies are also proposed in DiReCtX. The experiments show that the proposed CNN profiling models can achieve 94.6% and 97.1% accuracy for CNN model resource consumption and inference accuracy estimation. Meanwhile, the proposed reconfiguration scheme of DiReCtX can achieve at most 44.44% computation acceleration, 31.69% memory reduction, and 32.39% energy saving, respectively. On field-tests with state-of-the-art smartphones, DiReCtX can adapt CNN models to various resource constraints in mobile application scenarios with optimal real-time performance.
Fuxun Yu, Zhuwei Qin, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 REIN the RobuTS: Robust DNN-Based Image Recognition in Autonomous Driving Systems
abstract
In recent years, the neural network (NN) has shown its great potential in image recognition tasks of autonomous driving systems, such as traffic sign recognition, pedestrian detection, etc. However, theoretically well-trained NNs usually fail their performance when facing real-world scenarios. For example, adverse real-world conditions, e.g., bad weather and lighting conditions, can introduce different physical variations and cause considerable accuracy degradation. As for now, the generalization capability of NNs is still one of the most critical challenges for the autonomous driving system. To facilitate the robust image recognition tasks, in this work, we build the RobuTS dataset: a comprehensive Robust Traffic Sign Recognition dataset, which includes images with different environmental variations, e.g., rain, fog, darkening, and blurring. Then to enhance the NN's generalization capability, we propose two generalization-enhanced training schemes: 1) REIN for robust training without data in adverse scenarios and 2) Self-Teaching (ST) for robust training with unlabeled adverse data. The great advantages of such two training schemes are they are data-free (REIN) and label-free (ST), thus effectively reducing the huge human efforts/cost of on-road driving data collection, as well as the expensive manual data annotation. We conduct extensive experiments to validate our methods' performance on both classification and detection tasks. For classification tasks, our proposed training algorithms could consistently improve model performance by +15%-25% (REIN) and +16%-30% (ST) in all adverse scenarios of our RobuTS datasets. For detection tasks, our ST could also improve the detector's performance by +10.1 mean average precision (mAP) on Foggy-Cityscapes, outperforming previous state-of-the-art works by +2.2 mAP.
Fuxun Yu, Zhuwei Qin, Di Wang 0003, Xiang Chen 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2020 DC-CNN: Computational Flow Redefinition for Efficient CNN through Structural Decoupling
abstract
Recently Convolutional Neural Networks (CNNs) are widely applied into novel intelligent applications and systems. However, the CNN computation performance is significantly hindered by its computation flow, which computes the model structure sequentially by layers with massive convolution operations. Such a layer-wise sequential computation flow can cause certain performance issues, such as resource under-utilization, huge memory overhead, etc. To solve these problems, we propose a novel CNN structural decoupling method, which could decouple CNN models into "critical paths" and eliminate the inter-layer data dependency. Based on this method, we redefine the CNN computation flow into parallel and cascade computing paradigms, which can significantly enhance the CNN computation performance with both multi-core and single-core CPU processors. Experiments show that, our DC-CNN framework could reduce 24% to 33% latency on multi-core CPUs for CIFAR and ImageNet. On small-capacity mobile platforms, cascade computing could reduce the latency by average 24% on ImageNet and 42% on CIFAR10. Meanwhile, the memory reduction could also reach average 21% and 64%, respectively.
Fuxun Yu, Zhuwei Qin, Di Wang 0003, Ping Xu 0002, Zhi Tian, Xiang Chen 0010
DATE2
2020 Exploring Decentralized Collaboration in Heterogeneous Edge Training
abstract
Recent progress in deep learning techniques enabled collaborative edge training, which usually deploys identical neural network models globally on multiple devices for aggregating parameter updates over distributed data collection. However, as more and more heterogeneous edge devices are involved in practical training, the identical model deployment over collaborative edge devices cannot be guaranteed: On one hand, the weak edge devices with less computation resources may not catch up stronger ones' training progress, and appropriate local model training customization is necessary to balance the collaboration. On the other hand, a particular local edge device may have specific learning task preference, while the global identical model would exceed the practical local demand and cause unnecessary computation cost. Therefore, we explored the collaborative learning with heterogeneous convolutional neural networks (CNNs) in this work, expecting to address aforementioned real problems. Specifically, we proposed a novel decentralized collaborative training method by decoupling a training target CNN model into independently trainable sub-models correspond to a sub-set of learning tasks for each edge device. After sub-models are well-trained on edge nodes, the model parameters for individual learning tasks can be harvested from local models on every edge device and ensemble the global training model back to a single piece. Experiments demonstrate that, for the AlexNet and VGG on the CIFAR10, CIFAR100 and KWS dataset, our decentralized training method can save up to 11.8× less computation load while achieve central sever test accuracy.
Xiang Chen 0010, Zhuwei Qin
SEC2
2020 Towards Context-aware Distributed Learning for CNN in Mobile Applications
abstract
Intelligent mobile applications have been ubiquitous on mobile devices. These applications keep collecting new and sensitive data from different users while being expected to have the ability to continually adapt the embedded machine learning model to these newly collected data. To improve the quality of service while protecting users' privacy, distributed mobile learning (e.g., Federated Learning (FedAvg) [1]) has been proposed to offload model training from the cloud to the mobile devices, which enables multiple devices collaboratively train a shared model without leaking the data to the cloud. However, this design becomes impracticable when training the machine learning model (e.g., Convolutional Neural Network (CNN)) on mobile devices with diverse application context. For example, in conventional distributed training schemes, different devices are assumed to have integrated training datasets and train identical CNN model structures. Distributed collaboration between devices is implemented by a straightforward weight average of each identical local models. While, in mobile image classification tasks, different mobile applications have dedicated classification targets depending on individual users' preference and application specificity. Therefore, directly averaging the model weight of each local model will result in a significant reduction of the test accuracy. To solve this problem, we proposed CAD: a context-aware distributed learning framework for mobile applications, where each mobile device is deployed with a context-adaptive submodel structure instead of the entire global model structure.
Zhuwei Qin, Hao Jiang 0014
SEC1
2020 Enabling efficient ReRAM-based neural network computing via crossbar structure adaptive optimization
abstract
Resistive random-access memory (ReRAM) based accelerators have been widely studied to achieve efficient neural network computing in speed and energy. Neural network optimization algorithms such as sparsity are developed to achieve efficient neural network computing on traditional computer architectures such as CPU and GPU. However, such computing efficiency improvement is hindered when deploying these algorithms on the ReRAM-based accelerator because of its unique crossbar-structural computations. And a specific algorithm and hardware co-optimization for the ReRAM-based architecture is still in a lack. In this work, we propose an efficient neural network computing framework that is specialized for the crossbar-structural computations on the ReRAM-based accelerators. The proposed framework includes a crossbar specific feature map pruning and an adaptive neural network deployment. Experimental results show our design can improve the computing accuracy by 9.1% compared with the state-of-the-art sparse neural networks. Based on a famous ReRAM-based DNN accelerator, the proposed framework demonstrates up to 1.4× speedup, 4.3× power efficiency, and 4.4× area saving.
Fuxun Yu, Zhuwei Qin, Xiang Chen 0010
ISLPED3
2019 CAPTOR: a class adaptive filter pruning framework for convolutional neural networks in mobile applications
abstract
Nowadays, the evolution of deep learning and cloud service significantly promotes neural network based mobile applications. Although intelligent and prolific, those applications still lack certain flexibility: For classification tasks, neural networks are generally trained online with vast classification targets to cover various utilization contexts. However, only partial classes are practically tested due to individual mobile user preference and application specificity. Thus the unneeded classes cause considerable computation and communication cost. In this work, we propose CAPTOR - a class-level reconfiguration framework for Convolutional Neural Networks (CNNs). By identifying the class activation preference of convolutional filters through feature interest visualization and gradient analysis, CAPTOR can effectively cluster and adaptively prune the filters associated with unneeded classes. Therefore, CAPTOR enables class-level CNN reconfiguration for network model compression and local deployment on mobile devices. Experiment shows that, CAPTOR can reduce computation load for VGG-16 by up to 40.5% and 37.9% energy consumption with ignored loss of accuracy. For AlexNet, CAPTOR also reduces computation load by up to 42.8% and 37.6% energy consumption with less than 3% loss in accuracy.
Zhuwei Qin, Fuxun Yu, Xiang Chen 0010
ASP-DAC1
2019 Functionality-Oriented Convolutional Filter Pruning
Zhuwei Qin, Fuxun Yu, Xiang Chen 0010
BMVC1
2019 Interpreting and Evaluating Neural Network Robustness
abstract
Recently, adversarial deception becomes one of the most considerable threats to deep neural networks. However, compared to extensive research in new designs of various adversarial attacks and defenses, the neural networks' intrinsic robustness property is still lack of thorough investigation. This work aims to qualitatively interpret the adversarial attack and defense mechanisms through loss visualization, and establish a quantitative metric to evaluate the model's intrinsic robustness. The proposed robustness metric identifies the upper bound of a model's prediction divergence in the given domain and thus indicates whether the model can maintain a stable prediction. With extensive experiments, our metric demonstrates several advantages over conventional testing accuracy based robustness estimation: (1) it provides a uniformed evaluation to models with different structures and parameter scales; (2) it over-performs conventional accuracy based robustness evaluation and provides a more reliable evaluation that is invariant to different test settings; (3) it can be fast generated without considerable testing cost.
Fuxun Yu, Zhuwei Qin, Liang Zhao 0002, Yanzhi Wang 0001, Xiang Chen 0010
IJCAI2
2018 DiReCt: Resource-Aware Dynamic Model Reconfiguration for Convolutional Neural Network in Mobile Systems
abstract
Although Convolutional Neural Networks (CNNs) have been widely applied in various applications, their deployment in resource-constrained mobile systems remains a significant concern. To overcome the computation resource constraints, such as limited memory and energy capacity, many works are proposed for mobile CNN optimization. However, most of them lack a comprehensive modeling analysis of the CNN computation consumption and merely focus on static optimization schemes regardless of different mobile computation scenarios. In this work, we proposed DiReCt -- a resource-aware CNN reconfiguration system. Leveraging accurate CNN computation consumption modeling and mobile resource constraint analysis, DiReCt can reconfigure a CNN with different accuracy and resource consumption levels to adapt to various mobile computation scenarios. The experiment results show that: the proposed computation consumption models in DiReCt can well estimate the CNN computation consumption with 94.1% accuracy, and DiReCt achieves at most 34.9% computation acceleration, 52.7% memory reduction, and 27.1% energy saving. Eventually, DiReCt can effectively adapt CNNs to dynamic mobile usage scenarios for optimal performance.
Zhuwei Qin, Fuxun Yu, Xiang Chen 0010
ISLPED2
2017 AdaLearner: An adaptive distributed mobile learning system for neural networks
abstract
Neural networks hold a critical domain in machine learning algorithms because of their self-adaptiveness and state-of-the-art performance. Before the testing (inference) phases in practical use, sophisticated training (learning) phases are required, calling for efficient training methods with higher accuracy and shorter converging time. Many existing studies focus on the training optimization on high-performance servers or computing clusters, e.g. GPU clusters. However, training neural networks on resource-constrained devices, e.g. mobile platforms, is an important research topic barely touched. In this paper, we implement AdaLearner-an adaptive distributed mobile learning system for neural networks that trains a single network with heterogenous mobile resources under the same local network in parallel. To exploit the potential of our system, we adapt neural networks training phase to mobile device-wise resources and fiercely decrease the transmission overhead for better system scalability. On three representative neural network structures trained from two image classification datasets, AdaLearner boosts the training phase significantly. For example, on LeNet, 1.75-3.37× speedup is achieved when increasing the worker nodes from 2 to 8, thanks to the achieved high execution parallelism and excellent scalability.
Jiachen Mao, Zhuwei Qin, Kent W. Nixon, Xiang Chen 0010, Hai Li 0001, Yiran Chen 0001
ICCAD2
2017 VoCaM: Visualization oriented convolutional neural network acceleration on mobile system: Invited paper
abstract
Convolutional Neural Networks (CNNs) have been widely investigated as some of the most promising solution for various computer vision tasks. However, CNNs introduce massive computing overhead due to their complex network computing flow, resulting in significantly reduced applicability and performance, especially in the mobile devices. Various optimization schemes have been proposed mainly based on both model compression and stacked external computing resources. While these schemes have been proven effective, methods which take into account mobile-specific context-aware optimization approaches have been largely overlooked. One such opportunity is the feasible CNN computing flow simplification to the under-test objects with distinguish features, which can be efficiently pre-analyzed inside the mobile sensor system. Hence, we propose VoCaM, a visualization oriented CNN acceleration framework on mobile devices for image classification tasks. VoCaM takes advantage of the mobile camera system, where the comprehensive pre-analysis can be conducted to reveal the color composition of the under-test images without incurring any additional overhead. Also, the visualization analysis of VoCaM reveals that, certain color-specific filters may have very trivial result impact when the under-test images have mismatching primary color components. Then a set of approximate computing methods is applied to these insignificant filters to replace the intensive convolutional operation, and greatly accelerate the computing process. With ignorable overhead, VoCaM can significantly optimize the computation load of the convolutional layers, with very small impact on the overall classification accuracy.
Zhuwei Qin, Qide Dong, Yiran Chen 0001, Xiang Chen 0010
ICCAD1