EDBT 2026 Demo / reviewers in the wild / expert
Jun Xia 0003
dblp:22/3650-3
· DBLP profile ↗
21ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0003-0245-8499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective reinforcement learning-based dynamic flexible job shop scheduling using two-stage dispatching
Jiepin Ding, Jun Xia 0003, Yutong Ye 0001, Mingsong Chen 0001 |
J. Syst. Archit. | 2 |
| 2026 | Rethinking fairness in medical imaging: Maximizing group-specific performance with application to skin disease diagnosisabstractRecent efforts in medical image computing have focused on improving fairness by balancing it with accuracy within a single, unified model. However, this often creates a trade-off: gains for underrepresented groups can come at the expense of reduced accuracy for groups that were previously well-served. In high-stakes clinical contexts, even minor drops in accuracy can lead to serious consequences, making such trade-offs highly contentious. Rather than accepting this compromise, we reframe the fairness objective in this paper as maximizing diagnostic accuracy for each patient group by leveraging additional computational resources to train group-specific models. To achieve this goal, we introduce SPARE, a novel data reweighting algorithm designed to optimize performance for a given group. SPARE evaluates the value of each training sample using two key factors: utility, which reflects the sample's contribution to refining the model's decision boundary, and group similarity, which captures its relevance to the target group. By assigning greater weight to samples that score highly on both metrics, SPARE rebalances the training process-particularly leveraging the value of out-of-group data-to improve group-specific accuracy while avoiding the traditional fairness-accuracy trade-off. Experiments on two skin disease datasets demonstrate that SPARE significantly improves group-specific performance while maintaining comparable fairness metrics, highlighting its promise as a more practical fairness paradigm for improving clinical reliability. Gelei Xu, Yuying Duan, Jun Xia 0003, Ching-Hao Chiu, Michael Lemmon 0001, Wei Jin 0009, Yiyu Shi 0001 |
Medical Image Anal. | 3 |
| 2026 | NeFT: Negative Feedback Training to Improve Robustness of Compute-in-Memory DNN AcceleratorsabstractCompute-in-memory accelerators built upon non-volatile memory devices excel in energy efficiency and latency when performing deep neural network (DNN) inference, thanks to their in-situ data processing capability. However, the stochastic nature and intrinsic variations of non-volatile memory devices often result in performance degradation during DNN inference. Introducing these non-ideal device behaviors in DNN training enhances robustness, but drawbacks include limited accuracy improvement, reduced prediction confidence, and convergence issues. This arises from a mismatch between the deterministic training and non-deterministic device variations, as such training, though considering variations, relies solely on the model’s final output. In this work, inspired by control theory, we propose Negative Feedback Training (NeFT)—a novel concept supported by theoretical analysis—to more effectively capture the multi-scale noisy information throughout the network. We instantiate this concept with two specific instances, oriented variational forward (OVF) and intermediate representation snapshot (IRS). Based on device variation models extracted from measured data, extensive experiments show that our NeFT outperforms existing state-of-the-art methods with up to a 45.08% improvement in inference accuracy while reducing epistemic uncertainty, boosting output confidence, and improving convergence probability. These results underline the generality and practicality of our NeFT framework for increasing the robustness of DNNs against device variations. The source code for these two instances is available at https://github.com/YifanQin-ND/NeFT_CIM. Zheyu Yan, Dailin Gan, Jun Xia 0003, Zixuan Pan, Wujie Wen, Xiaobo Sharon Hu, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | Multi-Exit Class Activation Map Guided Feature Masking for Unsupervised Out-of-Distribution Detection in Medical ImagingabstractOut-of-distribution (OOD) detection is crucial for ensuring the safety and reliability of deep learning models in high-stakes domains such as medical imaging. However, existing methods often struggle to detect subtle or localized anomalies, which are common in clinical settings. We hypothesize that such challenges stem in part from a limited understanding of how models focus on different image regions under ID and OOD inputs. To investigate this, we analyze the behavior of deep models under different inputs, and observe that class activation maps (CAMs) for in-distribution (ID) data typically emphasize regions that are highly relevant to the prediction of a model, whereas OOD data often lacks such focused activations. Building on this, we find that masking input images with inverted CAMs induces larger shifts in feature representations for ID than OOD data, a signal that can be leveraged for robust detection. Based on this insight, we propose Multi-Exit Class Activation Map (MECAM), a novel unsupervised OOD detection framework that integrates aggregated multi-exit CAMs and CAM-guided feature masking. By combining CAMs from multiple network depths, our method captures both global and local feature representations, thereby enhancing the robustness of OOD detection. We evaluate MECAM on two ID datasets, including ISIC19 and PathMNIST, and test its performance against three medical OOD datasets, RSNA Pneumonia, COVID-19, and HeadCT, and one natural image OOD dataset, iSUN. Comprehensive experiments demonstrate that MECAM consistently outperforms state-of-theart OOD detection methods, validating its effectiveness. These findings highlight the potential of multi-exit architectures and CAM-guided feature masking in advancing unsupervised OOD detection for medical imaging, paving the way for more reliable and interpretable models in clinical practice. The source code is available at https://github.com/zx-pan/MECAM-OOD. Zixuan Pan, Jun Xia 0003, Max Ficco, Jianxu Chen 0001, Tsung-Yi Ho, Yiyu Shi 0001 |
BIBM | 3 |
| 2025 | Rethinking Medical Anomaly Detection in Brain MRI: An Image Quality Assessment PerspectiveabstractReconstruction-based methods, particularly those leveraging autoencoders, have been widely adopted for anomaly detection task in brain MRI. Unlike most existing works try to improve the task accuracy through architectural or algorithmic innovations, we tackle this task from image quality assessment (IQA) perspective, an under-explored direction in the field. Due to the limitations of conventional metrics such as £1 in capturing the nuanced differences in reconstructed images for medical anomaly detection, we propose fusion quality, a novel metric that wisely integrates the structure-level sensitivity of Structural Similarity Index Measure (SSIM) with the pixel-level precision of £1. The metric offers a more comprehensive assessment of reconstruction quality, considering intensity (subtractive property of l1and divisive property of SSIM), contrast, and structural similarity. Furthermore, the proposed metric makes subtle regional variations more impactful in the final assessment. Thus, considering the inherent divisive properties of SSIM, we design an average intensity ratio (AIR)-based data transformation that amplifies the divisive discrepancies between normal and abnormal regions, thereby enhancing anomaly detection. By fusing the aforementioned two components, we devise the IQA approach. Experimental results on two distinct brain MRI datasets show that our IQA approach significantly enhances medical anomaly detection performance when integrated with state-of-the-art baselines. Code is provided here. Zixuan Pan, Jun Xia 0003, Zheyu Yan, Guoyue Xu, Yawen Wu, Zhenge Jia, Jianxu Chen 0001, Yiyu Shi 0001 |
BIBM | 2 |
| 2025 | Enabling Memory-Efficient On-Device Learning via Dataset CondensationabstractUpon deployment to edge devices, it is often desirable for a model to further learn from streaming data to improve accuracy. However, learning from such data is challenging because it is typically unlabeled, non-independent and identically distributed (non-i.i.d), and only seen once, which can lead to potential catastrophic forgetting. A common strategy to mitigate this issue is to maintain a small data buffer on the edge device to select and retain the most representative data for rehearsal. However, the selection process leads to significant information loss since most data is either never stored or quickly discarded. This paper proposes a framework that addresses this issue by condensing incoming data into informative synthetic samples. Specifically, to effectively handle unlabeled incoming data, we propose a pseudo-labeling technique designed for on-device learning environments. We also develop a dataset condensation technique tailored for on-device learning scenarios, which is significantly faster compared to previous methods. To counteract the effects of noisy labels during the condensation process, we further utilize a feature discrimination objective to improve the purity of class data. Experimental results indicate substantial improvements over existing methods, especially under strict buffer limitations. For instance, with a buffer capacity of just one sample per class, our method achieves a 56.7% relative increase in accuracy compared to the best existing baseline on the CORe50 dataset. Gelei Xu, Ningzhi Tang, Jun Xia 0003, Ruiyang Qin, Wei Jin 0009, Yiyu Shi 0001 |
DATE | 3 |
| 2025 | FedGraft: Memory-Aware Heterogeneous Federated Learning via Model GraftingabstractAlthough Federated Learning (FL) is good at collaborative learning among devices without compromising their data privacy, it suffers from the problem of large-scale deployment in Mobile Edge Computing (MEC) applications. This is mainly because the varying memory sizes of edge devices inevitably result in limited sizes of their hosting models. According to the Cannikin Law, when dealing with heterogeneous devices with different memory sizes, the learning capability of existing homogeneous FL schemes is greatly restricted by the weakest device. Worse still, although existing heterogeneous FL methods enable a MEC application to involve numerous devices equipped with heterogeneous models, their knowledge aggregation processes require either extra training data or architecture similarity of models. To address the above issues, this paper presents a novel FL method named FedGraft that enables effective knowledge sharing among heterogeneous device models of different sizes without imposing unrealistic assumptions. In FedGraft, all the device models are grafted to a common rootstock based on our proposed model partitioning and grafting mechanism, facilitating knowledge sharing among heterogeneous models on top of a tree-like global model. Meanwhile, using our proposed device selection strategy, the reassembled submodels extracted from the global model can be reasonably dispatched to corresponding devices with sufficient memory, thus enhancing the overall FL performance. Comprehensive experimental results show that, compared with state-of-the-art heterogeneous FL methods, FedGraft can improve inference accuracy by up to 17% in various memory-constrained scenarios. Ruixuan Liu, Ming Hu 0003, Zeke Xia, Xiaofei Xie, Jun Xia 0003, Yihao Huang 0001, Mingsong Chen 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | Energy-Efficient Shop Scheduling Using Space-Cooperation Multi-Objective OptimizationabstractSince Industry 5.0 emphasizes that manufacturing enterprises should raise awareness of social contribution to achieve sustainable development, more and more meta-heuristic algorithms are investigated to save energy in manufacturing systems. Although non-dominated sorting-based meta-heuristics have been recognized as promising multi-objective optimization methods for solving the energy-efficient flexible job shop scheduling problem (EFJSP), it is hard to guarantee the quality of the Pareto front (e.g., total energy consumption, makespan) due to the lack of population diversity. This is mainly because an improper individual comparison inevitably reduces population diversity, thus limiting exploration and exploitation abilities during population updates. To achieve efficient population evolution, this paper introduces a novel space-cooperation multi-objective optimization (SCMO) method that can effectively solve EFJSP to obtain scheduling schemes with better trade-offs. By cooperatively evaluating the similarity among individuals in both the decision space and objective space, we propose a space-cooperation population update method based on a three-vector representation that can accurately eliminate repetitive individuals to derive higher-quality Pareto solutions. To further improve search efficiency, we propose a difference-driven local search, which selectively changes the positions of operations with higher differences to search for neighbors effectively. Based on the Taguchi method, we conduct experiments to obtain a suitable parameter combination of SCMO. Comprehensive experimental results show that, compared to state-of-the-art methods, our SCMO method achieves the highest HV and NR and the lowest IGD, with an average of 0.990, 0.952, and 0.001, respectively. Meanwhile, compared to traditional local search approaches, our difference-driven local search obtains twice the HV on instance Mk12 and reduces the solving time from 1521 s to 475 s. Jiepin Ding, Jun Xia 0003, Yaning Yang, Junlong Zhou, Mingsong Chen 0001, Keqin Li 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2024 | Enabling On-Device Large Language Model Personalization with Self-Supervised Data Selection and SynthesisabstractAfter a large language model (LLM) is deployed on edge devices, it is desirable for these devices to learn from user-generated conversation data to generate user-specific and personalized responses in real-time. However, user-generated data usually contains sensitive and private information, and uploading such data to the cloud for annotation is not preferred if not prohibited. While it is possible to obtain annotation locally by directly asking users to provide preferred responses, such annotations have to be sparse to not affect user experience. In addition, the storage of edge devices is usually too limited to enable large-scale fine-tuning with full user-generated data. It remains an open question how to enable on-device LLM personalization, considering sparse annotation and limited on-device storage. In this paper, we propose a novel framework to select and store the most representative data online in a self-supervised way. Such data has a small memory footprint and allows infrequent requests of user annotations for further fine-tuning. To enhance fine-tuning quality, multiple semantically similar pairs of question texts and expected responses are generated using the LLM. Our experiments show that the proposed framework achieves the best user-specific content-generating capability (accuracy) and fine-tuning speed (performance) compared with vanilla baselines. To the best of our knowledge, this is the very first on-device LLM personalization framework. Ruiyang Qin, Jun Xia 0003, Zhenge Jia, Meng Jiang 0001, Ahmed Abbasi, Peipei Zhou 0001, Jingtong Hu, Yiyu Shi 0001 |
DAC | 2 |
| 2024 | Towards Energy-Aware Federated Learning via MARL: A Dual-Selection Approach for Model and ClientabstractAlthough Federated Learning (FL) is promising in knowledge sharing for heterogeneous Artificial Intelligence of Thing (AIoT) devices, their training performance and energy efficacy are severely restricted in practical battery-driven scenarios due to the "wooden barrel effect" caused by the mismatch between homogeneous model paradigms and heterogeneous device capability. As a result, due to various kinds of differences among devices, it is hard for existing FL methods to conduct training effectively in energy-constrained scenarios, such as the battery constraints of devices. To tackle the above issues, we propose an energy-aware FL framework named DR-FL, which considers the energy constraints in both clients and heterogeneous deep learning models to enable energy-efficient FL. Unlike Vanilla FL, DR-FL adopts our proposed Muti-Agents Reinforcement Learning (MARL)-based dual-selection method, which allows participated devices to make contributions to the global model effectively and adaptively based on their computing capabilities and energy capacities in a MARL-based manner. Experiments conducted with various widely recognized datasets demonstrate that DR-FL has the capability to optimize the exchange of knowledge among diverse models in large-scale AIoT systems while adhering to energy limitations. Additionally, it improves the performance of each individual heterogeneous device's model. Jun Xia 0003, Yi Zhang 0165, Yiyu Shi 0001 |
ICCAD | 1 |
| 2024 | WaveAttack: Asymmetric Frequency Obfuscation-based Backdoor Attacks Against Deep Neural NetworksabstractDue to the increasing popularity of Artificial Intelligence (AI), more and more backdoor attacks are designed to mislead Deep Neural Network (DNN) predictions by manipulating training samples or processes. Although backdoor attacks have been investigated in various scenarios, they still suffer from the problems of both low fidelity of poisoned samples and non-negligible transfer in latent space, which make them easily identified by existing backdoor detection algorithms. To overcome this weakness, this paper proposes a novel frequency-based backdoor attack method named WaveAttack, which obtains high-frequency image features through Discrete Wavelet Transform (DWT) to generate highly stealthy backdoor triggers. By introducing an asymmetric frequency obfuscation method, our approach adds an adaptive residual to the training and inference stages to improve the impact of triggers, thus further enhancing the effectiveness of WaveAttack. Comprehensive experimental results show that, WaveAttack can not only achieve higher effectiveness than state-of-the-art backdoor attack methods, but also outperform them in the fidelity of images (i.e., by up to 28.27\% improvement in PSNR, 1.61\% improvement in SSIM, and 70.59\% reduction in IS). Our code is available at https://github.com/BililiCode/WaveAttack. Jun Xia 0003, Zhihao Yue, Yingbo Zhou 0001, Zhiwei Ling, Yiyu Shi 0001, Xian Wei, Mingsong Chen 0001 |
NeurIPS | 1 |
| 2023 | Model-Contrastive Learning for Backdoor EliminationabstractDue to the popularity of Artificial Intelligence (AI) techniques, we are witnessing an increasing number of backdoor injection attacks that are designed to maliciously threaten Deep Neural Networks (DNNs) causing misclassification. Although there exist various defense methods that can effectively erase backdoors from DNNs, they greatly suffer from both high Attack Success Rate (ASR) and a non-negligible loss in Benign Accuracy (BA). Inspired by the observation that a backdoored DNN tends to form a new cluster in its feature spaces for poisoned data, in this paper, we propose a novel two-stage backdoor defense method, named MCLDef, based on Model-Contrastive Learning (MCL). MCLDef can purify the backdoored model by pulling the feature representations of poisoned data towards those of their clean data counterparts. Due to the shrunken cluster of poisoned data, the backdoor formed by end-to-end supervised learning can be effectively eliminated. Comprehensive experimental results show that, with only 5% of clean data, MCLDef significantly outperforms state-of-the-art defense methods by up to 95.79% reduction in ASR, while in most cases, the BA degradation can be controlled within less than 2%. Our code is available at https://github.com/Zhihao151/MCL. Zhihao Yue, Jun Xia 0003, Zhiwei Ling, Ming Hu 0003, Ting Wang 0001, Xian Wei, Mingsong Chen 0001 |
ACM Multimedia | 2 |
| 2023 | GitFL: Uncertainty-Aware Real-Time Asynchronous Federated Learning Using Version ControlabstractAs a promising distributed machine learning paradigm that enables collaborative training without compromising data privacy, Federated Learning (FL) has been increasingly used in large-scale A IoT (Artificial Intelligence of Things) system design. However, due to the lack of efficient management of straggling devices, existing FL methods greatly suffer from the problems of long response time (e.g., training and communication latency) and low inference accuracy. Things become even worse when taking various uncertain factors (e.g., network delays, performance variances caused by process variation) existing in AIoT scenarios into account. To address this issue, this paper proposes a novel asynchronous FL framework named GitFL, whose implementation is inspired by the famous version control system Git. Unlike traditional FL, the cloud server of GitFL maintains a master model (i.e., the global model) together with a set of branch models indicating the trained local models committed by selected devices, where the master model is updated based on both all the pushed branch models and their version information, and only the branch models after the pull operation are dispatched to devices. By using our proposed Reinforcement Learning (RL)-based device selection mechanism, a pulled branch model with an older version will be more likely to be dispatched to a faster and less frequently selected device for the next round of local training. In this way, GitFL enables both effective controls of model staleness and adaptive load balance of versioned models among straggling devices, thus avoiding performance deterioration while ensuring real-time performance. Comprehensive experimental results on well-known models and datasets show that, compared with state-of-the-art asynchronous and synchronous FL methods, GitFL can achieve up to 2.64X training acceleration and 7.88 % inference accuracy improvements in various uncertain scenarios. Ming Hu 0003, Zeke Xia, Dengke Yan, Zhihao Yue, Jun Xia 0003, Yihao Huang 0001, Yang Liu 0003, Mingsong Chen 0001 |
RTSS | 5 |
| 2023 | Efficient Federated Learning for AIoT Applications Using Knowledge DistillationabstractAs a promising distributed machine learning paradigm, federated learning (FL) trains a central model with decentralized data without compromising user privacy, which makes it widely used by Artificial Intelligence Internet of Things (AIoT) applications. However, the traditional FL suffers from model inaccuracy, since it trains local models only using hard labels of data while useful information of incorrect predictions with small probabilities is ignored. Although various solutions try to tackle the bottleneck of the traditional FL, most of them introduce significant communication overhead, making the deployment of large-scale AIoT devices a great challenge. To address the above problem, this article presents a novel distillation-based FL (DFL) method that enables efficient and accurate FL for AIoT applications. By using knowledge distillation (KD), in each round of FL training, our approach uploads both the soft targets and local model gradients to the cloud server for aggregation, where the aggregation results are then dispatched to AIoT devices for the next round of local training. During the DFL local training, in addition to hard labels, the model predictions approximate soft targets, which can improve model accuracy by leveraging the knowledge of soft targets. To further improve our DFL model performance, we design a dynamic adjustment strategy of loss function weights for tuning the ratio of KD and FL, which can maximize the synergy between soft targets and hard labels. Comprehensive experimental results on well-known benchmarks show that our approach can significantly improve the model accuracy of FL without introducing significant communication overhead. Tian Liu 0005, Jun Xia 0003, Zhiwei Ling, Xin Fu 0001, Shui Yu 0001, Mingsong Chen 0001 |
IEEE Internet Things J. | 2 |
| 2023 | Automated Synthesis of Safe Timing Behaviors for Requirements Models Using CCSLabstractAs a promising requirement-level specification language for timing behavior modeling, the clock constraint specification language (CCSL) has become popular in the model-driven design community for safety-critical embedded systems. However, due to the skyrocketing design complexity, in practice, it is hard for requirement engineers to accurately construct requirement models with expected timing behaviors using CCSL, especially, for safe timing behaviors. Although more and more CCSL synthesis approaches are designed to facilitate the generation of CCSL specifications, most of them cannot be used directly for the synthesis of requirements models. This is because existing CCSL synthesis methods: 1) focus on filling the holes of CCSL constraints rather than completing requirements models and 2) rely heavily on limited observations of system behaviors, while the (temporal) safety properties of target systems are neglected. To address these issues, this article proposes a novel method that enables the automated synthesis of safe timing behaviors for requirements models. By specifying the safety timing properties of target systems using safely-LTL, our approach adopts CCSL as an intermediate representation of requirement synthesis, where incomplete requirements models coupled with safely-LTL-based properties are encoded into CCSL constraints with holes. Guided by the samples (expected behaviors) provided by requirement engineers, our approach can automatically figure out the complete version of incomplete requirements models. Comprehensive experimental results on two complex case studies demonstrate that our approach can not only quickly and efficiently synthesize requirement models but also guarantee that the synthesized models satisfy specified safety properties in Safely-LTL form. Ming Hu 0003, Jun Xia 0003, Min Zhang 0002, Xiaohong Chen 0007, Frédéric Mallet, Mingsong Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Accelerated synthesis of neural network-based barrier certificates using collaborative learningabstractMost of existing Neural Network (NN)-based barrier certificate synthesis methods cannot deal with high-dimensional continuous systems, since a large quantity of sampled data may easily result in inaccurate initial models coupled with slow convergence rate. To accelerate the synthesis of NN-based barrier certificates, this paper presents an effective two-stage approach named CL-BC, which fully exploits the parallel processing capability of underlying hardware to enable quick search for a barrier certificate. Unlike existing NN-based methods that adopt a random initial model for barrier certificate synthesis, in the first stage CL-BC pre-trains an initial model based on a small subset of sampling data. In this way, an approximate barrier certificate in an NN form can be quickly achieved with little overhead. Based on our proposed collaborative learning scheme, in the second stage CL-BC conducts the parallel learning on partitioned domains, where the learned knowledge from different partitions can be aggregated to accelerate the convergence of a global NN model for barrier certificate synthesis. In this way, the overall synthesis time of an NN-based barrier certificate can be drastically reduced. Experimental results show that our approach can not only drastically reduce barrier synthesis time, but also can synthesize barrier certificates for complex systems that cannot be handled by state-of-the-art. Jun Xia 0003, Ming Hu 0003, Mingsong Chen 0001 |
DAC | 1 |
| 2022 | Eliminating Backdoor Triggers for Deep Neural Networks Using Attention Relation Graph DistillationabstractDue to the prosperity of Artificial Intelligence (AI) techniques, more and more backdoors are designed by adversaries to attack Deep Neural Networks (DNNs). Although the state-of-the-art method Neural Attention Distillation (NAD) can effectively erase backdoor triggers from DNNs, it still suffers from non-negligible Attack Success Rate (ASR) together with lowered classification ACCuracy (ACC), since NAD focuses on backdoor defense using attention features (i.e., attention maps) of the same order. In this paper, we introduce a novel backdoor defense framework named Attention Relation Graph Distillation (ARGD), which fully explores the correlation among attention features with different orders using our proposed Attention Relation Graphs (ARGs). Based on the alignment of ARGs between teacher and student models during knowledge distillation, ARGD can more effectively eradicate backdoors than NAD. Comprehensive experimental results show that, against six latest backdoor attacks, ARGD outperforms NAD by up to 94.85% reduction in ASR, while ACC can be improved by up to 3.23%. Jun Xia 0003, Ting Wang 0001, Jiepin Ding, Xian Wei, Mingsong Chen 0001 |
IJCAI | 1 |
| 2022 | PervasiveFL: Pervasive Federated Learning for Heterogeneous IoT SystemsabstractFederated learning (FL) has been recognized as a promising collaborative on-device machine learning method in the design of Internet of Things (IoT) systems. However, most existing FL methods fail to deal with IoT applications that contain a variety of IoT devices equipped with different types of neural network (NN) models. This is because traditional FL methods assume that local models on devices should have the same architecture as the global model on cloud. To address this problem, we propose a novel framework named PervasiveFL that enables efficient and effective FL among heterogeneous IoT devices. Without modifying original local models, PervasiveFL installs one lightweight NN model named modellet on each device. By using the deep mutual learning (DML) and our entropy-based decision gating (EDG) method, modellets and local models can selectively learn from each other through soft labels using locally captured data. Meanwhile, since modellets are of the same architecture, the learned knowledge by modellets can be shared among devices in a traditional FL manner. In this way, PervasiveFL can be pervasively applied to any heterogeneous IoT system. Comprehensive experimental results on four well-known datasets show that PervasiveFL can not only pervasively enable FL among heterogeneous devices within a large-scale IoT system, but also significantly enhance the inference accuracy of heterogeneous IoT devices with low communication overhead. Jun Xia 0003, Tian Liu 0005, Zhiwei Ling, Ting Wang 0001, Xin Fu 0001, Mingsong Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Efficient Federated Learning for Cloud-Based AIoT ApplicationsabstractAs a promising method for central model training on decentralized device data without compromising user privacy, federated learning (FL) is becoming more and more popular in Internet-of-Things (IoT) design. However, due to limited computing and memory resources of devices that restrict the capabilities of hosted deep learning models, existing FL approaches for artificial intelligence IoT (AIoT) applications suffer from inaccurate prediction results. To address this problem, this article presents a collaborativeBig.Littlebranch architecture to enable efficient FL for AIoT applications. Inspired by the architecture of BranchyNet which has multiple prediction branches, our approach deploys deep neural network (DNN) models across both cloud and AIoT devices. OurBig.Littlebranch model has two branches, where the big branch is deployed on cloud for strengthened prediction accuracy, and the little branches are used to fit for AIoT devices. When AIoT devices cannot make the prediction with high confidence using local little branches, they will resort to the big branch for further inference. To increase both prediction accuracy and early exit rate ofBig.Littlebranch model, we propose a two-stage training and coinference scheme, which considers the local characteristics of AIoT scenarios. Comprehensive experiment results obtained from a real AIoT environment demonstrate the efficiency and effectiveness of our approach in terms of prediction accuracy and average inference time. Xinqian Zhang, Ming Hu 0003, Jun Xia 0003, Tongquan Wei, Mingsong Chen 0001, Shiyan Hu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | A PYNQ-compliant Online Platform for Zynq-based DNN DevelopersabstractThe Zynq heterogeneous SoC from Xilinx is able to supporting software/hardware co-designing in one single chip, making it possible to take advantage of software flexibility and hardware acceleration at the same time. PYNQ project from Xilinx is trying to take advantage of high performance and low power consumption of Zynq while improve its programmability. In order to improve the ecosystem of PYNQ and help more embedded AI applications use the Zynq-based high-efficiency computational engine, this paper proposes a PYNQ-compliant online platform (OpenHEC-PYNQ) that integrates all necessary factors for the Zynq-based DNN developer. This platform makes HDL/HLS designers able to access all resources they needed via the Internet and finish all jobs one-stop. To show effectiveness of this platform, a YOLOv2 FPGA acceleration library is implemented based on OpenHEC-PYNQ. Jun Xia 0003, Wenmin Yang, ZhiLei Chai |
FPGA | 2 |
| 2018 | Decouple and Stretch: A Boost to Channel PruningabstractDeep Neural Networks (DNNs) have shown superior performance on a variety of artificial intelligence problems. Reducing the resource usage of DNN is critical to adding intelligence on Internet of Things (IoT) devices. Channel pruning based network compression shows effective reduction simultaneously on storage, memory and computation without specialized software on general platforms. But limited by pruning flexibility, channel pruning methods have relatively low compression rate for a given target performance. In this paper, we demonstrate that channel pruning becomes more robust to decision errors by reducing the granularity of filters. Then we propose a Decouple and Stretch (DS) scheme to enhance channel pruning. Under this scheme, each filter in a specific layer is decoupled into two small spatial-wise filters, and the spatial-wise filters are stretched into two successive convolutional layers. Our scheme obtains up to 49% improvement on compression and 35% improvement on acceleration. To further demonstrate hardware compatibility, we deploy pruned networks on the FPGA, and the network produced by Decouple and Stretch scheme is more hardware-friendly with latency reduced by 42%. Zhen Chen 0013, Sen Liu 0001, Jun Xia 0003, Weiping Li 0003 |
IPCCC | 4 |