Yi Jin 0007

dblp:38/4674-7 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-4335-6691ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Self-aware collaborative edge inference with embedded devices for IIoT
Zhuoquan Yu, Yi Jin 0007, Christine Mwase, Zhuo Zou, Lirong Zheng 0001
Future Gener. Comput. Syst.3
2024 DAI-NET: Toward communication-aware collaborative training for the industrial edge
abstract
The industrial edge generates an abundance of spatially distributed and dynamic data that needs to remain on-site for privacy and security reasons. Collaborative training at the edge can leverage this data to refine pre-trained models locally for specific industrial tasks and environments and have them adapt to local changes for enhanced performance, agility, and resilience. However, communication between the devices during training is a key bottleneck and is not modelled by existing frameworks such as MxNet, PyTorch and TensorFlow. This paper introduces DAI-NET, a co-simulation framework for examining communication and its associated costs, and provides results from an implementation using Python, OMNET++ and INET. To validate it and showcase its utility, the developed platform is applied in the analysis of (i) the performance and cost of collaboratively training a Multilayer Perceptron model, and (ii) the influence of computational heterogeneity. Communication costs generated during the training are captured at the device and system levels. In computationally heterogeneous clusters , the root cause of stragglers is exposed. In addition, the key performance contributors are identified to be a cluster’s computation capability and the variation in the relative computation capabilities of its devices. This study is particularly useful for Artificial Intelligence of Things (AIoT) systems, whose bandwidth and energy resources are limited. It lends the way for more practical research on communication-efficient algorithms, network protocols and architectures for the AIoT edge.
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.2
2023 Self-aware Collaborative Edge Inference with Embedded Devices for Task-oriented IIoT
abstract
The computing and communication resources of embedded devices are constrained and heterogeneous, resulting in a low quality-of-experience for compute-intensive applications in task-oriented industrial Internet of Things (IIoT), such as edge inference. To address these challenges, we first propose a model partitioning-based self-aware collaborative edge inference framework. Furthermore, the throughput-aware collaborative inference algorithm is designed for typical IIoT scenario, stacking tasks. Via jointly optimizing the partition layer and collaborative device selection, the optimal inference efficiency, maximum inference throughput, can be obtained. Finally, the performance of our proposal is demonstrated by extensive simulations and tests based on 10 Raspberry Pi 4Bs and popular models. Specifically, with the proposed algorithm, our platform reaches up to 14.77× throughput speed up for stacking tasks, which indicates that the the proposed design can improve the inference efficiency.
Zhuoquan Yu, Christine Mwase, Yi Jin 0007, Lirong Zheng 0001, Zhuo Zou
VTC Fall4
2022 Communication-efficient distributed AI strategies for the IoT edge
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.2
2022 Edge-Based Collaborative Training System for Artificial Intelligence-of-Things
abstract
The descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension.
Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Ind. Informatics1
2021 Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou
Future Gener. Comput. Syst.1
2018 TMR Group Coding Method for Optimized SEU and MBU Tolerant Memory Design
abstract
This work proposes a fault tolerant memory design using the method of Triple Module Redundancy (TMR) group coding to tolerant the Single-Event Upset (SEU) and Multi-Bit Upset (MBU) influence on memory devices in space environment. The group coding method uses different models to partition and code each word line in memory with Hamming code to achieve best performance. TMR group coding method further increases the capability of self-correction for the errors occurred in parity bits. The evaluation results show that the suggested approach can obtain improved correctness for the memory output with optimized tradeoff between reliability and cost. At 5% error rate, the probability of correct output reaches 70.78% with small cost increment. To achieve 90% reliability, the accuracy improvement is 31.9% compared to TMR with 9% increased area. This solution proposed is evaluated on the memory rich micro-coded processor, but can be further extended to other memory-based processors that need high reliability for the SEU and MBU influence in aerospace applications.
Yi Jin 0007, Yuxiang Huan, Haoming Chu, Zhuo Zou, Lirong Zheng 0001
ISCAS1
2018 A Design of Autonomous Error-Tolerant Architectures for Massively Parallel Computing
Lizheng Liu, Yi Jin 0007, Yi Liu 0027, Yuxiang Huan, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.2