Xin Niu 0001

dblp:121/7606-1 · DBLP profile ↗
← Back
7ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-4173-0822ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Optimized Scheduling of Dependent Tasks and Idle Computational Resources for Edge Intelligence
abstract
Due to the dependencies among different computing tasks, edge devices must await the completion of preceding computing tasks in order to continue with the current computing task. As a result, there is a long waiting time. In the past, reducing waiting time was often achieved by optimizing the scheduling of edge devices to execute computing tasks more efficiently. However, this approach incurs a certain level of communication overhead. Additionally, the waiting process for edge devices results in a waste of their computational resources. In order to tackle the challenge of excessive waiting time, we propose a wait-time compression scheme (WCS) based on model early-exit. The WCS selects early exit points for computing tasks based on the user-tolerated latency and reduces the waiting time without scheduling computing tasks or edge devices. Furthermore, to optimize the use of idle computational resources of edge devices during the waiting process, we introduce an idle-resource-based computing task scheduling scheme (ICTS). A significant number of experiments demonstrate that compared to other computing tasks scheduling schemes, WCS achieves a latency reduction of up to 10.7%, while ICTS enhances computational resource utilization by as much as 28.8%.
Xin Niu 0001, Wang Chen 0004, Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers1
2026 Efficient Cluster-Based Knowledge Distillation for Deep Face Recognition
abstract
Knowledge distillation has been widely used to improve the performance of small compact models for face recognition. However, selecting key knowledge and effectively transferring it from teacher to student remains a challenging problem. In this work, we propose an efficient Cluster-based Knowledge Distillation (CKD) dedicated to aligning the student model with the teacher model in terms of both sample relations and class centers. Specifically, CKD first determines the key sample relations based on the similarities between the sample features extracted by the teacher and their cluster centers generated by existing clustering algorithms. Then, CKD effectively transfers the knowledge of the above relations from the teacher to the student by designing a cluster-based relation distillation loss. Finally, CKD further improves the quality of the student's class centers by constructing a center loss between the above representative cluster centers and the student's class centers. We validate the proposed CKD on multiple face benchmarks. For example, CKD improves the baseline student performance from 91.95% to 94.20% on MegaFace and consistently outperforms recent competitive distillation methods on multiple benchmarks. These results demonstrate the effectiveness and superiority of CKD.
Xianwei Lv 0001, Haibo Mi, Xin Niu 0001, Wang Chen 0004, Kun Wang 0059, Chen Yu 0003
IEEE Trans. Sustain. Comput.3
2025 Computing Tasks Saving Schemes Through Early Exit in Edge Intelligence-Assisted Systems
abstract
Edge intelligence (EI) is a promising paradigm where end devices collaborate with edge servers to provide artificial intelligence services to users. In most realistic scenarios, end devices often move unconsciously, resulting in frequent computing migrations. Moreover, a surge in computing tasks offloaded to edge servers significantly prolongs queuing latency. These two issues obstruct the timely completion of computing tasks in EI-assisted systems. In this paper, we formulate an optimization problem aiming to maximize computing task completion under latency constraints. To address this issue, we first categorize computing tasks into new computing tasks (NCTs) and partially completed computing tasks (PCTs). Subsequently, based on model partitioning, we design a new computing task saving scheme (NSS) to optimize early exit points for NCTs and computing tasks in the queuing queue. Furthermore, we propose a partially completed computing task saving scheme (PSS) to set early exit points for PCTs during computing migrations. Numerous experiments show that computing saving schemes can achieve at least 90% computing task completion rate and up to 61.81% latency reduction compared to other methods.
Xin Niu 0001, Xianwei Lv 0001, Wang Chen 0004, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers1
2025 RLA: A Low-Latency and High-Smoothness Path Planning System Based on Interpolation and Velocity Control
abstract
Local planning is a key issue in the field of unmanned delivery. Unmanned delivery requires high delivery efficiency and lower equipment maintenance costs, which pose challenges to the latency and smoothness of path planning algorithms. After investigation of the work on optimizing the latency and smoothness of local planning, we proposed a Robotic-Look-Ahead approach based on Look Ahead approach. It consists of four parts: calculating the conjunction speed, circular arc interpolation, the Look-Ahead method, and path modification. The experiment showed that with different paths, different running memory, and different maximum running speeds, latency decreased by an average of 90% compared to the benchmark, and smoothness improved by an average of 40%. Under different loads, the average energy consumption decreases by 4%.
Xupeng Zhu, Xin Niu 0001, Wang Chen 0004, Chen Yu 0003
IEEE Trans. Sustain. Comput.2
2024 Game-Based Adaptive FLOPs and Partition Point Decision Mechanism With Latency and Energy-Efficient Tradeoff for Edge Intelligence
abstract
As the product of the combination of edge computing and artificial intelligence, edge intelligence (EI) not only solves the problem of insufficient computing capacity of the end device, but also can provide users with various types of intelligent services. However, offline and online model partitioning methods respectively have problems of poor adaptability to the real computing environment and delayed feedback. In addition, previous work on optimizing energy consumption through model partitioning often ignores the latency of intelligent services. Similarly, the energy consumption of end devices and edge servers is usually not considered when optimizing latency. Therefore, we propose game-based adaptive floating-point operations and partition point decision mechanism (GAFPD) to efficiently find the optimal partition point that reduces latency and improves energy efficiency simultaneously in a dynamically changing computing environment. Numerous simulation experiments and robot-based EI system experiments show that GAFPD can simultaneously reduce the latency of intelligent services and improve the energy efficiency of edge devices, while exhibiting strong adaptability to bandwidth changes.
Xin Niu 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers1
2022 CRSM: Computation Reloading Driven by Spatial-Temporal Mobility in Edge-Assisted Automated Industrial Cyber-Physical Systems
abstract
Edge Computing, as an emerging computing mode, transfers computing capacity from the cloud center to the edge of new generation automation sensor networks. However, the surge in the number of sensor devices has resulted in the edge server overload. Fortunately, with the development of hardware technology, the computing capacity of sensor devices has improved. Therefore, we introduce a new concept of computation reloading, that is, resource-rich servers allocate tasks to sensor devices with stronger computing capacity. Most previous real-time EC researches have not considered sensor devices’ higher spatial–temporal mobility, which results in continuous computation migration and high delay. In this article, we expose a computation reloading scheme driven by spatial–temporal mobility (CRSM), for single time slice and multiple time slices cases, computation tasks are allocated in advance by predicting sensor devices locations. The experiments demonstrate that our scheme outperforms other methods in terms of shorting industrial cyber-physical systems delays.
Xin Niu 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Ind. Informatics1
2019 A Differential Private Mechanism to Protect Trajectory Privacy in Mobile Crowd-Sensing
abstract
With the fast development of smart mobile devices, the mobile crowd-sensing (MCS) has been witnessed as a new data collection paradigm. In this paper, we consider a scenario that an MCS server tries to collect trajectories from participants. In order to protect the participants' location privacy from their own side, we let participants submit noisy data to the server. In addition, we assume that the data collection is delay tolerant which means each participant is allowed to submit his trajectory in a bundle instead of submitting locations one by one. Based on this assumption, we regard each trajectory as a vector in the high dimension space and design a trajectory protection algorithm to perturb the true trajectory before submission. We use the differential privacy (DP) as the privacy model so we can estimate the amount of noise given a privacy level. To evaluate our mechanism, we use real world traffic data collected from Shanghai taxis and compare it with existing work. The results show that our mechanism not only guarantees privacy protection, but also preserves trajectories' utility.
Hongyu Huang 0001, Xin Niu 0001, Chao Chen 0004, Chunqiang Hu
WCNC2