Xianwei Lv 0001

dblp:178/6902-1 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-4864-314XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Optimized Scheduling of Dependent Tasks and Idle Computational Resources for Edge Intelligence
abstract
Due to the dependencies among different computing tasks, edge devices must await the completion of preceding computing tasks in order to continue with the current computing task. As a result, there is a long waiting time. In the past, reducing waiting time was often achieved by optimizing the scheduling of edge devices to execute computing tasks more efficiently. However, this approach incurs a certain level of communication overhead. Additionally, the waiting process for edge devices results in a waste of their computational resources. In order to tackle the challenge of excessive waiting time, we propose a wait-time compression scheme (WCS) based on model early-exit. The WCS selects early exit points for computing tasks based on the user-tolerated latency and reduces the waiting time without scheduling computing tasks or edge devices. Furthermore, to optimize the use of idle computational resources of edge devices during the waiting process, we introduce an idle-resource-based computing task scheduling scheme (ICTS). A significant number of experiments demonstrate that compared to other computing tasks scheduling schemes, WCS achieves a latency reduction of up to 10.7%, while ICTS enhances computational resource utilization by as much as 28.8%.
Xin Niu 0001, Wang Chen 0004, Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers3
2026 Efficient Cluster-Based Knowledge Distillation for Deep Face Recognition
abstract
Knowledge distillation has been widely used to improve the performance of small compact models for face recognition. However, selecting key knowledge and effectively transferring it from teacher to student remains a challenging problem. In this work, we propose an efficient Cluster-based Knowledge Distillation (CKD) dedicated to aligning the student model with the teacher model in terms of both sample relations and class centers. Specifically, CKD first determines the key sample relations based on the similarities between the sample features extracted by the teacher and their cluster centers generated by existing clustering algorithms. Then, CKD effectively transfers the knowledge of the above relations from the teacher to the student by designing a cluster-based relation distillation loss. Finally, CKD further improves the quality of the student's class centers by constructing a center loss between the above representative cluster centers and the student's class centers. We validate the proposed CKD on multiple face benchmarks. For example, CKD improves the baseline student performance from 91.95% to 94.20% on MegaFace and consistently outperforms recent competitive distillation methods on multiple benchmarks. These results demonstrate the effectiveness and superiority of CKD.
Xianwei Lv 0001, Haibo Mi, Xin Niu 0001, Wang Chen 0004, Kun Wang 0059, Chen Yu 0003
IEEE Trans. Sustain. Comput.1
2025 Computing Tasks Saving Schemes Through Early Exit in Edge Intelligence-Assisted Systems
abstract
Edge intelligence (EI) is a promising paradigm where end devices collaborate with edge servers to provide artificial intelligence services to users. In most realistic scenarios, end devices often move unconsciously, resulting in frequent computing migrations. Moreover, a surge in computing tasks offloaded to edge servers significantly prolongs queuing latency. These two issues obstruct the timely completion of computing tasks in EI-assisted systems. In this paper, we formulate an optimization problem aiming to maximize computing task completion under latency constraints. To address this issue, we first categorize computing tasks into new computing tasks (NCTs) and partially completed computing tasks (PCTs). Subsequently, based on model partitioning, we design a new computing task saving scheme (NSS) to optimize early exit points for NCTs and computing tasks in the queuing queue. Furthermore, we propose a partially completed computing task saving scheme (PSS) to set early exit points for PCTs during computing migrations. Numerous experiments show that computing saving schemes can achieve at least 90% computing task completion rate and up to 61.81% latency reduction compared to other methods.
Xin Niu 0001, Xianwei Lv 0001, Wang Chen 0004, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers2
2023 A Feedback-Driven DNN Inference Acceleration System for Edge-Assisted Video Analytics
abstract
With the proposal of edge computing, lots of intelligence applications have made significant progress. For enormous video analysis, how to further accelerate the process is still a major challenge. To overcome the challenge, researchers propose various video frame filtering systems to reduce the data transmission. In this work, we propose a Feedback-Driven DNN Inference Acceleration system (FDDIA). FDDIA is committed to further reducing the latency of DNN inference and the transmission according to the feedback information. Specifically, on the device side, FDDIA first uses the detection results of prior frames as feedback to determine the key candidate regions and uses the inter-frame difference to determine the new object regions. Those regions containing large objects are down-sampled to further reduce the frame. On the edge side, FDDIA first integrates all candidate regions into a smaller new image. Then it remaps the detection result for the new image back to the original frame and returns the result to the device. We evaluate FDDIA on different video benchmarks for three object detection tasks. The results show FDDIA improves the average end-to-end latency by 44% and average bandwidth usage by 41% than the existing advanced method while maintaining a high accuracy.
Xianwei Lv 0001, Qianqian Wang 0017, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Computers1
2022 HQ2CL: A High-Quality Class Center Learning System for Deep Face Recognition
abstract
Benefited from the proposals of function losses margin-based, face recognition has achieved significant improvements in recent years. Those losses aim to increase the margin between the different identities to enhance the discriminability. Ideally, the class center of different identities is far from each other, and face samples are compact around the corresponding class center. Hence, it's very vital to produce a high-quality class center. However, the distribution of training sets determines the class center. With low-quality samples being in the majority, the class center would be close to the samples with little identity information. As a result, it would impair the discriminability of the learned model for those unseen samples. In this work, we propose a High-Quality Class Center Learning system (HQ2CL). This is an effective system and guides the class center to approach the high-quality samples to keep the discriminability. Specifically, HQ2CL introduces a quality-aware scale and margin layer for the identification loss and constructs a new high-quality center loss. We implement the proposed system without additional burden. And we present the experimental evaluation over different face benchmarks. The experimental results show the superiority of our proposed HQ2CL over the state-of-the-arts.
Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001
IEEE Trans. Image Process.1
2022 Cost Efficient Sensor Positions Determination For Human Activity Recognition
abstract
Human activity recognition(HAR) is one of the most active topics in the field of ubiquitous computing. Multi-sensor based HAR has attracted extensive interest because of its high recognition accuracy. Correspondingly, the number of body sensors in terms of hardware cost, the overload on communications, the storage and the computational complexity will be very high. In this paper, we propose a novel approach which can optimize cost-efficient sensors positions to save all the costs while maintaining high recognition performance. The tradeoff among the sensor positions, the target category, and the redundancy is considered. We also propose a data set D that contains acceleration sensor data for seventeen positions of the human body by simulating the worker actions in a factory assembly lines. The experimental results show that only six sensors can maintain high activity recognition accuracy out of seventeen sensors by using the proposed method.
Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001, Ruiguo Zhang
IEEE Trans. Sustain. Comput.1